Paying for Speed: A Two-Sample Test of Parcel Carrier Delivery Times at a Direct-to-Consumer Retailer
Author Name
School of Business and Information Technology, Purdue Global
GB513: Business Analytics
Unit 4 Assignment
Professor Name
January 26, 2026
Decision Problem, Data, and Assumptions
Halstead Outdoor Supply ships about 240,000 parcels a year to consumers from a single distribution center. Both parcel contracts are up for renewal, and Carrier A has quoted a rate that runs $1.15 per parcel above Carrier B, or $276,000 a year if all volume moves. Sales leadership argues that Carrier A delivers faster and that speed protects repeat purchase in a channel where delivery expectations keep rising (U.S. Census Bureau, 2026). Operations argues that the premium buys hours nobody notices. The argument is settled only if the speed claim is treated as a measurable difference rather than a preference, so the business question becomes a statistical one: does mean delivery time differ between the two carriers, and by enough to price?
The data set is a shipment file covering the 61,400 parcels sent between October 1 and December 31, 2025. A simple random sample of 120 shipments per carrier was drawn, giving n = 240. Each record holds ship date, delivery scan date, destination zone, service level, and promise date. Delivery cycle time is defined as calendar days from carrier acceptance scan to delivery scan. Fourteen records with missing delivery scans were dropped before sampling, and the sample was limited to ground service so that the comparison is not contaminated by mixed service levels. Carrier A averaged 3.42 days (SD = 0.94, Mdn = 3.2) and Carrier B averaged 3.86 days (SD = 1.12, Mdn = 3.6).
Four assumptions carry the analysis, and each is checked rather than asserted. Independence holds because shipments were selected at random and limited to one parcel per order, so no customer appears twice. Normality of the sampling distribution is reasonable at 120 per group under the central limit theorem, which matters because both distributions carry a mild right skew. Equality of variance does not hold, with a variance ratio of 1.42 and a Levene test at p = .04, so the Welch correction is used instead of the pooled estimate (Field, 2018). Zone mix is comparable, with 55 percent of Carrier A shipments and 57 percent of Carrier B shipments in zones five through eight.
Method and Results
The primary method is an independent samples t test with the Welch correction, two-tailed, at alpha = .05. The null hypothesis states that mean delivery time is equal for the two carriers; the alternative states that it is not. The test returned t(231) = 3.30, p = .001, with a mean difference of 0.44 days favoring Carrier A and a 95 percent confidence interval of 0.18 to 0.70 days. The null hypothesis is rejected. Effect size is Cohen's d = 0.43, which sits between a small and a medium effect and should be read as a real but modest separation rather than a dramatic one (Cohen, 1988). The interval is the number worth carrying forward, because it states how small the advantage might plausibly be.
A second measure was tested because average speed is not what a customer experiences. Promise adherence was defined as delivery on or before the date shown at checkout. Carrier A missed 8 of 120 promises, or 6.7 percent, and Carrier B missed 17 of 120, or 14.2 percent. A two-proportion z test returned z = 1.90, p = .057, so the null of equal miss rates is not rejected at the stated alpha. That result is reported as it stands. A p value of .057 is not evidence that the two carriers perform alike, and neither is it license to claim the gap is established; the honest reading is that this sample is too small to settle the question (Wasserstein & Lazar, 2016).
Two checks test whether the main finding depends on choices made along the way. Three shipments held by a regional weather embargo ran past nine days, and removing them narrows the difference only slightly, to 0.41 days with t = 3.24 and p = .001, so the result is not an artifact of outliers. A distribution-free Mann-Whitney comparison run as a backup agrees in direction and significance at p = .002, which is reassuring given the skew noted earlier (Black, 2019). Power is the last check: at 120 per group the design detects an effect of d = 0.36, roughly a third of a day, about 80 percent of the time, so smaller advantages would likely have been missed.
From Days to Dollars: Interpretation and Recommendation
A difference of 0.44 days has no meaning until it is priced. Carrier A charges $1.15 more per parcel, so the company would pay between $1.64 and $6.39 for each parcel-day of speed, depending on whether the true difference sits at the upper or the lower end of the confidence interval. Moving all 240,000 parcels to Carrier A costs $276,000 a year and buys, at the conservative bound, about four hours per parcel. No margin structure in outdoor equipment retail supports that price across the whole book. The finding is statistically clear and commercially weak at full scale, and treating those two statements as one is the most common error in this kind of report.
The recommendation therefore segments rather than converts. Of annual volume, 38 percent, or roughly 91,200 parcels, carries a dated promise at checkout, and those are the shipments where a late arrival costs a refund, a service contact, or a lost repeat order. Routing that segment to Carrier A costs about $104,880 a year and leaves the remaining 148,800 parcels with Carrier B, which is $171,120 less than full conversion. If the observed 7.5 point gap in promise misses holds, the segment would see roughly 6,840 fewer late deliveries a year, an estimate that rests on a difference this sample did not confirm at the stated alpha and should be presented to the contract committee with that caveat attached.
Three limitations bound the recommendation. The data are observational rather than randomized, so the analysis supports a statement about association between carrier and delivery time and not a claim that one carrier causes faster delivery. The sample covers one quarter that includes peak season, when networks behave differently than they do in spring. Rates and zone mix are both renegotiated annually, so the price side of the comparison has a shelf life of one contract term. The next step follows from the weakest result: a 90-day randomized assignment of promise-dated parcels, sampled at roughly 260 shipments per carrier, would give the miss-rate comparison enough power to settle what this analysis could only point toward.
References
Black, K. (2019). Business statistics: For contemporary decision making (10th ed.). Wiley.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE Publications.
U.S. Census Bureau. (2026). Quarterly retail e-commerce sales: Fourth quarter 2025. https://www.census.gov/retail/index.html
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108
How this GB 513 Unit 4 example is structured
Purdue Global does not publish a deliverable name for each unit of Business Analytics, so this master's level example is written to the genre the unit almost certainly wants. In many sections this unit asks you to apply an inferential method to a supplied business file and report the result in decision terms; your classroom's instructions decide the exact form. The GB513 Unit 4 example is ordered so the numbers can be checked by a reader who never sees the file. The first body sheet states the decision, describes the data, and lists the assumptions before any test is run. The second reports hypotheses, statistics, interval, effect size, and a sensitivity check. The third converts days into dollars, names the recommendation, and states what would overturn it.
GB513 Unit 4 questions, answered
What does GB513 Unit 4 usually ask for?
In many sections this unit sits where the term turns from describing data to drawing inferences, so it asks you to apply a test or model to a supplied business file and report what it means for a decision. Your classroom's instructions and rubric decide the exact method and format, so confirm them before mirroring this example.
Do I have to include the software output in the paper?
Reported statistics matter more than screenshots. Give the test, degrees of freedom, exact p value, the difference, its confidence interval, and an effect size, formatted in APA style. Many classrooms also want the worksheet or output attached as an appendix. Check the instructions, and never paste output in place of writing what it means.
What if my result is not statistically significant?
Report it plainly and keep the paper. A result at p = .057, as in the second test here, is not proof that two options perform alike; it usually means the sample was too small to settle the question. Say what the sample could and could not detect, then recommend the size that would answer it.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Purdue University Global document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.