The disagreement
Two real studies, thirty times apart
Take a store with 100,000 visits a month, a 2% conversion rate and an $85 average order. That is $170,000 a month in revenue. Ask the calculators what one tenth of a second is worth to it and they answer $31,234 a month, $374,805 a year. Ask the randomised experiments and the answer is $1,020 a month.
Both figures come from real studies. The larger one compounds coefficients from a correlational study of 37 brands. The smaller one applies the result of an experiment in which 10% of Bing users were deliberately slowed by 100 milliseconds and another 10% by 250, for two weeks. On the same quantity the two are 30.6 times apart.
That is not a rounding difference. It is the difference between measuring an effect and noticing a correlation, and no speed calculator mentions it. This page exists so you can decide which of the two to carry into a budget meeting.
Grading
What each kind of source may claim
The tiers below are not a ranking of how interesting a study is. They decide what a figure taken from it is allowed to claim.
| Tier | Design | What a figure from it may claim |
|---|---|---|
| Tier A | Randomised controlled experiment. Speed was deliberately changed for a randomly assigned share of users. | A causal claim: making this faster would gain you X. |
| Tier B | Large-scale observation. Speed moved on its own and outcomes were correlated with it. | A cross-check and nothing more. It cannot separate a site getting faster from a fast visit being the kind of visit that buys. |
| Tier C | No primary source survives. Quoted everywhere, traceable to nothing that publishes a method. | Nothing. Listed so it can be refused by name instead of drifting back in from a blog post. |
Convergence
Three organisations, one scale
Every figure below is revenue per 100 milliseconds, converted onto a single scale so the four can be read against each other.
| Source | Design | Baseline it was measured from | Revenue per 100 ms |
|---|---|---|---|
| Vodafone, 2021 | Randomised controlled experiment | LCP 8.3 s improved to 5.7 s | 0.31% |
| Bing (Microsoft), KDD '13 | Randomised controlled experiment, peer reviewed | Sub-second at the 95th percentile | 0.60% |
| Zalando, 2018 | Randomised controlled experiment | Not published | 0.70% |
| Deloitte with Google, 2020 | Observational | Not published | 18.37% |
The first three were run by three unrelated organisations on three different businesses: web search, fashion retail, and a telecoms landing page. They share no design, no population and no authors, and they land inside a factor of 2.3 of each other. That agreement is why the calculator carries its money figure as a band across these three rather than as one number.
The fourth sits 26 to 60 times above all of them. It is also the only one of the four in which nobody decided who got the slow version.
Read closely, the spread inside the first three is a finding, not noise. Bing and Zalando report the marginal value of the first 100 ms at a fast baseline. Vodafone reports the average value of 100 ms across a 2.6 second repair that began at 8.3 s. A smaller average across a much larger improvement is what diminishing returns looks like when it is measured instead of assumed. Three points from three different businesses cannot establish a curve, so treat this as consistency and not as proof.
Extrapolation
How far a straight line may be drawn
A figure per 100 ms is only useful if it can be multiplied. Every calculator in this category multiplies it, and none of them says on whose authority. There is an authority, and it sets a limit.
Kohavi, Deng, Longbotham and Xu ran Bing slowdown experiments at several delay amounts and reported that a linear approximation was very reasonable for Bing, projecting their own 250 ms result to 500 ms in the same paragraph. In the same paper they state that the effect holds in both directions, slowdown and speedup. That second sentence is what allows an experiment in which users were deliberately slowed to say anything at all about the value of speeding them up, and it is the step no calculator justifies.
The same paper sets the ceiling. Earlier Bing tests, with delays of up to two seconds, had a smaller revenue impact than the line predicts, and it reports that without the numbers behind it.
So the line is honest to about 0.5 s and no further. Past that this model prices nothing and reports the remainder in seconds instead. A gap three times the validated range and a gap thirteen times it produce the same money figure, and the difference between them is published as seconds left unpriced. Where the evidence stops, the tool stops.
The shape, measured at five doses
The numbers behind that ceiling were published, four years earlier, by the same company. Eric Schurman of Bing and Jake Brutlag of Google presented their delay experiments at Velocity in June 2009. O'Reilly still serves the deck from the conference CDN, and the Bing results are a table inside it.
It is the only source in this field that reports revenue at more than one dose. One effect size tells you how much. Five tell you what shape the response has, which is the thing every speed calculator assumes and none of them measures.
| Server delay | Distinct Queries/User | Query Refinement | Revenue/User | Any Clicks | Satisfaction | Time to Click (increase in ms) |
|---|---|---|---|---|---|---|
| 50ms | none | none | none | none | none | none |
| 200ms | none | none | none | -0.3% | -0.4% | +500 |
| 500ms | none | -0.6% | -1.2% | -1.0% | -0.9% | +1200 |
| 1000ms | -0.7% | -0.9% | -2.8% | -1.9% | -1.6% | +1900 |
| 2000ms | -1.8% | -2.1% | -4.3% | -4.4% | -3.8% | +3100 |
Read down the revenue column. Nothing at 50 ms. Nothing at 200 ms. Then 1.2% at half a second, 2.8% at a second, and 4.3% at two. The step from half a second to a second costs more than the first half second did. The step from one second to two costs less than the step before it, despite being twice as long.
In the unit this page works in, the marginal cost of 100 ms runs 0.24% up to half a second, 0.32% between half a second and one, then 0.15% past one second. It roughly doubles, then halves. That is diminishing returns arriving at a particular place, and the place is a second.
The authors put it on their conclusion slide, in those words: "Delays under half a second impact business metrics". This model stops where they stopped.
Not every millisecond is the same millisecond
A delay in the critical path and a delay below the fold are not the same event, and the difference has been measured. Bing delayed the elements in its right-hand pane by 250 milliseconds across almost 20 million users. If there was an impact on key metrics, it was not detectable.
So a second is not simply a second. A calculator that takes one figure from a field report and prices every millisecond as though it sat in the critical path is pricing something it never measured.
Worked example
The whole calculation, in the open
The same business as above, run against one target from three starting speeds. Nothing is hidden inside a helper: the arithmetic is written out under the table.
Inputs
- 100,000 visits a month
- 2% conversion rate
- $85 average order value
- Retail, because that is a vertical where the source study counted a purchase
- Target LCP 2.5 s, which is the threshold Google grades as good
| Current LCP | Gap | Priced | Left unpriced | Per year: low, central, high |
|---|---|---|---|---|
| 2.8 s | 0.3 s | 0.3 s | 0.0 s | $18,831 · $36,720 · $42,840 |
| 4.1 s | 1.6 s | 0.5 s | 1.1 s | $31,385 · $61,200 · $71,400 |
| 9.0 s | 6.5 s | 0.5 s | 6.0 s | $31,385 · $61,200 · $71,400 |
The middle row, end to end: $170,000 a month, times 0.60% for each 100 ms, times 5 steps of 0.1 s. That is the central figure, and the bounds substitute 0.31% and 0.70% for the same arithmetic. Nothing else happens to the number, and there is no compounding anywhere in it.
The bottom two rows are identical on purpose. One site is 1.6 s from its target and the other is 6.5 s away, and the evidence prices the same half second in both. The rest is reported as distance, in seconds, because distance is what is actually known about it.
The outlier
Why the market figure is thirty times larger
Milliseconds Make Millions, published by Deloitte with Google in 2020, is a real study: 37 retail, travel, luxury and lead-generation brands across Europe and the US, four weeks, mobile only. It is where the 8.4% conversion and 9.2% order value figures in every speed calculator come from. Four things about it decide how far those figures can travel.
Speed was not assigned.
The report states that fluctuations in speed all occurred naturally and were not artificially created on any of the sites. That makes it correlational, so its figures cannot support the claim that a repair produces the gain.
Null results were removed.
Where an increase in speed had minimal or zero effect, the report says those observations were not included. The coefficients are conditioned on an effect having been found, which leans them high by an amount nobody outside the study can estimate.
The curve has no origin.
The report fits a logarithmic regression but never publishes the baseline speed the 0.1 s was measured from. Without that operating point the curve cannot be placed, and calibrated against plausible baselines the same data produces anything from +39% to +413% for a single worked target.
It did not measure LCP.
The 0.1 s is the cumulative movement of four metrics: Server Latency, Estimated Input Latency, First Meaningful Paint and Observer Load. LCP is not one of them, and it did not exist in Lighthouse when the study ran.
Then there is the arithmetic. Every calculator compounds the coefficient: 1.084 raised to the number of 0.1 s steps. On the source document's own example, LCP 1.6 s better and therefore 16 steps, that comes out as 263% more conversions and 309% more order value at the same time. It turns $170,000 of monthly revenue into $2,526,269 and reports the difference, $2,356,269 a month, as money the business is losing. 14 times its entire revenue.
The mistake is structural rather than a slip in a formula. A logarithmic regression has diminishing returns. Exponentiation has accelerating returns. They are opposite shapes, and the report says plainly which of the two it fitted.
There is one more thing worth knowing about that 8.4%, and it is not a point about causality. Shopify published the same quantity in April 2026, measured across its own merchant base on LCP specifically: 3.50% of conversion per 100 ms. That is 2.4 times apart from the figure every calculator uses, and both are observational studies of stores. The observational side does not agree with itself either.
Refusals
Numbers this page will not use
Five numbers, and between them they are most of what an owner has been told about this subject. Four of the five are what Google's own AI Overview answers with today. Each was followed to its source, and what the source says is set out beside what gets quoted from it.
Every 100 ms of latency cost Amazon 1% in sales.
The most quoted number in the field. Followed to its source it ends at a 2006 slide deck, and in one peer-reviewed paper it is cited to a Forbes column. Amazon has never published the experiment: no method, no sample, no duration, no definition of revenue. It appears here only so it can be refused by name.
A one-second delay costs up to 7% of conversions.
It traces to Aberdeen Group, “The Performance of Web Applications: Customers are Won or Lost in One Second”, November 2008. Read in the archive, because the brief itself is behind a login and Aberdeen's own page for it does not contain the figure at all. What that page does say is the method: a survey of over 160 organisations about web applications in the enterprise. Not a measurement of shoppers, not a measurement of anything. The only number it states publicly is that performance problems might cost up to 9% of corporate revenue, at companies averaging 1.3 billion dollars of it. Eighteen years later this is quoted as a law about a checkout.
On mobile, every second of delay costs up to 20% of conversions.
Credited on every page that carries it to “Google Analytics data”, a “Google and SOASTA study” or simply “research from Google”. It is in no Google document that can be retrieved, including the two that the other mobile figures come from, and those two say something else: 123% more bounce between one second and ten, and 53% of visits abandoned past three. Nothing published supports a 20% conversion figure per second. It is refused the same way the Amazon number is.
Going from one second to three raises bounce probability 32%.
Real research, misquoted. Google published it in February 2017 and revised it in 2018, and the text gives one second to TEN seconds and 123%, not one to three and 32%. The 32% is from a chart beside that paragraph. The method matters more than either figure: both are the output of a deep neural network trained on bounce and conversion data at 90% prediction accuracy. A model's prediction over observational data, not a measurement and not an experiment. Google has since taken the page down, so it is in the register through an archived copy.
53% of mobile visitors leave if a page takes over three seconds.
This one is real, it is Google's, and almost everything carried with it is wrong. Its own footnote says: aggregated Google Analytics data from mobile sites that opted into sharing benchmark data, n=3.7K, global, March 2016. Observational, a sample that selected itself, ten years old, published by an advertising product for publishers, and counting abandoned visits rather than lost sales. It is quoted today as a causal fact about a shop's customers. Google has removed that page too.
This list had a second entry until the slides turned up.
Schurman and Brutlag, Velocity 2009, sat here because the only accounts of it were second hand and they disagreed about the size of the Bing delay: Arapakis and colleagues say 1.5 seconds in 2021, most secondary accounts say 2. O'Reilly still serves the original deck from the conference CDN. The Bing results inside it are a vector drawing whose text can be read back with its coordinates, and the row carrying the disputed figures is 2000 ms. The secondary accounts are the correct ones, Arapakis and colleagues are not, and the study is now the best dose-response evidence in this register rather than a refusal.
Against
The case against this page
A page that collects only the studies agreeing with it is marketing. So here is the strongest version of the case against the number this one produces. Two of the four objections come out of the register's own arithmetic, and neither is resolved anywhere above.
Bing does not agree with Bing
The central figure here, 0.60% of revenue per 100 ms, is Bing in 2013 at doses of 100 ms and 250 ms. Over 250 ms that implies 1.50%. The same company in 2009, running the same kind of experiment on the same product, measured nothing statistically significant at 200 ms and only 1.20% at 500 ms, which is twice the delay. Both are tier A. The likely reconciliation is that the 2009 test could not detect a small effect and that Bing was faster by 2013, so the same delay was a larger proportional change. That is a guess. Nothing published settles it, and the figure this page leads with is the one being contradicted.
The shape is not the one the band assumes
This model prices a gap at a flat rate per 100 ms. The only dose-response table in the field says the rate is not flat and does not fall off smoothly either: 0.24% per 100 ms over the first 0–500 ms, then 0.32% over 500–1000, then 0.15% over 1000–2000. It rises before it falls. So the three randomised results are not necessarily three independent estimates of one quantity; they may be one curve sampled at three starting speeds, and Zalando never published a starting speed at all. Read that way the band is a measure of how little is known rather than a confidence interval, and the error from using a flat rate depends on how slow the reader's site is, which this page does not correct for.
Four times someone measured nothing
Etsy delayed users by 200 ms and found it did not matter for them. Bing delayed the right-hand pane by 250 ms and found no detectable effect across nearly 20 million users. Rakuten ran the best-designed commercial experiment in this register and LCP, the metric this calculator reads, is not among the metrics that moved. A cosmetics retailer in Blue Triangle's own whitepaper spent three months getting 1.5 s faster and reported no result in conversion or revenue. None of these removes the effect. Together they say it is conditional on things nobody has enumerated.
Look at who published the evidence
One source in this register is peer reviewed. Of the randomised commercial results the band is built from, Zalando's is an engineering blog post, Vodafone's is a case study on Google's own site, and neither publishes a sample size, a duration or an assignment procedure. The observational study the market runs on states that brands where the effect was not significant were left out of the report, which means its coefficients are conditioned on an effect being found. A reader who discounts vendor-published results and excluded nulls is not being unreasonable, and this page would have very little left.
What would falsify this: a randomised speed experiment on a store, published with its method, its sample and its baseline, that finds an effect materially below 0.31% per 100 ms. There is no such experiment in the public record in either direction, which is the point of the two open questions below.
None of this is an argument for using a larger number instead. Every objection above cuts against the market figure harder than against this one, because the market figure is a single observational coefficient compounded across seconds it was never measured over. It is an argument for treating the result as a floor with a stated provenance, which is what the calculator says it is.
Limits
What this cannot tell you
These sit beside every result the calculator produces, not in a footnote under it.
- Every causal source measured server delay. This tool reads LCP, because LCP is what field data exposes. That substitution is the largest single weakness in the exercise.
- Bing's figure is from web search, Zalando's from fashion retail, Vodafone's from a telecoms landing page. None of them is your site.
- The elasticity is local. The same authors write that if a site's performance or its audience changes, the impact may be different.
- Not every millisecond counts the same. A 250ms delay to below-the-fold elements at Bing produced no detectable effect across nearly 20 million users.
- A linear approximation is validated to roughly half a second. Past that the curve bends, and the doses Bing published show the second second costing about half what the first one did. Anything beyond that half second is reported as distance, never as money.
- Etsy ran the same kind of experiment and found a 200ms delay did not matter for its users. The effect is real but not universal.
Two things still open
Both are known holes. Neither is closed by care on this page.
- Not one of the 15 sources publishes an interval on the figure quoted from it. 0 of them. The register records this per source and the answer never changes: a single number, no standard error, nothing a reader could use to say how sure anyone is. The 2009 slides come closest by marking which rows were not significant, and even there the dose-response table is readable as a shape and not weightable as a measurement. Every figure on this page is a point estimate, and the band around it is a spread between sources rather than a confidence interval.
- There is no randomised commerce experiment in the public record published with its method. Zalando is the closest and gives it one sentence. That absence is the largest gap in this field, and it is why the central figure on this page comes from web search and not from a shop.
Protocol
How this register was assembled
Everything above rests on which sources exist, so here is where they were looked for. This matters most for the things that are not here: a page can only claim an experiment has never been published if it says where it looked.
| Where | What was looked for | What it returned |
|---|---|---|
| ACM Digital Library: KDD, CHIIR, SIGIR | Randomised latency or speed experiments reported with a method | Kohavi et al. 2013, Kohavi et al. 2014, Arapakis et al. 2021, and Etsy's null result reported inside the first of those |
| NBER, SSRN, arXiv, Harvard and MIT working papers | Economics or business-school work on page speed and commercial outcomes | Nothing. This question has no literature in those venues |
| Management Science, Marketing Science, Information Systems Research | The same, in the journals that would publish it if it existed | Nothing |
| Engineering blogs at retailers and commerce platforms | A/B tests on store speed, published by the team that ran them | Zalando 2018 and Shopify 2026. Zalando gives its experiment one sentence |
| web.dev case studies | Commercial speed results Google chose to publish | Vodafone 2021 and Rakuten 2022 |
| Conference archives and web.archive.org | Sources quoted everywhere whose page no longer exists | The Velocity 2009 slides on O'Reilly's CDN, two Google articles Google has removed, and the Aberdeen 2008 brief |
| Vendor whitepapers | Published null results, which nobody has an incentive to write | One, in Blue Triangle's own document |
| Three separate attempts, 2026-09-25 to 2026-10-04 | A slowdown or speed-up experiment on a shop, published with its method, sample and baseline | Nothing, in any venue. This is the largest gap in the field and the reason the central figure here comes from web search |
What a source had to meet
To enter the register at all, a figure has to be readable in a document that can be opened. Not quoted in a summary, not attributed in a blog post: read, with the sentence carrying it reproduced verbatim. That rule is what the 15 entries have in common and it is why four of them are here through an archived copy.
The tier is then decided by design alone, not by how interesting or how recent the result is. Speed assigned to users is tier A, speed that moved on its own is tier B, and a figure with no retrievable method is tier C and carries no money anywhere on this page.
A source is excluded only for being unreadable, never for disagreeing. That is why the null results and the study that halves this page's own central figure are both in the register rather than left out of it.
What this is not
It is not a systematic review in the formal sense. The searching was not logged as it happened, so there is no count of records screened and no two-reader agreement, and this table was written from what the search returned rather than from a protocol fixed in advance. Treat it as a reproducible trail rather than as a replication: every row can be re-run, and a reader who finds a source that belongs here and is missing has found a real defect.
The register
All 15 sources, in full
Each entry carries its design, where it was published, who was measured, how the quotes were checked, and what the study cannot support. Open an entry to read the sentences the figures come from.
Quotes appear in the language they were published in. A translated quote is no longer verbatim, and a reader checking it against the source would not find the sentence.
Online Controlled Experiments at Large Scale
Kohavi, Deng, Frasca, Walker, Xu, Pohlmann (Microsoft), 2013
- Published in
- KDD '13, ACM SIGKDD
- Who was measured
- Microsoft Bing users; two treatment arms, 10% slowed by 100ms and 10% by 250ms, over two weeks
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- A single number, with no interval and no error term
Quotes and limits
In the authors' words
We recently ran a slowdown experiment where we slowed 10% of users by 100msec (milliseconds) and another 10% by 250msec for two weeks. The results showed that performance absolutely matters a lot today: every 100msec improves revenue by 0.6%.Bing's server performance is now sub-second at the 95th percentile.This quantification may change over time, as the site's performance and bandwidth standards improve.What it cannot support
- Web search, not a store. The transaction is a click on a result, not a checkout, so the number transfers to retail only as an analogy.
- Measured against a sub-second baseline. The authors say directly that the quantification changes as a site's performance standards change.
- Server delay specifically, not page weight, render time or layout stability.
- The same authors report a contradicting result: Etsy found a 200ms delay did not matter for its users.
- A single dose pair, 100ms and 250ms. The dose-response shape comes from the same company's 2009 slides, not from here.
Seven Rules of Thumb for Web Site Experimenters
Kohavi, Deng, Longbotham, Xu (Microsoft), 2014
- Published in
- KDD '14, ACM SIGKDD
- Who was measured
- Bing users across several slowdown experiments at different delay amounts
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- A single number, with no interval and no error term
Quotes and limits
In the authors' words
a 250msec delay at the server impacts revenue at about 1.5% and clickthrough-rate by 0.25%. While this is a massive impact, 500msec would impact revenue about 3% not 20%, and clickthrough-rate would drop by 0.50%, not 20% (assuming a linear approximation is reasonable).Earlier tests at Bing [32] had similar click impact and smaller revenue impact with delays of up to two seconds.Using a linear approximation (1st-order Taylor expansion), we can assume that the impact of the metric is similar in both directions (slowdown and speedup) ... By running slowdown experiments with different slowdown amounts, we have confirmed that a linear approximation is very reasonable for Bing.The slowdown quantifies the impact on the metric of interest at the point today ... If the site performance changes (e.g., site is faster), or the audience changes (e.g., more international users) the impact may be different.A recent slowdown controlled experiment was run ... delaying when the right pane elements were shown by 250 milliseconds. If there was an impact on key metrics, it was not detectible, despite the experiment size of almost 20 million users.What it cannot support
- The linear approximation is validated for Bing, at Bing's speed, for Bing's audience. The authors say in the same paper that a different speed or audience may give a different impact.
- The two-second observation is reported here without its numbers. They are in the Velocity 2009 slides from the same company, and they say the marginal cost of 100ms roughly halves beyond one second.
- Still web search.
Loading Time Matters
Zalando engineering, 2018
- Published in
- Zalando Engineering Blog (not peer reviewed)
- Who was measured
- Zalando customers; correlation across the journey, then an A/B test for confirmation
- How the quotes were checked
- Read in the source blog post
- What it says about precision
- A single number, with no interval and no error term
- Source
- engineering.zalando.com
Quotes and limits
In the authors' words
100 msec loading time improvement led to a 0.7% uplift in revenue per session.We analyzed the correlations of loading time and revenue per session across every step of the user journey and for every device.An A/B test brought the final confirmation.What it cannot support
- No baseline loading time, sample size, duration or assignment detail is published, so the experiment cannot be inspected.
- The headline analysis is explicitly correlational; the A/B test is mentioned in a single sentence without numbers of its own.
- Fashion retail in Europe. One business, one catalogue, one audience.
Vodafone: A 31% improvement in LCP increased sales by 8%
Vodafone, published by Google on web.dev, 2021
- Published in
- web.dev case study (not peer reviewed)
- Who was measured
- 50/50 traffic split on a landing page, roughly 100K clicks and 34K visits per arm per day
- How the quotes were checked
- Read in the source case study
- What it says about precision
- A single number, with no interval and no error term
- Source
- web.dev
Quotes and limits
In the authors' words
50% of the traffic was sent to the optimized landing page (version A), and 50% was sent to the baseline page (version B).LCP improved by 31% (from 8.3s to 5.7s)Sales increased by 8%Lead-to-visit rate improved by 15%Cart-to-visit rate improved by 11%What it cannot support
- The treatment changed server-side rendering AND images together, so speed is not isolated from the other changes. It is a randomised comparison of two page builds, not of two speeds.
- Published by the vendor whose metric it validates, on Google's own site.
- No duration or total sample is given beyond daily traffic.
- 0.31% of revenue per 100ms here is an average across 2.6s from a slow baseline, and must not be compared like-for-like with a marginal first-100ms figure.
Performance Related Changes and their User Impact (The User and Business Impact of Server Delays, Additional Bytes, and HTTP Chunking in Web Search)
Schurman (Microsoft Bing), Brutlag (Google), 2009
- Published in
- Velocity 2009 (conference presentation, no paper)
- Who was measured
- Bing and Google web search users, each site experimenting separately. Bing: a small number of users per delay level. Google: a small percentage of search traffic, randomly assigned, run for four to six weeks per level
- How the quotes were checked
- Read in the original slides
- What it says about precision
- Significance is marked, but no interval is given
- Source
- cdn.oreillystatic.com
Quotes and limits
In the authors' words
Determine impact of server delaysDelay before sending resultsDifferent experiments with different delaysRandomly assign users to the experiment and control groups (A/B testing)Server-side delay: Emulates additional server processing timeVaried type of delay, magnitude (in ms), and duration (number of weeks)Strong negative impactsRoughly linear changes with increasing delayTime to Click changed by roughly double the delay- Means no statistically significant changeDelays under half a second impact business metricsThe cost of delay increases over time and persistsWhat it cannot support
- Slides and a talk, never a paper. The design is described well enough to understand and not well enough to replicate: no sample sizes, no confidence intervals, no statistical tests beyond a dash meaning not significant.
- Bing in 2009, at a 2009 baseline that is nowhere stated. Its marginal rates run from 0.15% to 0.32% of revenue per 100ms, below the 0.6% the same company measured in 2013 at a sub-second baseline, which is what an elasticity measured at a different operating point looks like.
- Web search, not commerce.
- Resolved here, and worth recording: Arapakis et al. (2021) attribute the -1.8% queries, -4.3% revenue and -3.8% satisfaction to a delay of 1.5 seconds. In the original slides that row is 2000ms. The secondary accounts saying two seconds are the correct ones.
A 200ms delay did not measurably matter at Etsy
McKinley (Etsy), reported in Kohavi et al. 2013, 2013
- Published in
- Reported in KDD '13
- Who was measured
- Etsy users
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- Nothing. No figure precision is stated at all
Quotes and limits
In the authors' words
we were surprised to see Etsy's Dan McKinley [2] claim that a 200msec delay did not matter. It is possible that for Etsy users, performance is not criticalWhat it cannot support
- Reported second-hand inside another team's paper, by authors who doubted it.
- A null result, so it bounds how universal the effect is rather than measuring one.
A cosmetics retailer got 1.5 seconds faster and saw no result
Blue Triangle Technologies, 2020
- Published in
- Vendor whitepaper (not peer reviewed)
- Who was measured
- One anonymous cosmetics retailer, named nowhere in the document
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- Nothing. No figure precision is stated at all
- Source
- bluetriangle.com
Quotes and limits
In the authors' words
They spent 3 months improving performance by 1.5 seconds on average, implementing A/B testing to ensure the site was production-ready. Upon launching the new “faster site”, they noticed no real results in conversion rates and top-line revenue because attention was given to the wrong webpages.To support the argument that there is not a “one-size-fits-all” page speed goal to best support user experience improvements and overall sales, three different eCommerce retailers were analyzed using Blue Triangle data in this study.What it cannot support
- The A/B test it mentions was run to check the new store was production-ready, not to measure the speed effect. So this is a before and after across a relaunch, with nothing holding the rest of the site constant.
- No sample, no duration of measurement, and no definition of “no real results”. The retailer is anonymous under a confidentiality agreement.
- Published by a performance vendor, and the explanation it offers for the null result is the thing it sells.
- It carries the Amazon 1% claim forward citing Linden 2006, which is the slide deck this register refuses that number for.
Rakuten 24's investment in Core Web Vitals increased revenue per visitor by 53.37%
Hayoung Lee, Linh Duong, Ryunosuke Akiba, Shogo Kashiwase (Rakuten, published by Google on web.dev), 2022
- Published in
- web.dev case study (not peer reviewed)
- Who was measured
- Rakuten 24 landing page, concurrent 50/50 traffic split for one month; no sample size published
- How the quotes were checked
- Read in the source case study
- What it says about precision
- A single number, with no interval and no error term
- Source
- web.dev
Quotes and limits
In the authors' words
During the test duration, 50% of the traffic was sent to the optimized landing page (version A), and 50% was sent to the original page (version B).The only difference between version A and version B was that version A was optimized for Core Web Vitals and there were no other functional or visual differences.53.37% increase in revenue per visitor33.13% increase in conversion rateCLS improved by 92.72%FID improved by 7.95%FCP improved by 8.45%TTFB improved by 18.03%What it cannot support
- LCP is not among the metrics reported to have improved. CLS, FID, FCP and TTFB are. So this is the value of a bundle dominated by layout stability, and it cannot be converted into a figure per 100ms of LCP.
- A bundle, not a speed manipulation. It is a randomised comparison of two page builds, like Vodafone, except that here the largest movement is in a metric that is not speed.
- No sample size, no confidence intervals, no statistical tests are published.
- One landing page, one month, one retailer, published by the vendor of the metric it validates on that vendor's own site.
- The widely quoted LCP figures from this page (61% better conversion under one second) are an observational segmentation of field data, not the experiment. Quoting them as the experiment's result is the same citation drift this register exists to catch.
Impact of Response Latency on User Behaviour in Mobile Web Search
Arapakis (Telefónica Research), Park (Telefónica Research), Pielot (Google), 2021
- Published in
- CHIIR '21, ACM SIGIR Conference on Human Information Interaction and Retrieval
- Who was measured
- Mobile web search users in a controlled study with assigned latency levels
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- A single number, with no interval and no error term
- Source
- arxiv.org · doi:10.1145/3406522.3446038
Quotes and limits
In the authors' words
Our preliminary results indicate that mobile web search users are four times more tolerant to response latency reported for desktop web search users. However, when exceeding a certain threshold of 7-10 sec, the delays have a sizeable impact and users report feeling significantly more tensed, tired, terrible, frustrated and sluggish, all which contribute to a worse subjective user experience.What it cannot support
- Measures subjective experience and search behaviour, not revenue.
- The authors call the results preliminary.
- A threshold at 7-10s is specific to mobile search; it is evidence that a threshold exists, not a threshold to reuse elsewhere.
Response time in man-computer conversational transactions
Miller (1968); extended by Card, Robertson & Mackinlay (1991), 1968
- Published in
- AFIPS Fall Joint Computer Conference, vol. 33, pp. 267-277
- Who was measured
- Human response-time perception, replicated across five decades of HCI work
- How the quotes were checked
- Reached through a paper citing it
- What it says about precision
- Nothing. No figure precision is stated at all
Quotes and limits
In the authors' words
0.1 second is about the limit for having the user feel that the system is reacting instantaneously1.0 second is about the limit for the user's flow of thought to stay uninterrupted10 seconds is about the limit for keeping the user's attention focused on the dialogueWhat it cannot support
- The three limits are quoted from Nielsen's codification, which cites Miller 1968 and Card et al. 1991. The 1968 proceedings are paywalled and were not read first-hand.
- Perceptual limits, not revenue. They explain why a response curve has steps in it; they do not price one.
Milliseconds Make Millions
Deloitte with Google, 2020
- Published in
- Industry report (not peer reviewed)
- Who was measured
- 37 retail, travel, luxury and lead-generation brands across Europe and the US, four weeks, mobile only
- How the quotes were checked
- Read in the source PDF
- What it says about precision
- A single number, with no interval and no error term
- Source
- www.thinkwithgoogle.com
Quotes and limits
In the authors' words
An 8.4% increase in conversions with retail consumers was observed, and an increase in average order value of 9.2%.A 10.1% increase in conversions with travel consumers was observed, and a slight increase in average order value of 1.9%.Once data had been collected and fed into a [l]ogarithmic regression model, we were in a position to analyse the data and extract meaning.Fluctuations in speed all occurred naturally and were not artificially created on any of the sites.The 0.1 second improvement was the cumulative impact of all four speed metrics.In some instances, the increase in speed was observed to have minimal or zero effect on funnel progression and/or the KPIs for the vertical. These observations were not included in the report.What it cannot support
- Correlational. Speed was not assigned, so the figures cannot support a causal claim that a fix produces the gain.
- Null results were excluded from publication, so the coefficients are conditioned on an effect being found and lean high.
- The baseline speed the 0.1s was measured from is never published, which leaves the curve unidentifiable. Calibrated against plausible baselines the same data yields anything from +39% to +413% for one worked example.
- The 0.1s is the cumulative movement of four metrics (Server Latency, Estimated Input Latency, First Meaningful Paint, Observer Load). LCP is not among them and did not exist in Lighthouse at the time.
- Mobile only. The authors looked at desktop and 'found a lot of contradicting parameters'.
- Luxury conversion is defined as adding to basket or clicking 'contact us', so the study never measured luxury revenue.
Store Speed and Conversion: What the Data Shows
Shopify, 2026
- Published in
- Shopify Enterprise blog (not peer reviewed)
- Who was measured
- Actively-selling Shopify stores, real-user performance data over a 28-day window at the turn of January and February 2026, with the slowest 5% excluded
- How the quotes were checked
- Read in the source blog post
- What it says about precision
- A single number, with no interval and no error term
- Source
- www.shopify.com
Quotes and limits
In the authors' words
for every 100 milliseconds slower a store loads, conversion tends to be about 3.5% lower.stores with 2.5 second LCP report roughly 30% lower conversion than stores with 1.5 second LCP.To prevent outliers from skewing the results, we excluded the slowest 5% of stores from the analysis.Largest Contentful Paint (LCP) measures how quickly your main content appears.At the individual store level, conversion depends on many factors beyond speed-product-market fit, pricing, marketing, and more.What it cannot support
- Observational. Stores were bucketed by the speed they already had, not assigned one, so the figure cannot support a claim that making a store faster produces the gain.
- Grouping stores to average out product fit, pricing and marketing dilutes a confounder, it does not remove one. A store that invests in being fast is plausibly a store that invests in everything else.
- Conversion, not revenue. It says nothing about order value, which is half of the other observational study's figure.
- Published by the platform whose merchants it measures, with no sample size given.
- It disagrees with the other observational study by a factor of 2.4 on the same quantity: 3.5% per 100ms here against 8.4% per 0.1s in Milliseconds Make Millions.
53% of mobile site visits are abandoned after three seconds
Shellhammer (Google / DoubleClick), 2016
- Published in
- DoubleClick by Google, “The need for mobile speed” (not peer reviewed)
- Who was measured
- Mobile sites that opted into sharing Google Analytics benchmark data, n=3.7K, global, March 2016
- How the quotes were checked
- Read in an archived copy, the original having been taken down
- What it says about precision
- A single number, with no interval and no error term
- Source
- web.archive.org
Quotes and limits
In the authors' words
In our new study, “The Need for Mobile Speed”, we found that 53% of mobile site visits are abandoned if pages take longer than 3 seconds to load.Google Data, Aggregated, anonymized Google Analytics data from a sample of mWeb sites opted into sharing benchmark data, n=3.7K, Global, March 2016What it cannot support
- Observational. Nothing was assigned, so it cannot say that making a page faster would keep anyone.
- The sample selected itself: sites that chose to share benchmark data with Google.
- Mobile site visits, published by an advertising product for publishers. A visit abandoned is not a sale lost.
- March 2016 data, on connection speeds and devices that no longer exist.
- Google has removed the page. Only the archive copy remains.
Bounce probability rises 123% between one and ten seconds
An (Google), 2017
- Published in
- Think with Google, “Mobile page speed: new industry benchmarks” (not peer reviewed)
- Who was measured
- Mobile landing pages analysed by Google, revised February 2018
- How the quotes were checked
- Read in an archived copy, the original having been taken down
- What it says about precision
- Only the accuracy of the model that produced it
- Source
- web.archive.org
Quotes and limits
In the authors' words
We also trained a deep neural network—a computer system modeled on the human brain and nervous system—with a large set of bounce and conversions data. The neural net, which had a 90% prediction accuracy, found that as page load time goes from one second to 10 seconds, the probability of a mobile site visitor bouncing increases 123%.What it cannot support
- A neural network's prediction over observational data, not a measurement and not an experiment.
- The widely quoted “one to three seconds, 32%” is not in the text. The text gives one to ten seconds and 123%.
- Bounce is not a purchase. Nothing here measures revenue.
- No sample size, no confidence interval and no description of the training data.
- Google has removed the page. Only the archive copy remains.
Every 100ms of latency cost Amazon 1% in sales
Attributed to Amazon, via Greg Linden's slides and secondary press, 2006
- Published in
- Never written up
- Who was measured
- Unknown. No method, sample size or date has been published.
- How the quotes were checked
- Not verifiable
- What it says about precision
- Nothing. No figure precision is stated at all
- Source
- www.exp-platform.com
Quotes and limits
In the authors' words
Greg Linden [40 p. 15] noted that 100msec slowdown at Amazon impacted revenue by 1%What it cannot support
- No published method, sample, duration or definition of revenue.
- Traced through citations it ends at a 2006 slide deck and, in one peer-reviewed paper, at a Forbes column.
- Excluded from every calculation here. It appears in the UI only as an example of a number that cannot be checked.
Citing this page: Quote it together with the study a figure rests on, not with this page alone. Every number below names its own source and links to it.
Let’s Create an Amazing Project Together!
Questions people ask about this
Is it true that 100 ms of latency cost Amazon 1% in sales?
There is no way to check it. Followed to its source the figure ends at a 2006 slide deck, and one peer-reviewed paper cites it to a Forbes column. The experiment behind it was never published, so there is no method to read, no sample to weigh and no definition of what counted as revenue. This page lists the claim so it can be refused by name, and takes nothing from it.
Is it true that 53% of mobile visitors leave after three seconds?
The sentence is real and it is Google's, from 2016. What travels with it is not. Its own footnote says the figure is aggregated Analytics data from mobile sites that opted in to sharing it, n=3,700, March 2016. So it is observational, the sample selected itself, it counts abandoned visits rather than lost sales, and it is ten years old. Google has since taken the page down.
Does a faster site always make more money?
No, and four published results say so. A 200 ms delay at Etsy produced no measurable effect on its own users. Bing delayed a side panel by 250 ms and saw nothing across nearly 20 million users. In the best-designed commercial experiment here, LCP is not among the metrics that moved at all. And one retailer spent three months getting 1.5 seconds faster and reported no change in conversion or revenue.
Which number should I take to a finance team?
The lower one, with the experiment named beside it. A figure from a randomised test survives being checked; a figure from a correlation does not, and the gap between them here is roughly thirty times. The calculator produces the first kind and prices only the part of a gap the evidence covers.
Why is there no confidence interval anywhere on this page?
Because not one source publishes one. Every figure in the register is a single number with no standard error, which is recorded source by source rather than mentioned once. The 2009 slides come closest by marking which results were not significant. The band shown with a result is the spread between separate studies, not a confidence interval.