Skip to content

Method and evidence

What a tenth of a second is actually worth

The public record holds 15 sources that try to answer it. One of them sits behind almost every page-speed calculator on the market, and of the ones that measure money it is among the weakest. This page lists them all, says what each can and cannot support, and shows the arithmetic that separates them.

Run these numbers on your own traffic

Evidence reviewed: 5 October 2026 · Compiled by: Yevhen Samkov

The disagreement

Two real studies, thirty times apart

Take a store with 100,000 visits a month, a 2% conversion rate and an $85 average order. That is $170,000 a month in revenue. Ask the calculators what one tenth of a second is worth to it and they answer $31,234 a month, $374,805 a year. Ask the randomised experiments and the answer is $1,020 a month.

Both figures come from real studies. The larger one compounds coefficients from a correlational study of 37 brands. The smaller one applies the result of an experiment in which 10% of Bing users were deliberately slowed by 100 milliseconds and another 10% by 250, for two weeks. On the same quantity the two are 30.6 times apart.

That is not a rounding difference. It is the difference between measuring an effect and noticing a correlation, and no speed calculator mentions it. This page exists so you can decide which of the two to carry into a budget meeting.

Grading

What each kind of source may claim

The tiers below are not a ranking of how interesting a study is. They decide what a figure taken from it is allowed to claim.

TierDesignWhat a figure from it may claim
Tier ARandomised controlled experiment. Speed was deliberately changed for a randomly assigned share of users.A causal claim: making this faster would gain you X.
Tier BLarge-scale observation. Speed moved on its own and outcomes were correlated with it.A cross-check and nothing more. It cannot separate a site getting faster from a fast visit being the kind of visit that buys.
Tier CNo primary source survives. Quoted everywhere, traceable to nothing that publishes a method.Nothing. Listed so it can be refused by name instead of drifting back in from a blog post.

Convergence

Three organisations, one scale

Every figure below is revenue per 100 milliseconds, converted onto a single scale so the four can be read against each other.

Ascending by effect size. The first three assigned speed to users. The fourth did not.
SourceDesignBaseline it was measured fromRevenue per 100 ms
Vodafone, 2021Randomised controlled experimentLCP 8.3 s improved to 5.7 s0.31%
Bing (Microsoft), KDD '13Randomised controlled experiment, peer reviewedSub-second at the 95th percentile0.60%
Zalando, 2018Randomised controlled experimentNot published0.70%
Deloitte with Google, 2020ObservationalNot published18.37%

The first three were run by three unrelated organisations on three different businesses: web search, fashion retail, and a telecoms landing page. They share no design, no population and no authors, and they land inside a factor of 2.3 of each other. That agreement is why the calculator carries its money figure as a band across these three rather than as one number.

The fourth sits 26 to 60 times above all of them. It is also the only one of the four in which nobody decided who got the slow version.

Read closely, the spread inside the first three is a finding, not noise. Bing and Zalando report the marginal value of the first 100 ms at a fast baseline. Vodafone reports the average value of 100 ms across a 2.6 second repair that began at 8.3 s. A smaller average across a much larger improvement is what diminishing returns looks like when it is measured instead of assumed. Three points from three different businesses cannot establish a curve, so treat this as consistency and not as proof.

Extrapolation

How far a straight line may be drawn

A figure per 100 ms is only useful if it can be multiplied. Every calculator in this category multiplies it, and none of them says on whose authority. There is an authority, and it sets a limit.

Kohavi, Deng, Longbotham and Xu ran Bing slowdown experiments at several delay amounts and reported that a linear approximation was very reasonable for Bing, projecting their own 250 ms result to 500 ms in the same paragraph. In the same paper they state that the effect holds in both directions, slowdown and speedup. That second sentence is what allows an experiment in which users were deliberately slowed to say anything at all about the value of speeding them up, and it is the step no calculator justifies.

The same paper sets the ceiling. Earlier Bing tests, with delays of up to two seconds, had a smaller revenue impact than the line predicts, and it reports that without the numbers behind it.

So the line is honest to about 0.5 s and no further. Past that this model prices nothing and reports the remainder in seconds instead. A gap three times the validated range and a gap thirteen times it produce the same money figure, and the difference between them is published as seconds left unpriced. Where the evidence stops, the tool stops.

The shape, measured at five doses

The numbers behind that ceiling were published, four years earlier, by the same company. Eric Schurman of Bing and Jake Brutlag of Google presented their delay experiments at Velocity in June 2009. O'Reilly still serves the deck from the conference CDN, and the Bing results are a table inside it.

It is the only source in this field that reports revenue at more than one dose. One effect size tells you how much. Five tell you what shape the response has, which is the thing every speed calculator assumes and none of them measures.

Bing, Velocity 2009. "None" is where the authors printed a dash, their own notation for no statistically significant change.
Server delayDistinct Queries/UserQuery RefinementRevenue/UserAny ClicksSatisfactionTime to Click (increase in ms)
50msnonenonenonenonenonenone
200msnonenonenone-0.3%-0.4%+500
500msnone-0.6%-1.2%-1.0%-0.9%+1200
1000ms-0.7%-0.9%-2.8%-1.9%-1.6%+1900
2000ms-1.8%-2.1%-4.3%-4.4%-3.8%+3100

Read down the revenue column. Nothing at 50 ms. Nothing at 200 ms. Then 1.2% at half a second, 2.8% at a second, and 4.3% at two. The step from half a second to a second costs more than the first half second did. The step from one second to two costs less than the step before it, despite being twice as long.

In the unit this page works in, the marginal cost of 100 ms runs 0.24% up to half a second, 0.32% between half a second and one, then 0.15% past one second. It roughly doubles, then halves. That is diminishing returns arriving at a particular place, and the place is a second.

The authors put it on their conclusion slide, in those words: "Delays under half a second impact business metrics". This model stops where they stopped.

Not every millisecond is the same millisecond

A delay in the critical path and a delay below the fold are not the same event, and the difference has been measured. Bing delayed the elements in its right-hand pane by 250 milliseconds across almost 20 million users. If there was an impact on key metrics, it was not detectable.

So a second is not simply a second. A calculator that takes one figure from a field report and prices every millisecond as though it sat in the critical path is pricing something it never measured.

Worked example

The whole calculation, in the open

The same business as above, run against one target from three starting speeds. Nothing is hidden inside a helper: the arithmetic is written out under the table.

Inputs

  • 100,000 visits a month
  • 2% conversion rate
  • $85 average order value
  • Retail, because that is a vertical where the source study counted a purchase
  • Target LCP 2.5 s, which is the threshold Google grades as good
Target LCP 2.5 s. Low is the Vodafone figure, central is Bing, high is Zalando.
Current LCPGapPricedLeft unpricedPer year: low, central, high
2.8 s0.3 s0.3 s0.0 s$18,831 · $36,720 · $42,840
4.1 s1.6 s0.5 s1.1 s$31,385 · $61,200 · $71,400
9.0 s6.5 s0.5 s6.0 s$31,385 · $61,200 · $71,400

The middle row, end to end: $170,000 a month, times 0.60% for each 100 ms, times 5 steps of 0.1 s. That is the central figure, and the bounds substitute 0.31% and 0.70% for the same arithmetic. Nothing else happens to the number, and there is no compounding anywhere in it.

The bottom two rows are identical on purpose. One site is 1.6 s from its target and the other is 6.5 s away, and the evidence prices the same half second in both. The rest is reported as distance, in seconds, because distance is what is actually known about it.

The outlier

Why the market figure is thirty times larger

Milliseconds Make Millions, published by Deloitte with Google in 2020, is a real study: 37 retail, travel, luxury and lead-generation brands across Europe and the US, four weeks, mobile only. It is where the 8.4% conversion and 9.2% order value figures in every speed calculator come from. Four things about it decide how far those figures can travel.

Speed was not assigned.

The report states that fluctuations in speed all occurred naturally and were not artificially created on any of the sites. That makes it correlational, so its figures cannot support the claim that a repair produces the gain.

Null results were removed.

Where an increase in speed had minimal or zero effect, the report says those observations were not included. The coefficients are conditioned on an effect having been found, which leans them high by an amount nobody outside the study can estimate.

The curve has no origin.

The report fits a logarithmic regression but never publishes the baseline speed the 0.1 s was measured from. Without that operating point the curve cannot be placed, and calibrated against plausible baselines the same data produces anything from +39% to +413% for a single worked target.

It did not measure LCP.

The 0.1 s is the cumulative movement of four metrics: Server Latency, Estimated Input Latency, First Meaningful Paint and Observer Load. LCP is not one of them, and it did not exist in Lighthouse when the study ran.

Then there is the arithmetic. Every calculator compounds the coefficient: 1.084 raised to the number of 0.1 s steps. On the source document's own example, LCP 1.6 s better and therefore 16 steps, that comes out as 263% more conversions and 309% more order value at the same time. It turns $170,000 of monthly revenue into $2,526,269 and reports the difference, $2,356,269 a month, as money the business is losing. 14 times its entire revenue.

The mistake is structural rather than a slip in a formula. A logarithmic regression has diminishing returns. Exponentiation has accelerating returns. They are opposite shapes, and the report says plainly which of the two it fitted.

There is one more thing worth knowing about that 8.4%, and it is not a point about causality. Shopify published the same quantity in April 2026, measured across its own merchant base on LCP specifically: 3.50% of conversion per 100 ms. That is 2.4 times apart from the figure every calculator uses, and both are observational studies of stores. The observational side does not agree with itself either.

Refusals

Numbers this page will not use

Five numbers, and between them they are most of what an owner has been told about this subject. Four of the five are what Google's own AI Overview answers with today. Each was followed to its source, and what the source says is set out beside what gets quoted from it.

Every 100 ms of latency cost Amazon 1% in sales.

The most quoted number in the field. Followed to its source it ends at a 2006 slide deck, and in one peer-reviewed paper it is cited to a Forbes column. Amazon has never published the experiment: no method, no sample, no duration, no definition of revenue. It appears here only so it can be refused by name.

A one-second delay costs up to 7% of conversions.

It traces to Aberdeen Group, “The Performance of Web Applications: Customers are Won or Lost in One Second”, November 2008. Read in the archive, because the brief itself is behind a login and Aberdeen's own page for it does not contain the figure at all. What that page does say is the method: a survey of over 160 organisations about web applications in the enterprise. Not a measurement of shoppers, not a measurement of anything. The only number it states publicly is that performance problems might cost up to 9% of corporate revenue, at companies averaging 1.3 billion dollars of it. Eighteen years later this is quoted as a law about a checkout.

On mobile, every second of delay costs up to 20% of conversions.

Credited on every page that carries it to “Google Analytics data”, a “Google and SOASTA study” or simply “research from Google”. It is in no Google document that can be retrieved, including the two that the other mobile figures come from, and those two say something else: 123% more bounce between one second and ten, and 53% of visits abandoned past three. Nothing published supports a 20% conversion figure per second. It is refused the same way the Amazon number is.

Going from one second to three raises bounce probability 32%.

Real research, misquoted. Google published it in February 2017 and revised it in 2018, and the text gives one second to TEN seconds and 123%, not one to three and 32%. The 32% is from a chart beside that paragraph. The method matters more than either figure: both are the output of a deep neural network trained on bounce and conversion data at 90% prediction accuracy. A model's prediction over observational data, not a measurement and not an experiment. Google has since taken the page down, so it is in the register through an archived copy.

53% of mobile visitors leave if a page takes over three seconds.

This one is real, it is Google's, and almost everything carried with it is wrong. Its own footnote says: aggregated Google Analytics data from mobile sites that opted into sharing benchmark data, n=3.7K, global, March 2016. Observational, a sample that selected itself, ten years old, published by an advertising product for publishers, and counting abandoned visits rather than lost sales. It is quoted today as a causal fact about a shop's customers. Google has removed that page too.

This list had a second entry until the slides turned up.

Schurman and Brutlag, Velocity 2009, sat here because the only accounts of it were second hand and they disagreed about the size of the Bing delay: Arapakis and colleagues say 1.5 seconds in 2021, most secondary accounts say 2. O'Reilly still serves the original deck from the conference CDN. The Bing results inside it are a vector drawing whose text can be read back with its coordinates, and the row carrying the disputed figures is 2000 ms. The secondary accounts are the correct ones, Arapakis and colleagues are not, and the study is now the best dose-response evidence in this register rather than a refusal.

Against

The case against this page

A page that collects only the studies agreeing with it is marketing. So here is the strongest version of the case against the number this one produces. Two of the four objections come out of the register's own arithmetic, and neither is resolved anywhere above.

Bing does not agree with Bing

The central figure here, 0.60% of revenue per 100 ms, is Bing in 2013 at doses of 100 ms and 250 ms. Over 250 ms that implies 1.50%. The same company in 2009, running the same kind of experiment on the same product, measured nothing statistically significant at 200 ms and only 1.20% at 500 ms, which is twice the delay. Both are tier A. The likely reconciliation is that the 2009 test could not detect a small effect and that Bing was faster by 2013, so the same delay was a larger proportional change. That is a guess. Nothing published settles it, and the figure this page leads with is the one being contradicted.

The shape is not the one the band assumes

This model prices a gap at a flat rate per 100 ms. The only dose-response table in the field says the rate is not flat and does not fall off smoothly either: 0.24% per 100 ms over the first 0–500 ms, then 0.32% over 500–1000, then 0.15% over 1000–2000. It rises before it falls. So the three randomised results are not necessarily three independent estimates of one quantity; they may be one curve sampled at three starting speeds, and Zalando never published a starting speed at all. Read that way the band is a measure of how little is known rather than a confidence interval, and the error from using a flat rate depends on how slow the reader's site is, which this page does not correct for.

Four times someone measured nothing

Etsy delayed users by 200 ms and found it did not matter for them. Bing delayed the right-hand pane by 250 ms and found no detectable effect across nearly 20 million users. Rakuten ran the best-designed commercial experiment in this register and LCP, the metric this calculator reads, is not among the metrics that moved. A cosmetics retailer in Blue Triangle's own whitepaper spent three months getting 1.5 s faster and reported no result in conversion or revenue. None of these removes the effect. Together they say it is conditional on things nobody has enumerated.

Look at who published the evidence

One source in this register is peer reviewed. Of the randomised commercial results the band is built from, Zalando's is an engineering blog post, Vodafone's is a case study on Google's own site, and neither publishes a sample size, a duration or an assignment procedure. The observational study the market runs on states that brands where the effect was not significant were left out of the report, which means its coefficients are conditioned on an effect being found. A reader who discounts vendor-published results and excluded nulls is not being unreasonable, and this page would have very little left.

What would falsify this: a randomised speed experiment on a store, published with its method, its sample and its baseline, that finds an effect materially below 0.31% per 100 ms. There is no such experiment in the public record in either direction, which is the point of the two open questions below.

None of this is an argument for using a larger number instead. Every objection above cuts against the market figure harder than against this one, because the market figure is a single observational coefficient compounded across seconds it was never measured over. It is an argument for treating the result as a floor with a stated provenance, which is what the calculator says it is.

Limits

What this cannot tell you

These sit beside every result the calculator produces, not in a footnote under it.

  • Every causal source measured server delay. This tool reads LCP, because LCP is what field data exposes. That substitution is the largest single weakness in the exercise.
  • Bing's figure is from web search, Zalando's from fashion retail, Vodafone's from a telecoms landing page. None of them is your site.
  • The elasticity is local. The same authors write that if a site's performance or its audience changes, the impact may be different.
  • Not every millisecond counts the same. A 250ms delay to below-the-fold elements at Bing produced no detectable effect across nearly 20 million users.
  • A linear approximation is validated to roughly half a second. Past that the curve bends, and the doses Bing published show the second second costing about half what the first one did. Anything beyond that half second is reported as distance, never as money.
  • Etsy ran the same kind of experiment and found a 200ms delay did not matter for its users. The effect is real but not universal.

Two things still open

Both are known holes. Neither is closed by care on this page.

  1. Not one of the 15 sources publishes an interval on the figure quoted from it. 0 of them. The register records this per source and the answer never changes: a single number, no standard error, nothing a reader could use to say how sure anyone is. The 2009 slides come closest by marking which rows were not significant, and even there the dose-response table is readable as a shape and not weightable as a measurement. Every figure on this page is a point estimate, and the band around it is a spread between sources rather than a confidence interval.
  2. There is no randomised commerce experiment in the public record published with its method. Zalando is the closest and gives it one sentence. That absence is the largest gap in this field, and it is why the central figure on this page comes from web search and not from a shop.

Protocol

How this register was assembled

Everything above rests on which sources exist, so here is where they were looked for. This matters most for the things that are not here: a page can only claim an experiment has never been published if it says where it looked.

Searched between 2026-09-25 and 2026-10-04.
WhereWhat was looked forWhat it returned
ACM Digital Library: KDD, CHIIR, SIGIRRandomised latency or speed experiments reported with a methodKohavi et al. 2013, Kohavi et al. 2014, Arapakis et al. 2021, and Etsy's null result reported inside the first of those
NBER, SSRN, arXiv, Harvard and MIT working papersEconomics or business-school work on page speed and commercial outcomesNothing. This question has no literature in those venues
Management Science, Marketing Science, Information Systems ResearchThe same, in the journals that would publish it if it existedNothing
Engineering blogs at retailers and commerce platformsA/B tests on store speed, published by the team that ran themZalando 2018 and Shopify 2026. Zalando gives its experiment one sentence
web.dev case studiesCommercial speed results Google chose to publishVodafone 2021 and Rakuten 2022
Conference archives and web.archive.orgSources quoted everywhere whose page no longer existsThe Velocity 2009 slides on O'Reilly's CDN, two Google articles Google has removed, and the Aberdeen 2008 brief
Vendor whitepapersPublished null results, which nobody has an incentive to writeOne, in Blue Triangle's own document
Three separate attempts, 2026-09-25 to 2026-10-04A slowdown or speed-up experiment on a shop, published with its method, sample and baselineNothing, in any venue. This is the largest gap in the field and the reason the central figure here comes from web search

What a source had to meet

To enter the register at all, a figure has to be readable in a document that can be opened. Not quoted in a summary, not attributed in a blog post: read, with the sentence carrying it reproduced verbatim. That rule is what the 15 entries have in common and it is why four of them are here through an archived copy.

The tier is then decided by design alone, not by how interesting or how recent the result is. Speed assigned to users is tier A, speed that moved on its own is tier B, and a figure with no retrievable method is tier C and carries no money anywhere on this page.

A source is excluded only for being unreadable, never for disagreeing. That is why the null results and the study that halves this page's own central figure are both in the register rather than left out of it.

What this is not

It is not a systematic review in the formal sense. The searching was not logged as it happened, so there is no count of records screened and no two-reader agreement, and this table was written from what the search returned rather than from a protocol fixed in advance. Treat it as a reproducible trail rather than as a replication: every row can be re-run, and a reader who finds a source that belongs here and is missing has found a real defect.

The register

All 15 sources, in full

Each entry carries its design, where it was published, who was measured, how the quotes were checked, and what the study cannot support. Open an entry to read the sentences the figures come from.

Quotes appear in the language they were published in. A translated quote is no longer verbatim, and a reader checking it against the source would not find the sentence.

Tier A, causalRandomised controlled experimentPeer reviewed: Yes

Online Controlled Experiments at Large Scale

Kohavi, Deng, Frasca, Walker, Xu, Pohlmann (Microsoft), 2013

Published in
KDD '13, ACM SIGKDD
Who was measured
Microsoft Bing users; two treatment arms, 10% slowed by 100ms and 10% by 250ms, over two weeks
How the quotes were checked
Read in the source PDF
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

We recently ran a slowdown experiment where we slowed 10% of users by 100msec (milliseconds) and another 10% by 250msec for two weeks. The results showed that performance absolutely matters a lot today: every 100msec improves revenue by 0.6%.
Bing's server performance is now sub-second at the 95th percentile.
This quantification may change over time, as the site's performance and bandwidth standards improve.

What it cannot support

  • Web search, not a store. The transaction is a click on a result, not a checkout, so the number transfers to retail only as an analogy.
  • Measured against a sub-second baseline. The authors say directly that the quantification changes as a site's performance standards change.
  • Server delay specifically, not page weight, render time or layout stability.
  • The same authors report a contradicting result: Etsy found a 200ms delay did not matter for its users.
  • A single dose pair, 100ms and 250ms. The dose-response shape comes from the same company's 2009 slides, not from here.
Tier A, causalRandomised controlled experimentPeer reviewed: Yes

Seven Rules of Thumb for Web Site Experimenters

Kohavi, Deng, Longbotham, Xu (Microsoft), 2014

Published in
KDD '14, ACM SIGKDD
Who was measured
Bing users across several slowdown experiments at different delay amounts
How the quotes were checked
Read in the source PDF
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

a 250msec delay at the server impacts revenue at about 1.5% and clickthrough-rate by 0.25%. While this is a massive impact, 500msec would impact revenue about 3% not 20%, and clickthrough-rate would drop by 0.50%, not 20% (assuming a linear approximation is reasonable).
Earlier tests at Bing [32] had similar click impact and smaller revenue impact with delays of up to two seconds.
Using a linear approximation (1st-order Taylor expansion), we can assume that the impact of the metric is similar in both directions (slowdown and speedup) ... By running slowdown experiments with different slowdown amounts, we have confirmed that a linear approximation is very reasonable for Bing.
The slowdown quantifies the impact on the metric of interest at the point today ... If the site performance changes (e.g., site is faster), or the audience changes (e.g., more international users) the impact may be different.
A recent slowdown controlled experiment was run ... delaying when the right pane elements were shown by 250 milliseconds. If there was an impact on key metrics, it was not detectible, despite the experiment size of almost 20 million users.

What it cannot support

  • The linear approximation is validated for Bing, at Bing's speed, for Bing's audience. The authors say in the same paper that a different speed or audience may give a different impact.
  • The two-second observation is reported here without its numbers. They are in the Velocity 2009 slides from the same company, and they say the marginal cost of 100ms roughly halves beyond one second.
  • Still web search.
Tier A, causalRandomised controlled experimentPeer reviewed: No

Loading Time Matters

Zalando engineering, 2018

Published in
Zalando Engineering Blog (not peer reviewed)
Who was measured
Zalando customers; correlation across the journey, then an A/B test for confirmation
How the quotes were checked
Read in the source blog post
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

100 msec loading time improvement led to a 0.7% uplift in revenue per session.
We analyzed the correlations of loading time and revenue per session across every step of the user journey and for every device.
An A/B test brought the final confirmation.

What it cannot support

  • No baseline loading time, sample size, duration or assignment detail is published, so the experiment cannot be inspected.
  • The headline analysis is explicitly correlational; the A/B test is mentioned in a single sentence without numbers of its own.
  • Fashion retail in Europe. One business, one catalogue, one audience.
Tier A, causalRandomised controlled experimentPeer reviewed: No

Vodafone: A 31% improvement in LCP increased sales by 8%

Vodafone, published by Google on web.dev, 2021

Published in
web.dev case study (not peer reviewed)
Who was measured
50/50 traffic split on a landing page, roughly 100K clicks and 34K visits per arm per day
How the quotes were checked
Read in the source case study
What it says about precision
A single number, with no interval and no error term
Source
web.dev
Quotes and limits

In the authors' words

50% of the traffic was sent to the optimized landing page (version A), and 50% was sent to the baseline page (version B).
LCP improved by 31% (from 8.3s to 5.7s)
Sales increased by 8%
Lead-to-visit rate improved by 15%
Cart-to-visit rate improved by 11%

What it cannot support

  • The treatment changed server-side rendering AND images together, so speed is not isolated from the other changes. It is a randomised comparison of two page builds, not of two speeds.
  • Published by the vendor whose metric it validates, on Google's own site.
  • No duration or total sample is given beyond daily traffic.
  • 0.31% of revenue per 100ms here is an average across 2.6s from a slow baseline, and must not be compared like-for-like with a marginal first-100ms figure.
Tier A, causalRandomised controlled experimentPeer reviewed: No

Performance Related Changes and their User Impact (The User and Business Impact of Server Delays, Additional Bytes, and HTTP Chunking in Web Search)

Schurman (Microsoft Bing), Brutlag (Google), 2009

Published in
Velocity 2009 (conference presentation, no paper)
Who was measured
Bing and Google web search users, each site experimenting separately. Bing: a small number of users per delay level. Google: a small percentage of search traffic, randomly assigned, run for four to six weeks per level
How the quotes were checked
Read in the original slides
What it says about precision
Significance is marked, but no interval is given
Quotes and limits

In the authors' words

Determine impact of server delays
Delay before sending results
Different experiments with different delays
Randomly assign users to the experiment and control groups (A/B testing)
Server-side delay: Emulates additional server processing time
Varied type of delay, magnitude (in ms), and duration (number of weeks)
Strong negative impacts
Roughly linear changes with increasing delay
Time to Click changed by roughly double the delay
- Means no statistically significant change
Delays under half a second impact business metrics
The cost of delay increases over time and persists

What it cannot support

  • Slides and a talk, never a paper. The design is described well enough to understand and not well enough to replicate: no sample sizes, no confidence intervals, no statistical tests beyond a dash meaning not significant.
  • Bing in 2009, at a 2009 baseline that is nowhere stated. Its marginal rates run from 0.15% to 0.32% of revenue per 100ms, below the 0.6% the same company measured in 2013 at a sub-second baseline, which is what an elasticity measured at a different operating point looks like.
  • Web search, not commerce.
  • Resolved here, and worth recording: Arapakis et al. (2021) attribute the -1.8% queries, -4.3% revenue and -3.8% satisfaction to a delay of 1.5 seconds. In the original slides that row is 2000ms. The secondary accounts saying two seconds are the correct ones.
Tier A, causalRandomised controlled experimentPeer reviewed: No

A 200ms delay did not measurably matter at Etsy

McKinley (Etsy), reported in Kohavi et al. 2013, 2013

Published in
Reported in KDD '13
Who was measured
Etsy users
How the quotes were checked
Read in the source PDF
What it says about precision
Nothing. No figure precision is stated at all
Quotes and limits

In the authors' words

we were surprised to see Etsy's Dan McKinley [2] claim that a 200msec delay did not matter. It is possible that for Etsy users, performance is not critical

What it cannot support

  • Reported second-hand inside another team's paper, by authors who doubted it.
  • A null result, so it bounds how universal the effect is rather than measuring one.
Tier B, observationalObservationalPeer reviewed: No

A cosmetics retailer got 1.5 seconds faster and saw no result

Blue Triangle Technologies, 2020

Published in
Vendor whitepaper (not peer reviewed)
Who was measured
One anonymous cosmetics retailer, named nowhere in the document
How the quotes were checked
Read in the source PDF
What it says about precision
Nothing. No figure precision is stated at all
Quotes and limits

In the authors' words

They spent 3 months improving performance by 1.5 seconds on average, implementing A/B testing to ensure the site was production-ready. Upon launching the new “faster site”, they noticed no real results in conversion rates and top-line revenue because attention was given to the wrong webpages.
To support the argument that there is not a “one-size-fits-all” page speed goal to best support user experience improvements and overall sales, three different eCommerce retailers were analyzed using Blue Triangle data in this study.

What it cannot support

  • The A/B test it mentions was run to check the new store was production-ready, not to measure the speed effect. So this is a before and after across a relaunch, with nothing holding the rest of the site constant.
  • No sample, no duration of measurement, and no definition of “no real results”. The retailer is anonymous under a confidentiality agreement.
  • Published by a performance vendor, and the explanation it offers for the null result is the thing it sells.
  • It carries the Amazon 1% claim forward citing Linden 2006, which is the slide deck this register refuses that number for.
Tier A, causalRandomised controlled experimentPeer reviewed: No

Rakuten 24's investment in Core Web Vitals increased revenue per visitor by 53.37%

Hayoung Lee, Linh Duong, Ryunosuke Akiba, Shogo Kashiwase (Rakuten, published by Google on web.dev), 2022

Published in
web.dev case study (not peer reviewed)
Who was measured
Rakuten 24 landing page, concurrent 50/50 traffic split for one month; no sample size published
How the quotes were checked
Read in the source case study
What it says about precision
A single number, with no interval and no error term
Source
web.dev
Quotes and limits

In the authors' words

During the test duration, 50% of the traffic was sent to the optimized landing page (version A), and 50% was sent to the original page (version B).
The only difference between version A and version B was that version A was optimized for Core Web Vitals and there were no other functional or visual differences.
53.37% increase in revenue per visitor
33.13% increase in conversion rate
CLS improved by 92.72%
FID improved by 7.95%
FCP improved by 8.45%
TTFB improved by 18.03%

What it cannot support

  • LCP is not among the metrics reported to have improved. CLS, FID, FCP and TTFB are. So this is the value of a bundle dominated by layout stability, and it cannot be converted into a figure per 100ms of LCP.
  • A bundle, not a speed manipulation. It is a randomised comparison of two page builds, like Vodafone, except that here the largest movement is in a metric that is not speed.
  • No sample size, no confidence intervals, no statistical tests are published.
  • One landing page, one month, one retailer, published by the vendor of the metric it validates on that vendor's own site.
  • The widely quoted LCP figures from this page (61% better conversion under one second) are an observational segmentation of field data, not the experiment. Quoting them as the experiment's result is the same citation drift this register exists to catch.
Tier A, causalControlled user studyPeer reviewed: Yes

Impact of Response Latency on User Behaviour in Mobile Web Search

Arapakis (Telefónica Research), Park (Telefónica Research), Pielot (Google), 2021

Published in
CHIIR '21, ACM SIGIR Conference on Human Information Interaction and Retrieval
Who was measured
Mobile web search users in a controlled study with assigned latency levels
How the quotes were checked
Read in the source PDF
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

Our preliminary results indicate that mobile web search users are four times more tolerant to response latency reported for desktop web search users. However, when exceeding a certain threshold of 7-10 sec, the delays have a sizeable impact and users report feeling significantly more tensed, tired, terrible, frustrated and sluggish, all which contribute to a worse subjective user experience.

What it cannot support

  • Measures subjective experience and search behaviour, not revenue.
  • The authors call the results preliminary.
  • A threshold at 7-10s is specific to mobile search; it is evidence that a threshold exists, not a threshold to reuse elsewhere.
Tier A, causalControlled user studyPeer reviewed: Yes

Response time in man-computer conversational transactions

Miller (1968); extended by Card, Robertson & Mackinlay (1991), 1968

Published in
AFIPS Fall Joint Computer Conference, vol. 33, pp. 267-277
Who was measured
Human response-time perception, replicated across five decades of HCI work
How the quotes were checked
Reached through a paper citing it
What it says about precision
Nothing. No figure precision is stated at all
Quotes and limits

In the authors' words

0.1 second is about the limit for having the user feel that the system is reacting instantaneously
1.0 second is about the limit for the user's flow of thought to stay uninterrupted
10 seconds is about the limit for keeping the user's attention focused on the dialogue

What it cannot support

  • The three limits are quoted from Nielsen's codification, which cites Miller 1968 and Card et al. 1991. The 1968 proceedings are paywalled and were not read first-hand.
  • Perceptual limits, not revenue. They explain why a response curve has steps in it; they do not price one.
Tier B, observationalObservationalPeer reviewed: No

Milliseconds Make Millions

Deloitte with Google, 2020

Published in
Industry report (not peer reviewed)
Who was measured
37 retail, travel, luxury and lead-generation brands across Europe and the US, four weeks, mobile only
How the quotes were checked
Read in the source PDF
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

An 8.4% increase in conversions with retail consumers was observed, and an increase in average order value of 9.2%.
A 10.1% increase in conversions with travel consumers was observed, and a slight increase in average order value of 1.9%.
Once data had been collected and fed into a [l]ogarithmic regression model, we were in a position to analyse the data and extract meaning.
Fluctuations in speed all occurred naturally and were not artificially created on any of the sites.
The 0.1 second improvement was the cumulative impact of all four speed metrics.
In some instances, the increase in speed was observed to have minimal or zero effect on funnel progression and/or the KPIs for the vertical. These observations were not included in the report.

What it cannot support

  • Correlational. Speed was not assigned, so the figures cannot support a causal claim that a fix produces the gain.
  • Null results were excluded from publication, so the coefficients are conditioned on an effect being found and lean high.
  • The baseline speed the 0.1s was measured from is never published, which leaves the curve unidentifiable. Calibrated against plausible baselines the same data yields anything from +39% to +413% for one worked example.
  • The 0.1s is the cumulative movement of four metrics (Server Latency, Estimated Input Latency, First Meaningful Paint, Observer Load). LCP is not among them and did not exist in Lighthouse at the time.
  • Mobile only. The authors looked at desktop and 'found a lot of contradicting parameters'.
  • Luxury conversion is defined as adding to basket or clicking 'contact us', so the study never measured luxury revenue.
Tier B, observationalObservationalPeer reviewed: No

Store Speed and Conversion: What the Data Shows

Shopify, 2026

Published in
Shopify Enterprise blog (not peer reviewed)
Who was measured
Actively-selling Shopify stores, real-user performance data over a 28-day window at the turn of January and February 2026, with the slowest 5% excluded
How the quotes were checked
Read in the source blog post
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

for every 100 milliseconds slower a store loads, conversion tends to be about 3.5% lower.
stores with 2.5 second LCP report roughly 30% lower conversion than stores with 1.5 second LCP.
To prevent outliers from skewing the results, we excluded the slowest 5% of stores from the analysis.
Largest Contentful Paint (LCP) measures how quickly your main content appears.
At the individual store level, conversion depends on many factors beyond speed-product-market fit, pricing, marketing, and more.

What it cannot support

  • Observational. Stores were bucketed by the speed they already had, not assigned one, so the figure cannot support a claim that making a store faster produces the gain.
  • Grouping stores to average out product fit, pricing and marketing dilutes a confounder, it does not remove one. A store that invests in being fast is plausibly a store that invests in everything else.
  • Conversion, not revenue. It says nothing about order value, which is half of the other observational study's figure.
  • Published by the platform whose merchants it measures, with no sample size given.
  • It disagrees with the other observational study by a factor of 2.4 on the same quantity: 3.5% per 100ms here against 8.4% per 0.1s in Milliseconds Make Millions.
Tier B, observationalObservationalPeer reviewed: No

53% of mobile site visits are abandoned after three seconds

Shellhammer (Google / DoubleClick), 2016

Published in
DoubleClick by Google, “The need for mobile speed” (not peer reviewed)
Who was measured
Mobile sites that opted into sharing Google Analytics benchmark data, n=3.7K, global, March 2016
How the quotes were checked
Read in an archived copy, the original having been taken down
What it says about precision
A single number, with no interval and no error term
Quotes and limits

In the authors' words

In our new study, “The Need for Mobile Speed”, we found that 53% of mobile site visits are abandoned if pages take longer than 3 seconds to load.
Google Data, Aggregated, anonymized Google Analytics data from a sample of mWeb sites opted into sharing benchmark data, n=3.7K, Global, March 2016

What it cannot support

  • Observational. Nothing was assigned, so it cannot say that making a page faster would keep anyone.
  • The sample selected itself: sites that chose to share benchmark data with Google.
  • Mobile site visits, published by an advertising product for publishers. A visit abandoned is not a sale lost.
  • March 2016 data, on connection speeds and devices that no longer exist.
  • Google has removed the page. Only the archive copy remains.
Tier B, observationalObservationalPeer reviewed: No

Bounce probability rises 123% between one and ten seconds

An (Google), 2017

Published in
Think with Google, “Mobile page speed: new industry benchmarks” (not peer reviewed)
Who was measured
Mobile landing pages analysed by Google, revised February 2018
How the quotes were checked
Read in an archived copy, the original having been taken down
What it says about precision
Only the accuracy of the model that produced it
Quotes and limits

In the authors' words

We also trained a deep neural network—a computer system modeled on the human brain and nervous system—with a large set of bounce and conversions data. The neural net, which had a 90% prediction accuracy, found that as page load time goes from one second to 10 seconds, the probability of a mobile site visitor bouncing increases 123%.

What it cannot support

  • A neural network's prediction over observational data, not a measurement and not an experiment.
  • The widely quoted “one to three seconds, 32%” is not in the text. The text gives one to ten seconds and 123%.
  • Bounce is not a purchase. Nothing here measures revenue.
  • No sample size, no confidence interval and no description of the training data.
  • Google has removed the page. Only the archive copy remains.
Tier C, no primary sourceNo published methodPeer reviewed: No

Every 100ms of latency cost Amazon 1% in sales

Attributed to Amazon, via Greg Linden's slides and secondary press, 2006

Published in
Never written up
Who was measured
Unknown. No method, sample size or date has been published.
How the quotes were checked
Not verifiable
What it says about precision
Nothing. No figure precision is stated at all
Quotes and limits

In the authors' words

Greg Linden [40 p. 15] noted that 100msec slowdown at Amazon impacted revenue by 1%

What it cannot support

  • No published method, sample, duration or definition of revenue.
  • Traced through citations it ends at a 2006 slide deck and, in one peer-reviewed paper, at a Forbes column.
  • Excluded from every calculation here. It appears in the UI only as an example of a number that cannot be checked.

Citing this page: Quote it together with the study a figure rests on, not with this page alone. Every number below names its own source and links to it.

Let’s Create an Amazing Project Together!

Contact Me

Questions people ask about this

Is it true that 100 ms of latency cost Amazon 1% in sales?

There is no way to check it. Followed to its source the figure ends at a 2006 slide deck, and one peer-reviewed paper cites it to a Forbes column. The experiment behind it was never published, so there is no method to read, no sample to weigh and no definition of what counted as revenue. This page lists the claim so it can be refused by name, and takes nothing from it.

Is it true that 53% of mobile visitors leave after three seconds?

The sentence is real and it is Google's, from 2016. What travels with it is not. Its own footnote says the figure is aggregated Analytics data from mobile sites that opted in to sharing it, n=3,700, March 2016. So it is observational, the sample selected itself, it counts abandoned visits rather than lost sales, and it is ten years old. Google has since taken the page down.

Does a faster site always make more money?

No, and four published results say so. A 200 ms delay at Etsy produced no measurable effect on its own users. Bing delayed a side panel by 250 ms and saw nothing across nearly 20 million users. In the best-designed commercial experiment here, LCP is not among the metrics that moved at all. And one retailer spent three months getting 1.5 seconds faster and reported no change in conversion or revenue.

Which number should I take to a finance team?

The lower one, with the experiment named beside it. A figure from a randomised test survives being checked; a figure from a correlation does not, and the gap between them here is roughly thirty times. The calculator produces the first kind and prices only the part of a gap the evidence covers.

Why is there no confidence interval anywhere on this page?

Because not one source publishes one. Every figure in the register is a single number with no standard error, which is recorded source by source rather than mentioned once. The 2009 slides come closest by marking which results were not significant. The band shown with a result is the spread between separate studies, not a confidence interval.