Performance

CrUX and Lighthouse: understanding field and laboratory differences

Pingsivo · 1506 words · Updated 7 October 2026

English editorial adaptation prepared with AI assistance. Hypothetical examples are labelled. Technical sources accompany the relevant explanations.

Lighthouse investigates a page under a selected test configuration. CrUX aggregates eligible real-user experiences over a reporting period. Differences between them can be informative rather than contradictory. Check URL versus origin scope, device context, collection dates and metric definitions before comparing values or claiming that a correction has reached every visitor.

Two reports answer different questions

A laboratory run helps reproduce conditions and inspect diagnostics. It can be repeated while changing one factor, which makes it useful for investigating causes and verifying local improvements.

Field data reflects a wider set of real experiences represented by the source. It includes variation in devices, conditions and visits that one controlled run cannot reproduce completely.

The CrUX methodology describes eligibility and aggregation. Keep these limits visible instead of treating a field value as a census of every visitor or a lab run as a universal prediction.

Check URL and origin scope

A URL-level response concerns the requested page where eligible data is available. Origin-level information combines a broader scope. A fallback from one to the other changes the meaning of the result.

Do not use an origin aggregate to certify a particular page. A site can contain several templates and experiences, and their behaviour may differ. Display the scope prominently in reports and comparisons.

Pingsivo exposes the available scope rather than silently replacing missing URL evidence with a page-specific claim. Preserve that distinction when sharing or copying values into another document.

Read the collection period

Field values describe a period, not only the moment you click the button. A newly deployed correction may coexist in the reporting window with visits from before the release.

Record the release date and the report's collection dates. Compare appropriate periods and avoid declaring failure merely because a rolling aggregate did not change immediately.

A local run can provide quicker evidence about the changed implementation, but it cannot prove that the same benefit has already reached every represented visitor. Use each source for the question it can answer.

Understand percentiles

A percentile is not an arithmetic average. A p75 value describes a position in the distribution under the source's methodology. Relabelling it as “the average visitor” changes its meaning.

Keep metric, percentile, scope and period together. Comparing a lab median with a field p75 without explanation mixes different summaries and populations.

The distribution matters too. A central-looking headline can conceal a less favourable portion of experiences. Avoid presenting one rounded value as a complete description of every user journey.

Use local diagnostics to investigate causes

When field evidence suggests a problem, select representative pages and reproduce important scenarios. Inspect main content, layout behaviour and interactions rather than trying to repair an aggregate directly.

The LCP guide, INP guide and CLS guide explain different dimensions. A loading audit does not measure every interaction simply because it includes a performance score.

Record a hypothesis and test one correction. Keep the original report and method so that changes can be compared honestly.

Treat missing data correctly

No CrUX result does not mean zero delay, perfect performance or a failed website. Eligibility and sufficient data are required. Preserve an explicit unavailable state rather than fabricating a substitute value.

Use laboratory observations and real user feedback where field data is unavailable, with their own limits stated. Do not convert a simulated metric into a field metric merely to fill an empty card.

Configuration matters too. Pingsivo's CrUX integration requires the operator's Google API configuration. A missing key or external request error is different from a valid response indicating insufficient data.

Follow a correction over time

Track the changed pages, deployment date, local validation and subsequent field periods. Record other changes that could affect the interpretation, such as a redesigned template or altered content.

Avoid attributing every movement in an aggregate to one release. Visitor mix and conditions can also change. Use repeated evidence and a plausible relationship between the correction and the metric.

A useful report distinguishes what was reproduced locally, what the field source currently shows and what remains uncertain. This is more actionable than forcing two unlike numbers to agree.

Build a release follow-up record with separate evidence columns

For each important release, record the affected routes, deployment time and the behaviour the change was intended to improve. Add one column for local test evidence and another for field observations. In the local column, keep the tool, profile, repeated results and relevant diagnostics. In the field column, keep the source, metric, percentile, URL or origin scope and collection period. This separation prevents a report from accidentally combining a controlled experiment with a rolling aggregate and presenting the combination as a single directly measured outcome.

Consider a hypothetical article template whose main image becomes discoverable earlier after a code change. Repeated local tests can help establish whether the changed implementation improves that template under the selected profile. The first field response after deployment may still represent many earlier experiences. Record both observations without forcing an immediate agreement. The local result supports a scoped implementation claim; the field result describes its own period and population. A careful release note can explain why both are useful and what later observation will help assess the broader effect.

Check scope before comparing routes. If one URL has eligible data and another only returns origin-level information, the two cards do not describe the same population. Do not rank the individual pages as though they both have page-specific field measurements. Keep the unavailable state where appropriate and use representative laboratory investigations to examine the second page. An origin aggregate can still provide context, but its label should remain visible in screenshots, exports and any summary passed to someone who did not see the original interface.

Metric definitions also need to survive copying. A percentile should retain its name, and a laboratory measure should not be renamed to match a field metric merely because the units are similar. In particular, a loading audit's blocking observation is not a direct substitute for interaction data collected from real experiences. A table that preserves these distinctions may look less complete than one filled with invented equivalents, but it gives readers a reliable basis for deciding what is known and which additional investigation is justified.

When field data is unavailable, identify the reason as far as the interface genuinely knows. An external request error, missing operator configuration and insufficient eligible data are different states. They suggest different next steps and should not all become a red performance failure. Keep error messages understandable and avoid exposing private keys or internal details in shared reports. The purpose is to explain the evidence gap, not to manufacture a negative score for the website because the measurement source could not provide an answer.

For trend interpretation, note other releases or material content changes in the same period. An aggregate can move while several things change, and the represented mix of visits can also differ. This does not make field evidence useless; it means that causal attribution needs care. Use the implementation hypothesis, controlled reproduction and repeated field periods together. Avoid a precise claim that one small edit caused the entire observed change unless the evidence actually supports that level of attribution.

A practical review schedule should identify who checks the later field window and what decision follows. Perhaps the change can be retained after local verification while broader tracking continues. Perhaps an important user complaint requires further interaction testing even though the aggregate looks favourable. Define the next action from the actual product risk rather than waiting for two unlike tools to display matching numbers. Their differences can reveal useful questions about devices, states or journeys that the initial laboratory scenario did not cover.

The final report should therefore explain the correction, the local evidence, the current field scope and the remaining uncertainty. This gives stakeholders an accurate view of progress without promising an immediate universal improvement. It also creates a reusable record for future regressions: another maintainer can see which experiment demonstrated the implementation change and which later data described visitor experience, rather than trying to reconstruct the story from a pair of unexplained scores.

FAQ: comparing field and lab evidence

Why can Lighthouse look good while CrUX is less favourable?

The sources cover different conditions, populations and periods. Investigate scope and representative scenarios before treating the difference as a contradiction.

Does missing CrUX data mean my site is bad?

No. The source needs sufficient eligible observations. An unavailable result is not a performance verdict.

Can origin data certify a page?

No. Origin-level data covers a broader scope and may combine different templates. Label it explicitly and avoid page-specific claims it cannot support.

Is p75 the average visitor?

No. It is a percentile, not an arithmetic mean or a literal representative person. Preserve the statistical definition when explaining the result.

How long should I wait after a correction?

Read the source's actual collection period and follow comparable reporting windows. Local checks can provide immediate implementation evidence, while field aggregation reflects experiences over time.

Related guides and next steps