IN SUMMARY
Primary data as a scalpel, not a blanket
Primary data as a scalpel, not a blanket
Life cycle assessment used to be an analytical exercise: commission a study, get a result, answer a question. Now organisations embed LCA into procurement, product design, compliance and reporting, which means it has to function as infrastructure: reproducible, explainable and able to withstand scrutiny. When that happens, outcomes depend less on the impact method and more on whether the underlying data reflects how products are actually made and sourced. This paper sets out why a deliberate primary data strategy, focused where it changes decisions, is what separates defensible results from a false sense of rigour.
Life cycle assessment used to be an analytical exercise: commission a study, get a result, answer a question. Now organisations embed LCA into procurement, product design, compliance and reporting, which means it has to function as infrastructure: reproducible, explainable and able to withstand scrutiny. When that happens, outcomes depend less on the impact method and more on whether the underlying data reflects how products are actually made and sourced. This paper sets out why a deliberate primary data strategy, focused where it changes decisions, is what separates defensible results from a false sense of rigour.
Life cycle assessment used to be an analytical exercise: commission a study, get a result, answer a question. Now organisations embed LCA into procurement, product design, compliance and reporting, which means it has to function as infrastructure: reproducible, explainable and able to withstand scrutiny. When that happens, outcomes depend less on the impact method and more on whether the underlying data reflects how products are actually made and sourced. This paper sets out why a deliberate primary data strategy, focused where it changes decisions, is what separates defensible results from a false sense of rigour.
Impact is concentrated. Around 20% of processes typically drive the majority of a product's footprint, so investing in primary data everywhere adds cost without leverage; targeting the few upstream, energy-intensive or highly variable processes is where results actually change.
Impact is concentrated. Around 20% of processes typically drive the majority of a product's footprint, so investing in primary data everywhere adds cost without leverage; targeting the few upstream, energy-intensive or highly variable processes is where results actually change.
Impact is concentrated. Around 20% of processes typically drive the majority of a product's footprint, so investing in primary data everywhere adds cost without leverage; targeting the few upstream, energy-intensive or highly variable processes is where results actually change.
Stable results can mislead. When a model's numbers don't move year on year, that often reflects unchanged assumptions and ageing data rather than a robust, accurate picture of reality.
Stable results can mislead. When a model's numbers don't move year on year, that often reflects unchanged assumptions and ageing data rather than a robust, accurate picture of reality.
Stable results can mislead. When a model's numbers don't move year on year, that often reflects unchanged assumptions and ageing data rather than a robust, accurate picture of reality.
Regulatory hotspots are a floor, not a strategy. Predefined hotspots reflect where impacts tend to sit across a market, not within a specific supply chain, so true hotspots shift when operating conditions differ.
Regulatory hotspots are a floor, not a strategy. Predefined hotspots reflect where impacts tend to sit across a market, not within a specific supply chain, so true hotspots shift when operating conditions differ.
Regulatory hotspots are a floor, not a strategy. Predefined hotspots reflect where impacts tend to sit across a market, not within a specific supply chain, so true hotspots shift when operating conditions differ.
From analysis to infrastructure
For most of its history, LCA was a tool for answering a specific question with a one-off study, where bespoke assumptions and one-off datasets were tolerable. That no longer fits how organisations use it. Regulation, customer requirements, investor scrutiny and internal decision-making now demand product-level data applied consistently across portfolios and updated over time. Infrastructure cannot tolerate the inconsistency an analytical study could; it must be reproducible and able to withstand audit. As LCA becomes infrastructure, the centre of gravity shifts away from calculation, since standards and impact methods are a stable foundation, and toward the quality, specificity and traceability of the underlying data, especially for the upstream processes that dominate impacts.
For most of its history, LCA was a tool for answering a specific question with a one-off study, where bespoke assumptions and one-off datasets were tolerable. That no longer fits how organisations use it. Regulation, customer requirements, investor scrutiny and internal decision-making now demand product-level data applied consistently across portfolios and updated over time. Infrastructure cannot tolerate the inconsistency an analytical study could; it must be reproducible and able to withstand audit. As LCA becomes infrastructure, the centre of gravity shifts away from calculation, since standards and impact methods are a stable foundation, and toward the quality, specificity and traceability of the underlying data, especially for the upstream processes that dominate impacts.
Why generic and modelled data mislead decisions
Secondary and AI-generated datasets are not wrong; teams need them to build complete models at scale. The problem is using them to support decisions that hinge on differences between processes, suppliers or operating conditions. By design, these datasets smooth variability and average performance, which makes them stable but poorly suited to how a specific site or supply chain actually behaves. Three failure modes recur: hotspots get mis-ranked because the data embeds geographies, technologies and conditions that may not match reality; false confidence builds because clean, consistent numbers look precise even when assumptions are never revisited; and meaningful differences between suppliers get hidden by averaging, which matters most where regulatory thresholds determine market access. Introducing process-specific data at the right points breaks these patterns: hotspots shift, risk profiles change, and the differences between options become visible.
Secondary and AI-generated datasets are not wrong; teams need them to build complete models at scale. The problem is using them to support decisions that hinge on differences between processes, suppliers or operating conditions. By design, these datasets smooth variability and average performance, which makes them stable but poorly suited to how a specific site or supply chain actually behaves. Three failure modes recur: hotspots get mis-ranked because the data embeds geographies, technologies and conditions that may not match reality; false confidence builds because clean, consistent numbers look precise even when assumptions are never revisited; and meaningful differences between suppliers get hidden by averaging, which matters most where regulatory thresholds determine market access. Introducing process-specific data at the right points breaks these patterns: hotspots shift, risk profiles change, and the differences between options become visible.
Using primary data where it changes the answer
The case for primary data usually fails when teams frame it as all-or-nothing, either collecting everything or defaulting to generic datasets. Neither works. The effective approach is selective and iterative: build a baseline model with secondary and modelled data, use it to identify candidate hotspots and areas of high uncertainty, then target primary data exchange at exactly those points. Generic data guides direction, primary data resolves uncertainty, and the updated model refines priorities. Specificity alone is not enough, though; poor-quality primary data can mislead more than robust secondary data, so process-specific inputs need a clear boundary, a stated time and location, internal plausibility and traceability to the operator. Crucially, regulatory requirements for primary data at predefined hotspots should be treated as a minimum baseline, not the strategy, since compliance defines where data is required while strategy defines where it creates value.
The case for primary data usually fails when teams frame it as all-or-nothing, either collecting everything or defaulting to generic datasets. Neither works. The effective approach is selective and iterative: build a baseline model with secondary and modelled data, use it to identify candidate hotspots and areas of high uncertainty, then target primary data exchange at exactly those points. Generic data guides direction, primary data resolves uncertainty, and the updated model refines priorities. Specificity alone is not enough, though; poor-quality primary data can mislead more than robust secondary data, so process-specific inputs need a clear boundary, a stated time and location, internal plausibility and traceability to the operator. Crucially, regulatory requirements for primary data at predefined hotspots should be treated as a minimum baseline, not the strategy, since compliance defines where data is required while strategy defines where it creates value.
Making it scale: tools, suppliers and transparency
Selective primary data only works if organisations can exchange, reuse and govern it across complex supply chains, and this is where most efforts stall. Manual questionnaires and one-off studies do not scale; they create friction for suppliers and produce data that cannot be reused. The shift needed is from data extraction to structured, permissioned, repeatable data exchange, supported by a supplier operating model built on a minimal viable ask, clear governance and genuine incentives. A common blocker is fear of losing control over sensitive information, but effective transparency is not full disclosure; it is sharing methods, boundaries, data quality indicators and explanations of change while keeping raw values confidential, separating evidence from disclosure. Minviro's Primary Data Share metric helps make progress visible, indicating the proportion of total impact modelled with quality-assured process-specific data. Tools are what determine whether all this becomes repeatable rather than ad hoc, which is the role XYCLE is built to play: connecting exchanged primary data across standards and platforms, preserving metadata, and turning it into consistent, auditable, decision-grade models.
Selective primary data only works if organisations can exchange, reuse and govern it across complex supply chains, and this is where most efforts stall. Manual questionnaires and one-off studies do not scale; they create friction for suppliers and produce data that cannot be reused. The shift needed is from data extraction to structured, permissioned, repeatable data exchange, supported by a supplier operating model built on a minimal viable ask, clear governance and genuine incentives. A common blocker is fear of losing control over sensitive information, but effective transparency is not full disclosure; it is sharing methods, boundaries, data quality indicators and explanations of change while keeping raw values confidential, separating evidence from disclosure. Minviro's Primary Data Share metric helps make progress visible, indicating the proportion of total impact modelled with quality-assured process-specific data. Tools are what determine whether all this becomes repeatable rather than ad hoc, which is the role XYCLE is built to play: connecting exchanged primary data across standards and platforms, preserving metadata, and turning it into consistent, auditable, decision-grade models.





