Method for establishing spatial LCA data set

By clearly defining the dataset objectives, collecting geographic location information, generating a grid system, and synthesizing the data, the problem of the lack of spatial resolution in LCA datasets was solved, achieving efficient construction of spatial LCA datasets and improving the scientific rigor and applicability of the datasets.

CN121682256APending Publication Date: 2026-03-17QINGDAO INST OF BIOENERGY & BIOPROCESS TECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511693857.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing LCA datasets lack spatial resolution, making it difficult to reflect the environmental impact of different geographical locations, thus limiting their application in region-specific analysis and policy making.

Method used

By clearly defining the dataset objectives, basic unit process data containing geographic location information is collected. A grid system is generated using spatial partitioning rules, and data matching and synthesis are performed to generate spatially representative synthetic unit process data. Finally, the data is stored in a standardized format.

Benefits of technology

It improves the spatial representation capabilities and applicability of the LCA dataset, supports multi-scale spatial analysis, ensures the scientific validity and usability of the data, and facilitates querying and visualization applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121682256A_ABST
    Figure CN121682256A_ABST
Patent Text Reader

Abstract

The invention provides a method for establishing a spatial LCA data set, and relates to the technical field of life cycle evaluation and environment management, and the method comprises the steps: determining a data set target and a corresponding product / service, and defining a system boundary, a region boundary, a time boundary and a data quality requirement; then collecting basic unit process data including input / output streams and geographic positions; making a space division rule (including a space range, a grid shape and granularity) according to a region boundary and an application scene, and generating a grid system; attributing the basic unit data to the corresponding grids through space matching; aggregating the data in the grids according to a weighted average algorithm and other synthesis rules to generate synthesis unit process data; associating the grid geographic position identification to form grid unit process data; and finally, storing the original data and processing records in standard formats such as ILCD / EcoSpold2 and the like. The time-space consistency, traceability and multi-scene adaptability of data are ensured through a systematic process, and the method is suitable for the fields of environmental evaluation, carbon accounting and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of life cycle assessment and environmental management, and particularly relates to a method for establishing a spatial LCA dataset. BACKGROUND

[0002] Life cycle assessment (LCA) is a widely used tool for evaluating the environmental impact of products, services or systems throughout their entire life cycle. Traditional LCA datasets are usually based on global or regional average data, lacking spatial resolution and being difficult to reflect the impact of resource distribution, environmental conditions, technology level and policy differences in different geographical locations. For example, the same product may have significant differences in energy consumption, emission characteristics and resource efficiency in the production process in different regions. Existing LCA data integration methods have many shortcomings in spatial data integration, such as scattered data sources, unclear spatial boundaries, inconsistent data granularity, difficult spatial matching, and missing synthesis rules, etc. This leads to weak spatial representation of existing data sets, limiting the application of LCA in regional-specific analysis, policy making and fine-grained environmental management. Therefore, there is an urgent need in the field for a method that can systematically and standardized establish LCA datasets with spatial resolution. SUMMARY

[0003] The application provides a method for establishing a spatial LCA dataset to solve one of the above technical problems.

[0004] The technical scheme adopted by the application is: The application provides a method for establishing a spatial LCA dataset, comprising: S1, clearly defining the target of the dataset to be established, determining the products or services contained according to the target, and determining the data requirements corresponding to the target; the data requirements include system boundary, regional boundary, time boundary, technical representativeness requirement and data quality requirement; S2, based on the data requirements, collecting basic unit process data corresponding to the products or services; the basic unit process data at least includes description information, input flow, output flow, and geographical location information of the basic unit process; S3, according to the regional boundary and the expected application in the data requirements, determining the spatial division rule, the spatial division rule including spatial range, grid shape and spatial granularity; based on the spatial division rule, dividing the regional boundary to generate a grid system covering the spatial range, the expected application being a specific scenario in which the dataset is planned to be used; S4, spatially matching the geographical location information of each basic unit process data with the grid system, so that each basic unit process data is attributed to at least one grid; S5. For each grid, aggregate and calculate all the basic unit process data belonging to it according to a predetermined data synthesis rule to generate the synthesized unit process data of the grid; the data synthesis rule includes a weighted average algorithm for input and output streams; S6. Associate the geographical location identifier of the grid to which each synthesized unit process data belongs to form the grid unit process data; S7. Store the basic unit process data and / or the grid unit process data together with the description information and data processing records in a specified data format.

[0005] According to an embodiment of the present application, in step S2, the data granularity of the basic unit process data includes process-level data, enterprise-level data or regional-level data, and the regional-level data includes provincial, municipal and county-level data; the type of the geographical location information includes point data, line data or surface data; and the expression of the geographical location information includes geographical coordinates based on a standard coordinate system, geographical coding or a self-defined geographical coordinate system.

[0006] According to an embodiment of the present application, in step S3, the spatial granularity includes administrative unit granularity, regular grid granularity or self-defined grid granularity.

[0007] According to an embodiment of the present application, in step S3, the grid shape is a regular grid or an irregular grid.

[0008] According to an embodiment of the present application, in step S3, the grid system is digitally expressed and stored in a vector format or a raster format.

[0009] According to an embodiment of the present application, in step S4, the spatial matching includes spatial calculation; when the geographical location of a basic unit process data is located at the boundary of a grid, the final belonging grid is determined according to a preset grid belonging rule.

[0010] According to an embodiment of the present application, in step S5, the predetermined data synthesis rule further includes a blank grid filling rule; for a blank grid without basic unit process data, at least one of the following methods is used to generate filling data to establish a spatial LCA data set: using the average value of the upper-level regional data to fill, using the data of the spatial adjacent grid to fill, or using a spatial interpolation method to calculate; and the filling method used is recorded.

[0011] According to an embodiment of the present application, in step S5, before data synthesis, a step of standardizing the basic unit process data is further included to ensure that all data are synthesized under the unified dimension and system boundary.

[0012] According to one embodiment of the present application, the description information includes at least one of the following: name, location, technical description, data processing method, data source, time representation description, spatial division purpose, division rule, applicable range, applicable scene suggestion, and synthesis relationship between basic unit process data and grid unit process data.

[0013] According to one embodiment of the present application, in step S7, the specified data format is ILCD format, EcoSpold2 format, JSON-LD format or database table structure.

[0014] As the above technical solutions are adopted, the present application has the following beneficial effects: The present application defines the data set target and defines the system boundary, regional boundary, time boundary, technical representation and data requirement in step S1, ensures that the construction of the data set has clear guidance and consistency, avoids data redundancy and boundary ambiguity from the source, and improves the pertinence and reliability of the data set.

[0015] Based on the data requirement, the basic unit process data containing geographic location information is collected, which not only covers the input and output flow and description information, but also introduces the spatial dimension, provides a data basis for subsequent spatial analysis and gridding processing, and enhances the spatial representation ability of the data set.

[0016] According to the regional boundary and the expected application, the spatial division rule (such as spatial range, grid shape and spatial granularity) is determined, and a grid system covering the whole region is generated, so that the data set can adapt to the needs of different application scenarios, support multi-scale spatial analysis, and improve the flexibility and applicability of the data set.

[0017] By spatially matching the geographic location of the basic unit process data with the grid system, it is ensured that each data unit belongs to at least one grid, realizing the accurate association of data and spatial units, providing a structural basis for subsequent data aggregation, and reducing spatial errors.

[0018] Using a predetermined data synthesis rule (such as a weighted average algorithm), the basic unit process data within the grid is aggregated and calculated to generate synthesized unit process data, ensuring the representativeness and consistency of the data within the spatial unit, avoiding information distortion, and improving the scientificity and usability of the data set.

[0019] Associating the geographic location identifier of each synthesized unit process data with the grid to which it belongs to form grid unit process data, realizing the standardized binding of data and spatial location, facilitating subsequent spatial query, analysis and visualization application.

[0020] Storing basic unit process data and / or grid unit process data, along with descriptive information and data processing records, in a standardized format not only supports long-term data preservation and sharing but also ensures the transparency and traceability of the data processing process, which is beneficial for data quality control and auditing. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating a method for establishing a spatial LCA dataset, as provided in an embodiment of this application. Detailed Implementation

[0022] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.

[0024] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

[0025] Example 1 like Figure 1 As shown, a method for establishing a spatial LCA dataset includes: S1. Define the objectives of the dataset to be built, determine the products or services included based on the objectives, and determine the data requirements corresponding to the objectives; the data requirements include system boundaries, regional boundaries, time boundaries, technical representativeness requirements, and data quality requirements.

[0026] As mentioned above, this step is the basis and starting point of the entire method, and its core is to establish a clear framework and constraints for all subsequent data operations through the principle of "target-oriented". Specifically, "defining the target of the established data set" means defining the final use and core problem to be solved of the spatial LCA data set, which determines the construction direction and refinement level of the data set. Subsequently, "determining the products or services contained according to the target" is to concretize the abstract target into one or more functional units that can be described by a life cycle model. Further, "determining the data requirements corresponding to the target" is a series of technical specifications set for achieving the target and ensuring the quality of the data set. Among them: System boundaries specify which unit processes should be included or excluded in the life cycle stages from raw material acquisition to final disposal ("cradle to grave") or to the factory gate ("cradle to gate").

[0027] Regional boundaries define the geographical scope covered by the data set, which is the physical basis for the entire spatial division.

[0028] Time boundaries clarify the time period or time point represented by the data, ensuring the temporal representativeness of the data.

[0029] Technical representativeness requirements specify the technical level that the collected data should reflect, such as representing the best available technology, average technical level, or specific process technology.

[0030] Data quality requirements are quantitative or qualitative indicators for the uncertainty, reliability, completeness, etc. of the data itself, such as the priority, age, geographical representativeness of the data source, etc.

[0031] S2, based on the data requirements, collect basic unit process data corresponding to the product or service; the basic unit process data at least includes description information, output stream, and geographical location information of the basic unit process.

[0032] As mentioned above, this step is the entity data foundation step of building the spatial LCA dataset, and its core is to carry out directional and selective data collection work according to the programmatic document of "data requirements" established in step S1. This process converts abstract data specifications into specific data entities, namely "basic unit process data". Each basic unit process data represents a minimum and indivisible activity link in the product life cycle chain (such as the mining of a raw material, the manufacturing of a component, the transportation of a specific distance, etc.). Its content must include: 1) description information, which is used to identify and explain the basic properties of the process; 2) input flow, which lists all resources, raw materials and energy (such as iron ore, electricity, natural gas) entering the process; 3) output flow, which lists all products, by-products and various substances discharged into the environment (such as steel, waste heat, carbon dioxide, wastewater) produced by the process; 4) geographic location information, which is the key to spatializing the dataset, and accurately records the geographic location of the unit process. This step ensures that the collected data not only meets the technical requirements of the target, but also has the ability to be recognized and matched by the subsequent grid system in the spatial dimension.

[0033] S3, according to the region boundary and the expected application in the data requirements, determine the spatial division rule, the spatial division rule includes spatial range, grid shape and spatial granularity; divide the region boundary based on the spatial division rule, generate a grid system covering the spatial range, the expected application is a specific scenario in which the dataset plan is used.

[0034] As mentioned above, this step is the key conversion step of realizing the LCA dataset from "having spatial information" to "real spatialization", and its core is to convert the macroscopic "region boundary" defined in S1 into a structured and calculable "grid system". This conversion is not arbitrary, but strictly follows the "spatial division rule" driven by the "expected application". The "spatial division rule" is a set of technical decisions, mainly including: 1) spatial range, which is the specific geographic area covered by the grid system, this range is directly inherited from the region boundary in the data requirements, and ensures that the grid system can completely cover all potential data geographic locations; 2) grid shape, which defines the geometric shape of the basic spatial unit that constitutes the grid system, and its selection directly affects the convenience of spatial analysis and the intuitiveness of the results; 3) spatial granularity, which defines the size of each grid unit, namely the spatial resolution, which determines the degree of spatial detail that the dataset can reveal. Finally, based on this set of explicit rules, the region boundary is discretized through spatial computing technology, generating a seamless and digital grid system covering the entire spatial range, laying the foundation for subsequent allocation of discrete data geographic locations to a unified spatial framework.

[0035] S4. Spatial matching of the geographical location information of each basic unit process data with the grid system, so that each basic unit process data belongs to at least one grid.

[0036] As described above, this step is the core spatial computation step in constructing the spatial LCA dataset. Its purpose is to establish a precise spatial association between the discretely distributed basic unit process data and the regularized grid system constructed in step S3. This process is achieved by performing "spatial matching," which utilizes spatial relationship calculations in the geographic information system to compare and determine the location and relationship of the geographic location information (whether it is a point, line, or area) carried by each basic unit process data with each grid cell in the grid system. Its core output is to assign one or more explicit grid identifiers to each basic unit process data, thereby realizing that "each basic unit process data belongs to at least one grid." This step systematically organizes the chaotic raw data, which only has its own location information, into a unified and structured spatial framework, laying an indispensable foundation for subsequent data aggregation based on grid units. In particular, when the geographic location of a basic unit process data happens to be located on a grid boundary, this step predefines explicit "grid affiliation rules" to determine its final affiliation, ensuring the consistency and repeatability of the matching process. When the geographic location is a regional range or linear coordinate, it may cross grids, which also requires rule processing.

[0037] S5. For each grid, aggregate and calculate all the basic unit process data to which it belongs according to a predetermined data synthesis rule to generate the synthesized unit process data of that grid; the data synthesis rule includes a weighted average algorithm for input and output streams.

[0038] As described above, this step is the core data processing step in constructing a spatialized dataset. Its purpose is to merge discrete "basic unit process data" belonging to the same grid into a spatially representative "synthetic unit process data." This process is not a simple data accumulation, but a scientific aggregation and calculation based on "predetermined data synthesis rules." Its core logic is to treat a grid as a "virtual composite unit process," representing the average level or typical characteristics of all similar activities within that specific geographical area. The "data synthesis rules" are crucial for ensuring the scientific validity, consistency, and comparability of the synthesis results. Their core includes, but is not limited to, a "weighted average algorithm for input and output flows." This algorithm identifies key attributes (such as product output, output value, number of employees, etc.) that contribute differently to the synthesis results in each basic unit process data, and uses these attributes as weights to perform a weighted average calculation on various input (such as raw materials, energy) and output (such as products, pollutants) flows of that unit process, thereby generating normalized synthetic unit process data for that grid under a unified functional unit. This step enables the transformation of data from individual to regional, from heterogeneous to homogeneous, and from dispersed to integrated, and is key to generating standardized data units that can be directly used for spatial LCA calculations.

[0039] S6. Associate each of the synthesized unit process data with the geographic location identifier of its respective grid to form grid unit process data.

[0040] As mentioned above, this step is the encapsulation and identification stage for constructing the spatial LCA dataset. Its core function is to assign a clear and unique spatial identity to the "synthetic unit process data" generated in step S5, which is still at the logical level. This process is achieved through a "association" operation, that is, binding each piece of synthetic unit process data to the "geographical location identifier" of its source grid. This association is not a simple data parallelism, but rather the establishment of an inherent link relationship that can be recognized and queried by computers. The resulting "grid unit process data" becomes a complete and independent data entity, which includes both the lifecycle inventory data (i.e., input / output stream) after aggregating all basic data within the grid, and its "registration" information in the spatial grid system. This step fundamentally solves the problem of the traditional LCA data and spatial information being "two separate entities," enabling data users to directly identify, call, and compare environmental performance data from different spatial locations without performing complex spatial calculations. It is the final step in generating standardized, spatially explicit data products that can be directly applied to spatial LCA models.

[0041] S7. Store the basic unit process data and / or the grid unit process data, together with their description information and data processing records, in accordance with the prescribed data format.

[0042] As described above, in step S7, storing the basic unit process data and / or the grid unit process data, along with their descriptive information and data processing records, in a prescribed data format is a crucial step in ensuring that the spatial LCA dataset possesses standardization, reusability, and full lifecycle traceability. Specifically, this step requires systematically archiving the raw data, synthetic data, and their additional information processed in steps S1-S6 in a structured format that conforms to industry standards, thereby meeting the core requirements of LCA analysis for data consistency, transparency, and interoperability.

[0043] According to one embodiment of this application, in step S2, the data granularity of the basic unit process data includes process-level data, enterprise-level data, or regional-level data, and the regional-level data includes provincial, municipal, and county-level data; wherein, the type of the geographic location information includes point data, line data, or area data; the expression method of the geographic location information includes geographic coordinates based on a standard coordinate system, geographic coding, or a custom geographic coordinate system.

[0044] As described above, in step S2, the data granularity of the basic unit process data specifically refers to the degree of refinement of the data at the spatial or organizational level. Specifically, process-level data describes the resource consumption and emission characteristics of specific operational steps in production activities (such as casting, welding, and painting), and is suitable for micro-level process optimization scenarios; enterprise-level data reflects the environmental load of a specific enterprise's overall production process and is typically used for enterprise carbon footprint accounting or cleaner production assessment; regional-level data is divided according to administrative levels. For example, provincial-level data can characterize the energy structure characteristics of a province, city-level data can be used for urban circular economy planning, and district / county-level data can be refined to the street-level pollution source distribution analysis.

[0045] For geographic location information, its type and expression method need to be flexibly adapted according to the data application scenario: point data (such as the coordinates of a company's sewage outlet) is suitable for accurately locating a single object; line data (such as highway transportation routes or river courses) is used to describe the spatial distribution of linear elements; and area data (such as administrative boundaries or ecological functional zoning) is used to characterize the attribute features within a region. In terms of expression methods, geographic coordinates based on standardized coordinate systems (such as WGS84 latitude and longitude) ensure the universality of data within GIS systems; geocoding is a coding system designed to identify the location and attributes of points, lines, and areas. It quantifies and records entity attributes and geometric coordinates in computer storage devices through a pre-defined classification system, which helps in the design of geographic information matching algorithms; custom coordinate systems (such as relative coordinates with the project's starting point as the origin) are suitable for internal data management in specific engineering or research areas. The combination of the above granularity classification and geographic information expression methods can support the needs of multi-scale environmental impact analysis from micro-processes to macro-regions, while ensuring the traceability and interoperability of data in the spatial dimension.

[0046] According to one embodiment of this application, in step S3, the spatial granularity includes administrative unit granularity, regular grid granularity, or custom grid granularity.

[0047] As described above, in step S3, the spatial granularity division methods specifically include the following three categories: Administrative unit granularity: This refers to using the existing administrative boundaries of a country or region (such as provinces, cities, districts, counties, etc.) as the basic unit of spatial division. This type of granularity is usually directly related to government statistical systems and environmental management policies, and is suitable for scenarios that need to be integrated with administrative management systems, such as regional carbon emission accounting and ecological function zoning assessment.

[0048] Regular grid granularity: This refers to dividing a target area into spatial units using regular grids of fixed sizes (such as 1km×1km, 500m×500m square or hexagonal grids). For example, in environmental monitoring, air quality data for a city can be divided into 1km×1km grid cells, with the average pollutant concentration calculated within each grid. This type of granularity is suitable for analyses requiring uniform spatial resolution and facilitates spatial overlay, trend prediction, or model calculations using GIS tools.

[0049] Custom grid granularity: This refers to an irregular grid division method designed according to actual needs. For example, a high-density monitoring grid (e.g., 50m × 50m) can be drawn around an industrial park, while a larger grid (e.g., 1km × 1km) can be used in areas far from the industrial zone, forming a hybrid grid system with dynamically adjusted spatial resolution. This type of granularity is often used in scenarios that require a balance between accuracy and efficiency, such as refined modeling around key pollution sources or differentiated analysis of cross-regional ecological corridors.

[0050] The three spatial granularities mentioned above can be flexibly selected according to the data application: administrative unit granularity emphasizes consistency with management boundaries, regular grid granularity focuses on the standardization of spatial analysis, and custom grid granularity focuses on refined coverage of key areas. The conversion and fusion between different granularities (such as mapping administrative unit data to regular grids) can further support the needs of multi-scale spatial analysis.

[0051] According to one embodiment of this application, in step S3, the mesh shape is a regular mesh or an irregular mesh.

[0052] As mentioned above, the grid shape division method includes two categories: regular grids and irregular grids. The regular grid refers to grid cells with uniform shape and area, such as squares, rectangles, or regular hexagons. Its advantage lies in its regular structure, which facilitates spatial calculations, statistical analysis, and visualization. It is especially suitable for scenarios that require uniform spatial resolution for model calculations (such as environmental diffusion models) or for direct overlay with raster data such as remote sensing data.

[0053] The irregular grid refers to grid cells with variable shape and size, typically based on existing administrative boundaries (such as provincial, municipal, and county boundaries) or natural geographical boundaries (such as watersheds and ecological functional zones). Its advantage lies in its high degree of alignment with existing socio-economic statistical data and management policy units, facilitating data acquisition and analysis-based decision-making within administrative units.

[0054] According to one embodiment of this application, in step S3, the grid system is digitally represented and stored using a vector format or a raster format.

[0055] As described above, the digital representation and storage format of the grid system specifically includes the following two methods: Vector format: This refers to a digital method that describes the grid structure using geometric objects such as points, lines, and surfaces, and their topological relationships. For example, a regular hexagonal grid system can be stored as vector polygons, where each hexagon is defined by its vertex coordinates (e.g., WGS84 latitude and longitude), and metadata such as grid number, area, and center coordinates is recorded through an attribute table. The advantages of vector format are high data precision, support for lossless scaling, and suitability for scenarios requiring precise spatial positioning and complex geometric operations (such as the overlay analysis of administrative boundaries and grids, and the intersection calculation of road networks and grids). Its storage efficiency is also high, making it particularly suitable for scenarios with a small number of grid cells but requiring frequent updates or dynamic adjustments (such as gridded land use zoning in urban planning).

[0056] Raster format: This refers to a digitization method that uses a regularly arranged matrix of pixels to represent a grid system. For example, a 1km × 1km square grid system can be stored in GeoTIFF format, where each pixel value corresponds to an attribute value of a grid cell (such as population density or pollution emissions), along with geographic coordinate information (such as top-left corner coordinates, resolution, and projection parameters). The advantage of raster format lies in its high spatial analysis efficiency, making it suitable for batch processing of continuously distributed data (such as remote sensing image classification and DEM terrain modeling). It also offers large storage space, making it particularly suitable for high-resolution grids (such as 50m × 50m) or scenarios covering large areas (such as nationwide climate data gridding).

[0057] The choice between the two formats needs to be weighed based on actual requirements: vector format is suitable for scenarios requiring precise geometric description and dynamic management, while raster format is suitable for scenarios requiring continuous data modeling and efficient spatial computation. In terms of technical implementation, the two formats can be converted to each other through algorithms (e.g., interpolation filling is used when converting vector grids to raster, and boundaries are extracted and topological relationships are reconstructed when converting raster data to vector), thereby supporting multi-source data fusion and cross-platform data sharing (e.g., overlaying and analyzing vector format administrative division grids with raster format remote sensing data).

[0058] According to one embodiment of this application, in step S4, the spatial matching includes spatial calculation; when the geographical location of a basic unit process data is located at the grid boundary, its final grid affiliation is determined according to a preset grid affiliation rule.

[0059] As mentioned above, the core of spatial matching lies in determining the correspondence between basic unit process data and the grid system through spatial computation, specifically including the following two parts: Spatial computation refers to matching the geographical locations (such as point coordinates, line features, and area boundaries) of basic unit process data with a grid system through geometric operations or spatial analysis algorithms. For example, when the coordinates of a company's emission point (such as WGS84 latitude and longitude) need to be mapped to a 1km × 1km regular grid, the grid cell to which it belongs is determined by calculating whether the coordinates fall within the geometric range of the target grid (such as determining whether the point is within a polygon of a hexagonal or square grid). For linear or area data (such as highway transportation routes or industrial zone boundaries), spatial intersection algorithms are used for assignment determination. For example, for linear data, the assignment or data allocation can be based on the proportion of the length of the part intersecting with the grid to the total length of the line; for area data, the assignment or data allocation can be based on the proportion of the area of ​​the part intersecting with the grid to the total area of ​​the area.

[0060] Grid Assignment Rules: When the geographical location of basic unit process data happens to be located at a grid boundary (e.g., point coordinates fall on the boundary line between two adjacent grids), assignment conflicts need to be resolved using preset rules. For example, if the coordinates of a factory's discharge outlet are located at the boundary between grids A and B, the assignment can be determined according to the following rules: Administrative boundary priority principle: If the grid division is consistent with the administrative division (such as the grid of a municipal district), then the grid to which the point belongs is determined by the administrative division to which it belongs (such as prioritizing the matching of the administrative grid of the street to which it belongs). Priority principle of natural geographical features: If the grid division is based on natural elements (such as rivers and mountains), then the boundary line of the geographical features shall be used as the basis for classification (such as classifying points into the grid on the east side of the river). Algorithm allocation rules: Dynamic allocation is achieved through mathematical methods, such as allocating weights according to the proportion of the distance from a point to the center of an adjacent grid (e.g., if the point is closer to the center of grid A, it is assigned to grid A), or according to the proportion of the length of the boundary line where the point is located (e.g., if the point is located in the middle of the boundary line between two grids, it is allocated to the two grids at a 50% ratio).

[0061] According to an embodiment of this application, in step S5, the predetermined data synthesis rule further includes a blank grid filling rule; for blank grids that do not have basic unit process data, at least one of the following methods for establishing a spatial LCA dataset is used to generate filling data: filling with the average value of the data of the previous level region, filling with the data of spatially adjacent grids, or calculating using a spatial interpolation method; and the filling method used is recorded.

[0062] As described above, the pre-defined data synthesis rules include the following three methods for filling blank grids, and the methods used must be recorded: Filling with the average value of data from the next higher level region: When a grid cell lacks basic unit process data (such as no enterprise emission data collected in the grid), but its parent region (such as city or province) has complete data, the average value of all grid data in the parent region can be calculated and used as the filling value for the blank grid.

[0063] Spatially adjacent grid data filling is primarily used to address the problem of missing grid-level data due to incomplete data collection, i.e., the absence of any basic unit process data within a certain grid. In this case, the following methods can be used to generate filling data for that grid: identify neighboring grids with valid data around the blank grid, and calculate their representative values ​​as the filling result based on the principle of spatial proximity. For example, the "nearest neighbor method" can be used to directly retrieve the synthetic unit process data from the nearest neighboring grid, or the "distance-weighted average method" can be used to calculate weights inversely proportional to the distance between the blank grid and multiple neighboring grids, and then perform a weighted average of the synthetic data from the neighboring grids. This method is suitable for scenarios where product data distributions in local areas are similar (such as areas with similar technical levels, processes, or time periods), but the proximity range (such as the search radius or the number of neighboring grids) and weight calculation rules need to be predefined.

[0064] In addition, another dimension of data incompleteness may be encountered during the data synthesis stage, namely, missing input or output stream data items within the basic unit process data. For such problems, they should be filled in according to predefined rules during the standardization process or synthesis calculation in step S5, such as using the average or median of similar data, or calculated values ​​based on the process model.

[0065] Spatial interpolation methods are used for calculation: the values ​​of blank grids are extrapolated based on existing data points using mathematical models. Common methods include Kriging, inverse distance weighted interpolation (IDW), or spline function interpolation.

[0066] The three methods described above can be flexibly combined according to the actual data characteristics and application scenarios. For example, spatial interpolation can be used to generate preliminary filled values ​​first, and then outliers can be corrected using neighboring grid data. Simultaneously, the filling method and parameters used (such as interpolation model type, weighting coefficients, and proximity range) must be clearly recorded in the data storage or metadata to ensure traceability and transparency during subsequent data retrieval. Furthermore, the selection of the filling method can be extended to dynamic adjustment strategies (such as automatically switching the filling method based on the data missing rate) or hybrid models (such as combining administrative boundaries with spatial interpolation results), all of which fall within the scope of this solution.

[0067] According to one embodiment of this application, in step S5, before data synthesis, a step of standardizing the basic unit process data is further included to ensure that all data are synthesized under a unified dimension and system boundary.

[0068] As described above, the specific content of the standardization processing of basic unit process data includes the following two aspects: Unit unification: This refers to converting data from different sources or formats into a unified unit of measurement and expression to eliminate calculation errors caused by unit differences. For example, if a company's electricity consumption data is expressed in "kWh" while another company's energy consumption data is expressed in "GJ", then unit conversion (1kWh=3.6MJ) is needed to unify the two into the same unit (e.g., converting them all to "GJ").

[0069] System boundary consistency adjustment: This refers to ensuring that the technological scope described by all basic unit process data belonging to the same grid remains consistent before data synthesis. Each basic unit process has its own system boundary, which clearly defines the activities that start and end the process, which core production units, auxiliary facilities, and internal transportation are included, and which activities are excluded. If the boundaries are inconsistent, direct aggregation will lead to distorted results.

[0070] For example, for the unit process of "steelmaking," Company A's data definition begins with "molten iron entering the furnace" and ends with "molten steel exiting the furnace," with inputs including molten iron, alloys, and electrical energy, and outputs of molten steel and slag. Company B's data definition, however, begins with "scrap steel entering the furnace" and includes the energy consumption and emissions of the plant's "waste gas treatment" unit. Both describe processes with the same name but different scopes. Before synthesis, adjustments must be made: either supplementing Company A's data with typical data for the "waste gas treatment" unit, or removing the "waste gas treatment" portion from Company B's data, so that all "steelmaking" process data to be aggregated have unified start and end activity boundaries. This adjustment is usually based on a standardized reference unit process definition, which clearly specifies, in text or flowchart form, a list of activities that a standard unit process should include and exclude.

[0071] According to one embodiment of this application, the descriptive information includes at least one of the following: name, location, technical description, data processing method, data source, time representativeness description, purpose of spatial division, division rules, scope of application, suggestions for applicable scenarios, and the synthesis relationship between basic unit process data and grid unit process data.

[0072] As described above, the specific content and examples of the descriptive information are as follows: Name: The title used to uniquely identify the dataset, which should reflect the data subject, time range, and spatial scale.

[0073] Location: Clearly define the geographical scope and specific boundaries of the data coverage.

[0074] Technical Description: Explain the core technologies and algorithms used in data processing.

[0075] Data Sources: Specify the channels and providers of the original data. For example: "The basic data comes from the Ministry of Ecology and Environment's enterprise pollution discharge permit database, the National Bureau of Statistics' annual report on industrial energy consumption, real-time monitoring data from environmental monitoring stations in the Yangtze River Delta region, and emission coefficients from publicly available literature."

[0076] Time representativeness description: Define the time range, update frequency, and representativeness of the data. For example: "The data covers the entire production cycle of 2022, is updated quarterly, and is suitable for reflecting the annual average emission level; because it does not include extreme weather events (such as abnormal shutdowns caused by typhoons), it is not suitable for short-term fluctuation analysis."

[0077] Spatial partitioning purpose: Explain the design goals of grid partitioning. For example: "To support the allocation of quotas in the regional carbon trading market, industrial emission data needs to be refined to the district and county-level administrative units; to assess pollutant diffusion paths, a 500m×500m regular grid is used to match the spatial resolution of the atmospheric diffusion model."

[0078] Division Rules: A detailed explanation of the methods and basis for grid division. For example: "Administrative unit grids are based on the provincial administrative boundaries in 2020. Regular grids are divided using 30°×30° latitude and longitude. Custom grids are dynamically adjusted according to the distribution of industrial parks, with priority given to increasing grid density in areas with dense pollution sources."

[0079] Scope of application: Limits the conditions and restrictions for using the data. For example: "This dataset is suitable for regional environmental impact assessment, policy planning and academic research, but not for precise accounting of individual enterprises; because it does not include groundwater resource consumption data, it is not suitable for hydrogeological analysis."

[0080] Recommended application scenarios and precautions: For example, "When used for urban air quality simulation, it is recommended to combine meteorological data to correct diffusion parameters; when used for carbon emission accounting, it is recommended to correct for the impact of differences in the proportion of energy types on the synthesis results."

[0081] The synthesis relationship between basic unit process data and grid unit process data: clearly define the data aggregation logic and calculation rules. For example: "Use a spatial matching algorithm to map enterprise-level emission data to corresponding grid units. Blank grids are filled by inverse distance weighting interpolation between adjacent grids. The final synthesis result is calculated by weighted average of grid area to determine the emission intensity per unit area."

[0082] The aforementioned descriptive information must be stored along with the dataset in a structured format (such as JSON, XML, or tables), and its retrievability must be ensured through a metadata management system. For example, when this dataset is accessed in an LCA analysis tool, the system can automatically read the "Applicable Scenario Recommendation" field and prompt the user to supplement transportation phase data if a lifecycle assessment is required. The completeness of this descriptive information directly determines the reusability and compliance of the dataset and is a key aspect of the protection provided by this solution.

[0083] According to one embodiment of this application, in step S7, the specified data format is ILCD format, EcoSpold2 format, JSON-LD format, or a database table structure.

[0084] As mentioned above, the specified data formats specifically include the following four types, and the selection of each format needs to be determined based on the data application scenario, compatibility requirements, and storage method: ILCD Format: ILCD stands for International Reference Life Cycle Data System. It is a complete LCA methodology and data format system developed and maintained by the Institute for Environment and Sustainable Development, a division of the Joint Research Centre (JRC) of the European Commission. At its core is an XML-based data format specification for storing and exchanging Life Cycle Inventory (LCI) and Impact Assessment (LCIA) data. Its goal is to create a globally applicable, neutral, and high-quality data format to support policies and applications such as Environmental Footprint (EF) and Environmental Product Declaration (EPD).

[0085] EcoSpold2 Format: EcoSpold2 is a standardized electronic format for storing and exchanging Lifecycle Inventory (LCI) data in the Lifecycle Assessment (LCA) field. It is an upgrade of its predecessor, EcoSpold1, and was developed under the leadership of the Swiss-based Ecoinvent database. It has become the de facto standard in the global LCA community, particularly among communities associated with Ecoinvent. Simply put, EcoSpold2 is a "universal language and container for LCA data," ensuring that data from diverse sources can be consistently understood, processed, and used.

[0086] JSON-LD format: A lightweight semantic data format based on JSON, which defines the semantic association of data attributes through the "@context" field (such as using schema.org or a custom vocabulary), and is suitable for data sharing and machine readability needs in the Web environment.

[0087] Database table structure: refers to the structured storage method defined by relational databases (such as MySQL and PostgreSQL) or NoSQL databases (such as MongoDB). Its advantages lie in supporting efficient querying, transaction management, and large-scale data processing.

[0088] The choice of the four formats mentioned above should be based on actual needs: ILCD and EcoSpold2 are suitable for standardized data exchange in the LCA field, JSON-LD is suitable for open data and semantic interoperability on the web, and database table structures are suitable for scenarios requiring high-performance storage and complex queries. Furthermore, conversion tools between formats (such as importing EcoSpold2 data into a database and exporting ILCD data to JSON-LD) are also within the scope of this solution to ensure data interoperability across different systems.

[0089] For any parts not mentioned in this application, existing technologies may be used or referenced.

[0090] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0091] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method of establishing a spatial LCA dataset, characterized by, The method comprises the following steps: S1, determining the target of the built data set, determining the contained product or service according to the target, and determining the data requirements corresponding to the target; the data requirements include system boundary, regional boundary, time boundary, technical representation requirement and data quality requirement; S2, collecting basic unit process data corresponding to the product or service based on the data requirements; the basic unit process data at least includes description information, input flow, output flow, and geographic location information of the basic unit process; S3, determining the spatial division rule according to the regional boundary and the expected application in the data requirements, the spatial division rule including spatial range, grid shape and spatial granularity; dividing the regional boundary based on the spatial division rule, generating a grid system covering the spatial range, and the expected application being a specific scenario in which the data set plan is used; S4, spatially matching the geographic location information of each basic unit process data with the grid system, so that each basic unit process data is attributed to at least one grid; S5, for each grid, aggregating and calculating all basic unit process data attributed to it according to a predetermined data synthesis rule to generate the synthesis unit process data of the grid; the data synthesis rule includes a weighted average algorithm for input and output flow; S6, associating each synthesis unit process data with the geographic location identifier of the grid to which it belongs to form grid unit process data; S7, storing the basic unit process data and / or the grid unit process data together with its description information and data processing record in a specified data format.

2. The method of claim 1, wherein, In step S2, the data granularity of the basic unit process data includes process level data, enterprise level data or regional level data, and the regional level data includes province, city and district level data; wherein the type of geographic location information includes point data, line data or area data; the expression of geographic location information includes geographic coordinates based on standard coordinate system, geographic code, or self-defined geographic coordinate system.

3. The method of claim 1, wherein, In step S3, the spatial granularity includes administrative unit granularity, regular grid granularity or self-defined grid granularity.

4. The method of claim 1 or 3, wherein, In step S3, the grid shape is regular grid or irregular grid.

5. The method of claim 1, wherein, In step S3, the grid system is digitally expressed and stored in vector format or raster format.

6. The method of claim 1, wherein, In step S4, the spatial matching includes spatial calculation; when the geographic location of a basic unit process data is located at the grid boundary, the final attribution grid is determined according to the preset grid attribution rule.

7. The method of claim 1, wherein, In step S5, the predetermined data synthesis rule further includes blank grid filling rule; for the blank grid without basic unit process data, at least one of the following methods for establishing spatial LCA data set is used to generate filling data: using the average value of the upper level regional data to fill, using the data of the spatial adjacent grid to fill, or using the spatial interpolation method to calculate; and the filling method used is recorded.

8. The method of claim 1, wherein, In step S5, before data synthesis, a step of standardizing the basic unit process data is further included to ensure that all data is synthesized under unified dimensions and system boundaries.

9. The method of claim 1, wherein, The description information includes at least one of the following: name, location, technical description, data processing method, data source, time representation description, spatial division purpose, division rule, applicable range, applicable scene suggestion, and synthesis relationship between the basic unit process data and the grid unit process data.

10. The method of claim 1, wherein, In step S7, the specified data format is ILCD format, EcoSpold2 format, JSON-LD format or database table structure.