Data screening and evaluation method and system for supporting development of photovoltaic product carbon footprint background data set based on environmental assessment file
By building a multi-level structured labeling system and a nonlinear scoring mechanism, quality scoring and data review are carried out on environmental impact assessment documents, which solves the efficiency and accuracy issues of data screening and evaluation in the development of photovoltaic product carbon footprint background data sets, achieves efficient and precise data quality control, improves data consistency and comparability, and provides high-quality data support for carbon accounting in the photovoltaic industry.
Patent Information
- Application Number
- CN202510810538.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the development of existing photovoltaic product carbon footprint background data sets, the efficiency of environmental impact assessment document screening is low, data quality assessment is inaccurate, data processing standardization is poor, and there is a lack of multivariate cross-validation mechanism, resulting in insufficient refinement of data quality control and difficulty in meeting the industry's carbon accounting needs.
Construct a multi-level structured labeling system to mark and store environmental impact assessment documents and data, establish a quality evaluation index system based on four first-level indicators: document reliability, representativeness, information comprehensiveness, and data readability, use a nonlinear confidence enhancement scoring function and a weighted geometric mean method, combined with a harmonic enhancement normalization function and an asymmetric confidence enhancement deviation function to review and verify data, and screen high-quality data through a key indicator elimination mechanism.
It achieves standardized management and efficient retrieval of environmental impact assessment documents, improves data query and extraction efficiency, ensures the accuracy and consistency of data quality, provides a high-quality carbon footprint background data set, and provides reliable data support for carbon footprint accounting in the photovoltaic industry.
Smart Images

Figure CN120706967A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of environmental assessment big data processing and carbon footprint calculation technology, and in particular to environmental assessment data screening and quality evaluation algorithm technology in the construction of photovoltaic product carbon footprint background data sets. Background Art
[0002] As global climate change becomes increasingly prominent, developing a green, low-carbon economy has become a widely held international consensus. Photovoltaic products also generate a carbon footprint over their lifecycle, and accurately calculating their carbon emissions is crucial for developing emission reduction strategies and improving the industry's environmental performance.
[0003] Currently, carbon footprint accounting for photovoltaic products primarily relies on establishing a comprehensive and reliable carbon footprint background dataset. Environmental Impact Assessment (EIA) documents, as a key data source, contain a wealth of data on raw materials, energy, and pollution emissions, offering unique advantages in supporting the development of carbon footprint background datasets. However, the sheer volume and diverse formats of EIA documents, coupled with varying data quality, present numerous challenges in data screening and evaluation.
[0004] Traditional EIA document screening methods rely primarily on manual judgment, which is inefficient and highly subjective. For data quality assessment, existing methods often employ simple linear weighted scoring models, which fail to fully reflect the multidimensional characteristics of the data and significantly impact the accuracy and reliability of the assessment results. Furthermore, the lack of unified data classification standards and standardized processing procedures during the extraction and processing of EIA document data leads to large data conversion errors and poor comparability, significantly impacting the accuracy of carbon footprint accounting.
[0005] Traditional methods for data auditing and verification primarily rely on single-threshold screening, making it difficult to effectively identify abnormal data and lacking multivariate cross-validation mechanisms, resulting in insufficiently refined data quality control. Furthermore, existing data quality assessment systems fail to consider the specific characteristics of the photovoltaic industry, making them incapable of meeting the actual needs of carbon accounting within the industry.
[0006] In summary, the development of existing photovoltaic product carbon footprint background datasets faces many problems, such as low efficiency in environmental impact assessment document screening, inaccurate data quality assessment, poor data processing standardization, and lack of quality control measures. A systematic and refined data screening and evaluation method is urgently needed to improve the screening efficiency and quality of data sources, improve data consistency and comparability, and provide high-quality data support for carbon footprint accounting in the photovoltaic industry. Summary of the Invention
[0007] The purpose of this application is to provide a data screening and evaluation method and system based on environmental impact assessment documents to support the development of a carbon footprint background dataset for photovoltaic products, so as to solve the problems raised in the above background technology.
[0008] This application discloses a data screening and evaluation method for developing a photovoltaic product carbon footprint background dataset based on environmental impact assessment documents, comprising:
[0009] Step A: Construct a multi-level structured label system for the environmental impact assessment documents and the data therein, and perform annotation and storage;
[0010] Step B: Establish a quality evaluation index system based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability, and calculate the original score of each second-level indicator. ij The following nonlinear confidence enhancement scoring function is used for processing:
[0011]
[0012] Where α>1 and β∈(0, 0.1) are adjustable parameters to obtain enhanced scores And evaluate the quality of environmental impact assessment documents based on the quality evaluation index system and classify the environmental impact assessment documents into different quality levels;
[0013] Step C: Based on the quality evaluation results of Step B, extract data from the environmental impact assessment documents with a high-quality or good rating, classify the extracted data into raw materials, energy, auxiliary materials, and pollution categories, and perform unit standardization processing;
[0014] Step D: Review and verify the data extracted and standardized in Step C, and calculate the data quality score;
[0015] Step E: Filter the data based on the data quality score obtained in step D to obtain a data pool for establishing a photovoltaic product carbon footprint background dataset.
[0016] In a preferred embodiment, the environmental impact assessment document labeling system in step A is a five-level structured labeling system, which includes a product label, a time label, a geographic label, a property label, and a quality label in sequence; the data labeling system is a four-level structured labeling system, which includes a document label, a category label, a name label, and a quality label in sequence;
[0017] After step B, the enhanced scores are normalized and weighted. Summarized as the first-level indicator enhancement score S i , and use the weighted geometric mean method to obtain the comprehensive quality score Q of the environmental impact assessment document, and then divide Q into four levels: high quality, good, poor and invalid;
[0018] In step D, the review and verification specifically include: converting the data into relative units using a harmonically enhanced normalization function; calculating the deviation between the data and the industry benchmark using an asymmetric confidence-enhanced deviation function; performing balance verification by combining element balance, mass balance, and water balance; constructing a quality scoring indicator system based on three dimensions: data accuracy, data consistency, and data reliability, and obtaining a data quality score using a comprehensive scoring method that integrates a multi-level weighting mechanism, a nonlinear scoring enhancement function, and a key indicator elimination mechanism;
[0019] In step E, data with a quality score of 100 are directly included in the data pool, data with a quality score of 50 to 100 are used as reference data after correction, and data with a quality score of 0 to 50 are eliminated.
[0020] In a preferred embodiment, the unit normalization process in step C adopts the following formula:
[0021] X std =X raw × C unit (5)
[0022] Among them, X std is the data value after conversion to standard unit, X raw is the original extracted data, C unit is the unit conversion factor.
[0023] In a preferred embodiment, the formula for converting the data to relative units by harmonizing the enhanced normalization function in step D is:
[0024]
[0025] Among them, X rel is the normalized relative data value, X std The original data after unit standardization, Q prod is the product output corresponding to the data, ∈ is the minimum value (used to avoid division by zero), and δ is the normalized harmonic parameter (used to enhance the robustness of small-output data). By introducing logarithmic terms and harmonic parameters, this function can improve the stability of low-output samples and enhance the comparability of data.
[0026] In a preferred embodiment, the formula for calculating the deviation between the data and the industry benchmark using the asymmetric confidence enhancement deviation function in step D is:
[0027] RBD = (1-exp(-α×RE))×100 (8)
[0028] Among them, RBD is the enhanced industry deviation score, ranging from 0 to 100, and the lower the score, the greater the deviation; RE is the original relative error, calculated as
[0029] X bench is the industry benchmark value of the data item; α is the adjustment coefficient (used to control the strictness of the deviation constraint); this function introduces nonlinear exponential decay to impose stricter penalties on high-deviation data, thereby improving the sensitivity of deviation assessment.
[0030] In a preferred embodiment, the comprehensive scoring method using the zero-score elimination mechanism for key indicators in both step B and 400 is:
[0031] Q=min(Q, Q cut-off ) (4)
[0032] Among them, Q is the comprehensive quality score, Q cut-off is the elimination threshold; in step B, when the original score of any key secondary indicator among the key data missing rate, data format structured degree, data and information readability is 0, the elimination mechanism is triggered, and the elimination threshold Q cut-off Set to 29.9; in step D, when the original score of any key secondary indicator in the deviation from the benchmark value and the cross-check consistency is 0, the elimination mechanism is triggered and the elimination threshold is set to 49.9.
[0033] In a preferred embodiment,
[0034] When the comprehensive quality score Q of the environmental assessment document is between 70 and 100, it is defined as a high-quality grade;
[0035] When Q is between 50 and 70, it is defined as a good grade;
[0036] When Q is between 30 and 50, it is defined as a poor grade;
[0037] When Q is between 0 and 30, it is defined as an invalid level.
[0038] In a preferred embodiment, the weighted geometric mean method in step B is used to calculate the comprehensive score Q of the environmental impact assessment document using the following formula:
[0039]
[0040] Among them, S i is the enhanced score of the i-th first-level indicator, W i is the global weight of the i-th first-level indicator and satisfies ∑W i =1, n is the number of primary indicators and n=4; this geometric mean method ensures that if the score of any indicator is too low, the overall score will be significantly lowered, reflecting the synergistic importance of each indicator.
[0041] In a preferred example, the balance check in step D includes calculations of the element balance deviation rate, the mass balance deviation rate, and the water balance deviation rate. The deviation rate thresholds are all set at 2%. Exceeding the thresholds indicates that there may be problems with the data. These are used to verify the input and output balance of key elements (such as silicon, chlorine, and fluorine), conservation of material mass, and input and output balance of water resources in the production process of photovoltaic products.
[0042] In a preferred example, the category labels in step A include raw materials, energy, auxiliary materials and pollution, wherein the raw materials category identifies materials with a silicon purity ≥99.9999% (wt), the energy category identifies energy in electricity measurement units (kWh) or steam pressure parameters, the auxiliary materials category identifies materials with a solder strip wire diameter ≤0.3mm or a backplane transmittance ≥93% or a cover glass thickness of 2mm, and the pollution category identifies pollutants with an emission volume of t / a.
[0043] In a preferred example, the method is implemented through a digital system, which includes a data source acquisition and storage module, a data source quality evaluation module, a data extraction and screening module, a data quality scoring module, and a data management and export module.
[0044] This application also discloses a data screening and evaluation system for developing a photovoltaic product carbon footprint background dataset based on environmental impact assessment documents. The system comprises:
[0045] Data labeling and storage module, used to construct a multi-level structured labeling system for environmental impact assessment documents and the data therein, and to label and store them;
[0046] The quality evaluation module is used to establish a quality evaluation index system based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability, and to evaluate the original scores of each second-level indicator. ij The following nonlinear confidence enhancement scoring function is used for processing:
[0047] Where α>1 and β∈(0, 0.1) are adjustable parameters to obtain enhanced scores And evaluate the quality of environmental impact assessment documents based on the quality evaluation index system and classify the environmental impact assessment documents into different quality levels;
[0048] The data extraction and classification module is used to extract data from environmental impact assessment documents with high-quality or good grades based on the quality evaluation results of the quality evaluation module, classify the extracted data into raw materials, energy, auxiliary materials and pollution categories, and perform unit standardization processing;
[0049] The data review and verification module is used to review and verify the data extracted and standardized in the data extraction and classification module, and calculate the data quality score;
[0050] The data screening module is used to screen the data based on the data quality score obtained by the data review and verification module to obtain a data pool for establishing a photovoltaic product carbon footprint background data set.
[0051] The data screening and evaluation method and system proposed in this application, which is based on environmental impact assessment documents to support the development of a carbon footprint background dataset for photovoltaic products, has the following significant technical effects:
[0052] This application has built a structured tagging system to achieve standardized management and efficient retrieval of environmental impact assessment documents. It categorizes and stores environmental impact assessment documents through a five-level structured tagging system (including product tags, time tags, geographic tags, property tags, and quality tags), while also annotating data with a four-level structured tagging system (including file tags, category tags, name tags, and quality tags). This significantly improves the efficiency of querying and extracting environmental impact assessment documents and their data.
[0053] Based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability (I i ) constructs a multi-dimensional quality evaluation index system and scoring matrix, which provides a systematic standard for the quality assessment of environmental impact assessment document data sources. The comprehensive scoring method for data source quality proposed in this application integrates a multi-level weighting mechanism, nonlinear scoring enhancement and a key indicator elimination mechanism. It processes the original score through a nonlinear confidence enhancement scoring function (Formula 1), effectively avoiding the problem of inflated scores for inferior data sources. This function can more accurately reflect the actual status of each quality dimension by adjusting the parameters α (score nonlinear amplification factor) and β (score confidence release factor). It calculates the comprehensive quality score Q in combination with the weighted geometric mean method (Formula 3), and ensures the screening of high-quality data sources through the key indicator zero-score elimination mechanism (Formula 4).
[0054] For the verification of extracted data, this application proposes a multivariate cross-validation method that combines industry benchmark verification with balance verification. By harmonizing the enhanced normalization function (Formula 6), the relative unit conversion of data is achieved, enhancing the comparability of production data of different scales. The asymmetric confidence-enhanced deviation function (Formula 8) is used to calculate deviations from industry benchmark values, effectively screening data that deviates from industry standards. At the same time, combining element balance, mass balance, and water balance verification mechanisms, this method overcomes the flaws of traditional single-threshold screening methods that are prone to misjudgment, improving the accuracy and consistency of data extraction and evaluation.
[0055] Ultimately, this application achieved integrated management of EIA documents and their data collection and storage, data source quality assessment, data extraction, data review and verification, and data quality scoring and screening through a digital integrated system. Based on OCR recognition and natural language processing technologies, the system automatically parses and extracts various types of data from unstructured EIA documents, significantly improving the efficiency and reliability of the development of carbon footprint background datasets for photovoltaic products, providing high-quality support for carbon footprint accounting and compliance applications.
[0056] Through the combined effect of the above-mentioned technical effects, this application solves the current problems in the development of photovoltaic product carbon footprint background data sets, such as scattered data acquisition, inconsistent formats, uneven quality, and lack of systematic screening and verification methods, providing a strong data foundation support for promoting the low-carbon development of the photovoltaic industry.
[0057] The specification of this application records a large number of technical features, which are distributed in various technical solutions. If all possible combinations of technical features of this application (i.e., technical solutions) are to be listed, the specification will be too lengthy. In order to avoid this problem, the various technical features disclosed in the above-mentioned invention content of this application, the various technical features disclosed in the various embodiments and examples below, and the various technical features disclosed in the accompanying drawings can be freely combined with each other to form various new technical solutions (these technical solutions are all deemed to have been recorded in this specification), unless such a combination of technical features is technically infeasible. For example, in one example, feature A+B+C is disclosed, and in another example, feature A+B+D+E is disclosed. Features C and D are equivalent technical means that play the same role. Technically, only one of them can be used, and it is impossible to use them at the same time. Feature E can be technically combined with feature C. Then, the solution of A+B+C+D should not be considered as having been recorded because it is technically infeasible, while the solution of A+B+C+E should be considered as having been recorded. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flow chart of a data screening and evaluation method for developing a photovoltaic product carbon footprint background data set based on environmental impact assessment documents according to the first embodiment of the present application.
[0059] Figure 2 It is a structural diagram of a data screening and evaluation system developed based on environmental impact assessment documents to support the carbon footprint background data set of photovoltaic products according to the second embodiment of the present application.
[0060] Figure 3 This is a flowchart of a specific example of a data screening and evaluation method for developing a photovoltaic product carbon footprint background data set based on environmental impact assessment documents according to the first embodiment of the present application. DETAILED DESCRIPTION
[0061] In the following description, many technical details are provided to help readers better understand this application. However, those skilled in the art will understand that even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can be implemented.
[0062] Description of some concepts:
[0063] A multi-level structured labeling system refers to a system architecture that standardizes the labeling of data according to preset hierarchies and classification rules. In this application, the environmental impact assessment documents use a five-level labeling system (product label, time label, geographic label, property label, and quality label), and the data uses a four-level labeling system (document label, category label, name label, and quality label), achieving hierarchical and classified management of documents and data.
[0064] Nonlinear confidence enhancement scoring function: refers to a function that processes the original score through nonlinear mathematical transformation.
[0065] Harmonic enhancement normalization function: refers to the mathematical function used for relative unit conversion of data. By introducing logarithmic terms and harmonic parameters δ, it improves the stability of low-yield samples and enhances the comparability of production data of different scales.
[0066] Asymmetric confidence-enhanced deviation function: refers to a function used to calculate the deviation of data from the industry benchmark value. It imposes stricter penalties on high-deviation data through an exponential decay mechanism, thereby enhancing the sensitivity of deviation assessment.
[0067] Key indicator elimination mechanism: refers to the quality control mechanism that is forcibly activated when the original score of the preset key secondary indicator is 0, which limits the comprehensive score to below the elimination threshold to ensure the effective exclusion of low-quality data sources.
[0068] Weighted geometric mean method: refers to the method of using the geometric mean and combining the weight coefficient to calculate the comprehensive score. This method has a significant penalty effect on data sources with extremely low scores on any indicator.
[0069] Balance verification: refers to a verification method that verifies data consistency through element balance, mass balance and water balance, respectively verifying the input and output balance of key elements (such as silicon, chlorine, and fluorine) in the production process, conservation of material mass, and input and output balance of water resources. The deviation rate threshold is set at 2%.
[0070] Data quality score: refers to the comprehensive score calculated based on the quality scoring system established based on the three dimensions of data accuracy, data consistency, and data reliability, which is used to quantitatively evaluate the availability and credibility of extracted data.
[0071] The following is a summary of some of the innovative features of this application:
[0072] In general, in the complex technical scenario of photovoltaic product carbon footprint background data set development, this application has achieved deep coupling optimization of environmental impact assessment document data source quality assessment and data extraction verification by constructing a synergistic enhancement mechanism for multi-level heterogeneous data quality evaluation and screening. Specifically, this application creatively parameterizes and coordinates the nonlinear amplification factor α and the confidence release factor β in the nonlinear confidence enhancement scoring function (Formula (1)) so that the original score s ij After the composite operation of power function transformation and exponential decay function, the enhanced score S is formed. ij This scoring method can suppress the phenomenon of inflated scores of low-quality data sources while amplifying the distinction between high-quality and low-quality data sources through a nonlinear mapping mechanism, thereby forming a cascade amplification effect with the multi-level weight distribution mechanism in the weighted geometric mean method (Formula (3)), ensuring that the comprehensive quality score Q can keenly capture the coordinated changes of various quality dimensions.
[0073] More importantly, this application solves the nonlinear distortion problem of production data of different scales in the relative unit conversion process by introducing the logarithmic term enhancement mechanism and the output harmonic parameter δ in the harmonic enhancement normalization function (Formula (6)). This function and the exponential decay penalty mechanism in the asymmetric confidence enhancement deviation function (Formula (8)) form a dual nonlinear correction system, so that the data deviation assessment not only has the characteristic of enhanced sensitivity to high-deviation data, but also can achieve adaptive matching of the strictness of the deviation constraint through dynamic regulation of the adjustment coefficient α.
[0074] In addition, the key indicator zero score elimination mechanism (Formula (4)) is implemented by the minimum function min(Q, Q cut-off )’s rigid constraint characteristics, together with the aforementioned multiple nonlinear scoring enhancement mechanism, form a rigid and flexible quality control architecture. Differentiated elimination thresholds (29.9 and 49.9) are set in steps 200 and 400 respectively, realizing a progressive and strict screening of the environmental impact assessment document quality evaluation and data quality score, ensuring high confidence and high consistency of the data in the final data pool.
[0075] This technical concept of deep integration of multi-dimensional nonlinear scoring enhancement and multi-stage quality control, through the heterogeneous labeling mechanism of five-level and four-level structured labeling systems, combined with the triple balance verification constraints of element balance, mass balance and water balance, forms a full-chain quality assurance system from the environmental impact assessment document data source to the carbon footprint background dataset. It fundamentally solves the technical bottlenecks of insufficient sensitivity and lack of discrimination of traditional linear scoring methods in complex data quality assessment, and provides a systematic technical breakthrough path for improving the data credibility of carbon footprint accounting in the photovoltaic industry.
[0076] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0077] The first embodiment of the present application relates to a data screening and evaluation method for developing a photovoltaic product carbon footprint background data set based on environmental impact assessment documents. The process is as follows: Figure 1 As shown, the method includes the following steps:
[0078] Step 100: Construct a multi-level structured label system for the environmental impact assessment documents and the data therein and perform labeling and storage.
[0079] Step 200: Establish a quality evaluation index system based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability, and assign the original score of each second-level indicator s ij The following nonlinear confidence enhancement scoring function is used for processing:
[0080]
[0081] Where α>1 and β∈(0, 0.1) are adjustable parameters to obtain enhanced scores And the environmental impact assessment documents are scored for quality based on the quality evaluation index system, and the environmental impact assessment documents are divided into different quality levels.
[0082] Step 300: Based on the quality evaluation results of step 200, extract data from the environmental impact assessment documents of high quality or good grade, classify the extracted data into raw materials, energy, auxiliary materials and pollution categories, and perform unit standardization.
[0083] Step 400: Review and verify the data extracted and standardized in step 300, and calculate the data quality score.
[0084] Step 500: Filter the data according to the data quality score obtained in step 400 to obtain a data pool for establishing a photovoltaic product carbon footprint background data set.
[0085] Optionally, the environmental impact assessment document labeling system in step 100 is a five-level structured labeling system, which includes a product label, a time label, a geographic label, a property label, and a quality label in sequence; the data labeling system is a four-level structured labeling system, which includes a document label, a category label, a name label, and a quality label in sequence;
[0086] After step 200, the enhanced scores are normalized and weighted. Summarized as the first-level indicator enhancement score S i , and use the weighted geometric mean method to obtain the comprehensive quality score Q of the environmental impact assessment document, and then divide Q into four levels: high quality, good, poor and invalid;
[0087] In step 400, the review and verification specifically include: converting the data into relative units using a harmonically enhanced normalization function; calculating the deviation between the data and the industry benchmark using an asymmetric confidence-enhanced deviation function; performing balance verification by combining element balance, mass balance, and water balance; constructing a quality scoring indicator system based on three dimensions: data accuracy, data consistency, and data reliability, and obtaining a data quality score using a comprehensive scoring method that integrates a multi-level weighting mechanism, a nonlinear scoring enhancement function, and a key indicator elimination mechanism.
[0088] In step 500, data with a quality score of 100 is directly included in the data pool for use, data with a quality score of 50 to 100 is used as reference data after correction, and data with a quality score of 0 to 50 is eliminated.
[0089] Optionally, the category labels in step 100 include raw materials, energy, auxiliary materials and pollution, where the raw materials category identifies materials with silicon purity ≥99.9999% (wt), the energy category identifies energy in electricity measurement units (kWh) or steam pressure parameters, the auxiliary materials category identifies materials with solder ribbon wire diameter ≤0.3mm or backplane transmittance ≥93% or cover glass thickness 2mm, and the pollution category identifies pollutants with emissions of t / a.
[0090] Optionally, the weighted geometric mean method in step 200 calculates the comprehensive score Q of the environmental impact assessment document using the following formula:
[0091]
[0092] Among them, S i is the enhanced score of the i-th first-level indicator, W i is the global weight of the i-th first-level indicator and satisfies ∑W i =1, n is the number of primary indicators and n=4; this geometric mean method ensures that if the score of any indicator is too low, the overall score will be significantly lowered, reflecting the synergistic importance of each indicator.
[0093] Optionally, the unit normalization process in step 300 uses the following formula:
[0094] X std =X raw × C unit (5)
[0095] Among them, X std is the data value after conversion to standard unit, X raw is the original extracted data, C unit is the unit conversion factor.
[0096] Optionally, the formula for performing relative unit conversion on the data by harmonizing the enhanced normalization function in step 400 is:
[0097]
[0098] Among them, X rel is the normalized relative data value, X std The original data after unit standardization, Q prod is the product output corresponding to the data, ∈ is the minimum value (used to avoid division by zero), and δ is the normalized harmonic parameter (used to enhance the robustness of small-output data). By introducing logarithmic terms and harmonic parameters, this function can improve the stability of low-output samples and enhance the comparability of data.
[0099] Optionally, in step 400, the formula for calculating the deviation between the data and the industry benchmark using the asymmetric confidence enhancement deviation function is:
[0100] RBD = (1-exp(-α×RE) ) ×100 (8)
[0101] Among them, RBD is the enhanced industry deviation score, ranging from 0 to 100, and the lower the score, the greater the deviation; RE is the original relative error, calculated as
[0102] X bench is the industry benchmark value of the data item; α is the adjustment coefficient (used to control the strictness of the deviation constraint); this function introduces nonlinear exponential decay to impose stricter penalties on high-deviation data, thereby improving the sensitivity of deviation assessment.
[0103] Optionally, the comprehensive scoring method using the zero-score elimination mechanism for key indicators in both steps 200 and 400 is:
[0104] Q=min(Q, Q cut-off ) (4)
[0105] Among them, Q is the comprehensive quality score, Q cut-off is the elimination threshold; in step 200, when the original score of any key secondary indicator among the key data missing rate, data format structured degree, data and information readability is 0, the elimination mechanism is triggered, and the elimination threshold Q cut-off Set to 29.9; in step 400, when the original score of any key secondary indicator in the deviation from the benchmark value and the cross-check consistency is 0, the elimination mechanism is triggered and the elimination threshold is set to 49.9.
[0106] Optionally, when the comprehensive quality score Q of the environmental assessment document is between 70 and 100, it is defined as a high-quality grade;
[0107] When Q is between 50 and 70, it is defined as a good grade;
[0108] When Q is between 30 and 50, it is defined as a poor grade;
[0109] When Q is between 0 and 30, it is defined as an invalid level.
[0110] Optionally, the balance check in step 400 includes calculations of element balance deviation rate, mass balance deviation rate, and water balance deviation rate, with the deviation rate thresholds all set at 2%. Exceeding the thresholds indicates that there may be problems with the data; these are used to verify the input and output balance of key elements (such as silicon, chlorine, and fluorine), conservation of material mass, and input and output balance of water resources in the production process of photovoltaic products.
[0111] Optionally, the method is implemented through a digital system, which includes a data source acquisition and storage module, a data source quality evaluation module, a data extraction and screening module, a data quality scoring module, and a data management and export module.
[0112] In order to better understand the technical solution of the present application, a specific example is provided below for illustration. The details listed in the example are mainly for ease of understanding and are not intended to limit the scope of protection of the present application.
[0113] In this example, a structured labeling system is constructed to achieve standardized management and efficient retrieval of environmental impact assessment documents. Based on multidimensional quality evaluation indicators and a scoring matrix, a comprehensive data source quality scoring method is proposed that integrates multi-level weighting, nonlinear scoring enhancement, and a key indicator elimination mechanism to accurately screen high-quality data sources. A multivariate cross-validation method combining industry benchmark value verification and balance verification is proposed to improve the accuracy and consistency of data extraction and evaluation, overcoming the defects of traditional single-threshold screening methods that are prone to misjudgment. Finally, through digital integration, integrated management of data extraction, screening, and verification is achieved, improving the development efficiency and credibility of the photovoltaic product carbon footprint background dataset, and providing high-quality support for carbon footprint accounting and compliance applications.
[0114] This example focuses on four key aspects: data source acquisition, quality screening, data verification, and automated management. It aims to improve the efficiency and quality of developing a photovoltaic product carbon footprint background dataset based on environmental impact assessment document data sources, and address many current challenges in data acquisition, screening, and verification. Specifically, it includes:
[0115] (1) Constructing a structured tagging system to achieve standardized management of the storage, retrieval, and extraction of photovoltaic product environmental impact assessment documents and data, and improving the efficiency of querying and extracting the data required for carbon footprint dataset development: Existing environmental impact assessment documents and their data are scattered, formatted in different ways, and difficult to quickly retrieve and extract, which affects the efficiency of constructing the background dataset. This example proposes a structured tagging system to classify, store, annotate, and visualize environmental impact assessment documents and their data, and establish a standardized data storage and retrieval mechanism to significantly improve data query and extraction efficiency.
[0116] (2) Construct a dual quality evaluation index system for data sources and data of enterprise environmental impact assessment reports, propose a comprehensive quality evaluation method, and improve the reliability of photovoltaic product carbon footprint background data: Existing research lacks a quality evaluation system for environmental impact assessment documents as data sources, resulting in uneven reliability of background data. This example constructs a multi-dimensional data source quality evaluation system and scoring matrix based on core indicators such as document reliability, time representativeness, information comprehensiveness, and data integrity. It also proposes a comprehensive data source quality scoring method that integrates a multi-level weighting mechanism, a nonlinear scoring enhancement function, and a key indicator elimination mechanism to quantitatively evaluate the integrity, representativeness, reliability, and readability of the data source to ensure high-quality screening of background data.
[0117] (3) Establish a standardized process for data review and verification of photovoltaic product environmental impact assessment documents to improve the quality of photovoltaic product carbon footprint inventory data: Traditional data review relies on manual experience or simple threshold methods (such as eliminating data that exceeds the industry average by ±20%). However, the photovoltaic industry chain is complex, and a single threshold is prone to mistakenly deleting valid data, affecting data integrity. This example proposes to build a multi-level cross-verification mechanism, including: ① Comparison of multiple data sources of similar data, that is, establishing a data benchmark value, determining the benchmark value based on industry standards, reports, literature or existing databases, and performing deviation testing on the extracted data; ② Verification of related data from the same data source, namely mass balance, element balance, water balance, etc., to ensure the accuracy, consistency and representativeness of the data.
[0118] (4) Construct a digital system for extracting and screening environmental impact assessment document data to improve the efficiency of photovoltaic product carbon footprint data processing: Existing methods still rely on manual parsing, screening, and verification of environmental impact assessment document data, which is labor-intensive and susceptible to human error. This example constructs a digital management system for data extraction and screening to achieve: ① Automated parsing of environmental impact assessment documents and accurate identification of key data fields; ② Standardized data processing procedures, including data structured storage, label assignment, quality scoring, calibration verification, etc.; ③ Intelligent data screening, automatically screening out low-confidence data based on the quality scoring system, improving data processing efficiency and reliability.
[0119] This example proposes a data screening and evaluation method based on environmental impact assessment documents to support the development of a photovoltaic product carbon footprint background dataset. Figure 3 , including the following specific steps:
[0120] 1. Data labeling system construction
[0121] 1.1 Data Source Labeling System Construction
[0122] Collect and record information on dimensions such as the type of environmental impact assessment document (draft for comments, draft for review, draft for approval), nature of construction (new construction, reconstruction, expansion, technological transformation), product plan (product, specification, composition, process), release time, geographic location, target data location (such as chapters or charts in the environmental impact assessment document), data units and unit conversion relationships, and use a five-level structured labeling system to assign labels to environmental impact assessment documents for file storage, so as to improve the efficiency of multi-condition file retrieval.
[0123] (1) Level 1 / Product Label: Descriptive information on the product, production process, specifications, or ingredients in the environmental impact assessment document;
[0124] (2) Secondary / time label: date of preparation of the EIA document;
[0125] (3) Level 3 / geographic tag: the province or city where the construction project is located;
[0126] (4) Level 4 / Nature Label: EIA document category (draft for comments, draft for review, draft for approval) and construction nature (new construction, reconstruction, expansion, technical transformation);
[0127] (5) Five levels / quality labels: high quality, good, average, poor, and invalid data source.
[0128] 1.2 Data labeling system construction
[0129] A four-level structured labeling system is used to label and store each available data in the environmental impact assessment document.
[0130] (1) Level 1 / File tag: the storage tag of the data source file;
[0131] (2) Secondary / category labels: raw materials, energy, auxiliary materials, pollution;
[0132] Table 1 Category label description
[0133]
[0134] (3) Level 3 / Name Tag: This is the standardized name of the data item. For example, high-purity polysilicon, solar-grade polysilicon, and purified polysilicon are all recorded and stored using the standardized name “silicon material.”
[0135] (4) Four levels / quality labels: high quality, low quality, invalid data;
[0136] 1.3 Examples of environmental impact assessment documents and data storage labels
[0137] (1) Data source storage tag example: Silicon material_Modified Siemens method_Solar grade_2023_Urumqi city_Draft for approval_New construction_Good data
[0138] (2) Data storage tag example: Data source tag (Silicon material_Modified Siemens method_Solar grade_2023_Urumqi_Draft for approval_Newly built_Good data source)_Energy_Electricity_High-quality data
[0139] 2. Data Source Quality Rating
[0140] 2.1 Comprehensive scoring method for data source quality
[0141] To efficiently screen environmental impact assessment data sources suitable for developing a photovoltaic product carbon footprint background dataset, this example proposes a comprehensive data source quality scoring method that integrates a multi-level weighting mechanism, a nonlinear score enhancement function, and a key indicator elimination mechanism. This method introduces a nonlinear confidence enhancement function to enhance the weighting of high-quality dimensions, while also introducing a zero-score elimination mechanism for key indicators. This effectively avoids the problem of inflated scores for low-quality data sources, significantly improving the efficiency and reliability of data collection in the background database. The method includes the following steps:
[0142] 2.1.1 Building a grading indicator system
[0143] This method is based on four main dimensions: document reliability, document representativeness, information comprehensiveness, and data readability (first-level indicator I i ) constructed 8 specific indicators (secondary indicator S ij ), let the global weight of the first-level indicator be W i , satisfying ∑W i =1, for each first-level indicator, the relative weight of all its subordinate second-level indicators is w ij , satisfying ∑w ij =1.
[0144] Table 2 Quality evaluation index system and weights of environmental impact assessment documents
[0145] To make the quality evaluation of environmental impact assessment documents more standardized and quantifiable, the scoring criteria for the secondary indicators were further designed into a scoring matrix, so that they can be directly referenced for scoring in practical applications. The scoring matrix uses a five-level scoring method and develops detailed scoring criteria based on the characteristics of different indicators, as shown in Table 3.
[0146] Table 3 Secondary indicator scoring matrix of data source quality evaluation index system
[0147] 2.1.2 Nonlinear Confidence Enhancement Scoring Function
[0148] In order to eliminate the misleading effect of the mean in the linear scoring mode, the original score s of each secondary indicator is ij , introduce the confidence enhancement function and calculate the enhancement score:
[0149]
[0150] s ij ——The original score of the j-th secondary indicator obtained from the scoring matrix;
[0151] α>1——Adjustable nonlinear amplification factor of the score;
[0152] β∈(0, 0.1) — Adjustable score confidence mitigation factor. Usually set to α=1.2 and β=0.02 to balance score sensitivity and stability.
[0153] 2.1.3 Calculating the Enhanced Score of the First-Level Indicator
[0154] Based on the confidence enhancement function processing results, calculate each first-level index I i Enhancement score S i , that is, S is obtained by normalizing and weighting the enhanced scores of all secondary indicators under it i :
[0155]
[0156] S i ——the enhanced score of the i-th first-level indicator;
[0157] w ij ——The weight of the j-th secondary indicator under the primary indicator to which it belongs, ∑w ij =1;
[0158] 2.1.4 Calculation of comprehensive total score
[0159] The comprehensive score Q of the environmental impact assessment document data source is calculated using the weighted geometric mean method. This method has a significant penalty effect on data sources with extremely low scores on any indicator, thereby improving the rigor and accuracy of the screening mechanism. The formula is as follows:
[0160]
[0161] Si——enhancement score of the i-th first-level indicator;
[0162] Wi——global weight of the i-th first-level indicator, satisfying ∑W i =1;
[0163] n——number of first-level indicators, i.e. n=4
[0164] 2.1.5 Elimination Mechanism for Zero Points in Key Indicators
[0165] To prevent low-quality environmental impact assessment documents from being masked by the mean, key data missing rate, data format structured degree, and data and information readability are set as key secondary indicators. If the original score of any key secondary indicator is 0, the elimination mechanism is triggered, and the comprehensive total score Q of the data source is forced to be:
[0166] Q = min(Q, Q cut-off )
Formula 4
[0167] Among them, the elimination threshold Q cut-off Set to 29.9.
[0168] 2.1.6 Data Source Selection Rules
[0169] Based on the final comprehensive score, Q, the score threshold and grading range for the selection of environmental impact assessment documents are set: high-quality data sources (70 ≤ Q ≤ 100), good data sources (50 ≤ Q < 70), poor data sources (30 ≤ Q < 50), and invalid data sources (0 ≤ Q < 30). Among them, high-quality data sources indicate that the environmental impact assessment documents can directly proceed to the next step of data extraction, verification, and scoring. Good data sources are recommended to be manually verified, while poor and invalid data sources indicate that the environmental impact assessment documents cannot be used for background data set development.
[0170] 2.2 Data Source Quality Scoring Example
[0171] The following illustrates the practical application of the comprehensive data source quality scoring method described in this example, using a specific environmental impact assessment case study. To compile a dataset for "N-type monocrystalline silicon photovoltaic cells, TOPCon, 210mm×210mm," we used the publicly available "Trina Solar (Dongtai) 10GW Annual High-Efficiency Solar Cell Project Environmental Impact Report (Draft for Review)" as an example to assess the data source quality of this environmental assessment document.
[0172] Table 4 Scoring examples and explanations (1) Calculate the nonlinear confidence enhancement score for each secondary indicator (only part is shown):
[0173]
[0174] The confidence enhancement function is used to obtain They are 86.466, 10.639, 86.466, 86.466, 61.061 and 61.061 respectively.
[0175] (2) Calculate the enhanced score of each first-level indicator:
[0176]
[0177] (3) Calculation of comprehensive total score:
[0178] Q = 100 × (0.6216) 0.2 ×(0.713) 0.1 ×(0.8647) 0.4 ×(0.6106) 0.3 =71.53
[0179] (4) Data source selection and judgment
[0180] According to the grading standards, 71.53 is between 70≤Q≤100, and none of the key secondary indicators are 0. Therefore, the data source of the environmental impact assessment document belongs to the "high-quality data source" and can be directly used for the next step of data extraction, verification and scoring.
[0181] 3. Data review and verification
[0182] Data is extracted from environmental impact assessment document data sources that are classified as high-quality and good, and each data is classified into four categories: raw materials, energy, auxiliary materials, and pollution, and corresponding data review and verification steps are carried out.
[0183] 3.1 Data Verification Method
[0184] 3.1.1 Unit Standardization
[0185] To ensure consistency across different data sources, all raw data must first be converted to standard units. The conversion rules are based on the International System of Units (SI) and incorporate common industry practices, such as converting "kilowatt-hours" to "kWh," converting "tons" to "kg," and converting natural gas consumption to Nm. 3 The unit conversion formula is as follows:
[0186] X std =X raw ×C unit [Formula 5]
[0187] X std ——Data value after conversion to standard units;
[0188] X raw — original extracted data;
[0189] C - unit conversion factor (for example, the factor for converting t to kg is 1000)
[0190] 3.1.2 Industry benchmark verification
[0191] Industry benchmark verification is primarily used to validate key parameters in photovoltaic product datasets, such as silicon and power consumption in wafer production, silicon wafer usage, power consumption, and paste usage in cell production, and cell, glass, and frame material consumption in module production. This method sets reference benchmark values based on standards, industry reports, and literature, and implements cross-validation of similar data from multiple sources by introducing ratio harmonization and confidence deviation enhancement functions. The specific steps are as follows:
[0192] (1) Numerical standardization (relative unit conversion)
[0193] Taking into account the impact of different production scales and process differences, relative units are used for data standardization, and a harmonically enhanced normalization function is introduced to improve stability for low-yield samples and enhance comparability. For example, the silicon consumption (in kg) needs to be divided by the square bar production (in kg) to obtain the silicon consumption per kg of square bar (kg / kg).
[0194]
[0195] X rel ——Normalized relative data value (such as kg silicon / kg square rod)
[0196] X std ——Raw data after unit standardization
[0197] Q prod ——Product output corresponding to this data
[0198] ∈——Minimum value, used to avoid division by zero (recommended setting ∈=10 -5 )
[0199] δ is the normalized harmonic parameter (recommended to be 1% of the mean yield), used to enhance the robustness of small yield data.
[0200] (2) Industry benchmark deviation verification
[0201] After normalization, the data will be compared with the industry benchmark for deviation. To improve calculation sensitivity and screening capabilities, an asymmetric confidence-enhanced deviation function is introduced for difference discrimination.
[0202] Original relative error rate RE (relative error):
[0203]
[0204] RE——original relative error;
[0205] X bench ——Industry benchmark value of the data item;
[0206] In order to enhance the penalty for high-deviation data, that is, to reflect stricter deviation constraints, a nonlinear enhancement function is introduced:
[0207] RBD=(1-exp(-α×RE))×100
Formula 8
[0208] RBD – Enhanced Industry Bias Score (range 0–100, with lower scores indicating greater bias)
[0209] α——Adjustment coefficient (recommended value 2 to 5, set according to industry tolerance)
[0210] 3.1.3 Balance Check
[0211] This method cross-checks related data from the same data source. This method is mainly used for raw material and auxiliary material data, and can be analyzed through methods such as element balance, mass balance, and water balance.
[0212] (1) Element balance test
[0213] This test verifies that the input and output of key elements (such as silicon, chlorine, and fluorine) in photovoltaic product production conform to the principle of element conservation and calculates the element balance deviation (EBD). The EBD threshold is set at 2%. If the threshold is exceeded, the data may be problematic.
[0214] Si element balance deviation rate during silicon wafer production
[0215]
[0216] m si,in ——The input amount of Si element in silicon material (kg), which can be converted according to the purity of silicon material
[0217] m wafer ——Mass of Si element in the output silicon wafer (kg)
[0218] m scrap ——Si element in process loss (scraps, waste rods, cutting loss, etc.) (kg)
[0219] Si element balance deviation rate during cell production
[0220]
[0221] m wafer,Si ——Mass of Si element put into silicon wafer (kg)
[0222] m cell ——Mass of Si element in the output cell (kg)
[0223] m scrap ——Mass of Si element in waste battery cells and scraps (kg)
[0224] Si element balance deviation rate in photovoltaic module production process
[0225]
[0226] m cell,Si ——Mass of Si element in the cell used for packaging (kg)
[0227] m module,Si ——Mass of Si element retained in the final component (kg)
[0228] m loss ——Mass of Si element loss caused by silicon loss or fragmentation during the packaging process (kg)
[0229] (2) Mass balance test
[0230] For key raw materials, the Mass Balance Deviation (MBD) rate is used to assess the degree of quality conservation of key raw materials in the production process, that is, the deviation between input and output quality. The MBD threshold is set at 2%. Exceeding the threshold indicates that the data may be erroneous. The yield rate parameter (Yield Rate) is introduced into the formula. The yield range threshold can be set based on industry experience. If the data deviates from this range, it indicates that further verification is required. The introduction of the yield rate parameter can more accurately reflect the ability of input materials to be converted into qualified outputs, and enhance the data screening ability to identify abnormal yield or incomplete loss data. The following explains the mass balance deviation rate formulas for the three product stages of silicon wafers, solar cells, and photovoltaic modules.
[0231] Mass balance deviation rate during silicon wafer production
[0232] The general process path is multi-crystalline / single-crystalline silicon material → pulling rods → slicing → finished silicon wafers, and the main focus is on the conversion efficiency and loss of silicon material during the slicing process.
[0233]
[0234] m si,in ——Mass of silicon material fed (kg);
[0235] m wafer ——Qualified silicon wafer mass produced (kg);
[0236] Ywafer ——Silicon wafer yield rate (e.g. 95%, use 0.95);
[0237] --Calculate the theoretically required silicon raw material quality, taking into account defective wafers and losses;
[0238] Mass balance deviation rate in the cell production process
[0239] The general process path is silicon wafer → cleaning and texturing → diffusion → printing → finished cell.
[0240]
[0241] m wafer,in ——Mass of silicon wafers fed (kg);
[0242] m cell ——Qualified battery cell mass produced (kg);
[0243] Y cell ——Battery cell yield rate;
[0244] ——Theoretical required silicon wafer quality, taking into account defective wafers and broken wafers during the process
[0245] Mass balance deviation rate in photovoltaic module production
[0246] The general process path is battery cell → string welding → lamination → finished component.
[0247]
[0248] m cell,in ——Mass of battery cells used for packaging (kg);
[0249] m module,cell ——The mass of the battery cell finally packaged into the module (kg);
[0250] Y module ——Component packaging yield (e.g. component soldering defects, broken components, etc.);
[0251] ——Theoretical required cell mass
[0252] (3) Water balance test
[0253] The Water Balance Deviation (WBD) rate measures whether water input and output during product production adhere to conservation principles, thereby identifying data anomalies or missing data. The WBD threshold is set at 2%. Exceeding the threshold indicates possible data errors. The following describes the calculation formula for the Water Balance Deviation rate for the three production stages of silicon wafers, solar cells, and photovoltaic modules.
[0254] Water balance deviation rate in silicon wafer production
[0255] The general process path is silicon material → crystal pulling → slicing → cleaning → finished silicon wafers. Water usage scenarios include cleaning (deionized water / ultrapure water), cooling water, etc. Output water scenarios include process wastewater, evaporation loss, and water contained in products or sludge.
[0256]
[0257] W in =W upw +W tap [Formula 16]
[0258] W out =W effluent +W evap +W product +W sludge [Formula 17]
[0259] W in ——Input water volume (such as tap water, ultrapure water, etc.), unit is m 3
[0260] W out ——Output water volume, including discharged water, evaporation loss, and residual water, in m 3
[0261] Water balance deviation rate in the cell production process
[0262] The general process path is texturing → pickling → diffusion → cleaning → printing. Water usage scenarios include pickling / alkaline washing water, cleaning water, etc. Output water scenarios include neutralization wastewater discharge, evaporation, and partial residue in the tank / sludge, etc.
[0263]
[0264] W process,in ——Total water consumption in production processes such as pickling and cleaning
[0265] W wastewater ——Amount of discharged wastewater
[0266] W evap ——Equipment operating evaporation capacity
[0267] W residue ——Residual water (estimated value of sediment, equipment entrainment, etc.)
[0268] Water balance deviation rate in photovoltaic module production
[0269] The general process path is component lamination → testing → cleaning. Water usage scenarios include cleaning glass, battery cells, water cooling, testing water, etc., and output water scenarios include wastewater discharge, evaporation, equipment entrainment, etc.
[0270]
[0271] W module,in ——Total amount of water added in each stage of component production
[0272] W discharge ——Wastewater discharge from component section
[0273] W evap - Natural evaporation losses during heating or exposure
[0274] W residue - Water is not completely discharged from the equipment or enters the waste
[0275] 3.2 Data Quality Scoring System
[0276] Based on industry benchmarks and balance test results, a data quality scoring index system was established from three dimensions: data accuracy, data consistency, and data reliability. This system covers three secondary indicators: deviation from the benchmark value, raw data credibility, and cross-check consistency. Each dimension was assigned the same weight to reflect the equal importance of the indicators in each dimension.
[0277] Table 3 Data quality evaluation index system
[0278]
[0279] Secondary indicators are designed using a scoring matrix to facilitate direct reference and scoring in practical applications. The matrix employs a three-level scoring approach, with specific scoring criteria established based on the characteristics of each indicator (see Table 4). Based on actual data availability, at least two primary indicators should be evaluated.
[0280] Table 4. Secondary indicator scoring matrix of data quality evaluation index system
[0281]
[0282] Finally, a comprehensive quality scoring method that combines a multi-level weighted mechanism consistent with the data source quality scoring method, a non-linear scoring enhancement function, and a key indicator elimination mechanism is used to calculate the data quality score Q'. The deviation from the benchmark value and the cross-check consistency are set as critical secondary indicators. If the original score of any critical secondary indicator is 0, the elimination mechanism is triggered, and the elimination threshold Q cut-off ′ is set to 49.9.
[0283] When Q' = 100, the data quality is high and can be directly included in the data pool for use; when 50 < Q' < 100, the data quality is average and can be used for reference, but calibration is required before use; when 0 ≤ Q' < 50, the data is considered unusable and new data needs to be collected. [[ID=⑥]] [[ID=⑦]]
[0284] [[ID=⑧]]3.3 Examples of Data Extraction, Verification, and Quality Scoring [[ID=⑨]] [[ID=⑩]]
[0285] [[ID=⑪]]The following combines a specific environmental impact assessment case to illustrate the actual application process of the data extraction, unit standardization, verification, and scoring methods described in this example. Taking the "Environmental Impact Report of the 5GW Solar Cell Production and Construction Project of Sany Silicon Energy (Zhuzhou) Co., Ltd." as an example, the project annually produces 625 million TOPCon double-sided monocrystalline cell wafers (182×182mm). Extract the data related to "waste cell wafers" from this environmental impact assessment document, classify it as pollution data, and carry out standardization processing, verification, and quality scoring according to the method of this example. [[ID=⑫]] [[ID=⑬]]
[0286] [[ID=⑭]](1) Unit Standardization [[ID=⑮]] [[ID=⑯]]
[0287] [[ID=⑰]]According to the environmental impact assessment document, the consumption of monocrystalline silicon wafers is 625 million wafers / year, and the production of waste cell wafers is 23.3t / year. The mass of a single cell is about 5g. After conversion, it is: [[ID=⑱]] [[ID=⑲]]
[0288] [[ID=⑳]] [[ID=㉑]] [[ID=㉒]]
[0289] [[ID=㉓]](2) Verification of Industry Benchmark Values [[ID=㉔]] [[ID=㉕]]
[0290] [[ID=㉖]]According to the photovoltaic industry research report in 2022, the average good product rate of TOPCon cell wafers is about 94%, and the corresponding waste product rate is 0.06 pieces / piece, which is set as the industry benchmark value. Taking α = 1.5, calculate the relative error: [[ID=㉗]] [[ID=㉘]]
[0291] [[ID=㉙]]X[[ID=㉚]] bench [[ID=㉛]]=(1 - 94%) = 0.06 pieces / piece [[ID=㉜]] [[ID=㉝]]
[0292] [[ID=㉞]] [[ID=㉟]] [[ID=㊱]]
[0293] [[ID=㊲]]RBD=(1 - exp(-2 * 0.8752))×100 = 82.63% [[ID=㊳]] [[ID=㊴]]
[0294] This value is significantly higher than the recommended threshold (30%), indicating a large data deviation and low accuracy.
[0295] (3) Mass balance verification
[0296] According to the environmental impact assessment document, all the purchased monocrystalline silicon wafer raw materials are qualified products. Calculate the mass deviation rate between the input of monocrystalline silicon wafers and the output of waste battery chips.
[0297]
[0298] The result is lower than the deviation threshold of 2%, indicating that the mass conservation relationship of monocrystalline silicon wafers is basically established.
[0299] (4) Data quality scoring
[0300] According to the data quality scoring indicators and methods proposed in this example, calculate the final score Q' using the comprehensive quality scoring method of the non-linear scoring enhancement function and the key indicator elimination mechanism. Set the adjustable scoring non-linear amplification factor α = 1.1 and the adjustable scoring confidence relief factor β = 0.05. For the monocrystalline silicon wafer data, the cross-check consistency index score is 100, and the original data credibility index score is 50.
[0301]
[0302] Q = 100×(0.9933) 0.5 ×(0.4282) 0.5 = 65.22
[0303] The final score of the monocrystalline silicon wafers is between 50 < Q' < 100, and it is recommended to use after correction;
[0304] For the waste battery chip data, the deviation index score from the reference value is 0, and the original data credibility index score is 50.
[0305]
[0306] Q = 100×(0) 0.5 ×(0.4282) 0.5 = 0
[0307] Since the original score of the key secondary indicator "deviation from the reference value" is 0, the elimination mechanism is triggered, that is, the waste battery chip data cannot be directly used for the development of the background data set.
[0308] 4. Construction of a digital system for data evaluation and screening
[0309] The process of data evaluation and screening is as Figure 1 shown. This digital system adopts a modular hierarchical architecture and specifically includes the following five core modules:
[0310] (1) Data source collection and storage module: crawling environmental impact assessment documents in PDF, Excel, Word and other formats, classifying them according to the five-level structured labeling system (such as draft for comments, draft for review, and draft for approval), building a data source pool, and storing the data sources with standardized file names in the data source pool to facilitate fast and efficient data source extraction;
[0311] (2) Data source quality evaluation module: Establish and embed a data source quality evaluation index system, comprehensive quality evaluation method, and Q value range. By building a data source evaluation engine based on logical rules and machine learning models, the Q value of each environmental impact assessment document data source is automatically calculated, and the data source is automatically classified and removed from the data source pool or data is extracted;
[0312] (3) Data Extraction and Screening Module: Based on technologies such as OCR recognition and natural language processing (NLP), the module automatically parses and extracts various types of data from unstructured environmental impact assessment documents and matches them to corresponding categories (raw materials, energy, auxiliary materials, and pollution). The system uses built-in industry standard values to compare the extracted data, cross-check related data from the same data source, and automatically calculate deviations.
[0313] (4) Data quality scoring module: This module establishes and integrates a data quality scoring indicator system and Q' value range. Based on the rule engine, the Q' value of each data point is automatically calculated, and the score is used to determine whether the data can be directly used for dataset development.
[0314] (5) Data management and export module: High-quality data is automatically stored in the data pool and supports subsequent queries and calls; it provides visual charts such as data quality distribution and scoring to support decision analysis; it also supports export in multiple formats (CSV, JSON, Excel) for the production and development of background data sets.
[0315] The second embodiment of the present application relates to a data screening and evaluation system based on environmental impact assessment documents to support the development of a photovoltaic product carbon footprint background data set, the structure of which is as follows: Figure 2 As shown in the figure, the data screening and evaluation system developed based on the environmental impact assessment documents to support the development of the photovoltaic product carbon footprint background dataset includes:
[0316] Data labeling and storage module, used to construct a multi-level structured labeling system for environmental impact assessment documents and the data therein, and to label and store them;
[0317] The quality evaluation module is used to establish a quality evaluation index system based on the four first-level indicators of file reliability, file representativeness, information comprehensiveness and data readability. The original score sij of each second-level indicator is processed using the following nonlinear confidence enhancement scoring function:
[0318] Where α>1 and β∈(0,0.1) are adjustable parameters to obtain enhanced scores And evaluate the quality of environmental impact assessment documents based on the quality evaluation index system and classify the environmental impact assessment documents into different quality levels;
[0319] The data extraction and classification module is used to extract data from environmental impact assessment documents with high-quality or good grades based on the quality evaluation results of the quality evaluation module, classify the extracted data into raw materials, energy, auxiliary materials and pollution categories, and perform unit standardization processing;
[0320] The data review and verification module is used to review and verify the data extracted and standardized in the data extraction and classification module, and calculate the data quality score;
[0321] The data screening module is used to screen the data based on the data quality score obtained by the data review and verification module to obtain a data pool for establishing a photovoltaic product carbon footprint background data set.
[0322] The first embodiment is a method embodiment corresponding to the present embodiment. The technical details in the first embodiment can be applied to the present embodiment, and the technical details in the present embodiment can also be applied to the first embodiment.
[0323] The above embodiments have the following technical effects:
[0324] The above-mentioned embodiment establishes a structured tagging system, enabling standardized management and efficient retrieval of environmental impact assessment documents. EIA documents are categorized and stored using a five-level structured tagging system (including product tags, time tags, geographic tags, property tags, and quality tags), while data is annotated using a four-level structured tagging system (including file tags, category tags, name tags, and quality tags). This significantly improves the efficiency of querying and extracting EIA documents and their data.
[0325] The multidimensional quality evaluation index system and scoring matrix constructed based on the four first-level indicators (Ii) of document reliability, document representativeness, information comprehensiveness, and data readability provide a systematic standard for the quality assessment of environmental impact assessment document data sources. The above embodiment proposes a comprehensive scoring method for data source quality that integrates a multi-level weighting mechanism, nonlinear scoring enhancement, and a key indicator elimination mechanism. The method processes the original score through a nonlinear confidence enhancement scoring function (Formula 1), effectively avoiding the problem of inflated scores for inferior data sources. This function can more accurately reflect the actual status of each quality dimension by adjusting the parameters α (score nonlinear amplification factor) and β (score confidence release factor). It calculates the comprehensive quality score Q in combination with the weighted geometric mean method (Formula 3), and ensures the screening of high-quality data sources through a key indicator zero-score elimination mechanism (Formula 4).
[0326] To verify the extracted data, the above embodiment proposes a multivariate cross-validation method that combines industry benchmark verification with balance verification. By using the harmonic enhancement normalization function (Formula 6) to achieve relative unit conversion of data, the comparability of production data of different scales is enhanced. By using the asymmetric confidence enhancement deviation function (Formula 8) to calculate deviations from industry benchmark values, data that deviates from industry standards can be effectively screened. Simultaneously, combining element balance, mass balance, and water balance verification mechanisms, this method overcomes the flaws of traditional single-threshold screening methods that are prone to misjudgment, thereby improving the accuracy and consistency of data extraction and evaluation.
[0327] Ultimately, this embodiment, through a digitally integrated system, achieves integrated management of EIA documents and their data collection and storage, data source quality assessment, data extraction, data review and verification, and data quality scoring and screening. Based on OCR recognition and natural language processing technologies, the system can automatically parse and extract various types of data from unstructured EIA documents, significantly improving the efficiency and reliability of the development of carbon footprint background datasets for photovoltaic products, providing high-quality support for carbon footprint accounting and compliance applications.
[0328] Through the combined effect of the above technical effects, the above embodiment solves the current problems in the development of photovoltaic product carbon footprint background data sets, such as scattered data acquisition, inconsistent formats, uneven quality, and lack of systematic screening and verification methods, and provides a strong data foundation support for promoting the low-carbon development of the photovoltaic industry.
[0329] It should be noted that those skilled in the art should understand that the implementation functions of the various modules shown in the above-mentioned embodiments of the data screening and evaluation system for supporting the development of a carbon footprint background dataset for photovoltaic products based on environmental assessment documents can be understood with reference to the relevant description of the data screening and evaluation method for supporting the development of a carbon footprint background dataset for photovoltaic products based on environmental assessment documents. The functions of the various modules shown in the above-mentioned embodiments of the data screening and evaluation system for supporting the development of a carbon footprint background dataset for photovoltaic products based on environmental assessment documents can be implemented by a program (executable instruction) running on a processor, or by a specific logic circuit. If the data screening and evaluation system for supporting the development of a carbon footprint background dataset for photovoltaic products based on environmental assessment documents in the embodiment of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0330] Accordingly, an embodiment of the present application further provides a computer storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the various method embodiments of the present application are implemented.
[0331] In addition, the embodiment of the present application also provides a data screening and evaluation system for the development of a carbon footprint background data set for photovoltaic products based on environmental impact assessment documents, which includes a memory for storing computer-executable instructions, and a processor; the processor is used to implement the steps in the above-mentioned method implementation methods when executing the computer-executable instructions in the memory. Among them, the processor can be a central processing unit (Central Processing Unit, referred to as "CPU"), or other general-purpose processors, digital signal processors (Digital Signal Processor, referred to as "DSP"), application-specific integrated circuits (Application Specific Integrated Circuit, referred to as "ASIC"), etc. The aforementioned memory can be a read-only memory (read-only memory, referred to as "ROM"), random access memory (random access memory, referred to as "RAM"), flash memory (Flash), hard disk or solid-state drive, etc. The steps of the method disclosed in each embodiment of the present invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0332] It should be noted that in this patent application, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element specified by the phrase "comprising a" does not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element. In this patent application, reference to performing an action in accordance with an element means performing the action in accordance with at least that element, including two situations: performing the action in accordance with that element alone, and performing the action in accordance with that element and other elements. Expressions such as "plurality," "multiple times," and "many" include "two," "twice," "two kinds," and "more than two," "more than two times," and "more than two kinds."
[0333] All documents mentioned in this application are considered to be included in their entirety in the disclosure of this application so that they can be used as a basis for modification when necessary. In addition, it should be understood that after reading the above disclosure of this application, those skilled in the art may make various changes or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.
Claims
1. A data screening and evaluation method for developing a photovoltaic product carbon footprint background dataset based on environmental impact assessment documents, characterized in that: The method comprises: Step A: Construct a multi-level structured label system for the environmental impact assessment documents and the data therein, and perform annotation and storage; Step B: Establish a quality evaluation index system based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability, and calculate the original score of each second-level indicator. ij The following nonlinear confidence enhancement scoring function is used for processing: Where α>1 and β∈(0, 0.1) are adjustable parameters to obtain enhanced scores And evaluate the quality of environmental impact assessment documents based on the quality evaluation index system and classify the environmental impact assessment documents into different quality levels; Step C: Based on the quality evaluation results of Step B, extract data from the environmental impact assessment documents with a high-quality or good rating, classify the extracted data into raw materials, energy, auxiliary materials, and pollution categories, and perform unit standardization processing; Step D: Review and verify the data extracted and standardized in Step C, and calculate the data quality score; Step E: Filter the data based on the data quality score obtained in step D to obtain a data pool for establishing a photovoltaic product carbon footprint background dataset.
2. The method according to claim 1, wherein: The environmental impact assessment document labeling system in step A is a five-level structured labeling system, which includes a product label, a time label, a geographic label, a property label, and a quality label in sequence; the data labeling system is a four-level structured labeling system, which includes a document label, a category label, a name label, and a quality label in sequence; After step B, the enhanced scores are normalized and weighted. Summarized as the first-level indicator enhancement score S i , and use the weighted geometric mean method to obtain the comprehensive quality score Q of the environmental impact assessment document, and then divide Q into four levels: high quality, good, poor and invalid; In step D, the review and verification specifically include: performing relative unit conversion on the data by harmonizing and enhancing the normalization function; An asymmetric confidence-enhanced deviation function is used to calculate the deviation between the data and the industry benchmark value; a balance check is performed by combining element balance, mass balance, and water balance; a quality scoring indicator system is constructed based on the three dimensions of data accuracy, data consistency, and data reliability, and a comprehensive scoring method is used to obtain the data quality score, integrating a multi-level weighting mechanism, a nonlinear scoring enhancement function, and a key indicator elimination mechanism. In step E, data with a quality score of 100 are directly included in the data pool, data with a quality score of 50 to 100 are used as reference data after correction, and data with a quality score of 0 to 50 are eliminated.
3. The method according to claim 1, characterized in that The unit normalization process in step C uses the following formula: X std =X raw × C unit (5) Among them, X std is the data value after conversion to standard unit, X raw is the original extracted data, C unit is the unit conversion factor.
4. The method according to claim 2, characterized in that The formula for converting the data to relative units by harmonizing the enhanced normalization function in step D is: Among them, X rel is the normalized relative data value, X std The original data after unit standardization, Q prod is the product output corresponding to the data, ∈ is the minimum value (used to avoid division by zero), and δ is the normalized harmonic parameter (used to enhance the robustness of small-output data). By introducing logarithmic terms and harmonic parameters, this function can improve the stability of low-output samples and enhance the comparability of data.
5. The method according to claim 2 or 4, characterized in that The formula for calculating the deviation between the data and the industry benchmark using the asymmetric confidence enhancement deviation function in step D is: RBD = (1-exp(-α×RE))×100 (8) Among them, RBD is the enhanced industry deviation score, ranging from 0 to 100, and the lower the score, the greater the deviation; RE is the original relative error, calculated as X bench is the industry benchmark value of the data item; α is the adjustment coefficient (used to control the strictness of the deviation constraint); this function introduces nonlinear exponential decay to impose stricter penalties on high-deviation data, thereby improving the sensitivity of deviation assessment.
6. The method according to claim 2, characterized in that The comprehensive scoring method using the zero-score elimination mechanism for key indicators in both Step B and 400 is: Q=min(Q, Q cut-off ) (4) Among them, Q is the comprehensive quality score, Q cut-off is the elimination threshold; in step B, when the original score of any key secondary indicator among the key data missing rate, data format structured degree, data and information readability is 0, the elimination mechanism is triggered, and the elimination threshold Q cut-off Set to 29.9; in step D, when the original score of any key secondary indicator in the deviation from the benchmark value and the cross-check consistency is 0, the elimination mechanism is triggered and the elimination threshold is set to 49.
9.
7. The method according to claim 2, wherein: When the comprehensive quality score Q of the environmental assessment document is between 70 and 100, it is defined as a high-quality grade; When Q is between 50 and 70, it is defined as a good grade; When Q is between 30 and 50, it is defined as a poor grade; When Q is between 0 and 30, it is defined as an invalid level.
8. The method according to claim 2, 6 or 7, characterized in that: The weighted geometric mean method in step B is used to calculate the comprehensive score Q of the environmental impact assessment document using the following formula: Among them, S i is the enhanced score of the i-th first-level indicator, W i is the global weight of the i-th first-level indicator and satisfies ∑W i =1, n is the number of primary indicators and n=4; this geometric mean method ensures that if the score of any indicator is too low, the overall score will be significantly lowered, reflecting the synergistic importance of each indicator.
9. The method according to claim 2 or 4, characterized in that The balance check in step D includes calculations of the element balance deviation rate, mass balance deviation rate, and water balance deviation rate. The deviation rate thresholds are all set at 2%. Exceeding the thresholds indicates that there may be problems with the data. These are used to verify the input and output balance of key elements (such as silicon, chlorine, and fluorine), conservation of material mass, and input and output balance of water resources in the photovoltaic product production process.
10. The method according to claim 2, characterized in that The category labels in step A include raw materials, energy, auxiliary materials and pollution. The raw materials category identifies materials with silicon purity ≥99.9999% (wt), the energy category identifies energy in electricity measurement units (kWh) or steam pressure parameters, the auxiliary materials category identifies materials with solder ribbon wire diameter ≤0.3mm or backplane transmittance ≥93% or cover glass thickness 2mm, and the pollution category identifies pollutants with emission volume t / a.
11. The method according to claim 1, wherein The method is implemented through a digital system, which includes a data source acquisition and storage module, a data source quality evaluation module, a data extraction and screening module, a data quality scoring module, and a data management and export module.
12. A data screening and evaluation system based on environmental impact assessment documents to support the development of a photovoltaic product carbon footprint background dataset, characterized in that: The system comprises: Data labeling and storage module, used to construct a multi-level structured labeling system for environmental impact assessment documents and the data therein, and to label and store them; The quality evaluation module is used to establish a quality evaluation index system based on the four first-level indicators of document reliability, document representativeness, information comprehensiveness and data readability, and to evaluate the original scores of each second-level indicator. ij The following nonlinear confidence enhancement scoring function is used for processing: Where α>1 and β∈(0, 0.1) are adjustable parameters to obtain enhanced scores And evaluate the quality of environmental impact assessment documents based on the quality evaluation index system and classify the environmental impact assessment documents into different quality levels; The data extraction and classification module is used to extract data from environmental impact assessment documents with high-quality or good grades based on the quality evaluation results of the quality evaluation module, classify the extracted data into raw materials, energy, auxiliary materials and pollution categories, and perform unit standardization processing; The data review and verification module is used to review and verify the data extracted and standardized in the data extraction and classification module, and calculate the data quality score; The data screening module is used to screen the data based on the data quality score obtained by the data review and verification module to obtain a data pool for establishing a photovoltaic product carbon footprint background data set.
Citation Information
Patent Citations
Crystalline silicon photovoltaic module raw material silica sand carbon footprint analysis method and system
CN115907464A
Environmental assessment report auxiliary writing method and system
CN117113943A
Method for calculating product carbon footprint accounting data quality
CN118626768A
Calculation method for enterprise carbon emission accounting data quality
CN118708873A
Calculation method for project carbon emission accounting data quality
CN118797232A
Cited By
Authentication method and system for automobile collision digital human body model
CN121302564A