An enterprise ESG index determination method and system based on data fusion

The data fusion method addresses data conflicts in ESG index determination by evaluating and reconciling data sources, ensuring accurate and consistent ESG index assessments.

CN119477059BActive Publication Date: 2025-07-15CHINA NAT INST OF STANDARDIZATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411526280.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-07-15
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

In the prior art, data conflicts exist when multiple data sources determine the enterprise ESG index, resulting in inaccurate assessment and insufficient transparency.

Method used

By obtaining the data set to be analyzed for each data source, data processing is carried out to obtain conflict evaluation data, calculate the data conflict evaluation index, and compare it with the threshold to obtain the data to be used, and conduct change trend analysis with the comparison data set to judge whether to check the data to ensure the consistency and accuracy of the data.

Benefits of technology

It improves the accuracy and transparency of the enterprise ESG index, reduces data processing errors, improves processor execution efficiency, and reduces the error rate determined by ESG index.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477059B_ABST
    Figure CN119477059B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for determining an enterprise ESG index based on data fusion, which relates to the technical field of enterprise ESG index determination supervision. The method for determining an enterprise ESG index based on data fusion includes the following steps: data acquisition, data conflict evaluation, and data change analysis. The present invention obtains conflict evaluation data from the to-be-analyzed data set of the to-be-analyzed index, then compares the data conflict evaluation index obtained based on the conflict evaluation data with the data conflict evaluation threshold and obtains the corresponding standby data. Then, it obtains the control data set of the to-be-analyzed index, and obtains the change trend analysis index based on the standby data and the control data set. Finally, it compares the change trend analysis index with the change trend judgment threshold to determine whether to perform data verification, improving the accuracy of the enterprise ESG index determined based on data fusion and solving the problem of data conflicts in determining the enterprise ESG index from multiple data sources in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enterprise ESG index determination and supervision, and particularly relates to a method and system for determining an enterprise ESG index based on data fusion. Background Art

[0002] With the continuous deepening of the concept of global sustainable development, the environmental (E), social (S), and governance (G) performance of enterprises has become an important indicator for measuring their long-term value and risks. Since ESG data comes from diverse sources and includes structured and unstructured information such as financial reports, social public opinions, policies and regulations, and environmental monitoring, data fusion technology plays a key role in the determination of ESG indexes. Through data fusion, data from different sources, formats, and types can be integrated to achieve effective unification and analysis of information. In addition, an ESG system integrating technologies such as artificial intelligence, natural language processing, and big data processing can dynamically capture changes in enterprises' environmental, social responsibility, and corporate governance aspects. Establishing an ESG index system based on data fusion can not only improve the comprehensiveness and real-time nature of evaluations, but also provide a scientific basis for investment decisions, risk management, and corporate strategies, promoting a more transparent and responsible sustainable development of the market.

[0003] Existing methods for determining enterprise ESG indexes mainly rely on a single data source or limited data dimensions, such as financial data, sustainable development reports disclosed by enterprises, etc. These methods usually have problems such as incomplete data, poor timeliness, and insufficient ability to process multi-dimensional information. In addition, most existing systems adopt static evaluation models and are difficult to capture real-time changes in enterprises' environmental, social, and governance aspects. At the same time, the accuracy of natural language processing and sentiment analysis technologies still has limitations when processing unstructured data (such as news, social media). The heterogeneity, cross-domain nature, and complexity of data further increase the difficulty of information integration, resulting in insufficient transparency, dynamics, and accuracy of existing ESG evaluation methods, and it is difficult to meet the real-time, comprehensive, and reliable ESG evaluation needs of investors and regulatory agencies.

[0004] For example, the method for determining an enterprise ESG index based on data completion and related products announced in the invention patent with the announcement number: CN113313362B includes: obtaining enterprise data of the enterprise to be evaluated within a preset time period; in the case where M is less than N, based on the enterprise data of the enterprise to be evaluated within the preset time period and the enterprise data of multiple first enterprises within the preset time period, completing the disclosure data related to evaluation indicator A of the enterprise to be evaluated to obtain the disclosure data related to evaluation indicator A of the enterprise to be evaluated at N historical moments; obtaining the disclosure data related to each evaluation indicator of the enterprise to be evaluated at N historical moments according to the disclosure data related to evaluation indicator A of the enterprise to be evaluated at N historical moments; performing an ESG evaluation on the enterprise to be evaluated according to the disclosure data related to each evaluation indicator of the enterprise to be evaluated at N historical moments to obtain the ESG index of the enterprise to be evaluated.

[0005] For example, the method for determining an enterprise ESG index and related products announced in the invention patent with the announcement number: CN113240272B includes: obtaining a set of evaluation indicators; obtaining the historical ESG evaluation data of multiple enterprises and the historical returns of multiple enterprises; using the historical ESG evaluation data of multiple enterprises as training samples and using the historical returns of multiple enterprises as training labels to perform model training to obtain the actual weights of each evaluation indicator in each evaluation dimension; determining the initial weights of each evaluation indicator in each evaluation dimension according to the historical ESG evaluation data of multiple enterprises; fusing the actual weights and initial weights of each evaluation indicator in each evaluation dimension to obtain the target weights of each evaluation indicator in each evaluation dimension; obtaining the ESG index of the enterprise to be evaluated according to the enterprise data of the enterprise to be evaluated and the target weights of each evaluation indicator in each evaluation dimension.

[0006] However, in the process of implementing the technical solution of the invention in the embodiments of the present application, it is found that the above technology has at least the following technical problems:

[0007] In the prior art, after data collection and standardization processing, since the data is collected from various data sources, the data disclosed by each data source for the same evaluation indicator may be different, and there is a problem of data conflict in determining the enterprise ESG index with multiple data sources. Summary of the Invention

[0008] By providing a method and system for determining an enterprise ESG index based on data fusion in the embodiments of the present application, the problem of data conflict in determining the enterprise ESG index with multiple data sources in the prior art is solved, and the accuracy of the enterprise ESG index determined based on data fusion is improved.

[0009] An embodiment of the present application provides a method for determining an enterprise ESG index based on data fusion, comprising the following steps: obtaining a data set to be analyzed of an indicator to be analyzed through various data sources, and obtaining conflict assessment data based on the data set to be analyzed; obtaining a data conflict assessment index based on the conflict assessment data, comparing the data conflict assessment index with a data conflict assessment threshold obtained from a preset database and obtaining corresponding stand-by data, wherein the data conflict assessment index is used to quantify the degree of data conflict of the data set to be analyzed of the indicator to be analyzed; obtaining a control data set of the indicator to be analyzed from a preset enterprise historical database, obtaining a change trend analysis index based on the stand-by data and the control data set, comparing the change trend analysis index with a change trend judgment threshold and judging whether to perform data verification, wherein the change trend analysis index is used to quantify the degree of conformity of the stand-by data with the historical change trend.

[0010] Furthermore, the specific process of obtaining the conflict assessment data is as follows: the conflict assessment data is obtained by processing the data set to be analyzed, the data set to be analyzed represents a data set obtained from each data source for evaluating the corresponding indicator to be analyzed, and the indicator to be analyzed is used to determine the enterprise ESG index; the conflict assessment data includes variance, standard deviation, range, interquartile range, mean absolute deviation and median absolute deviation; the conflict assessment data is obtained, and before that, data source extraction monitoring is also performed: the specific steps of data source extraction monitoring are as follows: A1, obtaining extraction data through a network detection tool, the extraction data includes extraction time, data packet size, extraction transmission rate, data loading delay and connection retry number; A2, obtaining reference extraction data from a preset database, the reference extraction data includes Average extraction time, average data packet size, minimum network transmission rate, maximum data loading delay and maximum number of connection retries; A3, number the extracted data according to the data type. If the extracted data meets the extraction condition, the data extraction analysis index is obtained based on the extracted data and the reference extracted data of the corresponding data type, otherwise the data extraction analysis index is recorded as 1, and the extraction condition indicates that the extraction transmission rate is greater than the minimum network transmission rate and the data loading delay is less than the maximum data loading delay and the number of connection retries is less than the maximum number of connection retries; A4, compare the data extraction analysis index with the data extraction threshold: if the data extraction analysis index is greater than the data extraction threshold, extract the data of the next data source, otherwise perform secondary extraction on the data of the corresponding data source until the data of the next data source is extracted.

[0011] Furthermore, the specific process of obtaining the data conflict assessment index based on the conflict assessment data is as follows: numbering the indicators to be analyzed and obtaining the conflict assessment data; performing logarithmic processing on the variance and standard deviation of the data set to be analyzed to obtain the discrete degree assessment value of the data set to be analyzed; performing hyperbolic tangent processing on the range of the data set to be analyzed to obtain the difference degree assessment value of the data set to be analyzed; performing hyperbolic sine processing on the interquartile range, mean absolute deviation and median absolute deviation of the data set to be analyzed to obtain the aggregation degree assessment value of the data set to be analyzed; and obtaining the data conflict assessment index of the data set to be analyzed by performing square root processing and hyperbolic secant processing on the sum of the discrete degree assessment value, the difference degree assessment value and the aggregation degree assessment value.

[0012] Furthermore, the specific acquisition process of the stand-by data is as follows: a data conflict assessment threshold is obtained from a preset database, and compared with a data conflict assessment index: if the data conflict assessment index is less than the data conflict assessment threshold, the data in the data set to be analyzed is averaged to obtain the stand-by data; if the data conflict assessment index is not less than the data conflict assessment threshold, the data to be analyzed is combined with the assessment weights of each data source to obtain the stand-by data; the assessment weights include a first-level assessment weight mapping set and a second-level assessment weight mapping set, the first-level assessment weight mapping set represents a set of weights assigned to each data source according to the type and authority of each data source, and the second-level assessment weight mapping set represents a set of weights corresponding to the industry status corresponding to the same type of data source.

[0013] Furthermore, the specific method for obtaining the change trend analysis index is as follows: obtain a reference data set and number them according to the order of the data sequence; perform data processing on the obtained standby data and the reference data set to obtain the change trend analysis index; the restriction expression of the change trend analysis index is as follows:

[0014]

[0015] In the formula, d represents the number of the index to be analyzed, d = 1, 2, ..., D, D represents the total number of the index to be analyzed, q represents the sequence number of the data in the control data set, q = 1, 2, ..., Q, Q represents the total number of data in the control data set, ρ d represents the standby data of the dth indicator to be analyzed, δ q represents the qth data in the control data set, δ q+1 represents the q+1th data in the control data set, ω d It represents the change trend analysis index of the dth indicator to be analyzed, and e represents the natural constant.

[0016] Further, the specific process of determining whether to perform data verification is as follows: If the change trend analysis index is greater than the change trend judgment threshold, data verification is not performed; if the change trend analysis index is not greater than the change trend judgment threshold, a preset staff member is prompted to perform data verification; the change trend judgment threshold is used to judge the degree of deviation between the data to be used and the change trend of the control data set.

[0017] Further, the specific process of the data verification is as follows: Step 1, send a first determination button to the preset staff member. If the preset staff member selects "yes", it is determined that the data to be used is for the evaluation of the index to be analyzed; otherwise, step 2 is executed. Step 2, send the data set to be analyzed and the corresponding data sources to the preset staff member and a data verification prompt instruction. Step 3, send a second determination button to the preset staff member. If the preset staff member selects "correct", a data source self-check prompt instruction is sent; otherwise, the data is re-extracted from each data source and the data set to be analyzed is updated. Step 4, re-obtain the data conflict evaluation index according to the data set to be analyzed changed by the preset staff member until data verification is not performed.

[0018] Further, after sending the data source self-check prompt instruction, it also includes obtaining the feedback result of the data source self-check prompt instruction; sending a data source self-check feedback instruction to the preset staff member. If the preset staff member feedbacks that the data source data is incorrect, the data source is re-obtained and the data set to be analyzed is updated; if the preset staff member feedbacks that the data source data is correct, the updated data conflict evaluation index is obtained according to the data set to be analyzed, and the conflict evaluation similarity coefficient is obtained based on the absolute value of the difference between the updated data conflict evaluation index and the original data conflict evaluation index and the operation result of the ratio of the updated data conflict evaluation index to the original data conflict evaluation index; compare the conflict evaluation similarity coefficient with the conflict similarity evaluation threshold and judge whether there is an error in the enterprise ESG index determination method; the original data conflict evaluation index represents the data conflict evaluation index obtained according to the data set to be analyzed without change; the updated data conflict evaluation index represents the data conflict evaluation index obtained according to the updated data set to be analyzed.

[0019] Further, the specific process of judging whether there is an error in the enterprise ESG index determination method is as follows: If the conflict evaluation similarity coefficient is not greater than the conflict similarity evaluation threshold, it indicates that the enterprise ESG index determination method is correct; if the conflict evaluation similarity coefficient is greater than the conflict similarity evaluation threshold, an error warning is issued and the programmer is requested to debug the program; the conflict similarity evaluation threshold is used to judge whether there is a program error in the enterprise ESG index determination method.

[0020] An embodiment of the present application provides an enterprise ESG index determination system based on data fusion, characterized in that it includes a data acquisition module, a data conflict assessment module and a data change analysis module; wherein the data acquisition module is used to acquire a data set to be analyzed of an indicator to be analyzed through various data sources, and obtain conflict assessment data based on the data set to be analyzed; the data conflict assessment module is used to obtain a data conflict assessment index based on the conflict assessment data, compare the data conflict assessment index with a data conflict assessment threshold obtained from a preset database and obtain corresponding stand-by data, and the data conflict assessment index is used to quantify the degree of data conflict of the data set to be analyzed of the indicator to be analyzed; the data change analysis module is used to acquire a control data set of the indicator to be analyzed from a preset enterprise historical database, obtain a change trend analysis index based on the stand-by data and the control data set, compare the change trend analysis index with the change trend judgment threshold and determine whether to perform data verification, and the change trend analysis index is used to quantify the degree of conformity of the stand-by data with the historical change trend.

[0021] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0022] 1. Obtain conflict assessment data by obtaining the data set to be analyzed of the indicator to be analyzed, and then obtain the data conflict assessment index based on the conflict assessment data, compare the data conflict assessment index with the data conflict assessment threshold and obtain the corresponding stand-by data, then obtain the reference data set of the indicator to be analyzed, and obtain the change trend analysis index based on the stand-by data and the reference data set, and then compare the change trend analysis index with the change trend judgment threshold to determine whether to perform data verification, so as to more reasonably confirm the data used to evaluate the ESG index, thereby achieving an improvement in the accuracy of the enterprise ESG index determined based on data fusion, and effectively solving the problem of data conflict in determining the enterprise ESG index from multiple data sources in the prior art.

[0023] 2. The extracted data is numbered by data type. Then, if the extracted data meets the extraction conditions, the data extraction analysis index is obtained based on the obtained extracted data and the reference extracted data. Otherwise, the data extraction analysis index is recorded as 1. Finally, the data extraction analysis index is compared with the data extraction threshold to determine whether to extract data from the next data source, thereby reducing the waste of processor resources caused by data extraction errors for subsequent data processing, thereby improving the execution efficiency of the processor.

[0024] 3. By obtaining the reference data set and numbering it in order of the data sequence, the obtained stand-by data and the reference data set are processed to obtain the change trend analysis index, thereby effectively evaluating whether the stand-by data conforms to the change trend, and thus achieving a more accurate determination of the ESG index.

[0025] 4. By sending the data source self-check feedback instruction to the preset staff, if the preset staff feedback that the data source data is incorrect, the data source is re-acquired and the data set to be analyzed is updated. If the preset staff feedback that the data source data is correct, the updated data conflict assessment index is obtained according to the data set to be analyzed, and the conflict assessment similarity coefficient is obtained based on the original data conflict assessment index, which is helpful to judge whether there is an error in the method of determining the enterprise ESG index, thereby reducing the error rate of obtaining the ESG index. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A flowchart of a method for determining an enterprise ESG index based on data fusion provided in an embodiment of the present application;

[0027] Figure 2 A schematic diagram of the change of the conflict assessment similarity coefficient provided in the embodiment of the present application;

[0028] Figure 3 A structural diagram of a system for determining an enterprise ESG index based on data fusion provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The embodiment of the present application solves the problem of data conflict in determining the enterprise ESG index from multiple data sources in the prior art by providing a method and system for determining the enterprise ESG index based on data fusion. The method obtains the data set to be analyzed of the indicator to be analyzed through each data source, obtains conflict assessment data based on the data set to be analyzed, and then obtains the data conflict assessment index based on the conflict assessment data. The data conflict assessment index is compared with the data conflict assessment threshold and the corresponding stand-by data is obtained. Then, a reference data set of the indicator to be analyzed is obtained from a preset enterprise historical database, and the reference data set is obtained and numbered in the order of the data sequence. The obtained stand-by data and the reference data set are processed to obtain a change trend analysis index. Then, the change trend analysis index is compared with the change trend judgment threshold and it is judged whether to perform data verification, thereby improving the accuracy of the enterprise ESG index determined based on data fusion.

[0030] The technical solution in the embodiment of the present application is to solve the problem of data conflict in determining the ESG index of an enterprise from multiple data sources. The overall idea is as follows:

[0031] The conflict assessment data is obtained by acquiring the data set to be analyzed, and then the data conflict assessment index obtained based on the conflict assessment data is compared with the data conflict assessment threshold to obtain the corresponding stand-by data, and the change trend analysis index is obtained based on this and the control data set of the indicator to be analyzed. Finally, the change trend analysis index is compared with the change trend judgment threshold to determine whether data verification is performed, thereby achieving the effect of improving the accuracy of the enterprise ESG index determined based on data fusion.

[0032] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0033] like Figure 1 As shown, it is a flow chart of a method for determining an enterprise ESG index based on data fusion provided by an embodiment of the present application, the method comprising the following steps: obtaining a data set to be analyzed of the indicator to be analyzed through various data sources, and obtaining conflict assessment data based on the data set to be analyzed; obtaining a data conflict assessment index based on the conflict assessment data, comparing the data conflict assessment index with a data conflict assessment threshold obtained from a preset database and obtaining corresponding stand-by data, the data conflict assessment index being used to quantify the degree of data conflict of the data set to be analyzed of the indicator to be analyzed; obtaining a control data set of the indicator to be analyzed from a preset enterprise historical database, obtaining a change trend analysis index based on the stand-by data and the control data set, comparing the change trend analysis index with a change trend judgment threshold and judging whether to perform data verification, the change trend analysis index being used to quantify the degree of conformity of the stand-by data with the historical change trend, and data verification means re-verifying the data in the data set to be analyzed.

[0034] In this embodiment, the data set to be analyzed represents a collection of digital data used to determine the relevant indicators of the enterprise ESG index. The data set to be analyzed includes three aspects: environment, society and governance. The indicators related to the environment include energy consumption, water resource usage and total waste, etc., the indicators related to society include employee turnover rate, safety accident rate and customer complaint rate, etc., and the indicators related to governance include the proportion of women on the board of directors, shareholder voting participation rate and tax transparency, etc.; and the data sources for obtaining the relevant indicators of the enterprise ESG index usually include internal data sources of the enterprise (annual report, financial system, environmental management system and enterprise operation log, etc.), supply chain and partner data sources (supplier report, supply chain management system and third-party supplier audit, etc.), external Third-party data sources (government agencies and public databases, academic reports and market analysis reports, and third-party rating agencies, etc.), social media and public opinion data (social media, news websites, and survey reports, etc.), and industry benchmarks and peer comparison data (industry benchmark reports and peer ESG reports, etc.); through a multi-step evaluation and analysis of data conflicts and changing trends, it is possible to identify anomalies or conflicts in the data set to be analyzed, thereby screening out more reliable data to be used and ensuring data accuracy and consistency; and, by comparing with data in historical databases, it is possible to ensure that new data is consistent with historical trends and prevent data bias from affecting analysis results, thereby improving the accuracy of data management and the accuracy of the corporate ESG index determined based on data fusion.

[0035] It should be noted that when determining the enterprise ESG index, the indicator to be analyzed can represent various indicator data, including total water consumption, wastewater discharge, total energy consumption, proportion of renewable energy use, carbon intensity, greenhouse gas emissions, employee turnover rate, gender ratio of employees, ethnic ratio of employees, workplace accident rate, employee training duration, customer satisfaction, proportion of independent directors, gender ratio of board members, and financial transparency, etc. Subsequently, the wastewater discharge will be taken as an example for explanation and analysis.

[0036] Furthermore, the specific process of obtaining the conflict assessment data is as follows: Extract the corresponding data from each data source according to the indicator to be analyzed and perform statistics to obtain the dataset to be analyzed; Perform data processing on the dataset to be analyzed to obtain the conflict assessment data. The dataset to be analyzed represents the data set obtained from each data source for evaluating the corresponding indicator to be analyzed, and the indicator to be analyzed is used to determine the enterprise ESG index; The conflict assessment data includes variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation; Before obtaining the conflict assessment data, it also includes performing data source extraction monitoring: The specific steps of data source extraction monitoring are as follows: A1, Obtain the extracted data through a network detection tool. The extracted data includes extraction time, data packet size, extraction transmission rate, data loading delay, and connection retry times; A2, Obtain the reference extraction data from a preset database. The reference extraction data includes average extraction time, average data packet size, minimum network transmission rate, maximum data loading delay, and maximum connection retry times; A3, Number the extracted data according to the data type. If the extracted data meets the extraction conditions, obtain the data extraction analysis index based on the extracted data and the reference extraction data of the corresponding data type, otherwise record the data extraction analysis index as 1. The extraction conditions mean that the extraction transmission rate is greater than the minimum network transmission rate, the data loading delay is less than the maximum data loading delay, and the connection retry times are less than the maximum connection retry times; A4, Compare the data extraction analysis index with the data extraction threshold: If the data extraction analysis index is greater than the data extraction threshold, extract the data of the next data source, otherwise perform secondary extraction on the data of the corresponding data source until the data extraction analysis index is greater than the data extraction threshold, and extract the data of the next data source.

[0037] The numerical expression of the data extraction analysis index is specifically as follows:

[0038]

[0039] In the formula, n represents the number of the data type, n = 1, 2,..., N, and N represents the total number of data types. T n represents the extraction time of the nth type of data, D n represents the data packet size of the nth type of data, S n represents the extraction transmission rate of the nth type of data, L nIndicates the data loading delay of the nth type of data, C n Indicates the number of connection retries for the nth type of data Indicates the average extraction time Indicates the average packet size, S min Indicates the minimum network transmission rate, L max Indicates the maximum data loading delay, C max Indicates the maximum number of connection retries, DE n Indicates the data extraction analysis index of the nth type of data, where e represents the natural constant

[0040] In this embodiment, data on the wastewater discharge of the to-be-analyzed indicators is obtained from each data source to obtain a to-be-analyzed data set corresponding to the wastewater discharge. By performing variance operation, standard deviation operation, range operation, interquartile range operation, mean absolute deviation operation, and median absolute deviation operation on the to-be-analyzed data set, variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation are obtained respectively. By precisely quantifying the data conflict degree between each data source, it provides a basis for subsequent data screening, and can ensure the accuracy and reliability of ESG indicator data, thereby enhancing the credibility of the evaluation results of the enterprise in terms of environment, society, and governance (ESG); at the same time, by analyzing the discreteness and central tendency of the data related to the to-be-analyzed indicators provided by each data source, it helps to quickly evaluate the differences in data sources and ensure the coordination and consistency after data integration

[0041] Specifically, the algorithm of this embodiment comprehensively analyzes the extracted data and the reference extracted data to obtain the data extraction analysis index. In the formula, when the extraction time is closer to the average extraction time and the packet size is closer to the average packet size, the larger the data extraction analysis index indicates the smaller the possibility of data extraction error. Similarly, when the extraction transmission rate is greater than the minimum network transmission rate, the data loading delay is less than the maximum data loading delay, and the number of connection retries is less than the maximum number of connection retries, the value of the data extraction analysis index is also larger, indicating the smaller the possibility of data extraction error. And in the ideal situation where the network state is normal and the data source is stable, the extraction time should be proportional to the packet size and inversely proportional to the extraction transmission rate, and the data loading delay and the number of retries will decrease; by comparing the relationships between the above physical quantities, it helps to judge whether there are abnormalities or other transmission problems in the data extraction process, and at the same time is beneficial to reducing the waste of processor resources caused by data extraction errors for subsequent data processing and improving the execution efficiency of the processor

[0042] Specifically, the reference extraction data is obtained from a preset database. Among them, the average extraction time is obtained by calculating the mean of the historical extraction times of the corresponding type of data, and the average data packet size is obtained by calculating the mean of the historical data packet sizes of the corresponding type of data; the minimum network transmission rate, the maximum data loading delay, and the maximum number of connection retries are set according to the historical data of the corresponding data type. For example, if the minimum network transmission rate of error-free data extraction in the same network environment in historical data is 60 bps, then the minimum network transmission rate is set to 60 bps; if the maximum data loading delay of error-free data extraction in historical data is 10 ms, then the maximum data loading delay is set to 10 ms; if the maximum number of connection retries of error-free data extraction in historical data is 3 times, then the maximum number of connection retries is set to 3 times.

[0043] Specifically, the data extraction threshold is obtained from a preset database. In a specific embodiment, the extraction data corresponding to all cases of data extraction errors in the historical data is substituted into the numerical expression of the data extraction analysis index to obtain a data set of the data extraction analysis index, and the result of calculating the mean of the data set is set as the data extraction threshold.

[0044] Further, the specific process of obtaining the data conflict evaluation index based on the conflict evaluation data is as follows: number the index to be analyzed and obtain the conflict evaluation data; perform logarithmic processing on the variance and standard deviation of the data set to be analyzed to obtain the evaluation value of the dispersion degree of the data set to be analyzed; perform hyperbolic tangent processing on the range of the data set to be analyzed to obtain the evaluation value of the difference degree of the data set to be analyzed; perform hyperbolic sine processing on the interquartile range, the mean absolute deviation, and the median absolute deviation of the data set to be analyzed to obtain the evaluation value of the aggregation degree of the data set to be analyzed; obtain the data conflict evaluation index of the data set to be analyzed by performing square root processing and hyperbolic secant processing on the sum of the evaluation value of the dispersion degree, the evaluation value of the difference degree, and the evaluation value of the aggregation degree.

[0045] The specific limit expression of the data conflict evaluation index is as follows:

[0046]

[0047] In the formula, d represents the number of the index to be analyzed, d = 1, 2,..., D, and D represents the total number of indexes to be analyzed. represents the variance of the d-th index to be analyzed. represents the standard deviation of the d-th index to be analyzed. represents the range of the d-th index to be analyzed. represents the interquartile range of the d-th index to be analyzed. represents the mean absolute deviation of the d-th index to be analyzed. Denote the median absolute deviation of the d-th metric to be analyzed. Denote the data conflict assessment index of the d-th metric to be analyzed.

[0048] In this embodiment, the algorithm comprehensively analyzes the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation to obtain the data conflict assessment index. Taking the wastewater discharge as an example of the metric to be analyzed, the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation of the dataset to be analyzed corresponding to the wastewater discharge are all negatively correlated with the data conflict assessment index. When the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation of the dataset to be analyzed corresponding to the wastewater discharge are larger, the data conflict assessment index is smaller, indicating that the data differences in the dataset to be analyzed of the wastewater discharge may be larger, and thus the possibility of data conflict is also larger. And the standard deviation is the square root of the variance, they are both calculated based on the mean and have the same direction of change. The mean absolute deviation is another form of the variance, which calculates the absolute deviation of the data points from the mean. The range only considers the maximum and minimum values, so it is easily affected by extreme values. The interquartile range focuses on the middle 50% of the data by removing the extreme values in the data. The interquartile range and the median absolute deviation are not affected by extreme values and are both robust measures of dispersion. The interquartile range focuses on the middle 50% of the data, while the median absolute deviation focuses on the deviation of all data points from the median. The data change table of the data conflict assessment index is obtained by combining the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation, as shown in Table 1 specifically:

[0049] Table 1 Data change table of the data conflict assessment index

[0050]

[0051] As can be seen from Table 1, when the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation of the dataset to be analyzed corresponding to the wastewater discharge are all 0, the data conflict assessment index is the largest (with a value of 1), indicating that there is no data conflict in the wastewater discharge, and the data provided by each data source is consistent. If the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation of the dataset to be analyzed corresponding to the wastewater discharge are larger, the data conflict assessment index is smaller, indicating a greater possibility of data conflict in the wastewater discharge. For example, when the variance in the third row of the table is 24.3, the standard deviation is 4.93, the range is 9, the interquartile range is 9, the mean absolute deviation is 4.32, and the median absolute deviation is 10, each independent variable reaches its maximum, and at this time, the data conflict assessment index for the wastewater discharge approaches 0, indicating that the possibility of data conflict reaches the maximum. Therefore, through the comprehensive analysis of the variance, standard deviation, range, interquartile range, mean absolute deviation, and median absolute deviation, it is possible to effectively identify the data conflicts regarding the wastewater discharge and take corresponding measures in a timely manner to intervene and improve the accuracy of the determined data to be used.

[0052] Furthermore, the specific process for obtaining the data to be used is as follows: Obtain the data conflict assessment threshold from the preset database and compare it with the data conflict assessment index. If the data conflict assessment index is less than the data conflict assessment threshold, perform a mean operation on the data in the dataset to be analyzed to obtain the data to be used. If the data conflict assessment index is not less than the data conflict assessment threshold, combine the data to be analyzed with the evaluation weights of each data source to obtain the data to be used. The evaluation weights include a first-level evaluation weight mapping set and a second-level evaluation weight mapping set. The first-level evaluation weight mapping set represents a set of weights assigned to each data source based on the type and authority of each data source, and the second-level evaluation weight mapping set represents a set of weights corresponding to the industry status of data sources of the same type.

[0053] In this embodiment, both the first-level evaluation weight mapping set and the second-level evaluation weight mapping set are obtained from the preset database. Specifically, the first-level evaluation weight mapping set is constructed based on the type of each data source and its corresponding weight, with a one-to-one correspondence between one type of data source and one weight. The second-level evaluation weight mapping set is constructed based on the weights corresponding to the industry status of data sources of the same type, also with a one-to-one correspondence between one data source and one weight. By using the first-level evaluation weight mapping set and the second-level evaluation weight mapping set, hierarchical processing of different data sources is achieved, and data sources with high authority and high industry status are given priority, ensuring that the finally obtained data to be used is more reliable, improving the data quality and credibility. At the same time, according to the comparison result of the data conflict assessment threshold, the corresponding calculation method (mean operation or weight operation) is selected, avoiding unnecessary complex processing and improving the efficiency of data integration.

[0054] Specifically, the first-level evaluation weight mapping set is the weight corresponding to different types of data sources in the preset database, which represents the numerical value of the influence degree of different types of data sources on the data to be used. When used, the weight corresponding to different types of data sources can be directly obtained from the preset database, and its corresponding relationship can be a pre-set mapping relationship. For example, the first-level evaluation weights corresponding to different types of data sources form a mapping set with the weights corresponding to the data to be used preset in the preset database. By inputting different types of data sources into the mapping set, the corresponding weights can be obtained, and the mapping relationship therein is one-to-one. In this example, its value range is [0, 1].

[0055] Specifically, the second-level evaluation weight mapping set is the weight corresponding to different data sources of the same type in the preset database, which represents the numerical value of the influence degree of different data sources of the same type on the first-level evaluation weight. When used, the weight corresponding to different data sources of the same type can be directly obtained from the preset database, and its corresponding relationship can be a pre-set mapping relationship. For example, the second-level evaluation weights corresponding to different data sources of the same type form a mapping set with the weights corresponding to the first-level evaluation weights preset in the preset database. By inputting different data sources of the same type into the mapping set, the corresponding weights can be obtained, and the mapping relationship therein is one-to-one. In this example, its value range is [0, 1].

[0056] Specifically, the data conflict evaluation threshold is obtained from the preset database. In a specific embodiment, the data set with data conflicts is obtained from the historical data, and the corresponding conflict evaluation data is obtained according to the data set. These data conflict evaluation data are substituted into the specific limit expression of the data conflict evaluation index to obtain the corresponding data conflict evaluation index, and the mean operation is performed on the obtained data conflict evaluation index to obtain the corresponding average value, and this average value is recorded as the data conflict evaluation threshold.

[0057] Further, the specific method for obtaining the change trend analysis index is as follows: Obtain the control data set and number it in the order of the data sequence; perform data processing on the obtained data to be used and the control data set to obtain the change trend analysis index; the limit expression of the change trend analysis index is specifically as follows:

[0058]

[0059] In the formula, d represents the number of the index to be analyzed, d = 1, 2,..., D, D represents the total number of indexes to be analyzed, q represents the sequence number of the data in the control data set, q = 1, 2,..., Q, Q represents the total number of data in the control data set, ρ d represents the data to be used for the d-th index to be analyzed, δ q represents the q-th data in the control data set, δ q+1 represents the (q + 1)-th data in the control data set, ω dThe change trend analysis index for the d-th index to be analyzed is denoted as, and e represents the natural constant.

[0060] In this embodiment, taking the wastewater discharge volume as an example of the index to be analyzed, the algorithm comprehensively analyzes the data to be used for the wastewater discharge volume and the control data set of the wastewater discharge volume to obtain the change trend analysis index. In the formula, represents the mean value of the control data set, represents the displacement mean value of the control data set. As the difference between the data to be used for the wastewater discharge volume and the mean value of the control data set increases, the change trend analysis index becomes smaller, indicating a greater possibility that the data to be used for the wastewater discharge volume does not conform to the change trend of the historical data. Among them, the control data set represents the historical data set of the index to be analyzed in the historical database of the enterprise; similarly, as the difference between the displacement amount of the data to be used for the wastewater discharge volume and the displacement mean value of the data in the last control data set increases, it indicates that the degree to which the displacement amount of the data to be used deviates from the average data change level of the last control data set for the wastewater discharge volume index to be analyzed is higher, then the possibility that the data to be used for the wastewater discharge volume is abnormal is greater, and the mean data types of the data to be used and the control data set are the same but have no influence on each other, and at the same time, there is no influence between the displacement amount of the data to be used and the displacement mean value; therefore, by analyzing and comparing the data to be used for the wastewater discharge volume index to be analyzed with the control data set, it helps to improve the usability of the data to be used, and at the same time improves the reliability of the determined enterprise ESG index.

[0061] Furthermore, the specific process of determining whether to perform data verification is as follows: If the change trend analysis index is greater than the change trend judgment threshold, no data verification is performed; if the change trend analysis index is not greater than the change trend judgment threshold, a preset staff is prompted to perform data verification; the change trend judgment threshold is used to judge the degree of deviation between the data to be used and the change trend of the control data set.

[0062] In this embodiment, automatically determining whether data verification is required reduces unnecessary manual intervention, optimizes the work process and improves efficiency, and through the preset change trend judgment threshold, intelligent judgment can be made according to the data fluctuation situation, realizing the automated management of the verification process and making the entire data processing process more intelligent; at the same time, the use of the change trend analysis index ensures that a verification prompt is triggered only when the data changes abnormally, focusing on key data to ensure that the data to be used is more accurately verified.

[0063] Specifically, the change trend judgment threshold is obtained from a preset database. In a specific embodiment, a data set of the index to be analyzed with no abnormal data changes is obtained from historical data, and the corresponding change trend analysis index is obtained based on the data set, and the mean operation is performed on the obtained change trend analysis indexes to obtain the corresponding average value, then this average value is denoted as the change trend judgment threshold.

[0064] Furthermore, the specific process of data verification is as follows: Step 1, send the first determination button to the preset staff. If the preset staff selects "Yes", it is determined that the data to be used is for the evaluation of the indicators to be analyzed. Otherwise, execute Step 2. The first determination button is used to determine whether the current abnormal change trend is real; Step 2, send the dataset to be analyzed and the corresponding data sources to the preset staff and the data verification prompt instruction; Step 3, send the second determination button to the preset staff. If the preset staff selects "Correct", send the data source self-check prompt instruction. Otherwise, extract data from each data source again and update the dataset to be analyzed. The second determination button is used to determine whether there is an error in data extraction; Step 4, re-obtain the data conflict evaluation index according to the dataset to be analyzed changed by the preset staff until data verification is not performed.

[0065] In this embodiment, data is extracted from each data source through programming languages (such as Python, Java, SQL) or data processing tools (such as the ETL tool Talend, Apache NiFi) for data extraction, and then connected to the API or database, send requests, download data files or perform web scraping. Finally, use stream processing tools (such as Apache Kafka, Flink) to process real-time data to update the dataset to be analyzed; determining whether the current abnormal change trend is real means checking whether the enterprise has corresponding development events that cause such abnormal changes; through the two determination buttons and the verification of the data sources, the reliability and accuracy of the data are guaranteed, and through the manual confirmation of the data by the preset staff, the misjudgment that may be generated by automatic detection can be reduced, ensuring that the trend reflected by the data is the actual situation; and the automated data extraction, verification, and update process reduces the frequency of manual intervention, while ensuring that the data always remains up-to-date, improving the efficiency of data analysis.

[0066] Further, send a self-check prompt instruction for the data source, and then it also includes obtaining the feedback result of the self-check prompt instruction for the data source; send a self-check feedback instruction for the data source to the preset staff. If the preset staff feedback that the data source data is incorrect, re-obtain the data source and update the dataset to be analyzed; if the preset staff feedback that the data source data is correct, obtain an updated data conflict assessment index based on the dataset to be analyzed, and obtain a conflict assessment similarity coefficient based on the absolute value of the difference between the updated data conflict assessment index and the original data conflict assessment index and the operation result of the ratio of the updated data conflict assessment index to the original data conflict assessment index; compare the conflict assessment similarity coefficient with the conflict similarity assessment threshold and determine whether there is an error in the enterprise ESG index determination method; the original data conflict assessment index represents the data conflict assessment index obtained according to the dataset to be analyzed without change; the updated data conflict assessment index represents the data conflict assessment index obtained according to the updated dataset to be analyzed.

[0067] The limiting expression of the conflict assessment similarity coefficient is specifically as follows:

[0068]

[0069] In the formula, d represents the number of the index to be analyzed, d = 1, 2,..., D, and D represents the total number of the indexes to be analyzed. represents the original data conflict assessment index. represents the updated data conflict assessment index, α represents the conflict assessment similarity coefficient, and e represents the natural constant.

[0070] In this embodiment, taking the wastewater discharge volume as an example of the index to be analyzed, the algorithm comprehensively analyzes the original data conflict assessment index and the updated data conflict assessment index of the wastewater discharge volume to obtain the data conflict assessment index. In the formula, when the values of the original data conflict assessment index and the updated data conflict assessment index of the wastewater discharge volume are closer, the value of the conflict assessment similarity index is smaller, indicating that the possibility of an error in the enterprise ESG index determination method is smaller. If the value of the conflict assessment similarity index is larger, the possibility of an error in the enterprise ESG index determination method is larger. Specifically, as Figure 2 shown is the change schematic diagram of the conflict assessment similarity coefficient provided by the embodiment of the present application. From Figure 2It can be intuitively seen that when the values of the original data conflict assessment index and the updated data conflict assessment index of the wastewater discharge are closer, the image shows a downward trend, and the change trend is not affected by the magnitudes of the values of the original data conflict assessment index and the updated data conflict assessment index respectively. The change trend is first downward and then upward. Moreover, regardless of whether the updated data conflict assessment index of the wastewater discharge increases or decreases, the domain of the original data conflict assessment index of the wastewater discharge is positive real numbers, and the value range of the corresponding conflict assessment similarity index is always [0, 1]. At the same time, as the updated data conflict assessment index of the wastewater discharge increases, although the change trend remains the same, the change speed gradually decreases, and the change curve gradually flattens, indicating that when the value of the updated data conflict assessment index is smaller, the change of the conflict assessment similarity coefficient is more rapid, and the change curve is steeper; Through the analysis and comparison of the original data conflict assessment index and the updated data conflict assessment index of the wastewater discharge, it helps to analyze whether there are still deficiencies in the enterprise ESG index determination method and promptly investigate and solve the problems that occur, reduce errors in subsequent data analysis, and improve the efficiency of enterprise ESG index determination.

[0071] Furthermore, the specific process of determining whether there is an error in the enterprise ESG index determination method is as follows: If the conflict assessment similarity coefficient is not greater than the conflict similarity assessment threshold, it indicates that the enterprise ESG index determination method is correct; If the conflict assessment similarity coefficient is greater than the conflict similarity assessment threshold, it indicates that there is an error in the enterprise ESG index determination method, then an error warning is issued and the programmer is requested to debug the program; The conflict similarity assessment threshold is used to determine whether there is a program error in the enterprise ESG index determination method.

[0072] In this embodiment, after the programmer understands the problem background, uses the debugging function of the IDE (such as PyCharm, VS Code), sets breakpoints, executes the code step by step, and adds debugging statements such as print() or logging in the script to output the intermediate results, then conducts step-by-step troubleshooting such as logical errors, data integrity, and algorithm correctness, and at the same time checks exception handling (division by zero error, null value handling, and data type error), and finally conducts unit tests and data verification to ensure that each functional module of the algorithm is correct; By analyzing the conflict assessment similarity coefficient, potential problems in the ESG index calculation process can be quickly identified, avoiding the continuous impact of incorrect data on subsequent analysis, and improving the accuracy of data processing; And when the conflict assessment similarity coefficient exceeds the threshold, an error warning will be automatically issued and the programmer will be requested to debug, ensuring that the problem can be promptly processed and reducing the long-term impact of errors on the entire ESG analysis process. At the same time, the automation of the error detection and debugging process improves the robustness of the method, and continuous monitoring and rapid response also ensure the stability and reliability of the ESG index determination method.

[0073] Specifically, the conflict similarity evaluation threshold is obtained from a preset database. In a specific embodiment, a data set corresponding to an index to be analyzed with a program error in the enterprise ESG index determination method is obtained from historical data, and a corresponding conflict similarity evaluation coefficient is obtained based on the data set. Then, the average value of the obtained conflict similarity evaluation coefficients is calculated to obtain the corresponding average value, and this average value is recorded as the conflict similarity evaluation threshold.

[0074] As Figure 3 shown, it is a schematic structural diagram of an enterprise ESG index determination system based on data fusion provided by an embodiment of the present application. The enterprise ESG index determination system based on data fusion provided by an embodiment of the present application includes: a data acquisition module, a data conflict evaluation module, and a data change analysis module. Among them, the data acquisition module is used to obtain a data set to be analyzed for an index to be analyzed through each data source, and obtain conflict evaluation data based on the data set to be analyzed. The data conflict evaluation module is used to obtain a data conflict evaluation index based on the conflict evaluation data, compare the data conflict evaluation index with the data conflict evaluation threshold obtained from the preset database, and obtain the corresponding data to be used. The data conflict evaluation index is used to quantify the data conflict degree of the data set to be analyzed for the index to be analyzed. The data change analysis module is used to obtain a control data set for the index to be analyzed from the preset enterprise historical database, obtain a change trend analysis index based on the data to be used and the control data set, compare the change trend analysis index with the change trend judgment threshold, and determine whether to perform data verification. The change trend analysis index is used to quantify the degree of conformity between the data to be used and the historical change trend. Data verification means re-verifying the data in the data set to be analyzed.

[0075] In this embodiment, by quantifying data conflicts and change trends, the system can more comprehensively detect data quality problems and ensure data consistency between different data sources. At the same time, through the cooperation of multiple modules, the system provides a full-process automated processing mechanism from data acquisition to data evaluation and verification, significantly improving the accuracy, consistency, and timeliness of data, thereby ensuring the scientificity and reliability of enterprise ESG index evaluation, reducing human intervention, and optimizing process efficiency.

[0076] In summary, in the embodiments of the present application, conflict evaluation data is obtained from the dataset to be analyzed of the index to be analyzed, and then the data conflict evaluation index is obtained based on the conflict evaluation data. The data conflict evaluation index is compared with the data conflict evaluation threshold to obtain the corresponding data to be used. Then, the control dataset of the index to be analyzed is obtained, and the change trend analysis index is obtained based on the data to be used and the control dataset. Then, the change trend analysis index is compared with the change trend judgment threshold to determine whether to perform data verification, so as to more reasonably confirm the data for evaluating the ESG index, and further improve the accuracy of the enterprise ESG index determined based on data fusion, effectively solving the problem of data conflict in determining the enterprise ESG index with multiple data sources in the prior art.

[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0078] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks

[0080] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.

[0081] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0082] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for determining an enterprise ESG index based on data fusion, characterized in that, The following steps are involved: Obtaining a data set to be analyzed of the indicator to be analyzed through various data sources, and obtaining conflict assessment data based on the data set to be analyzed; Based on the conflict assessment data, a data conflict assessment index is obtained, the data conflict assessment index is compared with a data conflict assessment threshold obtained from a preset database, and corresponding stand-by data is obtained, wherein the data conflict assessment index is used to quantify the degree of data conflict of the data set to be analyzed of the indicator to be analyzed; Obtaining a reference data set of the indicator to be analyzed from a preset enterprise historical database, obtaining a change trend analysis index based on the standby data and the reference data set, comparing the change trend analysis index with a change trend judgment threshold and determining whether to perform data verification, wherein the change trend analysis index is used to quantify the degree of conformity between the standby data and the historical change trend; The specific process of obtaining the standby data is as follows: Get the data conflict assessment threshold from the preset database and compare it with the data conflict assessment index: If the data conflict evaluation index is less than the data conflict evaluation threshold, the data in the data set to be analyzed are averaged to obtain the standby data; If the data conflict evaluation index is not less than the data conflict evaluation threshold, the data to be analyzed is combined with the evaluation weights of each data source to obtain the data to be used; The evaluation weights include a first-level evaluation weight mapping set and a second-level evaluation weight mapping set. The first-level evaluation weight mapping set represents a set of weights assigned to each data source according to its type and authority, and the second-level evaluation weight mapping set represents a set of weights corresponding to the industry status corresponding to the same type of data source.

2. The method for determining an enterprise ESG index based on data fusion according to claim 1, wherein: The specific process of obtaining the conflict assessment data is as follows: Processing the data set to be analyzed to obtain conflict assessment data, wherein the data set to be analyzed represents a data set obtained from various data sources for evaluating corresponding indicators to be analyzed, and the indicators to be analyzed are used to determine the ESG index of the enterprise; The conflict assessment data include variance, standard deviation, range, interquartile range, mean absolute deviation and median absolute deviation; Obtaining the conflict assessment data also includes extracting and monitoring the data source: The specific steps of data source extraction monitoring are as follows: A1, obtaining extraction data through a network detection tool, wherein the extraction data includes extraction time, data packet size, extraction transmission rate, data loading delay, and connection retry times; A2, obtaining reference extraction data from a preset database, wherein the reference extraction data includes average extraction time, average data packet size, minimum network transmission rate, maximum data loading delay, and maximum number of connection retries; A3, numbering the extracted data according to the data type, and obtaining a data extraction analysis index based on the extracted data and the reference extracted data of the corresponding data type if the extracted data meets the extraction condition, otherwise the data extraction analysis index is recorded as 1, and the extraction condition indicates that the extraction transmission rate is greater than the minimum network transmission rate and the data loading delay is less than the maximum data loading delay and the number of connection retries is less than the maximum number of connection retries; A4, compare the data extraction analysis index with the data extraction threshold: If the data extraction and analysis index is greater than the data extraction threshold, extract the data from the next data source; otherwise, perform secondary extraction on the data of the corresponding data source until the data from the next data source is extracted.

3. The method for determining an enterprise ESG index based on data fusion according to claim 2, wherein: The specific process of obtaining the data conflict evaluation index based on the conflict evaluation data is as follows: Number the metrics to be analyzed and obtain conflict evaluation data; Perform logarithmic processing on the variance and standard deviation of the data set to be analyzed to obtain the evaluation value of the dispersion degree of the data set to be analyzed; Perform hyperbolic tangent processing on the range of the data set to be analyzed to obtain the evaluation value of the difference degree of the data set to be analyzed; Perform hyperbolic sine processing on the interquartile range, mean absolute deviation, and median absolute deviation of the data set to be analyzed to obtain the evaluation value of the aggregation degree of the data set to be analyzed; Obtain the data conflict evaluation index of the data set to be analyzed by performing square root processing and hyperbolic secant processing on the sum of the evaluation value of the dispersion degree, the evaluation value of the difference degree, and the evaluation value of the aggregation degree.

4. The method for determining an enterprise ESG index based on data fusion according to claim 1, wherein: The specific method of obtaining the change trend analysis index is as follows: Obtain the control data set and number it in the order of the data sequence; Perform data processing on the obtained data to be used and the control data set to obtain the change trend analysis index; The limit expression of the change trend analysis index is specifically as follows: ; Where d represents the number of the index to be analyzed, , represents the total number of the indexes to be analyzed, q represents the sequence number of the data in the control dataset, , represents the total number of the data in the control dataset, represents the data to be used for the d-th index to be analyzed, represents the q-th data in the control dataset, represents the -th data in the control dataset, represents the change trend analysis index of the d-th index to be analyzed, and e represents the natural constant.

5. The method for determining an enterprise ESG index based on data fusion according to claim 4, characterized in that: The specific process of determining whether to perform data verification is as follows: If the change trend analysis index is greater than the change trend judgment threshold, do not perform data verification; If the change trend analysis index is not greater than the change trend judgment threshold, prompt the preset staff to perform data verification; The change trend judgment threshold is used to judge the deviation degree of the change trend between the data to be used and the control data set.

6. The method for determining an enterprise ESG index based on data fusion according to claim 5, characterized in that: The specific process of the data verification is as follows: Step 1, send the first judgment button to the preset staff. If the preset staff selects "yes", determine that the data to be used is for the evaluation of the metrics to be analyzed; otherwise, execute Step 2; Step 2, send the data set to be analyzed and the corresponding data sources to the preset staff and the data verification prompt instruction; Step 3, send the second judgment button to the preset staff. If the preset staff selects "correct", send the data source self-check prompt instruction; otherwise, extract data from each data source again and update the data set to be analyzed; Step 4, re-obtain the data conflict evaluation index according to the data set to be analyzed changed by the preset staff until data verification is not performed.

7. The method for determining an enterprise ESG index based on data fusion according to claim 6, wherein: After sending the data source self-check prompt instruction, it also includes obtaining the feedback result of the data source self-check prompt instruction; Send the data source self-check feedback instruction to the preset staff. If the preset staff feedbacks that the data of the data source is incorrect, re-obtain the data source and update the data set to be analyzed; If the preset staff feedbacks that the data of the data source is correct, obtain the updated data conflict evaluation index according to the data set to be analyzed, and obtain the conflict evaluation similarity coefficient based on the absolute value of the difference between the updated data conflict evaluation index and the original data conflict evaluation index and the operation result of the ratio of the updated data conflict evaluation index to the original data conflict evaluation index; Compare the conflict evaluation similarity coefficient with the conflict similarity evaluation threshold and judge whether there is an error in the enterprise ESG index determination method; The original data conflict evaluation index represents the data conflict evaluation index obtained based on the dataset to be analyzed without changes. The updated data conflict evaluation index represents the data conflict evaluation index obtained based on the updated dataset to be analyzed.

8. The method for determining an enterprise ESG index based on data fusion according to claim 7, wherein: The specific process for determining whether there is an error in the enterprise ESG index determination method is as follows: If the conflict evaluation similarity coefficient is not greater than the conflict similarity evaluation threshold, it indicates that the enterprise ESG index determination method is correct. If the conflict evaluation similarity coefficient is greater than the conflict similarity evaluation threshold, an error warning is issued and the programmer is requested to debug the program. The conflict similarity evaluation threshold is used to determine whether there is a program error in the enterprise ESG index determination method.

9. An enterprise ESG index determination system based on data fusion, characterized in that: It includes a data acquisition module, a data conflict evaluation module, and a data change analysis module. Among them, the data acquisition module is used to obtain the dataset to be analyzed for the index to be analyzed through each data source, and obtain conflict evaluation data based on the dataset to be analyzed. The data conflict evaluation module is used to obtain the data conflict evaluation index based on the conflict evaluation data, compare the data conflict evaluation index with the data conflict evaluation threshold obtained from the preset database, and obtain the corresponding data to be used. The data conflict evaluation index is used to quantify the data conflict degree of the dataset to be analyzed for the index to be analyzed. The data change analysis module is used to obtain the control dataset for the index to be analyzed from the preset enterprise historical database, obtain the change trend analysis index based on the data to be used and the control dataset, compare the change trend analysis index with the change trend judgment threshold, and determine whether to perform data verification. The change trend analysis index is used to quantify the degree of conformity between the data to be used and the historical change trend. The specific process for obtaining the data to be used is as follows: Obtain the data conflict evaluation threshold from the preset database and compare it with the data conflict evaluation index: If the data conflict evaluation index is less than the data conflict evaluation threshold, perform a mean operation on the data in the dataset to be analyzed to obtain the data to be used. If the data conflict evaluation index is not less than the data conflict evaluation threshold, combine the data to be analyzed with the evaluation weights of each data source to obtain the data to be used. The evaluation weights include a first-level evaluation weight mapping set and a second-level evaluation weight mapping set. The first-level evaluation weight mapping set represents the set of weights assigned to each data source according to the type and authority of each data source. The second-level evaluation weight mapping set represents the set of weights corresponding to the industry status of the same type of data source.

Citation Information

Patent Citations

  • Corporate ESG index determination method and related products

    CN113240272B

  • Enterprise ESG index determination method and related products based on data completion

    CN113313362B

  • Tunnel construction comprehensive risk evaluation method and device based on multi-source data fusion

    CN117035418A

  • System and method that rank businesses in environmental, social and governance (ESG)

    US20220343433A1