Enterprise digital transformation analysis decision method and system based on information error driving

Through an information error-driven method, real-time collection and analysis of enterprise data, correct data errors and build a decision prediction model, the data inconsistency and accuracy problems in enterprise digital transformation are solved, and the scientificity and reliability of decisions are improved.

CN120047014APending Publication Date: 2025-05-27BEIJING UNION UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510521281.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Enterprises face data inconsistency and accuracy problems during the digital transformation process. Traditional data analysis methods are difficult to identify and correct data deviations in a timely manner, which affects the accuracy and reliability of decisions.

Method used

Using an information error-driven method, data in the enterprise-related business system is collected and standardized in real time, data deviation coefficients are calculated through consistency analysis and logical analysis, information errors are quantified, and data errors are corrected based on this, and an enterprise decision prediction model is constructed for risk assessment.

Benefits of technology

Real-time, in-depth analysis and accurate correction of enterprise data is achieved, ensuring the accuracy and consistency of data in multiple dimensions, improving the scientificity and reliability of the decision-making process, and optimizing enterprise resource allocation and strategic implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047014A_ABST
    Figure CN120047014A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital data analysis and information monitoring, and particularly discloses an enterprise digital transformation analysis decision method and system based on information error driving, data in a related business system of an enterprise is collected in real time, the collected data is subjected to standardization processing, and through consistency analysis and logicality analysis, the enterprise digital transformation analysis decision method and system based on information error driving are obtained. The accuracy of the data is evaluated, information errors are calculated and quantified, and a consistency deviation coefficient and a logicality deviation coefficient are generated; based on the deviation coefficient, correcting and optimizing errors in the data so as to improve the accuracy of the data; and constructing an enterprise decision prediction model based on the optimized data, training a decision tree model and calculating a risk coefficient by taking the consistency deviation coefficient, the logical deviation coefficient and the correction coefficient characteristics in the corrected data set and the original data set as input, and determining the enterprise decision prediction by comparing the predicted risk coefficient with a preset threshold value. The enterprise transformation is divided into two modes of transformation proceeding and transformation pause.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital data analysis and information monitoring, and in particular to an enterprise digital transformation analysis and decision-making method and system driven by information errors. Background Art

[0002] As digital transformation becomes an important strategy for enterprise development, more and more enterprises are beginning to rely on data-driven decision-making methods in the face of market competition and a changing environment. In the process of digital transformation, enterprises need to rely on a large amount of business operation data, financial data, market data, and equipment operation data to support decision-making. However, with the sharp increase in data volume and the diversification of sources, enterprises face challenges in data accuracy, data consistency, and data quality in the decision-making process. Traditional data analysis methods often fail to identify deviations in data in a timely manner, resulting in insufficient decision support, which affects the effectiveness and efficiency of transformation. Therefore, how to accurately analyze and correct data through scientific methods to improve the accuracy of transformation decisions has become a key issue that needs to be urgently addressed in the digital transformation of enterprises.

[0003] The prior art has the following deficiencies: Existing technologies use traditional data analysis methods, such as rule-based data cleaning and basic statistical analysis, to deal with data problems in enterprise transformation, but these methods show obvious deficiencies when faced with complex data structures and multi-dimensional data sources. First, existing analysis methods often cannot effectively deal with consistency and logical deviations in data, especially in the case of multi-source heterogeneous data, where data consistency cannot be guaranteed; second, existing technologies have limited data correction capabilities and cannot timely discover and correct potential errors in data, resulting in low accuracy and reliability of transformation decisions; finally, existing technologies often lack flexible dynamic adjustment mechanisms when conducting enterprise transformation risk assessments, cannot respond to market changes in a timely manner, and are difficult to provide accurate decision support. Therefore, existing technologies have failed to fully solve how to comprehensively analyze and correct the inconsistency and inaccuracy of data in the process of enterprise transformation, nor can they provide enterprises with more accurate and dynamic risk assessments, affecting the ability of enterprises to make scientific decisions during the transformation process. Summary of the invention

[0004] The purpose of the present invention is to provide an enterprise digital transformation analysis and decision-making method and system driven by information errors to solve the problems in the above background.

[0005] The purpose of the present invention can be achieved through the following technical solutions: The enterprise digital transformation analysis and decision-making method based on information error driving includes the following steps: S1: Collect and standardize the data in the enterprise-related business systems in real time, process the missing values, duplicate data, and outliers in the data to ensure the integrity of the data structure; The data in the enterprise-related business systems includes: business operation data, financial data, market data, and equipment operation data; S2: Conduct consistency analysis and logical analysis on the collected data, evaluate the accuracy of the data, calculate the data accuracy deviation coefficient, quantify the information error, and evaluate the accuracy of the data; S3: Based on the accuracy evaluation results, correct and optimize the errors in the data; S4: Based on the optimized data, construct an enterprise decision prediction model to predict the risks of enterprise transformation. According to the risk prediction results of enterprise transformation, divide the enterprise transformation into two modes: proceed with transformation and suspend transformation; S5: Based on the enterprise decision to proceed with transformation, re-collect and evaluate the data in the enterprise-related business systems to further judge the accuracy of the enterprise decision to proceed with transformation.

[0006] As a further solution of the present invention: The conduct of consistency analysis and logical analysis on the collected data, evaluation of the accuracy of the data, calculation of the data accuracy deviation coefficient, quantification of the information error, and evaluation of the accuracy of the data specifically include: Collect the data in the enterprise-related business systems in real time, conduct consistency analysis on the data, and calculate the consistency deviation coefficient according to the analysis results for evaluating the consistency of the data; Collect the data in the enterprise-related business systems in real time, conduct logical accuracy analysis on the data, and calculate the logical deviation coefficient according to the analysis results for evaluating the logical accuracy of the data; Obtain the logical deviation coefficient and consistency deviation coefficient of the data, perform normalization calculation processing on the logical deviation coefficient and consistency deviation coefficient, and comprehensively calculate the accuracy deviation coefficient; Judge whether the accuracy deviation coefficient of the data is greater than or equal to the preset threshold. If so, it indicates that the corresponding data is inaccurate. If not, it indicates that the corresponding data is accurate.

[0007] As a further solution of the present invention: The process of obtaining the consistency deviation coefficient is: Collect the data in the enterprise-related business systems in real time, perform standardization processing on the data collected in real time in the enterprise-related business systems, so that features with different dimensions can be compared consistently; Use the Euclidean distance to calculate the distance between each pair of data points, and construct a distance matrix according to the distances between all data pairs ; Among them, the calculation expression for using the Euclidean distance to calculate the distance between each pair of data points is: ; In the formula, represents one of the data points, represents the other data point, represents the number of rows and columns of the distance matrix, represents one of the data after standardization, represents the other data after standardization, represents the th data point, represents the total number of data points, the distance matrix , represents the th type of data at the th data point after standardization, represents the th type of data at the th data point after standardization; Based on the distance matrix, the data is clustered using the hierarchical clustering algorithm. The minimum distance method is adopted to calculate the distance between clusters. The calculation expression is: ; In the formula, and are two clusters in the clustering process; For each cluster in the clustering result, the average value of the distances between all data points in the cluster and the cluster center is calculated as the consistency deviation coefficient. The consistency deviation coefficient of each cluster is calculated. The calculation expression is: ; Among them, represents the th cluster, represents the center of the th cluster, is the number of data points in cluster , is the th data point and the Euclidean distance from the cluster center ; Among them, The calculation expression of ; According to the deviation coefficient of each cluster, the consistency deviation coefficient of the entire data set is calculated. The calculation expression is: ; In the formula, represents the consistency deviation coefficient of the entire data set, represents the total number of clusters.

[0008] As a further solution of the present invention: the process of obtaining the logical deviation coefficient is as follows: Standardize the data in the enterprise-related business system collected in real time; Fit the probability distribution of the standardized data through maximum likelihood estimation and expectation maximization to obtain the parameters ; wherein, represents the weight of the th Gaussian distribution, represents the number of Gaussian distributions, represents the mean vector of the th Gaussian distribution, represents the covariance matrix of the th Gaussian distribution; Calculate the posterior probability that each data point belongs to the th Gaussian distribution. The calculation expression is: ; In the formula, c represents the number of data points, represents the posterior probability that the cth data point belongs to the ath Gaussian distribution, represents the total number of Gaussian distributions; Update the parameters according to the result of the posterior probability; According to the updated new parameters , calculate its log-likelihood value. The calculation expression is: ; In the formula, represents the log-likelihood value of the th data point; Calculate the logical deviation coefficient of each data point according to the log-likelihood value of each data point. The calculation expression is: ; In the formula, represents the logical deviation coefficient of the th data point, represents the maximum log-likelihood value of all data points; Calculate the average logical deviation coefficient of the entire data set. The calculation expression is: ; In the formula, represents the total number of data points, represents the average logical deviation coefficient of the entire data set.

[0009] As a further solution of the present invention: based on the accuracy evaluation result, correcting and optimizing the errors in the data, specifically including: Obtain inaccurate data and calculate a data correction coefficient for correcting the accuracy of the data; The process of obtaining the data correction coefficient is as follows: Obtain the logical deviation coefficient and consistency deviation coefficient of the data set, construct a prior distribution of the correction coefficient, and satisfy the normal distribution ; Among them, represents the expected value of the prior correction coefficient, represents the variance of the prior correction coefficient; The calculation expression of the prior distribution is: ; In the formula, represents the prior distribution of the correction coefficient; Construct a likelihood function, and the calculation expression is: ; In the formula, represents the likelihood function of the logical deviation coefficient and consistency deviation coefficient; represents the prior distribution of the logical deviation coefficient calculated through the calculation expression of the prior distribution, represents the prior distribution of the consistency deviation coefficient calculated through the calculation expression of the prior distribution; Calculate the posterior distribution of the correction coefficient, and the calculation expression is: ; In the formula, represents the posterior distribution of the correction coefficient; represents the evidence term, which is calculated by integration, and the calculation expression is: ; Calculate the correction coefficient through the expected value of the posterior distribution, and the calculation expression is: ; In the formula, represents the correction coefficient.

[0010] As a further solution of the present invention: based on the optimized data, construct an enterprise decision-making prediction model, specifically including: Perform a product calculation on the correction coefficient and the original data set to obtain a corrected accurate data set, and extract the consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature from the corrected accurate data set and the original accurate data set; Take the consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature as input features and input them into the enterprise decision prediction model; Use decision trees to train the enterprise decision prediction model; The input features of each decision tree are the consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature, and the output is the predicted risk coefficient; The structure of each tree is based on historical data, which is used as training data, and decisions are made through feature selection and splitting rules; Train each tree by randomly selecting subsets of data and subsets of features to increase diversity; The splitting nodes of each tree select the method that can maximize the information gain; The final risk coefficient is obtained by averaging the outputs of all trees; The calculation expression for the prediction result of the risk coefficient is: ; In the formula, represents the decision tree in the random forest, represents the total number of decision trees in the random forest, represents the th decision tree's prediction result, represents the prediction result of the risk coefficient.

[0011] As a further solution of the present invention: for the risk prediction of enterprise transformation, according to the risk prediction result of enterprise transformation, the enterprise transformation is divided into two modes: proceeding with transformation and suspending transformation, specifically including: According to the enterprise decision prediction model, output the risk coefficient, compare the risk coefficient with a preset threshold, and determine whether the risk coefficient is greater than or equal to the preset threshold. If not, it is recorded as proceeding with transformation; if so, it is recorded as suspending transformation.

[0012] As a further solution of the present invention: based on the enterprise decision of proceeding with transformation, re-collect and evaluate the data in the enterprise-related business system to further judge the accuracy of the enterprise decision of proceeding with transformation, specifically including: By preprocessing the data collected in the enterprise business system, calculate the consistency deviation coefficient, logical deviation coefficient, and correction coefficient, and then construct a comprehensive risk assessment model. According to the constructed comprehensive risk assessment model, output the risk coefficient, calculate the difference between the new risk coefficient and the historical risk coefficient to obtain the risk coefficient deviation value, and compare the risk coefficient deviation value with a preset threshold; Judge whether the risk coefficient deviation value is greater than or equal to the preset threshold. If so, the enterprise decision of proceeding with transformation is inaccurate; if not, the enterprise decision of proceeding with transformation is accurate.

[0013] Enterprise digital transformation analysis and decision-making system driven by information error, comprising: Data acquisition module, which acquires and standardizes the data in the enterprise-related business systems in real time, processes the missing values, duplicate data and outliers in the data, and ensures the integrity of the data structure; Data accuracy evaluation module, which conducts consistency analysis and logical analysis on the acquired data, evaluates the accuracy of the data, calculates the data accuracy deviation coefficient, quantifies the information error, and evaluates the accuracy of the data; Data correction module, which corrects and optimizes the errors in the data based on the accuracy evaluation results; Risk prediction and decision-making module, which constructs an enterprise decision-making prediction model based on the optimized data, predicts the risks of enterprise transformation, and divides the enterprise transformation into two modes: proceeding with transformation and suspending transformation according to the risk prediction results of enterprise transformation; Decision accuracy evaluation module, which re-acquires and evaluates the data in the enterprise-related business systems based on the enterprise decision of proceeding with transformation, and further judges the accuracy of the enterprise decision of proceeding with transformation.

[0014] Advantages of the present invention: (1) By integrating multi-source data acquisition, standardized processing and deviation coefficient calculation, the present invention solves the problems of data inconsistency and accuracy faced by enterprises in digital transformation. Using the consistency deviation coefficient, logical deviation coefficient and correction coefficient, the present invention realizes in-depth analysis and precise correction of real-time data in enterprise business systems, ensuring the accuracy and consistency of data in multiple dimensions. Through consistency analysis, the deviations between different data sources are identified, and further through logical analysis, it is ensured that the data conforms to business rules and preset transformation standards. The present invention can not only timely correct potential data errors, avoid decision-making errors caused by inaccurate data, but also significantly improve the scientificity and reliability of the decision-making process, thereby optimizing enterprise resource allocation, enhancing the effect of strategic implementation, ensuring that enterprises can efficiently and accurately achieve transformation goals in a complex transformation environment, and ultimately enhancing the market competitiveness and sustainable development ability of enterprises.

[0015] (2) By constructing an enterprise decision-making prediction model based on optimized data and combining advanced machine learning algorithms such as random forest, the present invention solves the complex risk assessment problems faced by enterprises in the process of digital transformation. By accurately predicting various risk factors that may occur during the enterprise transformation process, the present invention provides an effective early warning mechanism for enterprises, helps enterprises identify and avoid potential risks in a timely manner, and ensures the scientific nature of the decision-making process and the accuracy of implementation. The risk assessment model not only integrates historical data and real-time data, but also can dynamically adjust the risk prediction results to adapt to the rapidly changing market environment. The present invention reduces the decision-making uncertainty that enterprises may encounter during the transformation process, provides a more reliable decision-making basis for enterprises through a real-time risk feedback mechanism, helps enterprises formulate more accurate strategies and tactics in a complex and ever-changing market environment, and ultimately enhances the competitiveness of enterprises and the probability of successful transformation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present invention will be further described below with reference to the accompanying drawings.

[0017] Figure 1 is a specific step flow block diagram of the enterprise digital transformation analysis and decision-making method based on information error driving of the present invention; Figure 2 is a flow block diagram of the enterprise digital transformation analysis and decision-making system based on information error driving in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Please refer to Figure 1 As shown, the present invention is an enterprise digital transformation analysis and decision-making method based on information error driving, including the following steps: S1: Real-time collect and standardize the data in the enterprise-related business systems, process the missing values, duplicate data and outliers in the data, and ensure the integrity of the data structure; The data in the enterprise-related business systems includes: business operation data, financial data, market data and equipment operation data; S2: Conduct consistency analysis and logical analysis on the collected data, evaluate the accuracy of the data, calculate the data accuracy deviation coefficient, quantify the information error, and evaluate the accuracy of the data; S3: Based on the accuracy evaluation results, correct and optimize the errors in the data; S4: Based on the optimized data, construct an enterprise decision-making prediction model to predict the risks of enterprise transformation. According to the risk prediction results of enterprise transformation, divide enterprise transformation into two modes: proceeding with transformation and suspending transformation; S5: Based on the enterprise decision-making of proceeding with transformation, re-collect and evaluate the data in the enterprise-related business systems to further judge the accuracy of the enterprise decision-making of proceeding with transformation.

[0020] In S1: Real-time collect and standardize the data in the enterprise-related business systems, process the missing values, duplicate data and outliers in the data to ensure the integrity of the data structure; The data in the enterprise-related business systems includes: business operation data, financial data, market data and equipment operation data, specifically including: In the data collection process, first obtain multi-source data in real time from the enterprise-related business systems, including business operation data, financial data, market data and equipment operation data; It should be noted that: the content involved in each data source is different. For example, business operation data includes order processing, inventory management and production scheduling information, financial data involves the enterprise's income, expenditure, assets and liabilities, market data covers customer behavior, sales channels and advertising effects, and equipment operation data includes machine performance, failure rate and maintenance records; Through the API interface or data collection tool, capture the data in real time into the data platform to ensure the timeliness and integrity of the information; Subsequently, perform standardization processing on the collected data to ensure that the formats of each data source are consistent, and clean the missing values, duplicate data and outliers in the data.

[0021] In S2, perform consistency analysis and logical analysis on the collected data, evaluate the accuracy of the data, calculate the accuracy deviation coefficient of the data, quantify the information error, and evaluate the accuracy of the data, specifically including: Real-time collect the data in the enterprise-related business systems, perform consistency analysis on the data, and according to the analysis results, calculate the consistency deviation coefficient for evaluating the consistency of the data; Real-time collect the data in the enterprise-related business systems, perform logical accuracy analysis on the data, and according to the analysis results, calculate the logical deviation coefficient for evaluating the logical accuracy of the data; Obtain the logical deviation coefficient and consistency deviation coefficient of the data, perform normalization calculation processing on the logical deviation coefficient and consistency deviation coefficient, and comprehensively calculate the accuracy deviation coefficient; Judge whether the accuracy deviation coefficient of the data is greater than or equal to the preset threshold. If so, it means that the corresponding data is inaccurate. If not, it means that the corresponding data is accurate; The process of obtaining the consistency deviation coefficient is: Collect data from the enterprise's relevant business systems in real time, and perform standardization processing on the data collected in real time from the enterprise's relevant business systems, so that features with different dimensions can be compared consistently; Use the Euclidean distance to calculate the distance between each pair of data points, and construct a distance matrix based on the distances between all data pairs ; Among them, the calculation expression for using the Euclidean distance to calculate the distance between each pair of data points is: ; In the formula, represents one of the data points, represents another of the data points, represents the number of rows and columns of the distance matrix, represents one of the standardized data, represents another of the standardized data, represents the th data point, represents the total number of data points, and the distance matrix , represents the standardized value of the th type of data at the th data point, represents the standardized value of the th type of data at the th data point; Based on the distance matrix, use the hierarchical clustering algorithm to cluster the data, adopt the minimum distance method, and calculate the distance between clusters. The calculation expression is: ; In the formula, and are two clusters in the clustering process; For each cluster in the clustering result, calculate the average value of the distances between all data points in the cluster and the cluster center as the consistency deviation coefficient, and calculate the consistency deviation coefficient of each cluster. The calculation expression is: ; Among them, represents the th cluster, represents the center of the th cluster, is the number of data points in cluster , is the th data point and the Euclidean distance from the cluster center ; Among them, The calculation expression is: ; According to the deviation coefficient of each cluster, calculate the consistency deviation coefficient of the entire data set. The calculation expression is: ; In the formula, represents the consistency deviation coefficient of the entire data set, represents the total number of clusters; The process of obtaining the logical deviation coefficient is as follows: Standardize the data in the enterprise-related business system collected in real time; Use the maximum likelihood estimation and expectation maximization to fit the probability distribution of the standardized data to obtain the parameter ; Among them, represents the weight of the th Gaussian distribution, represents the number of Gaussian distributions, represents the mean vector of the th Gaussian distribution, represents the covariance matrix of the th Gaussian distribution; Calculate the posterior probability that each data point belongs to the th Gaussian distribution. The calculation expression is: ; In the formula, c represents the number of data points, represents the posterior probability that the cth data point belongs to the ath Gaussian distribution, represents the total number of Gaussian distributions; Update the parameter according to the result of the posterior probability; According to the updated new parameter , calculate its log-likelihood value. The calculation expression is: ; In the formula, represents the log-likelihood value of the th data point; According to the log-likelihood value of each data point, calculate the logical deviation coefficient of each data point. The calculation expression is: ; In the formula, represents the logical deviation coefficient of the th data point, represents the maximum log-likelihood value of all data points; Calculate the average logical deviation coefficient of the entire dataset. The calculation expression is as follows: ; In the formula, represents the total number of data points, represents the average logical deviation coefficient of the entire dataset; The calculation expression for the accuracy deviation coefficient is as follows: ; In the formula, represents the accuracy deviation coefficient, and represent preset proportionality coefficients, and and are both greater than 0; It should be noted that: The accuracy deviation coefficient reflects the data accuracy of the dataset, and when the accuracy deviation coefficient is larger, the accuracy of the corresponding dataset is lower.

[0022] In S3, based on the accuracy evaluation result, correct and optimize the errors in the data, specifically including: Obtain inaccurate data and calculate the data correction coefficient for correcting the data accuracy; The process of obtaining the data correction coefficient is as follows: Obtain the logical deviation coefficient and consistency deviation coefficient of the dataset, construct the prior distribution of the correction coefficient, and satisfy the normal distribution ; Among them, represents the expected value of the prior correction coefficient, represents the variance of the prior correction coefficient; The calculation expression for the prior distribution is as follows: ; In the formula, represents the prior distribution of the correction coefficient; Construct the likelihood function. The calculation expression is as follows: ; In the formula, represents the likelihood function of the logical deviation coefficient and consistency deviation coefficient; represents the prior distribution of the logical deviation coefficient calculated through the calculation expression of the prior distribution, represents the prior distribution of the consistency deviation coefficient calculated through the calculation expression of the prior distribution; Calculate the posterior distribution of the correction coefficient. The calculation expression is as follows: ; In the formula, represents the posterior distribution of the correction coefficient; represents the evidence term, which is calculated by integration, and the calculation expression is: ; The correction coefficient is calculated through the expected value of the posterior distribution, and the calculation expression is: ; In the formula, represents the correction coefficient; The original data set is multiplied by the correction coefficient to obtain the corrected accurate data set.

[0023] In S4, based on the optimized data, an enterprise decision prediction model is constructed to predict the risk of enterprise transformation. According to the risk prediction results of enterprise transformation, the enterprise transformation is divided into two modes: transformation and suspension of transformation, which specifically include: The original data set is multiplied by the correction coefficient to obtain the corrected accurate data set. The consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature are extracted from the corrected accurate data set and the original accurate data set; The consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature are used as input features and input into the enterprise decision prediction model; A decision tree is used to train the enterprise decision prediction model; The input features of each decision tree are the consistency deviation coefficient feature, logical deviation coefficient feature, and correction coefficient feature, and the output is the predicted risk coefficient; The structure of each tree is based on historical data, which is used as training data, and decisions are made through feature selection and splitting rules; Each tree is trained by randomly selecting subsets of data and subsets of features to increase diversity; The splitting node of each tree selects the method that can maximize the information gain; The final risk coefficient is obtained by averaging the outputs of all trees; The calculation expression for the prediction result of the risk coefficient is: ; In the formula, represents the decision tree in the random forest, represents the total number of decision trees in the random forest, represents the prediction result of the th decision tree, Perform risk prediction on the enterprise transformation. According to the risk prediction results of the enterprise transformation, divide the enterprise transformation into two modes: proceed with transformation and suspend transformation, specifically including: According to the enterprise decision prediction model, output the risk coefficient, compare the risk coefficient with the preset threshold, and determine whether the risk coefficient is greater than or equal to the preset threshold. If not, it is recorded as proceeding with transformation; if so, it is recorded as suspending transformation.

[0024] It should be noted that the risk coefficient reflects the possibility of risks occurring in enterprise data during the transformation. Moreover, the larger the value of the risk coefficient, the greater the possibility of risks occurring in the corresponding enterprise data during the transformation.

[0025] In S5, based on the enterprise decision to proceed with transformation, re-collect and evaluate the data in the enterprise-related business systems, and further judge the accuracy of the enterprise decision to proceed with transformation, specifically including: By preprocessing the data collected from the enterprise business systems, calculate the consistency deviation coefficient, logical deviation coefficient, and correction coefficient, and then construct a comprehensive risk assessment model. According to the comprehensive risk assessment model constructed, output the risk coefficient, calculate the difference between the new risk coefficient and the historical risk coefficient to obtain the risk coefficient deviation value, and compare the risk coefficient deviation value with the preset threshold; Judge whether the risk coefficient deviation value is greater than or equal to the preset threshold. If so, the enterprise decision to proceed with transformation is inaccurate; if not, the enterprise decision to proceed with transformation is accurate.

[0026] Please refer to Figure 2 As shown, the enterprise digital transformation analysis and decision-making system based on information error drive includes: A data collection module that collects and standardizes the data in the enterprise-related business systems in real time, processes the missing values, duplicate data, and outliers in the data, and ensures the integrity of the data structure; A data accuracy evaluation module that conducts consistency analysis and logical analysis on the collected data, evaluates the accuracy of the data, calculates the data accuracy deviation coefficient, quantifies the information error, and evaluates the accuracy of the data; A data correction module that corrects and optimizes the errors in the data based on the accuracy evaluation results; A risk prediction and decision-making module that constructs an enterprise decision prediction model based on the optimized data, performs risk prediction on the enterprise transformation, and divides the enterprise transformation into two modes: proceed with transformation and suspend transformation according to the risk prediction results of the enterprise transformation; Decision accuracy evaluation module, which re-collects and evaluates the data in the enterprise-related business systems based on the enterprise decisions for transformation, and further determines the accuracy of the enterprise decisions for transformation.

[0027] The working principle of the present invention: Optimize enterprise transformation decisions through accurate data analysis and risk assessment. The method includes real-time collection of multi-source data in enterprise-related business systems, such as business operation data, financial data, market data, and equipment operation data, and standardizes the collected data to ensure the integrity of the data structure. During the data processing process, special attention is paid to missing values, duplicate data, and outliers in the data. Through consistency analysis and logical analysis, the accuracy of the data is evaluated, the information error is calculated and quantified, and a consistency deviation coefficient and a logical deviation coefficient are generated. Based on these deviation coefficients, the errors in the data are corrected and optimized to improve the accuracy of the data.

[0028] Furthermore, based on the optimized data, an enterprise decision prediction model is constructed, and the risk of enterprise transformation is predicted through the random forest algorithm. Specifically, the consistency deviation coefficient, logical deviation coefficient, and correction coefficient features in the corrected data set and the original data set are used as inputs to train the decision tree model and calculate the risk coefficient. Finally, by comparing the predicted risk coefficient with a preset threshold, the enterprise transformation is divided into two modes: transformation or suspension of transformation. During the decision-making process, the data collected in the enterprise business system is evaluated again, and the deviation value is calculated by combining the new risk coefficient and the historical risk coefficient to further determine the accuracy of the enterprise decision, thereby realizing dynamic and scientific decision support. The present invention can help enterprises reduce risks, improve the accuracy of decisions, and the feasibility of implementation during the transformation process.

[0029] All the above formulas are calculated by removing the dimension and taking their numerical values. The formula is obtained by software simulation of collecting a large amount of data to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0030] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more sets of available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0031] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0032] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0033] The above has described in detail one embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. An enterprise digital transformation analysis and decision-making method based on information error driving, characterized by: The following steps are involved: S1: Collect and standardize data from the enterprise's relevant business systems in real time, process missing values, duplicate data and abnormal values ​​in the data, and ensure the integrity of the data structure; The data in the enterprise-related business system includes: business operation data, financial data, market data and equipment operation data; S2: Conduct consistency analysis and logic analysis on the collected data, evaluate the accuracy of the data, calculate the data accuracy deviation coefficient, quantify the information error, and evaluate the accuracy of the data; S3: Based on the accuracy assessment results, correct and optimize the errors in the data; S4: Based on the optimized data, an enterprise decision-making prediction model is constructed to predict the risks of enterprise transformation. According to the risk prediction results of enterprise transformation, the enterprise transformation is divided into two modes: transformation and suspension of transformation; S5: Based on the corporate decision to transform, re-collect and evaluate the data in the company's relevant business systems to further determine the accuracy of the corporate decision to transform.

2. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 1 is characterized in that: The consistency analysis and logic analysis of the collected data, evaluation of the accuracy of the data, calculation of the data accuracy deviation coefficient, quantification of information errors, and evaluation of the accuracy of the data specifically include: Collect data from the enterprise's relevant business systems in real time, perform consistency analysis on the data, and calculate the consistency deviation coefficient based on the analysis results to evaluate the consistency of the data; Collect data from the enterprise's relevant business systems in real time, analyze the data for logical accuracy, and calculate the logical deviation coefficient based on the analysis results to evaluate the logical accuracy of the data; Obtain the logical deviation coefficient and consistency deviation coefficient of the data, perform normalized calculation on the logical deviation coefficient and consistency deviation coefficient, and comprehensively calculate the accuracy deviation coefficient; Determine whether the accuracy deviation coefficient of the data is greater than or equal to the preset threshold. If so, it means that the corresponding data is inaccurate. If not, it means that the corresponding data is accurate.

3. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 2 is characterized in that: The process of obtaining the consistency deviation coefficient is as follows: Collect data from the enterprise's relevant business systems in real time, and standardize the data collected in real time from the enterprise's relevant business systems so that features of different dimensions can be compared consistently; Use Euclidean distance to calculate the distance between each pair of data points, and construct a distance matrix based on the distances between all data pairs. ; The calculation expression for calculating the distance between each pair of data points using the Euclidean distance is: ; In the formula, represents one of the data points, represents another data point, represents the number of rows and columns of the distance matrix, Represents one of the data after standardization, Represents another type of data after standardization, Indicates data points, Represents the total number of data points, distance matrix , Indicates The data in The standardized value of the data points, Indicates The data in Normalized values ​​at data points; Based on the distance matrix, a hierarchical clustering algorithm is used to cluster the data. The minimum distance method is used to calculate the distance between clusters. The calculation expression is: ; In the formula, and Two clusters in the clustering process; For each cluster in the clustering result, the average value of the distance between all data points in the cluster and the cluster center is calculated as the consistency deviation coefficient. The consistency deviation coefficient of each cluster is calculated using the following expression: ; in, Indicates Clusters, Indicates The center of the cluster, It is a cluster The number of data points in It is Data points With cluster center The Euclidean distance of in, The calculation expression is: ; According to the deviation coefficient of each cluster, the consistency deviation coefficient of the entire data set is calculated. The calculation expression is: ; In the formula, represents the consistency deviation coefficient of the entire data set, Represents the total number of clusters.

4. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 3 is characterized in that: The process of obtaining the logical deviation coefficient is as follows: Standardize the data collected in real time from the enterprise-related business systems; The standardized data is fitted with the probability distribution of the data through maximum likelihood estimation and expectation maximization to obtain the parameters ; in, Indicates The weights of a Gaussian distribution, represents the number of Gaussian distributions, Indicates The mean vector of a Gaussian distribution, Indicates The covariance matrix of the Gaussian distribution; Calculate the number of each data point that belongs to The calculation expression for the delay probability of a Gaussian distribution is: ; In the formula, c represents the number of data points, represents the posterior probability that the cth data point belongs to the ath Gaussian distribution, represents the total number of Gaussian distributions; Update the parameters based on the posterior probability results ; According to the updated new parameters , calculate its log-likelihood value, the calculation expression is: ; In the formula, Indicates The log-likelihood of the data points; According to the log-likelihood value of each data point, the logical deviation coefficient of each data point is calculated, and the calculation expression is: ; In the formula, Indicates The logistic deviation coefficient of the data points is represents the maximum log-likelihood value of all data points; Calculate the average logical deviation coefficient of the entire data set. The calculation expression is: ; In the formula, represents the total number of data points, Represents the average logistic deviation coefficient of the entire data set.

5. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 1 is characterized in that: The correction and optimization of errors in the data based on the accuracy evaluation results specifically include: Obtain inaccurate data and calculate data correction coefficients to correct the accuracy of the data; The process of obtaining the data correction coefficient is as follows: Obtain the logical deviation coefficient and consistency deviation coefficient of the data set, construct the prior distribution of the correction coefficient, and satisfy the normal distribution ; in, represents the expected value of the a priori correction coefficient, represents the variance of the prior correction coefficient; The calculation expression of the prior distribution is: ; In the formula, represents the prior distribution of the correction coefficient; Construct the likelihood function and calculate the expression as follows: ; In the formula, Likelihood function representing the logistic bias coefficient and the consistency bias coefficient; represents the prior distribution of the logical deviation coefficient calculated by the calculation expression of the prior distribution, represents the prior distribution of the consistency deviation coefficient calculated by the calculation expression of the prior distribution; Calculate the posterior distribution of the correction coefficient, the calculation expression is: ; In the formula, represents the posterior distribution of the correction coefficient; Represents the evidence term, and through integral calculation, the calculation expression is: ; The correction coefficient is calculated by the expected value of the posterior distribution, and the calculation expression is: ; In the formula, Indicates the correction factor.

6. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 1 is characterized in that: The enterprise decision-making prediction model is constructed based on the optimized data, specifically including: A corrected accurate data set is obtained by multiplying the correction coefficient with the original data set, and consistency deviation coefficient features, logic deviation coefficient features and correction coefficient features are extracted from the corrected accurate data set and the original accurate data set; The consistency deviation coefficient feature, the logic deviation coefficient feature and the correction coefficient feature are used as input features and input into the enterprise decision-making prediction model; Use decision trees to train enterprise decision prediction models; The input features of each decision tree are consistency deviation coefficient features, logical deviation coefficient features and correction coefficient features, and the output is the predicted risk coefficient; The structure of each tree is based on historical data as training data, and decisions are made through feature selection and splitting rules; Each tree is trained by randomly selecting a subset of the data and a subset of the features to increase diversity; The splitting node of each tree is selected in a way that maximizes information gain; The final risk factor is obtained by averaging the outputs of all trees; The calculation expression of the prediction result of the risk coefficient is: ; In the formula, represents a decision tree in a random forest, represents the total number of decision trees in the random forest, Indicates The prediction results of a decision tree, Represents the predicted result of risk factor.

7. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 1 is characterized in that: The risk prediction of enterprise transformation is carried out, and according to the risk prediction results of enterprise transformation, enterprise transformation is divided into two modes: transformation and suspension of transformation, which specifically include: According to the enterprise decision-making prediction model, the risk coefficient is output, and the risk coefficient is compared with the preset threshold to determine whether the risk coefficient is greater than or equal to the preset threshold. If not, it is recorded as transformation, and if so, it is recorded as suspension of transformation.

8. The enterprise digital transformation analysis and decision-making method based on information error drive according to claim 1 is characterized in that: The enterprise decision to transform is based on the re-collection and evaluation of data in the enterprise's relevant business systems to further determine the accuracy of the enterprise decision to transform, specifically including: By preprocessing the data collected in the enterprise business system, calculating the consistency deviation coefficient, logical deviation coefficient and correction coefficient, and then building a comprehensive risk assessment model, outputting the risk coefficient based on the comprehensive risk assessment model, calculating the difference between the new risk coefficient and the historical risk coefficient, and obtaining the risk coefficient deviation value, and comparing the risk coefficient deviation value with the preset threshold value; Determine whether the risk coefficient deviation value is greater than or equal to the preset threshold. If so, the enterprise decision to transform is inaccurate; if not, the enterprise decision to transform is accurate.

9. The enterprise digital transformation analysis and decision-making system based on information error driving is characterized by: The enterprise digital transformation analysis and decision-making method based on information error driving as described in any one of claims 1 to 8 comprises: A data collection module, which collects and standardizes data from the enterprise's relevant business systems in real time, processes missing values, duplicate data, and abnormal values ​​in the data, and ensures the integrity of the data structure; A data accuracy assessment module, which performs consistency analysis and logic analysis on the collected data, assesses the accuracy of the data, calculates the data accuracy deviation coefficient, quantifies information errors, and assesses the accuracy of the data; A data correction module, which corrects and optimizes errors in the data based on the accuracy evaluation results; A risk prediction and decision-making module, which builds an enterprise decision-making prediction model based on the optimized data, predicts the risks of enterprise transformation, and divides the enterprise transformation into two modes: transformation and suspension of transformation according to the risk prediction results of the enterprise transformation; A decision accuracy assessment module re-collects and assesses the data in the enterprise's relevant business systems based on the enterprise's decision to undergo transformation, and further determines the accuracy of the enterprise's decision to undergo transformation.

Citation Information

Patent Citations

  • Oil and gas field investment risk prediction and early warning method and system

    CN118840206A

  • Infrastructure multi-source heterogeneous data fusion security situation holographic sensing control method

    CN119130202A

  • Enterprise service data analysis method and system based on cloud computing

    CN119441201A

  • Information security protection method and system applied to electronic information platform

    CN119691759A

  • Data preparation device and data preparation method

    JP2014044536A