Enterprise data optimization processing method based on cloud computing data fusion
By collecting and storing data from multiple data sources on the cloud computing platform, conducting consistency and logical analysis, and correcting inaccurate data, the problems of data inconsistency and logical errors in the prior art are solved, and the accuracy and reliability of data are improved.
Patent Information
- Application Number
- CN202510131389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing cloud-based data processing methods lack effective consistency checks in the data fusion process of multiple data sources, resulting in inconsistency of data, affecting the accuracy of analysis, and ignoring the in-depth analysis of whether the data complies with business logic, resulting in logical errors not being discovered in time.
Data is collected from multiple data sources through the cloud platform interface, and the storage resources provided by the cloud computing platform are used for hierarchical storage and encryption processing, consistency detection and logical analysis are performed, the consistency coefficient and logical coefficient of each piece of data are calculated, the data accuracy is comprehensively evaluated, and inaccurate data are corrected.
It improves the quality and credibility of data, enhances the data reliability of enterprises in decision support, business analysis and prediction models, and ensures that data-driven decisions are more scientific and accurate.
Smart Images

Figure CN120068146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enterprise data optimization in cloud computing, and particularly relates to a method for optimizing and processing enterprise data based on cloud computing data fusion. Background Art
[0002] With the continuous advancement of informatization and digitalization, enterprises generate a large amount of business data in their daily operations. This data comes from different systems and devices, such as multiple fields including sales, finance, inventory, and supply chain. As the data volume increases, how to effectively collect, store, and process this heterogeneous data has become an important challenge for enterprise decision-making support and business optimization. Cloud computing, as a powerful technical platform, can provide flexible computing and storage resources to help enterprises process massive data and support cross-departmental and cross-regional data sharing and analysis. However, with the diversification of data sources, how to ensure the consistency, logic, and accuracy of data and avoid decision-making errors caused by data quality problems remains an urgent problem to be solved in the process of enterprise data optimization.
[0003] The existing technologies have the following deficiencies:
[0004] Most of the existing cloud computing-based data processing methods focus on data collection, storage, and preliminary analysis, but there are still obvious deficiencies in data quality assurance. First, in the process of data fusion from multiple data sources, the existing technologies often lack effective consistency checks, resulting in conflicts or inconsistencies in data from different sources, which affects the accuracy of data analysis. Second, although logical data analysis is a key link in data processing, many existing solutions ignore the in-depth analysis of whether the data conforms to business logic, resulting in logical errors not being discovered in a timely manner. In addition, existing methods often only focus on static analysis of data and lack a dynamic accuracy evaluation mechanism, unable to provide real-time and comprehensive accuracy feedback for enterprises, further affecting the effect of decision-making support. Therefore, there are obvious technical gaps in the existing technologies in ensuring data quality, improving data consistency and logic, and an innovative method that can comprehensively evaluate data consistency, logic, and calculate accuracy scores is urgently needed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for optimizing and processing enterprise data based on cloud computing data fusion to solve the problems in the above background.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A method for optimizing and processing enterprise data based on cloud computing data fusion, S1: Collect structured, semi-structured, and unstructured data from internal and external data sources through a cloud platform interface, and the data sources include: sales systems, financial systems, supply chain management systems, and market data sources;
[0008] S2: Hierarchically store the collected data using the storage resources provided by the cloud computing platform, and encrypt the stored data to ensure data security;
[0009] S3: Perform consistency detection on the stored data. By comparing the same data in different data sources, identify the differences between the data. According to the identification results, calculate the consistency coefficient of each piece of data, and calculate the consistency score of the overall data set to evaluate the consistency level of data storage;
[0010] S4: Perform logical analysis on the stored data to detect whether the data conforms to logical rules. According to the logical detection results, calculate the logical coefficient of each piece of data, and calculate the logical score for the overall data set to evaluate the logical level of the stored data;
[0011] S5: Comprehensively evaluate the accuracy of each day's data according to the consistency coefficient and logical coefficient of each piece of data, and through normalization processing, calculate the accuracy score of each piece of data to determine whether each piece of data is accurate;
[0012] S6: According to the judgment results, extract the inaccurate data. According to the consistency coefficient and logical coefficient of the inaccurately stored data, calculate the consistency correction coefficient and logical correction coefficient of the stored data, correct the inaccurately stored data, and based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to judge the overall accuracy of the stored data.
[0013] As a further solution of the present invention: The specific steps for calculating the consistency coefficient of each piece of data include:
[0014] Obtain the same data items from different data sources and convert them into probability distributions, including:
[0015] If there are two data sources D 1 and data source D 2 , and there is the same data item x in both of these data sources;
[0016] where x represents each piece of data;
[0017] Convert the data item x into probability distributions, denoted as P(x) and Q(x);
[0018] where P(x) represents the probability distribution of the data item x in data source D 1 , and Q(x) represents the probability distribution of the data item x in data source D 2 ;
[0019] Calculate the difference value between the probability distributions of the two data sources. The calculation expression is:
[0020]
[0021] Where D KL (P∥Q) indicates the data source D 1 and data source D 2 The difference between
[0022] According to data source D 1 and data source D 2 The difference between them is used to calculate the consistency coefficient of the data item x. The calculation expression is:
[0023]
[0024] In the formula, C(x) is the consistency coefficient of data item x, which represents the consistency coefficient of each data.
[0025] As a further solution of the present invention: the calculation of the consistency score of the entire data set is used to evaluate the consistency level of data storage, specifically including:
[0026] Obtain the consistency coefficient of each piece of data, and calculate the average of the consistency coefficients of all data in a data source to obtain the consistency score of the entire data set;
[0027] Determine whether the consistency score of the entire data set is greater than or equal to a preset threshold. If so, the corresponding data is stored consistently; if not, the corresponding data is stored inconsistently.
[0028] As a further solution of the present invention: the calculation of the logic coefficient of each piece of data specifically includes:
[0029] Get each field in the data set as a node of the Bayesian network;
[0030] Define a conditional probability table for each node in the Bayesian network;
[0031] The conditional probability table represents the probability distribution of a node under the condition of a given parent node;
[0032] Perform rule detection on each piece of data to verify whether it meets the set of logical rules predetermined by the enterprise;
[0033] The Bayesian inference formula is used to calculate the logical conditional probability for each piece of data. The calculation expression is:
[0034]
[0035] In the formula, x j represents the jth data, j represents the number of data, P(LogicValid) represents the prior probability that the data is logically qualified, P(x j|LogicValid) represents the likelihood probability of data under logically qualified conditions, P(x j ) represents the marginal probability of the j-th data, P(LogicValid|x j ) represents the logical conditional probability of the j-th data;
[0036] Based on the Bayesian inference result, calculate the logical coefficient of each data, and the calculation expression is:
[0037] L(x j ) = P(LogicValid|x j );
[0038] In the formula, L(x j ) represents the logical coefficient of the j-th data.
[0039] As a further solution of the present invention: calculating the logical score of the overall data set for evaluating the logical level of the stored data, specifically including:
[0040] Obtain the logical coefficient of each data, and calculate the logical score of the overall data set through the mean calculation formula;
[0041] Judge whether the logical score of the overall data set is greater than or equal to a preset threshold. If so, the overall data set conforms to the logical rules. If not, the overall data set does not conform to the logical rules.
[0042] As a further solution of the present invention: calculating the accuracy score of each data, specifically including:
[0043] Obtain the consistency coefficient and logical coefficient of each data, perform normalized comprehensive processing on the consistency coefficient and logical coefficient of each data, and calculate the accuracy score of each data according to the processing result through the calculation expression of the accuracy coefficient.
[0044] As a further solution of the present invention: judging whether each data is accurate, specifically including:
[0045] Judge whether the accuracy score of each data is greater than or equal to a preset threshold. If so, the corresponding data is accurate. If not, the corresponding data is inaccurate.
[0046] As a further solution of the present invention: calculating the consistency correction coefficient and logical correction coefficient of the stored data, specifically including:
[0047] Obtain the consistency coefficient C(x j ) and logical coefficient L(x j ) of inaccurate data during the storage process;
[0048] Among them, x jDenote the j-th data, where j represents the number of data;
[0049] Use the support vector random algorithm to train the inaccurate data in the dataset, generate a consistency correction function, and calculate the consistency correction coefficient. The calculation expression is:
[0050] δ C (x j ) = f C (x j ) - C(x j );
[0051] In the formula, f C (x j ) represents the consistency correction value calculated by the support vector random algorithm, δ C (x j ) represents the consistency correction coefficient of the j-th data, and C(x j ) represents the consistency coefficient of the j-th data;
[0052] Use the support vector random algorithm to generate a logical correction function, calculate the logical correction coefficient. The calculation expression is:
[0053] δ L (x j ) = f L (x j ) - L(x j );
[0054] In the formula, f L (x j ) represents the logical correction value calculated by the support vector random algorithm, L(x j ) represents the logical index of the j-th data, and δ L (x j ) represents the logical correction coefficient of the j-th data.
[0055] As a further solution of the present invention: The correction of the inaccurately stored data specifically includes:
[0056] Based on the consistency correction coefficient and the logical correction coefficient, respectively correct the consistency coefficient and the logical coefficient of each data to obtain the corrected data consistency coefficient and logical coefficient. The calculation expression is:
[0057] C new (x j ) = C(x j ) + δ C (x j );
[0058] L new (x j) = L(x j ) + δ L (x j );
[0059] In the formula, C new (x j ) represents the corrected data consistency coefficient, C(x j ) represents the consistency coefficient of the j-th piece of data, δ C (x j ) represents the consistency correction coefficient of the j-th piece of data, L new (x j ) represents the corrected data logic coefficient, L(x j ) represents the logic coefficient of the j-th piece of data, δ L (x j ) represents the logic correction coefficient of the j-th piece of data.
[0060] As a further solution of the present invention: based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to judge the overall accuracy of the stored data, specifically including:
[0061] According to the corrected data consistency coefficient and logic coefficient, recalculate the accuracy coefficient of each piece of data respectively, and calculate the average value of the accuracy coefficients of all data through the average value calculation formula to obtain the comprehensive accuracy score. According to the comparison between the comprehensive accuracy score and the preset threshold, re-judge whether the stored data is accurate after correction.
[0062] The beneficial effects of the present invention:
[0063] (1) The present invention provides an enterprise data optimization processing method based on cloud computing data fusion, aiming to solve the problems of multi-source data inconsistency, data logic conflicts, and data security faced by enterprises in the big data environment. Through the cloud platform interface, the system can collect structured, semi-structured, and unstructured data in real time from various data sources such as internal sales systems, financial systems, and supply chain management systems, and adopt a hierarchical storage strategy according to different data types to ensure the efficient access and utilization of data in different business scenarios. At the same time, strict encryption processing is carried out during the data storage process to ensure that sensitive data will not be leaked or tampered with during storage and transmission. Further through consistency detection and logical analysis, the system can accurately calculate the consistency coefficient and logical coefficient of each piece of data, and combine the evaluation results of the two to intelligently correct the data. Especially when there are inconsistencies or logical errors in the data, through the combined calculation of the consistency correction coefficient and the logical correction coefficient, the unreasonable or non-standard parts in the data can be effectively repaired, thereby optimizing the overall accuracy of the data. This technical solution not only improves the quality and credibility of the data, but also enhances the data reliability of enterprises in decision support, business analysis, and prediction models, ensuring that data-driven decisions are more scientific and accurate.
[0064] (2) By combining support vector machines, it is possible to automatically and accurately calculate the consistency correction coefficient and logical correction coefficient of each piece of data based on the consistency coefficient and logical coefficient of inaccurate stored data, and intelligently correct the data by combining historical data and business rules. Through continuous training and optimization, the correction coefficients can be continuously adjusted to adapt to different data quality problems and business scenarios, ensuring that the correction process has high flexibility and adaptability. This innovative technology not only significantly improves the overall accuracy of the data, but also can finely adjust the data correction strategy according to the actual needs of the enterprise, thereby effectively improving the data processing efficiency, reducing human intervention, and providing more accurate and reliable data support for enterprise decision-making, ultimately promoting the improvement of the enterprise's data-driven decision-making ability in the dynamic market environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The present invention will be further described below with reference to the accompanying drawings.
[0066] Figure 1 is the specific step flow block diagram of an enterprise data optimization processing method based on cloud computing data fusion of the present invention;
[0067] Figure 2 is the flow block diagram of the evaluation and correction steps of data accuracy in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0069] Please refer to Figure 1 As shown, the present invention is an enterprise data optimization processing method based on cloud computing data fusion, including the following steps:
[0070] S1: Collect structured, semi-structured, and unstructured data from internal and external data sources through the cloud platform interface. The data sources include: sales systems, financial systems, supply chain management systems, and market data sources;
[0071] S2: Use the storage resources provided by the cloud computing platform to hierarchically store the collected data, and encrypt the stored data to ensure data security;
[0072] S3: Perform consistency detection on the stored data. By comparing the same data in different data sources, identify the differences between the data. According to the identification results, calculate the consistency coefficient of each piece of data, and calculate the consistency score of the overall data set to evaluate the consistency level of data storage;
[0073] S4: Perform logical analysis on the stored data to detect whether the data conforms to logical rules. According to the logical detection results, calculate the logical coefficient of each piece of data, and calculate the logical score for the overall data set to evaluate the logical level of the stored data;
[0074] S5: Comprehensively evaluate the accuracy of each piece of data according to the consistency coefficient and logical coefficient of each piece of data, and through normalization processing, calculate the accuracy score of each piece of data to determine whether each piece of data is accurate;
[0075] S6: According to the judgment results, extract inaccurate data. According to the consistency coefficient and logical coefficient of the inaccurately stored data, calculate the consistency correction coefficient and logical correction coefficient of the stored data, correct the inaccurately stored data, and based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to determine the overall accuracy of the stored data.
[0076] In S1: Collect structured, semi-structured, and unstructured data from internal and external data sources through the cloud platform interface. The data sources include: sales systems, financial systems, supply chain management systems, and market data sources, specifically including:
[0077] Through the cloud platform interface, the system collects various types of data required by the enterprise in real time from multiple internal and external data sources;
[0078] Internal data sources include sales systems, financial systems, and supply chain management systems, etc. Structured data such as order information, financial statements, and inventory data are stored in these systems;
[0079] At the same time, external data sources include market data sources, such as market analysis data, social media data, or industry research reports from third-party platforms;
[0080] This data may be in semi-structured or unstructured form, such as JSON format API data, text files, or image data;
[0081] During the data collection process, the system uses the API interface or data synchronization tools provided by the cloud platform to ensure the real-time and integrity of the data;
[0082] For structured data, it is extracted in batches or in real time by directly connecting to the database;
[0083] For semi-structured and unstructured data, the system uses data parsing and conversion tools and uniformly formats them into a structure suitable for storage and analysis for subsequent processing and analysis;
[0084] The whole process ensures that the data converges to the cloud platform smoothly and efficiently, supporting the enterprise to conduct large-scale data analysis and decision-making.
[0085] In S2, the storage resources provided by the cloud computing platform are used to store the collected data in a hierarchical manner, and the stored data is encrypted to ensure data security, specifically including:
[0086] After the data collection is completed, the system uses the storage resources provided by the cloud computing platform to store different types and levels of data in a hierarchical manner;
[0087] Structured data such as financial statements and order records are usually stored in a high-performance relational database for quick query and analysis; semi-structured data such as JSON files and log data are stored in object storage or NoSQL databases for convenient and efficient processing and access;
[0088] Unstructured data such as pictures, videos, and documents are stored using file storage or a distributed file system to support the storage and backup of large-scale data;
[0089] According to the access frequency and importance of the data, the system automatically selects the appropriate storage method through policies and allocates resources;
[0090] To ensure data security, the system encrypts the stored data;
[0091] All sensitive data, such as customer information, financial data, etc., will be encrypted through an encryption algorithm before being written to cloud storage to ensure that the data will not be leaked or tampered with during storage;
[0092] Industry-standard encryption algorithms, such as AES-256, are adopted to ensure the security of the encrypted data; at the same time, the system will also set strict access control policies for the stored encryption keys, and only authorized personnel and applications can decrypt and access sensitive data;
[0093] The data transmission process will also be encrypted through security protocols such as TLS to further enhance the overall security of the data.
[0094] In S3, consistency detection is performed on the stored data. By comparing the same data in different data sources, the differences between the data are identified. According to the identification results, the consistency coefficient of each piece of data is calculated, and the consistency score of the overall data set is calculated to evaluate the consistency level of data storage;
[0095] Among them, the calculation of the consistency coefficient of each piece of data specifically includes:
[0096] Obtain the same data items from different data sources and convert them into probability distributions, including:
[0097] If there are two data sources D 1 and data source D 2 , and there is the same data item x in both of these data sources;
[0098] Among them, x represents each piece of data;
[0099] Convert the data item x into probability distributions, denoted as P(x) and Q(x);
[0100] Among them, P(x) represents the probability distribution of the data item x in data source D 1 , and Q(x) represents the probability distribution of the data item x in data source D 2 ;
[0101] Calculate the difference value between the probability distributions of the two data sources. The calculation expression is:
[0102]
[0103] In the formula, D KL (P∥Q) represents the difference value between data source D 1 and data source D 2 ;
[0104] According to data source D 1and data source D 2 Calculate the difference value between them, and calculate the consistency coefficient of data item x. The calculation formula is:
[0105]
[0106] In the formula, C(x) is the consistency coefficient of data item x, representing the consistency coefficient of each piece of data;
[0107] Among them, calculating the consistency score of the overall data set is used to evaluate the consistency level of data storage, specifically including:
[0108] Obtain the consistency coefficient of each piece of data, and calculate the consistency score of the overall data set by calculating the average value of the consistency coefficients of all data in a data source;
[0109] Judge whether the consistency score of the overall data set is greater than or equal to the preset threshold. If so, the corresponding data storage is consistent; if not, the corresponding data storage is inconsistent.
[0110] It should be noted that: The consistency coefficient measures the matching degree of data between different data sources or at different time points, that is, whether the data remains consistent.
[0111] In S4, perform logical analysis on the stored data, detect whether the data conforms to the logical rules, and calculate the logical coefficient of each piece of data according to the logical detection result, and calculate the logical score for the overall data set to evaluate the logical level of the stored data;
[0112] Among them, calculating the logical coefficient of each piece of data specifically includes:
[0113] Obtain each field in the data set as a node of the Bayesian network;
[0114] Define a conditional probability table for each node in the Bayesian network;
[0115] The conditional probability table represents the probability distribution of the node under the condition of the given parent node;
[0116] Perform rule detection on each piece of data to verify whether it meets the set of logical rules predetermined by the enterprise;
[0117] Calculate the logical conditional probability for each piece of data using the Bayesian inference formula. The calculation formula is:
[0118]
[0119] In the formula, x j represents the j-th piece of data, j represents the number of data, P(LogicValid) represents the prior probability that the data is logically qualified, P(x j|LogicValid) represents the likelihood probability of the data under logically qualified conditions, P(x j ) represents the marginal probability of the j-th data, P(LogicValid|x j ) represents the logical conditional probability of the j-th data;
[0120] Based on the Bayesian inference results, calculate the logical coefficient of each data, and the calculation expression is:
[0121] L(x j ) = P(LogicValid|x j );
[0122] In the formula, L(x j ) represents the logical coefficient of the j-th data;
[0123] Among them, calculating the logical score for the overall data set is used to evaluate the logical level of the stored data, specifically including:
[0124] Obtain the logical coefficient of each data, and calculate the logical score of the overall data set through the mean calculation formula;
[0125] Judge whether the logical score of the overall data set is greater than or equal to the preset threshold. If so, the overall data set conforms to the logical rules; if not, the overall data set does not conform to the logical rules.
[0126] It should be noted that: the logical coefficient measures whether the data conforms to the predetermined business logic; it indicates whether the data is logically reasonable, that is, whether the value of the data conforms to the business rules, constraints, and conditions.
[0127] Please refer to Figure 2 As shown, in S5, according to the consistency coefficient and logical coefficient of each data, comprehensively evaluate the accuracy of each data, and through normalization processing, calculate the accuracy score of each data to determine whether each data is accurate, specifically including:
[0128] The calculation of the accuracy score of each data specifically includes:
[0129] Obtain the consistency coefficient and logical coefficient of each data, perform normalization comprehensive processing on the consistency coefficient and logical coefficient of each data, and according to the processing result, calculate the accuracy score of each data through the calculation expression of the accuracy coefficient;
[0130] Among them, the calculation expression of the accuracy score of each data is:
[0131]
[0132] In the formula, x jDenote the j-th data, where j represents the number of data, C(x j ) is the consistency coefficient of the j-th data, and L(x j ) represents the logical coefficient of the j-th data. a 1 and a 2 are preset proportionality coefficients, and both a 1 and a 2 are greater than 0;
[0133] Judging whether each piece of data is accurate specifically includes:
[0134] Judging whether the accuracy score of each piece of data is greater than or equal to a preset threshold. If so, the corresponding data is accurate; if not, the corresponding data is inaccurate.
[0135] It should be noted that: The accuracy coefficient comprehensively considers the consistency and logic of the data and measures the overall accuracy of the data.
[0136] In S6, according to the judgment result, extract the inaccurate data. According to the consistency coefficient and logical coefficient of the inaccurately stored data, calculate the consistency correction coefficient and logical correction coefficient of the stored data, correct the inaccurately stored data, and based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to judge the overall accuracy of the stored data;
[0137] Among them, calculating the consistency correction coefficient and logical correction coefficient of the stored data specifically includes:
[0138] Obtain the consistency coefficient C(x j ) and logical coefficient L(x j ) of the inaccurate data during the storage process;
[0139] Among them, x j represents the j-th data, and j represents the number of data;
[0140] Use the support vector random algorithm to train the inaccurate data in the data set, generate a consistency correction function, and calculate the consistency correction coefficient. The calculation expression is:
[0141] δ C (x j ) = f C (x j ) - C(x j );
[0142] In the formula, f C (x j ) represents the consistency correction value calculated by the support vector random algorithm, and δ C (x j) represents the consistency correction coefficient of the j-th data, C(x j ) represents the consistency coefficient of the j-th data;
[0143] Use the support vector random algorithm to generate a logical correction function, calculate the logical correction coefficient, and the calculation expression is:
[0144] δ L (x j ) = f L (x j ) - L(x j );
[0145] In the formula, f L (x j ) represents the logical correction value calculated by the support vector random algorithm, L(x j ) represents the logical index of the j-th data, δ L (x j ) represents the logical correction coefficient of the j-th data;
[0146] The correction of inaccurate stored data specifically includes:
[0147] Based on the consistency correction coefficient and the logical correction coefficient, respectively correct the consistency coefficient and the logical coefficient of each data to obtain the corrected data consistency coefficient and logical coefficient, and the calculation expression is:
[0148] C new (x j ) = C(x j ) + δ C (x j );
[0149] L new (x j ) = L(x j ) + δ L (x j );
[0150] In the formula, C new (x j ) represents the corrected data consistency coefficient, C(x j ) represents the consistency coefficient of the j-th data, δ C (x j ) represents the consistency correction coefficient of the j-th data, L new (x j ) represents the corrected data logical coefficient, L(x j ) represents the logical coefficient of the j-th data, δ L (x j ) represents the logical correction coefficient of the j-th data;
[0151] Optimize the correction process according to the corrected data consistency and logic coefficients, dynamically adjust the parameters of the support vector random algorithm to improve the correction accuracy; through historical data analysis, automatically optimize the weights of consistency and logic correction to adapt to the correction requirements of different data sets;
[0152] Based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to judge the overall accuracy of the stored data, specifically including:
[0153] According to the corrected data consistency coefficient and logic coefficient, recalculate the accuracy coefficient of each piece of data respectively. Through the mean value calculation formula, calculate the mean value of the accuracy coefficients of all data to obtain the comprehensive accuracy score. According to the comparison between the comprehensive accuracy score and the preset threshold, re-judge whether the stored data is accurate after correction.
[0154] The working principle of the present invention: Through multi-dimensional data consistency and logic analysis, realize the accuracy evaluation and correction of enterprise stored data, including: Collect structured, semi-structured and unstructured data from multiple internal and external data sources through the cloud platform interface, and use the storage resources provided by the cloud platform for hierarchical storage and encryption processing to ensure data security; Perform consistency detection and logic analysis on the stored data, calculate the consistency coefficient and logic coefficient of each piece of data respectively, and comprehensively evaluate the accuracy of the data based on these two coefficients; Calculate the accuracy score of each piece of data through normalization processing, calculate the consistency correction coefficient and logic correction coefficient for inaccurate data, and use the support vector random algorithm for correction. Finally, calculate the accuracy score of the corrected data to judge the overall accuracy of the stored data. The present invention improves the quality of enterprise data through efficient data correction and optimization strategies, ensures the reliability and consistency of data, and provides a solid foundation for enterprise decision support and data analysis.
[0155] The above formulas are all calculated by taking the numerical value without dimension. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0156] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more sets of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0157] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0158] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0159] The above has described in detail one embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A method for optimizing enterprise data based on cloud computing data fusion, characterized in that: The following steps are involved: S1: Collect structured, semi-structured and unstructured data from internal and external data sources through the cloud platform interface, including sales system, financial system, supply chain management system and market data source; S2: Use the storage resources provided by the cloud computing platform to store the collected data in a hierarchical manner, and encrypt the stored data to ensure data security; S3: Perform consistency check on the stored data, identify the differences between the data by comparing the same data in different data sources, calculate the consistency coefficient of each data according to the identification results, and calculate the consistency score of the entire data set to evaluate the consistency level of data storage; S4: Perform logical analysis on the stored data to detect whether the data conforms to logical rules. According to the logical test results, calculate the logical coefficient of each piece of data and calculate the logical score of the entire data set to evaluate the logical level of the stored data. S5: Based on the consistency coefficient and logic coefficient of each piece of data, the accuracy of each piece of data is comprehensively evaluated, and the accuracy score of each piece of data is calculated through normalization processing to determine whether each piece of data is accurate; S6: Based on the judgment result, extract the inaccurate data, calculate the consistency correction coefficient and the logic correction coefficient of the stored data based on the consistency coefficient and the logic coefficient of the inaccurate stored data, correct the inaccurate stored data, and based on the corrected stored data, comprehensively calculate the comprehensive accuracy score of the entire data set to judge the overall accuracy of the stored data.
2. According to the enterprise data optimization processing method based on cloud computing data fusion according to claim 1, it is characterized in that: The calculation of the consistency coefficient of each piece of data specifically includes: Get the same data item from different data sources and convert it into a probability distribution, including: If there are two data sources D1 and D2, and there is the same data item x in both data sources; Among them, x represents each piece of data; Convert the data item x into a probability distribution, denoted as P(x) and Q(x); Where P(x) represents the probability distribution of data item x in data source D1, and Q(x) represents the probability distribution of data item x in data source D2; Calculate the difference in probability distribution between two data sources. The calculation expression is: Where D KL (P||Q) represents the difference value between data source D1 and data source D2; According to the difference between data source D1 and data source D2, the consistency coefficient of data item x is calculated. The calculation expression is: In the formula, C(x) is the consistency coefficient of data item x, which represents the consistency coefficient of each data.
3. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The calculation of the consistency score of the overall data set is used to evaluate the consistency level of data storage, specifically including: Obtain the consistency coefficient of each piece of data, and calculate the average of the consistency coefficients of all data in a data source to obtain the consistency score of the entire data set; Determine whether the consistency score of the entire data set is greater than or equal to a preset threshold. If so, the corresponding data is stored consistently; if not, the corresponding data is stored inconsistently.
4. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The calculation of the logic coefficient of each piece of data specifically includes: Get each field in the data set as a node of the Bayesian network; Define a conditional probability table for each node in the Bayesian network; The conditional probability table represents the probability distribution of a node under a given parent node condition; Perform rule detection on each piece of data to verify whether it meets the set of logical rules predetermined by the enterprise; The Bayesian inference formula is used to calculate the logical conditional probability for each piece of data. The calculation expression is: In the formula, x j represents the jth data, j represents the number of data, P(LogicValid) represents the prior probability that the data is logically qualified, P(x j |LogicValid) represents the likelihood probability of the data under the logical qualification condition, P(x j ) represents the marginal probability of the jth data, P(LogicValid|x j ) represents the logical conditional probability of the jth data; Based on the Bayesian inference results, the logical coefficient of each data is calculated, and the calculation expression is: L(x j )=P(LogicValid∣x j ); In the formula, L(x j ) represents the logical coefficient of the j-th data.
5. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The logic score of the whole data set is calculated to evaluate the logic level of the stored data, specifically including: Get the logic coefficient of each piece of data, and calculate the logic score of the entire data set through the mean calculation formula; It is determined whether the logic score of the entire data set is greater than or equal to a preset threshold. If so, the entire data set complies with the logic rule; if not, the entire data set does not comply with the logic rule.
6. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The calculation of the accuracy score of each piece of data specifically includes: Obtain the consistency coefficient and logic coefficient of each piece of data, normalize and comprehensively process the consistency coefficient and logic coefficient of each piece of data, and calculate the accuracy score of each piece of data based on the processing results through the calculation expression of the accuracy coefficient.
7. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The determination of whether each piece of data is accurate specifically includes: Determine whether the accuracy score of each piece of data is greater than or equal to a preset threshold. If so, the corresponding data is accurate; if not, the corresponding data is inaccurate.
8. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The calculation of the consistency correction coefficient and the logic correction coefficient of the stored data specifically includes: Obtain the consistency coefficient C(x j ) and the logical coefficient L(x j ); Among them, x j represents the jth piece of data, where j represents the number of data; The support vector random algorithm is used to train the inaccurate data in the data set, generate a consistency correction function, and calculate the consistency correction coefficient. The calculation expression is: δ C (x j )=f C (x j )-C(x j ); In the formula, f C (x j ) represents the consistency correction value calculated by the support vector random algorithm, δ C (x j ) represents the consistency correction coefficient of the jth data, C(x j ) represents the consistency coefficient of the jth data; Use the support vector random algorithm to generate the logic correction function and calculate the logic correction coefficient. The calculation expression is: δ L (x j )=f L (x j )-L(x j ); In the formula, f L (x j ) represents the logical correction value calculated by the support vector random algorithm, L(x j ) represents the logical index of the jth data, δ L (x j ) represents the logical correction coefficient of the jth data.
9. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The correction of inaccurately stored data specifically includes: Based on the consistency correction coefficient and the logic correction coefficient, the consistency coefficient and the logic coefficient of each data are corrected respectively to obtain the corrected data consistency coefficient and the logic coefficient. The calculation expression is: C new (x j )=C(x j )+δ C (x j ); L new (x j )=L(x j )+δ L (x j ); In the formula, C new (x j ) represents the corrected data consistency coefficient, C(x j ) represents the consistency coefficient of the jth data, δ C (x j ) represents the consistency correction coefficient of the jth data, L new (x j ) represents the corrected data logic coefficient, L(x j ) represents the logical coefficient of the jth data, δ L (x j ) represents the logical correction coefficient of the j-th data.
10. The enterprise data optimization processing method based on cloud computing data fusion according to claim 1 is characterized in that: The comprehensive accuracy score of the entire data set is calculated based on the corrected stored data to determine the overall accuracy of the stored data, specifically including: According to the revised data consistency coefficient and logic coefficient, the accuracy coefficient of each data is recalculated respectively. The mean of the accuracy coefficients of all data is calculated through the mean calculation formula to obtain a comprehensive accuracy score. Based on the comparison between the comprehensive accuracy score and the preset threshold, it is re-determined whether the stored data is accurate after correction.
Citation Information
Patent Citations
Driving situation reasoning method based on metadata driving and causal analysis theory
CN117217314A
Application fusion system oriented to big data analysis
CN117331995A
Carbon asset management method and system based on knowledge graph
CN119067254A
Data quality evaluation method and system based on big data analysis
CN119271657A
Methods, systems, and storage media for information service of industrial internet of things (IIOT) based on cloud platforms
US20250039268A1
Cited By
Bill business online financing management system and method based on Internet platform
CN120543285A
Internet platform-based bill business online financing management system and method
CN120543285B