Enterprise data management system based on big data

By designing a big data-based enterprise data management system, the problem of ignoring the data flow nodes in the existing technology is solved, and efficient management and evaluation of enterprise data is achieved, and the accuracy and efficiency of enterprise decision-making are supported.

CN119940711AInactive Publication Date: 2025-05-06ZOUPING KEHUI INFORMATION CONSULTING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510002794.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the enterprise data management, the existing technology is easy to ignore the flow, evaluation and allocation of data on each node during the flow, resulting in the inability to discover problems in the data flow and reduce the efficiency and accuracy of enterprise decision-making.

Method used

A big data-based enterprise data management system is designed, including data reception module, data analysis module, node analysis module, achievement evaluation module and achievement evaluation module. Through these modules, enterprise data is preprocessed, analyzed, node consistency inspection, achievement evaluation and comprehensive evaluation report generation.

Benefits of technology

It realizes accurate reception, analysis and evaluation of enterprise data, can promptly discover bottlenecks and problems in data circulation, improve the efficiency and accuracy of data circulation, supports enterprises to scientifically and impartially evaluate the value of data results, and provides comprehensive support for decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940711A_ABST
    Figure CN119940711A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to an enterprise data management system based on big data, and the system comprises a data receiving module which is used for obtaining enterprise data uploaded by a data site; preprocessing the uploaded enterprise data according to the uploaded data site, and storing the preprocessed enterprise data; the data analysis module is used for analyzing the enterprise data, acquiring a plurality of processing nodes of the enterprise data in enterprise collaboration, determining circulation information of the enterprise data on the plurality of processing nodes, and obtaining a data evaluation result related to the circulation information; the node analysis module is used for checking consistency and collaboration among the processing nodes and acquiring node analysis results related to the processing nodes; the achievement evaluation module is used for evaluating achievements and projects in the enterprise data and obtaining an achievement distribution function related to enterprise data access; and the efficiency and the accuracy of data management are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular to an enterprise data management system based on big data. Background Art

[0002] With the rapid development of big data, enterprises have an increasing demand for data management. However, traditional data management systems often have many shortcomings, such as low data processing efficiency, poor data flow, and lack of scientific standards for results evaluation. These problems not only affect the efficiency and accuracy of enterprise decision-making, but also restrict the innovation and development of enterprises. Therefore, there is an urgent need for a new system that can comprehensively and accurately manage enterprise data.

[0003] For example, Chinese patent publication number CN115860679A discloses a data management system and method for enterprise projects based on big data, wherein the system includes a duration risk analysis module, which analyzes the correlation between various project nodes in different projects based on historical data, and obtains the duration risk values ​​corresponding to different project nodes in the enterprise project to be tested in combination with the workload corresponding to each associated project node. The present invention not only manages the project files in the enterprise project and the preset duration corresponding to each project file, but also takes into account the impact of the duration change of the project file on the duration of subsequent projects, and the self-adjustment ability during the project execution process, thereby realizing accurate prediction of the comprehensive duration of the enterprise project.

[0004] Existing technologies manage enterprise data by considering the relationship between nodes and construction period risks. However, this approach easily overlooks the flow, evaluation and allocation of enterprise data at each node during its flow, resulting in the inability to discover problems in the data flow, reducing the efficiency and accuracy of enterprise decision-making. Summary of the invention

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: an enterprise data management system based on big data, including: a data receiving module, used to obtain enterprise data uploaded by a data site; pre-processing the uploaded enterprise data according to the uploaded data site, and storing the pre-processed enterprise data.

[0006] The data analysis module is used to analyze enterprise data, obtain enterprise data in multiple processing nodes of enterprise collaboration, determine the flow information of enterprise data on multiple processing nodes, and obtain data evaluation results related to the flow information.

[0007] The node analysis module is used to check the consistency and coordination between various processing nodes and obtain node analysis results related to the processing nodes.

[0008] The achievement evaluation module is used to evaluate the achievements and projects in the enterprise data and obtain the achievement allocation functions related to the enterprise data access.

[0009] The results evaluation module is used to evaluate the operation status of enterprise data based on the results allocation function, data evaluation results, and node analysis results, and generate an evaluation report.

[0010] The beneficial effects of the present invention are: 1. The present invention can receive enterprise data from different data sites, and pre-process and store them according to the data sites; this helps to ensure the accuracy and completeness of the data, and provides a reliable basis for subsequent data analysis.

[0011] 2. The present invention can analyze enterprise data, obtain data flow information on multiple processing nodes, and generate data evaluation results; this helps enterprises to promptly discover and solve bottlenecks and problems in data flow, and improve the efficiency and accuracy of data flow.

[0012] 3. The present invention can check the consistency and coordination between various processing nodes and generate node analysis results, which helps enterprises understand the working status and coordination efficiency of each processing node and provides a basis for optimizing the data processing process.

[0013] 4. The present invention can evaluate the achievements and projects in the enterprise data and generate an achievement allocation function; this helps the enterprise to scientifically and impartially evaluate the value of data achievements and provide a basis for the reasonable allocation of achievements.

[0014] 5. The present invention can generate an evaluation report based on the outcome allocation function, data evaluation results and node analysis results, which helps enterprises to fully understand the overall situation of data management and provide support for decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0016] Figure 1 It is a system framework diagram of an enterprise data management system based on big data.

[0017] Figure 2 It is a flow chart of an enterprise data management system based on big data.

[0018] Figure 3 It is a flow chart of a node analysis module of an enterprise data management system based on big data.

[0019] Figure 4 It is a flow chart of the achievement evaluation module of an enterprise data management system based on big data. DETAILED DESCRIPTION

[0020] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. If no specific techniques or conditions are specified in the embodiments, the techniques or conditions described in the literature in the art or the product specifications are used.

[0021] See also Figure 1 , an enterprise data management system based on big data, including: a data receiving module, a data analysis module, a node analysis module, an achievement evaluation module, and an achievement judgment module.

[0022] The data receiving module is used to obtain the enterprise data uploaded by the data site, and perform preprocessing and storage; and send the stored data to the data analysis module; the data analysis module analyzes the enterprise data and transmits the analyzed data to the node analysis module; the node analysis module checks each processing node and transmits the checked data to the results evaluation module; the results evaluation module evaluates the results and projects in the enterprise data and outputs the evaluated values ​​to the results judgment module; the results judgment module conducts a comprehensive evaluation of the operation of the enterprise data and generates an evaluation report.

[0023] like Figure 2 As shown, the data receiving module is used to obtain the enterprise data uploaded by the data site; pre-process the uploaded enterprise data according to the uploaded data site, and store the pre-processed enterprise data.

[0024] The data analysis module is used to analyze enterprise data, obtain enterprise data in multiple processing nodes of enterprise collaboration, determine the flow information of enterprise data on multiple processing nodes, and obtain data evaluation results related to the flow information.

[0025] The node analysis module is used to check the consistency and coordination between various processing nodes and obtain node analysis results related to the processing nodes.

[0026] The achievement evaluation module is used to evaluate the achievements and projects in the enterprise data and obtain the achievement allocation functions related to the enterprise data access.

[0027] The results evaluation module is used to evaluate the operation status of enterprise data based on the results allocation function, data evaluation results, and node analysis results, and generate an evaluation report.

[0028] In one embodiment of the present invention, the data receiving module is used to obtain the enterprise data uploaded by the data site; pre-process the uploaded enterprise data according to the uploaded data site, and store the pre-processed enterprise data.

[0029] When acquiring enterprise data, the first thing to do is to determine the content corresponding to the enterprise. For example, enterprise data can be divided into project data, management data, audit data, and process data, etc. The enterprise data is divided according to the set classification, and then the enterprise data is preprocessed. The preprocessing method includes obtaining the business object corresponding to the enterprise data, identifying the first target section and the first target node corresponding to the business object; the first target section represents the content of the corresponding project in the enterprise data, and is used to obtain the project that needs to be identified and processed. The first target node is used to identify the data site where the enterprise uploads the data to determine the data uploaded by different sites at this time.

[0030] The enterprise data is extracted according to the first target plate, and the first target plate is centrally processed according to the address information of the current data site, and it is determined whether the enterprise data of the first target plate matches the data attribute corresponding to the first target node. When the data matches, the address of the first target plate is stored based on the first target node. Otherwise, the address index set for the first target plate is obtained, and the enterprise data of the first target plate is stored in the default storage area to complete the storage of the enterprise data.

[0031] At the same time, when the main enterprise data to be processed at this time is from multiple enterprises or multiple departments of an enterprise, when the enterprise data exists in multiple sources in a single project or multiple projects and the data in the sources are different, it is necessary to verify the enterprise data according to the uploaded data nodes, and store the corresponding data after verification. At the same time, distribution analysis and detection analysis of the enterprise data are also required to prevent corresponding problems in the current storage and processing of the enterprise data, and the multiple processing nodes existing in the enterprise data need to meet the preset standard relationship between projects to prevent conflicts and analysis errors in the currently identified enterprise data when the projects are centrally processed.

[0032] In one embodiment of the present invention, a data analysis module is used to analyze enterprise data, obtain enterprise data in multiple processing nodes of enterprise collaboration, determine the flow information of enterprise data on multiple processing nodes, and obtain data evaluation results related to the flow information.

[0033] The identification methods include node identification, analyzing the flow paths and processing nodes of enterprise data within the enterprise, ensuring the flow between various processing nodes, and the data transmission situation, and managing the processing tasks and permissions set for each processing node to analyze the current situation of enterprise data under collaborative control; finally, obtaining data evaluation results related to the objects existing in the enterprise data and the distribution of enterprise data. This result is used to indicate whether the current allocation management of enterprise data is correct, and to evaluate these distributions and settings.

[0034] The processing nodes at this time include: physical nodes, such as servers: physical servers in the data center, responsible for processing computing tasks, storing data, etc.; logical nodes, such as nodes responsible for executing batch data processing tasks, small nodes for each business function when executing enterprise management; business function nodes, such as order processing nodes, inventory management nodes, financial processing nodes and other function implementation nodes in business processes; at the same time, the flow path represents the process of data flowing between various processing nodes, and this process and the data of the processing nodes together constitute the flow information.

[0035] Therefore, when analyzing enterprise data and obtaining data evaluation results related to flow information, it is first necessary to identify the processing nodes and flow paths in the enterprise data, and obtain the flow information based on the processing nodes and flow paths, so as to determine whether the allocated storage corresponding to the current enterprise data corresponds.

[0036] For example, methods for obtaining data evaluation results include: analyzing the first correlation between each processing node based on historical data, the first correlation is used to obtain the correlation between projects implemented on the processing nodes; combining the working hours and working frequency of each processing node to obtain the second correlation corresponding to each processing node, the second correlation is used to describe the possible risk situations of each processing node when performing processing; recording the working hours corresponding to the first correlation and the second correlation, determining the flow prediction value on each processing node, selecting the flow path according to the obtained flow prediction value, and outputting the flow path as the data evaluation result.

[0037] In this implementation process, the flow information includes the corresponding data of the flow path and processing nodes. The use of flow information at this time is used to indicate that when the enterprise processes the relevant content, it needs to analyze its data storage nodes and the nodes for pre-analysis of the data to determine whether the overall data processing, storage and verification process is accurate, and regard this part of the verification as the flow information. At the same time, it is also necessary to obtain data evaluation results based on this flow information to complete the evaluation of the current data.

[0038] The above-mentioned first correlation is judged by extracting the features of the processing nodes and calculating the intersection of the processing nodes to determine the correlation in the current processing node; for example, the intersection between adjacent processing nodes is divided by the union to determine the first correlation between the processing nodes at this time, and the first correlation is traversed to obtain a group of processing nodes with the maximum value of the first correlation.

[0039] The above-mentioned second correlation is calculated by the failure probability of the processing node, that is, the failure probability of the processing node is compared with the average value of the failure probability of the processing node in the historical data, and is calculated using the Pearson correlation coefficient; at this time, it is also necessary to output a group of processing nodes with the largest second correlation value.

[0040] After obtaining two groups of processing nodes, the working time of the two groups of processing nodes is recorded to obtain the flow prediction value, that is, the first working time of the processing node corresponding to the maximum value of the first correlation and the second working time of the processing node corresponding to the maximum value of the second correlation are obtained to obtain the flow prediction value.

[0041] Among them, FR represents the flow prediction value, w1 represents the first correlation, w2 represents the second correlation, T1 represents the first working time, and T2 represents the second working time. When calculating the first working time and the second working time at this time, the first working time and the second working time are normalized so that the finally calculated flow prediction value can represent a more specific numerical value, and according to the calculated value of the flow prediction value, the flow prediction values ​​on all processing nodes are compared, and the flow path is selected according to the processing node corresponding to the maximum value of the flow prediction value, and the flow path is output as the data evaluation result at this time; selecting the flow path at this time is essentially selecting the flow path where this processing node is located, and using this flow path and corresponding data as the data evaluation result at this time, where the data evaluation result may include the flow prediction value, the processing node corresponding to the maximum value of the flow prediction value, the flow path and other data.

[0042] The purpose of using the flow prediction value for processing at this time is that the enterprise can effectively select the optimal data flow path to ensure efficient flow and consistent evaluation of data between various processing nodes; the processing node will obtain it from the corresponding flow path and determine whether the use of the flow path at this time is reasonable, so as to assist subsequent enterprises in improving processing efficiency when processing certain data.

[0043] In one embodiment of the present invention, the node analysis module is used to check the consistency and coordination between various processing nodes and obtain node analysis results related to the processing nodes.

[0044] At this time, verify each processing node, as well as the specific situation of the currently set enterprise data under the collaborative control of multiple targets, and the consistency of the collaborative scheme and access control set when the enterprise data is accessed, to obtain the node analysis results.

[0045] Specifically, the synergy between each processing node can be evaluated by the negative prediction rate and positive prediction rate of the processing nodes, and the consistency between each processing node can be evaluated by the unit capacity ratio and the unit reliability ratio, so as to determine which processing node performs poorly during the processing process, and which processing node resources are likely to be allocated to and thus reduce resource utilization, and at the same time evaluate whether the current processing node is prone to risks due to incorrect predictions.

[0046] This module mainly verifies whether the processing nodes set for each decision implementation in the process of enterprise management decision-making are reasonable. For example, the unit capacity ratio will be set according to the data existing in the processing node. The unit capacity ratio at this time indicates the proportion of a certain type of data processed by the current processing node to the total data, and the unit reliability ratio indicates the proportion of the normal operation time of the processing node to the total time in the entire project or data processing flow. At this time, the consistency of the processed data can be judged by extracting the features related to the processing node from the enterprise data and calculating the unit capacity ratio and unit reliability ratio corresponding to each feature.

[0047] The negative prediction rate and the positive prediction rate represent a binary classification problem. At this time, the processing process of the processing node is represented by binary classification. For example, data with defects is a positive class, and data without defects is a negative class. Or, the processing node has low efficiency as a positive class, and the processing node has high efficiency as a negative class. The negative prediction rate and the positive prediction rate represent the prediction probability at this time, completing the diagnosis of the current enterprise data to optimize resource allocation and reduce business risks.

[0048] like Figure 3 As shown, the implementation method of obtaining the node analysis results includes: obtaining the negative prediction rate, positive prediction rate, unit capacity ratio and unit reliability ratio of each processing node.

[0049] Based on the negative prediction rate and the positive prediction rate of each processing node, the synergy coefficient of the processing node with respect to synergy is obtained.

[0050] Based on the unit capacity ratio and unit reliability ratio of each processing node, the consistency coefficient of the processing node with respect to consistency is obtained.

[0051] The consistency coefficient and the synergy coefficient are combined to obtain the output node analysis results.

[0052] The synergy coefficient can be expressed as follows: obtain the confusion matrix corresponding to the processing node, the confusion matrix includes the number of samples TP that are actually positive and correctly predicted to be positive, the number of samples FN that are actually positive but incorrectly predicted to be negative, the number of samples FP that are actually negative but incorrectly predicted to be positive, and the number TN that are actually negative and correctly predicted to be negative; calculate the accuracy, recall and harmonic mean corresponding to the negative prediction rate and the positive prediction rate to obtain the synergy coefficient; at this time, the negative prediction rate represents the probability value of the result being the negative class, and the positive prediction rate represents the result of the positive class.

[0053] Among them, Accuracy represents the accuracy corresponding to the negative prediction rate and the positive prediction rate.

[0054] Among them, Recall represents the recall rate corresponding to the negative prediction rate and the positive prediction rate.

[0055] Among them, CM represents the harmonic mean of the negative prediction rate and the positive prediction rate, NPR represents the negative prediction rate, and PPV represents the positive prediction rate; the positive prediction rate is the ratio of the number of samples correctly predicted as positive to the total number of positive samples, and the negative prediction rate is the ratio of the number of samples correctly predicted as negative to the total number of negative samples.

[0056] Therefore, the synergy coefficient is expressed as follows.

[0057] Among them, χ1 represents the synergy coefficient, e represents the exponential constant, α1, α2, α3, and α4 represent weight coefficients, which can be set to 0.2, 0.3, 0.25, and 0.25 respectively.

[0058] At this point, we can know the relative comprehensive situation of the data on the current processing node after the binary classification process to determine whether the data set on the current processing node is reasonable.

[0059] The consistency coefficient can be expressed as taking the weighted average of the unit capacity ratio and the unit reliability ratio of the processing node as the consistency coefficient; the weights set at this time can be set to 0.6 and 0.4 in the order of the unit capacity ratio and the unit reliability ratio.

[0060] The method of combining the consistency coefficient and the synergy coefficient to obtain the output node analysis result is to compare the obtained consistency coefficient and synergy coefficient with the preset coefficient. When they are greater than the preset coefficient, the data corresponding to the consistency coefficient and the synergy coefficient are output as the output node analysis result.

[0061] The preset coefficients for the consistency coefficient and the synergy coefficient can be set to 0.8 and 0.9 respectively; and the node analysis result at this time also needs to output the composition of the corresponding processing nodes and the corresponding data on the processing nodes.

[0062] The consistency coefficient and the synergy coefficient can also be combined to obtain a comprehensive result as the output node analysis result. For example, the node analysis result can be expressed as: Among them, χ3 represents the node analysis result, χ1 represents the synergy coefficient, and χ2 represents the consistency coefficient. represents the weight coefficient of the synergy coefficient, It is represented as the weight coefficient of the consistency coefficient, χ′1 is represented as the standard value of the synergy coefficient, and the standard value of the synergy coefficient can adopt the standard value in the historical data; the weights of the consistency coefficient and the synergy coefficient can be set to 0.5 and 0.5 respectively.

[0063] In one embodiment of the present invention, the achievement evaluation module is used to evaluate the achievements and projects in the enterprise data and obtain the achievement allocation function related to the enterprise data access.

[0064] When evaluating outcomes and projects in enterprise data, outcomes usually refer to specific results or achievements achieved through specific projects or activities. These outcomes can be quantitative, such as financial indicators, or qualitative, such as customer satisfaction, depending on the nature and objectives of the project; at this time, projects are used to represent the grouping of current enterprise data to comprehensively evaluate the content in the enterprise data.

[0065] For example, the current enterprise data is separated by nodes, and the achievement allocation function is used to describe the achievement allocation to quantify the specific implementation of the current enterprise. The achievement allocation function will process the importance of the current processing node and each data in the processing node. The implementation method of the achievement allocation function is expressed as follows: obtain the importance of the processing node, take the corresponding data in the processing node as the independent variable, take the importance of the processing node as the dependent variable, the importance of the processing node is obtained by the number of processing nodes connected to each processing node, use linear regression analysis to analyze the changes in enterprise data, and mark the data with linear relationships in the enterprise data; calculate the linear regression coefficient of each processing node in the enterprise data, and use the t test to perform significance analysis on the linear regression coefficient to obtain the significance coefficient corresponding to each processing node; use the significance coefficient to judge the project in the enterprise data, obtain the strain parameter item related to the project, and combine the strain parameter item and the significance coefficient to obtain the achievement allocation function related to the enterprise data at this time. The achievement allocation function can be the product of the strain parameter item and the significance coefficient to indicate whether there are achievements in the current enterprise data and the specific situation of the corresponding achievements.

[0066] The linear regression coefficient is the weight of each independent variable in the linear regression model, which indicates the influence of the independent variable on the dependent variable. The main purpose of linear regression calculation is to find a set of parameters that minimizes the error between the importance of the processing node and the data of the processing node. At the same time, it is necessary to verify whether the linear regression coefficient is less than the preset significance level. At this time, the obtained linear regression coefficient can be verified by t-test. At this time, a value of 0.05 is generally used to judge the value tested during the t-test, and the value of the data with a significance level after the t-test is used as the significance coefficient; then the significance coefficient is compared with the items in the enterprise data to determine whether the significance coefficient exists in the corresponding project at this time. If so, the data item in the current project is used as the strain parameter item at this time, and finally the values ​​on these strain parameter items are multiplied and summed with the value of the significance coefficient to obtain the result distribution function; when performing the product summation, if the value of the calculated strain parameter item has a dimension, it is normalized so that the value calculated in the result distribution function obtained at this time is dimensionless and has a uniform numerical range.

[0067] In one embodiment of the present invention, the achievement evaluation module is used to evaluate the operation status of enterprise data based on the achievement allocation function, data evaluation results, and node analysis results, and generate an evaluation report.

[0068] In this module, the indicator values ​​corresponding to the outcome allocation function, data evaluation results, and node analysis results are combined in turn to obtain the relationship between the three contents, and when analyzing the processing nodes, when the number and corresponding positions of the processing nodes change, whether these three contents will produce corresponding difference values. Finally, these contents are summarized to form a comprehensive evaluation result, which is the generated evaluation report.

[0069] The index value of the outcome distribution function refers to the sum of the product of the value on the strain parameter item and the value of the significance coefficient. The index value of the node analysis result refers to the combined value of the consistency coefficient and the synergy coefficient. The index value of the data evaluation result is the maximum value of the flow prediction value on the corresponding flow path.

[0070] like Figure 4As shown, the implementation method of the evaluation report includes obtaining the index values ​​of the achievement allocation function, the data evaluation results, and the node analysis results; calculating the first correlation coefficient between the achievement allocation function and the data evaluation results, the second correlation coefficient between the achievement allocation function and the node analysis results, and the third correlation coefficient between the data evaluation results and the node analysis results in sequence; wherein the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient all use the Pearson correlation coefficient method to verify whether there is a linear relationship between the achievement allocation function, the data evaluation results, and the node analysis results; the first correlation coefficient is calculated by taking the index values ​​of the achievement allocation function and the data evaluation results as input, and comparing them with the average value of the index values ​​of the achievement allocation function and the data evaluation results in the historical data to obtain the first correlation coefficient at this time; the calculation method of the second correlation coefficient and the third correlation coefficient is the same as the calculation method of the first correlation coefficient, the second correlation coefficient takes the index values ​​of the achievement allocation function and the node analysis results as input, and the third correlation coefficient takes the index values ​​of the data evaluation results and the node analysis results as input, and finally compares them with the average value of the index values ​​of the achievement allocation function, the data evaluation results, and the node analysis results in the historical data to obtain the calculated value.

[0071] And obtain the difference between the first correlation coefficient, the second correlation coefficient, the third correlation coefficient and the standard correlation coefficient in the historical data, and calculate the evaluation report coefficient corresponding to the evaluation report; at this time, the standard correlation coefficient adopts the weighted average of the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient in the historical data; at this time, the weighted coefficient will be set to 0.3, 0.3, and 0.4 respectively according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient in the historical data to complete the weighted processing of the historical data.

[0072] According to the value range of the evaluation report coefficient, an evaluation report corresponding to the current management system is generated; at this time, the evaluation report will be described according to the value of the evaluation report coefficient, explaining the specific situation of the current enterprise data management, and the multiple values ​​calculated at this time will be output, thereby completing the management of the enterprise data.

[0073] The evaluation report coefficient can be expressed as: combining the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient to obtain a correlation coefficient set, and obtaining the difference between the correlation coefficient set and the standard correlation coefficient in the historical data to obtain the evaluation report coefficient.

[0074] Among them, τ C represents the evaluation report coefficient, r i represents the value of the i-th correlation coefficient in the correlation coefficient set. The values ​​of i are 1, 2, and 3, corresponding to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, respectively. i′ represents the difference between the i-th correlation coefficient in the correlation coefficient set and the standard correlation coefficient in the historical data, and e represents the exponential constant; when calculating the evaluation report coefficient, the value range of the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient is (0, 1); at this time, the denominator will not be 0.

[0075] After obtaining the evaluation report coefficient, the current situation can be described based on the values ​​corresponding to the evaluation report coefficient and multiple output results. For example, the evaluation report coefficient can show the comprehensive situation of the current evaluation. When the evaluation report coefficient is large, it means that the overall evaluation of the system is relatively consistent and stable. When the value of the achievement allocation function is small, it means that there are deficiencies in the achievement evaluation, and the set processing nodes need to be analyzed separately. By judging the values ​​of these numerical values, the management accuracy and efficiency of enterprise data can be improved.

[0076] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention and they are still covered by the protection scope of the present invention.

Claims

1. An enterprise data management system based on big data, characterized in that: include: Data receiving module, used to obtain enterprise data uploaded by data sites; Pre-process the uploaded enterprise data according to the uploaded data site, and store the pre-processed enterprise data; The data analysis module is used to analyze the enterprise data, obtain the enterprise data in multiple processing nodes of the enterprise collaboration, determine the flow information of the enterprise data on multiple processing nodes, and obtain the data evaluation results related to the flow information; The node analysis module is used to check the consistency and coordination between various processing nodes and obtain the node analysis results related to the processing nodes; The achievement evaluation module is used to evaluate the achievements and projects in the enterprise data and obtain the achievement allocation function related to the enterprise data access; The results evaluation module is used to evaluate the operation status of enterprise data based on the results allocation function, data evaluation results, and node analysis results, and generate an evaluation report.

2. The enterprise data management system based on big data according to claim 1, characterized in that: The preprocessing method includes obtaining the business object corresponding to the enterprise data, identifying the first target section and the first target node corresponding to the business object; The enterprise data is extracted according to the first target plate, and the first target plate is centrally processed according to the address information of the current data site, and it is determined whether the enterprise data of the first target plate matches the data attribute corresponding to the first target node. When the data matches, the address of the first target plate is stored based on the first target node. Otherwise, the address index set for the first target plate is obtained, and the enterprise data of the first target plate is stored in the default storage area to complete the storage of the enterprise data.

3. The enterprise data management system based on big data according to claim 1, characterized in that: The methods for obtaining data evaluation results include: According to historical data, the first correlation between each processing node is analyzed, and the first correlation is used to obtain the correlation between the projects implemented on the processing nodes; in combination with the working time and working frequency of each processing node, the second correlation corresponding to each processing node is obtained, and the second correlation is used to describe the possible risk situation of each processing node when processing; the working time corresponding to the first correlation and the second correlation is recorded, the flow prediction value on each processing node is determined, the flow path is selected according to the obtained flow prediction value, and the flow path is output as the data evaluation result.

4. The enterprise data management system based on big data according to claim 3, characterized in that: The flow prediction value is expressed as: obtaining the first working time of the processing node corresponding to the maximum value of the first correlation and the second working time of the processing node corresponding to the maximum value of the second correlation, and obtaining the flow prediction value; Among them, FR represents the flow prediction value, w1 represents the first correlation, w2 represents the second correlation, T1 represents the first working time, and T2 represents the second working time.

5. The enterprise data management system based on big data according to claim 1, characterized in that: The implementation methods of node analysis results include: Obtain the negative prediction rate, positive prediction rate, unit capacity ratio and unit reliability ratio of each processing node; Based on the negative prediction rate and the positive prediction rate of each processing node, a synergy coefficient of the processing node with respect to synergy is obtained; Based on the unit capacity ratio and unit reliability ratio of each processing node, the consistency coefficient of the processing node with respect to consistency is obtained; The consistency coefficient and the synergy coefficient are combined to obtain the output node analysis results.

6. The enterprise data management system based on big data according to claim 5, characterized in that: The synergy coefficient can be expressed as follows: obtain the confusion matrix corresponding to the processing node, which includes the number of samples TP that are actually positive and correctly predicted as positive, the number of samples FN that are actually positive but incorrectly predicted as negative, the number of samples FP that are actually negative but incorrectly predicted as positive, and the number TN that are actually negative and correctly predicted as negative; calculate the accuracy, recall, and harmonic mean corresponding to the negative prediction rate and the positive prediction rate to obtain the synergy coefficient; Among them, χ1 represents the synergy coefficient, e represents the exponential constant, α1, α2, α3, and α4 represent weight coefficients; NPR represents the negative prediction rate, PPV represents the positive prediction rate; Accuracy represents the accuracy corresponding to the negative prediction rate and the positive prediction rate; Recall represents the recall corresponding to the negative prediction rate and the positive prediction rate; CM represents the harmonic mean corresponding to the negative prediction rate and the positive prediction rate; The consistency coefficient is expressed as the weighted average of the unit capacity ratio and the unit reliability ratio of the processing nodes as the consistency coefficient.

7. The enterprise data management system based on big data according to claim 6, characterized in that: The node analysis results can be expressed as: Among them, χ3 represents the node analysis result, χ1 represents the synergy coefficient, and χ2 represents the consistency coefficient. represents the weight coefficient of the synergy coefficient, It is expressed as the weight coefficient of the consistency coefficient, and χ′1 is expressed as the standard value of the synergy coefficient.

8. The enterprise data management system based on big data according to claim 1, characterized in that: The implementation method of the achievement allocation function is as follows: obtain the importance of the processing node, take the corresponding data in the processing node as the independent variable, take the importance of the processing node as the dependent variable, calculate the linear regression coefficient of each processing node in the enterprise data, and use the t-test to perform significance analysis on the linear regression coefficient to obtain the significance coefficient corresponding to each processing node; The significance coefficient is used to judge the items in the enterprise data to obtain the strain parameter items related to the items. The result allocation function related to the enterprise data at this time is obtained by combining the strain parameter items and the significance coefficient.

9. The enterprise data management system based on big data according to claim 1, characterized in that: The implementation method of the evaluation report includes obtaining the index values ​​of the achievement allocation function, the data evaluation result, and the node analysis result; sequentially calculating the first correlation coefficient between the achievement allocation function and the data evaluation result, the second correlation coefficient between the achievement allocation function and the node analysis result, and the third correlation coefficient between the data evaluation result and the node analysis result; Obtain the difference between the first correlation coefficient, the second correlation coefficient, the third correlation coefficient and the standard correlation coefficient in the historical data, and calculate the evaluation report coefficient corresponding to the evaluation report; According to the value range of the evaluation report coefficient, an evaluation report corresponding to the current management system is generated.

10. The enterprise data management system based on big data according to claim 9, characterized in that: The evaluation report coefficient can be expressed as: combining the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient to obtain a correlation coefficient set, and obtaining the difference between the correlation coefficient set and the standard correlation coefficient in the historical data to obtain the evaluation report coefficient; Among them, τ C represents the evaluation report coefficient, r i represents the value of the i-th correlation coefficient in the correlation coefficient set. The values ​​of i are 1, 2, and 3, corresponding to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, respectively. i ′ represents the difference between the i-th correlation coefficient in the correlation coefficient set and the standard correlation coefficient in the historical data, and e represents the exponential constant.