Enterprise work data flow path monitoring and evaluation method

By constructing a baseline link and calculating multi-dimensional feature parameters, the problem of the evaluation results being out of touch with business scenarios in existing technologies has been solved. This enables a comprehensive and accurate evaluation of the data flow path of enterprise operations, improving the efficiency of path optimization and anomaly tracing.

CN121967268AInactive Publication Date: 2026-05-01BEIJING QIZHONG SOFTWARE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING QIZHONG SOFTWARE TECHNOLOGY CO LTD
Filing Date
2026-02-08
Publication Date
2026-05-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies fail to fully consider the inherent characteristics of each node when assessing the data flow path of enterprise work, resulting in a disconnect between the assessment results and the business scenario. They cannot accurately reflect the actual flow capacity of the nodes, and cannot accurately assess the operational status of different flow stages during statistics.

Method used

By constructing a baseline link for data flow, we extract characteristic parameters such as data transmission rate, data storage capacity, number of personnel, work efficiency, and data task complexity of each node, calculate the receiving flow efficiency and processing flow efficiency, and combine the similarity of similar nodes to calculate the deviation between the expected receiving time and processing time, thus obtaining the data flow path evaluation result.

Benefits of technology

It enables a comprehensive and accurate assessment of node flow capabilities, improves the relevance and accuracy of assessment results, provides reliable data support for path optimization and anomaly tracing, and enhances the efficiency of enterprise work data flow path management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967268A_ABST
    Figure CN121967268A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data flow path monitoring and evaluation, and particularly relates to an enterprise work data flow path monitoring and evaluation method, which comprises the following steps of: firstly, calculating receiving flow efficiency and processing flow efficiency according to multi-dimensional characteristic parameters and receiving or processing duration of each node; secondly, calculating predicted receiving duration and predicted processing duration of the current node in combination with the transfer efficiency of each node, calculating similarity between the current node and each node through the characteristic parameters, screening out similar nodes, finally quantifying a deviation value between the predicted duration and a similar node mean value, and obtaining an evaluation result in combination with the characteristic parameters. According to the method, through multi-dimensional feature fusion, similar node reference benchmark construction and accurate deviation quantification, the actual flow capability of the nodes is comprehensively reflected, the evaluation accuracy is improved, reliable data support is provided for data flow path optimization and abnormal traceability, and the enterprise data flow management efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of data flow path monitoring and evaluation, and specifically relates to a method for monitoring and evaluating enterprise work data flow paths. Background Technology

[0002] The flow of enterprise work data runs through the entire business chain, including procurement, production, sales, finance, and human resources. The accuracy of the data flow path determines the efficiency of business operations. In order to improve the efficiency of business operations, it is necessary to monitor and evaluate the data flow path.

[0003] Existing data flow path assessments mostly consider technical performance and process timeliness indicators of each node, such as transmission delay, processing time, and storage usage. They do not take into account the inherent characteristics of each node in the data flow path, resulting in various indicators being monitored and evaluated separately. This makes it impossible to accurately reflect the actual flow capacity of the nodes, leading to a disconnect between the assessment results and the business scenario, and failing to provide data support for subsequent path optimization and tracing the cause of anomalies.

[0004] Furthermore, existing technologies often rely on average or instantaneous data for judgment during the statistical reception and processing stages, which makes it impossible to comprehensively and accurately assess the operational status of nodes at different stages of the process and thus hinders subsequent path optimization. Summary of the Invention

[0005] In view of this, in order to solve the above problems, a method for monitoring and evaluating the flow path of enterprise work data is proposed.

[0006] The objective of this invention can be achieved through the following technical solution: This invention provides a method for monitoring and evaluating the data flow path of an enterprise. The method includes: constructing a baseline link for data flow based on the enterprise data flow nodes, extracting the data transmission rate, data storage capacity, number of personnel, work efficiency and data task complexity of each node, and recording them as feature parameters, synchronously extracting the receiving time and processing time of the corresponding node, and calculating the receiving flow efficiency and processing flow efficiency respectively.

[0007] Extract the feature parameters of the current node, and combine them with the receiving and processing efficiency of each node to calculate the expected receiving time and the expected processing time.

[0008] Based on the feature parameters, the similarity between the current node and each other is calculated, and the similarity of each node is stratified to obtain the similar nodes of the current node. The average reception time and average processing time of the similar nodes are also calculated.

[0009] The deviations between the expected reception time and the average reception time, and between the expected processing time and the average processing time are calculated separately. The maximum value is selected as the final deviation value of the current node. Combined with the characteristic parameters of the current node, the data flow path evaluation result of the current node is obtained.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention combines data transmission rate, data storage capacity, number of personnel, work efficiency and data task complexity, reception time and processing time to calculate reception flow efficiency and processing flow efficiency respectively, thereby comprehensively and accurately reflecting the actual flow capacity of nodes in the reception and processing links, making the evaluation results fit the business scenario of enterprise work data flow, and providing a reliable data foundation for subsequent path optimization and anomaly tracing.

[0011] (2) By calculating the expected reception time and the expected processing time separately, and combining the average reception time and the average processing time of similar nodes, the present invention accurately quantifies the deviation value of each node in different stages of reception and processing, improves the pertinence and accuracy of monitoring and evaluation, and effectively supports the subsequent optimization of data flow path.

[0012] (3) This invention calculates node similarity based on multi-dimensional feature parameters, selects similar nodes that match the current node flow characteristics as reference benchmarks, calculates the final deviation value and obtains the data flow path evaluation result, and can further locate abnormal feature parameters, thereby providing effective data support for path optimization and anomaly tracing, and improving the efficiency and practicality of enterprise work data flow path management. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram of the overall implementation process of the present invention.

[0015] Figure 2 This is a schematic diagram illustrating the calculation process for the receiving and circulating efficiency of this invention.

[0016] Figure 3 This is a schematic diagram illustrating the similarity calculation process of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] For details, please refer to [link / reference]. Figure 1 As shown, this invention provides a method for monitoring and evaluating the data flow path of an enterprise. The method includes: S1. Constructing a baseline link for data flow based on the enterprise's data flow nodes. Based on the enterprise's historical operational data for the past T days, and using H hours as the extraction period, extracting in real time the data transmission rate, data storage capacity, number of personnel, work efficiency, and data task complexity of each node, and recording these as feature parameters. Simultaneously extracting the receiving duration and processing duration of the corresponding nodes, and calculating the receiving flow efficiency and processing flow efficiency respectively. Wherein, T and H are both positive integers, T is greater than or equal to 7 days and less than or equal to 90 days, and H is greater than or equal to 1 hour and less than or equal to 24 hours. For example, when the business data is the enterprise's daily sales data, T is 15 days and H is 4 hours; when the business data is the enterprise's quarterly financial data, T is 60 days and H is 24 hours.

[0019] Specifically, the process of constructing the baseline link includes: extracting data names, data types, storage locations, data ownership, business types, staff configuration information, and basic data on data transmission and processing from various business systems of the enterprise, and integrating them to form a data asset list.

[0020] Based on the data asset inventory, data assets are classified according to the enterprise's business process type. The corresponding flow nodes for each category, from generation, transmission, processing, storage, sharing to destruction, are extracted. A baseline link is constructed based on the sequential interaction order of these flow nodes. Preferably, the aforementioned flow nodes are nodes with a data interaction frequency of three times or more per day.

[0021] Specifically, the calculation process of the receiving flow efficiency includes: counting the number of data received by each node and the single data storage capacity corresponding to each data, calculating the sum of all single data storage capacities to obtain the total storage capacity, and performing minimum-maximum linear normalization processing on the total storage capacity and the number of received data to obtain the baseline occupancy rate and number correction coefficient of the total storage capacity.

[0022] It should be added that the maximum and minimum values ​​of the total storage capacity and the number of received data are selected from the total storage capacity and the number of received data of all nodes, and are used as the maximum and minimum values ​​required for the normalization, respectively. The minimum-maximum linear normalization process is an existing technology and will not be described in detail in this invention.

[0023] The baseline occupancy rate is adjusted using a number correction factor to obtain the final occupancy rate. This final occupancy rate is then multiplied by a preset adjustment factor threshold and the data transmission rate to obtain the actual data transmission rate. The above adjustment involves multiplying the number correction factor by the baseline occupancy rate.

[0024] It should be noted that the specific process for obtaining the above-mentioned preset adjustment coefficient threshold is as follows: First, multiple sets of historical data are collected. Each set of data includes the baseline occupancy rate, the number correction coefficient, the data transmission rate, and the actual reasonable transmission rate. Based on each set of historical data, the adjustment coefficient of each set of data is calculated.

[0025] The formula for calculating the above adjustment coefficient is as follows: In the formula, To adjust the coefficient, Based on the baseline occupancy rate, This is a number correction factor. For data transmission rate, For a practical and reasonable transmission rate, This represents the theoretical baseline transmission capacity after load correction under the current data load, data transmission speed, and storage occupancy conditions. A higher actual reasonable transmission rate results in a lower baseline occupancy rate and a smaller number of correction coefficients; conversely, a lower data transmission rate results in a larger adjustment coefficient.

[0026] Then, the 3σ criterion is used to calculate the mean and standard deviation of all adjustment coefficients. Values ​​that exceed the range of mean ± 3 times the standard deviation are identified as abnormal adjustment coefficients and removed. Finally, the average value of the remaining adjustment coefficients is calculated as the preset adjustment coefficient threshold. For example, the preset adjustment coefficient threshold is 0.8.

[0027] Calculate the ratio of the actual data transmission rate to the reception time, and perform minimum-maximum linear normalization to obtain the reception throughput efficiency.

[0028] It should be added that the selection method for the maximum and minimum values ​​required to normalize the ratio of the actual data transmission rate to the reception time is the same as the selection method for the maximum and minimum values ​​required to normalize the total storage capacity.

[0029] Specifically, please refer to Figure 2As shown, the calculation process of the processing efficiency includes: performing minimum-maximum linear normalization on the number of personnel, work efficiency and data task complexity respectively, to obtain the personnel coefficient, the per capita efficiency coefficient and the complexity coefficient in turn.

[0030] The selection method for the maximum and minimum values ​​required to normalize the above-mentioned number of personnel, work efficiency, and data task complexity is the same as the selection method for the maximum and minimum values ​​required to normalize the total storage capacity.

[0031] The unit load processing capacity coefficient is calculated based on the personnel coefficient, the per capita efficiency coefficient, the complexity coefficient, and the final occupancy rate.

[0032] The formula for calculating the unit load handling capacity coefficient is: In the formula, This is the unit load processing capacity coefficient. For personnel coefficient, The efficiency coefficient per capita. The complexity coefficient is... For the final occupancy rate, This represents the space that adapts to the processing capabilities of data processing tasks. This represents the basic processing potential of a node in terms of personnel configuration, per capita efficiency, and task complexity. The larger the personnel coefficient, the larger the per capita efficiency coefficient, and the smaller the complexity coefficient, the smaller the final occupancy rate, and the larger the unit load processing capacity coefficient.

[0033] The ratio of unit load processing capacity coefficient to processing time is calculated, standardized, and denoted as processing flow coefficient. Then, it is subjected to minimum-maximum linear normalization to obtain processing flow efficiency.

[0034] It should be added that the above standardization process is implemented using the Z-score standardization method. The purpose is to remove the dimension from the ratio of the unit load processing capacity coefficient to the processing time. The Z-score standardization method is existing technology and will not be described in detail in this invention.

[0035] The method for selecting the maximum and minimum values ​​required to normalize the processing turnover coefficient is the same as the method for selecting the maximum and minimum values ​​required to normalize the total storage capacity.

[0036] The transmission rate and data storage capacity directly determine whether data can arrive at the node in a timely and stable manner. The number of personnel, work efficiency, and data task complexity directly determine whether data can be efficiently processed and analyzed. If only a single efficiency indicator is used for evaluation, it is impossible to distinguish the operating status of the two links. Therefore, this invention incorporates multiple dimensions such as data transmission rate, storage capacity, number of personnel, work efficiency, task complexity, duration, and load into the calculation of receiving and processing efficiency, thereby comprehensively covering the flow links and making the evaluation results more consistent with the actual operating status of the node, avoiding misjudgments such as high transmission rate and low efficiency or low transmission speed and high efficiency.

[0037] S2. Extract the feature parameters of the current node, and calculate the expected reception time and expected processing time by combining the reception flow efficiency and processing flow efficiency of each node.

[0038] Specifically, the calculation process of the expected reception time includes: calculating the actual data transmission rate and final occupancy rate of the current node in the same way as calculating the actual data transmission rate and final occupancy rate of each node.

[0039] Based on historical enterprise data, extract the historical reception duration of each time series of the current node to form a set of historical reception durations for that node, and calculate the average value of all historical reception durations in that set.

[0040] For each two adjacent time points in the time series, the relative deviation value of the later time point relative to the previous time point is calculated to obtain multiple relative deviation values ​​of reception duration. Each relative deviation value is used as the rate of change of reception duration. After outlier removal of the rate of change of reception duration using the 3σ criterion, the average value of the remaining rate of change of reception duration is calculated as the duration correction coefficient. The historical average reception duration is then corrected to obtain the final average reception duration.

[0041] It should be noted that the specific process of outlier removal includes: first, calculating the mean and standard deviation of all reception duration change rates, and then identifying values ​​that exceed the mean ± 3 times the standard deviation as outliers and removing them.

[0042] It should be noted that the formula for correcting the final average reception time is as follows: In the formula, This represents the final average reception time. This represents the historical average reception duration. This is the duration correction factor.

[0043] The ratio of the actual data transmission rate to the average reception time is standardized to obtain the baseline reception flow coefficient of the current node, and then subjected to minimum-maximum linear normalization to obtain the baseline reception flow efficiency.

[0044] The standardization process for the ratio of the actual data transmission rate to the average reception time is the same as the standardization process for the ratio of the unit load processing capacity coefficient to the processing time.

[0045] The selection method for the maximum and minimum values ​​required to normalize the reference receiving flow coefficient is the same as the selection method for the maximum and minimum values ​​required to normalize the total storage capacity.

[0046] The expected reception time is calculated based on the actual data transmission rate of the current node, the baseline reception flow efficiency and final occupancy rate, and the reception flow coefficient of each node.

[0047] The formula for calculating the expected reception time is: In the formula, For the expected reception time, This represents the final occupancy rate of the current node. To preset the adjustment coefficient threshold, This represents the actual data transmission rate of the current node. This serves as the baseline receiving and processing efficiency for the current node. This is the maximum value among all the receiving flow coefficients of all nodes. It is the minimum value among all the receiving flow coefficients of all nodes. This represents the actual transmittance rate that can be carried after being dynamically constrained by the final resource occupancy rate of the current node. This represents the maximum fluctuation range of the data reception and transfer coefficient across all nodes, reflecting the overall fluctuation range of the enterprise's data reception and transfer efficiency. The current node's actual reception throughput efficiency is calibrated after the maximum fluctuation amplitude. The higher the final occupancy rate of the current node, the higher the actual data transmission rate of the current node, the lower the baseline reception throughput efficiency of the current node, the higher the maximum fluctuation amplitude of the reception throughput coefficient of all nodes, and the lower the minimum value of the reception throughput coefficient of all nodes, the longer the expected reception time.

[0048] This invention directly represents the receiving efficiency of a node per unit time by calculating the ratio of the actual data transmission rate to the average historical reception duration with dynamic correction. This avoids interference from instantaneous fluctuations and abnormal data, making the calculated baseline receiving efficiency more consistent with the actual receiving capability of the node and improving the accuracy of the evaluation results.

[0049] Specifically, the calculation process of the expected processing time includes: calculating the unit load processing capacity coefficient and the benchmark processing efficiency of the current node in the same way as calculating the unit load processing capacity coefficient of each node and the benchmark receiving flow efficiency of the current node.

[0050] The estimated processing time is calculated based on the unit load processing capacity coefficient of the current node, the baseline processing efficiency, and the processing efficiency coefficient of each node.

[0051] The formula for calculating the expected processing time is: In the formula, To estimate processing time, This represents the unit load processing capacity coefficient of the current node. This represents the baseline processing throughput efficiency for the current node. This is the maximum value among all the processing flow coefficients of all nodes. This is the minimum processing flow coefficient among all nodes. To handle the maximum fluctuation amplitude of the turnover coefficient for all nodes, This represents the actual processing efficiency of the current node after calibration for the maximum fluctuation amplitude. The larger the unit load processing capacity coefficient of the current node, the smaller the baseline processing efficiency of the current node. The smaller the maximum fluctuation amplitude of the processing efficiency coefficient of all nodes, and the smaller the minimum value among the processing efficiency coefficients of all nodes, the longer the expected processing time.

[0052] S3. Based on the feature parameters, calculate the similarity between the current node and each node, and perform stratified processing on the similarity of each node to obtain the similar nodes of the current node, and calculate the average reception time and average processing time of the similar nodes.

[0053] Specifically, please refer to Figure 3 As shown, the similarity calculation process includes: recording the baseline receiving flow efficiency and baseline processing flow efficiency of the current node as the receiving flow efficiency and processing flow efficiency of the current node, respectively.

[0054] The data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, receiving and processing efficiency of the current node and each other are combined to form a feature parameter dataset.

[0055] For each feature parameter in the feature parameter dataset, perform min-max linear normalization to obtain the standardized feature parameters of the current node and all other nodes.

[0056] The method for selecting the maximum and minimum values ​​required to normalize each feature parameter in the feature parameter dataset is the same as the method for selecting the maximum and minimum values ​​required to normalize the total storage capacity.

[0057] Using seven standardized feature parameters as dimensions, and following the order of data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, receiving and processing efficiency, feature vectors of the current node and each node are constructed respectively.

[0058] The cosine similarity algorithm is used to calculate the cosine similarity value between the feature vector of the current node and the feature vectors of each node, thus obtaining the similarity between the current node and each node.

[0059] The formula for calculating cosine similarity is: In the formula, This is the feature vector of the current node. For each node's feature vector, For vector dot product, Let be the vector magnitude of the feature vector of the current node. Let be the vector magnitude of the feature vectors of each node. The range of the cosine similarity value is [-1, 1]. The closer the cosine similarity is to 1, the more similar the feature parameters of the current node are to the corresponding nodes. Conversely, the greater the difference, the less similar the cosine similarity is to 1.

[0060] Specifically, the calculation process of the similar nodes includes: marking nodes whose similarity is greater than or equal to a preset similarity stratification threshold as similar nodes.

[0061] The calculation process of the preset similarity stratification threshold is as follows: collect node similarity samples under different loads and task types, calculate the mean and standard deviation of the similarity samples, and use the 3σ criterion to calculate the preset similarity stratification threshold. For example, the preset similarity stratification threshold is 0.8.

[0062] If no nodes have a similarity greater than or equal to a preset similarity stratification threshold, then the top few nodes with the highest similarity are selected as similar nodes. Preferably, the top few nodes are three or more.

[0063] To comprehensively reflect the data flow performance of enterprise work nodes, seven characteristics are used: data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, receiving and processing efficiency, and processing efficiency. The role of similar nodes is to provide a reference benchmark that matches the flow characteristics of the current node. Only when the reference node and the current node are similar in multiple dimensions can their average receiving and processing times reflect the flow level of the current node. This avoids distortion in subsequent deviation calculations or misjudgments due to significant differences between the characteristics of the reference node and the current node.

[0064] This invention calculates flow deviation based on reference benchmarks of similar nodes, making the anomaly judgment results more consistent with the actual flow level of the current node. At the same time, by screening similar nodes through multi-dimensional feature matching, it can provide data support for subsequent anomaly tracing and path optimization.

[0065] S4. Calculate the deviation between the expected reception time and the average reception time, and the deviation between the expected processing time and the average processing time, respectively. Select the maximum value as the final deviation value of the current node. Combine the characteristic parameters of the current node to obtain the data flow path evaluation result of the current node.

[0066] Specifically, the calculation process of the final deviation value includes: calculating the absolute deviation between the expected reception duration and the average reception duration to obtain the reception deviation value.

[0067] The processing deviation value is calculated similarly to the method used to calculate the receiving deviation value.

[0068] The larger of the received deviation value and the processed deviation value is taken as the final deviation value of the current node.

[0069] Specifically, the process of obtaining the data flow path evaluation result includes: if the final deviation value of the current node is less than the preset deviation threshold, then the data flow path of the current node is determined to be running normally.

[0070] The process of obtaining the preset deviation threshold is as follows: First, samples of the final deviation values ​​of the node under different loads and time periods are collected. Then, the mean and standard deviation of the samples are calculated. Finally, the preset deviation threshold is calculated using the 3σ criterion. For example, the preset deviation threshold is 45 milliseconds.

[0071] If the final deviation value of the current node is greater than or equal to the preset deviation threshold, the data flow path of the current node is determined to be abnormal.

[0072] The correlation coefficients between data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, and final deviation value were calculated using the Pearson correlation coefficient.

[0073] The formula for calculating the Pearson correlation coefficient is as follows: In the formula, The Pearson correlation coefficient is used. For the first The values ​​of the corresponding feature parameters in the group of samples. For the first The final deviation value in the group sample is taken as follows. The mean of the feature parameter samples. The mean of the final deviation values ​​is... Given the total number of samples, the Pearson correlation coefficient ranges from [-1, 1]. A value greater than 0 indicates that the characteristic parameter is positively correlated with the final deviation value. When the value is less than 0, it indicates that the characteristic parameter is negatively correlated with the final deviation value. When the value is equal to 0, it means that there is no linear correlation between the characteristic parameter and the final deviation value. The closer the absolute value of the correlation coefficient between a certain characteristic parameter and the final deviation value is to 1, the stronger the linear correlation between the characteristic parameter and the final deviation value, and the greater the influence of the characteristic parameter on the final deviation value.

[0074] Feature parameters with a correlation coefficient of 0 are removed. The remaining feature parameters are then sorted in ascending order of their absolute correlation coefficient values. The last few selected feature parameters are denoted as abnormal feature parameters. Preferably, in this invention, the number of these selected parameters is three.

[0075] The normal judgment result, or the combined abnormal judgment result and abnormal characteristic parameters, will be used as the evaluation result of the data flow path of the current node.

[0076] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A method for monitoring and evaluating the data flow path in enterprise operations, characterized in that, The method includes: Based on the enterprise data flow nodes, a baseline link for data flow is constructed. The data transmission rate, data storage capacity, number of personnel, work efficiency and data task complexity of each node are extracted and recorded as feature parameters. The receiving time and processing time of the corresponding nodes are extracted synchronously, and the receiving flow efficiency and processing flow efficiency are calculated respectively. Extract the feature parameters of the current node, and calculate the expected reception time and expected processing time by combining the reception flow efficiency and processing flow efficiency of each node. Based on the feature parameters, the similarity between the current node and each other is calculated, and the similarity of each node is processed in layers to obtain the similar nodes of the current node. The average reception time and average processing time of the similar nodes are also calculated. The deviations between the expected reception time and the average reception time, and between the expected processing time and the average processing time are calculated separately. The maximum value is selected as the final deviation value of the current node. Combined with the characteristic parameters of the current node, the data flow path evaluation result of the current node is obtained.

2. The method for monitoring and evaluating the flow path of enterprise work data as described in claim 1, characterized in that: The process of constructing the baseline link includes: Extract data names, data types, storage locations, data ownership, business types, staff configuration information, and basic data on data transmission and processing from various business systems of the enterprise, and integrate them to form a data asset list; Based on the data asset list, data assets are classified according to the enterprise's business process type, and the corresponding flow nodes of data assets in each category from generation, transmission, processing, storage, sharing to destruction are extracted. Based on the interaction order of the flow nodes, a baseline link is constructed.

3. The enterprise work data flow path monitoring and evaluation method as described in claim 1, characterized in that: The calculation process for the receiving and transfer efficiency includes: The number of data received by each node and the single data storage capacity corresponding to each data are counted. The sum of the single data storage capacities of all data is calculated to obtain the total storage capacity. The total storage capacity and the number of received data are subjected to minimum-maximum linear normalization to obtain the baseline occupancy rate and number correction coefficient of the total storage capacity. The baseline occupancy rate is corrected by a number correction factor to obtain the final occupancy rate. The final occupancy rate is then multiplied by the preset adjustment factor threshold and the data transmission rate to obtain the actual data transmission rate. Calculate the ratio of the actual data transmission rate to the reception time, and perform minimum-maximum linear normalization to obtain the reception throughput efficiency.

4. The enterprise work data flow path monitoring and evaluation method as described in claim 3, characterized in that: The calculation process for the processing efficiency includes: The number of personnel, work efficiency, and data task complexity are respectively subjected to minimum-maximum linear normalization to obtain the personnel coefficient, per capita efficiency coefficient, and complexity coefficient. The unit load processing capacity coefficient is calculated based on the personnel coefficient, the per capita efficiency coefficient, the complexity coefficient, and the final occupancy rate. The formula for calculating the unit load handling capacity coefficient is: In the formula, This is the unit load processing capacity coefficient. For personnel coefficient, The efficiency coefficient per capita. The complexity coefficient is... This represents the final occupancy rate. The ratio of unit load processing capacity coefficient to processing time is calculated, standardized, and denoted as processing flow coefficient. Then, it is subjected to minimum-maximum linear normalization to obtain processing flow efficiency.

5. The enterprise work data flow path monitoring and evaluation method as described in claim 3, characterized in that: The calculation process for the expected reception duration includes: Similarly, the actual data transmission rate and final occupancy rate of the current node are calculated using the same method as the calculation of the actual data transmission rate and final occupancy rate of each node. Based on historical enterprise data, extract each historical reception duration under the time series of the current node to form a set of historical reception durations for that node, and calculate the average value of all historical reception durations under that set. For each two adjacent time points in the time series, the relative deviation value of the later time point relative to the previous time point is calculated to obtain multiple relative deviation values ​​of reception duration. Each relative deviation value is used as the rate of change of reception duration. After removing outliers, the average value of the remaining rate of change of reception duration is calculated as the duration correction coefficient. The historical average reception duration is corrected to obtain the final average reception duration. The ratio of the actual data transmission rate to the average reception time is standardized to obtain the baseline reception flow coefficient of the current node, and then the minimum-maximum linear normalization is performed to obtain the baseline reception flow efficiency. The expected reception time is calculated based on the actual data transmission rate of the current node, the baseline reception flow efficiency and final occupancy rate, and the reception flow coefficient of each node. The formula for calculating the expected reception time is: In the formula, For the expected reception time, This represents the final occupancy rate of the current node. To preset the adjustment coefficient threshold, This represents the actual data transmission rate of the current node. This serves as the baseline receiving and processing efficiency for the current node. This is the maximum value among all the receiving flow coefficients of all nodes. It is the minimum value among all the receiving flow coefficients of all nodes.

6. The enterprise work data flow path monitoring and evaluation method as described in claim 5, characterized in that: The calculation process for the estimated processing time includes: Similarly, the unit load processing capacity coefficient and the baseline processing efficiency of the current node are calculated using the same method as the calculation of the unit load processing capacity coefficient of each node and the baseline receiving flow efficiency of the current node. The estimated processing time is calculated based on the unit load processing capacity coefficient of the current node, the baseline processing flow efficiency, and the processing flow coefficient of each node. The formula for calculating the expected processing time is: In the formula, To estimate processing time, This represents the unit load processing capacity coefficient of the current node. This represents the baseline processing throughput efficiency for the current node. This is the maximum value among all the processing flow coefficients of all nodes. It is the minimum value among all the processing flow coefficients of all nodes.

7. The method for monitoring and evaluating the flow path of enterprise work data as described in claim 6, characterized in that: The similarity calculation process includes: The baseline receiving efficiency and baseline processing efficiency of the current node are respectively denoted as the receiving efficiency and processing efficiency of the current node. The data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, receiving and processing efficiency of the current node and each other are combined to form a feature parameter dataset. For each feature parameter in the feature parameter dataset, perform min-max linear normalization to obtain the standardized feature parameters of the current node and each node. Using seven standardized feature parameters as dimensions, construct the feature vector of the current node and the feature vectors of each node respectively; The cosine similarity algorithm is used to calculate the cosine similarity value between the feature vector of the current node and the feature vectors of each node, thus obtaining the similarity between the current node and each node.

8. The enterprise work data flow path monitoring and evaluation method as described in claim 7, characterized in that: The calculation process for the similar nodes includes: Nodes with a similarity greater than or equal to a preset similarity stratification threshold are marked as similar nodes; If there are no nodes with a similarity greater than or equal to the preset similarity stratification threshold, then select the top few nodes with the highest similarity as similar nodes.

9. The method for monitoring and evaluating the flow path of enterprise work data as described in claim 1, characterized in that: The calculation process for the final deviation value includes: Calculate the absolute deviation between the expected reception time and the average reception time to obtain the reception deviation value; The processing deviation value is calculated similarly using the same method as the receiving deviation value. The larger of the received deviation value and the processed deviation value is taken as the final deviation value of the current node.

10. The method for monitoring and evaluating the flow path of enterprise work data as described in claim 1, characterized in that: The process of obtaining the data flow path evaluation results includes: If the final deviation value of the current node is less than the preset deviation threshold, the data flow path of the current node is determined to be running normally. If the final deviation value of the current node is greater than or equal to the preset deviation threshold, the data flow path of the current node is determined to be abnormal. The correlation coefficients between data transmission rate, data storage capacity, number of personnel, work efficiency, data task complexity, and final deviation value were calculated using the Pearson correlation coefficient. Remove feature parameters with a correlation coefficient of 0, sort the remaining feature parameters according to the absolute value of the correlation coefficient from smallest to largest, and select the last few feature parameters after sorting, which are recorded as abnormal feature parameters. The normal judgment result, or the combined abnormal judgment result and abnormal characteristic parameters, will be used as the evaluation result of the data flow path of the current node.