A method for intelligent analysis and traceability of enterprise data

By classifying and differentiating enterprise data, combining Excel document analysis and dynamic threshold setting, the problems of low recognition rate of complex formulas and inconsistent cross-system data are solved, and efficient and accurate enterprise data traceability and risk monitoring are achieved.

CN120256882BActive Publication Date: 2025-08-19XIAMEN MEIYA YIAN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510734018.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-19
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing enterprise data analysis traceability system has low accuracy when identifying multi-level nested formulas for complex typesettings, especially the insufficient recognition rate of handwritten formulas, which leads to easy breakage of the traceability chain, and inconsistent cross-system data standards, resulting in redundancy in ETL processes and serious time loss.

Method used

By classifying enterprise data, calculating variation coefficients and fluctuation parameters, dividing data analysis areas, calculating fluctuation factors using average deviations, combining Excel document hidden formula identification and use function call rate, differentiated management and step-by-step verification, dynamically setting abnormal thresholds to achieve accurate positioning and traceability of abnormal data.

Benefits of technology

It improves the accuracy and efficiency of enterprise data analysis, reduces audit costs, enhances financial risk monitoring capabilities, ensures the accuracy and reliability of abnormal traceability, and supports flexible adaptation to different industries' scales and risk preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256882B_ABST
    Figure CN120256882B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of enterprise data analysis, and in particular to a method for intelligent analysis and tracing of enterprise data, comprising: collecting enterprise data of a target management enterprise to obtain initial enterprise data, obtaining historical enterprise data of the target management enterprise and classifying it to obtain a number of enterprise data analysis areas, determining data fluctuation parameters based on the enterprise data within the current detection cycle of each type of enterprise data analysis area to determine the type of abnormal trend; selecting an analysis method for the corresponding enterprise data analysis area based on the abnormal trend type to determine abnormal enterprise data, performing enterprise data balance verification on the data source, determining whether the results of the current enterprise data statistical inspection are valid based on the verification result, and determining whether to trace the source of the abnormal data. The present invention significantly improves the accuracy and efficiency of data anomaly detection through the hierarchically classified intelligent analysis of enterprise data and a dynamic tracing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of enterprise data analysis, and in particular to a method for intelligent analysis and traceability of enterprise data. Background Art

[0002] In terms of formula recognition technology, current enterprise data analysis and traceability systems primarily rely on a fusion of optical character recognition (OCR) and natural language processing (NLP). Deep learning-based OCR models (such as PaddleOCR) can locate and extract mathematical formula symbols from scanned documents. Combined with symbol recognition algorithms (such as Mathpix), these formulas are converted into LaTeX structured data, which is then matched to accounting standards or business rules through a semantic parsing engine. However, when faced with complex, multi-level nested formulas (such as deferred income tax formulas in international tax calculations), existing models still exhibit errors in parsing operator precedence and the correlation between variable subscripts and subscripts. In particular, the recognition accuracy of handwritten formulas is generally below 70%, which can easily lead to breaks in the traceability chain at key calculation nodes.

[0003] The accuracy of formula recognition is limited by the diversity and quality of the training data. The annotation of specialized financial formulas is costly and faces industry barriers. Furthermore, time-sensitive traceability tasks require GPU clusters for accelerated computing, but hardware investment and energy costs increase exponentially. Furthermore, inconsistent data standards across systems lead to redundant ETL processes, further exacerbating time losses. Summary of the Invention

[0004] The purpose of the present invention is to provide an enterprise data intelligent analysis and tracing method, which can solve the problems of low hidden formula discovery rate and long audit time resulting in low enterprise data audit efficiency and low analysis accuracy.

[0005] To this end, the present invention provides a method for intelligent analysis and traceability of enterprise data, which includes:

[0006] Step S1, collecting enterprise data of the target management enterprise to obtain initial enterprise data and acquire historical enterprise data of the target management enterprise, wherein the initial enterprise data includes cost data output data;

[0007] Step S2: classify the enterprise data sources of the target management enterprise to obtain several types of enterprise data analysis areas, and obtain historical enterprise data corresponding to each type of enterprise data analysis area;

[0008] Step S3, determining data fluctuation parameters based on the enterprise data of each type of enterprise data analysis area within the current detection period to determine the abnormal trend type of the corresponding enterprise data analysis area;

[0009] Step S4, selecting an analysis method for the corresponding enterprise data analysis area based on the abnormal trend type to determine abnormal enterprise data, wherein the analysis method includes extracting some data sources for enterprise data statistical inspection and performing enterprise data statistical inspection on all data sources;

[0010] Step S5: perform enterprise data balance verification on the data source, and determine whether the result of the current enterprise data statistical check is valid based on the verification result, so as to determine whether to trace the source of abnormal data.

[0011] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, in step S2, several types of enterprise data analysis areas in the target management enterprise are determined, including:

[0012] Calculating the average cost data, standard deviation of cost data, output data and standard deviation of output data for each data source in each inspection period;

[0013] Calculate the cost coefficient of variation and output coefficient of variation for each data source separately;

[0014] The analysis type of the current enterprise data analysis area is determined according to whether the cost variation coefficient and the output variation coefficient are within a stable fluctuation range.

[0015] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, determining the type of the enterprise data analysis area includes:

[0016] If the cost variation coefficient and the output variation coefficient are both in a stable fluctuation range, the type of the current enterprise data analysis area is a stable type;

[0017] If the cost variation coefficient is not in a stable fluctuation range or the output variation coefficient is not in a stable fluctuation range, the type of the current enterprise data analysis area is a fluctuation type;

[0018] If the cost variation coefficient and the output variation coefficient are both not in a stable fluctuation range, the type of the current enterprise data analysis area is an explicit fluctuation type.

[0019] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, in step S3, determining the data fluctuation parameters determined by the historical enterprise data of each enterprise data analysis area includes:

[0020] Calculate the cost data average, cost data average deviation, output data average, and output data average deviation of the current enterprise data analysis area based on the historical enterprise data;

[0021] The ratio of the average deviation of the output data to the average value of the output data is calculated and recorded as the output fluctuation factor, and the ratio of the average deviation of the cost data to the average value of the cost data is calculated and recorded as the cost fluctuation factor;

[0022] The data fluctuation parameter is determined according to an average value of the output fluctuation factor and the cost fluctuation factor.

[0023] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, the abnormal trend type of the current enterprise data analysis area is determined based on the data fluctuation parameters, including:

[0024] If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a high abnormal trend type;

[0025] If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a low abnormal trend type.

[0026] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, in step S4, the analysis method corresponding to the enterprise data analysis area is selected according to the abnormal trend type to determine the abnormal enterprise data, including:

[0027] If the abnormal trend type is a high abnormal trend type, the analysis method is to conduct enterprise data statistical inspection on all data sources;

[0028] If the abnormal trend type is a low abnormal trend type, the analysis method is to extract part of the data sources to perform enterprise data statistical inspection.

[0029] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, in step S4, the enterprise data statistical inspection is performed on the data source, including

[0030] Obtain the Excel document of the data source, identify and obtain the total number of formulas in the Excel document;

[0031] Identify hidden formulas in Excel documents and record the total number of hidden formulas;

[0032] Calculate the infrequently used function call rate of the Excel document of the current data source by using the total number of hidden formulas and the total number of formulas;

[0033] If the infrequently used function call rate exceeds the call threshold, the current data source is determined to be a data source that needs to be traced.

[0034] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, in step S5, the process of the enterprise data balance verification includes:

[0035] Calculate the total output data minus the total cost data and record it as the first balance difference;

[0036] The first balance difference minus the total output data of all data sources is recorded as the enterprise data balance difference;

[0037] If the balance difference of the enterprise data is greater than a preset abnormality threshold, it is determined that the target management enterprise has an abnormality;

[0038] If the enterprise data balance difference is less than or equal to the preset abnormality threshold, it is determined that there is no abnormality in the target management enterprise.

[0039] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, determining whether the results of the current enterprise data statistical inspection are valid based on the verification results includes:

[0040] If the verification result meets the type consistency condition, the result of the enterprise data statistical check is judged to be valid, and the source of the abnormal data is traced based on the enterprise data statistical check;

[0041] If the verification result does not meet the type consistency condition, the result of the enterprise data statistical check is judged to be invalid, and the hidden formula in the enterprise data statistical check needs to be updated;

[0042] Among them, the type consistency condition is that the target management enterprise has an abnormality and the abnormal trend type is a high abnormal trend type, or the target management enterprise does not have an abnormality and the abnormal trend type is a low abnormal trend type.

[0043] The beneficial effects of the present invention are:

[0044] The present invention divides enterprise data analysis areas into multiple types based on the coefficient of variation of cost and output data, realizes differentiated management of data with different degrees of volatility, allocates analysis resources in a targeted manner, and improves the efficiency of anomaly identification while reducing audit costs; secondly, it adopts average deviation to calculate the volatility factor to reduce the interference of extreme values, and dynamically sets preset stable volatility parameters and anomaly thresholds based on industry benchmarks and historical data, taking into account both industry commonality and enterprise characteristics to avoid misjudgment of single indicators; thirdly, through the identification of hidden formulas in Excel documents and the analysis of uncommon function call rates, it accurately locates the source of potential abnormal data, and combines the step-by-step verification mechanism of balance difference to screen layer by layer from basic data fluctuations to global data balance to ensure the reliability of anomaly judgment; in addition, it verifies the validity of statistical inspection results through type consistency conditions, realizes anomaly tracing and dynamic update of hidden formula logic, provides enterprises with scalable financial risk monitoring capabilities, and effectively improves enterprise data stability, predictability and audit efficiency.

[0045] In particular, this invention divides enterprise data analysis areas into stable and fluctuating types, and further subdivides the degree of fluctuation, enabling more targeted financial management and risk control. For stable enterprise data areas, enterprises can continue to operate according to established strategies. For fluctuating areas, enterprises must invest appropriate resources in analysis based on the degree of fluctuation to reduce enterprise risk, improve the stability and predictability of enterprise data, and ensure sustainable development.

[0046] In particular, in the present invention, by calculating the average value and average deviation of cost data and output data, the output fluctuation factor and cost fluctuation factor are determined, and the data fluctuation parameter is determined based on their average value. The data fluctuation factors are considered from multiple aspects, making the determination of the data fluctuation parameter more comprehensive, more representative of the whole, and more accurately reflecting the stability of enterprise data. At the same time, the use of average deviation instead of standard deviation reduces the sensitivity to extreme values, avoids the excessive amplification of the overall data fluctuation due to individual abnormal data, more accurately grasps the central trend and stability of the data, and improves the accuracy of enterprise data anomaly identification. Based on the comparison of data fluctuation parameters with preset stable fluctuation parameters, the abnormal trend type is divided into high abnormality and low abnormality, and the analysis method of matching all data source inspections and partial data extraction inspections is matched respectively, realizing differentiated processing of enterprise data with different abnormality levels, which can not only efficiently discover potential problems, but also reasonably allocate analysis resources, improve the efficiency and pertinence of enterprise data analysis, and effectively monitor the enterprise data status accurately and respond to abnormalities.

[0047] In particular, in the present invention, the accuracy and efficiency of financial anomaly detection are significantly improved through balance difference calculation combined with a dynamic threshold judgment mechanism. A step-by-step verification mode is adopted to screen layer by layer from basic data to global balance, effectively reducing the risk of single indicator deviation. Thresholds are dynamically set based on industry benchmarks and historical data (such as twice the standard deviation), taking into account both horizontal industry commonalities and vertical enterprise characteristics to avoid one-size-fits-all misjudgments. At the same time, anomalies are quickly located through automated rules, significantly shortening the manual review cycle, and supporting flexible adaptation to different industry scales and risk preferences, providing enterprises with proactive and scalable data health monitoring and risk warning capabilities.

[0048] In particular, by setting type consistency or inconsistency conditions to determine the validity of results, it ensures that only accurately verified enterprise data statistical inspection results are used in subsequent abnormal data tracing, thereby improving the accuracy and reliability of abnormal tracing. At the same time, when the verification results do not meet the conditions, it explicitly requires updating the hidden formula. This not only helps to promptly correct potential data anomalies or logical errors, but also demonstrates the solution's dynamic optimization and self-improvement capabilities, further enhancing the accuracy and adaptability of intelligent enterprise data analysis, and providing solid technical support for enterprise data health monitoring and risk management. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Flowchart of the method for intelligent analysis and traceability of enterprise data in an embodiment of the present invention;

[0050] Figure 2 A flowchart of determining several types of enterprise data analysis areas in a target management enterprise in an embodiment of the present invention;

[0051] Figure 3 A logical diagram of abnormal trend types in the current enterprise data analysis area in an embodiment of the present invention;

[0052] Figure 4 This is a flow chart of performing enterprise data statistical inspection on data sources in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0054] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0055] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0056] See also Figure 1 As shown, it is a flow chart of an enterprise data intelligent analysis and traceability method in an embodiment of the present invention. The present invention provides an enterprise data intelligent analysis and traceability method, including:

[0057] Step S1, collecting enterprise data of the target management enterprise to obtain initial enterprise data and acquire historical enterprise data of the target management enterprise, wherein the initial enterprise data includes cost data output data;

[0058] Step S2: classify the enterprise data sources of the target management enterprise to obtain several types of enterprise data analysis areas, and obtain historical enterprise data corresponding to each type of enterprise data analysis area;

[0059] Step S3, determining data fluctuation parameters based on the enterprise data of each type of enterprise data analysis area within the current detection period to determine the abnormal trend type of the corresponding enterprise data analysis area;

[0060] Step S4, selecting an analysis method for the corresponding enterprise data analysis area based on the abnormal trend type to determine abnormal enterprise data, wherein the analysis method includes extracting some data sources for enterprise data statistical inspection and performing enterprise data statistical inspection on all data sources;

[0061] Step S5: perform enterprise data balance verification on the data source, and determine whether the result of the current enterprise data statistical check is valid based on the verification result, so as to determine whether to trace the source of abnormal data.

[0062] The present invention divides enterprise data analysis areas into multiple types based on the coefficient of variation of cost and output data, realizes differentiated management of data with different degrees of volatility, allocates analysis resources in a targeted manner, and improves the efficiency of anomaly identification while reducing audit costs; secondly, it uses average deviation to calculate the volatility factor to reduce the interference of extreme values, and dynamically sets preset stable fluctuation parameters and anomaly thresholds based on industry benchmarks and historical data, taking into account both industry commonality and enterprise characteristics to avoid misjudgment of single indicators; thirdly, through the identification of hidden formulas in Excel documents and the analysis of uncommon function call rates, it accurately locates the source of potential anomaly data, and combines the step-by-step verification mechanism of balance difference to screen layer by layer from basic data fluctuations to global data balance to ensure the reliability of anomaly judgment; in addition, the validity of statistical inspection results is verified through type consistency conditions, and the dynamic update of anomaly tracing and hidden formula logic is realized, providing enterprises with scalable data risk monitoring capabilities, effectively improving the stability, predictability and audit efficiency of enterprise data.

[0063] See also Figure 2 As shown, it is a flowchart of determining several types of enterprise data analysis areas in a target management enterprise in an embodiment of the present invention. In the step S2, determining several types of enterprise data analysis areas in the target management enterprise includes:

[0064] Calculating the average cost data, standard deviation of cost data, output data and standard deviation of output data for each data source in each inspection period;

[0065] Calculate the cost coefficient of variation and output coefficient of variation for each data source separately;

[0066] The analysis type of the current enterprise data analysis area is determined according to whether the cost variation coefficient and the output variation coefficient are within a stable fluctuation range.

[0067] In practice, the inspection cycle is preferably one month, but is not limited thereto. Those skilled in the art may adjust this value based on actual needs or the nature of the enterprise. When calculating the average and standard deviation, the cost data and output data for all days within the inspection cycle are obtained for calculation.

[0068] The coefficient of variation (CV) is calculated as follows: CV = mean / standard deviation × 100%;

[0069] In this embodiment, the stable fluctuation range is [10%, 20%].

[0070] Understandably, the coefficient of variation (CV) is widely used in enterprise data analysis to assess the stability of metrics like revenue and costs (e.g., in textbooks like "Financial Risk Management"). For example, the manufacturing industry typically requires a production cost CV of ≤10% to control supply chain fluctuations. Low volatility (CV ≤10%) is suitable for revenue data in mature markets with fixed costs, while high volatility (CV >20%) is common in startups or cyclical industries (such as commodities).

[0071] Specifically, determining the type of the enterprise data analysis area includes:

[0072] If the cost variation coefficient and the output variation coefficient are both in a stable fluctuation range, the type of the current enterprise data analysis area is a stable type;

[0073] If the cost variation coefficient is not in a stable fluctuation range or the output variation coefficient is not in a stable fluctuation range, the type of the current enterprise data analysis area is a fluctuation type;

[0074] If the cost variation coefficient and the output variation coefficient are both not in a stable fluctuation range, the type of the current enterprise data analysis area is an explicit fluctuation type.

[0075] In this embodiment, the stable type indicates that the cost data and output data of the current enterprise data analysis area remain in a stable state with a small fluctuation range. Within this range, the changes in enterprise data are relatively mild and may be affected by some normal and controllable internal and external factors, such as small fluctuations in raw material prices and slight changes in market demand. This type of enterprise data is less likely to have anomalies, and any anomalies in this type of enterprise data are easy to identify.

[0076] The presence of a fluctuation type indicates moderate fluctuations in cost data or output data within the enterprise data analysis area, indicating significant changes in enterprise data. This may be due to factors such as intensified market competition, rising production costs, adjustments to sales strategies, or minor statistical anomalies in the data. Closely monitor these changes, assess their impact on the enterprise's data status, and consider implementing appropriate measures to address potential data risks, such as optimizing cost controls, adjusting production plans, and expanding sales channels.

[0077] High volatility in the explicit volatility type indicates that the cost data and output data in the enterprise data analysis area fluctuate greatly, which means that the enterprise data is more abnormal and may be affected by major factors, such as sudden changes in the macroeconomic environment, adjustments in industry policies, major operational decision-making errors, or a large degree of statistical anomalies in data information.

[0078] By dividing enterprise data analysis areas into stable and fluctuating types, and further subdividing them by the degree of fluctuation, this invention enables more targeted data management and risk control. For stable enterprise data areas, enterprises can continue to operate according to established strategies. For fluctuating areas, enterprises need to invest appropriate resources in analysis based on the degree of fluctuation to reduce enterprise risks, improve the stability and predictability of enterprise data, and ensure sustainable development.

[0079] Specifically, in step S3, determining the data fluctuation parameters determined by the historical enterprise data of each enterprise data analysis area includes:

[0080] Calculate the cost data average, cost data average deviation, output data average, and output data average deviation of the current enterprise data analysis area based on the historical enterprise data;

[0081] The ratio of the average deviation of the output data to the average value of the output data is calculated and recorded as the output fluctuation factor, and the ratio of the average deviation of the cost data to the average value of the cost data is calculated and recorded as the cost fluctuation factor;

[0082] The data fluctuation parameter is determined according to an average value of the output fluctuation factor and the cost fluctuation factor.

[0083] During implementation, the cost data and output data of each day in the current detection cycle are obtained for calculation

[0084] Understandably, in actual financial analysis, fluctuations in corporate data may be affected by unique factors, such as seasonality or one-time events. The mean deviation is relatively less sensitive to individual extreme values and more accurately reflects the central tendency and stability of the data. However, the standard deviation amplifies the impact of extreme values during calculation, leading to an over-interpretation of overall data fluctuations.

[0085] See also Figure 3 As shown, it is a logic diagram of the abnormal trend type of the current enterprise data analysis area in an embodiment of the present invention. The abnormal trend type of the current enterprise data analysis area is determined according to the data fluctuation parameter, including:

[0086] If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a high abnormal trend type;

[0087] If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a low abnormal trend type.

[0088] In implementation, the preset stable fluctuation parameters of the smooth type are selected in the interval [0.05, 0.1], the preset stable fluctuation parameters of the existential fluctuation type are selected in the interval [0.08, 0.15], and the preset stable fluctuation parameters of the dominant fluctuation type are selected in the interval [0.12, 0.2].

[0089] Specifically, in step S4, selecting an analysis method for the corresponding enterprise data analysis area according to the abnormal trend type to determine abnormal enterprise data includes:

[0090] If the abnormal trend type is a high abnormal trend type, the analysis method is to conduct enterprise data statistical inspection on all data sources;

[0091] If the abnormal trend type is a low abnormal trend type, the analysis method is to extract part of the data sources to perform enterprise data statistical inspection.

[0092] During implementation, some data sources are extracted for statistical inspection, which is 5% to 15% of all data sources in the enterprise data analysis area, preferably 10%. There is no limit on the extraction method.

[0093] In the present invention, by calculating the average value and average deviation of cost data and output data, the output fluctuation factor and cost fluctuation factor are determined, and the data fluctuation parameter is determined based on their average value. The data fluctuation factors are considered from multiple aspects, making the determination of the data fluctuation parameter more comprehensive, more representative of the whole, and more accurately reflecting the stability of enterprise data. At the same time, the use of average deviation instead of standard deviation reduces the sensitivity to extreme values, avoids excessive amplification of the overall data fluctuation due to individual abnormal data, more accurately grasps the central trend and stability of the data, and improves the accuracy of enterprise data anomaly identification. Based on the comparison of data fluctuation parameters with preset stable fluctuation parameters, the abnormal trend type is divided into high abnormality and low abnormality, and the analysis method of matching all data source inspections and partial data extraction inspections is used for them respectively, to achieve differentiated processing of enterprise data with different abnormality levels, which can not only efficiently discover potential problems, but also reasonably allocate analysis resources, improve the efficiency and pertinence of enterprise data analysis, and effectively monitor the enterprise data status accurately and respond to abnormalities.

[0094] See also Figure 4 As shown, it is a flow chart of performing enterprise data statistical inspection on data sources in an embodiment of the present invention. In the step S4, performing enterprise data statistical inspection on the data sources includes:

[0095] Obtain the Excel document of the data source, identify and obtain the total number of formulas in the Excel document;

[0096] Identify hidden formulas in Excel documents and record the total number of hidden formulas;

[0097] Calculate the infrequently used function call rate of the Excel document of the current data source by using the total number of hidden formulas and the total number of formulas;

[0098] If the infrequently used function call rate exceeds the call threshold, the current data source is determined to be a data source that needs to be traced.

[0099] During implementation, the OpenXML SDK was used to parse Excel's XML metadata to identify hidden formulas. The call threshold was 10%. According to the definition of materiality levels in the International Standards on Auditing (ISA 240), if the proportion of hidden formulas exceeds 10%, it usually requires special review.

[0100] Obtain the Excel file that needs to be analyzed from the financial system and back it up. Use the OpenXML SDK to parse the Excel file into XML format, extract the metadata of related elements, parse the XML metadata, obtain structural information such as workbooks, worksheets, cells, and formulas, extract the formula content from the metadata, record the relevant locations and function composition, establish a library of uncommon functions, compare the extracted functions with the library, mark uncommon functions, analyze the logical structure of the formula, evaluate the complexity and rationality, and check whether the calculation results meet expectations.

[0101] Specifically, in step S5, the process of verifying the balance of enterprise data includes:

[0102] Calculate the total output data minus the total cost data and record it as the first balance difference;

[0103] The first balance difference minus the total output data of all data sources is recorded as the enterprise data balance difference;

[0104] If the balance difference of the enterprise data is greater than a preset abnormality threshold, it is determined that the target management enterprise has an abnormality;

[0105] If the enterprise data balance difference is less than or equal to the preset abnormality threshold, it is determined that there is no abnormality in the target management enterprise.

[0106] During implementation, the preset abnormality threshold is determined based on reference to data reports and industry analysis reports of other companies in the same industry, combined with the general situation of balance sheet balance in the industry and the common error range, or two times the standard deviation of the average value of balance difference of historical enterprise data is selected as the threshold.

[0107] In the present invention, the accuracy and efficiency of data anomaly detection are significantly improved through balanced difference calculation combined with a dynamic threshold judgment mechanism. A step-by-step verification mode is adopted to screen layer by layer from basic data to global balance, effectively reducing the risk of single indicator deviation. Thresholds are dynamically set based on industry benchmarks and historical data (such as twice the standard deviation), taking into account both horizontal industry commonalities and vertical enterprise characteristics to avoid one-size-fits-all misjudgments. At the same time, anomalies are quickly located through automated rules, significantly shortening the manual review cycle, and supporting flexible adaptation to different industry scales and risk preferences, providing enterprises with proactive and scalable data health monitoring and risk warning capabilities.

[0108] Specifically, determining whether the result of the current enterprise data statistical check is valid according to the verification result includes:

[0109] If the verification result meets the type consistency condition, the result of the enterprise data statistical check is judged to be valid, and the source of the abnormal data is traced based on the enterprise data statistical check;

[0110] If the verification result does not meet the type consistency condition, the result of the enterprise data statistical check is judged to be invalid, and the hidden formula in the enterprise data statistical check needs to be updated;

[0111] Among them, the type consistency condition is that the target management enterprise has an abnormality and the abnormal trend type is a high abnormal trend type, or the target management enterprise does not have an abnormality and the abnormal trend type is a low abnormal trend type.

[0112] In this invention, by setting type consistency or inconsistency conditions to determine whether the results are valid, it can ensure that only accurately verified enterprise data statistical inspection results will be used for subsequent abnormal data tracing work, thereby improving the accuracy and reliability of abnormal tracing. At the same time, when the verification results do not meet the conditions, it is clearly stated that the hidden formula needs to be updated. This not only helps to correct potential data anomalies or logical errors in a timely manner, but also reflects the dynamic optimization and self-improvement capabilities of the solution, further enhancing the accuracy and adaptability of enterprise data intelligent analysis, and providing solid technical support for enterprise data health monitoring and risk management.

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based device that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0114] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. A person skilled in the art would be able to make other variations or modifications based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for intelligent analysis and traceability of enterprise data, characterized in that: include; Step S1, collecting enterprise data of the target management enterprise to obtain initial enterprise data and acquire historical enterprise data of the target management enterprise, wherein the initial enterprise data includes cost data output data; Step S2: classify the enterprise data sources of the target management enterprise to obtain several types of enterprise data analysis areas, and obtain historical enterprise data corresponding to each type of enterprise data analysis area; Step S3, determining data fluctuation parameters based on the enterprise data of each type of enterprise data analysis area within the current detection period to determine the abnormal trend type of the corresponding enterprise data analysis area; Step S4, selecting an analysis method for the corresponding enterprise data analysis area based on the abnormal trend type to determine abnormal enterprise data, wherein the analysis method includes extracting some data sources for enterprise data statistical inspection and performing enterprise data statistical inspection on all data sources; Step S5: Perform enterprise data balance verification on the data source, and determine whether the results of the current enterprise data statistical check are valid based on the verification results, so as to determine whether to trace the source of abnormal data; In step S2, several types of enterprise data analysis areas in the target management enterprise are determined, including: Calculating the average cost data, standard deviation of cost data, output data and standard deviation of output data for each data source in each inspection period; Calculate the cost coefficient of variation and output coefficient of variation for each data source separately; The analysis type of the current enterprise data analysis area is determined according to whether the cost variation coefficient and the output variation coefficient are within a stable fluctuation range.

2. The enterprise data intelligent analysis and traceability method according to claim 1 is characterized in that: Determining the type of the enterprise data analysis area includes: If the cost variation coefficient and the output variation coefficient are both in a stable fluctuation range, the type of the current enterprise data analysis area is a stable type; If the cost variation coefficient is not in a stable fluctuation range or the output variation coefficient is not in a stable fluctuation range, the type of the current enterprise data analysis area is a fluctuation type; If the cost variation coefficient and the output variation coefficient are both not in a stable fluctuation range, the type of the current enterprise data analysis area is an explicit fluctuation type.

3. The enterprise data intelligent analysis and traceability method according to claim 2 is characterized in that: In step S3, determining the data fluctuation parameters determined by the historical enterprise data of each enterprise data analysis area includes: Calculate the cost data average, cost data average deviation, output data average, and output data average deviation of the current enterprise data analysis area based on the historical enterprise data; The ratio of the average deviation of the output data to the average value of the output data is calculated and recorded as the output fluctuation factor, and the ratio of the average deviation of the cost data to the average value of the cost data is calculated and recorded as the cost fluctuation factor; The data fluctuation parameter is determined according to an average value of the output fluctuation factor and the cost fluctuation factor.

4. The enterprise data intelligent analysis and traceability method according to claim 3 is characterized in that: Determining the abnormal trend type of the current enterprise data analysis area based on the data fluctuation parameters includes: If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a high abnormal trend type; If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a low abnormal trend type.

5. The enterprise data intelligent analysis and traceability method according to claim 4 is characterized in that: In step S4, selecting an analysis method for the corresponding enterprise data analysis area according to the abnormal trend type to determine abnormal enterprise data includes: If the abnormal trend type is a high abnormal trend type, the analysis method is to conduct enterprise data statistical inspection on all data sources; If the abnormal trend type is a low abnormal trend type, the analysis method is to extract part of the data sources to perform enterprise data statistical inspection.

6. The enterprise data intelligent analysis and traceability method according to claim 5 is characterized in that: In step S4, the data source is statistically checked for enterprise data, including Obtain the Excel document of the data source, identify and obtain the total number of formulas in the Excel document; Identify hidden formulas in Excel documents and record the total number of hidden formulas; Calculate the infrequently used function call rate of the Excel document of the current data source by using the total number of hidden formulas and the total number of formulas; If the infrequently used function call rate exceeds the call threshold, the current data source is determined to be a data source that needs to be traced.

7. The enterprise data intelligent analysis and traceability method according to claim 6 is characterized in that: In step S5, the process of enterprise data balance verification includes: Calculate the total output data minus the total cost data and record it as the first balance difference; The first balance difference minus the total output data of all data sources is recorded as the enterprise data balance difference; If the balance difference of the enterprise data is greater than the preset abnormality threshold, it is determined that the target management enterprise has an abnormality; If the enterprise data balance difference is less than or equal to the preset abnormality threshold, it is determined that there is no abnormality in the target management enterprise.

8. The enterprise data intelligent analysis and traceability method according to claim 7 is characterized in that: Determine whether the results of the current enterprise data statistical inspection are valid based on the verification results, including: If the verification result meets the type consistency condition, the result of the enterprise data statistical check is judged to be valid, and the source of the abnormal data is traced based on the enterprise data statistical check; If the verification result does not meet the type consistency condition, the result of the enterprise data statistical check is judged to be invalid, and the hidden formula in the enterprise data statistical check needs to be updated; Among them, the type consistency condition is that the target management enterprise has an abnormality and the abnormal trend type is a high abnormal trend type, or the target management enterprise does not have an abnormality and the abnormal trend type is a low abnormal trend type.

Citation Information

Patent Citations

  • Checking management system

    CN109242416A

  • Block chain-based job datamation template verification and traceability method

    CN117932674A

  • Method for grabbing target data in batch in mass data synchronization process

    CN118673034A