Enterprise data intelligent analysis and traceability method

By classifying enterprise data and setting dynamic thresholds, the abnormal data sources are accurately identified, and the problems of low recognition rate of complex formulas and inconsistent data across systems are solved, and efficient and reliable enterprise data traceability and risk management are achieved.

CN120256882AActive Publication Date: 2025-07-04XIAMEN MEIYA YIAN INFORMATION TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510734018.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing enterprise data analysis and traceability system has low accuracy when identifying complex typesetting multi-level nested formulas, especially the insufficient recognition rate of handwritten formulas, which leads to easy breakage of the traceability chain, and inconsistent cross-system data standards, resulting in redundancy in ETL processes and serious time loss.

Method used

By classifying enterprise data, calculating coefficients of variation and fluctuation parameters, dividing data analysis areas, using average deviation to reduce extreme value interference, combining Excel document hidden formula identification and extraordinary function call rate analysis, dynamically setting abnormal thresholds to achieve differentiated management and precise positioning of abnormal data sources.

Benefits of technology

It improves the accuracy and efficiency of enterprise data analysis, reduces audit costs, ensures the reliability of abnormal identification and global data balance, provides flexible risk monitoring capabilities, and supports adaptation in different industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256882A_ABST
    Figure CN120256882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of enterprise data analysis, in particular to an enterprise data intelligent analysis and traceability method, which comprises the following steps: performing enterprise data acquisition on a target management enterprise to obtain initial enterprise data, obtaining historical enterprise data of the target management enterprise and classifying the historical enterprise data to obtain a plurality of enterprise data analysis areas, according to data fluctuation parameters determined by the enterprise data in the current detection period of each type of enterprise data analysis area, determining an abnormal trend type; selecting and determining an analysis mode corresponding to the enterprise data analysis area according to the abnormal trend type so as to determine abnormal enterprise data, performing enterprise data balance verification on the data source, and determining whether a current enterprise data statistical inspection result is valid or not according to a verification result so as to determine whether the abnormal data source is traced or not. According to the invention, through a layered and classified enterprise data intelligent analysis and dynamic traceability mechanism, the accuracy and efficiency of data anomaly detection are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of enterprise data analysis, and particularly to an intelligent analysis and traceability method for enterprise data. Background Art

[0002] At the technical level of formula recognition, the current enterprise data analysis and traceability system mainly relies on the integration scheme of OCR (Optical Character Recognition) and natural language processing (NLP). The OCR model based on deep learning (such as PaddleOCR) can locate and extract mathematical formula symbols in scanned documents, convert the formula into LaTeX structured data in combination with a symbol recognition algorithm (such as Mathpix), and then match accounting standards or business rules through a semantic parsing engine. However, in the face of multi-level nested formulas with complex layouts (such as the deferred income tax formula in international tax calculations), there are still errors in the existing models' parsing of operator precedence and the relevance of variable subscripts and superscripts. In particular, the recognition accuracy of handwritten formulas is generally lower than 70%, resulting in the breakage of the traceability chain at key calculation nodes.

[0003] The accuracy of formula recognition is limited by the diversity and annotation quality of training data. The annotation cost of professional formulas in the financial field is high and there are industry barriers. On the other hand, time-sensitive traceability tasks rely on GPU clusters to accelerate calculations, but the hardware investment and energy consumption costs increase exponentially. In addition, the non-uniformity of cross-system data standards leads to redundant ETL processes, further exacerbating time consumption. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent analysis and traceability method for enterprise data, which can solve the problems of low discovery rate of hidden formulas, long audit time, low audit efficiency, and low analysis accuracy of enterprise data.

[0005] To this end, the present invention provides an intelligent analysis and traceability method for enterprise data, which includes: Step S1: Collect enterprise data of the target managed enterprise to obtain initial enterprise data, and acquire the historical enterprise data of the target managed enterprise, where the initial enterprise data includes cost data output data; Step S2: Classify the sources of enterprise data of the target managed enterprise to obtain several types of enterprise data analysis regions, and acquire the historical enterprise data corresponding to each type of enterprise data analysis region; Step S3: Determine the abnormal trend type of the corresponding enterprise data analysis region according to the data fluctuation parameters determined by the enterprise data within the current detection period of each type of enterprise data analysis region; Step S4, select and determine the analysis method for the enterprise data analysis area corresponding to the abnormal trend type to determine the abnormal enterprise data, where the analysis method includes extracting some data sources for enterprise data statistical inspection and performing enterprise data statistical inspection on all data sources; Step S5, perform enterprise data balance verification on the data sources, and determine whether the result of the current enterprise data statistical inspection is valid according to the verification result, so as to determine whether to trace the abnormal data source.

[0006] As an optimal technical solution of the enterprise data intelligent analysis and traceability method, in the step S2, determine several types of enterprise data analysis areas in the target management enterprise, including: Calculate the average value of cost data, the standard deviation of cost data, the output data, and the standard deviation of output data of each data source in each inspection period respectively; Calculate the cost variation coefficient and the output variation coefficient of each data source respectively; Determine the analysis type of the current enterprise data analysis area according to whether the cost variation coefficient and the output variation coefficient are within the stable fluctuation range.

[0007] As an optimal technical solution of the enterprise data intelligent analysis and traceability method, determine the types of the enterprise data analysis areas including: If both the cost variation coefficient and the output variation coefficient are within the stable fluctuation range, the type of the current enterprise data analysis area is the stable type; If the cost variation coefficient is not within the stable fluctuation range or the output variation coefficient is not within the stable fluctuation range, the type of the current enterprise data analysis area is the type with fluctuations; If both the cost variation coefficient and the output variation coefficient are not within the stable fluctuation range, the type of the current enterprise data analysis area is the obvious fluctuation type.

[0008] As an optimal technical solution of the enterprise data intelligent analysis and traceability method, in the step S3, determine the data fluctuation parameters of the historical enterprise data of each enterprise data analysis area, including: Calculate the average value of cost data, the average deviation of cost data, the average value of output data, and the average deviation of output data of the current enterprise data analysis area according to the historical enterprise data; Calculate the ratio of the average deviation of output data to the average value of output data and record it as the output fluctuation factor, and calculate the ratio of the average deviation of cost data to the average value of cost data and record it as the cost fluctuation factor; Determine the data fluctuation parameter according to the average value of the output fluctuation factor and the cost fluctuation factor.

[0009] As an optimal technical solution for the enterprise data intelligent analysis and traceability method, determining the abnormal trend type of the current enterprise data analysis area according to the data fluctuation parameter, including: If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a high abnormal trend type; If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a low abnormal trend type.

[0010] As an optimal technical solution for the enterprise data intelligent analysis and traceability method, in the step S4, selecting and determining the analysis method for the corresponding enterprise data analysis area according to the abnormal trend type to determine the abnormal enterprise data, including: If the abnormal trend type is a high abnormal trend type, the analysis method is to conduct enterprise data statistical inspection on all data sources; If the abnormal trend type is a low abnormal trend type, the analysis method is to extract part of the data sources for enterprise data statistical inspection.

[0011] As an optimal technical solution for the enterprise data intelligent analysis and traceability method, in the step S4, conducting enterprise data statistical inspection on the data source, including Obtaining the Excel document of the data source, identifying and obtaining the total number of formulas in the Excel document; Identifying the hidden formulas in the Excel document and recording the total number of hidden formulas; Calculating the call rate of uncommon functions in the current data source Excel document through the total number of hidden formulas and the total number of formulas; If the call rate of uncommon functions exceeds the call threshold, it is determined that the current data source is a data source to be traced.

[0012] As an optimal technical solution for the enterprise data intelligent analysis and traceability method, in the step S5, the process of enterprise data balance verification includes; Calculating the total output data minus the total cost data and recording it as the first balance difference; Subtracting the total output data of all data sources from the first balance difference, and recording it as the enterprise data balance difference; If the enterprise data balance difference is greater than the preset abnormal threshold, it is determined that there is an abnormality in the target management enterprise; If the enterprise data balance difference is less than or equal to the preset abnormal threshold, it is determined that there is no abnormality in the target management enterprise.

[0013] As a preferred technical solution for the enterprise data intelligent analysis and traceability method, determining whether the result of the current enterprise data statistical inspection is valid according to the verification result includes: If the verification result meets the type consistency condition, the result of the enterprise data statistical check is judged to be valid, and the source of the abnormal data is traced based on the enterprise data statistical check; If the verification result does not meet the type consistency condition, the result of the enterprise data statistical check is judged to be invalid, and the hidden formula in the enterprise data statistical check needs to be updated; Among them, the type consistency condition is that the target management enterprise has an abnormality and the abnormal trend type is a high abnormal trend type, or the target management enterprise does not have an abnormality and the abnormal trend type is a low abnormal trend type.

[0014] The beneficial effects of the present invention are: The present invention divides the enterprise data analysis area into multiple types based on the coefficient of variation of cost and output data, realizes differentiated management of data with different volatility degrees, allocates analysis resources in a targeted manner, and improves the efficiency of anomaly identification while reducing audit costs; secondly, the average deviation is used to calculate the volatility factor to reduce the interference of extreme values, and the preset stable volatility parameters and abnormal thresholds are dynamically set in combination with industry benchmarks and historical data, taking into account both industry commonality and enterprise characteristics to avoid misjudgment of a single indicator; thirdly, through the identification of hidden formulas in Excel documents and the analysis of uncommon function call rates, the potential source of abnormal data is accurately located, and combined with the step-by-step verification mechanism of the balance difference, it screens layer by layer from basic data fluctuations to global data balance to ensure the reliability of anomaly judgment; in addition, the validity of the statistical inspection results is verified through type consistency conditions, and the dynamic update of anomaly tracing and hidden formula logic is realized, providing enterprises with scalable financial risk monitoring capabilities, effectively improving the stability, predictability and audit efficiency of enterprise data.

[0015] In particular, the present invention can carry out financial management and risk control more targetedly by dividing the enterprise data analysis area into stable type and fluctuating type, and further subdividing the degree of fluctuation. For stable type enterprise data areas, enterprises can continue to operate according to established strategies. For areas with fluctuating types, enterprises need to invest corresponding resources for analysis according to the degree of fluctuation, so as to reduce enterprise risks, improve the stability and predictability of enterprise data, and ensure the sustainable development of enterprises.

[0016] In particular, in the present invention, by calculating the average values and average deviations of cost data and output data, the output fluctuation factor and the cost fluctuation factor are further determined, and the data fluctuation parameter is determined based on their average values. Considering the data fluctuation factors from multiple aspects makes the determination of the data fluctuation parameter more comprehensive, with stronger representativeness for the whole, and can more accurately reflect the stability of enterprise data. At the same time, using the average deviation instead of the standard deviation reduces the sensitivity to extreme values, avoids the over-amplification of the overall data fluctuation situation caused by individual abnormal data, more precisely grasps the central tendency and stability of the data, and improves the accuracy of enterprise data anomaly identification. Based on the comparison between the data fluctuation parameter and the preset stable fluctuation parameter, the abnormal trend types are divided into two types: high anomaly and low anomaly, and analysis methods of checking all data sources and extracting partial data are respectively matched for them, realizing the differential processing of enterprise data with different abnormal degrees, which can not only efficiently discover potential problems, but also reasonably allocate analysis resources, improve the efficiency and pertinence of enterprise data analysis, and effectively monitor and respond to anomalies in the enterprise data situation accurately.

[0017] In particular, in the present invention, through the balance difference calculation and in combination with the dynamic threshold determination mechanism, the accuracy and efficiency of financial anomaly detection are significantly improved. The step-by-step verification mode is adopted to screen layer by layer from basic data to global balance, effectively reducing the risk of single-index deviation; the threshold is dynamically set based on industry benchmarks and historical data (such as twice the standard deviation), taking into account both horizontal industry commonalities and vertical enterprise characteristics, avoiding one-size-fits-all misjudgments; at the same time, anomalies are quickly located through automated rules, greatly shortening the manual review cycle, and supporting flexible adaptation to different industry scales and risk preferences, providing the enterprise with the capabilities of automated and scalable data health monitoring and risk warning.

[0018] In particular, by setting conditions of consistent or inconsistent types to judge whether the result is valid, it can be ensured that only the statistical inspection results of enterprise data that have been accurately verified will be used for subsequent anomaly data tracing work, thereby improving the accuracy and reliability of anomaly tracing. At the same time, when the verification result does not meet the conditions, it is clearly proposed that the hidden formula needs to be updated, which not only helps to timely correct potential data anomalies or logical errors, but also reflects the dynamic optimization and self-improving ability of the solution, further enhancing the accuracy and adaptability of enterprise data intelligent analysis, and providing a solid technical support for the enterprise's data health monitoring and risk management. Description of the Drawings

[0019] Figure 1 It is a flowchart of the enterprise data intelligent analysis and tracing method in the embodiment of the present invention; Figure 2 It is a flowchart of determining several types of enterprise data analysis regions in the target management enterprise in the embodiment of the present invention; Figure 3It is a logic diagram of the abnormal trend type in the current enterprise data analysis area in the embodiment of the present invention; Figure 4 It is a flowchart for statistically checking enterprise data on the data source in the embodiment of the present invention. Specific implementation manner

[0020] Next, the technical solution of the present invention will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0022] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0023] Please refer to Figure 1 As shown, it is a flowchart of the enterprise data intelligent analysis and traceability method in the embodiment of the present invention. The present invention provides an enterprise data intelligent analysis and traceability method, including; Step S1, collect enterprise data for the target management enterprise to obtain initial enterprise data, and obtain the historical enterprise data of the target management enterprise. Among them, the initial enterprise data includes cost data output data; Step S2, classify the data sources of the enterprise data of the target management enterprise to obtain several types of enterprise data analysis areas, and obtain the historical enterprise data corresponding to each type of enterprise data analysis area; Step S3, determine the abnormal trend type of the corresponding enterprise data analysis area according to the data fluctuation parameters determined by the enterprise data in the current detection period of each type of enterprise data analysis area; Step S4, selecting an analysis method for determining a corresponding enterprise data analysis area according to the abnormal trend type to determine abnormal enterprise data, wherein the analysis method includes extracting part of the data sources for enterprise data statistical inspection and performing enterprise data statistical inspection on all data sources; Step S5, perform enterprise data balance verification on the data source, and determine whether the result of the current enterprise data statistical check is valid based on the verification result, so as to determine whether to trace the source of abnormal data.

[0024] The present invention divides the enterprise data analysis area into multiple types based on the coefficient of variation of cost and output data, realizes differentiated management of data with different volatility degrees, allocates analysis resources in a targeted manner, and improves the efficiency of anomaly identification while reducing audit costs; secondly, the average deviation is used to calculate the volatility factor to reduce the interference of extreme values, and the preset stable volatility parameters and abnormal thresholds are dynamically set in combination with industry benchmarks and historical data, taking into account both industry commonality and enterprise characteristics to avoid misjudgment of a single indicator; thirdly, through the identification of hidden formulas in Excel documents and the analysis of uncommon function call rates, the potential source of abnormal data is accurately located, and combined with the step-by-step verification mechanism of the balance difference, it screens layer by layer from basic data fluctuations to global data balance to ensure the reliability of anomaly judgment; in addition, the validity of the statistical inspection results is verified through type consistency conditions, and the dynamic update of anomaly tracing and hidden formula logic is realized, providing enterprises with scalable data risk monitoring capabilities, effectively improving the stability, predictability and audit efficiency of enterprise data.

[0025] See also Figure 2 As shown, it is a flowchart of determining several types of enterprise data analysis areas in a target management enterprise in an embodiment of the present invention. In the step S2, determining several types of enterprise data analysis areas in the target management enterprise includes: Calculate the cost data average, cost data standard deviation, output data and output data standard deviation of each data source in each inspection period respectively; Calculate the cost coefficient of variation and output coefficient of variation for each data source; The analysis type of the current enterprise data analysis area is determined according to whether the cost variation coefficient and the output variation coefficient are within a stable fluctuation range.

[0026] In practice, the inspection cycle is preferably one month, but the inspection cycle is not limited thereto, and those skilled in the art may also adjust the value according to actual needs or the nature of the enterprise. When calculating the mean and standard deviation, the cost data and output data of all single days in the inspection cycle are obtained for calculation.

[0027] The calculation formula of the coefficient of variation (CV) is CV = average value / standard deviation × 100%; In this embodiment, the stable fluctuation range is [10%, 20%].

[0028] It can be understood that in enterprise data analysis, the coefficient of variation CV is widely used to evaluate the stability of indicators such as revenue and cost (such as textbooks like "Financial Risk Management"). For example, the manufacturing industry usually requires the CV of production costs to be ≤ 10% to control supply chain fluctuations. Low fluctuations (CV ≤ 10%) are applicable to revenue data in mature markets with fixed costs, and high fluctuations (CV > 20%) are common in startups or cyclical industries (such as commodities).

[0029] Specifically, determining the type of the enterprise data analysis area includes: If both the cost coefficient of variation and the output coefficient of variation are within the stable fluctuation range, the type of the current enterprise data analysis area is the stable type; If the cost coefficient of variation is not within the stable fluctuation range or the output coefficient of variation is not within the stable fluctuation range, the type of the current enterprise data analysis area is the type with fluctuations; If both the cost coefficient of variation and the output coefficient of variation are not within the stable fluctuation range, the type of the current enterprise data analysis area is the type with obvious fluctuations.

[0030] In this embodiment, the stable type indicates that the cost data and output data of the current enterprise data analysis area are maintained in a stable state with a small fluctuation range. Within this range, the changes in enterprise data are relatively gentle and may be affected by some normal and controllable internal and external factors, such as slight fluctuations in raw material prices and slight changes in market demand. For enterprise data of this type, it is not easy to show abnormalities, and if there are abnormalities in the enterprise data of this type, they are easy to be identified; The type with fluctuations indicates that the cost data or output data of the enterprise data analysis area has a medium fluctuation range, indicating that there are relatively obvious changes in the enterprise data. It may be caused by factors such as intensified market competition, rising production costs, and adjustment of sales strategies, or there are minor data information statistical abnormalities. It is necessary to closely monitor these changes, evaluate the impact on the enterprise data status, and consider taking corresponding measures to address potential data risks, such as optimizing cost control, adjusting production plans, and expanding sales channels.

[0031] The type with obvious fluctuations and high fluctuations indicates that the cost data and output data of the enterprise data analysis area have a large fluctuation range, meaning that the abnormality degree of the enterprise data is relatively obvious and may be impacted by major factors, such as sudden changes in the macroeconomic environment, adjustment of industry policies, major business decision-making mistakes, or there are large data information statistical abnormalities.

[0032] By dividing the enterprise data analysis area into stable types and fluctuating types and further subdividing the degree of fluctuation, the present invention can perform data management and risk control more pertinently. For the stable type of enterprise data area, the enterprise can continue to operate according to the established strategy. For the areas with fluctuating types, the enterprise needs to invest corresponding resources for analysis according to the different degrees of fluctuation to reduce enterprise risks, improve the stability and predictability of enterprise data, and ensure the sustainable development of the enterprise.

[0033] Specifically, in the step S3, determining the data fluctuation parameters of the historical enterprise data of each of the enterprise data analysis areas includes: Calculating the average cost data value, the average cost data deviation, the average output data value, and the average output data deviation of the current enterprise data analysis area according to the historical enterprise data; Calculating the ratio of the average output data deviation to the average output data value and denoting it as the output fluctuation factor, and calculating the ratio of the average cost data deviation to the average cost data value and denoting it as the cost fluctuation factor; Determining the data fluctuation parameter according to the average value of the output fluctuation factor and the cost fluctuation factor.

[0034] In implementation, obtaining the cost data and output data of each day within the current detection period for calculation It can be understood that in actual financial analysis, the fluctuation of enterprise data may be affected by some special factors, such as seasonal factors or one-time events. The average deviation is relatively less sensitive to individual extreme values and can more accurately reflect the central tendency and stability of the data. While the standard deviation will amplify the influence of extreme values during the calculation process, resulting in overprocessing of the overall data fluctuation situation.

[0035] Please refer to Figure 3 As shown, it is the logic diagram of the abnormal trend type of the current enterprise data analysis area in the embodiment of the present invention. Determining the abnormal trend type of the current enterprise data analysis area according to the data fluctuation parameter includes: If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a high abnormal trend type; If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter of the corresponding enterprise data analysis area type, the abnormal trend type of the current enterprise data analysis area is a low abnormal trend type.

[0036] In implementation, the preset stable fluctuation parameter of the stable type is selected within the interval [0.05, 0.1], the preset stable fluctuation parameter of the existing fluctuation type is selected within the interval [0.08, 0.15], and the preset stable fluctuation parameter of the dominant fluctuation type is selected within the interval [0.12, 0.2].

[0037] Specifically, in the step S4, according to the abnormal trend type, an analysis method for determining the analysis area of enterprise data is selected to determine the abnormal enterprise data, including: If the abnormal trend type is a high abnormal trend type, the analysis method is to conduct a statistical inspection of enterprise data from all data sources; If the abnormal trend type is a low abnormal trend type, the analysis method is to extract part of the data sources for statistical inspection of enterprise data.

[0038] In implementation, extracting part of the data sources for statistical inspection means extracting 5% - 15% of all data sources in the enterprise data analysis area, preferably 10%, and the extraction method is not limited.

[0039] In the present invention, by calculating the average value and average deviation of cost data and output data, the output fluctuation factor and cost fluctuation factor are further determined, and the data fluctuation parameter is determined based on their average values. Considering data fluctuation factors from multiple aspects makes the determination of the data fluctuation parameter more comprehensive, has stronger representativeness for the whole, and can more accurately reflect the stability of enterprise data. At the same time, using the average deviation instead of the standard deviation reduces the sensitivity to extreme values, avoids the over-amplification of the overall data fluctuation situation caused by individual abnormal data, more accurately grasps the central tendency and stability of the data, and improves the accuracy of enterprise data anomaly recognition. According to the comparison between the data fluctuation parameter and the preset stable fluctuation parameter, the abnormal trend type is divided into two types: high anomaly and low anomaly, and analysis methods of checking all data sources and extracting part of the data are respectively matched for them, realizing differential processing of enterprise data with different abnormal degrees, which can not only efficiently discover potential problems, but also reasonably allocate analysis resources, improve the efficiency and pertinence of enterprise data analysis, and effectively monitor and respond to anomalies in the enterprise data situation.

[0040] Please refer to Figure 4 As shown, it is a flowchart of the statistical inspection of enterprise data for the data source in the embodiment of the present invention. In the step S4, the statistical inspection of enterprise data for the data source includes: Obtain the Excel document of the data source, identify and obtain the total number of formulas in the Excel document; Identify the hidden formulas in the Excel document and record the total number of hidden formulas; Calculate the call rate of infrequently used functions in the current data source Excel document based on the total number of hidden formulas and the total number of formulas; If the call rate of infrequently used functions exceeds the call threshold, determine that the current data source is a data source that needs to be traced.

[0041] In implementation, use the OpenXML SDK to transform and parse the XML metadata of Excel to identify hidden formulas; the call threshold is 10%. According to the definition of the materiality level in the International Standards on Auditing (ISA 240), generally, if the proportion of hidden formulas exceeds 10%, key verification is required.

[0042] Obtain the Excel file to be analyzed from the financial system and back it up. Use the OpenXML SDK to parse the Excel file into XML format, extract the metadata of relevant elements, parse the XML metadata, obtain the structural information of the workbook, worksheet, cells, and formulas, extract the formula content from the metadata, record the relevant positions and function compositions, establish a library of infrequently used functions, compare the extracted functions with the library, mark the infrequently used functions, analyze the logical structure of the formulas, evaluate the complexity and rationality, and check whether the calculation results meet the expectations.

[0043] Specifically, in the step S5, the process of verifying the enterprise data balance includes; Calculate the difference between the total output data and the total cost data and record it as the first balance difference; Subtract the total output data of all data sources from the first balance difference and record it as the enterprise data balance difference; If the enterprise data balance difference is greater than the preset abnormal threshold, determine that there is an abnormality in the target management enterprise; If the enterprise data balance difference is less than or equal to the preset abnormal threshold, determine that there is no abnormality in the target management enterprise.

[0044] In implementation, the preset abnormal threshold is determined based on referring to the data reports of other enterprises in the same industry and industry analysis reports, combined with the general situation of the balance sheet balance in the industry and the common error range, or by selecting twice the standard deviation of the average value of the historical enterprise data balance difference as the threshold.

[0045] In the present invention, through balance difference calculation and in combination with a dynamic threshold determination mechanism, the accuracy and efficiency of data anomaly detection are significantly improved. The step-by-step verification mode is adopted to screen layer by layer from basic data to global balance, effectively reducing the risk of deviation of a single indicator. The threshold is dynamically set based on industry benchmarks and historical data (such as twice the standard deviation), taking into account both horizontal industry commonalities and vertical enterprise characteristics to avoid one-size-fits-all misjudgments. At the same time, anomalies are quickly located through automated rules, greatly shortening the manual review cycle, and supporting flexible adaptation to different industry scales and risk preferences, providing enterprises with proactive, scalable data health monitoring and risk warning capabilities.

[0046] Specifically, determining whether the result of the current enterprise data statistics inspection is valid according to the verification result includes: If the verification result meets the type consistency condition, it is determined that the result of the enterprise data statistics inspection is valid, and the source of the abnormal data is traced according to the enterprise data statistics inspection; If the verification result does not meet the type consistency condition, it is determined that the result of the enterprise data statistics inspection is invalid, and the hidden formula in the enterprise data statistics inspection needs to be updated; Among them, the type consistency condition is that the target management enterprise has an anomaly and the anomaly trend type is a high anomaly trend type, or the target management enterprise has no anomaly and the anomaly trend type is a low anomaly trend type.

[0047] In the present invention, by setting type consistency or inconsistency conditions to determine whether the result is valid, it can ensure that only the accurately verified enterprise data statistics inspection results will be used for subsequent abnormal data traceability work, thereby improving the accuracy and reliability of abnormal traceability. At the same time, when the verification result does not meet the conditions, it is clearly proposed that the hidden formula needs to be updated, which not only helps to timely correct potential data anomalies or logical errors, but also reflects the dynamic optimization and self-improving ability of the solution, further enhancing the accuracy and adaptability of enterprise data intelligent analysis, and providing a solid technical support for enterprise data health monitoring and risk management.

[0048] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based apparatus that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0049] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An enterprise data intelligent analysis and traceability method, characterized in that, Include; Step S1: Collect enterprise data of the target managed enterprise to obtain initial enterprise data, and acquire the historical enterprise data of the target managed enterprise. Among them, the initial enterprise data includes cost data output data. Step S2: Classify the enterprise data sources of the target managed enterprise to obtain several types of enterprise data analysis regions, and acquire the historical enterprise data corresponding to each type of enterprise data analysis region. Step S3: Determine the abnormal trend type of the corresponding enterprise data analysis region according to the data fluctuation parameters determined by the enterprise data within the current detection period of each type of enterprise data analysis region. Step S4: Select and determine the analysis method for the corresponding enterprise data analysis region according to the abnormal trend type to determine the abnormal enterprise data. Among them, the analysis methods include extracting some data sources for enterprise data statistical inspection and conducting enterprise data statistical inspection on all data sources. Step S5: Conduct enterprise data balance verification on the data sources, and determine whether the result of the current enterprise data statistical inspection is valid according to the verification result to determine whether to trace the abnormal data sources.

2. The enterprise data intelligent analysis and traceability method according to claim 1, wherein, In the step S2, determine several types of enterprise data analysis regions in the target managed enterprise, including: Calculate the average cost data, standard deviation of cost data, output data, and standard deviation of output data of each data source in each inspection period respectively. Calculate the cost variation coefficient and output variation coefficient of each data source respectively. Determine the analysis type of the current enterprise data analysis region according to whether the cost variation coefficient and the output variation coefficient are within the stable fluctuation range.

3. The enterprise data intelligent analysis and traceability method according to claim 2, characterized in that Determine the types of the enterprise data analysis regions, including: If both the cost variation coefficient and the output variation coefficient are within the stable fluctuation range, the type of the current enterprise data analysis region is the stable type. If the cost variation coefficient is not within the stable fluctuation range or the output variation coefficient is not within the stable fluctuation range, the type of the current enterprise data analysis region is the type with fluctuations. If both the cost variation coefficient and the output variation coefficient are not within the stable fluctuation range, the type of the current enterprise data analysis region is the obvious fluctuation type.

4. The enterprise data intelligent analysis and traceability method according to claim 3, characterized in that In the step S3, determine the data fluctuation parameters determined by the historical enterprise data of each enterprise data analysis region, including: Calculate the average cost data, average deviation of cost data, average output data, and average deviation of output data of the current enterprise data analysis region according to the historical enterprise data. Calculate the ratio of the average deviation of output data to the average output data and record it as the output fluctuation factor, and calculate the ratio of the average deviation of cost data to the average cost data and record it as the cost fluctuation factor. Determine the data fluctuation parameter according to the average value of the output fluctuation factor and the cost fluctuation factor.

5. The enterprise data intelligent analysis and traceability method according to claim 4, characterized in that Determine the abnormal trend type of the current enterprise data analysis region according to the data fluctuation parameter, including: If the data fluctuation parameter is greater than the preset stable fluctuation parameter of the corresponding enterprise data analysis region type, the abnormal trend type of the current enterprise data analysis region is the high abnormal trend type. If the data fluctuation parameter is less than or equal to the preset stable fluctuation parameter corresponding to the enterprise data analysis region type, the abnormal trend type of the current enterprise data analysis region is the low abnormal trend type.

6. The enterprise data intelligent analysis and traceability method according to claim 5, wherein In the step S4, the analysis method corresponding to the enterprise data analysis region is selected according to the abnormal trend type to determine the abnormal enterprise data, including: If the abnormal trend type is the high abnormal trend type, the analysis method is to conduct enterprise data statistical inspection on all data sources; If the abnormal trend type is the low abnormal trend type, the analysis method is to extract part of the data sources to conduct enterprise data statistical inspection.

7. The enterprise data intelligent analysis and traceability method according to claim 6, characterized in that In the step S4, the enterprise data statistical inspection is carried out on the data sources, including Obtain the Excel document of the data source, identify and obtain the total number of formulas in the Excel document; Identify the hidden formulas in the Excel document and record the total number of hidden formulas; Calculate the call rate of infrequently used functions in the current data source Excel document through the total number of hidden formulas and the total number of formulas; If the call rate of infrequently used functions exceeds the call threshold, it is determined that the current data source is a data source to be traced.

8. The enterprise data intelligent analysis and traceability method according to claim 7, wherein In the step S5, the process of enterprise data balance verification includes; Calculate the output data total minus the cost data total and record it as the first balance difference; Subtract the output data total of all data sources from the first balance difference, and record it as the enterprise data balance difference; If the enterprise data balance difference is greater than the preset abnormal threshold, it is determined that the target management enterprise has an abnormality; If the enterprise data balance difference is less than or equal to the preset abnormal threshold, it is determined that the target management enterprise has no abnormality.

9. The enterprise data intelligent analysis and traceability method according to claim 8, wherein Determine whether the result of the current enterprise data statistical inspection is valid according to the verification result, including: If the verification result meets the type consistency condition, it is determined that the result of the enterprise data statistical inspection is valid, and the source of abnormal data is traced according to the enterprise data statistical inspection; If the verification result does not meet the type consistency condition, it is determined that the result of the enterprise data statistical inspection is invalid, and the hidden formulas in the enterprise data statistical inspection need to be updated; Among them, the type consistency condition is that the target management enterprise has an abnormality and the abnormal trend type is the high abnormal trend type, or the target management enterprise has no abnormality and the abnormal trend type is the low abnormal trend type.

Citation Information

Patent Citations

  • Intelligent operation and maintenance analysis method for enterprise information system

    CN106600115A

  • Checking management system

    CN109242416A

  • Power secondary equipment defect data mining method and system

    CN114297253A

  • Block chain-based job datamation template verification and traceability method

    CN117932674A

  • Method for grabbing target data in batch in mass data synchronization process

    CN118673034A