Method and equipment for analyzing and monitoring audit data

By employing multi-source data acquisition, distributed storage, and machine learning technologies, this approach addresses the issues of low data acquisition efficiency, poor fusion accuracy, and insufficient real-time monitoring in traditional audit data analysis. It enables efficient and accurate audit data analysis and real-time risk warning, adapting to the needs of enterprises of different sizes.

CN120929753APending Publication Date: 2025-11-11国网山东省电力公司日照供电公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511042538.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional audit data analysis methods suffer from low data collection efficiency, poor data fusion accuracy, and insufficient real-time monitoring capabilities when faced with massive and heterogeneous data, making it difficult to comprehensively obtain effective information and promptly identify potential risks.

Method used

It employs a multi-source data acquisition module, a distributed storage architecture, machine learning algorithms, and a real-time risk monitoring model, combined with semantic matching and DS evidence theory, to achieve data cleaning, fusion, analysis, and real-time early warning, and supports multi-dimensional data analysis and access control.

Benefits of technology

It improves the efficiency of data collection and analysis, enhances the accuracy of data fusion, enables real-time risk monitoring and early warning, improves the quality and value of audit work, and adapts to the needs of enterprises of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929753A_ABST
    Figure CN120929753A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audit data analysis, in particular to an audit data analysis monitoring method and device, and the method comprises the steps: S1, collecting multi-source data; s2, data cleaning and conversion; s3, data association based on semantics; s4, fusion algorithm optimization; s5, a multi-dimensional data analysis model is established; s6, machine learning auxiliary analysis; s7, real-time data acquisition and transmission; s8, carrying out real-time risk early warning; s9, a distributed storage architecture; and S10, performing data authority management. Through the multi-source data acquisition module and the optimized data acquisition mode, audit data can be quickly and comprehensively acquired. Through application of a distributed storage architecture and a cluster computing technology, the data storage and analysis efficiency is greatly improved, the time cost of data processing is reduced, an advanced data cleaning algorithm and a semantic-based data fusion technology are adopted, the data fusion accuracy is improved, and a reliable data basis is provided for subsequent analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audit data analysis technology, specifically to an audit data analysis and monitoring method and device. Background Technology

[0002] With the widespread application of information technology across various industries, the amount of data involved in auditing work has exploded, and the data types have become increasingly complex and diverse. Traditional audit data analysis methods have exposed many problems when faced with massive and heterogeneous data. For example, data collection efficiency is low, making it difficult to comprehensively obtain effective information from various data sources; data loss or mismatches are prone to occur during data fusion, affecting the accuracy of analysis results; and the ability to monitor audit data in real time is insufficient, making it impossible to detect potential risks and anomalies in a timely manner.

[0003] Some existing audit data processing methods and systems, such as the audit data collection and fusion method mentioned in application number 202310386857.0 (Audit Data Collection and Fusion Method, Device, Equipment and Storage Medium), while standardizing the data collection and fusion process to some extent, still lack flexibility and adaptability when facing complex and ever-changing data sources and dynamic data needs. The big data intelligent cloud audit method and system (Application number 201810097451.X) focuses on the overall cloud audit architecture construction, with relatively little in-depth research on refined analysis and real-time monitoring of audit data. The database security audit system, method and server (Application number 201810529452.7) mainly focuses on database-level security auditing, lacking comprehensive coverage in integrated audit data analysis and monitoring.

[0004] In response to the problems mentioned above, we propose an audit data analysis and monitoring method and equipment. Summary of the Invention

[0005] The purpose of this invention is to provide an audit data analysis and monitoring method and device to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: An audit data analysis and monitoring method includes the following steps: S1. Multi-source data acquisition: Utilizing a self-developed data acquisition module that includes components adapted to different data sources, raw audit data is obtained from various data sources, including relational databases, non-relational databases, file systems storing financial statements and business documents, and network logs recording server access and application operations, through customized data interface extraction, file reading and parsing, network traffic capture, and other adaptation methods. S2. Data Cleaning and Transformation: A data cleaning algorithm is used, which includes a sub-algorithm for identifying and removing noisy data, and a sub-algorithm for filling missing values ​​based on data type and business rules by selecting mean interpolation, linear interpolation, or machine learning prediction. This removes noisy data, fills in missing values, and converts text in different encoding formats to UTF-8 encoding and unifies different time formats into a standard time format, thereby achieving data format standardization. S3. Semantic-based data association: Construct an audit data semantic model that covers the meaning of data business and the annotation of relationships between data. Use semantic matching algorithms to semantically annotate data from different sources and achieve data fusion through semantic matching and association analysis. S4. Fusion Algorithm Optimization: An improved DS evidence theory fusion algorithm is adopted, which takes into account factors such as the historical accuracy of the data source and the data update frequency, assigns corresponding weights to each data source, handles data conflicts, and fuses multi-source data. S5. Multidimensional data analysis model: Establish a multidimensional data analysis model based on OLAP technology, which allows auditors to flexibly select multiple dimensions such as time, business department, and business type through the operation interface to perform slicing, dicing, drilling, and rotating operations on audit data, and generate and display the data analysis results in a real-time and visual manner. S6. Machine Learning-Assisted Analysis: Machine learning algorithms such as decision trees, random forests, and support vector machines are introduced. The model is trained using historical audit data, and the model parameters are optimized using methods such as cross-validation. The audit data is then classified, clustered, and analyzed predictively. S7. Real-time data acquisition and transmission: Build a real-time data acquisition and transmission framework based on message queue technology to realize the real-time acquisition and rapid transmission of audit data; S8. Real-time risk warning: Based on real-time collected data, a real-time risk monitoring model is established, and risk indicator thresholds are set. When the monitored data exceeds the threshold, an early warning mechanism containing detailed information on abnormal data and risk level is automatically triggered via SMS, email, or system pop-up. S9. Distributed storage architecture: It adopts a storage architecture that combines a distributed file system and a distributed database. Structured data is stored in the distributed database, and unstructured data is stored in the distributed file system and indexed. The distributed storage system is managed and maintained by the storage server. S10. Data Access Control: A role-based access control model is adopted to assign corresponding data access permissions to auditors based on their different roles and responsibilities, such as audit supervisors, auditors, and data analysts.

[0007] Preferably, the data acquisition module also has data acquisition task scheduling and monitoring functions, which can coordinate the work of each data source acquisition component according to a preset acquisition plan or a real-time triggering mechanism, and monitor the integrity and accuracy of data acquisition in real time.

[0008] Preferably, in the semantic-based data association step, the semantic annotation also includes metadata information such as the data source identifier and update time, in order to improve the accuracy and reliability of semantic matching.

[0009] Preferably, the improved DS evidence theory fusion algorithm also introduces a dynamic adjustment mechanism for data source credibility, which dynamically adjusts the weight of the data source based on its real-time performance to adapt to changes in the data environment.

[0010] Preferably, the multidimensional data analysis model also supports the function of customizing analysis dimensions and analysis indicators, allowing auditors to flexibly customize the data analysis model according to actual needs.

[0011] Preferably, the machine learning-assisted analysis step also includes a model performance evaluation and update mechanism, which periodically evaluates the performance of the trained machine learning model and automatically updates the model using new audit data when the model performance deteriorates.

[0012] Preferably, the real-time risk warning step also has a graded processing function for warning information, adopting different warning processing procedures and response mechanisms according to different risk levels.

[0013] Preferably, the distributed storage architecture also supports data backup and recovery functions, regularly backing up audit data and quickly recovering data when it is damaged or lost.

[0014] An audit data analysis and monitoring device, comprising: The data acquisition equipment consists of a high-performance server and customized data acquisition terminals. The server is used to coordinate and manage data acquisition tasks, and the data acquisition terminals include a database acquisition terminal equipped with a high-speed data extraction card, a file system acquisition terminal for file reading and parsing, and a network log acquisition terminal equipped with a high-performance network interface card. The data analysis equipment adopts a cluster computing architecture, consisting of multiple computing nodes equipped with high-performance CPUs, GPUs and large-capacity memory, used to run multidimensional data analysis models and machine learning algorithms; Data storage devices, based on a storage architecture that combines distributed file systems and distributed databases, include storage servers for managing and maintaining distributed storage systems, and storage arrays that provide large-capacity storage space; The data processing software integrates functional modules such as data acquisition, cleaning, fusion, analysis, and monitoring. It runs on the Linux operating system, has a user-friendly interface, and supports operation and management by auditors. Database management systems are used to efficiently manage and query structured data in distributed databases.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention, through a multi-source data acquisition module and optimized data acquisition methods, enables rapid and comprehensive acquisition of audit data. The application of distributed storage architecture and cluster computing technology significantly improves the efficiency of data storage and analysis, and reduces the time cost of data processing.

[0016] This invention employs advanced data cleaning algorithms and semantic-based data fusion technology to effectively solve problems such as data noise, missing values, and inconsistencies, thereby improving the accuracy of data fusion and providing a reliable data foundation for subsequent analysis.

[0017] This invention enables auditors to promptly detect anomalies in audit data and receive early warning information through a real-time data acquisition and transmission framework and a real-time risk monitoring model, thus helping to prevent risks in advance.

[0018] This invention provides auditors with a deeper and more comprehensive data analysis perspective through the application of multidimensional data analysis models and machine learning-assisted analysis techniques. It can discover potential problems and risks that are difficult to detect using traditional auditing methods, thereby improving the quality and value of auditing work.

[0019] The distributed storage architecture and system integration scheme of the present invention enable the devices and methods of the present invention to be easily expanded to meet the auditing needs of enterprises and institutions of different sizes, and to be seamlessly integrated with existing information systems, thus protecting the user's existing investment. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the process steps of the present invention; Figure 2 This is a schematic diagram of the system framework of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] like Figure 1 As shown, an audit data analysis and monitoring method is characterized by comprising the following steps: S1. Multi-source data acquisition: Utilizing a self-developed data acquisition module that includes components adapted to different data sources, raw audit data is obtained from various data sources such as relational databases, non-relational databases, file systems storing financial statements and business documents, and network logs recording server access and application operations through customized data interface extraction, file reading and parsing, network traffic capture, and other adaptation methods. Assume the data source collection is For data sources The amount of data it collects It can be represented as: Because different data sources (such as relational databases, file systems, network logs, etc.) have different data structures, storage methods, and access protocols, efficient and accurate data collection requires tailoring the approach for each data source. Design specific data acquisition functions This function comprehensively considers factors such as the data source interface type, data format, and collection frequency. Through customized data interface extraction, file reading and parsing, and network traffic capture adaptation methods, it transforms and extracts the data from the data source into a processable format, ultimately obtaining the collected data volume. .

[0023] The specific steps are as follows: First, select the corresponding data collection component based on the data source type (such as relational database, file system, network log, etc.); then, configure the parameters of the data collection component, such as database connection parameters, file read path, network port, etc.; finally, start the data collection component to collect data and record collection logs for subsequent troubleshooting.

[0024] S2. Data Cleaning and Transformation: A data cleaning algorithm is used, which includes a sub-algorithm for identifying and removing noisy data, and a sub-algorithm for filling missing values ​​based on data type and business rules by selecting mean interpolation, linear interpolation, or machine learning prediction. This removes noisy data, fills in missing values, and converts text in different encoding formats to UTF-8 encoding and unifies different time formats into a standard time format, thereby achieving data format standardization. S3. Semantic-based data association: Construct an audit data semantic model that covers the meaning of data business and the annotation of relationships between data. Use semantic matching algorithms to semantically annotate data from different sources and achieve data fusion through semantic matching and association analysis. Construct a semantic model for audit data, and set the data... The semantic annotation vector is For two data and their semantic similarity It can be calculated using cosine similarity: In audit data fusion, a vector space model is introduced to accurately determine the semantic relevance between data from different data sources. This involves... and Transform semantic annotations into semantic annotation vectors. and Each element in the vector and Represents the features of data in a specific semantic dimension. Cosine similarity measures the directional difference between two vectors by calculating the cosine of the angle between them. The smaller the angle, the closer the cosine value is to 1, indicating that the two data points are more similar in semantic features. (Numerator) The dot product of vectors reflects the degree of common characteristics of the vectors across all dimensions; the denominator Normalizing the vector length eliminates its influence on similarity calculation, thus obtaining accurate semantic similarity. .when When the data exceeds the set threshold, it is considered that the data is invalid. and It possesses semantic relevance and achieves data fusion through semantic matching and association analysis. The specific steps are as follows: First, based on the audit business needs and data characteristics, the rules and content of semantic annotation are determined, and each piece of collected data is semantically annotated to generate a semantic annotation vector; then, the semantic similarity between data is calculated, and data with similarity exceeding a threshold are initially associated; next, the initially associated data is manually or automatically reviewed to eliminate erroneous associations; finally, the reviewed and approved data are fused to generate a fused data set.

[0025] S4. Fusion Algorithm Optimization: An improved DS evidence theory fusion algorithm is adopted, which takes into account factors such as the historical accuracy of the data source and the data update frequency, assigns corresponding weights to each data source, handles data conflicts, and fuses multi-source data. An improved DS evidence theory fusion algorithm is adopted, assuming the data source set is... The basic probability allocation function for each data source is: The basic probability allocation function after fusion It can be calculated using the following formula: in, As a normalization constant, a corresponding weight is assigned to each data source based on factors such as the historical accuracy and data update frequency of the data source, to handle data conflicts and fuse multi-source data.

[0026] DS (Data Sources) evidence theory is an effective method for handling uncertain information fusion. In audit data fusion scenarios, each data source... The degree of support for different propositions (data attributes or categories) is determined by the basic probability assignment function. Indicated. In the formula This indicates that multiple data sources are used to represent the proposition. The degree of joint support needs to be determined by summing all possible joint support scenarios, since different data sources may have conflicting or duplicate support. This yields the sum of all joint support pairs that meet the criteria. However, this sum may become numerically abnormal due to data conflicts, so a normalization constant is introduced. Normalization is performed to ensure the basic probability allocation function after fusion. Satisfying the probability axiom (values ​​between 0 and 1, and the sum of the probabilities of all propositions being 1), a basic probability allocation function accurately reflecting the results of multi-source data fusion is obtained. The specific steps are as follows: First, collect historical data from each data source, analyze its accuracy and update frequency, and determine the initial weight of each data source; then, calculate the basic probability allocation function for each data item after fusion based on the improved DS evidence theory fusion algorithm formula; next, for conflicting data, conflict resolution is performed according to the weights and algorithm rules; finally, the accurate fused data is obtained.

[0027] S5. Multidimensional data analysis model: Establish a multidimensional data analysis model based on OLAP technology, which allows auditors to flexibly select multiple dimensions such as time, business department, and business type through the operation interface to perform slicing, dicing, drilling, and rotating operations on audit data, and generate and display the data analysis results in a real-time and visual manner. Establish a multidimensional data analysis model based on OLAP technology, and assume the audit data cube is... The dimension set is Through slicing operations, from the time dimension Data extraction It can be represented as: The multidimensional data analysis model abstracts audit data into data cubes in a multidimensional space. Each dimension This represents an attribute characteristic of the data (such as time, business department, business type, etc.). The essence of slicing is to divide data in a multi-dimensional space according to a specified dimension (such as the time dimension). ) specific values Select a subset of data that meets this condition. Mathematically, this is equivalent to selecting a multidimensional array (data cube). Apply a filter condition to retain those that meet the criteria. The data elements are used to obtain the sliced ​​data. Cutting, drilling, and rotating operations are similar; they are all based on manipulating and filtering different dimensions of multidimensional data to achieve data analysis from different perspectives.

[0028] The specific steps are as follows: First, determine the analysis dimensions and indicators based on the audit requirements and construct an audit data cube; then, receive the user's analysis operation instructions (such as slicing, dicing, drilling, rotating, etc.); next, perform the corresponding operations on the data cube according to the instructions to extract the required data; finally, calculate and process the extracted data and display it to the user in a visual way such as charts and reports.

[0029] S6. Machine Learning-Assisted Analysis: Machine learning algorithms such as decision trees, random forests, and support vector machines are introduced. The model is trained using historical audit data, and the model parameters are optimized using methods such as cross-validation. The audit data is then classified, clustered, and analyzed predictively. Introducing machine learning algorithms, taking the decision tree algorithm as an example, for the dataset... Its information entropy for: in For the number of categories, For category In the dataset The probability of occurrence is calculated. The optimal partitioning attribute is selected by calculating information gain, and a decision tree model is constructed to classify, cluster, and predict the audit data.

[0030] Information entropy is an important metric in information theory for measuring data uncertainty. In the problem of auditing data classification, the dataset... Information entropy Used to describe the degree of disorder in the distribution of categories in the dataset. For each category The probability of its occurrence The larger the value, the smaller its contribution to information entropy, because the more concentrated the data in that category, the lower the uncertainty; conversely, the smaller the probability, the higher the uncertainty. The negative sign in the formula ensures that the information entropy is non-negative, and the logarithmic function... This further amplifies the impact of probability differences on uncertainty. By calculating the information gain (the difference between the information entropy of the original dataset and the sum of the information entropies of the subsets after partitioning) after each attribute is used to divide the dataset, the attribute with the largest information gain is selected as the splitting attribute for the decision tree node. This approach can minimize data uncertainty and gradually build an effective decision tree model, enabling the classification, clustering, and predictive analysis of audit data.

[0031] The specific steps are as follows: First, the audit data is preprocessed and divided into training and testing sets. Then, a decision tree algorithm is applied to the training set to calculate the information gain of each attribute, and the attribute with the largest information gain is selected as the splitting attribute of the current node, gradually building a decision tree model. Next, the built model is evaluated using the testing set, and performance indicators such as accuracy and recall are calculated. Finally, the model is adjusted and optimized based on the evaluation results, and the optimized model is applied to the analysis of actual audit data for classification, clustering, or prediction.

[0032] S7. Real-time data acquisition and transmission: Build a real-time data acquisition and transmission framework based on message queue technology to realize the real-time acquisition and rapid transmission of audit data; Build a real-time data acquisition and transmission framework based on message queue technology (such as Kafka), assuming a data acquisition interval of . In the time interval The amount of data collected internally is Data transmission delay is ,satisfy: in, For data source The amount of data collected within this time interval enables real-time collection and rapid transmission of audit data.

[0033] In a real-time data acquisition and transmission system, the total amount of acquired data is the sum of the data acquired from multiple data sources within the same time interval. For each data source... Within a time interval \([t, t+\Deltat]\), a certain amount of data \(D_{i, t, t+\Deltat}\) is collected based on its own data generation patterns and collection mechanism. The total amount of data collected from all data sources within this time interval is summed to obtain the total amount of data collected by the entire system during that time period \(D_{t, t+\Deltat}\). This accurately describes the changes in data volume during the real-time data collection process, providing a foundation for subsequent data processing and analysis. The specific steps are as follows: First, deploy data collection agents at each data source end to collect data at set time intervals; then, encapsulate the collected data into message format and send it to a Kafka message queue; next, consume the data from the Kafka message queue at the data analysis end for subsequent processing; finally, monitor the data collection and transmission process to ensure the real-time performance and integrity of the data, and promptly address any delays or packet loss issues.

[0034] S8. Real-time risk warning: Based on real-time collected data, a real-time risk monitoring model is established, and risk indicator thresholds are set. When the monitored data exceeds the threshold, an early warning mechanism containing detailed information on abnormal data and risk level is automatically triggered via SMS, email, or system pop-up. Based on real-time data collection, a real-time risk monitoring model is established, with the risk indicator set as follows: Its threshold is ,when When this occurs, an early warning mechanism is triggered, and the warning information includes detailed information about the abnormal data and the risk level. Assuming risk indicators... Composed of multiple sub-indicators The weighted calculation yields:

[0035] in Sub-indicators The weight.

[0036] In auditing, a single risk sub-indicator often fails to fully reflect the risk profile of the audited entity. To achieve accurate risk assessment, multiple sub-indicators related to audit risk are combined. A comprehensive evaluation was conducted. Each sub-indicator... Different levels of contribution to overall risk are assigned corresponding weights. This reflects the differences in their importance. The weighting is determined based on audit expertise, historical data, and expert experience to ensure that significant risk factors carry a greater proportion in the risk indicator calculation. Through a weighted summation, the various sub-indicators are integrated into a comprehensive risk indicator. ,when Exceeding the preset threshold When this occurs, it indicates that the audited entity poses a high risk, thus triggering an early warning mechanism to promptly remind relevant personnel to take countermeasures. The specific steps are as follows: First, based on auditing expertise and historical data, determine the sub-indicators and weights of the risk monitoring model; then, collect data in real time and calculate the values ​​of each sub-indicator; finally, calculate the risk indicators according to the aforementioned formula. Finally, With threshold When comparing, In such cases, an alert is generated containing detailed information about the abnormal data and the risk level, and sent to relevant personnel via SMS, email, or system pop-ups.

[0037] S9. Distributed storage architecture: It adopts a storage architecture that combines a distributed file system and a distributed database (Cassandra). Structured data is stored in the distributed database, and unstructured data is stored in the distributed file system and indexed. The distributed storage system is managed and maintained by the storage server. A storage architecture combining a distributed file system (Ceph) and a distributed database (Cassandra) is adopted. Structured data is stored in the distributed database, and its data storage structure is as follows: Unstructured data is stored in a distributed file system and indexed, with the following index structure: The distributed storage system is managed and maintained by the storage server.

[0038] Audit data includes structured data (such as financial data and business records in database tables) and unstructured data (such as documents, images, and log files). Different types of data have different storage and access requirements. Distributed databases (such as Cassandra) offer high scalability, high availability, and flexible data models, making them suitable for storing structured data. This can be achieved by designing appropriate data table structures. It can efficiently store, query, and update structured data. Distributed file systems (such as Ceph) excel at handling large-scale unstructured data storage by establishing index structures. This allows for the rapid location and retrieval of unstructured data. Combining the two leverages their respective strengths to meet the diverse storage needs of audit data. Simultaneously, the distributed storage system is uniformly managed and maintained by the storage server, ensuring data security, consistency, and accessibility. The specific steps are as follows: First, the collected and processed data is categorized to determine whether it is structured or unstructured. Then, for structured data, it is inserted into the Cassandra distributed database according to the designed data table structure. For unstructured data, it is stored in the Ceph distributed file system, and corresponding indexes are created based on the data characteristics. Finally, the storage system is regularly maintained and managed, including data backup, disk space management, and index optimization.

[0039] S10. Data Access Control: A role-based access control model is adopted to assign corresponding data access permissions to auditors based on their different roles and responsibilities, such as audit supervisors, auditors, and data analysts.

[0040] Using a role-based access control (RBAC) model, let the set of roles be... The permission set is The role-permission mapping matrix is ​​as follows: ,like , indicating role Have permission Based on the different roles and responsibilities of audit supervisors, auditors, and data analysts, appropriate data access permissions are assigned to auditors.

[0041] In audit data management systems, precise control over access permissions for different users is necessary to ensure data security and compliance. The RBAC (Role-Based Access Control) model, as a mature access management solution, effectively reduces the complexity of access management by constructing a three-tier architecture of "user-role-permission." In this model, the set of roles...

[0042] It includes various roles defined in the system. The audit supervisor, as senior management, is responsible for overall planning of audit projects, approval of sensitive operations, and system configuration management. Auditors focus on performing specific audit tasks and have permissions such as data query and audit clue tracking. Data analysts focus on data mining and visualization and can perform multi-dimensional data analysis and report generation.

[0043] Permission set The system adopts a modular design, covering permissions throughout the entire data lifecycle: at the data operation level, it is subdivided into data query, data modification, and data import / export; at the business process level, it includes operation permissions such as audit plan preparation, audit report generation, and permission change application. To achieve precise mapping between roles and permissions, a role-permission mapping matrix is ​​introduced. Its dimensions are .when At that time, it indicates the role. Have permission ;when If the value is zero, it indicates that the permission is not granted. Taking the auditor role as an example, its corresponding data modification permission item in the matrix... And data query permission items The matrix clearly displays the permission groups that each role possesses.

[0044] In practical applications, system administrators assign appropriate roles to auditors based on their job descriptions and the access control logic of the RBAC model. A dynamic permission adjustment mechanism is also introduced; as an audit project enters different stages, the system can automatically adjust user roles according to preset rules. For example, during the audit report review stage, the lead auditor may be temporarily granted report signing permissions, which are automatically revoked upon completion of the review. This ensures efficient auditing while minimizing the risk of data leakage.

[0045] like Figure 2 As shown, an audit data analysis and monitoring device includes: The data acquisition equipment consists of a high-performance server and customized data acquisition terminals. The server is used to coordinate and manage data acquisition tasks, and the data acquisition terminals include a database acquisition terminal equipped with a high-speed data extraction card, a file system acquisition terminal for file reading and parsing, and a network log acquisition terminal equipped with a high-performance network interface card. The data analysis equipment adopts a cluster computing architecture, consisting of multiple computing nodes equipped with high-performance CPUs, GPUs and large-capacity memory, used to run multidimensional data analysis models and machine learning algorithms; Data storage devices, based on a storage architecture that combines distributed file systems and distributed databases, include storage servers for managing and maintaining distributed storage systems, and storage arrays that provide large-capacity storage space; The data processing software integrates functional modules such as data acquisition, cleaning, fusion, analysis, and monitoring. It runs on the Linux operating system, has a user-friendly interface, and supports operation and management by auditors. Database management systems are used to efficiently manage and query structured data in distributed databases.

[0046] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An audit data analysis and monitoring method, characterized in that, Includes the following steps: S1. Multi-source data acquisition: Utilizing a self-developed data acquisition module that includes components adapted to different data sources, raw audit data is obtained from various data sources, including relational databases, non-relational databases, file systems storing financial statements and business documents, and network logs recording server access and application operations, through customized data interface extraction, file reading and parsing, network traffic capture, and other adaptation methods. S2. Data Cleaning and Transformation: A data cleaning algorithm is used, which includes a sub-algorithm for identifying and removing noisy data, and a sub-algorithm for filling missing values ​​based on data type and business rules by selecting mean interpolation, linear interpolation, or machine learning prediction. This removes noisy data, fills in missing values, and converts text in different encoding formats to UTF-8 encoding and unifies different time formats into a standard time format, thereby achieving data format standardization. S3. Semantic-based data association: Construct an audit data semantic model that covers the meaning of data business and the annotation of relationships between data. Use semantic matching algorithms to semantically annotate data from different sources and achieve data fusion through semantic matching and association analysis. S4. Fusion Algorithm Optimization: An improved DS evidence theory fusion algorithm is adopted, which takes into account factors such as the historical accuracy of the data source and the data update frequency, assigns corresponding weights to each data source, handles data conflicts, and fuses multi-source data. S5. Multidimensional data analysis model: Establish a multidimensional data analysis model based on OLAP technology, which allows auditors to flexibly select multiple dimensions such as time, business department, and business type through the operation interface to perform slicing, dicing, drilling, and rotating operations on audit data, and generate and display the data analysis results in a real-time and visual manner. S6. Machine Learning-Assisted Analysis: Machine learning algorithms such as decision trees, random forests, and support vector machines are introduced. The model is trained using historical audit data, and the model parameters are optimized using methods such as cross-validation. The audit data is then classified, clustered, and analyzed predictively. S7. Real-time data acquisition and transmission: Build a real-time data acquisition and transmission framework based on message queue technology to realize the real-time acquisition and rapid transmission of audit data; S8. Real-time risk warning: Based on real-time collected data, a real-time risk monitoring model is established, and risk indicator thresholds are set. When the monitored data exceeds the threshold, an early warning mechanism containing detailed information on abnormal data and risk level is automatically triggered via SMS, email, or system pop-up. S9. Distributed storage architecture: It adopts a storage architecture that combines a distributed file system and a distributed database. Structured data is stored in the distributed database, and unstructured data is stored in the distributed file system and indexed. The distributed storage system is managed and maintained by the storage server. S10. Data Access Control: A role-based access control model is adopted to assign corresponding data access permissions to auditors based on their different roles and responsibilities, such as audit supervisors, auditors, and data analysts.

2. The audit data analysis and monitoring method according to claim 1, characterized in that, The data acquisition module also has data acquisition task scheduling and monitoring functions, which can coordinate the work of each data source acquisition component according to a preset acquisition plan or a real-time triggering mechanism, and monitor the integrity and accuracy of data acquisition in real time.

3. The audit data analysis and monitoring method according to claim 1, characterized in that, In the semantic-based data association step, semantic annotation also includes metadata information such as data source identifier and update time to improve the accuracy and reliability of semantic matching.

4. The audit data analysis and monitoring method according to claim 1, characterized in that, The improved DS evidence theory fusion algorithm also introduces a dynamic adjustment mechanism for data source credibility, which dynamically adjusts the weight of the data source based on its real-time performance to adapt to changes in the data environment.

5. The audit data analysis and monitoring method according to claim 1, characterized in that, The multidimensional data analysis model also supports the function of customizing analysis dimensions and indicators, allowing auditors to flexibly customize the data analysis model according to actual needs.

6. The audit data analysis and monitoring method according to claim 1, characterized in that, The machine learning-assisted analysis steps also include a model performance evaluation and update mechanism, which periodically evaluates the performance of the trained machine learning model and automatically updates the model using new audit data when the model performance deteriorates.

7. The audit data analysis and monitoring method according to claim 1, characterized in that, The real-time risk warning step also has a graded processing function for warning information, which adopts different warning processing procedures and response mechanisms according to different risk levels.

8. The audit data analysis and monitoring method according to claim 1, characterized in that, The distributed storage architecture also supports data backup and recovery functions, regularly backing up audit data and quickly recovering data in the event of damage or loss.

9. An audit data analysis and monitoring device, characterized in that, include: The data acquisition equipment consists of a high-performance server and customized data acquisition terminals. The server is used to coordinate and manage data acquisition tasks, and the data acquisition terminals include a database acquisition terminal equipped with a high-speed data extraction card, a file system acquisition terminal for file reading and parsing, and a network log acquisition terminal equipped with a high-performance network interface card. The data analysis equipment adopts a cluster computing architecture, consisting of multiple computing nodes equipped with high-performance CPUs, GPUs and large-capacity memory, used to run multidimensional data analysis models and machine learning algorithms; Data storage devices, based on a storage architecture that combines distributed file systems and distributed databases, include storage servers for managing and maintaining distributed storage systems, and storage arrays that provide large-capacity storage space; The data processing software integrates functional modules such as data acquisition, cleaning, fusion, analysis, and monitoring. It runs on the Linux operating system, has a user-friendly interface, and supports operation and management by auditors. Database management systems are used to efficiently manage and query structured data in distributed databases.

Citation Information

Patent Citations

  • Big data intelligent cloud auditing method and system

    CN108268656A

  • A database security auditing system, method and server

    CN108763957B

  • Audit data acquisition and fusion method and device, equipment and storage medium

    CN116596683A