Audit analysis method based on big data

Through distributed computing and in-memory computing technology, combined with multi-dimensional indicators and anomaly detection algorithm, a visual audit report is generated, which solves the problem of analyzing massive heterogeneous data, and realizes efficient and accurate data analysis and visual display of audit doubts, supporting enterprise compliance management.

CN120450858AActive Publication Date: 2025-08-08GUANGDONG POWER GRID CO LTD INFORMATION CENT

Patent Information

Application Number
CN202510385087.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-08
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the audit questionable data analysis scenario, how to effectively integrate and analyze massive heterogeneous data, identify abnormal points and violations, and clearly and intuitively display audit results on different dimensions to ensure the performance and efficiency of data processing and analysis.

Method used

The distributed computing framework is used for parallel processing, combined with memory computing technology for real-time analysis, multi-dimensional indicators and anomaly detection algorithm are used to identify potential doubts, generate visual audit reports, and support users' dynamic interaction through interactive technology, and automatically adjust algorithm parameters to optimize performance.

Benefits of technology

It improves audit efficiency and accuracy, enhances the visual presentation and application value of audit results, provides strong support for enterprise compliance management, and realizes closed-loop data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450858A_ABST
    Figure CN120450858A_ABST
Patent Text Reader

Abstract

The invention provides an audit analysis method based on big data, and the method comprises the steps: collecting the data of a marketing domain from the audit data penetration coverage conditions of a network, a province, a city and a district county, and obtaining the data conditions of doubtful audit points of each unit; for the audit data after the primary processing, performing real-time analysis on the data by adopting a memory computing technology, and extracting key indexes including a business type, a transaction amount and a transaction frequency to obtain a multi-dimensional index data set; according to the multi-dimensional index data set, a preset anomaly detection algorithm is adopted to analyze data, if the transaction amount or transaction frequency of a certain data point exceeds a preset threshold value, the data point is marked as a potential doubtful point, and a doubtful point candidate set is obtained; and for the illegal behavior detection result, displaying the audit result on different dimensions such as a map and a pie chart by adopting a preset visual display template, and generating a visual audit report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to an audit analysis method based on big data. Background Art

[0002] In the context of audit suspicion data analysis, the technical challenge is how to effectively integrate and analyze massive amounts of audit data. Audit data often originates from diverse business systems and databases, with varying formats and structures, and the data volume is enormous. Quickly and accurately extracting, cleaning, transforming, and loading this heterogeneous data, and conducting correlation analysis and mining, presents a complex technical challenge. Furthermore, the identification of audit suspicions requires comprehensive consideration of indicators across multiple dimensions. Designing appropriate algorithms and models to automatically identify suspicious anomalies and violations is also a pressing issue. Visualizing audit results also presents challenges. Clearly and intuitively presenting audit results across multiple dimensions, such as maps and pie charts, requires carefully designed visualization solutions and interactive technologies. In the context of massive data volumes, ensuring the performance and efficiency of data processing and analysis places higher demands on technical architecture and algorithm optimization. These intertwined technical challenges require in-depth research and innovation in multiple areas, including data processing, algorithm design, and visualization, to ultimately achieve efficient, accurate, and comprehensive audit suspicion data analysis. Summary of the Invention

[0003] The present invention provides an audit analysis method based on big data, which mainly includes: From the audit data penetration coverage of the network, province, city, district and county, collect data from the marketing domain and obtain the audit suspicious data of each unit; according to the preset distributed computing framework, divide the audit suspicious data into multiple data blocks and assign them to different computing nodes for parallel processing to obtain the preliminary processed audit data, which contains standardized and pre-processed transaction records, account information and business operation logs; for the preliminary processed audit data, use memory computing technology to perform real-time analysis on the data and extract key indicators, including business type, transaction amount, and transaction frequency, to obtain a multi-dimensional indicator data set; based on the multi-dimensional indicator data set, use the preset anomaly detection algorithm to analyze the data in the multi-dimensional indicator data set. If the transaction amount of a data point exceeds the preset threshold A, or the transaction frequency exceeds the preset threshold B, it is marked as a potential suspicious point to obtain a suspicious point candidate set; obtain the transaction data set, which contains transaction amount, transaction frequency, and transaction time. By setting the normal range of each indicator, if multiple indicators exceed the normal range at the same time , it is marked as a suspicious anomaly point, and the suspicious anomaly point is combined with the candidate set of suspicious points to generate the violation detection result, and the data processing time for violation detection is recorded; for the violation detection result, the preset visual display template is used to display the audit result on a map or pie chart to generate a visual audit report; based on the visual audit report, dynamic interaction between users and data is achieved through interactive technology. If the user selects a data point, the detailed information of the data point is displayed, including the type of violation, the amount involved, and the time of occurrence; if the data processing time for violation detection exceeds the preset threshold C, the algorithm parameters are automatically adjusted, including the number of trees or the anomaly score threshold D of the anomaly detection algorithm, to improve processing efficiency. Performance optimization runs through the entire data processing process to ensure the real-time responsiveness of the system; the preset interface specification is used to feed back the visual audit report to each business system to achieve closed-loop management of data. Closed-loop management includes pushing the violation detection results to the relevant business system, triggering the early warning or blocking mechanism, and recording the subsequent processing status to form a complete audit tracking chain.

[0004] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: The present invention discloses an audit analysis method based on big data. The method collects marketing domain data from multiple levels and uses a distributed computing framework to perform parallel processing and real-time analysis on suspicious data. A multi-dimensional indicator comprehensive model is used to detect violations, and a preset anomaly detection algorithm is combined to identify potential suspicious points. The present invention also integrates visual display and interactive technology to generate intuitive audit reports and support dynamic interaction between users and data. During the data processing process, the present invention can automatically adjust algorithm parameters according to performance requirements to ensure real-time responsiveness. Finally, by feeding back the analysis results to the business system, closed-loop management of audit data is achieved. This method not only improves audit efficiency and accuracy, but also enhances the visual presentation and application value of audit results, providing strong support for corporate compliance management. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 This is a flow chart of the big data-based audit analysis method of the present invention. DETAILED DESCRIPTION

[0006] To further understand the content of the present invention, the present invention is described in detail with reference to the accompanying drawings and examples. The present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.

[0007] like Figure 1 The audit analysis method based on big data in this embodiment may specifically include: S101. Collect data from the marketing domain based on the audit data penetration coverage of the network, province, city, district and county levels, and obtain the audit data of each unit.

[0008] An initial audit dataset is obtained from the network-level audit system. Data cleansing is performed on the initial audit dataset to remove duplicate records, address missing values, and unify the data format to form a cleaned dataset. For the cleaned dataset, a province-level filtering rule is pre-established, containing an identification field for each province. Based on the province-level filtering rule, audit data corresponding to each province is filtered from the cleaned dataset and linked with the network-level audit data to form a provincial-level linked dataset. The linking operation uses a database table join, with the linking field being the unit code. For the provincial-level linked dataset, a prefecture-level filtering rule is pre-established, containing an identification field for each prefecture-level city. Based on the prefecture-level filtering rule, audit data corresponding to each prefecture-level city is filtered from the provincial-level linked dataset. A determination is made as to whether the unit code field in the prefecture-level audit data matches the unit code field in the provincial-level linked dataset. If so, the prefecture-level audit data is merged with the provincial-level linked dataset to form a prefecture-level linked dataset. For the prefecture-level linked dataset, a district-level filtering rule is pre-established, containing an identification field for each district-level county. According to the district and county-level screening rules, the audit data corresponding to each district and county is screened out from the prefecture-level associated data set. The district and county-level audit data is verified to check whether the data is within the preset business rules. If it meets the preset business rules, a district and county-level data set is formed. For the district and county-level data set, marketing domain field rules are pre-established. The marketing domain field rules include field names related to marketing business. According to the marketing domain field rules, all marketing domain-related field data is screened out from the district and county-level data set to form a marketing domain data set. For the marketing domain data set, the interquartile range of each indicator of each unit is calculated. If the value of an indicator of a unit in the marketing domain data set exceeds the upper or lower threshold of the interquartile range, the indicator value of the unit is marked as suspicious data, forming a unit suspicious data set. The unit suspicious data set is grouped according to the unit code, and the suspicious data in each group is counted and summarized to generate an audit suspicious data report for each unit. The audit suspicious data status of each unit is obtained from the report.

[0009] For example, obtaining the initial audit dataset from the network-level audit system is the starting point of the entire audit process. This data may include various network device and system logs, as well as user activity records. For example, it may include router traffic data, firewall logs, and application access records. This raw data is often formatted in a variety of ways and may contain duplication, missing values, or errors. Data cleansing is a critical step in ensuring data quality. At this stage, you may encounter issues such as duplicate login records, missing IP addresses, and inconsistent timestamp formats. By removing duplicates, filling missing values, and standardizing the format, you can obtain a more standardized and reliable dataset. Establishing province-level filtering rules facilitates geographic classification of the data. Provinces may have unique identifiers, such as administrative division codes. This classification makes subsequent analysis more targeted and can identify regional issues or trends. Using database table joins to link provincial data with network-level data, using unit-coded fields, can enrich the data's content and context. For example, abnormal traffic patterns in a particular region may differ significantly from the overall trend, potentially indicating network security issues specific to that region.

[0010] Data screening at the prefecture, city, and district / county levels further refines data granularity. This multi-layered data structure enables auditors to comprehensively understand the situation from the macro to the micro level. For example, an unusually high increase in network traffic in a subregion may be related to a newly built data center. Data validation ensures data accuracy and compliance. Pre-established marketing domain field rules may include traffic thresholds and access frequency limits. The marketing domain dataset focuses on specific business areas. This may include metrics such as customer acquisition cost, customer retention rate, and marketing campaign effectiveness. Analyzing this data allows assessment of marketing strategy effectiveness and ROI (Return on Investment). Calculating the interquartile range for each unit's metrics helps identify outliers. For example, if a unit's customer acquisition cost is significantly higher than the upper interquartile range for the industry, this may indicate inefficient marketing or unusual spending. Data outside the interquartile range is identified as suspicious data, forming a unit suspicious data set. The unit suspicious data set is grouped by unit code, and the suspicious data within each group is counted and aggregated to generate a unit-specific audit suspicious data report. Finally, the generated audit suspicion data report provides management with a clear risk overview. This report may show that certain units have multiple abnormal indicators, such as unusually high network traffic, frequent security alerts, and excessive marketing expenditures, which may indicate a complex management problem or potential regulatory violations.

[0011] S102. According to a preset distributed computing framework, the audit suspicious data is divided into multiple data blocks, and distributed to different computing nodes for parallel processing to obtain preliminary processed audit data. The preliminary processed audit data includes standardized and pre-processed transaction records, account information, and business operation logs.

[0012] A distributed computing framework is used to segment suspicious audit data, generating multiple data blocks. Each block contains a specific range of transaction records, account information, and operation logs. These data blocks are distributed to different computing nodes through a pre-defined distribution mechanism, ensuring that each node processes a balanced amount of data. Each computing node performs standardization and preprocessing operations in parallel. The standardization process includes data format unification, field completion, and outlier handling. The preprocessing process includes data cleansing, deduplication, and feature extraction, resulting in processed data blocks. These processed data blocks are aggregated and integrated at a central node to generate a preliminary audit dataset containing standardized transaction records, account information, and business operation logs. This preliminary audit dataset undergoes a quality check. Any missing or anomaly data is traced back to the corresponding computing node for correction. Machine learning algorithms are used to further analyze this preliminary audit data, including clustering algorithms to identify unusual transaction patterns, classification algorithms to assess account risk, and association rule mining to analyze potential risks in operation logs, resulting in preliminary processed audit data.

[0013] For example, assume that the default distributed computing framework is Hadoop, and the suspicious data is a bank's transaction records for one year, totaling 10TB. The records contain fields such as transaction serial number, transaction time, transaction amount, account numbers of both parties, and transaction type. First, the 10TB transaction records are divided into 365 data blocks based on transaction date, each approximately 28GB in size. These blocks are stored using Hadoop's HDFS distributed file system. Parallel processing is then performed using the MapReduce programming model. In the Map phase, the 365 data blocks are distributed to 100 compute nodes in the cluster, with each node responsible for processing approximately three blocks. Each Map task first performs data cleansing. For example, records with a transaction amount less than 0 or greater than 100 million RMB are marked as suspicious, and records with an empty transaction type field are filtered out. Next, data normalization is performed, converting the transaction time field to a Unix timestamp format and the transaction account field to a fixed-length string format, with missing digits padded with zeros. For example, the account number "12345" is converted to "000000000012345." Next, data preprocessing is performed. For example, transactions are grouped by amount and the number and total amount of transactions in each group (e.g., 0-1,000 yuan, 1,001-10,000 yuan, 10,001-100,000 yuan, etc.) are calculated. This step can utilize algorithms such as fast group aggregation to perform local aggregation in the Map phase, reducing data transfer in the Reduce phase. In the Reduce phase, the output results of the 100 Map tasks are merged. For example, records marked as suspicious transactions from all Map tasks are aggregated to generate a list of suspicious transactions. Statistics for transaction amount groups from all Map tasks are merged to obtain the annual distribution of transaction amounts. Finally, preliminarily processed audit data is obtained, including business operation logs such as cleaned, standardized, and preprocessed suspicious transaction records, Unix timestamps and standardized account information for all transaction records, and annual transaction amount group statistics. After obtaining the preliminarily processed audit data, a quality check is performed. If any missing or abnormal data is found, the data is traced back to the corresponding compute node for correction. For example, business operation logs, including suspicious transaction records, Unix timestamps and standardized account information for all transaction records, and annual transaction amount grouping statistics, are reviewed to check for missing items or obvious anomalies. If any are found, the corresponding nodes in the data are re-analyzed. This data can be further used for subsequent audit analysis, such as analyzing suspicious transaction patterns through clustering algorithms and discovering unusual transaction relationships through association rule mining.

[0014] S103. For the audit data after preliminary processing, use in-memory computing technology to perform real-time analysis on the data, extract key indicators, including business type, transaction amount, and transaction frequency, and obtain a multi-dimensional indicator data set.

[0015] An in-memory computing framework is used to load the pre-processed audit data and perform data preprocessing, including data cleaning and formatting, to ensure data quality and obtain pre-processed audit data. Real-time analysis of the pre-processed audit data is performed using in-memory computing technology, and parallel processing capabilities are leveraged to rapidly extract key indicator data, including business type, transaction amount, and transaction frequency. Multi-dimensional aggregation operations are performed on the extracted key indicator data to generate a preliminary multi-dimensional indicator dataset. Based on this preliminary multi-dimensional indicator dataset, machine learning algorithms are used for pattern recognition to construct a transaction behavior model for predicting future transaction trends. The transaction behavior model is continuously optimized by combining historical and real-time data to obtain a multi-dimensional indicator dataset.

[0016] For example, assume that preliminarily processed audit data has been loaded into an in-memory database, such as Apache Spark or SAP HANA. Taking Apache Spark as an example, it can load preliminarily processed audit data into distributed memory, forming a Resilient Distributed Dataset (RDD). This approach significantly improves data processing speed and is particularly suitable for data that requires repeated access. During the data preprocessing phase, Spark's transformation operations can be used to cleanse and format data. For example, outliers in transaction records, such as negative or extremely large amounts, can be removed or corrected using filters. For inconsistent date and time formats, custom functions can be used to convert them to a standard format. These operations ensure data consistency and reliability, laying the foundation for subsequent analysis. Real-time analysis is a major advantage of in-memory computing. Using technologies such as Spark Streaming, pre-processed audit data can be processed in parallel to extract key metrics in real time. For example, a sliding window can be set to calculate the total transaction amount, average transaction amount, and transaction frequency over the past hour every five minutes. This approach can promptly detect abnormal transactions and improve audit efficiency. Multi-dimensional aggregation operations are a key step in data analysis. Using Spark SQL, you can perform complex aggregate queries on extracted indicator data to generate a preliminary multi-dimensional indicator dataset. For example, you can group transactions by business type, time period, and amount range, and calculate the number of transactions and total amount for each group. This allows you to observe transaction patterns from different perspectives and identify potential risk points. Machine learning algorithms play a key role in building transaction behavior models. Clustering algorithms in the Spark MLlib library, such as K-means, can be used to classify transaction data and identify distinct transaction patterns. For example, customers can be categorized as high-risk, medium-risk, or low-risk based on characteristics such as transaction frequency, amount, and time, facilitating targeted risk control. Predicting future transaction trends is a key aspect of model application. Time series analysis algorithms, such as the ARIMA model, can be used to predict future transaction volume and amount based on historical transaction data. This is crucial for fund management and risk assessment. Continuous model optimization is key to maintaining its effectiveness. Online learning algorithms, such as random forests, can be used to continuously incorporate new transaction data and update model parameters. At the same time, through regular cross-validation, we evaluate model accuracy and promptly adjust feature selection and algorithm parameters to adapt to the ever-changing trading environment. This in-memory computing-based audit data analysis method can significantly improve processing speed and analytical depth. Through real-time monitoring, multi-dimensional analysis, and intelligent prediction, potential risks can be identified earlier, improving the efficiency and accuracy of audit work. Furthermore, continuously optimized models can adapt to the ever-changing financial environment and provide reliable data support for audit decision-making.

[0017] S104. Based on the multi-dimensional indicator data set, a preset anomaly detection algorithm is used to analyze the data in the multi-dimensional indicator data set. If the transaction amount of a certain data point exceeds a preset threshold A, or the transaction frequency exceeds a preset threshold B, it is marked as a potential suspicious point, and a suspicious point candidate set is obtained.

[0018] Obtain a multi-dimensional indicator dataset containing key indicators such as business type, transaction amount, and transaction frequency. Analyze this multi-dimensional indicator dataset item by item using a pre-set anomaly detection algorithm. If the transaction amount of a data point exceeds a pre-set threshold A, mark the data point as a suspected amount anomaly. If the transaction frequency of a data point exceeds a pre-set threshold B, mark the data point as a suspected frequency anomaly. Aggregate potential suspicious points marked as suspected amount anomalies and suspected frequency anomalies to generate a candidate set of suspicious points.

[0019] For example, assume that the multi-dimensional indicator dataset described in the previous step has been obtained, containing information such as business type, transaction amount, transaction frequency, and number of unique users. Now, use the Isolation Forest algorithm for anomaly detection. For example, for business type "A001," its historical average transaction amount is 120,000 yuan, with a standard deviation of 15,000 yuan. Setting threshold A equal to 3 times the standard deviation, transactions exceeding 120,000 + 3 * 15,000 = 165,000 yuan or falling below 120,000 - 3 * 15,000 = 75,000 yuan will be flagged as potentially suspicious. Regarding transaction frequency, assume that user "user123" averaged 5 transactions over the past hour, with a standard deviation of 1. Setting threshold B equal to 3 times the standard deviation, transactions exceeding 5 + 3 * 1 = 8 or falling below 5 - 3 * 1 = 2 are considered anomalies. At the same time, combined with the rate of change of the number of independent users (UV), for example, if the growth rate of UV in the current time window (150) compared with the previous time window (135) exceeds the preset threshold (for example, 10%), that is, (150-135) / 135>10%, even if the transaction frequency of a single user does not exceed the threshold, there may be a risk of batch registration of fake accounts. Therefore, all transactions within these time windows are included as potential doubts. All marked potentially suspicious transactions (including anomalies detected based on amount, frequency, and UV change rate) are integrated into a suspicious candidate set in JSON format, for example: `{"suspicious_transactions": [{"user_id":"user123", "business_type":"A001", "amount": 180000, "timestamp":"2024-07-24 10:35:00", "reason":"Transaction amount exceeds the threshold"}, {"user_id":"user456","business_type":"B002", "count": 10, "timestamp":"2024-07-24 10:40:00","reason":"Transaction frequency exceeds the threshold"}, {"timestamp":"2024-07-24 10:30:00", "reason":"UV abnormal growth"}]}`.

[0020] S105. Obtain a transaction data set, which includes transaction amount, transaction frequency, and transaction time. Set a normal range for each indicator. If multiple indicators simultaneously exceed the normal range, they are marked as suspicious outliers. The suspicious outliers are combined with the suspicious point candidate set to generate violation detection results, and the data processing time for violation detection is recorded.

[0021] Obtain a transaction data set containing basic information including transaction amount, transaction frequency, and transaction time. Use interpolation to fill in missing values in the data in the transaction data set, and use the 3-times standard deviation method to eliminate outliers to ensure data quality. Set normal range thresholds for the transaction data in each transaction data set in the three dimensions of amount, frequency, and time. If the indicator values of multiple dimensions of a transaction record exceed the normal range threshold at the same time, it is marked as a suspicious anomaly. Further screen the marked suspicious anomalies, and use the random forest algorithm to predict violations based on the candidate set of suspicious points. If the prediction result is a violation, it is confirmed as a violation. Generate a violation detection result containing detailed information about the violation transaction record and the basis for violation determination, and record the data processing time for violation detection and store it in the database.

[0022] For example, acquiring a transaction dataset is the foundation of risk monitoring. This dataset typically contains key information such as transaction amount, frequency, and time. For example, a typical transaction record might include fields such as user ID, transaction amount, transaction timestamp, and transaction type. Data quality is crucial for accurate risk assessment. Interpolation methods are used to fill missing values, such as linear interpolation for missing transaction amounts in time series data. The three-standard-deviation rule is used to eliminate outliers, helping to eliminate extreme data points that could distort analytical results. For example, if a user's average transaction amount is 1,000 yuan and the standard deviation is 200 yuan, transactions exceeding 1,600 yuan or below 400 yuan might be flagged as abnormal and require further review. Determining normal range thresholds is a key step in detecting suspicious transactions. This typically involves analyzing historical transaction data and calculating the mean and standard deviation for each dimension (amount, frequency, and time). For example, if the normal transaction amount for a certain type of transaction is between 500 and 2,000 yuan, the normal frequency is 1 to 5 times per day, and the normal transaction time is between 9:00 a.m. and 10:00 p.m., then a transaction of 5,000 yuan occurring at 2:00 a.m. might be flagged as suspicious. Further screening of suspicious anomalies involves more complex analysis, involving factors such as a candidate set of suspicious transactions, transaction behavior characteristics, and historical transaction patterns. For example, if a user suddenly begins making a large number of small transactions, which is inconsistent with their historical transaction pattern, even if each individual transaction is within the normal range, it may be flagged as suspicious. Combining suspicious anomalies with the candidate set of suspicious transactions, the use of the random forest algorithm exemplifies the application of machine learning in risk detection. This algorithm constructs multiple decision trees to predict whether a transaction is a violation. Each tree may focus on different features, such as the degree of abnormality in transaction time, the deviation in transaction amount, or changes in recent user behavior. Ultimately, the algorithm combines the results of all decision trees to produce a probability of violation. Confirmation of the violation is a key output of the entire process. This includes not only the transaction record determined to be a violation, but also the basis for this determination. For example, a transaction might be deemed a violation due to its unusually high amount, unusual timing, and the user's recent history of similar suspicious behavior. Finally, the data processing time for each violation detection is recorded and stored in a database. Recording data processing time serves not only as a record but also provides a foundation for future analysis and model optimization, helping to evaluate and optimize the efficiency of the entire detection process. This complete process embodies a comprehensive risk management approach from data collection, cleaning, analysis, to decision-making, effectively combining statistical analysis and machine learning techniques to provide a strong guarantee for financial security.

[0023] S106. Based on the violation detection results, a preset visual display template is used to display the audit results on a map or pie chart to generate a visual audit report.

[0024] Obtain the violation detection results and perform data preprocessing, including data cleaning and data standardization. Through data cleaning, remove invalid and redundant data to ensure data quality; through data standardization, unify the data format to facilitate subsequent processing. Classify and aggregate the preprocessed data according to the preset visualization template requirements. For the map display dimension, classify the data according to geographic location information to generate a geographic distribution data set; for the pie chart display dimension, classify the data according to the violation type to generate a type distribution data set. Use a map generation algorithm to convert the geographic distribution data set into a map visualization image. The specific steps include loading the map base map, annotating data points, and rendering heat maps. Use a pie chart generation algorithm to convert the type distribution data set into a pie chart visualization image. The specific steps include drawing the pie chart, calculating the data percentage, and filling in the color. Integrate the generated map and pie chart visualization images according to the preset report format to generate a visual audit report.

[0025] For example, let's assume we have an audit dataset containing employee violations in different regions, including fields such as "region ID," "violation type," "violation amount," and "processing status." First, the system reads the dataset and uses predefined rules to perform a preliminary screening of violations. After identifying the violation records, the system generates a visualization report. For map visualization, the system calculates the total amount of violations for each region. For example, the total amount of violations in Region A is 125,800 yuan, Region B is 87,650 yuan, and Region C is 36,300 yuan. Then, using geographic information data, the violation amounts are mapped onto the corresponding map, with color depth representing the amount of violation. For example, using a Jet color ramp, Region A, with the highest violation amount, appears dark red, while regions with lower violation amounts appear light blue. For pie chart visualization, the system first groups the violation records based on the "Violation Type" field and calculates the percentage of each violation type. For example, "Violation Type 1" accounts for 45%, "Violation Type 2" accounts for 30%, "Violation Type 3" accounts for 15%, and "Other" accounts for 10%. Then, a pie chart is generated, where each sector represents a type of violation, and the size of the sector represents the proportion of that type of violation. After the pie chart is generated, the pie chart data needs to be associated with the map data. When the mouse hovers over a certain area on the map, the pie chart can be dynamically updated to show the distribution of violation types in that area. For example, when the mouse hovers over area A, a new pie chart pops up, showing the proportion of violation types in area A: "Violation type one" accounts for 60%, "Violation type two" accounts for 25%, "Violation type three" accounts for 10%, and "Other" accounts for 5%. Through this association analysis, it can be found that the problem of violation type one in area A is more prominent than the overall average level. Finally, the system integrates the generated map, pie chart, and related statistical tables into an HTML file to form a visual audit report.

[0026] S107. Based on the visual audit report, dynamic interaction between users and data is achieved through interactive technology. If the user selects a data point, detailed information about the data point will be displayed, including the type of violation, the amount involved, and the time of occurrence.

[0027] Visual audit report data is obtained and integrated interactive technology is used to make the chart interactive. Users can interact with the chart through clicks and hovering. If the user selects a data point, the detailed information display logic is triggered. The user's selected data point is determined and detailed information about the data point is extracted from the visual audit report data, including the type of violation, the amount involved, and the time of occurrence. A detailed information display interface is dynamically generated, presenting the extracted data point details on the interface, allowing users to quickly access information.

[0028] For example, the acquisition and preprocessing of audit report data forms the foundation of visual analysis. For example, during a company's annual audit, data on various types of violations may be collected. During preprocessing, data formats must be standardized, outliers removed, and field integrity ensured. For example, amounts can be standardized to ten thousand yuan, and time periods formatted as "year-month-day." The choice of visualization technology directly impacts the effectiveness of data presentation. For the distribution of violation types, pie charts or bar charts can intuitively display the proportion of each type. For trends in amounts over time, line charts are more suitable. If you want to simultaneously display violation type, amount, and time, a scatter plot is a good choice. The horizontal axis represents time, the vertical axis represents amount, and different colored dots represent different violation types. The integration of interactive features greatly enhances the user experience. For example, hovering over a scatter plot displays brief information, while clicking on it opens a detailed information window. This design allows users to quickly gain an overview of the overall situation while also delving deeper into specific cases of interest. After a user selects a data point, the system must quickly extract relevant information from the preprocessed data. For example, if a user clicks on a data point representing violation type 1, the system will immediately display detailed information about the incident, such as "Violation type: Violation type 1; Amount involved: 5 million yuan; Date of occurrence: July 15, 2023; Department involved: Business Department A; Processing status: Case filed for investigation." The design of the detailed information display interface should consider information hierarchy and readability. A card-style layout can be used to display different categories of information in separate blocks. Important information, such as violation type and amount, can be emphasized with larger font size or eye-catching colors, while less important information can be presented in smaller font size. Analyzing user interaction behavior is crucial for system optimization. By recording the violation types or amount ranges that users frequently focus on, data display priority can be adjusted. For example, if users are frequently viewing large-value violations, this information can be highlighted on the initial interface to improve user efficiency. This visual analysis system based on audit reports not only helps auditors identify issues more efficiently but also provides intuitive decision support for management. Through dynamic and interactive data presentation, complex audit information becomes easier to understand and analyze, thereby promoting the improvement of an enterprise's internal control system and enhancing risk management capabilities.

[0029] S108. If the data processing time for violation detection exceeds a preset threshold C, the algorithm parameters are automatically adjusted, including the number of trees or the anomaly score threshold D of the anomaly detection algorithm, to improve processing efficiency. Performance optimization runs through the entire data processing process to ensure the system's real-time responsiveness.

[0030] Obtain the data processing time for violation detection and compare it with the preset threshold C. If the data processing time for violation detection exceeds the preset threshold C, the parameter adjustment mechanism is triggered. Determine the anomaly detection algorithm parameters that need to be adjusted, including but not limited to the number of trees and the anomaly score threshold D of the isolation forest algorithm. According to the preset adjustment rules, calculate the new parameter values to ensure that the new parameters can significantly improve the processing efficiency. Apply the new parameter values to the anomaly detection algorithm and update the algorithm configuration. Monitor the operation of the updated algorithm in real time, obtain the new data processing time for violation detection, and compare it with the preset threshold C again. If the new data processing time for violation detection still exceeds the preset threshold C, repeat the above adjustment process until the real-time response requirements are met to ensure the real-time response capability of the system.

[0031] For example, assume the system has a preset data processing time threshold, C, of 500 milliseconds. A batch of 10,000 network traffic data items is being processed for anomaly detection. Initially, the Isolation Forest algorithm uses 100 trees and a threshold for the anomaly score, D, of 60. The system uses the performance monitoring module to monitor the data processing time for violation detection and discovers that the current batch of data is processing 650 milliseconds, exceeding the preset threshold, C, of 500 milliseconds. This triggers the automatic parameter adjustment mechanism. First, the system analyzes the relationship between historical data processing time and algorithm parameters and constructs a simple linear regression model: processing time = a * number of trees + b * anomaly score threshold + c. Training is performed using data from the past 10 batches (e.g., with numbers of trees of 80, 90, 110, and 120, and anomaly score thresholds of 55, 60, and 65, corresponding to processing times of 450, 520, 700, and 780 milliseconds, respectively). The resulting model parameters are a = 5, b = 100, and c = 50. Parameter optimization is then performed based on this model. To reduce processing time, the number of trees can be reduced or the anomaly score threshold D can be lowered (this may result in an increased false alarm rate, a trade-off). To simplify calculations and prioritize efficiency, the system decided to adjust the number of trees first. Setting a target processing time of 480 milliseconds (slightly below the threshold to allow for margin), the system substituted the following into the model formula: 480 Business Department A = Business Department A 5 Business Department A * Business Department A Number of New Trees Business Department A + Business Department A 100 Business Department A * Business Department A 6 Business Department A + Business Department A 50. The calculated number of new trees is approximately 74. The system automatically adjusted the number of trees in the Isolation Forest algorithm to 74. After this adjustment, the processing time for the next batch of tasks, also containing 10,000 network traffic data points, dropped to 490 milliseconds, meeting the preset threshold C requirement. At the same time, the system continuously monitors the false alarm rate. If it increases significantly (for example, exceeding the preset upper limit of 5%), it triggers an adjustment to the anomaly score threshold D. For example, the anomaly score threshold can be gradually increased from 60 to 62, 65, and so on. The impact on processing time and false alarm rate is observed, and a balance is selected. All of this data and model parameters are recorded in the system log and can be exported as a CSV file for subsequent offline analysis and model improvement.

[0032] S109. Use preset interface specifications to feed back visual audit reports to various business systems to achieve closed-loop management of data. Closed-loop management includes pushing violation detection results to relevant business systems, triggering early warning or blocking mechanisms, and recording subsequent processing to form a complete audit tracking chain.

[0033] After obtaining the visual audit report, data cleansing is performed to remove duplicate data and invalid fields, and missing values are interpolated to obtain cleaned data. The cleaned data is formatted, with standardized field types and encoding rules to ensure data quality and consistency, resulting in formatted data. This formatted data is then input into the audit analysis module, where SQL queries and Python scripts are used for classification and statistics, identifying abnormal transaction records and sending them to the relevant business systems. Upon receiving these abnormal transaction records, the relevant business systems determine the abnormality type of each record. If the abnormality type is high-risk, an alert is sent to the relevant responsible personnel via email and SMS. If the abnormality type is a serious violation, the relevant user account is immediately locked and their trading privileges are suspended, while a detailed log of the locking operation is recorded. Based on the detailed locking log, the business system generates a violation investigation task and assigns it to a designated risk control personnel. The risk control personnel, based on the task requirements, verify the violation and enter the verification results into the business system. Upon receiving the verification results, the business system updates the status of the abnormal transaction record, forming a complete audit trail that includes information such as the abnormality type, handling action, and verification results.

[0034] For example, the data processing process for audit reports begins with data cleansing after obtaining a visual report. This step aims to improve data quality and ensure the accuracy of subsequent analysis. For example, when processing a company's financial transaction records, duplicate numbers or invalid transactions with zero amounts may be discovered. By removing these data and interpolating missing transaction dates (for example, using the average of adjacent dates), data reliability can be significantly improved. Data formatting is a key step in ensuring data consistency. For example, cross-departmental transaction data may use different date formats or currency units. By uniformly converting dates to the "year-month-day" format and standardizing the currency unit to RMB, subsequent data analysis can be greatly simplified. During the audit analysis phase, the combination of SQL queries and Python scripts can effectively identify abnormal transactions. For example, SQL queries can be used to filter out transactions that exceed departmental budgets, and Python scripts can then be used to conduct in-depth analysis of these transactions, such as calculating the degree of deviation from historical averages to identify potential abnormal behavior. The push of abnormal transaction records demonstrates the collaborative work between systems. When the audit system detects a procurement transaction with an unusually large amount, it immediately pushes this record to the procurement management system. This real-time push mechanism significantly improves the efficiency of risk management. The process for handling exception records received by the business system demonstrates the importance of risk-tiered management. For example, a transaction that exceeds the authorized limit but only slightly might be classified as "high risk" by the system and notified by email and text message to the relevant supervisor for review. Conversely, a suspected large transfer might be classified as "serious violation," immediately freezing the account and notifying the risk control department. The review process by risk control personnel is a crucial link in the audit trail. For example, in the case of an abnormal account login, risk control personnel might review login logs, contact relevant personnel to verify the situation, and check for unauthorized operations. These findings are recorded in the system, forming a complete incident handling process. The entire process is designed to embody the principle of "preventing problems before they occur." Through automated data processing and analysis, combined with a multi-tiered risk response mechanism, the system can promptly detect and address anomalies before they escalate. This not only improves audit efficiency but also significantly enhances the company's risk management capabilities. Furthermore, a complete audit trail provides valuable data support for subsequent process optimization and risk assessment, helping companies continuously improve their internal control systems.

[0035] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations. In addition, the various different embodiments of the present invention can also be arbitrarily combined, as long as they do not violate the concept of the present invention, they should also be regarded as the content disclosed by the present invention.

Claims

1. The audit analysis method based on big data is characterized by: The method comprises: Collect data from the marketing domain and obtain the audit data of each unit based on the penetration coverage of audit data at the network, provincial, municipal, district and county levels; Based on the pre-set distributed computing framework, the audit data is divided into multiple data blocks and assigned to different computing nodes for parallel processing to obtain the preliminary processed audit data. The preliminary processed audit data includes standardized and pre-processed transaction records, account information, and business operation logs. For the audit data after preliminary processing, we use in-memory computing technology to conduct real-time analysis of the data, extract key indicators, including business type, transaction amount, and transaction frequency, and obtain a multi-dimensional indicator data set; Based on the multi-dimensional indicator data set, a preset anomaly detection algorithm is used to analyze the data in the multi-dimensional indicator data set. If the transaction amount of a certain data point exceeds the preset threshold A, or the transaction frequency exceeds the preset threshold B, it is marked as a potential suspicious point, and a candidate set of suspicious points is obtained; Obtain a transaction data set containing transaction amounts, transaction frequencies, and transaction times. Set a normal range for each indicator. If multiple indicators simultaneously exceed the normal range, mark them as suspicious outliers. Combine the suspicious outliers with the candidate set of suspicious points to generate violation detection results, and record the data processing time for violation detection. Based on the violation detection results, a preset visual display template is used to display the audit results on a map or pie chart to generate a visual audit report. Based on the visual audit report, interactive technology is used to enable dynamic interaction between users and data. If a user selects a data point, detailed information about that data point will be displayed, including the type of violation, the amount involved, and the time of occurrence. If the data processing time for violation detection exceeds the preset threshold C, the algorithm parameters are automatically adjusted, including the number of trees or the anomaly score threshold D of the anomaly detection algorithm, to improve processing efficiency. Performance optimization is carried out throughout the entire data processing process to ensure the system's real-time responsiveness. Visual audit reports are fed back to various business systems using preset interface specifications to achieve closed-loop management of data. Closed-loop management includes pushing violation detection results to relevant business systems, triggering early warning or blocking mechanisms, and recording subsequent processing to form a complete audit tracking chain.

2. The method according to claim 1, characterized in that The aforementioned audit data penetration coverage from the network, province, city, district and county levels is used to collect data from the marketing domain and obtain the audit data of each unit, including: Obtain the initial audit data set from the network-level audit system; Cleaning the initial audit data set, removing duplicate records, processing missing values, and unifying the data format to form a cleaned data set; For the cleaned data set, a province screening rule is pre-established, and the province screening rule includes an identification field for each province; According to the provincial screening rules, the audit data corresponding to each province is screened from the cleaned data set, and the audit data of each province is associated with the network-level audit data to form a provincial associated data set. The association operation uses a database table join operation, and the connection field is the unit code; For the provincial related data set, pre-establish prefecture-level city screening rules, which include an identification field for each prefecture-level city; According to the city-level screening rules, the audit data corresponding to each city is screened out from the provincial-level related data set; Determine whether the unit code field in the prefecture-level audit data is consistent with the unit code field in the provincial linked dataset. If they are consistent, merge the prefecture-level audit data with the provincial linked dataset to form a prefecture-level linked dataset. For the prefecture-level city-level associated data set, pre-establish district-level screening rules, which include an identification field for each district-level county; According to the district and county level screening rules, the audit data corresponding to each district and county is filtered out from the prefecture-level city-level associated data set; Verify the district and county audit data to see if it falls within the pre-set business rules. If it does, a district and county data set is generated. For the county-level dataset, pre-establish marketing domain field rules, which include field names related to marketing business; According to the marketing domain field rules, all marketing domain-related field data are filtered out from the district and county level data sets to form a marketing domain data set; For the marketing domain dataset, calculate the interquartile range of each indicator for each unit; If a certain indicator value of a unit in the marketing domain dataset exceeds the upper or lower threshold of the interquartile range, the indicator value of the unit is marked as suspicious data to form a unit suspicious data set; The unit suspicious data set is grouped according to the unit code, the suspicious data in each group is counted and summarized, an audit suspicious data report for each unit is generated, and the audit suspicious data situation of each unit is obtained from the report.

3. The method according to claim 1, characterized in that According to the preset distributed computing framework, the audit suspicious data is divided into multiple data blocks and assigned to different computing nodes for parallel processing to obtain the preliminary processed audit data. The preliminary processed audit data includes standardized and pre-processed transaction records, account information and business operation logs, including: Using a distributed computing framework, we segment audit data into multiple data blocks, each containing transaction records, account information, and operation logs within a specific range. The data blocks are distributed to different computing nodes through a preset distribution mechanism to ensure that the amount of data processed by each node is balanced; Each computing node performs standardization and preprocessing operations in parallel. The standardization process includes data format unification, field completion, and outlier processing. The preprocessing process includes data cleaning, deduplication, and feature extraction to obtain processed data blocks. Aggregating the processed data blocks to a central node for integration to generate a preliminary audit data set containing standardized transaction records, account information, and business operation logs; Perform a quality check on the preliminary audit data set. If any data is missing or abnormal, go back to the corresponding computing node for correction; The preliminary audit data set is further analyzed using machine learning algorithms, including clustering algorithms to identify abnormal transaction patterns, classification algorithms to assess account risks, and association rule mining to analyze potential risks in operation logs, to obtain preliminary processed audit data.

4. The method according to claim 1, wherein The audit data after preliminary processing is analyzed in real time using in-memory computing technology to extract key indicators, including business type, transaction amount, and transaction frequency, to obtain a multi-dimensional indicator data set, including: Use the in-memory computing framework to load the audit data after preliminary processing and perform data preprocessing, including data cleaning and data formatting, to ensure data quality and obtain preprocessed audit data; Performing real-time analysis on the pre-processed audit data through in-memory computing technology, and utilizing parallel processing capabilities to quickly extract key indicator data including business type, transaction amount, and transaction frequency; Perform multi-dimensional aggregation operations on the extracted key indicator data to generate a preliminary multi-dimensional indicator data set; Based on the preliminary multi-dimensional indicator dataset, machine learning algorithms are used to perform pattern recognition and construct a trading behavior model to predict future trading trends; By combining historical data and real-time data, the trading behavior model is continuously optimized to obtain a multi-dimensional indicator data set.

5. The method according to claim 1, wherein According to the multi-dimensional indicator data set, a preset anomaly detection algorithm is used to analyze the data in the multi-dimensional indicator data set. If the transaction amount of a certain data point exceeds a preset threshold A, or the transaction frequency exceeds a preset threshold B, it is marked as a potential suspicious point, and a candidate set of suspicious points is obtained, including: Obtain a multi-dimensional indicator data set, including key indicators such as business type, transaction amount, and transaction frequency; Using a preset anomaly detection algorithm, analyze the multi-dimensional indicator data set item by item; If the transaction amount of a data point exceeds the preset threshold A, the data point will be marked as an abnormal amount point; If the transaction frequency of a data point exceeds the preset threshold B, the data point will be marked as a suspected frequency anomaly; The potential suspicious points marked as abnormal amount suspicious points and abnormal frequency suspicious points are aggregated to generate a suspicious point candidate set.

6. The method according to claim 1, characterized in that The transaction data set is obtained, and the transaction data set includes transaction amount, transaction frequency, and transaction time. By setting a normal range for each indicator, if multiple indicators simultaneously exceed the normal range, they are marked as suspicious outliers. The suspicious outliers are combined with the suspicious point candidate set to generate violation detection results, and the data processing time for violation detection is recorded, including: Obtain a transaction dataset containing basic information such as transaction amount, transaction frequency, and transaction time; Interpolation was used to fill missing values in the transaction dataset, and the 3-times standard deviation method was used to eliminate outliers to ensure data quality. Set normal range thresholds for transaction data in each transaction data set in terms of amount, frequency, and time. If the indicator values of multiple dimensions of a transaction record simultaneously exceed the normal range thresholds, it will be marked as a suspicious anomaly. Further screening of the marked suspicious anomalies, combined with the candidate set of suspicious points, and using the random forest algorithm to predict violations; If the predicted result is a violation, it is confirmed as a violation; Generates violation detection results containing detailed information on violation transaction records and the basis for violation determination, records the data processing time for violation detection, and stores them in the database.

7. The method according to claim 1, characterized in that The aforementioned violation detection results are displayed on a map or pie chart using a preset visual display template to generate a visual audit report, including: Obtain violation detection results and perform data preprocessing, including data cleaning and data standardization; Through data cleaning, invalid and redundant data are removed to ensure data quality; Through data standardization, the data format is unified to facilitate subsequent processing; Classify and aggregate the pre-processed data according to the preset visualization template requirements; For map display dimensions, data is classified according to geographic location information to generate a geographic distribution dataset; For the pie chart display dimension, the data is classified according to the violation type to generate a type distribution data set; Use map generation algorithms to convert geographically distributed datasets into map visualization images. The specific steps include loading the map basemap, annotating data points, and rendering heat maps. A pie chart generation algorithm is used to convert the type distribution dataset into a pie chart visualization image. The specific steps include pie chart drawing, data proportion calculation, and color filling. The generated map and pie chart visualization images are integrated according to the preset report format to generate a visual audit report.

8. The method according to claim 1, characterized in that According to the visual audit report, dynamic interaction between users and data is achieved through interactive technology. If the user selects a data point, detailed information about the data point will be displayed, including the type of violation, the amount involved, and the time of occurrence, including: Obtain visual audit report data and use integrated interactive technology to make charts interactive. Users can interact with charts by clicking or hovering. If the user selects a data point, the detailed information display logic is triggered. Determine a data point selected by the user, and extract detailed information of the data point from the visual audit report data, including the type of violation, the amount involved, and the time of occurrence; Dynamically generate a detailed information display interface to present the extracted data point details on the interface, making it easier for users to quickly obtain information.

9. The method according to claim 1, characterized in that If the data processing time for violation detection exceeds a preset threshold C, the algorithm parameters are automatically adjusted, including the number of trees or the anomaly score threshold D of the anomaly detection algorithm, to improve processing efficiency. Performance optimization runs through the entire data processing process to ensure the system's real-time responsiveness, including: Obtaining the data processing time for violation detection and comparing it with the preset threshold C; If the data processing time for violation detection exceeds the preset threshold C, the parameter adjustment mechanism is triggered; Determine the anomaly detection algorithm parameters that need to be adjusted, including but not limited to the number of trees and the anomaly score threshold D of the isolation forest algorithm; Calculate new parameter values based on preset adjustment rules to ensure that the new parameters can significantly improve processing efficiency; Applying the new parameter values to the anomaly detection algorithm to update the algorithm configuration; Monitor the updated algorithm's operation in real time, obtain the new data processing time for violation detection, and compare it with the preset threshold C again; If the data processing time for new violation detection still exceeds the preset threshold C, the above adjustment process is repeated until the real-time response requirements are met to ensure the real-time response capability of the system.

10. The method according to claim 1, characterized in that The preset interface specifications are used to feed back visual audit reports to various business systems to achieve closed-loop management of data. Closed-loop management includes pushing violation detection results to relevant business systems, triggering early warning or blocking mechanisms, and recording subsequent processing to form a complete audit tracking chain, including: After obtaining the visual audit report, perform data cleaning, remove duplicate data and invalid fields, and interpolate missing values to obtain cleaned data; Formatting the cleaned data, unifying field types and encoding rules, ensuring data quality and consistency, and obtaining formatted data; Input the formatted data into the audit analysis module, use SQL query and Python script to perform classification and statistics, identify abnormal transaction records, and push them to relevant business systems; After receiving the abnormal transaction records, the relevant business system determines the abnormal type of each record one by one. If the abnormal type is high risk, an early warning message is sent to the relevant responsible person via email and SMS; If the exception type is a serious violation, the relevant user account will be immediately locked and their trading permissions will be suspended. A detailed log of the locking operation will also be recorded. The business system generates a violation investigation task based on the detailed log of the locking operation and assigns it to the designated risk control personnel; The risk control personnel shall verify the violations according to the task requirements and enter the verification results into the business system; After receiving the verification results, the business system updates the status of the abnormal transaction record to form a complete audit tracking chain, which includes information such as the abnormality type, processing operation, and verification results.

Citation Information

Patent Citations

  • Digital auditing system and method based on process automation technology

    CN111461668A

  • Audit doubtful point identification method and device based on data analysis, equipment and medium

    CN119180718A

Cited By

  • Data processing method and device, equipment and medium

    CN121478753A