Big data-based discovery auditing clue system and auditing clue collecting method thereof

CN122656783APending Publication Date: 2026-08-28GUANGZHOU SUNSHINE NAITE ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610761300.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

这导致审计工作中仍存在线索发现滞后、风险预警缺失等问题,难以满足新时代审计工作对全覆盖、精准性、时效性的要求

Benefits of technology

[0007] The beneficial effects of this invention are as follows: By employing a big data-based audit clue collection method, the audit process has been transformed from manual operation to a systematic and intelligent one, significantly improving the efficiency and accuracy of audit clue discovery. The system can automatically access and process multi-source heterogeneous electronic data, effectively reducing the time cost and labor intensity of manually sifting through massive amounts of data. Through the deep application of big data analysis engines and machine learning algorithms, the system enhances the ability to identify abnormal patterns and provide risk warnings, enabling audits to detect potential problems earlier. Simultaneously, this method reduces the omission of clues due to insufficient manual analysis experience or negligence, avoiding blind spots in audit coverage. The system supports in-depth economic responsibility audits and can be flexibly extended to performance audits, audits of assets involved in law enforcement activities, and other fields, achieving comprehensive coverage and efficient collaboration in audit operations, significantly improving the overall effectiveness of audit supervision and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122656783A_ABST
    Figure CN122656783A_ABST
Patent Text Reader

Abstract

The application provides a big data-based discovery audit clue system and an audit clue collection method, and belongs to the technical field of auditing.The method comprises the following steps: constructing an audit data integration platform, accessing a multi-source heterogeneous electronic data source of an audited unit, and generating a standardized audit data resource pool; through data cleaning and preprocessing technology, redundant data and error data in the resource pool are automatically corrected and filtered to generate a high-quality audit basic data set; through the big data-based audit clue collection method, the transformation of audit work from manual operation to systematization and intelligentization is realized, and the excavation efficiency and accuracy of audit clues are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes a big data-based system for discovering audit clues and a method for collecting audit clues, belonging to the field of audit technology. Background Technology

[0002] In the traditional audit supervision model, the discovery of audit leads relies heavily on manual methods, requiring auditors to spend a significant amount of time and energy reviewing and analyzing each page of paper documents submitted by the audited entities. This method is not only time-consuming and inefficient, but also demands extremely high levels of professional competence and experience from the analysts. Due to the scarcity of human resources, audit departments often have to concentrate their limited resources on economic responsibility audits, making it difficult to cover other important business areas such as performance audits and audits of assets involved in law enforcement activities, resulting in significant blind spots in audit coverage.

[0003] With the deepening of digital transformation, electronic data across industries is experiencing explosive growth, rendering traditional manual clue discovery methods inadequate for handling massive amounts of data. While some audit information tools exist, they largely remain at the level of data storage and simple querying, lacking the ability to deeply integrate multi-source heterogeneous electronic data, and failing to build an intelligent clue discovery system based on big data analytics. This results in problems such as delayed clue discovery and lack of risk warnings in audit work, making it difficult to meet the requirements of comprehensiveness, accuracy, and timeliness in audit work in the new era.

[0004] Against this backdrop, this solution achieves a breakthrough in the transformation from manual operation to systematic and intelligent auditing by building an audit clue collection system based on big data, providing an innovative solution for expanding the boundaries of audit business and improving the effectiveness of audit supervision. Summary of the Invention

[0005] This invention provides a big data-based system for discovering audit clues and a method for collecting audit clues, in order to solve the problems mentioned in the background section above: The present invention proposes an audit clue collection method based on big data for discovering audit clues, the method comprising: S1. Construct an audit data integration platform, connect to the multi-source heterogeneous electronic data sources of the audited entity, and generate a standardized audit data resource pool; through data cleaning and preprocessing technology, automatically correct and filter redundant and erroneous data in the resource pool to generate a high-quality audit basic dataset; S2. Based on the big data analysis engine, a full scan and feature extraction are performed on the audit basic dataset. An abnormal pattern, correlation and potential risk point in the data are identified through a distributed computing framework to generate an initial audit clue candidate set. At the same time, a clue labeling system is established to classify and label the candidate set in multiple dimensions. S3. Utilize machine learning algorithms to intelligently filter and prioritize the initial audit lead candidate set, combine it with the audit business rule base to verify the compliance of the leads, and generate a high-value audit lead set after eliminating invalid leads; use natural language processing technology to perform semantic parsing and structured transformation on the leads in unstructured data; S4. Push the collection of high-value audit leads to the audit business system to trigger an automated lead distribution process. Allocate the leads to different business modules according to the needs of the audit project. The different business modules include economic responsibility audit, performance audit, and financial audit of law enforcement activities. Simultaneously establish a full life cycle tracking mechanism for leads to record the lead processing status and feedback results. S5. The dynamic visualization dashboard displays the distribution and processing progress of audit leads in real time, supporting interactive analysis based on audit type, risk level, time dimension, and other conditions. The audit knowledge graph is automatically updated based on the lead processing results, optimizing the parameters of subsequent lead mining models and forming a data-driven closed-loop system for audit lead collection.

[0006] The present invention proposes a big data-based system for discovering audit clues, the system comprising: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are made to implement the method described in any one of the above.

[0007] The beneficial effects of this invention are as follows: By employing a big data-based audit clue collection method, the audit process has been transformed from manual operation to a systematic and intelligent one, significantly improving the efficiency and accuracy of audit clue discovery. The system can automatically access and process multi-source heterogeneous electronic data, effectively reducing the time cost and labor intensity of manually sifting through massive amounts of data. Through the deep application of big data analysis engines and machine learning algorithms, the system enhances the ability to identify abnormal patterns and provide risk warnings, enabling audits to detect potential problems earlier. Simultaneously, this method reduces the omission of clues due to insufficient manual analysis experience or negligence, avoiding blind spots in audit coverage. The system supports in-depth economic responsibility audits and can be flexibly extended to performance audits, audits of assets involved in law enforcement activities, and other fields, achieving comprehensive coverage and efficient collaboration in audit operations, significantly improving the overall effectiveness of audit supervision and resource utilization efficiency. Attached Figure Description

[0008] Figure 1 This is a diagram illustrating the steps of the method described in this invention. Detailed Implementation

[0009] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0010] One embodiment of the present invention, such as Figure 1 As shown, the audit clue collection method of the big data-based audit clue discovery system includes: S1. Construct an audit data integration platform, connect to the multi-source heterogeneous electronic data sources of the audited entity, and generate a standardized audit data resource pool; through data cleaning and preprocessing technology, automatically correct and filter redundant and erroneous data in the resource pool to generate a high-quality audit basic dataset; S2. Based on the big data analysis engine, a full scan and feature extraction are performed on the audit basic dataset. An abnormal pattern, correlation and potential risk point in the data are identified through a distributed computing framework to generate an initial audit clue candidate set. At the same time, a clue labeling system is established to classify and label the candidate set in multiple dimensions. S3. Utilize machine learning algorithms to intelligently filter and prioritize the initial audit lead candidate set, combine it with the audit business rule base to verify the compliance of the leads, and generate a high-value audit lead set after eliminating invalid leads; use natural language processing technology to perform semantic parsing and structured transformation on the leads in unstructured data; S4. Push the collection of high-value audit leads to the audit business system to trigger an automated lead distribution process. Allocate the leads to different business modules according to the needs of the audit project. The different business modules include economic responsibility audit, performance audit, and financial audit of law enforcement activities. Simultaneously establish a full life cycle tracking mechanism for leads to record the lead processing status and feedback results. S5. The dynamic visualization dashboard displays the distribution and processing progress of audit leads in real time, supporting interactive analysis based on audit type, risk level, time dimension, and other conditions. The audit knowledge graph is automatically updated based on the lead processing results, optimizing the parameters of subsequent lead mining models and forming a data-driven closed-loop system for audit lead collection.

[0011] The working principle of the above technical solution is as follows: This method first builds a dedicated integration platform to aggregate heterogeneous data generated by various business systems, unifies the data format, and completes the collection and integration. Then, through automated cleaning and preprocessing, redundant and erroneous content is removed, and the data is organized into basic audit data that can be uniformly accessed. Relying on a big data analysis engine, the entire organized data is comprehensively scanned and its features are extracted. With the help of a distributed computing framework, the inherent correlations and fluctuations in the data are deeply explored to identify abnormal behaviors and potential risks. After summarizing and sorting, an initial candidate set of clues is formed, and a hierarchical labeling system is built to complete multi-dimensional classification and labeling. Machine learning algorithms are introduced to quantitatively evaluate and rank the candidate clues, and compliance verification is completed according to industry audit standards to filter out invalid and redundant clues. At the same time, semantic parsing and structural transformation are performed on textual information such as reports and official documents to supplement and improve clue details and collect them into high-value clue resources. The selected high-quality clues are pushed to the business system, and the distribution is automatically completed according to the audit business category and job responsibilities. The status of each stage of clue flow, review and rectification is recorded throughout the process. By relying on visual carriers to intuitively present the overall distribution and processing progress of clues, by relying on multi-condition filtering to achieve layered and penetrating analysis, and by combining offline verification results to iteratively update the knowledge graph content, and simultaneously adjust the operating parameters of the mining model, and by relying on actual business data to continuously feed back into the system's computing power, the entire audit clue process, from data collection, risk mining, screening and distribution to iterative optimization, forms a complete operational chain, realizing intelligent closed-loop operation of the entire process.

[0012] The above technical solutions achieve the following effects: By unifying the access to and automatically cleaning multi-source heterogeneous data, the integrity and availability of audit data are effectively improved, reducing errors and time consumption caused by manual data processing. Through full-volume scanning and distributed computing using a big data engine, the efficiency of identifying abnormal patterns and potential risks is enhanced, reducing omissions during manual investigation. Intelligent filtering and sorting via machine learning accelerates the location of high-value leads, reducing the waste of audit resources on invalid leads. Semantic parsing of unstructured data broadens the range of lead sources, reducing the unusability of textual information. Automated distribution and full lifecycle tracking improve the efficiency of lead flow and processing, preventing lead stagnation and unclear responsibilities. Through iterative optimization of visual dashboards and knowledge graphs, the real-time nature of audit supervision is maintained while continuously improving the accuracy of model mining, preventing lead identification failures due to untimely data updates, and overall enhancing the coverage and intelligence level of audit work.

[0013] In one embodiment of the present invention, S1 includes: S11. Establish a dedicated data integration platform for public security audits, connecting multiple heterogeneous electronic data sources of the audited entities. These heterogeneous electronic data sources include financial systems, asset management systems, law enforcement and case handling systems, and police cloud platforms, enabling real-time connection and batch import of all data. S12. Perform unified format conversion on the accessed raw data, unify field naming, data type and storage structure, complete cross-system, cross-business and cross-department data field mapping and data fusion, and form a standardized audit data resource pool; S13. Start the automated data cleaning engine, identify problematic data in the resource pool, including duplicate records, missing fields, abnormal values ​​and logical conflicts, and perform duplicate data merging, blank field completion and abnormal data filtering operations. S14. Perform integrity, consistency and accuracy checks on the cleaned data, repair conflicting data, supplement key business information and improve data availability; output a high-quality audit basic dataset covering all scenarios, including financial revenue and expenditure, law enforcement assets, economic responsibility and key projects.

[0014] The working principle and effects of the above technical solution are as follows: By building a dedicated data integration platform for public security auditing and connecting multiple business systems, it can stably complete the real-time docking and batch import of full data, improving the completeness and timeliness of data collection and reducing the time consumption and errors caused by manual collection. Through unified data format and field mapping integration, it enhances the universality and compatibility of multi-source data, reducing adaptation barriers in cross-system data calls. Through automated data cleaning and processing, it quickly removes abnormal and conflicting data, improving the regularity of basic data and preventing dirty data from affecting subsequent audit analysis. Through multiple verifications and information supplementation, it enhances the credibility and usability of data, avoiding audit judgment bias due to missing or contradictory data. It can cover the data needs of all scenarios, including financial revenue and expenditure, law enforcement, property, economic responsibility, and key projects, and provides stable and reliable data support for subsequent clue mining, thus improving the overall quality of preliminary data preparation for audit work.

[0015] In one embodiment of the present invention, S2 includes: S21. Activate the distributed big data analysis engine to perform a full data scan on the high-quality audit basic dataset and extract core data features, including fund flow, approval nodes, business cycles, related entities and operation records. S22. Utilize a distributed computing framework to conduct cross-table association analysis, trend comparison analysis, and abnormal fluctuation analysis to identify abnormal patterns, hidden relationships, and high-risk behavior nodes that deviate from normal business logic. S23. Summarize all identified suspicious data fragments, risky behaviors and violations to form a complete initial audit clue candidate set covering the entire business scope; S24. Construct a multi-level clue tagging system, which includes risk level, business type, clue source, problem type and involved departments, and standardize the tag dimensions and labeling rules; S25. For each clue in the initial audit clue candidate set, complete the label assignment and data binding to form a structured candidate clue set.

[0016] The working principle and effects of the above technical solution are as follows: By conducting a full scan through a distributed big data analysis engine, core features such as fund flows and business behaviors can be comprehensively extracted, improving the comprehensiveness of data feature coverage and reducing the risk of omissions caused by manual extraction. Multi-dimensional analysis through a distributed computing framework can accurately identify abnormal patterns and hidden correlations, enhancing the depth of risk discovery and reducing hidden problems that are difficult to detect through manual verification. By aggregating suspicious data to form a complete candidate set, the scope of clue sources is expanded, preventing risk points from being overlooked. By constructing a multi-level tagging system and completing binding, the standardization of clue management and retrieval efficiency are enhanced, reducing the cost of subsequent screening and processing. It can quickly generate a structured clue set covering the entire business and provide clear classification for subsequent intelligent screening, thus improving the overall efficiency and accuracy of audit clue mining.

[0017] In one embodiment of the present invention, step S22 includes: S221. Utilize a distributed computing framework to conduct cross-table association analysis on core data features, generating multi-dimensional data association results; conduct trend comparison analysis on the multi-dimensional data association results, generating business trend change results; S222. Conduct abnormal fluctuation analysis on the results of business trend changes and generate data anomaly detection results; determine abnormal patterns that deviate from the normal business logic based on the data anomaly detection results; S223. Trace hidden relationships between data based on abnormal patterns; mark high-risk behavior nodes in the corresponding data based on hidden relationships.

[0018] The working principle and effects of the above technical solution are as follows: By conducting cross-table association and trend comparison analysis through a distributed computing framework, the inherent connections of business data can be fully restored, improving the comprehensiveness of risk identification and reducing judgment bias caused by single-dimensional analysis. By conducting abnormal fluctuation analysis on trend changes, abnormal data states can be accurately captured, improving the sensitivity of problem detection and reducing hidden risks that conventional verification cannot reach. Through anomaly pattern judgment and hidden correlation tracing, the depth of risk positioning is enhanced, avoiding the problem of scattered data failing to form effective clues. By marking high-risk behavior nodes, the risk direction is made clearer, avoiding deviations in audit direction due to ambiguous relationships. This approach can comprehensively uncover violations behind the data while significantly shortening manual investigation time, thus improving the overall accuracy and efficiency of audit clue discovery.

[0019] In one embodiment of the present invention, step S221 includes: The scheduled distributed computing framework performs cross-table data association operations on core data features and generates intermediate results of cross-table data association. Perform multi-dimensional data integration on intermediate results of cross-table data association to generate multi-dimensional data association results; Retrieve historical business data and the correlation results of multi-dimensional data to perform a periodic comparison calculation and generate periodic data comparison results; Continuous cycle trend extrapolation is performed based on the comparison results of data from the same period to generate results on changes in business trends.

[0020] The working principle and effects of the above technical solution are as follows: By conducting cross-table data association operations through a distributed computing framework, it can break down barriers between scattered data, improve the integrity of data association, and reduce analytical blind spots caused by data fragmentation. Through multi-dimensional data integration, it enhances the systematic nature of data presentation and reduces situations where fragmented data cannot be centrally analyzed. By introducing historical data for year-on-year comparison operations, it strengthens the rationality of data references and avoids making biased judgments based solely on current data. Through continuous periodic trend extrapolation, it improves the foresight of business change perception and avoids short-term fluctuations misleading risk assessments. It can fully present the inherent relationships and dynamic trends of data, and provide a stable and reliable analytical foundation for subsequent anomaly identification, comprehensively improving the accuracy and objectivity of audit data analysis.

[0021] In one embodiment of the present invention, S3 includes: S31. Load the trained machine learning classification and ranking model, perform feature quantification calculation on the initial audit clue candidate set, and automatically complete the clue risk scoring and importance assessment. S32. Based on the scoring results, complete the stratification and priority sorting of clues, distinguish between key clues for verification, general clues of concern and clues to be verified, and plan the order of clue processing. S33. Retrieve the public security audit business rule database, compare the content of each clue with the degree of matching of compliance clauses, audit standards, and supervision requirements, mark invalid clues that do not comply with the rules and remove them; S34. Integrate and supplement the information of the retained valid clues, improve the key contents such as the objects, time, amount and behavior involved in the clues, and form a set of high-value audit clues; S35. Use natural language processing technology to perform semantic analysis on unstructured data, including reports, official documents, contracts and records, extract key information and convert it into structured fields to supplement and improve the clue content.

[0022] The working principle and effects of the above technical solution are as follows: By using machine learning models to quantify features and score risks of leads, the objectivity of lead evaluation is improved, reducing biases caused by subjective human judgment. Lead stratification and prioritization optimize the allocation of verification resources, reducing the waste of audit effort on non-priority leads. Rule-by-rule comparison and verification accurately eliminates invalid leads, preventing invalid information from interfering with subsequent audit processes. Information integration and detail completion enhance the completeness and readability of leads, preventing information gaps from hindering the progress of verification. Natural language processing parses unstructured data, broadening the dimensions of lead sources and reducing the problem of unusable textual information. This approach efficiently filters high-value leads while significantly reducing lead preprocessing time, overall improving the accuracy and efficiency of audit lead selection.

[0023] In one embodiment of the present invention, S35 includes: Text splitting and paragraph segmentation are performed on unstructured data such as reports, official documents, contracts, and transcripts to generate text units to be parsed; Semantic feature extraction is performed on text units to extract information such as time, subject, amount, cause, behavior, and description from the text; The extracted information is categorized and formatted to generate standardized structured fields; these structured fields are then linked with the corresponding audit leads to supplement and improve the complete content of the leads. Once the supplemented clues are completed, they will be included in the high-value audit clue set, thus completing the transformation of unstructured data into usable clues.

[0024] The working principle and effects of the above technical solution are as follows: By performing text splitting and paragraph segmentation on unstructured data, the text processing units can be refined, improving the detail of subsequent analysis and reducing information omissions caused by directly parsing large blocks of text. Through semantic feature extraction, key business information can be accurately captured, enhancing the comprehensiveness of information acquisition and reducing the time spent on manual review and organization. Through field classification and format standardization, standardized structured content is generated, improving information reusability and preventing unstructured information from being unable to participate in system calculations. By associating and splicing structured fields with clues, clue details are improved, preventing incomplete information from affecting verification and judgment. By incorporating the supplemented clues into a high-value set, the source channels of clues are broadened, fully utilizing various text data resources and continuously enriching the dimensions of audit clues, thereby improving the overall depth and completeness of clue mining.

[0025] In one embodiment of the present invention, step S4 includes: S41. High-value audit leads are pushed to the public security audit supervision and management business system via system interface and enter the lead processing process. S42. Read the audit project configuration, personnel responsibility list and supervision point setting information, and classify the clues according to business types such as economic responsibility audit, performance audit, law enforcement activity financial audit, and key project audit; S43. Accurately match clues with processing personnel according to the jurisdiction, police department, and responsible position, generate a standardized clue allocation list, and start automatic distribution; S44. Generate a unique identifier for each clue, bind it to all nodes in the entire process from clue generation, allocation, review, processing to feedback, and establish an unalterable tracking link; S45. Collect the progress of clue processing, verification results, review opinions and rectification status in real time, write them to the system log and update the clue status dynamically to form a traceable record of the whole process.

[0026] The working principle and effects of the above technical solution are as follows: By strategically pushing high-value audit leads through system interfaces, the speed of lead flow is accelerated, the smoothness of lead handover is improved, and the risk of delays and loss caused by manual transmission is reduced. By categorizing leads according to business type, the level of organization in audit work is enhanced, reducing confusion caused by cross-processing of different business scenarios. By accurately matching processing personnel and automatically distributing leads, resource allocation is optimized, avoiding work shirking caused by unclear division of responsibilities. By generating a unique identifier for each lead and binding it to a node in the entire process, process control is strengthened, preventing omissions in lead processing. By collecting processing information and updating status in real time, the completeness of tracking and tracing is improved, avoiding difficulties in defining responsibility due to missing process records. This ensures both standardized and transparent lead processing throughout the entire process and significantly improves the efficiency and control of audit work.

[0027] In one embodiment of the present invention, step S5 includes: The S51 system features a three-tiered dynamic visualization interface, including a main hall dashboard, a provincial dashboard, and a personnel dashboard, which displays the total number of leads, risk distribution, processing progress, overdue status, and handling results in real time. S52 offers multi-condition filtering and drill-down analysis based on audit type, risk level, time dimension, local police department, and modeling unit, and supports chart switching and data drill-down viewing; S53 summarizes the results of completed clues, extracts effective identification rules, typical anomaly patterns and high-frequency problem characteristics, and enriches the basis for audit supervision; S54 will simultaneously update the newly added effective rules and features to the audit knowledge graph, and simultaneously optimize the parameter configuration of the big data analysis engine and clue mining model; S55 adjusts the model's identification threshold, feature weights, and classification rules to continuously improve the accuracy and effectiveness of clue discovery, forming a data-driven, continuously optimized closed-loop system for audit clue collection.

[0028] The working principle and effects of the above technical solution are as follows: By building a three-level dynamic visualization page, the full-dimensional operational status of clues is presented intuitively, improving the real-time nature of audit supervision and reducing the time-consuming process of manual statistical summarization. Multi-condition filtering and penetrating analysis enhance the flexibility of data retrieval, reducing unnecessary browsing and information search costs. By summarizing processing results and extracting rule features, the basis for supervision is enriched, avoiding a lack of actual data support for audit judgments. By synchronously updating the knowledge graph and model parameters, the system's analytical capabilities are continuously iterated, avoiding a decrease in recognition accuracy due to model lag. By dynamically adjusting model thresholds and weights, the stability of clue mining is improved. This not only makes the entire audit process visible and controllable but also forms a self-optimizing closed-loop operation mode, comprehensively improving the intelligence level of audit supervision and risk prevention capabilities.

[0029] In one embodiment of the present invention, S54 includes: The newly added effective rules will be integrated with the audit knowledge graph to expand the coverage of the graph's supervisory nodes; Update the internal connections of the audit knowledge graph to strengthen the correspondence between rules and risk scenarios; input the extracted typical anomaly patterns into the big data analysis engine and adjust the engine's calculation dimensions and detection paths; Input high-frequency problem features into the clue mining model to enrich the model's feature input sources and matching dimensions; The audit knowledge graph and model parameters are updated synchronously to maintain the consistency and coherence of the overall system operation logic.

[0030] The working principle and effects of the above technical solution are as follows: By integrating newly added effective rules with the audit knowledge graph, the coverage of supervisory nodes is expanded, the comprehensiveness of risk identification is improved, and the possibility of blind spots in supervisory scenarios is reduced. By updating the internal association paths of the graph, the correspondence between rules and risk scenarios is strengthened, the tightness of data association is enhanced, and the dispersion of analysis logic is reduced. By inputting typical anomaly patterns into the analysis engine, the calculation and detection paths are adjusted, improving the accuracy of anomaly capture and avoiding the omission of new types of violations. By injecting high-frequency features into the mining model, the input and matching dimensions are enriched, enhancing the model's ability to adapt to new risks. By synchronously updating the graph and model parameters, the system logic remains consistent and coherent, avoiding conflicts in analysis results. This allows the supervisory system to continuously adapt to actual business changes while maintaining a stable and efficient operating state for the entire audit system.

[0031] In one embodiment of the present invention, S55 includes: The recognition threshold of the clue mining model is segmented and corrected to generate a threshold configuration result that adapts to the latest data features; The weights of various features within the model are redistributed to generate a balanced feature weight configuration. The model classification rules are supplemented hierarchically to generate a classification rule system that covers more risk scenarios; the updated threshold weights and rules are used to perform model calculations to generate optimized clue mining output results. The optimized model output is compared and verified with historical data to form a closed-loop system for continuously iterating audit clue collection.

[0032] The working principle and effects of the above technical solution are as follows: By segmenting and correcting the recognition threshold of the clue mining model, a configuration adapted to the latest data features is generated, improving the fit of clue recognition and reducing false positives and false negatives caused by fixed thresholds. By reallocating the feature weight ratios within the model, a balanced configuration result is generated, enhancing the rationality of feature utilization and reducing the excessive influence of a single feature on the output. By hierarchically supplementing the model classification rules, the coverage of risk scenarios is expanded, preventing new types of violations from failing to be effectively identified. By using updated parameters for calculation, the accuracy of clue mining output is improved. By comparing and verifying with historical data, a closed-loop iteration is formed, ensuring the system continuously adapts to business changes. This allows the model to maintain a consistently efficient and stable operating state, continuously improving the accuracy and effectiveness of clue detection, and ensuring the long-term reliable operation of audit data collection.

[0033] One embodiment of the present invention provides a big data-based system for discovering audit clues, the system comprising: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are made to implement the method described in any one of the above.

[0034] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An audit clue collection method based on a big data-driven audit clue discovery system, characterized in that: The method includes: S1. Construct an audit data integration platform, connect to the multi-source heterogeneous electronic data sources of the audited entity, and generate a standardized audit data resource pool; through data cleaning and preprocessing technology, automatically correct and filter redundant and erroneous data in the resource pool to generate a high-quality audit basic dataset; S2. Based on the big data analysis engine, a full scan and feature extraction are performed on the audit basic dataset. An abnormal pattern, correlation and potential risk point in the data are identified through a distributed computing framework to generate an initial audit clue candidate set. At the same time, a clue labeling system is established to classify and label the candidate set in multiple dimensions. S3. Utilize machine learning algorithms to intelligently filter and prioritize the initial audit lead candidate set, combine it with the audit business rule base to verify the compliance of the leads, and generate a high-value audit lead set after eliminating invalid leads; use natural language processing technology to perform semantic parsing and structured transformation on the leads in unstructured data; S4. Push the collection of high-value audit leads to the audit business system to trigger an automated lead distribution process. Allocate leads to different business modules according to the needs of the audit project. Simultaneously establish a lead lifecycle tracking mechanism to record the lead processing status and feedback results. S5. The dynamic visualization dashboard displays the distribution and processing progress of audit leads in real time, automatically updates the audit knowledge graph based on the lead processing results, optimizes the parameters of subsequent lead mining models, and forms a data-driven closed-loop system for audit lead collection.

2. The audit clue collection method for the big data-based audit clue discovery system according to claim 1, characterized in that, S1 includes: S11. Establish a dedicated data integration platform for public security audits, connect multiple heterogeneous electronic data sources of the audited entities, and complete the real-time connection and batch import of all data. S12. Perform unified format conversion on the accessed raw data, unify field naming, data type and storage structure, complete cross-system, cross-business and cross-department data field mapping and data fusion, and form a standardized audit data resource pool; S13. Start the automated data cleaning engine, identify problematic data in the resource pool, and perform operations such as merging duplicate data, filling blank fields, and filtering abnormal data. S14. Perform integrity, consistency and accuracy checks on the cleaned data, repair data conflicts, supplement key business information, and output a high-quality audit basic dataset covering all scenarios.

3. The audit clue collection method for the big data-based audit clue discovery system according to claim 1, characterized in that, S2 includes: S21. Activate the distributed big data analysis engine to perform a full data scan on the high-quality audit basic dataset and extract core data features; S22. Utilize a distributed computing framework to conduct cross-table association analysis, trend comparison analysis, and abnormal fluctuation analysis to identify abnormal patterns, hidden relationships, and high-risk behavior nodes that deviate from normal business logic. S23. Summarize all identified suspicious data fragments, risky behaviors and violations to form a complete initial audit clue candidate set covering the entire business scope; S24. Construct a multi-level clue tagging system and standardize tag dimensions and labeling rules; S25. For each clue in the initial audit clue candidate set, complete the label assignment and data binding to form a structured candidate clue set.

4. The audit clue collection method for the big data-based audit clue discovery system according to claim 3, characterized in that, S22 includes: S221. Utilize a distributed computing framework to conduct cross-table association analysis on core data features, generating multi-dimensional data association results; conduct trend comparison analysis on the multi-dimensional data association results, generating business trend change results; S222. Conduct abnormal fluctuation analysis on the results of business trend changes and generate data anomaly detection results; determine abnormal patterns that deviate from the normal business logic based on the data anomaly detection results; S223. Trace hidden relationships between data based on abnormal patterns; mark high-risk behavior nodes in the corresponding data based on hidden relationships.

5. The audit clue collection method for the big data-based audit clue discovery system according to claim 4, characterized in that, S221 includes: The scheduled distributed computing framework performs cross-table data association operations on core data features and generates intermediate results of cross-table data association. Perform multi-dimensional data integration on intermediate results of cross-table data association to generate multi-dimensional data association results; Retrieve historical business data and the correlation results of multi-dimensional data to perform a periodic comparison calculation and generate periodic data comparison results; Continuous cycle trend extrapolation is performed based on the comparison results of data from the same period to generate results on changes in business trends.

6. The audit clue collection method for the big data-based audit clue discovery system according to claim 1, characterized in that, The S3 includes: S31. Load the trained machine learning classification and ranking model, perform feature quantification calculation on the initial audit clue candidate set, and automatically complete the clue risk scoring and importance assessment. S32. Based on the scoring results, complete the stratification and priority sorting of clues, distinguish between key clues for verification, general clues of concern and clues to be verified, and plan the order of clue processing. S33. Retrieve the public security audit business rule database, compare the content of each clue with the degree of matching of compliance clauses, audit standards, and supervision requirements, mark invalid clues that do not comply with the rules and remove them; S34. Integrate and supplement the information and details of the retained valid clues, improve the key content involved in the clues, and form a set of high-value audit clues. S35. Use natural language processing technology to perform semantic analysis on unstructured data, extract key information and transform it into structured fields to supplement and improve the clue content.

7. The audit clue collection method for the big data-based audit clue discovery system according to claim 6, characterized in that, The S35 includes: Unstructured data is split into text and paragraphs to generate text units to be parsed; Semantic feature extraction is performed on text units to extract information such as time, subject, amount, cause, behavior, and description from the text; The extracted information is categorized and formatted to generate standardized structured fields; these structured fields are then linked with the corresponding audit leads to supplement and improve the complete content of the leads. Once the supplemented clues are completed, they will be included in the high-value audit clue set, thus completing the transformation of unstructured data into usable clues.

8. The audit clue collection method for the big data-based audit clue discovery system according to claim 1, characterized in that, The S4 includes: S41. High-value audit leads are pushed to the public security audit supervision and management business system via system interface and enter the lead processing process. S42. Read the audit project configuration, personnel responsibility list and supervision point setting information, and classify the clues according to the business type; S43. Accurately match clues with processing personnel according to the jurisdiction, police department, and responsible position, generate a standardized clue allocation list, and start automatic distribution; S44. Generate a unique identifier for each clue, bind it to all nodes in the entire process from clue generation, allocation, review, processing to feedback, and establish an unalterable tracking link; S45. Collect the progress of clue processing, verification results, review opinions and rectification status in real time, write them to the system log and update the clue status dynamically to form a traceable record of the whole process.

9. The audit clue collection method for the big data-based audit clue discovery system according to claim 1, characterized in that, The S5 includes: The S51 system features a three-tiered dynamic visualization interface, including a main hall dashboard, a provincial dashboard, and a personnel dashboard, which displays the total number of leads, risk distribution, processing progress, overdue status, and handling results in real time. S52 offers multi-condition filtering and drill-down analysis based on audit type, risk level, time dimension, local police department, and modeling unit, and supports chart switching and data drill-down viewing; S53 summarizes the results of completed clues, extracts effective identification rules, typical anomaly patterns and high-frequency problem characteristics, and enriches the basis for audit supervision; S54 will simultaneously update the newly added effective rules and features to the audit knowledge graph, and simultaneously optimize the parameter configuration of the big data analysis engine and clue mining model; S55 adjusts the model's identification threshold, feature weights, and classification rules to continuously improve the accuracy and effectiveness of clue discovery, forming a data-driven, continuously optimized closed-loop system for audit clue collection.

10. A big data-based system for discovering audit clues, characterized in that: The system includes: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 9.