Work fund compliance auditing method and system based on large language model

By integrating multi-source data through a large language model and constructing a dynamic compliance rule base, the problems of data integration and regulatory matching in the audit of trade union funds have been solved, enabling intelligent audit analysis and efficient compliance determination.

CN121685182APending Publication Date: 2026-03-17GONGFU (BEIJING) TECH DEV CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511999889.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Auditing trade union funds faces challenges such as difficulties in integrating and verifying multi-source heterogeneous data, complex matching of regulations and business scenarios, and low efficiency and inaccuracy due to reliance on manual labor.

Method used

By employing a large language model-based approach, through data cleaning and preprocessing, in-depth audit analysis, construction of a dynamic compliance rule base, and a learning mechanism, we achieve automatic integration and intelligent auditing of multi-source data.

Benefits of technology

It enables intelligent analysis of the entire process of auditing trade union funds, automatically identifies abnormal behavior, generates accurate compliance judgments and differentiated improvement suggestions, thereby improving audit efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685182A_ABST
    Figure CN121685182A_ABST
Patent Text Reader

Abstract

The invention provides a work fund compliance auditing method and system based on a big language model, and relates to the technical field of compliance auditing, the method comprises the following steps: extracting an unstructured text, and using the big language model to carry out data cleaning, privacy desensitization and standardization processing; based on a large language model and a retrieval enhancement generation technology, constructing a work fund compliance knowledge graph, performing deep auditing analysis on standardized data, and identifying income and expenditure anomalies, asset vulnerabilities and compliance risks; building a compliance rule base by using a large language model, and performing automatic compliance judgment; generating a differentiated improvement suggestion, a rectification path and an audit report first draft, and continuously optimizing a model and a rule; the system comprises a multi-source data acquisition and preprocessing module, a deep audit analysis module, a compliance analysis module, a decision support and report generation module, and a learning and optimization module. According to the invention, through automation and intelligentization of the work fund auditing process, the auditing efficiency, the auditing accuracy and the self-adaptive capability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of compliance audit, in particular to a trade union fund compliance audit method and system based on a large language model. BACKGROUND

[0002] Trade union fund compliance audit is an important link to ensure the healthy operation of trade union organizations, and its core is to verify the legality and reasonableness of fund collection and expenditure, asset management and special fund use. However, in the current practice, trade union fund audit is facing extremely complex and difficult practical challenges, and its technical difficulty has far exceeded the effective response range of traditional manual or simple informationization means.

[0003] The current audit work mainly faces two big problems. One is the difficulty of integration and verification of multi-source heterogeneous data. Trade union fund data sources are extremely scattered, and core data exist in trade union's own financial software system, while verification data is scattered in bank flow, payment account book, paper expenditure vouchers, contract texts and a large number of bill scans. These data have great differences in format, granularity, timeliness and credibility. For example, the financial system records a "training fee" expenditure, and the corresponding bank flow can verify the transfer, but whether the expenditure is really used for compliant training activities needs to verify the non-structured files such as training notice, contract, invoice and participant list. At present, auditors have to spend a lot of time in manual cross comparison between multiple systems and paper archives, which is low in efficiency and easy to cause omission due to fatigue and negligence. More difficult is that when different sources of information are inconsistent, such as the inconsistency between the amount of bill and the financial record, the expenditure time logic contradiction, there is a lack of effective automatic tools for semantic consistency verification and contradiction resolution, which seriously depends on the personal experience judgment of auditors, resulting in difficulty in guaranteeing the objectivity and comprehensiveness of audit results. Another key problem is the dynamic matching of complex regulations and specific business scenarios. Trade union fund management is subject to multiple regulations such as "Trade Union Law" and "Trade Union Audit Regulations", and the interpretation of clauses often has flexibility. For example, the identification of "excessive expenditure", "false expenditure" or "special fund misappropriation" not only needs accurate numerical comparison, but also needs to understand the semantics of business background, such as a seemingly ordinary meal reimbursement, which needs to be combined with multiple dimensions of information such as reimbursement reason, participants, per capita standard, whether pre-approval, etc. for comprehensive reasoning. Traditional audit software based on fixed rules cannot handle such complex semantic association and context reasoning, and pure reliance on manual judgment also has the problem of non-uniform standards and different scale grasping. Especially in the face of some "invisible" irregularities that deliberately evade supervision, the identification difficulty is extremely great. Therefore, there is an urgent need for a new trade union fund compliance audit method and system based on a large language model. SUMMARY

[0004] The purpose of the present application is to provide a trade union fund compliance audit method and system based on a large language model, to solve the problems of scattered data sources, difficult verification, complex regulation and business scenario matching, and low efficiency and insufficient accuracy of relying on manual work in the prior art, and the specific technical solutions are as follows: The present application provides a trade union fund compliance audit method based on a large language model, comprising: S01, obtaining full-amount data from multi-source data, cleaning and preprocessing the data to obtain high-quality standardized audit data; S02, based on a large language model and a retrieval enhancement generation technology, performing deep audit analysis on the standardized audit data, identifying fund income and expenditure abnormalities, asset management vulnerabilities and potential compliance risks, and obtaining audit analysis results; S03, based on laws and regulations and internal management systems of trade unions, utilizing the semantic reasoning capability of the large language model to construct a dynamic compliance rule library, and performing fund use compliance determination to obtain compliance analysis results; S04, fusing the audit analysis results and the compliance analysis results, and generating differentiated improvement suggestions, rectification paths and audit report drafts that fit the actual scenarios of trade unions through the large language model, to provide customized decision support for auditors; S05, establishing a learning mechanism of user feedback and new data iteration, dynamically optimizing the semantic analysis model, the compliance rule library and the risk identification standard through the continuous learning capability of the large language model, and improving the audit accuracy.

[0005] Further, the multi-source data includes trade union financial systems, bank flow, payment accounts, expenditure vouchers, contract texts and bill scans; the data cleaning and preprocessing includes extracting text information in unstructured data using OCR technology, and performing entity recognition and semantic cleaning using a large language model to realize format unification and privacy desensitization processing of structured and unstructured data.

[0006] Further, the data cleaning and preprocessing further includes: Synchronizing data from multi-source data through API interface and database query; Applying semantic cleaning rules based on a large language model for entity recognition, semantic deduplication, format standardization and missing value filling; Based on the semantic consistency rules of the large language model, checking the logical relationship, numerical association, time sequence and format consistency of data in different data sources, and automatically correcting inconsistent data; Using a large language model weighted semantic fusion algorithm and a multi-source voting mechanism to fuse the semantic features and classification attributes of the audit data to generate standardized audit data.

[0007] Further, the deep audit analysis includes: Standardized audit data is converted into a semantic feature matrix using a large language model; Retrieval enhancement generation technology is used to retrieve relevant laws, regulations and historical cases from the knowledge graph of trade union funding compliance, thereby enhancing the reasoning ability of the large language model; Construct an anomaly detection module based on a large language model, calculate the reconstruction error or anomaly score of semantic features, and identify abnormal behavior in financial income and expenditure; The temporal semantic analysis method of large language model is used to predict the trend of fund revenue payment progress and expenditure budget execution, and to classify the rationality based on the deviation ratio; By combining indicators such as revenue integrity, expenditure compliance, consistency between accounts and physical assets, and the rate of dedicated use of special funds, a comprehensive risk score is calculated through a risk scoring function to identify high-risk audit areas.

[0008] Furthermore, the anomaly detection module uses a sequence autoencoder or an anomaly detector based on an attention mechanism to train the model using historical normal data, and uses the reconstruction error exceeding a preset threshold as the anomaly judgment criterion; the reasonableness classification includes a deviation ratio of less than or equal to 10% as reasonable, a deviation ratio of greater than 10% and less than or equal to 20% as the interest area, and a deviation ratio of greater than 20% as an anomaly.

[0009] Furthermore, building a dynamic compliance rule base includes: Collect legal and regulatory documents and internal management system texts, and use a large language model for word segmentation, entity recognition and relation extraction to build a legal knowledge base; We use a large language model to extract compliance rules from regulatory texts and represent them in the form of conditional statements. We assign dynamic weights to each rule based on authority and timeliness. Standardized audit data is converted into semantic feature vectors, compliance scores are calculated based on semantic matching degree, and compliance status is determined based on thresholds; The generated differentiated improvement suggestions include: Construct a multi-dimensional decision feature matrix based on risk scores and violation types; Generate improvement suggestion templates for different risk levels using a large language model; Utilize search enhancement generation techniques to retrieve solutions for similar scenarios from historical audit cases; Generate phased rectification paths and timelines by leveraging the logical reasoning capabilities of large language models; Automatically generate a draft of a structured audit report, including a summary, findings, risk analysis, improvement suggestions, and rectification requirements.

[0010] Furthermore, the learning mechanism includes: Collect feedback and evaluations from auditors regarding the audit results and recommendations; Leveraging new and feedback data, the semantic parsing model, compliance rule base, and risk identification standards are dynamically optimized through incremental learning algorithms of large language models. The system can self-iterate by adjusting compliance thresholds and weight parameters based on performance monitoring metrics.

[0011] This invention also provides a union fund compliance audit system based on a large language model, used to implement the method described above. The system includes: The multi-source data acquisition and preprocessing module is used to automatically acquire full data from multiple sources, and clean, preprocess and standardize the data to generate high-quality standardized audit data. The in-depth audit analysis module is used to perform in-depth analysis of standardized audit data based on large language models and retrieval enhancement generation technology, identify abnormal financial income and expenditure, asset control loopholes and potential compliance risks, and output audit analysis results. The compliance analysis module is used to construct a dynamic compliance rule base based on laws, regulations, and internal management systems of the trade union, and to determine the compliance of fund usage by utilizing the semantic reasoning capabilities of a large language model, and output compliance analysis results. The decision support and report generation module integrates audit analysis results with compliance analysis results, and generates differentiated improvement suggestions, rectification paths, and initial drafts of audit reports through a large language model, providing auditors with customized decision support. The learning and optimization module is used to establish a learning mechanism based on user feedback and new data iteration. Through the continuous learning capability of the large language model, it dynamically optimizes the semantic parsing model, compliance rule base, and risk identification standards to improve audit accuracy.

[0012] Furthermore, the deep audit analysis module includes: Semantic feature representation unit, used to convert standardized audit data into a semantic feature matrix through a large language model; The retrieval enhancement feature engineering unit is used to retrieve relevant rules and cases from the knowledge graph of trade union funding compliance, enhance the contextual understanding ability of the large language model, and extract key semantic features; Anomaly detection unit is used to build anomaly detection model based on a large language model, calculate the reconstruction error or anomaly score of semantic features, and identify abnormal behavior; The temporal semantic analysis unit is used to predict income and expenditure trends and classify rationality using the temporal semantic analysis method of large language models; The risk assessment unit is used to calculate a comprehensive risk score by combining multi-dimensional indicators through a risk scoring function. The compliance analysis module includes: The legal knowledge base construction unit is used to collect legal texts and preprocess them through a large language model to form structured rule data; The compliance rule extraction unit is used to extract compliance rules using a large language model and assign dynamic weights. Audit data characterization unit, used to convert standardized audit data into semantic feature vectors; The compliance scoring unit is used to calculate a compliance score based on semantic matching degree and set a threshold for compliance determination; The dynamic optimization unit is used to continuously learn new regulations and user feedback through a large language model and dynamically update the compliance rule base. The decision support and report generation module includes: A multi-dimensional feature fusion unit is used to construct a decision feature matrix that integrates risk scores, compliance scores, and business scenario features; The differentiated suggestion generation unit is used to generate improvement suggestion templates for different risk levels based on a large language model; The retrieval enhancement generation unit is used to retrieve solutions for similar scenarios from historical audit cases; The rectification path planning unit is used to generate phased rectification paths and time plans through a large language model; The audit report generation unit is used to automatically generate draft structured audit reports. The learning and optimization module includes: The feedback collection unit is used to collect user feedback and evaluations; The model optimization unit is used to dynamically optimize the semantic parsing model and compliance rule base using new data and feedback data. The standard update unit is used to adjust risk identification standards and compliance thresholds; The performance monitoring unit is used to monitor system performance indicators and trigger optimization cycles.

[0013] The present invention also relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0014] The present invention also relates to an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described thereon.

[0015] The beneficial effects of this invention are as follows: The union fund compliance audit method and system based on a large language model, which relates to this invention, achieves a leap from traditional manual verification to intelligent analysis of the entire process by introducing a large language model and retrieval enhancement generation technology. The system can automatically integrate and process heterogeneous data from multiple sources such as financial systems, bank statements, and contract texts, overcoming the problem of data silos. Utilizing the deep semantic understanding and reasoning capabilities of the large language model, it can not only accurately identify hidden violations such as fictitious expenditures and embezzlement of funds, but also make real-time and accurate compliance judgments based on a dynamically constructed compliance rule base. Furthermore, the system can automatically generate differentiated rectification suggestions and structured audit reports that are practical and highly operable, greatly improving audit efficiency and decision support. At the same time, the system has continuous learning and self-optimization capabilities, and can dynamically adapt to changes in regulations and business needs, thereby significantly improving the coverage, accuracy, automation level, and intelligence level of union fund audits.

[0016] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram illustrating the steps of a method for auditing the compliance of trade union funds based on a large language model, provided by the present invention. Figure 2 This is a schematic diagram of the structure of a trade union fund compliance audit system based on a large language model provided by the present invention. Detailed Implementation

[0018] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0019] Please see Figure 1 The embodiments of the present invention provide a method for auditing the compliance of trade union funds based on a large language model, including: S01: Obtain full data related to union dues income and expenditure, assets and liabilities, and the use of special funds from multiple sources, and clean and preprocess the data to obtain high-quality standardized audit data. Specifically, the multi-source data includes union financial systems, bank statements, payment ledgers, expenditure vouchers, contract texts, and scanned copies of invoices. Data related to financial auditing is automatically collected from multiple data sources within the union, including but not limited to financial system records, bank statement information, payment ledgers, expenditure vouchers, contract texts, and scanned copies of invoices. OCR technology is used to extract textual information from unstructured data, and a large language model is employed for entity recognition and semantic cleaning. This ensures format consistency and privacy anonymization of both structured and unstructured data, guaranteeing the quality, consistency, and security of the input data and providing a reliable foundation for subsequent audit analysis.

[0020] S02, based on large language models and retrieval enhancement generation technology, conducts in-depth audit analysis on standardized audit data to identify abnormal financial income and expenditure, asset control loopholes and potential compliance risks, and obtains audit analysis results; Specifically, a large language model is used to analyze relevant laws and regulations, historical audit cases, and internal management systems related to trade union funds. This constructs a multi-dimensional knowledge graph of trade union fund compliance, encompassing revenue compliance, expenditure standardization, asset control, and the use of special funds. Audit data is then linked to the knowledge graph through entity connections and relationship mapping. Standardized audit data is converted into semantic vectors via an embedded model. Based on a vector database, accurate semantic matching between revenue and expenditure data and the knowledge graph is achieved, thus opening up a channel for the correlation analysis of structured and unstructured data. An anomaly semantic recognition module driven by the large language model is constructed. By analyzing the semantic consistency between expenditure vouchers and actual business transactions, and the logical coherence of approval processes, it identifies hidden violations such as fictitious expenditures, misappropriation of funds, and embezzlement of public funds. The temporal semantic analysis capabilities of the large language model are used to predict trends in the progress of trade union fund revenue disbursement and expenditure budget execution, and to determine the reasonableness of revenue and expenditure fluctuations by comparing them with historical data from the same period. By combining core indicators such as revenue integrity, expenditure compliance, asset consistency, and the rate of dedicated use of special funds, a risk scoring function is defined. Through a large language model, multi-dimensional analysis results are integrated to calculate a comprehensive risk score, thereby identifying high-risk audit areas and achieving in-depth audit analysis.

[0021] S03, based on laws and regulations such as the Trade Union Law, the Trade Union Audit Regulations, and the Trade Union Funds Revenue and Expenditure Management Measures, as well as the internal management system of trade unions, uses the semantic reasoning capabilities of a large language model to construct a dynamic compliance rule base, conducts compliance judgments on the use of funds, and obtains compliance analysis results; Specifically, by combining the latest laws, regulations, and internal management norms of the trade union, the use and management of funds are assessed to determine whether they comply with the prescribed standards. Through the semantic reasoning capabilities of a large language model, a compliance rule base is dynamically constructed and updated, enabling the understanding and interpretation of complex regulatory texts. These rules are then applied to actual fund auditing activities, automatically determining compliance and ensuring the accuracy and timeliness of audit results.

[0022] S04 integrates audit analysis results with compliance analysis results, and generates differentiated improvement suggestions, rectification paths and initial draft audit reports tailored to the actual scenarios of trade unions through a large language model, providing customized decision support for auditors; Specifically, based on the results of in-depth audit and compliance analysis, a large language model is used to generate differentiated improvement suggestions, rectification paths, and draft audit reports tailored to the specific circumstances of the trade union. These suggestions may include optimizing the use of funds, adjusting processes, and implementing risk mitigation measures. Furthermore, feedback from auditors is taken into account to continuously optimize the relevance and practicality of the recommendations, providing comprehensive support for audit decision-making.

[0023] S05, Establish a learning mechanism based on user feedback and new data iteration, and dynamically optimize the semantic parsing model, compliance rule base and risk identification standards through the continuous learning capability of the large language model to improve audit accuracy; Specifically, through self-learning and optimization mechanisms, the semantic parsing model, compliance rule base, and risk identification standards can be continuously adjusted based on new data and user feedback, ensuring the continued effectiveness and adaptability of audit methods. This learning mechanism leverages the continuous learning capabilities of a large language model to achieve dynamic optimization, thereby improving the accuracy and efficiency of audits.

[0024] In this optional embodiment, full-volume data related to union dues income and expenditure, assets and liabilities, and the use of special funds are obtained from multiple sources of data, including union financial systems, bank statements, payment ledgers, expenditure vouchers, contract texts, and scanned copies of invoices. Unstructured text information is extracted using OCR technology, and entity recognition and semantic cleaning are performed using a large language model. This achieves unified structured and unstructured data formats and privacy-desensitizing processing, resulting in high-quality, standardized audit data, including: Based on the trade union multi-source data system, a set of data sources is defined, and financial audit data is synchronized through API interfaces and database queries to ensure comprehensive acquisition and preliminary preparation of audit data. Data cleaning rules driven by a large language model are used to perform entity recognition, semantic deduplication, and format standardization on synchronized financial audit data to obtain cleaned audit data. Based on the semantic consistency check rules of the large language model, the consistency of logical relationships, numerical associations, time order and semantic format of the audit data of funds from different data sources is checked to ensure that the audit data remains accurate after multi-source fusion. Based on the cleaned and consistent audit data, a semantic fusion algorithm with weighted large language model is used to process the semantic features of the audit data, and a multi-source voting mechanism is used to process the classification attributes of the audit data to obtain high-quality standardized audit data.

[0025] Specifically, multi-source data intelligent processing and standardization algorithms include: 1) Multi-source data identification and synchronization: Define the data source set D={D1, D2, ..., D...} for union dues audit. n},in Representing the i-th data source, including financial systems, bank statements, payment ledgers, expenditure vouchers, contract texts, scanned copies of invoices, etc.; through API interface calls and database queries, relevant data on income and expenditure, assets and liabilities, and special fund usage are synchronized from each data source in real time or periodically; for unstructured data such as scanned copies of invoices, OCR technology is used to extract text information to form processable semi-structured data.

[0026] 2) Large language model-driven data cleaning: For each data source Define semantic cleaning rules based on large language models. This includes entity recognition, semantic deduplication, format standardization, and missing value imputation; The large language model is used to perform entity recognition on text data, identifying key entities related to financial income and expenditure, such as payees, payers, amounts, times, and purposes. Through the semantic understanding capabilities of the large language model, records with semantic repetition are detected and removed, format errors are corrected, and missing key information is filled in through semantic reasoning. Privacy desensitization processing is performed, automatically masking or encrypting sensitive personal information (such as ID card numbers and bank card numbers) identified by the large language model.

[0027] 3) Semantic consistency check: Design semantic consistency rules based on a large language model. These rules are a set of predefined semantic constraints to ensure that the same or related data from different data sources maintain semantic consistency in dimensions such as logical relationships, numerical associations, and time order. Entity consistency rule: Key entity information (such as the parties to the transaction, amount, and time) for the same funding matter should be semantically consistent across different data sources; Logical consistency rule: The logical relationship between revenue and expenditure data should conform to business rules, such as expenditure amount not exceeding the budget, and special funds being used for designated purposes. Time consistency rule: Time-related data should conform to the logical order of time, such as the reimbursement time being later than the consumption time; Format consistency rule: The format of the same semantic field should be consistent across all data sources, such as monetary units and date formats; For semantically inconsistent data, automatic correction is performed using the semantic reasoning capabilities of large language models, or processing is carried out according to predefined priority rules.

[0028] 4) Semantic Data Fusion: For the semantic features of audit data, a semantic fusion algorithm weighted by a large language model is adopted. Where S represents the fused semantic feature vector; This represents the semantic weight of the i-th data source, reflecting the semantic importance of that data source. This represents the semantic confidence of the i-th data source, derived by the large language model based on data quality assessment; LLM embedding ( This indicates that the data from the i-th data source is processed through a large language model. Convert to a semantic vector; n represents the total number of data sources participating in the fusion.

[0029] For the classification attributes of audit data (such as expenditure type, compliance status, etc.), a multi-source voting mechanism enhanced by a large language model is used: From each data source Extracting semantically cleaned classification attribute values The large language model is used to verify whether vᵢ conforms to the predefined set of legal values ​​V; Calculate the semantic voting weight of each data source. Due to its basic weight and semantic confidence Joint decision: Perform semantically weighted majority voting. For each candidate value v∈V, calculate its total voting weight and select the candidate value with the highest semantic weight as the final result. Where II(·) is an indicator function, when The value is 1 if it is true, and 0 otherwise.

[0030] Through the above-mentioned intelligent processing and standardization algorithms for multi-source data, the format unification and semantic consistency of structured and unstructured data are achieved, providing a high-quality standardized audit data foundation for subsequent in-depth audit analysis based on a large language model.

[0031] In this optional embodiment, a deep audit analysis is performed on standardized audit data based on a large language model and retrieval enhancement generation technology to identify abnormal financial receipts and expenditures, asset control loopholes, and potential compliance risks. The audit analysis results include: Based on standardized audit data, a semantic feature matrix is ​​constructed, and semantic vectorization is performed through a large language model to obtain a feature representation of a unified semantic space. By utilizing retrieval-enhanced generation techniques, relevant rules and cases are retrieved from the knowledge graph of trade union funding compliance, thereby enhancing the reasoning ability of the large language model and extracting key semantic features. A large language model-driven anomaly detection module is constructed and trained using extracted key semantic features. The reconstruction error or anomaly score of each semantic feature is calculated, and the anomaly threshold is determined. Based on the progress of revenue and expenditure payments and the execution of expenditure budgets, the temporal semantic analysis method of large language model is used to predict revenue and expenditure trends and obtain the results of revenue and expenditure rationality classification. Based on indicators such as expenditure compliance and consistency between asset records and physical assets, determine whether there are any abnormalities in the use of funds; By combining multiple indicators such as revenue integrity, expenditure compliance, asset consistency, and the rate of dedicated use of special funds, a risk scoring function is defined. The comprehensive risk score is calculated by integrating the multi-dimensional analysis results through a large language model, thereby obtaining the identification results of potential audit risks.

[0032] In this optional embodiment, based on the progress of revenue and expenditure payments and the execution of expenditure budgets, a temporal semantic analysis method using a large language model is employed to predict revenue and expenditure trends, resulting in a revenue and expenditure rationality classification that includes: Based on historical revenue and expenditure data, obtain annual or monthly revenue payment progress and expenditure execution data, perform time-series semantic analysis, and obtain the baseline trend of revenue and expenditure. By using the baseline trend of revenue and expenditure, the progress deviation or execution deviation ratio at each time point is calculated to obtain the revenue and expenditure status, and the rationality is classified according to the deviation ratio and the baseline trend. Based on the obtained deviation ratio, a three-level classification standard is established, and the income and expenditure situation is divided into reasonable, attention area and abnormal category, to obtain the income and expenditure reasonableness classification results.

[0033] Specifically, the deep auditing algorithm for union funds based on large language models and retrieval enhancement includes: 1) Semantic Feature Representation: Standardized audit data (including structured data such as income and expenditure records, and unstructured data such as contract texts) are transformed into a semantic feature matrix using a large language model. , where m is the number of data points (such as transaction records, expenditure items), and d is the dimension of the semantic vector.

[0034] Semantic features are standardized to eliminate dimensional differences and ensure semantic consistency. The embedding layer of a large language model is used to convert text and data into vectors, using the formula: v=LLM. embed (date); where v is a semantic vector, LLM embed (date) represents the embedding function of the large language model.

[0035] 2) Retrieval Enhancement Feature Engineering: Apply retrieval enhancement generation technology to retrieve relevant laws and regulations, historical audit cases and internal systems from the knowledge graph of trade union funds compliance, thereby enhancing the contextual understanding ability of the large language model.

[0036] Knowledge Graph Construction: Using a large language model to parse texts such as the "Trade Union Law" and the "Trade Union Audit Regulations", a knowledge graph containing entities (such as income types and expenditure items) and relationships (such as compliance rules) is constructed.

[0037] Semantic retrieval: For each audit data point, relevant rules and cases in the knowledge graph are retrieved through vector similarity to form enhanced prompts, which are then input into a large language model to extract key semantic features.

[0038] Feature dimensionality reduction: Using autoencoders or principal component analysis (PCA) to reduce the dimensionality of high-dimensional semantic features, resulting in feature representations with high information content. (where k < d), providing a basis for anomaly detection.

[0039] 3) Anomaly Detection Model: Train an anomaly detection model based on a large language model, such as using a sequence autoencoder or an attention-based anomaly detector. Use the dimensionality-reduced semantic feature matrix X' as input.

[0040] Training data: The model is trained using historical normal data (without abnormal labels) to learn the distribution of normal funding patterns.

[0041] Calculate the reconstruction error: for each sample Obtained through model reconstruction Calculate the reconstruction error (e.g., mean square error): The abnormal threshold θ is determined based on the error distribution of normal samples (such as the 95th percentile).

[0042] Model evaluation: Accuracy and recall are calculated using a labeled test set to ensure model performance.

[0043] 4) Revenue and Expenditure Trend Analysis: Define revenue and expenditure rationality indicators, such as the ratio of revenue payment progress to planned progress and the ratio of expenditure execution rate to budget. Utilize the temporal semantic analysis capabilities of a large language model to predict revenue and expenditure trends and identify abnormal fluctuations. Data preparation: Collect monthly revenue payment progress and expenditure execution data for the most recent period t (default t=12 months): Income= [I1, I2,…, I t ], Expense = [E1, E2, …, E t ]; where I t E represents the actual income payment amount in month t. t This represents the actual expenditure amount in month t.

[0044] Calculate the baseline trend: Use a large language model to analyze historical data to obtain baseline values ​​or trend lines for revenue and expenditure. For example, the baseline value B for revenue payment progress. IIt can be calculated using moving average or trend fitting: Similarly, the benchmark value for expenditure execution .

[0045] Reasonableness classification: Calculate the deviation ratio D at each time point. t For example, regarding income: Establish a three-level classification standard: Reasonable category: Deviation ratio D t ≤10% indicates that income and expenditure are within the normal range.

[0046] Attention Zone: 10% < D t ≤20% requires attention but is not necessarily abnormal.

[0047] Exception class: D t >20%, there may be abnormal fluctuations.

[0048] The classification results are adjusted by semantic analysis of a large language model and in combination with context (such as seasonal factors).

[0049] 5) Analysis of expenditure compliance and asset consistency: Define key metrics: Expenditure compliance indicator C e The system uses a large language model to determine whether expenditures comply with regulations and calculates the compliance rate.

[0050] Asset book-physical consistency index C a Compare the book assets with the actual inventory data and calculate the consistency ratio.

[0051] Special funds used for their designated purpose rate C s The proportion of special funds used for designated purposes.

[0052] Anomaly detection: Utilize the logical reasoning capabilities of large language models to detect anomalies in metrics. For example, if C... e If < 95%, the expenditure is marked as non-compliant; if C a If the percentage is less than 90%, the assets are not properly labeled.

[0053] 6) Risk Assessment: Define a risk scoring function R(x) and calculate a comprehensive risk score based on multi-dimensional indicators. First, normalize each indicator: Then, a weighted risk score is calculated using the following formula: Where w1, w2, w3, w4 are weights, satisfying ∑w i =1, and the weights can be learned through a large language model or set by experts.

[0054] Set risk threshold θ R If R(x) > θR If so, it is marked as a high-risk audit area.

[0055] In this optional embodiment, based on laws and regulations such as the "Trade Union Law," the "Trade Union Audit Regulations," and the "Trade Union Funds Revenue and Expenditure Management Measures," as well as the internal management system of the trade union, a dynamic compliance rule base is constructed using the semantic reasoning capabilities of a large language model to determine the compliance of fund usage. The resulting compliance analysis includes: A legal knowledge base was constructed, collecting a set of texts of laws, regulations and internal management systems related to trade union funds. The text sets were preprocessed to obtain structured legal text data. Natural language processing techniques using large language models are used to extract keywords and semantic rules representing compliance rules from regulatory texts; The compliance rules are treated as a series of conditional statements. Each conditional statement defines the compliance standards for the use and management of funds, and each compliance rule is assigned a weight, which is dynamically adjusted based on the importance and timeliness of the rule. Standardized audit data is used as a set of feature vectors, and semantic features are performed based on a large language model to transform the audit data into semantic vectors, ensuring that each feature vector can comprehensively describe the use of funds. Based on the weights of compliance rules and the semantic matching degree of feature vectors, a compliance scoring model is constructed, and a compliance score is calculated for each feature vector to obtain the compliance score. Set a compliance score threshold. If the compliance score of a feature vector is less than the threshold, it is judged as non-compliant; otherwise, it is judged as compliant. Output a compliance score and decision result for each feature vector and generate a compliance report.

[0056] Specifically, dynamic compliance analysis algorithms based on large language models include: 1) Construction of a legal knowledge base: Collect a set of legal texts T = {T1, T2, …, T k} where T represents a legal or regulatory text. A large language model is used to preprocess the text, including word segmentation, entity recognition, and relation extraction, to form structured rule data, which is then stored in a vector database for easy semantic retrieval.

[0057] 2) Compliance Rule Extraction: Utilizing the semantic reasoning capabilities of a large language model, compliance rules are extracted from regulatory texts. For example, a rule might take the form: IF condition THEN Compliance, where conditions could include: "Expenditures exceeding 1000 yuan require a contract," or "Special funds may not be used for general expenditures," etc. A weight w is assigned to each rule. j The weights are dynamically calculated based on the authority of the rule source (e.g., laws take precedence over internal regulations) and the timeliness (e.g., the latest rules have higher weights).

[0058] 3) Audit data characterization: Standardized audit data is converted into semantic feature vectors using a large language model. , where h is the feature dimension. Features include expenditure type, amount, time, related entities, etc.

[0059] 4) Compliance scoring model: For each feature vector f, calculate the degree of matching with the compliance rules. Semantic similarity is used for calculation, such as cosine similarity. , where R j Let r be the vector representation of the j-th rule. j It is its semantic vector. Compliance score S c Weighted matching degree: , where n is the number of rules.

[0060] 5) Compliance determination: Set a threshold θ c If S c ≥θ c If the threshold is met, the audit is deemed compliant; otherwise, it is deemed non-compliant. The threshold can be adjusted according to the audit's rigor.

[0061] 6) Dynamic optimization: The compliance rule base continuously learns new regulations and user feedback through a large language model, and dynamically updates rules and weights.

[0062] Through the above steps, in-depth audit analysis and compliance analysis based on a large language model are achieved, providing a precise and adaptive solution for auditing trade union funds.

[0063] In this optional embodiment, a legal knowledge base is constructed, collecting a text set of laws, regulations, and internal management systems related to union funds. The text set is preprocessed to obtain structured legal text data, including: A knowledge base of laws and regulations concerning trade union funds was constructed, which collected a set of texts of laws, regulations and internal management systems related to the audit of trade union funds. The texts were then processed by removing redundant characters, unifying encoding and standardizing format to obtain cleaned legal texts. Using natural language processing techniques based on large language models, semantic segmentation, optimized segmentation of professional terms, and removal of general and domain-specific stop words are performed on the cleaned legal text to obtain high-quality legal text segmentation results. Based on the obtained word segmentation results of the regulatory text, a large language model is used to perform semantic role labeling and grammatical attribute analysis, extract key compliance phrases and deep semantic information, and obtain structured regulatory text data.

[0064] Specifically, a dynamic compliance analysis algorithm based on a large language model: 1) Construction of a legal knowledge base: Create a legal knowledge base on trade union funds, which includes a collection of texts of laws and regulations such as the "Trade Union Law", the "Trade Union Audit Regulations", and the "Measures for the Management of Trade Union Funds Income and Expenditure", as well as internal management systems of trade unions; The knowledge base is stored using a vector database, supporting semantic retrieval and dynamic updates.

[0065] 2) Text preprocessing driven by large language models: Deep semantic preprocessing is performed on the text in the legal knowledge base, including natural language processing steps such as semantic word segmentation, stop word filtering, and semantic role labeling. Remove redundant characters: Delete HTML tags, special symbols, illegal characters, and formatting marks; Unified encoding and format: Convert text to UTF-8 encoding in a unified manner to standardize the expression format of key information such as date and amount; Semantic word segmentation is performed using a large language model, and the segmentation accuracy of professional terms is optimized by combining a professional dictionary of trade union funds. Apply large language models to identify and filter general stop word lists and domain-specific stop words (such as redundant expressions in regulations like "Article X" and "according to regulations"). The semantic role labeling function of the large language model is used to identify semantic roles such as condition subject, behavior object, and restrictive conditions in legal texts.

[0066] 3) Semantic feature extraction: Keyword and semantic rule extraction: Utilizing the semantic understanding capabilities of large language models, keywords and deep semantic rules that can represent compliance rules are extracted from regulatory texts; Compliance phrase and clause association extraction: Extract compliance phrases and clause numbers from regulatory texts and establish logical relationships between clauses; Semantic vectorization: The extracted compliance rules are converted into semantic vectors through a large language model and stored in a vector database.

[0067] 4) Construction of a dynamic compliance rule base: The compliance rules are represented as a series of semantic conditional statements, each of which defines the compliance standards for the use and management of funds. Rules are expressed in the form of "semantic condition → compliant action", for example: Condition: The single expenditure exceeds 5,000 yuan and no contract is attached; Action: Marked as "Expenditure procedure non-compliant"; Each rule is assigned a dynamic weight, which can be dynamically adjusted based on the authority, timeliness, and historical audit results of the regulations.

[0068] 5) Semantic characterization of audit data: Standardized audit data is transformed into a set of semantic feature vectors using a large language model, where each vector xᵢ describes the complete semantic features of a fund usage item. Data semantic feature generation, which transforms audit data into semantic feature vectors, includes: Define a set of semantic features for union dues audit data, for example: Transaction ID, date of occurrence, amount, related party information, and transaction type; Completeness of the approval process, completeness of supporting documents, compliance with budget, and matching of specific purposes; By mapping each audit data point to a high-dimensional semantic space using a large language model, we can ensure that each feature vector can comprehensively describe how the funds were used.

[0069] 6) Compliance scoring model: Construct a compliance scoring model S:X→[0,1] based on a large language model. This model calculates a compliance score for each audit feature vector xᵢ. The formula for calculating the compliance score is defined as follows:

[0070] In the formula, Represents the dynamic weight of the j-th compliance rule; LLM match This represents the audit feature vector calculated using a large language model. With compliance rules r j The semantic matching degree; n represents the number of relevant compliance rules.

[0071] 7) Compliance-based intelligent decision-making: Set a compliance threshold τ, if the compliance score S for the use of funds is... If <τ, then the use of funds is considered non-compliant; For each use of funds, output its compliance score and decision result (compliant / non-compliant), and provide a specific explanation of the reasons for the violation; Detailed compliance analysis reports are generated based on large language models, providing targeted improvement suggestions and rectification basis for non-compliance items.

[0072] In this optional embodiment, the audit analysis results and compliance analysis results are integrated, and a large language model is used to generate differentiated improvement suggestions, rectification paths, and initial draft audit reports tailored to the actual scenarios of trade unions, providing auditors with customized decision support, including: Based on the risk scores and anomaly identification information in the audit analysis results, and combined with the violations and compliance status in the compliance analysis results, a multi-dimensional decision feature matrix is ​​constructed. Leveraging the semantic generation capabilities of large language models, a differentiated improvement suggestion template library is generated for scenarios combining different risk levels and violation types. By leveraging search-enhanced generation techniques, solutions for similar scenarios can be retrieved from historical audit cases and best practices, thereby enhancing the practicality and relevance of recommendations. Based on the specific business scenarios and resource constraints of the trade union, the logical reasoning capabilities of the large language model are used to generate feasible rectification paths and time plans. By integrating all analysis results and recommendations, and leveraging the structured report generation capabilities of the large language model, a complete draft audit report is automatically generated.

[0073] Specifically, intelligent decision support algorithms based on large language models: 1) Multi-dimensional feature fusion: Construct the decision feature matrix D∈R m×k Where m is the number of matters to be decided, and k is the feature dimension, including: Risk scoring characteristics: A comprehensive risk score R(x) derived from audit analysis; Compliance characteristics: Compliance score derived from compliance analysis and violation type coding; Business scenario characteristics: contextual features such as union size, funding size, and historical audit issues.

[0074] 2) Generation of differentiated suggestions: Design a suggestion generation template based on a large language model: For combinations of high risk and serious violations: generate "immediate rectification" recommendations, emphasizing timeliness and seriousness; For a combination of medium-risk and general violations: generate "improvement within a specified period" suggestions and provide specific improvement paths; For low-risk, procedural issues: generate "optimization and improvement" suggestions, focusing on process optimization.

[0075] 3) Enhanced search generation: Establish a vector database of historical audit cases to store the types of problems and solutions from past audits; For the current auditing issue, the most relevant historical cases are retrieved using vector similarity. The retrieved cases are used as context input into a large language model to generate more practical suggestions.

[0076] 4) Rectification Path Planning: By leveraging the logical reasoning capabilities of large language models, we can analyze the dependencies and resource requirements of rectification measures. Generate a phased rectification path: immediate measures (within 1 week) → short-term rectification (within 1 month) → long-term optimization (within 3 months); Consider the practical constraints imposed by the union, such as budget limitations, staffing, and business processes.

[0077] 5) Audit reports are automatically generated: Design a structured template for an audit report: Summary → Problems Identified → Risk Analysis → Improvement Suggestions → Rectification Requirements; By leveraging the text generation capabilities of large language models, the analysis results are naturally transformed into professional auditing language. It supports personalized customization of report content to meet the reading needs of auditors at different levels.

[0078] 6) Decision support optimization: Establish a feedback mechanism for the effectiveness of recommendations, and collect auditors' evaluations of the practicality of the generated recommendations; By continuously optimizing the suggestion generation quality of the large language model through feedback data, we can achieve continuous improvement in decision support.

[0079] The aforementioned method for auditing trade union funds compliance based on a large language model, through the deep integration of the large language model and retrieval-enhanced generation technology, achieves a paradigm shift in trade union fund auditing from passive verification to proactive intelligent early warning. This method can automatically integrate and process multi-source heterogeneous financial data, accurately identify hidden violations that are difficult to detect using traditional methods by utilizing the semantic understanding capabilities of the large language model; it achieves real-time and accurate regulatory compliance determination by constructing a dynamic compliance rule base, and generates actionable rectification paths and audit reports based on multi-dimensional analysis results; its continuous learning mechanism enables the system to continuously optimize audit accuracy and adaptability, providing accurate, practical, and personalized decision support for trade union fund auditing, significantly improving the efficiency and quality of audit work.

[0080] Please see Figure 2 The embodiments of the present invention also provide a trade union fund compliance audit system based on a large language model, including a multi-source data acquisition and preprocessing module, a deep audit analysis module, a compliance analysis module, a decision support and report generation module, and a learning and optimization module; The multi-source data acquisition and preprocessing module is used to automatically acquire all data related to income and expenditure, assets and liabilities, and the use of special funds from multi-source data of trade unions, and to clean, preprocess and standardize the data to generate high-quality standardized audit data. The multi-source data acquisition and preprocessing module includes: a data source identification and synchronization unit, which is used to define the data source set for union fund audit (such as financial system, bank statements, payment ledgers, expenditure vouchers, contract texts, scanned copies of invoices, etc.), synchronize data in real time or periodically through API interface and database query, and use OCR technology to extract text information from unstructured data (such as scanned copies of invoices) to form processable semi-structured data; The large language model-driven cleaning unit is used to apply semantic cleaning rules based on the large language model (including entity recognition, semantic deduplication, format standardization, missing value imputation, etc.) to perform entity recognition (such as payee, payer, amount, time, etc.) on synchronized data, and to achieve privacy desensitization processing (such as automatically blocking sensitive information). The semantic consistency check unit is used to detect the consistency of logical relationships, numerical associations and temporal order of data in different data sources based on semantic consistency rules (such as entity consistency, logical consistency, temporal consistency and format consistency) of large language models, and automatically correct inconsistent data through semantic reasoning. The semantic data fusion unit is used to fuse the semantic features and classification attributes of audit data using a large language model weighted semantic fusion algorithm and a multi-source voting mechanism to generate standardized audit data in a unified format, ensuring data quality and semantic consistency.

[0081] In this embodiment of the invention, the multi-source data acquisition and preprocessing module is used to automatically acquire, clean, and standardize multi-source data from trade unions. The system dynamically accesses various data sources through a data source identification and synchronization unit and uses OCR technology to process unstructured data. Subsequently, a large language model drives a cleaning unit to perform deep semantic cleaning on the data, ensuring data accuracy and privacy security. A semantic consistency check unit ensures the logical correctness of the multi-source data fusion through predefined rules. Finally, a semantic data fusion unit generates high-quality standardized data, providing a reliable foundation for subsequent audit analysis. This module overcomes the heterogeneity problem of traditional audit data and improves the automation level and accuracy of data processing.

[0082] The in-depth audit analysis module is used to perform in-depth analysis of standardized audit data based on large language models and retrieval enhancement generation technology, identify abnormal financial income and expenditure, asset control loopholes and potential compliance risks, and output audit analysis results. The in-depth audit analysis module includes: The semantic feature representation unit is used to convert standardized audit data (including structured and unstructured data) into a semantic feature matrix through a large language model, and to perform standardization processing to eliminate dimensional differences and ensure semantic consistency. The retrieval enhancement feature engineering unit is used to apply retrieval enhancement generation technology to retrieve relevant laws and regulations, historical cases and internal systems from the knowledge graph of trade union funding compliance, enhance the contextual understanding ability of the large language model, extract key semantic features, and obtain high information content feature representations through feature dimensionality reduction (such as autoencoders or PCA). Anomaly detection unit is used to build anomaly detection models based on large language models (such as sequence autoencoders), train the model using historical normal data, calculate the reconstruction error or anomaly score of semantic features, set anomaly threshold, and identify abnormal financial behavior (such as fictitious expenditures, misappropriation of funds, etc.). The temporal semantic analysis unit is used to predict revenue and expenditure trends, calculate the deviation ratio, and classify rationality (rational, concern area, abnormal) based on the progress of fund revenue payment and expenditure budget execution and the temporal semantic analysis method of large language model. The risk assessment unit combines multiple indicators such as revenue integrity, expenditure compliance, and asset consistency to define a risk scoring function. It then uses a large language model to integrate and analyze the results to calculate a comprehensive risk score and identify high-risk audit areas.

[0083] In this embodiment of the invention, the deep audit analysis module utilizes the semantic understanding and reasoning capabilities of a large language model to achieve intelligent auditing of trade union funds. The system first converts data into semantic features and enriches the context through retrieval enhancement technology to improve analytical accuracy. The anomaly detection unit identifies hidden violations using a machine learning model, the time-series semantic analysis unit dynamically monitors income and expenditure trends, and the risk assessment unit quantifies risk levels. This module achieves a leap from surface auditing to deep semantic analysis, improving the comprehensiveness and accuracy of the audit.

[0084] The compliance analysis module is used to construct a dynamic compliance rule base based on laws and regulations such as the Trade Union Law and the Trade Union Audit Regulations, and to make compliance judgments on the use of funds, and output compliance analysis results. The compliance analysis module includes: a legal knowledge base construction unit, which collects legal and regulatory documents and internal management system texts related to trade union funds, preprocesses them using a large language model (such as word segmentation, entity recognition, and relation extraction), forms structured rule data, and stores it in a vector database to support semantic retrieval; The compliance rule extraction unit is used to extract compliance rules (such as conditional statements) from regulatory texts using a large language model, and to assign dynamic weights (based on authority and timeliness) to each rule to build a dynamic compliance rule library; The audit data characterization unit is used to convert standardized audit data into semantic feature vectors through a large language model, comprehensively describing the use of funds (such as expenditure type, amount, time, etc.). The compliance scoring unit is used to calculate the matching score between audit data and compliance rules based on semantic matching degree (such as cosine similarity), combine weights to generate a compliance score, and set thresholds for compliance determination. The dynamic optimization unit is used to continuously learn new regulations and user feedback through a large language model, dynamically update the compliance rule base and weights, and ensure the accuracy and timeliness of judgments.

[0085] In this embodiment of the invention, the compliance analysis module transforms complex regulations into executable rules using a large language model, enabling automated compliance checks. The system constructs a regulatory knowledge base and dynamically extracts rules. Audit data is characterized and semantically matched with the rules to output compliance scores and decision results. This module replaces the traditional method of manually interpreting regulations, improving the efficiency and consistency of compliance audits.

[0086] The decision support and report generation module integrates audit analysis results with compliance analysis results, and generates differentiated improvement suggestions, rectification paths, and initial drafts of audit reports through a large language model, providing auditors with customized decision support. The decision support and report generation module includes a multi-dimensional feature fusion unit, which is used to construct a decision feature matrix and integrate multi-dimensional information such as risk scores, compliance scores, and business scenario features (such as union size and historical issues). The differentiated suggestion generation unit is used to generate improvement suggestion templates (such as immediate rectification, improvement within a time limit, and optimization) based on a large language model for different risk levels and violation types. The retrieval enhancement generation unit is used to retrieve solutions for similar scenarios from the historical audit case vector database, enhancing the practicality and relevance of the recommendations; The rectification path planning unit is used to analyze the dependencies and resource requirements of rectification measures through the logical reasoning capabilities of the large language model, and generate phased rectification paths and time plans. The audit report generation unit is used to automatically generate a complete draft audit report (including summary, findings, risk analysis, and improvement suggestions) through the structured report generation capabilities of the large language model.

[0087] In embodiments of this invention, the decision support and report generation module transforms analysis results into actionable decision support. The system integrates multi-dimensional features to generate personalized suggestions and rectification paths, and improves the quality of suggestions through retrieval enhancement technology. The report generation unit automatically outputs professional audit reports, significantly reducing the burden of manual drafting. This module achieves intelligent transformation of audit conclusions, improving the efficiency and quality of audit decision-making.

[0088] The learning and optimization module is used to establish a learning mechanism based on user feedback and new data iteration. Through the continuous learning capability of the large language model, it dynamically optimizes the semantic parsing model, compliance rule base, and risk identification standards to improve audit accuracy. The learning and optimization module includes: a feedback collection unit, used to collect feedback and evaluation from auditors on audit results, suggestions and reports, forming a feedback dataset; The model optimization unit is used to dynamically optimize the semantic parsing model, anomaly detection model, and compliance rule base by utilizing new data and feedback data through continuous learning algorithms of the large language model (such as incremental learning or fine-tuning); The standard update unit is used to adjust risk identification standards, compliance thresholds, and weight parameters based on the learning results, ensuring that the system adapts to regulatory changes and business needs. The performance monitoring unit is used to monitor the accuracy, recall, and other performance indicators of the auditing system, triggering optimization cycles and achieving self-iteration.

[0089] In embodiments of this invention, the learning and optimization module enables the audit system to have adaptive evolution capabilities. Through feedback and data-driven processes, the system continuously optimizes its core models and standards, avoiding model aging issues. This module ensures the long-term effectiveness and competitiveness of the audit system, achieving an upgrade from static auditing to dynamic intelligent auditing.

[0090] Through the collaborative work of the above modules, this system has achieved full-process automation, intelligence, and adaptability in the audit of trade union funds compliance, significantly improving audit efficiency, accuracy, and decision support capabilities.

[0091] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A union fund compliance auditing method based on a large language model, characterized by, Comprise: S01, obtaining full data from multi-source data, cleaning and preprocessing data to obtain high-quality standardized audit data; S02, based on large language model and retrieval enhancement generation technology, the standardized audit data is deeply analyzed, the abnormality of fund collection and payment, the vulnerability of asset management and the potential compliance risk are identified, and the audit analysis result is obtained; S03, according to the laws and regulations and the internal management system of trade union, the semantic reasoning ability of large language model is used to construct dynamic compliance rule library, and the compliance of fund use is judged, and the compliance analysis result is obtained; S04, fusion of audit analysis result and compliance analysis result, and through large language model, differential improvement suggestion, rectification path and audit report draft are generated according to the actual scene of trade union, and customized decision support is provided for auditors; S05, establish user feedback and new data iteration learning mechanism, through the continuous learning ability of large language model, dynamically optimize semantic analysis model, compliance rule library and risk identification standard, improve the accuracy of audit.

2. The method of claim 1, wherein, Multi-source data includes trade union financial system, bank flow, payment account, expenditure voucher, contract text and bill scanning; data cleaning and preprocessing includes using OCR technology to extract text information in unstructured data, and using large language model for entity recognition and semantic cleaning, realizing the format unification and privacy desensitization processing of structured and unstructured data.

3. The method of claim 2, wherein, Data cleaning and preprocessing also includes: Synchronize data from multi-source data through API interface and database query; Apply semantic cleaning rules based on large language model for entity recognition, semantic deduplication, format standardization and missing value filling; Based on the semantic consistency rules of large language model, check the logical relationship, numerical correlation, time sequence and format consistency of data in different data sources, and automatically correct inconsistent data; Adopt large language model weighted semantic fusion algorithm and multi-source voting mechanism to fuse the semantic features and classification attributes of audit data, and generate standardized audit data.

4. The method of claim 1, wherein, Deep audit analysis includes: Convert standardized audit data into semantic feature matrix through large language model; Use retrieval enhancement generation technology to retrieve relevant laws and regulations and historical cases from trade union fund compliance knowledge graph, enhance the reasoning ability of large language model; Build an anomaly detection module based on large language model, calculate the reconstruction error or anomaly score of semantic features, and identify abnormal behavior of fund collection and payment; Adopt time sequence semantic analysis method of large language model to predict the trend of fund income and payment progress and expenditure budget execution, and classify the rationality based on deviation ratio; Combine income integrity, expenditure compliance, asset account consistency and special fund special use rate index, calculate the comprehensive risk score through risk scoring function, and locate the high-risk audit field.

5. The method of claim 4, wherein, The anomaly detection module uses a sequence autoencoder or an attention mechanism-based anomaly detector, trains the model with historical normal data, and uses the reconstruction error exceeding a preset threshold as the anomaly determination standard; the rationality classification includes that the deviation ratio less than or equal to 10% is a reasonable class, the deviation ratio greater than 10% and less than or equal to 20% is a focus area, and the deviation ratio greater than 20% is an abnormal class.

6. The method of claim 1, wherein, The construction of the dynamic compliance rule library includes: Collecting legal regulations and internal management system texts, performing word segmentation, entity recognition and relationship extraction through a large language model, and constructing a regulation knowledge base; Using a large language model to extract compliance rules from regulation texts and representing them in the form of conditional statements, and assigning dynamic weights based on authority and timeliness to each rule; Converting standardized audit data into semantic feature vectors, calculating compliance scores through semantic matching degrees, and determining compliance status based on threshold values; Generating differentiated improvement suggestions includes: Building a multi-dimensional decision feature matrix based on risk scores and violation types; Generating improvement suggestion templates for different risk levels through a large language model; Using retrieval enhancement generation technology to retrieve similar scene solutions from historical audit cases; Generating phased rectification paths and time plans through the logical reasoning ability of the large language model; Automatically generating a structured audit report draft, including abstract, problem findings, risk analysis, improvement suggestions and rectification requirements.

7. The method of claim 1, wherein, The learning mechanism includes: Collecting feedback and evaluation of auditors on audit results and suggestions; Using new data and feedback data to dynamically optimize the semantic analysis model, compliance rule library and risk identification standard through the incremental learning algorithm of the large language model; Adjusting compliance thresholds and weight parameters based on performance monitoring indicators to achieve system self-iteration.

8. A union dues compliance auditing system based on large language models, for implementing the method of any one of claims 1-7, characterized in that, The system includes: A multi-source data acquisition and preprocessing module for automatically obtaining full data from multiple sources, cleaning, preprocessing and standardizing the data, and generating high-quality standardized audit data; A deep audit analysis module for deep analysis of standardized audit data based on a large language model and retrieval enhancement generation technology, identifying abnormal expenses, asset management vulnerabilities and potential compliance risks, and outputting audit analysis results; A compliance analysis module for determining the compliance of expenses based on legal regulations and internal management systems, using the semantic reasoning ability of the large language model to construct a dynamic compliance rule library, and outputting compliance analysis results; A decision support and report generation module for integrating audit analysis results and compliance analysis results, and generating differentiated improvement suggestions, rectification paths and audit report drafts through a large language model to provide customized decision support for auditors; A learning and optimization module for establishing a learning mechanism for user feedback and new data iteration, dynamically optimizing the semantic analysis model, compliance rule library and risk identification standard through the continuous learning ability of the large language model, and improving the accuracy of the audit.

9. The system of claim 8, wherein, The deep audit analysis module includes: A semantic feature representation unit for converting standardized audit data into a semantic feature matrix through a large language model; The retrieval enhancement feature engineering unit is configured to retrieve relevant rules and cases from the trade union fund compliance knowledge graph, enhance the contextual understanding ability of the large language model, and extract key semantic features. The anomaly detection unit is configured to build an anomaly detection model based on the large language model, calculate the reconstruction error or anomaly score of the semantic features, and identify abnormal behavior. The time-series semantic analysis unit is configured to use a time-series semantic analysis method of the large language model to predict the trend of the income and expenditure, and perform rationality classification. The risk assessment unit is configured to calculate a comprehensive risk score by a risk scoring function in combination with multi-dimensional indicators. The compliance analysis module includes: The regulation knowledge base construction unit is configured to collect legal and regulatory texts, and form structured rule data through preprocessing by the large language model. The compliance rule extraction unit is configured to extract compliance rules using the large language model, and assign dynamic weights. The audit data featureization unit is configured to convert standardized audit data into semantic feature vectors. The compliance scoring unit is configured to calculate a compliance score based on semantic matching degree, and set a threshold for compliance determination. The dynamic optimization unit is configured to dynamically update the compliance rule library by continuously learning new regulations and user feedback through the large language model. The decision support and report generation module includes: The multi-dimensional feature fusion unit is configured to build a decision feature matrix that integrates risk scores, compliance scores, and business scenario features. The differentiated suggestion generation unit is configured to generate improvement suggestion templates for different risk levels based on the large language model. The retrieval enhancement generation unit is configured to retrieve solutions for similar scenarios from historical audit cases. The rectification path planning unit is configured to generate a phased rectification path and time plan through the large language model. The audit report generation unit is configured to automatically generate a structured audit report draft. The learning and optimization module includes: The feedback collection unit is configured to collect user feedback and evaluations. The model optimization unit is configured to dynamically optimize the semantic analysis model and the compliance rule library using new data and feedback data. The standard update unit is configured to adjust risk identification standards and compliance thresholds. The performance monitoring unit is configured to monitor system performance indicators and trigger optimization cycles.

10. A computer-readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the method of any one of claims 1-7.

Citation Information

Cited By

  • Automatic discovery method for database mode association for penetrating audit

    CN122088658A