Business review method, system, electronic device and medium based on large model

By employing a large-model-based business auditing approach that leverages domain knowledge graphs and dynamic rule update mechanisms, the shortcomings in risk identification and traceability in business auditing are addressed, thereby improving the accuracy and reliability of business audits.

CN120875816BActive Publication Date: 2026-02-10STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511404930.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-02-10
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing business auditing methods are insufficient in identifying and tracing risks when dealing with dynamic evolution characteristics that are multi-dimensional and span multiple time periods, leading to an expansion of blind spots in compliance audits. This is especially true when there are changes in the company's operational structure, making it difficult to maintain information transparency and risk control.

Method used

A business auditing method based on a large model is adopted. By receiving preprocessed business data, target auditing features are extracted and mapped with a pre-set domain knowledge graph to generate associated feature vectors. The large model is used to load the current rule base and historical result records to generate a field missing bitmap, output follow-up instructions to fill in the missing fields, update the associated feature vectors, output the auditing results according to the rule set, and write the latest rule set to non-volatile storage medium.

Benefits of technology

It has enabled the maintenance of information integrity, rule timeliness, and decision consistency in the business audit process, improving the accuracy and reliability of the audit and ensuring the compliance and risk control of enterprise operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875816B_ABST
    Figure CN120875816B_ABST
Patent Text Reader

Abstract

The application relates to a large model-based commercial auditing method and system, electronic equipment and medium, which comprises the following steps: first, receiving pre-processed commercial data and extracting target auditing features, and mapping the target auditing features to an associated feature vector according to a preset field knowledge graph; second, loading auditing decision rule sets and historical auditing result records in a current rule library by a set large model, and generating a field missing bitmap to identify missing fields; if there is a set position, output a follow-up question instruction, update the associated feature vector by supplementing the missing fields through a data acquisition interface, otherwise output an auditing result according to the rule set; finally, appending the current auditing result to the historical record, and covering and writing the new auditing decision rule set into a non-volatile storage medium according to a quantitative evaluation index, for loading in the next auditing round. The method ensures the completeness of the auditing information and the timeliness of the decision rules through the missing field follow-up question and the rule set continuous updating, and improves the accuracy and reliability of the commercial auditing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data auditing technology, and in particular to a business auditing method, system, electronic device, and medium based on a large model. Background Technology

[0002] Business auditing is a management practice in which companies conduct a comprehensive and systematic review of their business activities, operating systems, and compliance. It aims to ensure that business operations comply with laws, regulations, industry standards, and internal systems, while also identifying risks, optimizing processes, and improving operational efficiency. With the exponential expansion of global business activity, the corporate operational chain has evolved from a linear "procurement-production-sales" model to a "networked" structure covering multiple regions, entities, and channels. In this structure, compliance requirements, business processes, funding channels, and ethical norms are like the main nodes, while various sub-businesses, subsidiaries, or external partners are like branch nodes. Whether the "interfaces" between the main and branch nodes can maintain information transparency and risk controllability becomes a crucial factor determining the overall reliability of corporate governance: once a compliance crack or process break occurs at a certain interface, local risks can spread rapidly along contract flows, data flows, or funding flows, triggering regulatory penalties, financial losses, and even brand crises.

[0003] However, after long-term tracking and research, the inventors found that existing commercial auditing methods still have significant limitations in dealing with such "interface" risks: mainstream practices either rely on manual sampling or use isolated algorithms to statically score single-point data, failing to fully couple the dynamic evolution characteristics of multiple dimensions and cross-time periods, resulting in insufficient ability to identify and trace the risk transmission path.

[0004] (1) When branch businesses frequently adjust their counterparties or settlement methods due to market changes, the compliance weight of the main process will be redistributed accordingly, which directly leads to transient conflicts in the audit rules at the "interface". Such conflicts are not limited to the interface itself, but may also spread to adjacent business units through channels such as shared databases and unified financial systems, forming cross-module compliance disturbances and further amplifying the audit blind spots.

[0005] (2) As the most intuitive indicator for measuring business health, the percentage of abnormal transactions is affected by multiple factors such as seasonal fluctuations, policy updates, and exchange rate changes, and exhibits a high degree of nonlinearity. If only cross-sectional data collected in a fixed period and by a fixed module is used, it is difficult to truly reflect the risk exposure level of the enterprise as a whole in the spatiotemporal dimension. Summary of the Invention

[0006] In view of the above-mentioned deficiencies or disadvantages, the present invention provides a business auditing method, system, electronic device and medium based on a large model, which can solve at least one of the above technical problems.

[0007] This invention provides a business auditing method based on a large model, comprising:

[0008] Receive preprocessed business data and extract target audit features from the business data;

[0009] The target review features are mapped to a pre-defined domain knowledge graph to obtain associated feature vectors.

[0010] Input the associated feature vector into the set large model, and the large model loads the review decision rule set and historical review result records stored in the current rule base to generate a field missing bitmap to identify the missing fields of the associated feature vector.

[0011] If a missing field bitmap exists, the large model outputs a follow-up instruction, which is then used by the data acquisition interface to fill in the missing field and write it into the cache to update the associated feature vector.

[0012] If no position is set, the audit result will be output according to the audit decision rule set to complete the business audit.

[0013] The large model appends the current audit results to the historical audit results record, and writes the new audit decision rule set to the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation indicators, so that it can be loaded in the next audit round.

[0014] According to a second aspect, this invention provides a large-scale business auditing system, comprising:

[0015] The audit feature extraction module is used to receive preprocessed business data and extract target audit features from the business data.

[0016] The association vector construction module is used to map the target review features to a preset domain knowledge graph to obtain the association feature vector.

[0017] The review result generation module is used to input the associated feature vector into a predefined large model. The large model loads the review decision rule set and historical review result records stored in the current rule base, generating a field missing bitmap to identify missing fields in the associated feature vector. If the field missing bitmap is set, the large model outputs a follow-up instruction, which fills in the missing fields via the data acquisition interface and writes them to the cache to update the associated feature vector. If the bitmap is not set, the review result is output according to the review decision rule set to complete the commercial review.

[0018] The large model is also used to append the current audit results to the historical audit result record, and to overwrite the new audit decision rule set into the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation indicators, so that it can be loaded in the next audit round.

[0019] According to a third aspect, the present invention provides an electronic device comprising:

[0020] At least one processor; and

[0021] The memory that is communicatively connected to the at least one processor;

[0022] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the large-model-based business auditing methods in the embodiments of the present invention.

[0023] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any of the large-model-based business auditing methods in the embodiments of the present invention.

[0024] The technical solution of this invention extracts target review features from preprocessed commercial data and obtains associated feature vectors by mapping them to a preset domain knowledge graph. Then, a large model is used to load the current review decision rule set and historical review result records to generate a field missing bitmap. When a missing field is identified, a follow-up instruction is immediately output, and the missing field is sequentially filled in and the associated feature vector is updated through the data acquisition interface. Otherwise, the review result is directly output according to the rule set. Subsequently, the review result of this round is appended to the historical record, and the new review decision rule set is overwritten into the non-volatile storage medium according to the quantitative evaluation indicators, so that the next round of review directly loads the latest rules. This ensures that information integrity, rule timeliness, and decision consistency are maintained throughout the entire process, thereby improving the accuracy and reliability of commercial review. Attached Figure Description

[0025] Figure 1 This is a flowchart of a business auditing method based on a large model according to an embodiment of the present invention;

[0026] Figure 2 This is a structural block diagram of a commercial auditing system based on a large model according to an embodiment of the present invention;

[0027] Figure 3 This is a block diagram of an electronic device used to implement embodiments of the present invention. Detailed Implementation

[0028] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0029] According to the first aspect of this invention, a business auditing method based on a large model is provided. This method can be applied to a business auditing system for intelligent terminal devices (hereinafter referred to as the "system"). The intelligent terminal device should be able to run the large model and rule base through local deployment or containerization to complete local business auditing, result output, and rule round updates. Specifically, the intelligent terminal device includes, but is not limited to, computer workstations, business servers, industrial PDAs (Personal Digital Assistants), financial POS (Point of Sale) machines, and embedded industrial control computers. The large model can be pre-set as any Transformer large model such as Qwen3-14B (Qwen 3 14B Parameters), GPT-4 (Generative Pre-trained Transformer 4), LLaMA3 (Large Language Model Meta AI 3), and ChatGLM-4 (Chat General Language Model 4).

[0030] like Figure 1 As shown, the method may include:

[0031] Step S110: Receive the preprocessed business data and extract the target audit features from the business data.

[0032] The preprocessed business data refers to a structured dataset that has undergone format standardization, anomaly removal, and dimensional expansion. The target audit features can be numerical vector features extracted using a multi-head attention mechanism to characterize the risk attributes of a business entity. For example, after extraction by the multi-head attention module, the system obtains a 128-dimensional floating-point vector, where the 0th bit (0.82) represents the enterprise's debt-to-equity ratio risk score, and the 5th bit (-0.31) reflects the weight of administrative penalty records. These values ​​sequentially constitute the input for subsequent mapping and rule comparison, thus enabling the extraction of target audit features.

[0033] The system can continuously receive pre-processed business data through a message queue, and then vectorize the pre-processed business data to obtain target audit features.

[0034] For example, the system can automatically retrieve the previous day's business registration change records, tax details, and contract images at 00:30 every day, and generate target audit features with 128 dimensions after parsing for subsequent mapping. Alternatively, when the system detects the arrival of network packets, it can immediately trigger a streaming feature operator to perform window aggregation on real-time transaction flows, output target audit features, and mark them with timestamps to ensure time sequence consistency.

[0035] Step S120: Map the target review features to the preset domain knowledge graph to obtain the associated feature vector.

[0036] Among them, the domain knowledge graph can be a network of business entity relationships pre-stored in a graph database, where nodes represent enterprises, personnel, and account entities, and edges represent equity, transaction, and guarantee relationships; the association feature vector can refer to the numerical expression of semantic association after entity alignment and embedding operations.

[0037] For example, the system can reduce the dimensionality of the target audit features to 64 dimensions, then retrieve candidate nodes in the domain knowledge graph using cosine similarity. After conflict resolution, the most matching set of target nodes is embedded to generate an associated feature vector for subsequent rule comparison. Furthermore, when different entities share the same company name, the system performs entity disambiguation based on both the unified social credit code and registered capital fields to obtain a uniquely corresponding set of target nodes, thus ensuring the accuracy of the associated feature vector. Conflict resolution refers to the process of uniquely selecting nodes in the candidate node set where the same feature vector matches multiple nodes, according to a domain rule priority table.

[0038] For example, the entity alignment operation described above can be obtained by calculating the semantic similarity between the audit features and the knowledge graph nodes. The semantic similarity can be calculated using the cosine distance formula shown below:

[0039] ;

[0040] in, To review the feature vectors, For knowledge graph node vectors; For vector dot product, and These are the L2 norms of the vectors, which are the square roots of the sum of the squares of the vector's elements and are used to measure the vector's "length" or "size". Further, the output... This is the semantic similarity value, and its range can be... This value is used to match the review features with candidate nodes in the knowledge graph, mapping the review features to the most semantically relevant nodes in the knowledge graph. This is then combined with feature embedding operations to generate associated feature vectors containing semantic relationships, providing input data with a structured representation of domain knowledge for the large model review module, thereby improving the semantic relevance and knowledge reasoning ability of the review features.

[0041] Step S130: Input the associated feature vector into the set large model, and let the large model load the review decision rule set and historical review result records stored in the current rule base, and generate a field missing bitmap to identify the missing fields of the associated feature vector; if the field missing bitmap is set, the large model outputs a follow-up instruction, fills in the missing fields through the data acquisition interface and writes them into the cache to update the associated feature vector; if there is no set, the review result is output according to the review decision rule set to complete the commercial review; the large model appends the review result output this time to the historical review result record, and writes the new review decision rule set to the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation index for loading in the next review round.

[0042] Among them, the field missing bitmap can be a Boolean array with the same length as the dimension of the associated feature vector, and setting a bit indicates that the corresponding field is missing; the follow-up instruction can be a string containing natural language questions to guide users to fill in the missing data; the preset quantitative evaluation index can refer to the set of values ​​that are pre-written into the rule record area to quantify the quality of the current review action, which is a linear combination of accuracy weight (e.g., 0.7) and throughput weight (e.g., 0.3), with a value range of 0–1. The larger the value, the better the review strategy, and it serves as the sole basis for the reinforcement learning controller to replace the review decision rule set.

[0043] The cache can be deployed in the DRAM (Dynamic Random-Access Memory) or a temporary partition of the onboard solid-state drive in the smart terminal device. It is directly managed by the data acquisition module of the smart terminal device and bidirectionally connected to the large model through the system bus. It is used to temporarily store missing fields and updated associated feature vectors, enabling millisecond-level read and write operations for single-round interactions. The non-volatile storage medium can be the smart terminal device's built-in NAND (Not-AND) flash memory array or a pluggable SSD (Solid State Drive), or it can be expanded into a dedicated logical volume of Network Attached Storage (NAS). The large model can complete the cross-round persistent optimization of the audit decision rule set through a batch write link from the cache to the non-volatile storage medium. During runtime, the model only performs real-time read and write operations on the cache. At the startup phase, the latest rules are loaded from the non-volatile storage medium all at once, thus forming a three-level closed loop of "cache-model-non-volatile storage medium" to ensure data consistency and rule timeliness.

[0044] Specifically, the system can generate follow-up questions through the multi-round interactive reasoning submodule built into the large model, and receive the missing fields through the data acquisition interface configured outside the model. Then, the missing fields are written into the cache and merged into the associated feature vector to complete the vector update.

[0045] For example, when the 0th and 5th bits of the missing bitmap are set to 1, the system outputs the follow-up instruction "Please provide the company's latest annual financial statements and the image of the legal representative's certificate". After being collected by the front-end form, the missing fields are filled in and written to the cache through the data collection interface configured outside the model, and the associated feature vector is recalculated for use in the next round of rule comparison.

[0046] Next, after the current review is completed, the system can append the review results to the historical review result record. At the same time, based on the reward value output by the reinforcement learning controller configured externally to the large model, the system replaces the content of the review decision rule set and writes the new review decision rule set to the rule record area of ​​the non-volatile storage medium. This allows the next review to directly load the latest rules, thereby maintaining information integrity, rule timeliness, and decision consistency throughout the entire process and improving the accuracy and reliability of commercial reviews. The reward value is the "instant score" given by the system after the current review, and the preset quantitative evaluation index converts this score into a single value between 0 and 1 by "accuracy weight 0.7 + throughput weight 0.3". In other words, the reward value provides raw feedback, and the quantitative evaluation index standardizes it into the weight basis required for rule replacement. The two have a linear mapping relationship.

[0047] Therefore, according to the above implementation method, the system can extract target review features from preprocessed business data, obtain associated feature vectors by mapping with a preset domain knowledge graph, and then use a large model to load the current review decision rule set and historical review result records to generate a field missing bitmap. When a missing field is identified, a follow-up instruction is immediately output and the missing field is filled in and the associated feature vector is updated sequentially through the data acquisition interface; otherwise, the review result is directly output based on the rule set. Subsequently, the review result of this round is appended to the historical record, and the new review decision rule set is overwritten into the non-volatile storage medium according to the quantitative evaluation indicators, so that the next round of review directly loads the latest rules, thereby continuously maintaining information integrity, rule timeliness and decision consistency throughout the entire process, and improving the accuracy and reliability of business review.

[0048] In some embodiments, the preprocessing steps for business data include:

[0049] Acquire multi-source business data.

[0050] Multi-source business data can refer to a collection of original business information from different business systems, different data formats, and different collection points, including but not limited to enterprise registration change records, tax details, contract images, and transaction records.

[0051] Normalize multi-source business data to obtain structured data with the same semantics.

[0052] The system can normalize field names, data types, and encoding formats from multi-source business data to a predefined semantic schema, forming a row- and column-aligned two-dimensional table structure. Specifically, the semantic schema refers to a data structure template that adds business semantic definitions to a general schema. It not only specifies field names, types, order, and constraints, but also unifies the business meaning, value range, and encoding rules of the fields. This ensures that data from different systems has consistent business interpretation and machine-readable semantics after normalization, facilitating subsequent feature extraction and knowledge graph mapping.

[0053] For example, the system unifies "registered address" and "business address" into "business entity address", and unifies the date format "2024-01-15" and "January 15, 2024" into the string "YYYY-MM-DD", so that subsequent algorithms can read it directly.

[0054] The isolated forest algorithm is used to identify anomalies in the structured data, and abnormal records are removed from the structured data based on the anomaly identification results.

[0055] The system's anomaly identification step involves constructing multiple random partitioning trees to calculate the path length of each record. Records with path lengths below a set threshold are marked as anomalies.

[0056] For example, if the "annual turnover" field in a company's tax record is negative and the path length corresponds to an anomaly score higher than 0.6, the system will remove the entire record and exclude it from subsequent modeling.

[0057] For example, the system identifies abnormal data using the following anomaly scoring formula:

[0058] ;

[0059] in, For the sample exist Anomaly scores in isolated trees, with values ​​ranging from 1 to 2. For the sample exist Average path length in the trees: for Average path length of trees; number of trees The value range is 50-200, and the preset tree depth for a single tree does not exceed [a certain value]. ,in This represents the number of samples.

[0060] Virtual samples are generated by expanding the dimensions of the structured data after removing outlier records using a configured generative adversarial network.

[0061] The generative adversarial network configured in this way can consist of a generator and a discriminator. The generator takes a random noise vector as input and outputs a virtual sample with the same dimension as the real sample. The discriminator performs binary classification on the real sample and the virtual sample. After alternating training, the generator can output a virtual sample that is consistent with the real distribution.

[0062] For example, in a scenario where the real sample contains only 5,000 loan records for micro and small enterprises, the generator outputs an additional 1,000 virtual samples, which expands the size of the subsequent model training set and improves the generalization ability.

[0063] For example, the system augments the feature dimensions of business data using a generative adversarial network (GAN). The generator loss function of the GAN is as follows:

[0064] ;

[0065] in, It is a random noise vector that follows a probability distribution. For generator; For discriminators; Used for noise vector The expectation operation; the generative adversarial network adopts an alternating training mechanism, and the discriminator loss function is as follows:

[0066] ;

[0067] in, This represents the actual data distribution.

[0068] The virtual sample is merged with the structured data and written into the cache to obtain the preprocessed business data.

[0069] The system can append virtual samples to the end of structured data by merging them column-aligned to form an extended dataset. For example, the system can concatenate 1,000 generated virtual samples with the original 5,000 real samples in the same field order to form an extended dataset of 6,000 rows. This extended dataset is then written in binary format to the temporary buffer of the DRAM in the smart terminal device for direct use in subsequent feature extraction steps.

[0070] Therefore, according to the above implementation method, the system can complete the entire process of preprocessing, including multi-source data normalization, anomaly removal, dimension expansion, and cache writing, on the local terminal, providing a commercial data foundation that is complete in fields, evenly distributed, and available in real time for subsequent large model review.

[0071] In some embodiments, the target review features are mapped to a preset domain knowledge graph to obtain an associated feature vector, including:

[0072] The configured multi-head attention mechanism module is used to extract features from the target review features to obtain the review features to be mapped.

[0073] Specifically, the multi-head attention mechanism module can be composed of multiple self-attention heads running in parallel. Each head independently calculates the query, key, and value vectors and outputs weighted features, which are then concatenated and linearly transformed to form the features to be mapped for review.

[0074] Assuming there are 8 heads and 256 output dimensions, it can simultaneously capture the relationships between three different subspaces—equity, liabilities, and transactions—in a single forward computation, forming a unified vector representation.

[0075] For example, the system is also configured with a feature extraction module, which internally includes a multi-head attention knowledge graph construction unit. This multi-head attention knowledge graph construction unit is used to extract review features through a multi-head attention mechanism, the calculation formula of which is as follows:

[0076] ;

[0077]

[0078] in, This is the query matrix, used to represent the query vectors from which the features to be extracted are located; This is the key matrix, used to store key information about feature associations; It is a value matrix used to store the specific numerical information of the features; The key vector dimension is used to scale the attention score; This represents the number of attention heads, used to account for the number of attention mechanisms involved in parallel computing. , is the key matrix The transpose of the query vector is used to calculate the similarity between the query vector and the key vector; This is a splicing operation used to integrate the outputs of various attention heads: The output projection matrix is ​​used to map the concatenated feature vectors to the target dimension.

[0079] The dimensionality of the features to be mapped and reviewed is reduced, and a knowledge graph node index is constructed based on the reduced dimensionality features.

[0080] The system can use Principal Component Analysis (PCA) to transform the features to be mapped and reviewed for dimensionality reduction, projecting the high-dimensional features to be mapped and reviewed into a low-dimensional principal component space, retaining the components with a cumulative variance contribution rate of not less than 95%; then, the system uses the dimensionality-reduced feature vectors as keys to construct an inverted list with a fixed step size, forming the knowledge graph node index, which is used to quickly locate candidate nodes.

[0081] For example, the system reduces a 256-dimensional vector to 64 dimensions, and then uses the 64-dimensional values ​​as keys to build a B+ tree index in the graph database, enabling millisecond-level range queries.

[0082] The cosine similarity algorithm is used to match candidate nodes in the knowledge graph node index, resulting in a set of candidate nodes.

[0083] For example, the system uses the cosine similarity algorithm to calculate the cosine of the angle between the dimensionality-reduced feature to be mapped and the index key, and selects nodes with a cosine value greater than a preset threshold of 0.8 as candidates. When the cosine value of the feature vector of the enterprise to be reviewed and the key of the "Unified Social Credit Code 91xxxxxx" node in the graph database is 0.85, the node is included in the candidate node set for subsequent conflict resolution.

[0084] The candidate node set is reconciled based on the preset domain rules to obtain the target node set.

[0085] The preset domain rules can refer to a uniqueness constraint table pre-written into the rule base, including dual verification items such as "unique unified social credit code" and "unique enterprise name + registered address". The system can traverse candidate nodes according to the priority order in the table. If multiple nodes correspond to the same enterprise identity, the nodes with completely matching credit codes and consistent registered addresses are retained, and the rest are discarded, thus forming the target node set. `i` is an index variable representing the i-th attention head. This represents one of many attention heads used in parallel computation. `h` is a constant representing the total number of attention heads. Therefore, Literally, it refers to the last attention head (i.e., the h-th head).

[0086] The dimensionality-reduced features to be mapped are embedded into the target node set to generate associated feature vectors.

[0087] When performing the above embedding operation, the system can concatenate the 64-dimensional feature vector after dimensionality reduction with the attribute vector of the target node, and then regress it to the original dimension through a linear mapping layer to obtain the associated feature vector that is aligned with the semantics of the graph.

[0088] For example, the system concatenates the enterprise's financial risk characteristics with the industry, shareholders, and historical default attributes in the graph nodes, and outputs a 256-dimensional associated feature vector through a fully connected layer for subsequent rule comparison.

[0089] Therefore, according to the above implementation method, the system can complete the entire process of feature extraction, dimensionality reduction indexing, similarity matching, conflict resolution and vector embedding on the local terminal, providing a semantically consistent and uniquely pointing related feature vector basis for subsequent large model review.

[0090] In some embodiments, the large model is configured with multi-round interactive inference submodules; the steps of the large model outputting follow-up questions, filling in missing fields via a data acquisition interface and writing them into a cache to update the associated feature vector include:

[0091] The multi-round interactive reasoning submodule generates semantic follow-up questions based on the field missing bitmap.

[0092] The multi-turn interactive reasoning submodule can be a sequence-to-sequence network embedded in a large model, with a field missing bitmap as input and a follow-up text in natural language form as output.

[0093] For example, when the zeroth and fifth bits of the bitmap are set to 1, the submodule outputs the follow-up instruction "Please supplement the company's latest annual financial statements and the image of the legal representative's certificate", with a text length not exceeding 256 characters, for direct rendering by the front-end form.

[0094] Send semantic follow-up commands to the data acquisition interface to retrieve missing fields.

[0095] The data acquisition interface can be a REST (Representational State Transfer) style API (Application Programming Interface), with a fixed interface path of " / api / v1 / missing" and using the HTTP (HyperText Transfer Protocol) POST method. After receiving a follow-up instruction, the front end pops up a file upload box. The user selects a PDF (Portable Document Format) financial statement and a JPG (Joint Photographic Experts Group) document image. The system encapsulates the file content into a JSON (JavaScript Object Notation) message using Base64 encoding and sends it back to the data acquisition interface. The timeout threshold for the return is 30 seconds.

[0096] The missing fields are filled in and written to the cache, and then merged into the associated feature vector to complete the vector update.

[0097] The cache can be a circular buffer in the DRAM of the terminal device, with a fixed size of 496kB (bytes). After receiving the missing field, the background service immediately decodes the Base64 string into a binary stream and writes it to the end of the circular buffer. Then, it calls the feature extraction submodule to vectorize the missing file to obtain a 64-dimensional missing feature vector. This vector is then concatenated with the original associated feature vector along the channel dimension to form an updated associated feature vector for use in the next round of rule comparison.

[0098] Therefore, according to the above implementation method, the system can immediately generate renderable follow-up instructions after the missing field is identified, and complete file return, cache writing and vector merging through standardized interfaces to realize closed-loop supplementation of missing information, providing a complete input basis for subsequent review decisions.

[0099] In some embodiments, the large model is further configured with a dynamic optimization module, which includes a reinforcement learning controller and an active learning sampler; the steps of appending the current review result to the historical review result record and overwriting the new review decision rule set into the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation index include:

[0100] By using a reinforcement learning controller, the target action is determined in the action space formed by the historical review results and the current associated feature vectors. The reward value corresponding to the target action is written into the quantitative evaluation index, and the content of the review decision rule set is replaced according to the quantitative evaluation index.

[0101] The state space can be formed by concatenating the most recent 100 records in the historical review results with the current associated feature vector, with a dimension of 320; the action space can include three discrete actions: pass, reject, and follow-up question; the reward value can be formed by a trilinear combination of an accuracy weight of 0.7 and a throughput weight of zero.

[0102] For example, when the target action is determined to be "follow-up question" and the review result changes from rejection to approval after the field is filled in, the system assigns a reward value of positive one and writes the reward value into the quantitative evaluation index. Then, the content of the review decision rule set is replaced according to the reward value. The replacement strategy is to move the upper and lower bounds of the original threshold range by 5% in the direction that increases the reward value.

[0103] By using an active learning sampler to perform Bayesian uncertainty sampling on historical review result records, high uncertainty samples are selected and added to the incremental training set. The incremental training set after adding the samples is then supplemented into the historical review result records.

[0104] In this context, Bayesian uncertainty sampling of the system can refer to performing ten forward inferences on the same input using the Monte Carlo Dropout method, and calculating the output entropy as a measure of uncertainty.

[0105] For example, if the output entropy of a corporate loan record is greater than 0.35 after ten inferences, it is judged as a high-uncertainty sample and added to the incremental training set. The incremental training set is a circular buffer with a fixed capacity of 500 records. When the buffer is full, the new sample overwrites the earliest sample, and then all the contents of the incremental training set are added to the historical review result record to form the updated historical review result record.

[0106] For example, an active learning sampler can be used to optimize training data through Bayesian uncertainty sampling, the sample uncertainty measure of which is as follows:

[0107] ;

[0108] in, The sample to be evaluated; For the model to sample Category The predicted probability, The highest predicted probability; For uncertainty scoring, the range of values ​​is... .

[0109] Furthermore, the reinforcement learning controller in the dynamic optimization module uses the current review status... and review actions As input, the formula is obtained through the Q-learning algorithm. Calculate the value function of state-to-action pairs ;

[0110] Among the rewards according to Throughput calculation, For learning rate, As a discount factor, The next review status is... For the next action, the value function is iteratively updated to dynamically adjust the review strategy of the large model.

[0111] The combination of reinforcement learning controller and active learning sampler is based on the review decision rule set and training data. The review strategy is dynamically adjusted through reinforcement learning, while uncertainty sampling is used to improve the quality of training data, forming a two-way optimization mechanism between review strategy and training data, which continuously improves the accuracy and adaptability of large model review.

[0112] The new set of audit decision rules, along with the historical audit results records supplemented by the incremental training set, are overwritten and written to the rule record area of ​​the non-volatile storage medium.

[0113] The system's overwrite can refer to the following: using page alignment, with each page being 4096 bytes in size, performing CRC (Cyclic Redundancy Check) verification before writing, and after the verification passes, the system writes the new audit decision rule set and the updated historical audit result records into the rule record area at once. After writing is completed, the system sends an ACK (Acknowledgement) signal to the main control chip, indicating that the latest rules and records can be loaded for the next round of audit.

[0114] Therefore, according to the above implementation method, the system can dynamically adjust the rule parameters based on actual feedback after each audit round, supplement high uncertainty samples, and persist the latest rules and records, thereby realizing the self-evolution of audit strategies and online adaptation of data distribution, and continuously improving the accuracy and robustness of commercial audits.

[0115] In some embodiments, the large model is also used before loading the current rule base:

[0116] The parameter transfer technique is used to adapt large models across domains. The parameter transfer technique adopts a prefix fine-tuning method, which freezes the main parameters of the large model and only updates the cross-domain adapter.

[0117] Among them, the parameter transfer technique mentioned above refers to the technique of making a large model quickly adapt to the target domain distribution by concatenating a trainable tensor sequence at the front end of the input sequence while keeping the source domain pre-trained weights unchanged; the prefix fine-tuning method refers to updating the gradient only on the tensor sequence while freezing all weights of the Transformer decoder; the cross-domain adapter can be a continuous vector of length 128 with the same dimension as the word embedding, stored in the adapter partition of the non-volatile storage medium.

[0118] For example, when the source domain is finance and the target domain is healthcare, the system appends a 128-dimensional trainable prefix to the input sequence. After three training steps, the cross-domain adapter parameters are updated, completing the domain transfer.

[0119] For example, a cross-domain adapter is used to transfer model parameters to cross-domain data using a prefix fine-tuning technique. The parameter optimization objective formula for the prefix fine-tuning technique is as follows:

[0120] ;

[0121] Where prefix is ​​a trainable prefix vector, which is concatenated to the beginning of the input sequence; The cross-entropy loss function; For large language models; For inputting business data; For the audit labels corresponding to the input business data.

[0122] Load the updated cross-domain adapter into the large model to complete the cross-domain adaptation of the large model.

[0123] The loading process described above can be executed by the model initialization submodule configured externally to the large model: After power-on reset, the system reads the latest cross-domain adapter from the adapter partition and splices it to the front end of the input embedding layer to form an extended input; subsequently, the frozen main parameters remain unchanged, the extended input enters the Transformer computation graph, and the output reflects the target domain features. Specifically, the adapter partition is located in non-volatile storage medium, belongs to the system storage area, and is managed by the dynamic optimization module; the large model is only read and loaded once at startup by the model initialization submodule and is not accessed again during runtime.

[0124] For example, after loading, the same company name is assigned a higher risk weight in the medical domain than in the financial domain, and the difference can be achieved without modifying the subject parameters.

[0125] Therefore, according to the above implementation method, the system can quickly complete cross-domain adaptation through prefix fine-tuning without increasing the training overhead of all parameters, ensuring that the large model takes effect immediately in new industry scenarios, and continuously works in collaboration with subsequent rule rounds of updates, thereby improving the migration efficiency and deployment flexibility of commercial review.

[0126] In some embodiments, the audit decision rule set includes threshold rules, logical rules, and priority rules;

[0127] Threshold rules are used to compare continuous fields in associated feature vectors with preset threshold ranges to generate primary pass or rejection labels.

[0128] The preset threshold range can refer to the upper and lower bound value pairs stored in the rule record area. For example, the threshold range of the continuous field "asset-liability ratio" is set to 60% to 80%. When the value of this field in the associated feature vector is 75%, it falls into the range, and the threshold rule generates a "pass" primary mark. Otherwise, it generates a "reject" primary mark.

[0129] Logical rules are used to perform Boolean combinations on primary tags and discrete fields to obtain the combination determination result.

[0130] Boolean combination refers to concatenating the primary marker with the discrete field using the logical operators "AND", "OR", and "NOT". For example, when the primary marker is "Pass" and the discrete field "Administrative Penalty Record" is equal to "None", the logical rule outputs the "Pass" combination judgment result; if the primary marker is "Pass" but the administrative penalty record is "Yes", then the "Reject" combination judgment result is output, thereby avoiding misjudgment by a single indicator.

[0131] Priority rules are used to select the final output action based on the domain priority table when there is a conflict in the combined judgment results, and write the final output action into the review results.

[0132] The domain priority can be represented as a two-dimensional array stored in non-volatile storage medium. The row number represents the rule number, and the column number represents the priority weight. The larger the weight value, the higher the priority. For example, when the "threshold rule" outputs "pass" and the "logic rule" outputs "reject", the system queries the domain priority table. If the weight corresponding to the logical rule is 0.8 and the weight corresponding to the threshold rule is 0.6, then the "reject" with the higher weight is selected as the final output action and written into the review result to ensure decision consistency.

[0133] Therefore, according to the above implementation method, the system can output a unique and traceable final review action through a three-level progressive judgment of threshold-logic-priority in the case of rule conflict, thereby improving the stability and interpretability of commercial review.

[0134] Figure 2 This is a structural block diagram of a commercial auditing system based on a large model according to an embodiment of the present invention.

[0135] like Figure 2 As shown, this large-model-based business auditing system includes:

[0136] The audit feature extraction module 210 is used to receive preprocessed business data and extract target audit features from the business data.

[0137] The association vector construction module 220 is used to map the target review features to a preset domain knowledge graph to obtain the association feature vector.

[0138] The review result generation module 230 is used to input the associated feature vector into the set large model, causing the large model to load the review decision rule set and historical review result records stored in the current rule base, and generate a field missing bitmap to identify the missing fields of the associated feature vector; if the field missing bitmap is set, the large model outputs a follow-up instruction, fills in the missing fields through the data acquisition interface and writes them into the cache to update the associated feature vector; if it is not set, the review result is output according to the review decision rule set to complete the commercial review.

[0139] The large model is also used to append the current audit results to the historical audit result record, and to overwrite the new audit decision rule set into the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation indicators, so that it can be loaded in the next audit round.

[0140] The specific functions and examples of each module and submodule of the device in this embodiment can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0141] According to embodiments of the present invention, the above-described method of the present invention can be applied to an electronic device and a readable storage medium.

[0142] Figure 3 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0143] like Figure 3As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0144] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0145] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as a large-model-based business auditing method. For example, in some embodiments, a large-model-based business auditing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of a large-model-based business auditing method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured, by any other suitable means (e.g., by means of firmware), to perform a business auditing method based on a large model.

[0146] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT or LCD monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0150] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0151] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0152] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.

Claims

1. A business review method based on a large model, characterized in that, The large model is configured with multi-round interactive inference sub-modules; the method includes: Receive preprocessed business data and extract target audit features from the business data; The target review features are mapped to a preset domain knowledge graph to obtain an associated feature vector; The associated feature vector is input into a set large model, which loads the review decision rule set and historical review result records stored in the current rule base to generate a field missing bitmap to identify the missing fields of the associated feature vector; If the missing bitmap of the field is set, the large model outputs a follow-up instruction, which fills in the missing field through the data acquisition interface and writes it into the cache to update the associated feature vector; If no setting exists, the audit result is output according to the audit decision rule set to complete the business audit; The large model appends the current audit result to the historical audit result record, and writes the new audit decision rule set into the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation indicators, so that it can be loaded in the next audit round. The steps of the large model outputting follow-up instructions, filling in missing fields via the data acquisition interface, and writing them into the cache to update the associated feature vector include: The multi-round interactive reasoning submodule generates semantic follow-up instructions based on the field missing bitmap. The semantic follow-up command is sent to the data acquisition interface to retrieve the missing fields; The missing fields are written into the cache, and the missing fields are merged into the associated feature vector to complete the vector update.

2. The method according to claim 1, characterized in that, The preprocessing steps for the business data include: Acquire multi-source business data; The multi-source business data is normalized to obtain structured data with the same semantics; The isolated forest algorithm is used to identify anomalies in the structured data, and abnormal records are removed from the structured data based on the anomaly identification results. The structured data after removing outlier records is augmented with dimensions using a configured generative adversarial network to generate virtual samples. The virtual sample is merged with the structured data and written into the cache to obtain the preprocessed business data.

3. The method according to claim 2, characterized in that, The step of mapping the target review features to a preset domain knowledge graph to obtain an associated feature vector includes: The configured multi-head attention mechanism module is used to extract features from the target review features to obtain the review features to be mapped; The dimensionality reduction process is performed on the features to be mapped for review, and a knowledge graph node index is constructed based on the dimensionality-reduced features to be mapped for review. The knowledge graph node index is matched with candidate nodes using the cosine similarity algorithm to obtain a candidate node set. The candidate node set is reconciled based on preset domain rules to obtain the target node set; The dimensionality-reduced features to be mapped are embedded into the target node set to generate the associated feature vector.

4. The method according to claim 3, characterized in that, The large model is also equipped with a dynamic optimization module, which includes a reinforcement learning controller and an active learning sampler; the steps of appending the current review result to the historical review result record and overwriting the new review decision rule set into the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation index include: The reinforcement learning controller determines the target action in the action space of pass, reject and follow-up questioning based on the state space formed by the historical review result record and the current associated feature vector. The reward value corresponding to the target action is written into the quantitative evaluation index, and the content of the review decision rule set is replaced according to the quantitative evaluation index. The active learning sampler performs Bayesian uncertainty sampling on the historical review result records, filters out high uncertainty samples and adds them to the incremental training set, and then supplements the historical review result records with the added samples in the incremental training set. The new set of audit decision rules, along with the historical audit result records supplemented by the incremental training set, are overwritten and written to the rule record area of ​​the non-volatile storage medium.

5. The method according to claim 1, characterized in that, Before loading the current rule base, the large model is also used for: The large model is adapted across domains using parameter transfer technology, which employs a prefix fine-tuning method to freeze the main parameters of the large model and only update the cross-domain adapter. The updated cross-domain adapter is loaded into the large model to complete the cross-domain adaptation of the large model.

6. The method according to claim 1, characterized in that, The set of review decision rules includes threshold rules, logical rules, and priority rules; The threshold rule is used to compare the continuous fields in the associated feature vector with a preset threshold range to generate a primary pass or rejection flag. The logical rules are used to perform Boolean combinations on the primary markers and discrete fields to obtain a combination determination result; The priority rule is used to select the final output action according to the domain priority table when the combined judgment results conflict, and to write the final output action into the review result.

7. A commercial review system based on a large model, characterized in that, The large model is configured with multi-turn interactive reasoning sub-modules. This commercial auditing system includes: The audit feature extraction module is used to receive preprocessed business data and extract target audit features from the business data. The association vector construction module is used to map the target review features to a preset domain knowledge graph to obtain an association feature vector; The review result generation module is used to input the associated feature vector into a set large model, causing the large model to load the review decision rule set and historical review result records stored in the current rule base, and generate a field missing bitmap to identify the missing fields of the associated feature vector; if the field missing bitmap is set, the large model outputs a follow-up instruction, fills in the missing fields through the data acquisition interface and writes them into the cache to update the associated feature vector; if no field is set, the review result is output according to the review decision rule set to complete the commercial review; The large model is also used to append the current audit results to the historical audit results record, and to write the new audit decision rule set to the rule record area of ​​the non-volatile storage medium according to the preset quantitative evaluation indicators, so that it can be loaded in the next audit round. The audit result generation module is also used to generate semantic follow-up instructions based on the field missing bitmap through the multi-round interactive reasoning submodule; The semantic follow-up command is sent to the data acquisition interface to retrieve the missing fields; The missing fields are written into the cache, and the missing fields are merged into the associated feature vector to complete the vector update.

8. An electronic device, characterized in that, include: At least one processor; as well as The memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Building construction scheme intelligent auditing system based on large model

    CN120598724A

  • Water conservancy industry electronic dark bidding document enterprise internal examination method and system based on artificial intelligence

    CN120707259A