Template automatic updating method and system

By standardizing the original policy information and using the semantic parsing and mapping layer of a pre-trained large language model for deep adaptation, the problem of insufficient correlation between policy and audit information is solved, and highly accurate template updates are achieved.

CN121979899APending Publication Date: 2026-05-05NANJING AUDIT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING AUDIT UNIV
Filing Date
2026-01-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies lack in-depth two-way adaptation and targeted processing mechanisms for correlation when handling policy-related information and audit-related content. This results in the integrated information failing to fully reflect the inherent connection between the two, leading to insufficient accuracy in the generated policy and audit correspondence results.

Method used

By responding to audit template update requests, the system obtains and standardizes the original policy information. It then uses a pre-trained, fine-tuned large language model with a policy audit semantic parsing and mapping layer and a mapping result calibration and optimization layer to perform semantic parsing and mapping, generating structured mapping correction data. The system also iteratively fine-tunes the model to improve accuracy when the target industry template does not meet the preset conditions.

Benefits of technology

It achieves a high degree of accuracy in linking policies and audit elements, ensuring the adaptability and accuracy of target industry templates to policy requirements, and improving the efficiency and quality of audit template updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979899A_ABST
    Figure CN121979899A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic template updating method and system, and relates to the technical field of auditing informationization, original policy information is standardized to obtain standardized corpus data by responding to an auditing template updating request, and a pre-training fine-tuning large language model comprising a policy auditing semantic analysis mapping layer and a mapping result calibration optimization layer is called to obtain a pre-training fine-tuning large language model; inputting the standardized corpus into a policy audit semantic analysis mapping layer to complete semantic analysis and mapping so as to output policy audit initial mapping data, verifying and correcting the initial mapping data through a mapping result calibration optimization layer to obtain structured mapping correction data, and matching a target industry template based on the structured data, and if the target industry template does not meet the preset template evaluation condition, performing fine adjustment on the fine adjustment large language model and circularly executing the mapping process until the template meets the condition, thereby ensuring the suitability and accuracy of the target industry template to policy requirements, and solving the problem of insufficient accuracy of a policy and audit corresponding result in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audit information technology, and in particular to a method and system for automatic template updating. Background Technology

[0002] With the rapid development of the digital economy, the informatization and intelligent transformation of auditing work has become a core trend in the industry. As a key carrier for conducting audit business, the efficiency and accuracy of audit templates in adapting to the latest industry rules are directly related to the compliance and professionalism of audit work. Against this background, the automated updating of audit templates based on pre-trained language models has become a key direction for technology research and development and industry application.

[0003] Currently, related technologies have attempted to apply large language models to audit template update scenarios, extracting key information related to policies and audit content to assist in the correspondence between the two. However, these technologies often employ simple integration or one-way association when processing these two types of information. They fail to conduct in-depth mutual adaptation based on the characteristics of the two types of information, nor do they perform targeted processing based on the degree of correlation between the information, thus failing to build a processing mechanism adapted to the characteristics of the two types of information. This limitation leads to the integrated information failing to fully reflect the inherent connection between policies and audit content, resulting in insufficient effective correlation and consequently, insufficient accuracy in the generated policy-audit correspondence results. Ultimately, this affects the accuracy and efficiency of audit template updates, failing to meet the actual needs of the financial and corporate sectors for rapid adaptation of audit templates to policy adjustments. Summary of the Invention

[0004] This invention provides a method and system for automatic template updating, which solves the technical problem in the prior art that the lack of a mechanism for deep two-way adaptation and targeted processing of correlation between policy-related information and audit-related content leads to the inability of integrated information to fully reflect the inherent relationship between the two, resulting in insufficient accuracy of the generated policy and audit correspondence results.

[0005] The first aspect of this invention provides a method for automatic template updating, comprising: In response to the audit template update request, the original policy information is obtained and standardized to obtain standardized corpus data; Obtain a pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer; The standardized corpus data is input into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and the initial policy audit mapping data is output. The initial mapping data from the policy audit is used as input to the calibration and optimization layer of the mapping result for verification and correction, generating structured mapping correction data; Based on the structured mapping, correct the data to match the target industry template; When the target industry template does not meet the preset template evaluation conditions, the fine-tuned large language model is fine-tuned, and the process jumps to the step of using the standardized corpus data to input the policy audit semantic parsing and mapping layer for semantic parsing and mapping, until the target industry template meets the preset template evaluation conditions.

[0006] Optionally, the response to the audit template update request, obtaining the original policy information and standardizing it to obtain standardized corpus data, includes: Respond to audit template update requests and obtain original policy information; The original policy information is cleaned to obtain policy text data; Key corpus extraction is performed on the policy text data to obtain core policy element data; Semantic noise reduction and unnecessary information removal are performed on the core policy element data to obtain key corpus data; The key corpus data is converted from non-standardized terminology to obtain standardized element data. The standardized element data is structured and encapsulated to obtain standardized corpus data.

[0007] Optionally, the policy audit semantic parsing and mapping layer includes a corpus preprocessing unit, a domain feature enhancement and extraction unit, a policy audit semantic fusion unit, and a mapping inference unit. The step of inputting the standardized corpus data into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and outputting initial policy audit mapping data, includes: The standardized corpus data is input into the corpus preprocessing unit, where format verification, noise filtering, and policy audit terminology alignment are performed sequentially, and standardized corpus data is output. The standardized corpus data is input into the domain feature enhancement extraction unit to perform dual feature extraction of policy semantics and audit elements, and output dual domain feature vectors; The dual-domain feature vectors are input into the policy audit semantic fusion unit for feature fusion and normalization processing, and the fused features are output. The fusion features are input into the mapping inference unit for mapping, and the initial mapping data for policy audit is output.

[0008] Optionally, the dual-domain feature vector includes a policy semantic feature vector and an audit element feature vector. The step of inputting the dual-domain feature vector into the policy audit semantic fusion unit for feature fusion and normalization processing, and outputting fused features, includes: Bidirectional semantic matching is performed between the policy semantic feature vector and the audit element feature vector, and the element-level similarity matrix between the vectors is calculated to obtain the association weight data. Based on the aforementioned correlation weight data, the policy semantic feature vector and the audit element feature vector are weighted respectively to obtain policy correlation features and audit correlation features; The policy-related features and the audit-related features are concatenated along the channel dimension, and the concatenated features are integrated using a 1×1 convolution to obtain the related feature data. The associated feature data and the dual-domain feature vector are element-weighted and fused according to preset weights to obtain initial fused features; The initial fusion features are subjected to layer normalization to obtain the fusion features.

[0009] Optionally, the step of using the initial mapping data from the policy audit as input to the mapping result calibration and optimization layer for verification and correction, generating structured mapping correction data, includes: The initial policy audit mapping data is subjected to compliance verification to obtain compliant mapping data and abnormal mapping data; The abnormal mapping data is corrected to obtain corrected mapping data; The compliance mapping data and the corrected mapping data are integrated and structured to obtain structured mapping corrected data.

[0010] Optionally, the matching of the target industry template based on the structured mapping correction data includes: Applicable industry tags and audit elements are extracted from the structured mapping correction data and used as template matching search keywords; The preset industry template set is searched using the search keywords, and multiple candidate industry templates that match the applicable industry are selected. The audit elements are semantically compared with the audit indicators of each candidate industry template, and the matching similarity is calculated. Based on the matching similarity ranking, the candidate industry template with the highest similarity is selected as the target industry template.

[0011] Optionally, the preset template evaluation condition specifically means that the coverage ratio of the audit indicators of the target industry template to the audit elements in the structured mapping correction data is not less than a preset coverage ratio threshold.

[0012] A second aspect of the present invention provides a template automatic update system, comprising: The policy collection and parsing module is used to respond to audit template update requests, obtain raw policy information and standardize it to obtain standardized corpus data. The model acquisition module is used to acquire a pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer. The policy semantic understanding and impact analysis module is used to input the standardized corpus data into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and output the initial policy audit mapping data. The verification and correction module is used to verify and correct the mapping result by inputting the initial mapping data of the policy audit into the calibration and optimization layer, and generate structured mapping correction data. The template generation and configuration module is used to match target industry templates based on the structured mapping correction data; The template push and feedback module is used to fine-tune the fine-tuning large language model when the target industry template does not meet the preset template evaluation conditions, and then jump to the step of using the standardized corpus data to input the policy audit semantic parsing and mapping layer for semantic parsing and mapping, until the target industry template meets the preset template evaluation conditions.

[0013] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the template automatic update method as described above.

[0014] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the template automatic update method as described above.

[0015] As can be seen from the above technical solutions, the present invention has the following advantages: (1) This invention provides a method and system for automatic template updating. In response to an audit template update request, the original policy information is first standardized to obtain standardized corpus data. Then, a pre-trained fine-tuned large language model containing a policy audit semantic parsing and mapping layer and a mapping result calibration and optimization layer is called. The standardized corpus is input into the policy audit semantic parsing and mapping layer to complete semantic parsing and mapping, outputting initial policy audit mapping data. The initial mapping data is then verified and corrected by the mapping result calibration and optimization layer to obtain structured mapping correction data. Based on this structured data, a target industry template is matched. If the target industry template does not meet the preset template evaluation conditions, the fine-tuned large language model is fine-tuned, and the above mapping process is repeated until the template meets the conditions. In this invention, the processing of standardized corpus can first eliminate redundancy and irregularities in the policy information. The model provides high-quality basic input for the correlation mapping, while the policy audit semantic parsing mapping layer can perform targeted deep adaptation processing on policy and audit-related content, making up for the lack of bidirectional adaptation and targeted correlation processing in existing technologies. This allows the inherent correlation information between policy and audit elements to be fully carried out, thus enabling the output policy audit initial mapping data to have a more accurate correspondence. Combined with the verification and correction of the mapping result calibration and optimization layer, data deviations can be further corrected and data reliability can be improved. At the same time, the model's cyclic fine-tuning mechanism can continuously optimize the mapping effect. Ultimately, through the synergistic effect of this series of links, the problem of insufficient accuracy of policy and audit correspondence results in existing technologies is effectively solved, ensuring the adaptability and accuracy of the target industry template to policy requirements, and simultaneously improving the efficiency and quality of audit template updates.

[0016] (2) The present invention constructs a dedicated four-unit collaborative architecture adapted to the policy audit field as a policy audit semantic parsing mapping layer. Unlike the generalized processing mode of the existing general semantic parsing module, the policy audit semantic parsing mapping layer integrates the format verification, noise filtering and policy audit terminology alignment dedicated purification mechanism of the corpus preprocessing unit, combined with the policy semantic and audit element dual feature targeted extraction function of the domain feature enhancement extraction unit, and combined with the domain adaptation processing logic of the policy audit semantic fusion unit and the mapping reasoning unit, to realize the targeted adaptation and collaborative linkage of each unit to the policy audit scenario, and solve the problem that the existing general module cannot adapt to the domain characteristics from the architecture level.

[0017] (3) In the policy audit semantic fusion unit, a dual-domain feature deep bidirectional adaptation technology system is adopted. Unlike the crude processing method of simple splicing or one-way matching in the existing technology, the bidirectional semantic matching mechanism is introduced to actively explore the intrinsic relationship between policy semantic feature vector and audit element feature vector. The element-level similarity matrix is ​​used to calculate the accurate association weight. Based on the weight, the dual feature vector is weighted and processed. At the same time, the feature optimization structure of channel dimension splicing and 1×1 convolution integration, the secondary weighted fusion of associated features and original dual-domain feature vectors and the layer normalization stabilization mechanism are combined to form a complete set of association degree targeted processing technology solutions. This solves the defect of existing technology that cannot fully bear the intrinsic relationship between the two domains from the feature fusion level. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating the steps of an automatic template updating method provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of an automatic template updating system provided in an embodiment of the present invention; Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0020] This invention provides a method and system for automatic template updating, which addresses the technical problem in the prior art where the processing of policy-related information and audit-related content lacks a mechanism for in-depth bidirectional adaptation and targeted processing of correlation, resulting in the integrated information failing to fully reflect the inherent relationship between the two, and consequently leading to insufficient accuracy in the generated policy and audit correspondence results.

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 , Figure 1The flowchart illustrates the steps of an automatic template updating method provided in an embodiment of the present invention.

[0023] This invention provides a method for automatic template updating, comprising: Step 101: Respond to the audit template update request, obtain the original policy information and standardize it to obtain standardized corpus data.

[0024] Audit template update request: refers to a request signal triggered by a user or system to apply for updating the existing audit template to adapt to the latest policy requirements. It includes core information such as the requesting entity, target industry identifier, and reason for the update.

[0025] Original policy information: refers to unprocessed original data such as policy documents, standards and norms related to auditing work, issued by official or industry authorities, including various formats such as text and tables.

[0026] Standardized corpus data refers to corpus data obtained after processing the original policy information through standardization of format, redundancy removal, and terminology standardization. It has the characteristics of consistent format, concise information, and adaptability to model processing.

[0027] In this embodiment of the invention, after responding to the audit template update request, the original policy information can be obtained from the policy release platform or the enterprise's internal policy database through a preset data interface. The standardization process is to complete the data format unification, redundant information removal and basic terminology standardization to ensure the adaptability of subsequent model processing.

[0028] Further, step 101 may include the following sub-steps: S11. Respond to the audit template update request and obtain the original policy information.

[0029] In this embodiment of the invention, the audit template update request can be initiated by the user through the visual operation interface of the audit management system, or it can be automatically triggered by the system according to the preset policy update monitoring cycle. After receiving the request, the system first verifies the legality of the request signal. The verification includes the authorization authentication of the request initiator and the completeness of the request format (such as whether it contains core fields such as target industry identifier and update requirement description). After the verification is passed, the key instruction information in the request is parsed, and the scope and source channel of the original policy information are determined based on this information. The original policy information can be obtained by calling the preset official open interface for policy release, connecting to the enterprise's internal policy document management database, etc. During the acquisition process, the system will filter the appropriate data source according to the type of policy information release entity (such as industry association, leading enterprise, etc.), and at the same time perform preliminary format recognition and completeness detection on the acquired original policy information to identify whether it is a common format such as text, table, PDF, etc., and detect whether there are problems such as missing information or damaged content. For original policy information with incompatible format or incomplete information, the system will trigger a completion mechanism to obtain complete information by re-calling the interface or sending a supplementary request to the data source management end, ensuring that the acquired original policy information can meet the basic requirements of standardized processing.

[0030] S12. Perform data cleaning on the original policy information to obtain policy text data.

[0031] Data cleaning refers to a series of processing steps performed on raw policy information, including format standardization, noise removal, text normalization, and integrity verification. The core purpose is to remove invalid and interfering information, correct data defects, improve the usability and standardization of policy information, and lay the foundation for subsequent processing steps.

[0032] Policy text data: refers to plain text data with a unified format and accurate information obtained after data cleaning of the original policy information. It retains the core policy content related to the audit work, removes redundant and interfering information, and can be directly used for subsequent key data extraction and other processing.

[0033] In this embodiment of the invention, based on the acquired complete original policy information, the data cleaning process is carried out with the core objective of "retaining audit-related core policy information and eliminating invalid and interfering information." First, a unified format conversion is performed on the different formats of the original policy information (text, tables, PDFs, etc.). A PDF parsing tool extracts the original policy information from PDF format into plain text, and a table-to-text tool converts the policy data in table format into text descriptions row by row, ensuring that original policy information from different sources is unified into a directly processable text format. Then, noise removal processing is initiated, using preset regular expressions to match and delete headers and footers, watermarks, irrelevant advertising pop-up residual text, and special symbols (such as those without specific meaning). The process involves removing placeholders, special separators, and whitespace characters from the cleaned text. Redundant content irrelevant to the audit (such as contact information and irrelevant attachment descriptions in policy announcements) is filtered using a keyword filtering mechanism. Next, text normalization is performed, standardizing line breaks and character encoding (e.g., UTF-8 encoding) and correcting typos and punctuation errors. For incomplete sentence segments, a semantic coherence assessment model is used to complete the expression. Finally, a second integrity check is performed on the cleaned text to ensure no core policy information is missing or key semantics are lost, resulting in a clear, accurate, and uniformly formatted policy text.

[0034] S13. Extract key data from policy text data to obtain core policy elements.

[0035] Key corpus extraction: refers to the process of locating, capturing and integrating core information directly related to the audit work from policy text data. The core objective is to eliminate redundant text, focus on core policy information, and provide a focused data source for subsequent accurate processing.

[0036] Policy core element data: refers to the collection of core policy information obtained after extracting key data, including audit scope, audit standards, compliance basis, etc., which are directly related to the audit template. It has the characteristics of structured initial preparation and information focus.

[0037] In the embodiments of the present invention, based on the obtained policy text data with unified format and accurate information, the key corpus extraction aims to capture the core policy information directly related to the audit work as the core goal. First, the policy text data is preprocessed. The text is disassembled into independent lexical units through a word segmentation tool, and at the same time, a preset stop word list for the policy audit field (including general stop words such as "de", "le", "ji", etc. and domain-specific general vocabulary unrelated to the core policy content such as "notice", "announcement", etc.) is loaded to eliminate redundant words and streamline the text dimension. Subsequently, key information identification is performed on the processed text. This model can accurately locate and extract the core element types in the text by learning the term collocations and semantic association rules in the policy audit field, including key corpus directly related to the adaptation of the audit template such as audit scope, audit criteria, audit process requirements, compliance judgment basis, responsible entity, etc. To further improve the extraction accuracy, an auxiliary screening is synchronously performed in combination with a preset regular expression library for policy core elements. For key corpus with fixed format features such as dates, amounts, percentages, and specific industry terms, accurate capture and supplementation are achieved through regular matching. After the extraction is completed, semantic relevance verification is performed on the initially obtained key corpus, eliminating marginal information that has no direct association with the audit scenario, and at the same time integrating scattered key corpus of the same type (such as summarizing the corpus related to "audit time requirements" mentioned in different paragraphs). Finally, policy core element data with a preliminary degree of structuring and complete core information is output.

[0038] S14. Perform semantic noise reduction and non-essential information screening on the policy core element data to obtain key corpus data.

[0039] Semantic noise reduction: refers to the process of eliminating semantically vague, ambiguous, and repetitive redundant synonymous expressions in the policy core element data through semantic clustering, semantic similarity calculation, etc. The core purpose is to improve the semantic accuracy of information and reduce ineffective semantic interference.

[0040] Non-essential information: refers to the auxiliary content in the policy core element data that has no direct association with the update of the audit template (such as policy background introduction, formulation basis explanation, etc.). Such information does not affect the adaptation design of the audit template and needs to be screened out.

[0041] Key corpus data: refers to the set of core policy information with accurate semantics and focused information obtained after semantic noise reduction and non-essential information screening, only retaining the key content directly related to the update of the audit template, which can be directly used for subsequent non-standardized term conversion.

[0042] In this embodiment of the invention, based on the obtained structured preliminary and information-focused policy core element data, this step aims to further refine audit-related core information and eliminate semantic interference and unnecessary content. First, the policy core element data is semantically analyzed, and semantically similar core element information is categorized and aggregated to facilitate rapid identification of redundant expressions and semantic interference items within the same type of information. Then, semantic noise reduction processing is initiated. Using a pre-set policy audit core semantic dictionary, semantic similarity calculation (such as cosine similarity algorithm) is used to determine the degree of association between each element information and the core semantics, eliminating semantically ambiguous information, ambiguous expressions, and repetitive synonymous expressions (such as...) with a correlation lower than a preset threshold. (Different wording versions of the same audit standard mentioned repeatedly); followed by the screening of unnecessary information. Based on the preset rules for determining unnecessary information, auxiliary information that is not directly related to the update of the audit template is accurately removed, such as background introductions attached to the core elements of the policy, explanations of the basis for policy formulation, and non-audit-related accountability details, while retaining core functional information such as the audit scope, audit standards, and compliance determination basis; to ensure the accuracy of the screening, a semantic integrity review is also performed after the screening is completed to confirm that no key audit-related semantics are lost in the removal process, and that the remaining information can fully support the subsequent standardized processing requirements, and finally output key corpus data that is semantically accurate, information-focused, and free of redundant interference.

[0043] S15. Perform non-standardized terminology conversion on the key corpus data to obtain standardized element data.

[0044] Non-standard terminology conversion: This refers to the process of converting non-standard expressions such as differential expressions, colloquialisms, abbreviations, and non-standard terms in key corpus data into standard terms in the policy audit field through methods such as domain dictionary matching and semantic similarity calculation. The core purpose is to eliminate terminological ambiguity and standardize expression.

[0045] Standardized element data: refers to the set of core policy elements that have been converted from non-standardized terminology and have standardized terminology and expression. It has the characteristics of being able to adapt to the semantic parsing and mapping processing of subsequent models and is the core intermediate data connecting key corpus data and subsequent model processing.

[0046] In this embodiment of the invention, based on the obtained semantically accurate and information-focused key corpus data, this step aims to unify the terminology in the policy audit field, eliminate terminological ambiguity, and adapt to the needs of subsequent model processing and template matching. First, a full-domain scan of the key corpus data is performed to accurately locate non-standardized terms. These terms include differentiated expressions of the same audit concept in different policy documents (such as "accounting verification," "financial review," and "accounting audit reconciliation"), industry slang, abbreviations (such as "internal audit" corresponding to "internal audit"), and non-standard expressions. Then, a pre-set standard terminology dictionary for the policy audit field is invoked. This dictionary covers common standard terms in the audit industry, standardized expressions in various sub-fields, and a terminology mapping table. The semantic similarity between non-standardized terms and standard terms in the dictionary is calculated (e.g., using a cosine similarity algorithm) to match... The system obtains the optimal standard terminology. For non-standard terms with ambiguity, a secondary judgment is made based on their contextual semantics to ensure the accuracy of the conversion (e.g., "special verification" must be clarified in context whether it is "special financial audit verification" or "special compliance audit verification" before the conversion is completed). For new non-standard terms not included in the standard terminology dictionary, the system triggers an unmatched term review mechanism, temporarily storing these terms in a pending confirmation list and pushing them to the backend management terminal. After review by domain experts, these terms are added to the standard terminology dictionary or their corresponding standard expressions are determined. The conversion is then completed retrospectively. After the conversion is completed, all term conversion results are checked for consistency to confirm that the terms in the key corpus data are all uniformly standardized expressions without any ambiguity. Finally, standardized element data with standardized terminology, unified expression, and adaptability to subsequent model processing is output.

[0047] S16. The standardized element data is structured and encapsulated to obtain standardized corpus data.

[0048] Structured encapsulation refers to the process of classifying and decomposing standardized element data into corresponding fields according to preset field specifications, establishing logical relationships between fields, and unifying data formats. The core purpose is to transform scattered standardized element data into structured data that is well-organized and adaptable to model processing.

[0049] Standardized corpus data refers to the set of core policy audit information that has been structured and encapsulated, with a regular structure, clear fields, and uniform format. It is fully compatible with the input requirements of the pre-trained and fine-tuned large language model and is the core input data for the subsequent semantic parsing and mapping process.

[0050] In this embodiment of the invention, based on the standardized element data with consistent terminology and unified expression, this step aims to construct structured data that is formatted, has clear fields, and can directly adapt to the input requirements of pre-trained, fine-tuned large language models. First, based on the model processing needs and template matching logic in the policy audit field, pre-defined structured encapsulation field specifications are established. These encapsulated fields include core dimensions such as audit element type (e.g., audit scope, audit standards, compliance judgment basis), standard terminology, policy effective date, applicable industry sub-identifiers, and core semantic summaries, ensuring that the fields cover the key requirements of subsequent semantic parsing mapping and template matching. Then, the data encapsulation process is initiated, classifying and filling the standardized element data according to the pre-defined fields. For example, "Financial audit scope of manufacturing: fixed asset verification, revenue account reconciliation" is filled into fields such as "Audit element type - audit scope," "Standard terminology," and "Applicable industry sub-identifier - manufacturing." Simultaneously, related labels are added... The system establishes logical connections between different fields using various methods (e.g., binding audit standards and compliance bases from the same policy source with the same association ID); during the encapsulation process, it simultaneously performs format standardization processing, unifying the data of each field to a preset data type (e.g., date fields are uniformly formatted as "YYYY-MM-DD", and identifier fields are uniformly formatted as strings), and performs length standardization and format validation on field content to avoid problems such as missing fields and incorrect formats; for field adaptation deviations that occur during the encapsulation process (e.g., some standardized element data cannot be directly matched with preset fields), the system will trigger a field adaptation correction mechanism, determine the data's belonging field through semantic parsing and complete accurate filling, or supplement custom fields to be compatible with special types of data; after encapsulation, a full-scale structured validation is performed to confirm that all standardized element data has been reasonably encapsulated, the logical connections between fields are accurate, and the format is completely unified, ultimately outputting standardized corpus data with a well-organized structure, clear fields, and format adaptation to the model processing.

[0051] Step 102: Obtain the pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer.

[0052] Pre-trained fine-tuned large language model: refers to a language model that is pre-trained on a general corpus and then fine-tuned and optimized using a small amount of policy audit domain corpus, and has the ability to understand semantics and make associations in the policy audit domain.

[0053] Policy audit semantic parsing and mapping layer: This refers to the functional layer in the pre-trained and fine-tuned large language model that is responsible for completing policy semantic parsing, audit element extraction, and the mapping between the two. It is the core module for generating initial mapping data.

[0054] Mapping Result Calibration and Optimization Layer: This refers to the functional layer in the pre-trained fine-tuning large language model responsible for compliance verification, deviation correction, and generating structured results for the initial mapping data, thereby improving the reliability and standardization of the mapping data.

[0055] In this embodiment of the invention, the pre-trained fine-tuned large language model is a basic model pre-trained based on a large amount of policy audit corpus. Its two functional layers respectively undertake the core tasks of semantic parsing and mapping and result verification and correction, realizing phased and precise control from data processing to result optimization.

[0056] Step 103: Input standardized corpus data into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and output the initial mapping data for policy audit.

[0057] Initial mapping data for policy audit: refers to the raw data output by the semantic parsing mapping layer of policy audit, which contains the preliminary correspondence between policy content and audit elements, without verification and correction processing.

[0058] In this embodiment of the invention, after the standardized corpus data is input, the policy audit semantic parsing and mapping layer first completes the domain adaptation preprocessing of the corpus, then extracts the policy semantics and audit element features and completes the association mapping, and finally outputs the initial mapping data containing the correspondence between policy content and audit elements.

[0059] Furthermore, the policy audit semantic parsing mapping layer includes a corpus preprocessing unit, a domain feature enhancement extraction unit, a policy audit semantic fusion unit, and a mapping inference unit. Step 103 may include the following sub-steps: S21. Standardized corpus data is input into the corpus preprocessing unit, which performs format verification, noise filtering, and policy audit terminology alignment in sequence, and outputs standardized corpus data.

[0060] Corpus Preprocessing Unit: This refers to the functional unit in the policy audit semantic parsing and mapping layer responsible for pre-optimizing the input standardized corpus data. Its core responsibility is to purify and adapt the data through format verification, noise filtering, and term alignment, providing high-quality input for subsequent feature extraction and mapping.

[0061] Format validation: This refers to the process of checking and correcting the field integrity, data type, and encoding format of standardized corpus data according to the model input specifications. The core purpose is to ensure that the corpus format meets the adaptation requirements of subsequent model processing.

[0062] Policy audit terminology alignment: This refers to the process of secondary verification and unified calibration of terms in the corpus based on a standard terminology dictionary in the policy audit field, to strengthen the consistency of terminology expression and avoid terminology deviations from affecting the accuracy of subsequent semantic analysis.

[0063] Standardized corpus data refers to corpus data that has been fully processed by the corpus preprocessing unit, and is formatted correctly, free of interference information, and with consistent terminology. It is the core data source for direct input into the domain feature enhancement extraction unit.

[0064] Model-specific fine-tuning: This refers to the secondary optimization of standardized data to meet the specific needs of fine-tuning large language models, such as input specifications, weight distribution, and semantic understanding habits, ensuring that the data fully matches the model's processing requirements.

[0065] In this embodiment of the invention, based on the obtained standardized corpus data oriented towards the entire process, unlike the global coarse processing in step 101, the core objective of the corpus preprocessing unit is to perform fine-tuning adaptation to the specific input requirements of the large language model. First, a format verification process is initiated. Here, verification no longer unifies the format of the original policy information (e.g., PDF to text), but strictly matches the model input specifications of the policy audit semantic parsing mapping layer. It checks whether the field structure of the standardized corpus data meets the tensor dimension requirements of the model input, whether core fields (e.g., audit element types, standard terminology expressions) are completely consistent with the field identifiers of the model pre-training, and whether special fields such as dates meet the model's numerical input format. If deviations are found, the field names or formats are automatically corrected to adapt to the model. Subsequently, noise filtering is performed. Here, the filtering target is not the headers and footers or irrelevant content in the original policy information, but rather model-sensitive noise that may remain after structured encapsulation, such as redundant separators. Noise such as blank placeholder fields and hidden symbols left over from character encoding does not affect the overall processing, but it can interfere with the semantic parsing accuracy of the model. Next, a policy audit terminology alignment operation is performed. Unlike the basic conversion of "non-standardized terms to standard terms" in step 101, this alignment is based on the domain terminology weight distribution pre-trained by the model to perform a secondary calibration of standardized terms. For example, if the model has a higher weight for "internal audit" than "internal audit", it will still be unified as "internal audit" even if the conversion has been completed in step 101. At the same time, it unifies the terminology expression deviation caused by contextual semantics (such as subtle wording differences of the same standard term in different fields). If new terms not included are found, a review mechanism linked to step 101 is triggered simultaneously. After all processing steps are completed, standardized corpus data that is fully adapted to the subsequent feature extraction requirements of the policy audit semantic parsing mapping layer is output, ensuring that the model input is "precisely adapted" data, rather than just "generally standardized" data.

[0066] S22. Using standardized corpus data as input, the domain feature enhancement extraction unit performs dual feature extraction of policy semantics and audit elements, and outputs dual domain feature vectors.

[0067] Domain Feature Enhancement Extraction Unit: This refers to the functional unit in the policy audit semantic parsing and mapping layer responsible for extracting and enhancing policy and audit features from standardized corpus data. Its core responsibility is to generate dual-domain feature vectors with high discriminative power and strong representational ability through a domain-adaptive extraction mechanism.

[0068] Policy semantic features: These are features extracted from standardized corpora that can represent key information such as the core requirements, applicable scenarios, and constraints of a policy. They are core data representations that reflect the core intent of a policy.

[0069] Audit element characteristics: These refer to the characteristics of core audit elements extracted from standardized corpora, corresponding to audit scope, audit indicators, compliance judgment standards, etc. They are the key foundation for supporting the mapping of subsequent policies and audits.

[0070] Dual-domain feature vector: refers to a feature set composed of policy semantic feature vector and audit element feature vector. The two vectors have a unified dimensional format and can be directly input into the policy audit semantic fusion unit for fusion processing.

[0071] Domain feature enhancement extraction mechanism: refers to a feature extraction optimization scheme that combines domain knowledge graphs, industry standard libraries and attention mechanisms to enhance the domain adaptability, discriminative power and representational ability of dual-domain features, thereby improving feature quality.

[0072] In this embodiment of the invention, based on the standardized corpus data with compliant format and unified terminology obtained in S21, the domain feature enhancement extraction unit aims to accurately capture the core features of the dual domains of policy and audit, and enhance the domain distinguishability and representational ability of the features. First, the standardized corpus data undergoes sentence vector encoding preprocessing. Using a RoBERTa model fine-tuned from a massive corpus of policy and audit data, the text content in the structured fields is transformed into initial semantic vectors. Simultaneously, a unique domain identifier (such as "policy-effective date" or "audit-verification standard") is added to each field to achieve initial differentiation between policy and audit-related content. Subsequently, the dual feature enhancement extraction mechanism is activated. For policy semantic features, the core requirements, applicable scenarios, and constraints of the policy are extracted. Feature weights are optimized by introducing a policy domain knowledge graph. The system strengthens the representation strength of the core policy semantics. For audit element characteristics, it accurately extracts key elements such as audit scope, audit indicators, and compliance judgment thresholds. These elements are then screened and enhanced using an audit industry standard process library to ensure the extracted audit elements are mappable and verifiable. During extraction, an attention mechanism is simultaneously employed to increase the weight of vector dimensions highly correlated with the core features of the two domains in the initial semantic vector, weakening interference from irrelevant vectors. Furthermore, through unified feature dimension processing, the policy semantic feature vector and the audit element feature vector are adjusted to a tensor format of the same dimension, facilitating subsequent fusion processing. After extraction, the dual feature vectors are validated to ensure no features are omitted, domain distinctions are clear, and representational capabilities meet standards. The final output is a dual-domain feature vector containing both policy semantic feature vectors and audit element feature vectors.

[0073] S23. Use dual-domain feature vectors as input to the policy audit semantic fusion unit for feature fusion and normalization processing, and output fused features.

[0074] Policy audit semantic fusion unit: refers to the functional unit in the policy audit semantic parsing and mapping layer responsible for correlation mining, fusion optimization and normalization of dual-domain feature vectors. Its core responsibility is to generate fusion features that have both core information of dual domains and stable representation characteristics.

[0075] Feature fusion refers to the process of integrating policy semantic feature vectors and audit element feature vectors into a unified feature carrier through a series of processes such as bidirectional semantic matching, weighting, dimension splicing, and convolutional aggregation. The core purpose is to mine and retain the inherent correlation information between the two domains.

[0076] Normalization processing: refers to the process of using algorithms such as layer normalization to uniformly map the values ​​of each dimension of the fused feature vector to a preset range. The core purpose is to eliminate differences in numerical scale, ensure the stability of feature representation, and avoid affecting the subsequent model inference effect.

[0077] Fusion features: These are feature carriers that have undergone feature fusion and normalization, possessing core related information from both policy and audit domains, as well as stable numerical characteristics. They are the core inputs that support the mapping reasoning unit in completing accurate correlation mapping.

[0078] Element-level similarity matrix: refers to matrix data that quantifies the correlation strength between policy semantic feature vector and audit element feature vector at each dimension. It is the core basis for determining correlation weights and achieving targeted weighted fusion.

[0079] In this embodiment of the invention, based on the obtained dual-domain feature vectors with clear domain distinctions and satisfactory representation capabilities, the policy audit semantic fusion unit aims to deeply mine the intrinsic relationship between policy semantics and audit elements, and generate fusion features that possess both core information from both domains and stable representational characteristics. First, a bidirectional semantic matching process is initiated between the input policy semantic feature vector and the audit element feature vector. Using a vector matching model optimized with policy audit domain corpus, the element-level similarity matrix between the two vectors is calculated to accurately quantify the correlation strength of features in different dimensions, thereby obtaining targeted correlation weight data. Subsequently, based on this correlation weight data, the policy semantic feature vector and the audit element feature vector are weighted and enhanced respectively, giving higher weights to feature dimensions with high correlation, strengthening the representation of core correlation information between the two domains, and obtaining policy-related features and audit-related features. Finally, the two types of related features are concatenated along the channel dimension. The process involves generating preliminary fusion features containing complete relational information. These features are then compressed and aggregated using a 1×1 convolutional layer to eliminate redundant features and improve feature compactness. To prevent over-fusion from causing the loss of original domain feature information, the aggregated relational features are further weighted and fused with the original dual-domain feature vectors according to preset weights, preserving the core basic information of both domains. Finally, a normalization process is initiated, using a layer normalization algorithm to numerically regularize the features after secondary fusion, mapping the values ​​of each dimension of the feature vector to a preset range. This eliminates numerical scale differences between different feature dimensions, preventing gradient imbalance during subsequent mapping inference. After all processing is complete, the fusion features are verified for relational integrity and representational stability to confirm that they fully carry the core relational information of policies and audits while possessing stable numerical characteristics. The final output is a fusion feature that can be directly input into the mapping inference unit.

[0080] Furthermore, the dual-domain feature vector includes a policy semantic feature vector and an audit element feature vector, and S23 may include the following sub-steps: Bidirectional semantic matching: refers to the process of constructing a bidirectional interaction framework through a domain adaptation model to mine the relationship between elements of policy semantic feature vectors and audit element feature vectors. The core purpose is to accurately quantify the element-level correlation strength of bidirectional quantities.

[0081] Channel dimension concatenation: refers to the operation of stacking and integrating two feature vectors of the same dimension along the channel direction (the direction in which the feature dimensions are parallel) to form a higher-dimensional concatenated feature, which is used to preserve the complete correlation information of the two features.

[0082] 1×1 convolution: refers to a convolution operation with a kernel size of 1×1. Its core function is to compress the dimension of features and aggregate information, optimizing feature representation without changing the spatial location information of features.

[0083] Element-weighted fusion: refers to a fusion method that performs element-wise weighted summation of feature vectors from different sources according to preset weights, so as to retain both core related information and basic feature information.

[0084] Layer normalization: refers to a normalization algorithm that normalizes the mean and variance of each element of the feature vector. Its core purpose is to eliminate the numerical scale differences of each dimension of the feature and improve the stability of the feature representation.

[0085] S231. Perform bidirectional semantic matching between the policy semantic feature vector and the audit element feature vector, calculate the element-level similarity matrix between the vectors, and obtain the association weight data.

[0086] In this embodiment of the invention, a bidirectional semantic matching framework is first constructed by calling a BERT model finely tuned from a massive pairwise corpus in the policy audit field, and the policy semantic feature vectors are then used. With audit element feature vector Input the matching frames respectively (where, (As the feature dimension), after obtaining the context-enhanced vectors of both through model encoding, the element-wise cosine similarity between the vectors is calculated to construct a similarity matrix. The calculation of the element-wise cosine similarity is as follows:

[0087] In the formula, For the policy semantic feature vector, the first The element and the feature vector of the audit element The similarity of elements, For the policy semantic feature vector, the first One element, For the audit element feature vector, the first One element, Let L2 norm represent the vectors; after obtaining the initial similarity matrix, the Softmax function is used to normalize each row of the matrix, converting the similarity values ​​into association weights in the 0-1 interval, ultimately yielding a matrix with dimension L2. Association weight data This enables precise quantification of the correlation strength between elements of a two-vector system.

[0088] S232. Based on the correlation weight data, the policy semantic feature vector and the audit element feature vector are weighted respectively to obtain the policy correlation feature and the audit correlation feature.

[0089] In this embodiment of the invention, the weighted processing adopts an element-wise weighted mode of correlation weight data and feature vector. For the policy semantic feature vector, it is weighted and summed with the row vector of correlation weight data to highlight the policy semantic dimension that is highly correlated with the audit elements. The calculation is as follows:

[0090] In the formula, This is a policy-related feature vector. For associated weight data The transpose of .

[0091] For the feature vector of audit elements, a weighted sum is performed using the column vector of the associated weighted data and the vector itself to strengthen the dimensions of audit elements that are highly correlated with policy semantics. The calculation is shown below:

[0092] In the formula, For audit-related feature vectors; Through the above weighting process, policy-related characteristics and audit-related characteristics focusing on core information in both domains are obtained, with both dimensions maintaining [the desired consistency / preservation]. Consistent with the original feature vector.

[0093] S233. The policy-related features and audit-related features are concatenated along the channel dimension, and the concatenated features are integrated using a 1×1 convolution to obtain the related feature data.

[0094] In this embodiment of the invention, the splicing is first completed by stacking channel dimensions to incorporate policy-related features. Characteristics related to audit Concatenate along the channel dimension (i.e., the parallel direction of the feature dimensions) to obtain a dimension of splicing features Subsequently, a 1×1 convolutional layer was constructed for feature integration, with the kernel parameters set to... The number of output channels is Step length The convolution calculation is as follows:

[0095] In the formula, The feature vector after convolution. It is a 1×1 convolution kernel weight matrix. For convolution bias terms, The ReLU activation function is used; 1×1 convolution is used to compress the dimension of the concatenated features (from...). Down to By aggregating information, eliminating redundant related information, and strengthening the representation of core related information, the final dimension is obtained. Related feature data .

[0096] S234. Perform element-weighted fusion of the associated feature data and the dual-domain feature vector according to preset weights to obtain the initial fused features.

[0097] In this embodiment of the invention, the preset weights are determined jointly through the experience of experts in the field of policy auditing and model training and verification. Let the weight coefficients of the associated feature data be... The weighting coefficient of the dual-domain feature vector (the concatenated vector of policy semantic feature vector and audit element feature vector) is: The concatenation method of the two-neighborhood feature vectors is the same as that of S233, resulting in... At the same time, in order to match the dimensions of associated feature data, for Dimensionality reduction using 1×1 convolution with the same parameters to ,get The weighted fusion calculation is shown below:

[0098] In the formula, For initial fusion features, This is a channel concatenation vector of the two-neighborhood feature vectors. The resulting dual-domain feature vectors are dimensionality-reduced. This fusion method preserves the core correlation information between the two domains while also taking into account the basic representation of the original features, avoiding information loss due to over-fusion. The final result is a feature vector with dimension [missing information]. The initial fusion characteristics.

[0099] S235. Perform layer normalization on the initial fusion features to obtain the fusion features.

[0100] In this embodiment of the invention, a layer normalization algorithm is used to eliminate the numerical scale differences between the dimensions of the initial fused features, ensuring the stability of the feature representation. The layer normalization calculation is as follows:

[0101] In the formula, For the final fusion feature, The mean of each element of the initial fusion feature. The variance of each element of the initial fused feature. For smoothing terms (values) (to avoid a denominator of 0) For scaling parameters, The translation parameters are used; through the above normalization process, the values ​​of each dimension of the initial fused features are uniformly mapped to the interval [-1, 1], eliminating the interference of numerical differences on subsequent mapping inference. The final output dimension is... It possesses the fusion characteristics of having both core related information from two domains and stable representation properties.

[0102] S24. The fusion feature input mapping reasoning unit is used for mapping, and the initial mapping data of policy audit is output.

[0103] Mapping Reasoning Unit: This refers to the functional unit in the policy audit semantic parsing mapping layer that is responsible for completing the association mapping between policies and audit elements based on fusion features. Its core responsibility is to establish a precise association between core policy elements and audit indicators and output the initial mapping results through a domain-adapted mapping model.

[0104] Dimension adaptation processing: refers to the process of adjusting the dimensions of the fused features to the input dimensions required by the mapping inference model. The core purpose is to ensure the compatibility between feature representation and model inference logic, and to ensure the smooth operation of mapping inference.

[0105] Policy-audit association rule base: refers to a pre-built set of rules containing logical constraints on the relationship between policy and audit elements, used to constrain the reasoning process of the mapping model and avoid logically conflicting mapping results.

[0106] Confidence score: refers to the quantitative score output by the mapping model used to characterize the reliability of the relationship between "core policy elements and audit indicators". The value range is [0,1]. The higher the score, the more reliable the relationship.

[0107] Initial policy audit mapping data: refers to the initial data output by the mapping inference unit, which contains the relationship between policy and audit core elements and is accompanied by a confidence level label. It has not been verified or corrected and is the basic input for subsequent calibration and optimization.

[0108] In this embodiment of the invention, based on the obtained fusion features that possess both core correlation information and stable representation characteristics in the policy and audit domains, the mapping inference unit aims to establish a precise correlation mapping between core policy elements and audit indicators and audit processes. First, it performs dimensional adaptation processing on the input fusion features, adjusting them to the input dimensions required by the mapping inference model through a preset feature dimension transformation layer, ensuring that the feature representation matches the model's inference logic. Then, it calls the Transformer mapping model, fine-tuned with pairwise correlation samples from the policy and audit domains. This model, by learning the correlation mapping patterns between massive amounts of policy elements and audit elements, can automatically match the corresponding audit indicator categories based on the bidirectional semantic correlation information contained in the fusion features. The audit implementation outlines key points and compliance criteria, while also outputting confidence scores for each relationship (used to characterize the reliability of the mapping results). During the mapping process, the model incorporates a pre-defined policy audit association rule base to constrain the mapping results and avoid logically conflicting results (such as mapping manufacturing policy elements to financial industry-specific audit indicators). After mapping, the original mapping results output by the model are initially standardized, removing low-confidence relationships with confidence scores below a pre-defined threshold (e.g., 0.6). The effective relationships are then preliminarily organized according to the structure of "policy core elements - audit indicators - association confidence scores," ultimately outputting initial policy audit mapping data that contains accurate relationships between policy and audit core elements, is formatted correctly, and includes confidence level indicators.

[0109] Step 104: Use the initial mapping data from the policy audit as input to the calibration and optimization layer to perform verification and correction, and generate structured mapping correction data.

[0110] Structured mapping correction data: refers to mapping data that has been corrected and optimized by the mapping result calibration layer and then standardized into a preset data structure. It has the characteristics of accurate relationships, standardized format, and can be directly used for template matching.

[0111] In this embodiment of the invention, the mapping result calibration and optimization layer will perform validity verification on the initial mapping data according to the preset policy audit compliance rules, correct the mapping relationship with deviation or incompleteness, and at the same time, organize the corrected result into a preset data structure to form structured mapping correction data to adapt to subsequent template matching requirements.

[0112] Furthermore, step 104 may include the following sub-steps: S31. Perform compliance verification on the initial mapping data of policy audit to obtain compliant mapping data and abnormal mapping data.

[0113] Compliance verification: refers to the process of comprehensively verifying the industry adaptability, indicator validity, correlation logic, and confidence level of the initial mapping data of policy audits based on the policy audit compliance rule library. The core purpose is to screen out mapping relationships that meet compliance requirements and identify and mark abnormal mapping relationships.

[0114] Compliance mapping data refers to a set of mapping relationships that, after compliance verification, simultaneously meet the requirements of confidence level, industry adaptability, effective audit indicators, and consistent policy-audit correlation logic, and possess the characteristics of meeting policy and audit compliance requirements.

[0115] Abnormal mapping data: refers to a set of mapping relationships that, after compliance verification, have issues such as insufficient confidence, industry mismatch, invalid audit indicators, or conflicting related logic. Specific abnormality type labels must be attached to support subsequent corrections.

[0116] In this embodiment of the invention, based on the output policy audit initial mapping data with attached confidence level labels, this step conducts verification based on compliance requirements, industry standards, and related logic in the policy audit field. First, a pre-set policy audit compliance rule library is loaded. This rule library covers core verification dimensions such as industry adaptation rules, audit indicator validity rules, policy-audit related logic constraint rules, and confidence level threshold standards. Then, the initial mapping data is structured and parsed to extract key information such as the core policy elements, corresponding audit indicators, applicable industry identifiers, related confidence levels, and policy effective time from each mapping relationship. The verification process first performs preliminary screening based on the confidence level threshold. Mapping relationships with confidence levels lower than the pre-set compliance threshold (e.g., 0.7, higher than the preliminary screening threshold in S24, achieving layered verification) are directly marked as candidate abnormal data. Then, deep compliance verification is performed on mapping relationships that meet the confidence level standard. First, the matching between audit indicators and applicable industry identifiers is verified through industry adaptation rules (e.g., prohibiting the inclusion of certain items in the audit indicator list). The process involves mapping manufacturing-specific audit indicators to financial industry policy elements, then verifying whether the corresponding audit indicators are currently valid standards based on audit indicator validity rules (excluding obsolete or outdated indicators), and finally verifying whether the requirements of the core policy elements are consistent with the audit indicator verification direction through association logic constraint rules (e.g., when the policy element is "supervision of special fund use," the corresponding audit indicators must focus on relevant dimensions such as fund flow and compliance of use; otherwise, it is judged as a logical conflict). After verification, mapping relationships that simultaneously meet the requirements of confidence level, industry fit, indicator validity, and logical consistency are classified as compliant mapping data. Mapping relationships that do not meet the requirements of confidence level, have industry mismatch, indicator failure, or logical conflict are classified as abnormal mapping data. Each abnormal mapping data is labeled with a specific abnormality type (e.g., "industry fit error," "audit indicator has been obsolete," "association logic conflict") to facilitate accurate processing in subsequent correction stages. The final output is compliant mapping data and abnormal mapping data with clear classification and abnormal labeling.

[0117] S32. Correct the abnormal mapping data to obtain corrected mapping data.

[0118] In this embodiment of the invention, based on anomaly mapping data labeled with specific anomaly types, the correction process aims at "accurately matching the causes of anomalies and achieving differentiated and precise correction based on a rule system." First, it loads a pre-set anomaly correction rule library, the latest policy audit industry indicator library, a high-confidence mapping template library, and a semantic association dictionary for the policy audit field. Then, it categorizes and organizes the anomaly mapping data according to the labeled anomaly types and executes targeted correction strategies. For anomaly data with "insufficient confidence," it extracts the key semantics of the corresponding policy core elements (such as policy requirements, applicable scenarios, and constraints) and the core attributes of the original audit indicators, and supplements the high correlation between the two using the domain semantic association dictionary. Semantic features are then used to match high-confidence mapping templates with high-quality mapping cases of similar policy elements. Multi-dimensional semantic similarity calculations (policy requirement matching, applicable scenario fit, and verification dimension consistency) are used to select the optimal matching audit indicators. The mapping relationship is then reconstructed and the confidence level is calculated to ensure the corrected confidence level meets the standards. For abnormal data with "industry mismatch," based on the applicable industry identifier in the abnormal data, a list of specific and valid audit indicators for that industry is accurately retrieved from the policy audit industry indicator library. Combined with industry indicator matching rules in the abnormality correction rule library, semantic matching is used to identify the indicators that best fit the core policy elements. After replacing the original mismatched indicators, the industry fit of the new mapping relationship is verified. For abnormal data indicating "audit indicators have been abolished," current valid indicators with the same function and verification dimensions as the abolished indicators are retrieved from the policy audit industry indicator library. The timeliness of the indicators is confirmed by considering the policy's effective date (ensuring the effective date covers the policy implementation cycle). After replacing the abolished indicators with compliant current indicators, the relationship between the policy and the indicators is re-verified. For abnormal data indicating "conflicting relationship logic," the logical constraint details in the abnormality correction rule library are invoked to break down the core requirements of the policy's core elements and the verification direction of the original audit indicators, clarifying the focus of the conflict. Indicators whose verification direction perfectly matches the policy requirements are then selected from the policy audit industry indicator library. If multiple indicators exist... When selecting indicators, the optimal solution is determined by sorting them according to the priority of the domain indicators preset in the anomaly correction rule base (core indicators > general indicators > auxiliary indicators). After all anomaly data is corrected, the corrected mapping data is uniformly included in the secondary compliance verification process. The correction effect is checked using compliance verification standards that are completely consistent with S31 (confidence level, industry adaptability, indicator validity, and logical consistency). For a small number of stubborn anomaly data that still do not meet the standards, they are directly pushed to the back-end management terminal for manual review and confirmation by domain experts. The corrected data that meets the standards is classified as compliant mapping data. Finally, the accurate correction and effective integration of anomaly mapping data are achieved, providing complete and compliant basic data for the subsequent generation of structured mapping correction data.

[0119] S33. Integrate and structure the compliance mapping data and the corrected mapping data to obtain structured mapping corrected data.

[0120] Integration and structuring: This refers to the comprehensive processing of compliant mapping data and corrected mapping data, including deduplication, cleaning, format unification, field filling, and logical association binding, ultimately transforming them into standardized structured data. The core purpose is to achieve effective integration and standardized output of the two types of data.

[0121] Structured mapping correction data: refers to the set of mapping data that has been integrated and structured, with complete fields, uniform format, clear logic, and traceability identifiers. It is the final output data of the policy audit semantic parsing mapping layer and can directly support the subsequent audit template adaptation and implementation application.

[0122] In this embodiment of the invention, based on the compliant mapping data obtained through screening and the corrected mapping data completed in S32, this step aims to achieve unified data format, clear logical association, and complete field specifications. First, both types of data are loaded and pre-processing data cleaning is performed. Duplicate mapping relationships are eliminated through field comparison (e.g., compliant substitutes corresponding to original abnormal data already covered in the corrected mapping data need to be deleted). Simultaneously, the basic formats of the two types of data are unified (e.g., policy effective time is unified to "YYYY-MM-DD" format, confidence scores are retained to two decimal places, and applicable industry identifiers use a unified industry coding standard). Then, based on the subsequent model processing requirements and template matching logic in the policy audit field, structured field specifications are preset, specifying core fields including policy core element ID, policy core element content, corresponding audit indicator ID, audit indicator name, audit indicator type (core / general / auxiliary), applicable industry code and name, associated confidence score, data type identifier (compliant / corrected), corrected abnormal type (only corrected data is filled in, compliant data is marked "none"), policy effective time, audit indicator validity period, and associated logic description, ensuring that the fields cover the data. The process encompasses the entire workflow, including traceability, template matching, and audit implementation. Next, the integration and field filling process is initiated, breaking down and filling compliance mapping data and corrected mapping data according to preset fields. For corrected mapping data, additional traceability fields such as "corrected anomaly type" and "correction basis (e.g., matched template case ID, semantic similarity score)" are added. For compliance mapping data, the "data type identifier" is uniformly labeled "compliant," and the "corrected anomaly type" is labeled "none." Simultaneously, logical connections between different mapping relationships under the same policy source are established by adding associated IDs (e.g., mapping relationships corresponding to multiple core elements in the same policy document are bound to the same policy source ID). After filling, a full-scale structured validation is performed to verify field completeness (ensuring no records with missing core fields), logical consistency (e.g., policy effective time must be within the validity period of audit indicators), and coding standardization (e.g., industry codes and names must correspond one-to-one). For data with missing fields or logical deviations discovered during validation, a lightweight correction mechanism is triggered (supplementing missing fields and correcting coding errors). Finally, structured mapping and corrected data with standardized fields, clear logic, traceability, and adaptability to subsequent model processing and audit template matching requirements are output.

[0123] Step 105: Match target industry templates based on structured mapping correction data.

[0124] Target industry template: refers to an industry-specific audit template selected from a set of preset industry templates that is compatible with the structured mapping correction data and meets the assessment criteria.

[0125] In this embodiment of the invention, the matching process uses industry tags and core audit elements in the structured mapping correction data as key bases, retrieves and filters suitable candidate templates in the preset industry template set, and determines the target industry template through semantic comparison of core elements.

[0126] Furthermore, step 105 may include the following sub-steps: Template matching search keywords: refers to the core information set extracted from structured mapping correction data for searching industry templates. It uses "applicable industry tags + audit elements" as the core combination form and has the function of accurately locating industry-suitable templates.

[0127] Preset industry template set: refers to a collection of audit templates for different policy scenarios in various industries, stored hierarchically according to the industry classification system. Each template is accompanied by industry tags, audit indicator list and other search identifiers, which are the core data foundation for template retrieval.

[0128] Candidate industry templates: These refer to the set of audit templates that are consistent with the applicable industry tags and have potential matching value after being screened by S42 search. They are the objects of subsequent semantic comparison and selection of the optimal template.

[0129] Semantic comparison: refers to the process of comparing the semantic relevance between audit elements and candidate template audit indicators from multiple dimensions, such as audit indicator names, verification directions, and applicable scenarios, by combining a domain semantic association dictionary.

[0130] Target industry template: This refers to the audit template that best matches the audit elements and meets the current policy scenario requirements after similarity ranking and compatibility verification. It serves as the core reference for subsequent audit implementation.

[0131] S41. Extract applicable industry tags and audit elements from the structured mapping correction data and use them as template matching search keywords.

[0132] In this embodiment of the invention, based on the obtained structured mapping correction data with standardized fields and clear logic, this step aims to extract accurate and valuable core information. First, key fields such as "applicable industry code and name," "audit indicator ID," "audit indicator name," and "audit indicator type" are located in the structured data. Applicable industry tags are then extracted (preferably using a unified industry code and industry name combination to ensure uniqueness, such as "01-Manufacturing"). Simultaneously, a set of audit elements is extracted, covering core information such as audit indicator name, audit verification direction, and audit indicator type (core / general / auxiliary). During the extraction process, information is standardized, and synonymous expressions in audit elements are unified according to preset standards (e.g., "accounting verification" is standardized to "financial audit verification"). Duplicate or redundant element information is removed (such as multiple duplicate records of the same audit indicator). The extracted applicable industry tags are then linked and bound to the audit elements to form a combined search keyword set of "industry tags + multi-dimensional audit elements." This ensures that the keywords cover industry attributes and accurately match core audit needs, providing accurate search basis for subsequent template searches.

[0133] S42. Use search keywords to search the preset industry template set and filter out multiple candidate industry templates that are consistent with the applicable industry.

[0134] In this embodiment of the invention, based on the generated combined search keywords, this step aims to quickly identify the range of industry-adapted templates. First, the organizational logic of the preset industry template set is clarified—this template set is stored hierarchically according to an industry classification system, with each industry category containing multiple audit templates adapted to different policy scenarios. Each template is accompanied by industry tags, a list of core audit indicators, and search identifiers such as applicable policy scenarios. During the search, the "applicable industry tags" are used as the core search condition to perform precise matching within the preset industry template set, quickly filtering out all templates whose industry tags match the search keywords, and initially excluding templates with industry mismatches. Subsequently, a lightweight verification is performed on the initial screening results to check the validity of the templates (removing updated or obsolete old templates) and the compatibility of the applicable policy scenarios with the core policy elements in the structured mapping correction data (e.g., excluding templates that are only applicable to specific policies and do not fit the current policy scenario). Finally, multiple candidate industry templates that meet the industry adaptation requirements and have potential matching value are obtained.

[0135] S43. Perform semantic comparison between the audit elements and the audit indicators of each candidate industry template, and calculate the matching similarity.

[0136] In this embodiment of the invention, based on the audit elements extracted in S41 and the candidate industry templates screened in S42, this step aims to quantify the degree of fit between audit elements and template indicators. First, a semantic association dictionary for the policy audit domain is loaded to clarify the synonyms, associations, and applicable scenarios of audit terms, providing a benchmark for semantic comparison. Then, for each candidate industry template, its built-in core audit indicator list and corresponding verification direction, applicable scenarios, and other information are extracted and compared with the audit elements extracted in S41 in a multi-dimensional semantic comparison. The comparison dimensions include the matching degree of audit indicator names, the consistency of verification direction, and the fit of applicable scenarios. The semantic similarity calculation adopts the cosine similarity algorithm. Combined with the domain semantic association dictionary, the audit elements and template indicators are vector-encoded. The single-dimensional similarity score is obtained by calculating the cosine value between the vectors. Then, the comprehensive matching similarity is calculated according to preset weights (such as the weight of indicator name matching degree 0.4, the weight of verification direction consistency 0.3, and the weight of applicable scenario fit 0.3). The comprehensive score range is [0,1]. The higher the score, the higher the degree of fit. Finally, a unique comprehensive matching similarity score is generated for each candidate industry template.

[0137] S44. Based on the matching similarity ranking, select the candidate industry template with the highest similarity as the target industry template.

[0138] In this embodiment of the invention, based on the calculated comprehensive matching similarity of each candidate industry template, this step aims to determine the optimal suitable template. First, all candidate industry templates are sorted in descending order of comprehensive matching similarity score, and the template with the highest score is selected as the initial target template. If multiple candidate templates have the same score and are all the highest values, auxiliary screening conditions are further introduced, including template update time (the most recently updated template is selected), expert rating (the template with a higher rating after review by domain experts is selected), and historical matching success rate (the template with a higher matching accuracy in past applications is selected). The unique optimal template is selected through auxiliary conditions. After selection, a final adaptability verification is required to check whether the template's audit indicator system fully covers the core audit elements extracted in S41, and whether the template's applicable scenarios conflict with the current policy core requirements. After confirmation, the template is determined as the target industry template. If the verification finds an adaptability deviation, the second highest-ranked candidate template is selected for re-verification until a target industry template that meets the requirements is determined.

[0139] Step 106: When the target industry template does not meet the preset template evaluation conditions, the fine-tuning large language model is fine-tuned, and the process jumps to the step of using standardized corpus data input policy audit semantic parsing and mapping layer to perform semantic parsing and mapping until the target industry template meets the preset template evaluation conditions.

[0140] Preset template evaluation conditions: These are pre-defined standards used to determine whether a target industry template meets policy adaptation requirements. The core is that the template audit indicators meet the coverage ratio of audit elements in the structured mapping correction data.

[0141] Adaptation deviation information: refers to the list of audit elements not covered by the target industry template and the corresponding core policy elements, as well as the preliminary analysis results of the reasons for the lack of coverage, which is the core basis for determining the direction of model fine-tuning.

[0142] Model fine-tuning direction: refers to the key points of model optimization based on the adaptation deviation information, focusing on correcting the deficiencies of the original model in feature extraction, association mapping and other aspects, so as to ensure that the optimized model can generate structured mapping correction data that better fits the template adaptation requirements.

[0143] A small number of domain samples: This refers to paired labeled data of policy and audit elements that cover the deviation points of the specific coverage. The number of samples does not need to be large (usually a few hundred), but they need to be highly accurate and representative for precise iterative optimization of model parameters.

[0144] Parameter iterative optimization: refers to the process of gradually adjusting the existing parameters of the model based on a small number of domain samples using algorithms such as mini-batch gradient descent. The core goal is to improve the model's performance at deviation points while avoiding destroying the model's original generalization ability.

[0145] Iterative loop: refers to the repeated process of "template evaluation - model fine-tuning - re-semantic parsing and mapping - template evaluation again", which gradually improves the model performance and template adaptability through multiple rounds of iteration until the preset evaluation conditions are met.

[0146] In this embodiment of the invention, based on the target industry template determined in S44 and the structured mapping correction data output in S33, this step aims to achieve accurate template adaptation through model iterative optimization. First, a verification process for the preset template evaluation conditions is executed to clarify the preset coverage ratio threshold (e.g., the preset threshold is 80%, which can be dynamically adjusted according to industry characteristics). Then, the actual coverage ratio of the audit indicators of the target industry template to the audit elements in the structured mapping correction data is calculated. First, the total set of audit elements after deduplication in the structured mapping correction data is extracted (denoted as set A, with N elements). Then, the set of audit indicators of the target industry template is extracted (denoted as set B). The number of elements in the intersection of the two sets is counted (denoted as M). The actual coverage ratio is calculated as "M / N×100%". If the calculation result is lower than the preset coverage ratio threshold, it is determined that the preset template evaluation conditions are not met, and the model fine-tuning process needs to be initiated. The model fine-tuning phase first determines the direction of fine-tuning based on adaptation bias information. This information specifically comprises a list of uncovered audit elements (the difference between set A and set B) and their corresponding core policy elements. The root causes of these biases are analyzed (e.g., insufficient feature extraction of these audit elements by the original model, logical deviations in the policy-audit association mapping, etc.). This clarifies the optimization directions to be focused on during fine-tuning (e.g., strengthening feature representation learning of uncovered audit elements, optimizing the association mapping rules between specific industry policies and audit elements). Next, a small number of domain-specific fine-tuning samples are prepared. These samples are sourced from paired-labeled data of policy-audit elements corresponding to uncovered audit elements and high-quality mapping cases supplemented by domain experts. The samples must contain complete information such as core policy elements, precisely matched audit elements, and applicable industry tags to ensure targeted coverage of adaptation bias points. Finally, a small-batch, stepped adjustment method is used. The degree descent method is used to iteratively optimize the parameters of the pre-trained fine-tuned large language model. A low learning rate (e.g., 1e-5) is set to avoid model overfitting. During the iteration process, the "mapping accuracy of uncovered audit elements" is used as the core monitoring indicator. Fine-tuning stops when the indicator tends to stabilize and meets the preset requirements. After fine-tuning, the process jumps to the step of "using standardized corpus data to input policy audit semantic parsing and mapping layer for semantic parsing and mapping" (i.e., regressing S21 and subsequent processes). The structured mapping correction data is regenerated, the target industry template is matched, and the coverage ratio is checked again. The above iterative process of "template evaluation - model fine-tuning - reprocessing" is repeated until the coverage ratio of the audit indicators of the target industry template to the audit elements in the structured mapping correction data is not lower than the preset coverage ratio threshold, that is, the preset template evaluation conditions are met, and the iterative process stops.

[0147] Furthermore, the preset template evaluation conditions specifically stipulate that the coverage ratio of the audit indicators of the target industry template to the audit elements in the structured mapping correction data is not lower than the preset coverage ratio threshold.

[0148] In this embodiment of the invention, the preset coverage ratio threshold is determined by combining the experience of experts in the policy audit field with industry application scenarios. Different industries can set differentiated thresholds (e.g., the financial industry has strict regulatory requirements, so the preset threshold is set to 90%; the general manufacturing industry has a preset threshold of 80%), and the threshold can be dynamically calibrated according to the actual application effect. The specific calculation process for the coverage ratio can be broken down as follows: First, deduplicate the audit elements in the structured mapping correction data, removing duplicate records to form a unique set of audit elements, and labeling the core attributes of each element (such as element type and verification priority); Second, extract the complete list of audit indicators built into the target industry template, clarifying the coverage scope and applicable scenarios of each indicator; Third, use semantic similarity matching combined with rule verification to determine the matching relationship between audit elements and template indicators—if the semantic similarity between the two is ≥ the preset matching threshold (e.g., 0.85) and the applicable scenarios are consistent, it is determined to be "covered"; otherwise, it is "not covered"; Fourth, count the number of covered audit elements (the number of successfully matched elements), divide it by the total number of deduplicated audit elements, and obtain the actual coverage ratio; Fifth, compare the actual coverage ratio with the preset coverage ratio threshold. If the actual ratio is ≥ the threshold, the target industry template is determined to meet the preset evaluation conditions and can be directly used for subsequent audit implementation; if the actual ratio is < the threshold, it is determined that the conditions are not met, triggering the subsequent model fine-tuning and iterative process. This assessment method ensures the template's suitability for audit needs by quantifying the coverage ratio, avoiding omissions of key points in audit implementation due to missing template indicators.

[0149] Step 107: When the target industry template meets the preset template evaluation conditions, the target industry template will be fed back to the users in the corresponding industry.

[0150] In this embodiment of the invention, after the template meets the evaluation conditions, the target industry template is fed back to the corresponding industry user through system message push or user-specific interface, while template update records and model processing logs are retained for subsequent traceability and management.

[0151] Please see Figure 2 , Figure 2 This is a structural block diagram of an automatic template update system provided in an embodiment of the present invention.

[0152] This invention provides an automatic template update system, comprising: The policy collection and parsing module 201 is used to respond to audit template update requests, obtain raw policy information and standardize it to obtain standardized corpus data; The model acquisition module 202 is used to acquire a pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer. The policy semantic understanding and impact analysis module 203 is used to input standardized corpus data into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and output the initial mapping data of policy audit. The verification and correction module 204 is used to verify and correct the mapping result calibration optimization layer by using the initial mapping data input from the policy audit, and to generate structured mapping correction data. Template generation and configuration module 205 is used to match target industry templates based on structured mapping correction data; The template push and feedback module 206 is used to fine-tune the large language model when the target industry template does not meet the preset template evaluation conditions, and then jump to execute the steps of semantic parsing and mapping using the standardized corpus data input policy audit semantic parsing mapping layer until the target industry template meets the preset template evaluation conditions.

[0153] Since the above is a system corresponding to one of the automatic template update methods, its implementation principle is the same as that of an automatic template update method. For the sake of convenience and brevity, those skilled in the art can clearly understand that the specific working process of the system and modules described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here.

[0154] Please see Figure 3 , Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of the present invention.

[0155] An electronic device according to an embodiment of the present invention includes: a memory 301 and a processor 302. The memory 301 stores a computer program. When the computer program is executed by the processor 302, the processor 302 executes the template automatic update method as described in the above embodiment.

[0156] Memory 301 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 301 has storage space 303 for program code 313 for performing any of the method steps described above. For example, storage space 303 for program code may include various program codes 313 for implementing the various steps in the methods described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When this code is run by a computing device, it causes the computing device to perform the various steps in the template auto-update method described above.

[0157] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the template automatic update method as described in the above embodiments.

[0158] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0159] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatically updating templates, characterized in that, include: In response to the audit template update request, the original policy information is obtained and standardized to obtain standardized corpus data; Obtain a pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer; The standardized corpus data is input into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and the initial policy audit mapping data is output. The initial mapping data from the policy audit is used as input to the calibration and optimization layer of the mapping result for verification and correction, generating structured mapping correction data; Based on the structured mapping, correct the data to match the target industry template; When the target industry template does not meet the preset template evaluation conditions, the fine-tuned large language model is fine-tuned, and the process jumps to the step of using the standardized corpus data to input the policy audit semantic parsing and mapping layer for semantic parsing and mapping, until the target industry template meets the preset template evaluation conditions.

2. The template automatic update method according to claim 1, characterized in that, The response to the audit template update request obtains the original policy information and standardizes it to obtain standardized corpus data, including: Respond to audit template update requests and obtain original policy information; The original policy information is cleaned to obtain policy text data; Key corpus extraction is performed on the policy text data to obtain core policy element data; Semantic noise reduction and unnecessary information removal are performed on the core policy element data to obtain key corpus data; The key corpus data is converted from non-standardized terminology to obtain standardized element data. The standardized element data is structured and encapsulated to obtain standardized corpus data.

3. The template automatic update method according to claim 1, characterized in that, The policy audit semantic parsing and mapping layer includes a corpus preprocessing unit, a domain feature enhancement and extraction unit, a policy audit semantic fusion unit, and a mapping inference unit. The standardized corpus data is input into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, outputting initial policy audit mapping data, including: The standardized corpus data is input into the corpus preprocessing unit, where format verification, noise filtering, and policy audit terminology alignment are performed sequentially, and standardized corpus data is output. The standardized corpus data is input into the domain feature enhancement extraction unit to perform dual feature extraction of policy semantics and audit elements, and output dual domain feature vectors; The dual-domain feature vectors are input into the policy audit semantic fusion unit for feature fusion and normalization processing, and the fused features are output. The fusion features are input into the mapping inference unit for mapping, and the initial mapping data for policy audit is output.

4. The template automatic update method according to claim 3, characterized in that, The dual-domain feature vector includes a policy semantic feature vector and an audit element feature vector. The process of inputting the dual-domain feature vector into the policy audit semantic fusion unit for feature fusion and normalization, and outputting fused features, includes: Bidirectional semantic matching is performed between the policy semantic feature vector and the audit element feature vector, and the element-level similarity matrix between the vectors is calculated to obtain the association weight data. Based on the aforementioned correlation weight data, the policy semantic feature vector and the audit element feature vector are weighted respectively to obtain policy correlation features and audit correlation features; The policy-related features and the audit-related features are concatenated along the channel dimension, and the concatenated features are integrated using a 1×1 convolution to obtain the related feature data. The associated feature data and the dual-domain feature vector are element-weighted and fused according to preset weights to obtain initial fused features; The initial fusion features are subjected to layer normalization to obtain the fusion features.

5. The template automatic update method according to claim 1, characterized in that, The process of using the initial mapping data from the policy audit as input to the calibration and optimization layer for verification and correction, generating structured mapping correction data, includes: The initial policy audit mapping data is subjected to compliance verification to obtain compliant mapping data and abnormal mapping data; The abnormal mapping data is corrected to obtain corrected mapping data; The compliance mapping data and the corrected mapping data are integrated and structured to obtain structured mapping corrected data.

6. The template automatic update method according to any one of claims 1-5, characterized in that, The matching of target industry templates based on the structured mapping correction data includes: Applicable industry tags and audit elements are extracted from the structured mapping correction data and used as template matching search keywords; The preset industry template set is searched using the search keywords, and multiple candidate industry templates that match the applicable industry are selected. The audit elements are semantically compared with the audit indicators of each candidate industry template, and the matching similarity is calculated. Based on the matching similarity ranking, the candidate industry template with the highest similarity is selected as the target industry template.

7. The template automatic update method according to claim 1, characterized in that, The preset template evaluation condition specifically means that the coverage ratio of the audit indicators of the target industry template to the audit elements in the structured mapping correction data is not less than a preset coverage ratio threshold.

8. A template automatic update system, characterized in that, include: The policy collection and parsing module is used to respond to audit template update requests, obtain raw policy information and standardize it to obtain standardized corpus data. The model acquisition module is used to acquire a pre-trained fine-tuned large language model, which includes a policy audit semantic parsing mapping layer and a mapping result calibration and optimization layer. The policy semantic understanding and impact analysis module is used to input the standardized corpus data into the policy audit semantic parsing and mapping layer for semantic parsing and mapping, and output the initial policy audit mapping data. The verification and correction module is used to verify and correct the mapping result by inputting the initial mapping data of the policy audit into the calibration and optimization layer, and generate structured mapping correction data. The template generation and configuration module is used to match target industry templates based on the structured mapping correction data; The template push and feedback module is used to fine-tune the fine-tuning large language model when the target industry template does not meet the preset template evaluation conditions, and then jump to the step of using the standardized corpus data to input the policy audit semantic parsing and mapping layer for semantic parsing and mapping, until the target industry template meets the preset template evaluation conditions.

9. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the template automatic update method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the automatic template update method as described in any one of claims 1-7.