A training method for a controllable and trustworthy official document generation model
The method addresses issues of dynamic constraints and adversarial sensitivity in public document generation by integrating a multi-source corpus and robust training, ensuring legal compliance and coherence, and enhancing model resilience through user feedback.
Patent Information
- Application Number
- CN202510474631.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing official document generation model has shortcomings in controllability, credibility and robustness, making it difficult to achieve real-time control of format elements and legal compliance, and lacks an effective closed-loop mechanism for user feedback, resulting in format errors, logical contradictions and semantic deviations in generated content.
Build a multi-source and multi-type corpus, adopt a pre-trained language model with BART-large architecture, integrates a controllability gating mechanism and a credibility evaluation network, and trains a basic controllable model, trustworthy enhancement model and robust enhancement model in parallel. Through format verification, legal compliance detection and logical contradiction detection, and optimizes model performance with user feedback closed-loop mechanism.
It significantly improves the legal compliance and semantic integrity of official document generation, enhances the adaptability and stability of the model, can effectively resist adversarial input, and ensures the credibility and continuous optimization of the generated content.
Smart Images

Figure CN119988648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a training method for a controllable and trustworthy official document generation model. Background Art
[0002] With the rapid development of natural language processing technology, text generation models based on deep learning have shown broad application prospects in the field of official document processing. As an important carrier for government agencies, enterprises and institutions to carry out official activities, the generation process of official documents needs to strictly follow format specifications, legal provisions and semantic logic. Traditional official document generation methods rely on template filling or rule engines, which can meet the basic format requirements, but have obvious limitations in terms of flexibility, semantic coherence and legal compliance. In recent years, pre-trained language models based on the Transformer architecture have significantly improved the naturalness and accuracy of text generation through large-scale corpus training, providing technical support for intelligent official document generation. Existing research attempts to apply these models to official document generation tasks and optimize the generation quality through domain adaptation and constraint mechanisms, but still faces multiple challenges in terms of controllability, trustworthiness and robustness.
[0003] Currently, mainstream official document generation models generally have three technical bottlenecks: First, the generation process lacks a dynamic constraint mechanism, making it difficult to control format elements and legal compliance in real time, resulting in problems such as format errors or invalid policy references in the generated content; second, the model is sensitive to adversarial inputs and is prone to logical contradictions or semantic deviations when the input contains noise or perturbations; third, the existing evaluation system focuses on language fluency, lacks quantitative evaluations of legal compliance and content consistency, and has not established an effective user feedback closed-loop mechanism, making it difficult to continuously optimize the model performance. In addition, key issues such as the fusion processing of multi-source heterogeneous corpora, the real-time verification of legal provisions, and the credibility traceability of generation results have not been effectively solved. In response to this, we propose a training method for a controllable and trustworthy official document generation model. Summary of the Invention
[0004] To solve the above technical problems, a training method for a controllable and trustworthy official document generation model is provided, and the technical solution solves the above problems.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A training method for a controllable and trustworthy official document generation model includes the following steps:
[0007] S1. Collect original official document data based on government agencies, public databases and legal document libraries, and establish a multi-source and multi-type corpus. The multi-source and multi-type corpus includes: requests for instructions, meeting minutes, notices and announcements, and policy documents;
[0008] S2. Dedup and standardize the original corpus data collected;
[0009] S3. Establish an evaluation system consisting of a format validator, a legal compliance detection module, and a semantic integrity analyzer to evaluate the data quality, and based on the evaluated data, divide the data into a training set and a test set;
[0010] S4. Build a pre-trained language model based on the BART-large architecture, integrating a controllability gating mechanism and a credibility evaluation network;
[0011] S5. Parallel train the basic controllable model, the credibility enhancement model, and the robustness enhancement model. Among them, the credibility enhancement model is connected to the laws and regulations database in real time for verification, and the robustness enhancement model injects adversarial noise. Dynamically monitor the controllable deviation and legal compliance of the generated content through the test set, and trigger an early stopping mechanism to prevent overfitting;
[0012] S6. Perform weighted parameter fusion on the performance of the three types of models after training to generate the final deployment model, and automatically trigger iterative optimization through the user feedback closed-loop mechanism;
[0013] The credibility evaluation network includes:
[0014] The legal clause checker verifies the timeliness of the cited clauses through the API of the national laws and regulations database, and automatically replaces the expired clauses with the latest version;
[0015] The logical contradiction detector constructs a time-event relationship graph to identify temporal conflicts in the generated content;
[0016] The source credibility scoring module calculates:
[0017] ;
[0018] In the formula, is the comprehensive credibility score, is the legal compliance score, is the content consistency score, is the weighting coefficient;
[0019] Among them, the calculation formula of the legal compliance score is:
[0020] ;
[0021] In the formula, is the legal compliance score, is the total number of generated clauses, is the th generated clause, is the legal knowledge base set, is the indicator function.
[0022] Preferably, the duplicate removal and standardization processes in step S2 specifically include:
[0023] Among them, the duplicate removal process specifically includes:
[0024] Generate a 64-bit document fingerprint using an improved SimHash algorithm, and define the calculation of document similarity as:
[0025] ;
[0026] In the formula, is the similarity between document and document , is the Hamming distance between document fingerprint and ;
[0027] When the similarity is greater than the preset threshold, it is determined as an approximately duplicate document, and the duplicate document is removed;
[0028] Among them, in the standardization process, forced conversion of the execution date format, conversion of the amount number to Chinese capitalization, and mapping of the full name of the institution are performed.
[0029] Preferably, the working process of the evaluation system is as follows:
[0030] The format validator matches the document structure elements through regular expressions, including the combination rules of the agency code, year, and serial number of the document number;
[0031] The legal compliance detection module matches the policy terms based on the knowledge graph and outputs a compliance score;
[0032] The semantic integrity analyzer uses the RoBERTa model to calculate the context coherence score.
[0033] Preferably, the controllable gating mechanism includes:
[0034] The rule constraint layer deploys a Drools rule engine containing document writing specifications to intercept in real time the missing agency code of the document number and the incorrect use of document types;
[0035] The dynamic gating unit calculates the gating weight in the decoder:
[0036] ;
[0037] In the formula, is the gating weight of the th step, is the Sigmoid activation function, is the weight matrix of the gating mechanism, is the current hidden state of the decoder, is the matching degree between the currently generated content and the preset template;
[0038] The backtracking correction mechanism performs local regeneration on the illegal paragraphs and retains the Top-3 compliant candidate sequences.
[0039] Preferably, the training of the robustness enhancement model includes: injecting Gaussian noise into the input layer, and the noise intensity follows ; adopting the FGM adversarial training strategy and adding a perturbation term:
[0040] ;
[0041] In the formula, are the parameters after adversarial training perturbation, are the original model parameters, is the perturbation intensity, is the gradient vector, is the L2 norm of the gradient;
[0042] Adding a robustness loss term:
[0043] ;
[0044] In the formula, is the robustness loss, is the expectation operation, is the input sampled from the data distribution and is the input noise, is the model function.
[0045] Preferably, the weighted parameter fusion of the validation set performance is specifically: performing exponential smoothing weighting on the validation set losses of the three types of models, and the weight formula is:
[0046] ;
[0047] In the formula, is the weight of the th model, corresponding to the three types of models, is the validation set loss of the th model, is the exponential smoothing term;
[0048] The encoder layer parameters are weighted and fused using cosine similarity, and the decoder layer implements a Top-k credibility voting mechanism to retain the generation head parameters of the model with the highest weight.
[0049] Preferably, the user feedback closed-loop mechanism includes: parsing feedback features into triples, where the triples include: error type, original content, and correction result, and establishing a feedback database; starting online incremental learning for high-frequency error types, and the update formula is:
[0050] ;
[0051] In the formula, is the updated model parameter, is the original model parameter, is the learning rate, is the gradient of the parameter ; is the generated loss, is the feedback loss, is the weighting coefficient of the feedback loss.
[0052] Preferably, for the final deployed model, its deployment includes: encapsulating the final model as a RESTful API service that supports template input, allowing users to configure a credibility threshold; implementing real-time credibility monitoring, and calculating a risk index for the generated paragraph:
[0053] ;
[0054] In the formula, is the risk index of the generated paragraph, is the risk item weighting coefficient, is the single-paragraph risk score, is the credibility item weighting coefficient, is the credibility score of the generated content;
[0055] Each generated result is accompanied by a model version number, a training data snapshot ID, and a credibility assessment report.
[0056] Preferably, the working process of the legal compliance detection module is: constructing a legal provision knowledge graph, decomposing the provisions into quadruples; using the BiDAF model for clause matching, calculating the maximum slice similarity between the query statement and the knowledge base entries; for the detected expired clauses, comprehensively sorting and recommending alternative solutions based on the edit distance and semantic similarity, and the recommendation formula is:
[0057] ;
[0058] In the formula, is the ranking score of the alternative clause ; is the original clause and the semantic similarity between the alternative clause ; is the edit distance, is the weighted coefficient of semantic similarity and edit distance.
[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0060] The training method of the official document generation model proposed by the present invention significantly improves the quality and efficiency of official document generation by deeply integrating natural language processing technology and official document generation specifications. By constructing a multi-source and multi-type corpus, the richness and diversity of official document content are ensured. At the same time, duplicate removal and standardization processing effectively avoid the problems of information redundancy and inconsistent formats. The introduction of the evaluation system strictly controls the data quality, provides a reliable guarantee for model training, realizes dynamic constraints and real-time monitoring of the official document generation process, and significantly improves the legal compliance and semantic integrity of the generated content. In addition, by parallel training the basic controllable model, the trustworthy enhancement model and the robustness enhancement model, the adaptability and stability of the model are further enhanced, effectively resisting the influence of adversarial inputs, establishing a user feedback closed-loop mechanism, being able to continuously collect and analyze user feedback, and continuously optimizing the model performance to ensure the continuous progress of official document generation. Description of the Drawings
[0061] Figure 1 is the step diagram of the training method of the present invention;
[0062] Figure 2 is the mind map of the training method of the present invention. Detailed Embodiments
[0063] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0064] Refer to Figure 1 As shown, a training method for a controllable and trustworthy official document generation model includes the following steps:
[0065] S1. Collect original official document data based on government agencies, public databases, and legal document libraries, and establish a multi-source and multi-type corpus, where the multi-source and multi-type corpus includes: requests for instructions, meeting minutes, notices and announcements, and policy documents;
[0066] S2. Perform duplicate removal and standardization processing on the collected original corpus data;
[0067] S3. Establish an evaluation system composed of a format validator, a legal compliance detection module, and a semantic integrity analyzer to evaluate the data quality, and based on the evaluated data, divide the data into a training set and a test set;
[0068] S4. Construct a pre-trained language model based on the BART-large architecture, integrating a controllability gating mechanism and a credibility evaluation network;
[0069] S5. Parallelly train the basic controllable model, the credibility enhancement model, and the robustness enhancement model. Among them, the credibility enhancement model is connected to the laws and regulations database in real time for verification, and the robustness enhancement model injects adversarial noise. Dynamically monitor the controllable deviation and legal compliance of the generated content through the test set, and trigger an early stopping mechanism to prevent overfitting;
[0070] S6. Perform weighted parameter fusion on the performance of the three types of models after training on the validation set to generate the final deployment model, and automatically trigger iterative optimization through the user feedback closed-loop mechanism.
[0071] In the actual implementation process of the training method of the controllable and credible official document generation model of the present invention, it is first necessary to construct an official document corpus covering multiple sources and types. Obtain regulatory documents through the government information disclosure platform, extract meeting minutes templates from public databases, and integrate standard clauses in legal document libraries to form an initial data set including categories such as requests for instructions, reports, notices, and announcements. In the data preprocessing stage, the improved SimHash algorithm is used to compare the fingerprints of documents, and a similarity threshold is set to automatically screen and eliminate duplicate official documents to ensure the uniqueness of the corpus content. In the standardization process, the date format is uniformly converted through a preset rule engine. For example, "2023.12.31" is standardized to "December 31, 2023". At the same time, the numbers involving amounts are converted to Chinese capital letters, and a mapping table between the full name and the standardized abbreviation of the institution name is established to ensure the consistency of the data format.
[0072] In terms of model architecture design, BART-large is selected as the basic pre-trained model, and a dynamic constraint of the generation process is realized by introducing a controllability gating mechanism. Specifically, in the decoder layer, a rule constraint module is embedded to monitor in real time whether the generated content conforms to the official document writing specifications. For example, it detects the correctness of the agency code in the document number and the accuracy of the document type used. When format errors or expired legal clause references are found, the system immediately triggers a backtracking correction mechanism to retain multiple compliant candidate sequences for local regeneration. At the same time, a credibility evaluation network is constructed to connect to the API interface of the national laws and regulations database to verify the timeliness of the policy clauses cited in the generated content and automatically replace the repealed clause versions. The logical contradiction detection module analyzes the temporal logic conflicts in the generated content by constructing a time-event relationship graph. For example, the contradiction situation where the meeting resolution time is earlier than the meeting convening time.
[0073] During the model training stage, a three-way parallel training strategy is adopted. The basic controllable model focuses on learning the official document format specifications. The trustworthy enhancement model strengthens the accuracy of clause citation through a real-time legal verification mechanism. The robustness enhancement model injects Gaussian noise into the input layer and adopts an adversarial training strategy to improve the anti-interference ability. During the training process, the legal compliance deviation degree on the test set is dynamically monitored. When it is detected that there are consecutive rounds of format errors or clause citation failures in the generated content, the early stopping mechanism is immediately triggered to prevent the model from overfitting. During the adversarial training process, a perturbation term is added during gradient update, forcing the model to maintain stable generation quality under noise interference and increasing the robustness loss function to constrain the consistency of the generation results.
[0074] During the model fusion stage, differential parameter fusion is implemented according to the performance on the validation set. The validation set losses of the basic controllable model, the trustworthy enhancement model, and the robustness enhancement model are calculated with exponential smoothing weighting. The encoder layer adopts a parameter fusion strategy weighted by cosine similarity, and the decoder layer implements a voting mechanism based on credibility scores to retain the optimal generation head parameters. The final deployed model is encapsulated as a configurable API service, supporting users to customize the credibility threshold and generating a paragraph risk index assessment report in real time. The system has a built-in version management function, and each generation result is accompanied by a training data snapshot identifier and a legal compliance detection log to ensure the traceability of the generation process.
[0075] The user feedback mechanism constructs a closed-loop optimization system, parsing the error types marked by users into structured triples and storing them in the feedback database. For frequently occurring format errors or legal clause citation problems, the system automatically starts online incremental learning to perform targeted optimization on specific error types while maintaining the performance of the original model. The feedback data is incorporated into the model parameter update process through a weighted loss function, forming a continuously iterative self-improving mechanism. At the same time, a dynamic update interface for the legal knowledge graph is established. When a new version of the laws and regulations database is detected, the model retraining process is automatically triggered to ensure that the generated content always meets the latest policy requirements. After the entire system is deployed, the generation quality indicators, including the format compliance rate, the timeliness of legal clause updates, and the user satisfaction score, are displayed in real time through a visual monitoring panel, providing data support for model optimization.
[0076] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A training method for a controllable and trustworthy official document generation model, characterized in that, The following steps are involved: S1. Collect original official document data from government agencies, public databases and legal document libraries to establish a multi-source and multi-type corpus, which includes: request reports, meeting minutes, notices and policy documents; S2, de-duplication and standardization of the collected corpus raw data; S3. Establish an evaluation system consisting of a format checker, a legal compliance detection module, and a semantic integrity analyzer to evaluate data quality, and divide the data into a training set and a test set based on the evaluated data; S4. Build a pre-trained language model based on the BART-large architecture, integrating the controllability gating mechanism and the credibility evaluation network; S5. Parallel training of the basic controllable model, the trustworthy enhancement model and the robustness enhancement model. The trustworthy enhancement model is connected to the legal and regulatory database in real time for verification. The robustness enhancement model injects adversarial noise and dynamically monitors the controllable deviation and legal compliance of the generated content through the test set, triggering the early stopping mechanism to prevent overfitting. S6. Perform validation set performance weighted parameter fusion on the three types of models that have been trained to generate the final deployment model, and automatically trigger iterative optimization through a user feedback closed-loop mechanism; The credibility evaluation network comprises: The legal clause checker verifies the timeliness of referenced clauses through the national laws and regulations database API, and automatically replaces expired clauses with the latest version; The logical contradiction detector constructs a time-event relationship graph to identify timing conflicts in generated content; Source credibility score module calculation: ; In the formula, is the comprehensive credibility score, is the legal compliance score, is the content consistency score, is the weighting coefficient; The calculation formula for the legal compliance score is: ; Wherein, is the legal compliance score, is the total number of generated clauses, is the -th generated clause, is the legal knowledge base set, is the indicator function.
2. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that, The deduplication and standardization process described in step S2 is as follows: include: The deduplication process specifically includes: The improved SimHash algorithm is used to generate 64-bit document fingerprints, and the document similarity calculation is defined as: ; In the formula, is the similarity between the document and the document , is the Hamming distance between the document fingerprint and ; When the similarity is greater than a preset threshold, it is determined to be a nearly duplicate document and the duplicate document is removed; Among them, the standardized processing includes mandatory conversion of date format, conversion of amount numbers into uppercase Chinese characters, and mapping of full name of institution.
3. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that The workflow of the evaluation system is as follows: The format checker uses regular expressions to match the structural elements of official documents, including the combination rules of the agency code, year and sequence number of the document number; The legal compliance detection module matches policy clauses based on the knowledge graph and outputs a compliance score; The semantic completeness analyzer uses the RoBERTa model to calculate the contextual coherence score.
4. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that The controllability gating mechanism comprises: The rule constraint layer deploys the Drools rule engine that includes official document writing specifications, which can intercept missing document number agency codes and document type usage errors in real time; The dynamic gating unit calculates the gating weights in the decoder: ; wherein, is the gating weight of the step, is the Sigmoid activation function, is the weight matrix of the gating mechanism, is the current hidden state of the decoder, is the matching degree between the currently generated content and the preset template; The backtracking correction mechanism performs local regeneration on the illegal paragraphs and retains the Top-3 compliant candidate sequences.
5. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that The training of the robustness enhancement model includes: injecting Gaussian noise into the input layer, and the noise intensity follows ; adopting the FGM adversarial training strategy and adding a perturbation term: ; Wherein, is the parameter after adversarial training perturbation, is the original model parameter, is the perturbation intensity, is the gradient vector, is the L2 norm of the gradient; Add robustness loss term: ; wherein, is the robustness loss, is the desired operation, is the input sampled from the data distribution and is the input noise, is the model function. 6. The training method of a controllable and trustworthy official document generation model according to claim 1, wherein The validation set performance weighted parameter fusion is specifically: exponential smoothing weighting is performed on the validation set losses of the three types of models, and the weight formula is: ; Wherein, is the weight of the th model, corresponding to three types of models, is the validation set loss of the th model, is the exponential smoothing term; The encoder layer parameters are weighted fused using cosine similarity, and the decoder layer implements a Top-k credibility voting mechanism to retain the generation head parameters of the highest weight model.
7. A training method for a controllable and trustworthy official document generation model according to claim 1, characterized in that The user feedback closed-loop mechanism includes: parsing feedback features into triples, where the triples include: error type, original content, and correction result, and establishing a feedback database; initiating online incremental learning for high-frequency error types, and the update formula is: ; In the formula, is the updated model parameter, is the original model parameter, is the learning rate, is the gradient of the parameter , is the generation loss, is the feedback loss, is the weighting coefficient of the feedback loss.
8. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that For the final deployed model, its deployment includes: encapsulating the final model into a RESTful API service that supports template input, allowing users to configure a credibility threshold; implementing real-time credibility monitoring, and calculating a risk index for the generated paragraphs: ; Wherein, is the risk index of the generated paragraph, is the risk item weighting coefficient, is the single paragraph risk score, is the credibility item weighting coefficient, is the credibility score of the generated content; Each generated result is accompanied by a model version number, a training data snapshot ID, and a credibility assessment report.
9. The training method of a controllable and trustworthy official document generation model according to claim 1, characterized in that, The working process of the legal compliance detection module is as follows: construct a knowledge graph of legal provisions and decompose the provisions into quadruples; use the BiDAF model for clause matching to calculate the maximum slice similarity between the query statement and the knowledge base entries; for the detected expired clauses, comprehensively sort and recommend alternative solutions based on the edit distance and semantic similarity, and the recommendation formula is: ; Wherein, is the ranking score of the alternative clause , is the original clause and the semantic similarity between the original clause and the alternative clause is the edit distance is the weighting coefficient of the semantic similarity and the edit distance.
Citation Information
Patent Citations
Intelligent legal document generation method and system based on large model
CN119669485A
Document generation method based on large model and knowledge graph fusion
CN119719386A