Document writing method and system based on government affair industry large model and medium

Through the official document writing method based on the government affairs industry big model, the historical government affairs official document data training model is used to generate a logically consistent official document outline and conduct a standardized review, the problems of inconsistent and difficult to guarantee the logical nature of official document generation in the existing technology are solved, and efficient, coherent and standardized official document generation is achieved.

CN120163159APending Publication Date: 2025-06-17SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510210589.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the process of generating long document, the model may lack logical consistency between different paragraphs, resulting in incoherence of context and difficulty in ensuring that language expression complies with official document specifications.

Method used

By obtaining historical government official documents data, training pre-training models, generating official document outlines, and combining paragraph summary and context constraints, it is expanded to the initial official document text. Perform preset normative review of the initial official document text to ensure that it complies with the official document specifications.

Benefits of technology

It improves the logical consistency and contextual coherence of official document generation, ensures that the content of official document complies with formal official document writing standards, and improves the formality and authority of official documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163159A_ABST
    Figure CN120163159A_ABST
Patent Text Reader

Abstract

The invention discloses an official document writing method and system based on a government affair industry large model and a medium, mainly relates to the technical field of official document writing, and is used for solving the problems that in the prior art, in a long official document generation process, a model may lack logic consistency among different paragraphs, contexts are not coherent, and expression does not conform to official document specifications. Comprising the steps of obtaining a trained pre-training model; inputting the obtained cue word into a trained pre-training model, outputting an official document outline, and extracting a keyword corresponding to the official document outline; determining whether the keyword and the cue word have semantic consistency, and when the keyword and the cue word have semantic consistency, determining that the obtained document outline is qualified; expanding the qualified official document outline into an initial official document text in a manner of combining paragraph abstracts and context constraints; and performing preset normative auditing on the initial official document text, and modifying contents which do not accord with the preset normative to contents which accord with the preset normative to obtain a final official document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of official document writing, and particularly to an official document writing method, system and medium based on a large model for the government affairs industry. Background Art

[0002] Government official documents are an important part of administrative management activities and the main carrier for government agencies to convey information, issue policies and implement decisions. Official document writing not only requires the accuracy and standardization of content, but also needs to demonstrate high logic and coherence. However, at present, government official document writing mainly relies on manual work and has many deficiencies: Low efficiency: In actual work, it often takes a lot of time to write a standard government official document. Especially when dealing with long official documents involving multi-department collaboration or complex topics, the efficiency problem is particularly prominent. It is difficult to fully meet the standard of standardization: Government official documents have fixed formats and strict language requirements. In the process of manual writing, problems such as non-standard terms and improper formats are likely to occur, increasing the workload of subsequent review and revision.

[0003] Although there are already some automated writing tools for official document writing, in the process of generating long official documents, the model may lack logical consistency between different paragraphs, resulting in discontinuous context; secondly, although the model has strong generation ability, there are challenges in ensuring that the language expression conforms to the official document standard and needs further optimization. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, this application provides an official document writing method, system and medium based on a large model for the government affairs industry to solve the problems that in the process of generating long official documents in the prior art, the model may lack logical consistency between different paragraphs, easily lead to discontinuous context and non-conformity of expressions to the official document standard.

[0005] In a first aspect, this application provides an official document writing method based on a large model for the government affairs industry. The method includes: Obtain historical government official document data, input the historical government official document data into a pre-training module to obtain a trained pre-training model; Input the obtained prompt words into the trained pre-training model to output an official document outline, and at the same time extract the keywords corresponding to the official document outline; wherein, the official document outline at least includes: a title and preset-level sub-titles; determine whether there is semantic consistency between the keywords and the prompt words. When there is semantic consistency between the keywords and the prompt words, determine that the obtained official document outline is qualified; combine the paragraph summary with the context constraint to expand the qualified official document outline into an initial official document text; Perform a preset normative review on the initial official document text, modify the content that does not meet the preset norms to meet the preset norms, and obtain the final official document; among them, the preset language norms include at least vocabulary usage, grammatical structure, and format requirements.

[0006] In one implementation manner of the present application, historical government official document data is obtained, and the historical government official document data is input into a pre-training module to obtain a trained pre-training model, specifically including: Obtain a preset number of historical government official document data, preprocess the data and then input it into the pre-training module for training; among them, the historical government official document data includes the historical government official document content part and the historical government official document outline part; During the training process, set the learning rate to 2e-5, the number of training rounds to 3, and the batch size to 32, and then obtain a trained model as the pre-training model.

[0007] In one implementation manner of the present application, input the obtained prompt words into the trained pre-training model, specifically including: Obtain the requirement data input by the user through a preset interface; Extract prompt words from the requirement data through a semantic analysis algorithm.

[0008] In one implementation manner of the present application, combine the paragraph summary and the context constraint to expand the qualified official document outline into the initial official document text, specifically including: When the natural language processing algorithm generates the content of each paragraph based on the official document outline, automatically generate the paragraph summary of the current paragraph content; Use the paragraph summary of the current paragraph content as the reference data for generating the next paragraph content to generate the next paragraph content; After generating a paragraph of content, perform a context consistency check through the natural language processing algorithm. When the check is successful, continue to generate the next paragraph content; when the check fails, the natural language processing algorithm uses the paragraph summary of the previous paragraph content as the reference data to generate the next paragraph content again until the check is successful.

[0009] In one implementation manner of the present application, before performing a preset normative review on the initial official document text, modifying the content that does not meet the preset norms to meet the preset norms, and obtaining the final official document, the method further includes: Obtain the running program corresponding to the preset normative review through a preset backend interface, and establish a trigger relationship between the running program and the preset normative review task; When the preset normative review task is triggered, call the running program.

[0010] In a second aspect, the present application provides an official document writing system based on a government industry large model, and the system includes: A model acquisition module, configured to acquire historical government official document data, input the historical government official document data into a pre-training module, and obtain a trained pre-training model; a determination module, configured to input the obtained prompt words into the trained pre-training model, output an official document outline, and simultaneously extract keywords corresponding to the official document outline; wherein, the official document outline at least includes: a title and sub-titles at a preset level; determine whether there is semantic consistency between the keywords and the prompt words, and when there is semantic consistency between the keywords and the prompt words, determine that the obtained official document outline is qualified; combine the paragraph summary and the context constraint combination method to expand the qualified official document outline into an initial official document text; an official document acquisition module, configured to perform a preset standardization review on the initial official document text, modify the content that does not meet the preset standardization to the content that meets the preset standardization, and obtain a final official document; wherein, the preset language standardization at least includes vocabulary usage, grammar structure, and format requirements.

[0011] In an implementation manner of the present application, the model acquisition module includes a training unit, configured to acquire a preset number of historical government official document data, and after data preprocessing, input it into the pre-training module for training; wherein, the historical government official document data includes a historical government official document content part and a historical government official document outline part; During the training process, set the learning rate to 2e-5, the number of training rounds to 3 rounds, and the batch size to 32, so as to obtain a trained model as the pre-training model.

[0012] In an implementation manner of the present application, the determination module includes a verification unit, configured to automatically generate a paragraph summary of the current paragraph content when the natural language processing algorithm generates each paragraph content based on the official document outline; Use the paragraph summary of the current paragraph content as reference data for generating the next paragraph content to generate the next paragraph content; After generating a paragraph of content, perform context consistency verification through the natural language processing algorithm. When the verification is successful, continue to generate the next paragraph of content; when the verification fails, the natural language processing algorithm uses the paragraph summary of the previous paragraph of content as reference data to generate the next paragraph of content again until the verification is successful.

[0013] In an implementation manner of the present application, the system further includes a calling module, configured to obtain a running program corresponding to the preset standardization review through a preset back-end interface, and establish a triggering relationship between the running program and the preset standardization review task; When the preset standardization review task is triggered, call the running program.

[0014] In a third aspect, the present application provides a non-volatile computer storage medium, on which computer instructions are stored, and when the computer instructions are executed, a document writing method based on a large model of the government affairs industry as described in any one of the above is implemented.

[0015] Those skilled in the art can understand that the present application has at least the following beneficial effects: The present application provides a document writing method, system and medium based on a large model of the government affairs industry. By obtaining historical government affairs document data and inputting it into a pre-training module to obtain a trained pre-training model, the understanding and generation ability of the model for government affairs documents is ensured. By determining whether there is semantic consistency between the keyword and the prompt word. When there is semantic consistency between the keyword and the prompt word, it is determined that the document outline is qualified. The consistency between the document outline and the user requirements is ensured. Combining the paragraph summary with the context constraint, the qualified document outline is expanded into an initial document text. It helps to maintain the logical coherence and context consistency of the document content. Performing a preset normative review on the initial document text, modifying the content that does not meet the preset normativity to the content that meets the preset normativity, and obtaining the final document. It improves the logical consistency of document generation: Through the document writing method based on the large model of the government affairs industry, combining the paragraph summary with the context constraint, the problem of logical inconsistency that may occur in the process of generating long documents in the prior art is effectively solved. By extracting keywords and performing semantic consistency judgment with the prompt words, the consistency between the document outline and the user requirements is ensured, thereby improving the coherence of the document content. At the same time, performing a preset normative review and modification on the initial document text ensures that the document meets the formal document writing norms, improving the formality and authority of the document. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a document writing method based on a large model of the government affairs industry provided by an embodiment of the present application.

[0018] Figure 2 It is a schematic internal structure diagram of a document writing system based on a large model of the government affairs industry provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Those skilled in the art should understand that the embodiments described below are only the preferred embodiments of the present disclosure, and do not mean that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are only used to explain the technical principles of the present disclosure and are not used to limit the protection scope of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts should still fall within the protection scope of the present disclosure.

[0020] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0021] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0022] The embodiment provides a document writing method based on a large model for the government affairs industry, as Figure 1 shown, the method provided by the embodiment of the present application mainly includes the following steps: Step 110: Obtain historical government affairs document data, input the historical government affairs document data into a pre-training module, and obtain a trained pre-training model.

[0023] It should be noted that in this step, by obtaining historical government affairs document data and inputting it into the pre-training module, a trained pre-training model is obtained, which provides a solid foundation for subsequent document outline generation and text expansion, and improves the accuracy and efficiency of document generation. As an example: Obtain a preset number of historical government affairs document data (including content and outline), and after data preprocessing, input it into the pre-training module for training. Set specific learning rates, training epochs, and batch sizes to obtain a trained model as the pre-training model.

[0024] It should be further noted that the pre-training model is a large model based on the transformers type, such as open-source models like qwen, llama, and Baichuan.

[0025] In some embodiments, obtaining historical government affairs document data, inputting the historical government affairs document data into a pre-training module, and obtaining a trained pre-training model can specifically be: A preset amount of historical government document data is obtained, and after data preprocessing, it is input into the pre-training module for training; wherein, the historical government document data includes the historical government document content part and the historical government document outline part; during the training process, the learning rate is set to 2e-5, the number of training rounds is 3 rounds, and the batch size is 32, thereby obtaining a trained model as a pre-training model.

[0026] It should be noted that this application can use historical government document data (content + outline) to train the model.

[0027] For example: If you input "Notice on Environmental Protection Policy", the model can automatically match the framework of similar historical documents (such as "Policy Background → Implementation Requirements → Division of Responsibilities").

[0028] In addition, this application can also perform pre-processing such as cleaning and formatting of historical government documents, reduce noise interference (such as typos and redundant information), and improve model training efficiency and the standardization of generated text.

[0029] Example: Convert documents in different formats (Word, PDF) into structured text and remove irrelevant tables or comments.

[0030] Through the ‌low learning rate‌ (2e-5), overfitting can be avoided and the generalization ability of the model to government data can be ensured. Through the ‌few rounds of training‌ (3 rounds), the computational cost is reduced while ensuring that the model fully learns government features. Through the ‌moderate batch size‌ (32), the training speed and memory usage can be balanced.

[0031] For example: If the number of training rounds is too large (such as 10 rounds), the model may over-memorize the training data and lack flexibility in the generated text.

[0032] Through joint training of content and outline, this application can simultaneously learn the details of official document content and the overall framework logic, and take into account paragraph connection and structural rationality when generating text.

[0033] For example: When generating a "work summary", the model can automatically follow the outline logic of "work review → existing problems → improvement plan".

[0034] Step 120: input the acquired prompt words into the trained pre-trained model, output the official document outline, and extract the keywords corresponding to the official document outline; determine whether there is semantic consistency between the keywords and the prompt words. When there is semantic consistency between the keywords and the prompt words, determine that the official document outline is qualified; combine the paragraph summary with the context constraint to expand the qualified official document outline into the initial official document text.

[0035] It should be noted that the official document outline includes at least: a title and preset subheadings. In this step, the official document outline is obtained by inputting prompt words, and keywords are extracted for semantic consistency judgment to ensure the qualification of the official document outline. Combining the paragraph summary with context constraints, the outline is expanded into an initial official document text, improving the coherence and logic of the official document content. Specifically, in this step, the demand data input by the user is obtained through a preset interface, and after extracting the prompt words, they are input into a trained pre-trained model to generate the official document outline. The paragraph summary of each paragraph is generated using natural language processing algorithms and used as reference data for the next paragraph, ensuring the coherence of the content through context consistency verification.

[0036] Among them, inputting the obtained prompt words into the trained pre-trained model can specifically be: Obtain the demand data input by the user through a preset interface; extract the prompt words from the demand data through semantic analysis algorithms.

[0037] Among them, the qualified official document outline is expanded into an initial official document text by combining the paragraph summary with context constraints, specifically including: When the natural language processing algorithm generates the content of each paragraph based on the official document outline, automatically generate the paragraph summary of the current paragraph content; Use the paragraph summary of the current paragraph content as reference data to generate the next paragraph content; After generating a paragraph of content, perform context consistency verification through the natural language processing algorithm. When the verification is successful, continue to generate the next paragraph content; when the verification fails, the natural language processing algorithm regenerates the next paragraph content based on the paragraph summary of the previous paragraph content as reference data until the verification is successful.

[0038] It should be noted that after generating the current paragraph in this application, the paragraph summary can be automatically extracted (such as "Policy Background: Reasons for the Introduction of the Environmental Protection Policy in a Certain City"). The summary is used as a reference for generating the next paragraph (such as when generating "Implementation Requirements", it is necessary to associate with the summary of "Policy Background"). It can avoid logical jumps between paragraphs (such as directly jumping from "Background" to "Summary").

[0039] Example: When generating an "Environmental Protection Policy Notice", if the current paragraph is "Policy Background", the next paragraph will be automatically associated and generated as "Scope and Objectives of Implementation".

[0040] Here, when generating the subsequent paragraph, verify its logical relevance to the previous text (such as whether the keywords of the previous paragraph are mentioned). If the verification fails, regenerate based on the summary of the previous paragraph (such as deleting redundant descriptions that deviate from the theme).

[0041] Reduce the number of manual corrections and improve the automation rate.

[0042] Example: If the key departments in the "Policy Background" are not mentioned in the "Responsibility Assignment" section, the system will generate them automatically.

[0043] This application enforces that each section of content inherits the core information of the previous section's summary (e.g., the "Implementation Requirements" section needs to include the policy name in the "Background"). This achieves a natural transition between paragraphs (e.g., "Background → Requirements → Assignment" forms a logical chain).

[0044] Example: When generating a "Work Summary", the "Problems Existed" section will automatically echo the deficiencies in the "Work Review" section.

[0045] Step 130: Conduct a preset normative review on the initial official document text, modify the content that does not meet the preset norms to content that meets the preset norms, and obtain the final official document.

[0046] Among them, the preset language norms at least include vocabulary usage, grammar structure, and format requirements.

[0047] It should be noted that conducting a preset normative review on the initial official document text ensures that the official document meets the normative standards such as vocabulary usage, grammar structure, and format requirements, and improves the formality and authority of the official document. The specific method can be: obtaining the running program corresponding to the preset normative review through the preset backend interface and establishing a trigger relationship. When the preset normative review task is triggered, the running program is called to review and modify the initial official document text to obtain the final official document.

[0048] Among them, before conducting a preset normative review on the initial official document text, modifying the content that does not meet the preset norms to content that meets the preset norms, and obtaining the final official document, the method can include: Through the preset backend interface, obtain the running program corresponding to the preset normative review, establish a trigger relationship between the running program and the preset normative review task; when the preset normative review task is triggered, call the running program.

[0049] It should be noted that this application binds the trigger relationship between the review program and the task through the preset backend interface (such as "automatically trigger the review after submitting the official document"). There is no need to manually start the review program, which improves the efficiency. It ensures that each official document must undergo a normative review.

[0050] For example: After the user uploads the initial official document, the system automatically triggers the review program without the need to click the "Start Review" button additionally.

[0051] This application dynamically binds different review programs (such as "Vocabulary Review" or "Format Review") through the preset backend interface here. It realizes the ability to match different review rules according to the official document type (notice, report). When the review rules change, only the backend program needs to be replaced without modifying the main system.

[0052] For example: When adding a new "Data Security Terminology Audit", the administrator can directly associate the new program through the interface without deactivating the original system.

[0053] In addition, this application Figure 2 The embodiment of the present application provides a document writing system based on a large government industry model. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes: The model acquisition module 210 is used to obtain historical government document data, input the historical government document data into the pre-training module, and obtain a trained pre-training model.

[0054] It should be noted that the model acquisition module 210 obtains the historical government document data and inputs it into the pre-training module to obtain a trained pre-trained model, which provides a solid foundation for the subsequent document outline generation and text expansion, and improves the accuracy and efficiency of document generation. As an example: a preset amount of historical government document data (including content and outline) is obtained, and after data pre-processing, it is input into the pre-training module for training. Set a specific learning rate, number of training rounds, and batch size to obtain a trained model as a pre-trained model.

[0055] The model acquisition module 210 includes a training unit, Used to obtain a preset amount of historical government document data, and input it into the pre-training module for training after data pre-processing; wherein the historical government document data includes the content part of the historical government document and the outline part of the historical government document; During the training process, the learning rate is set to 2e-5, the number of training rounds is 3, and the batch size is 32, and then a trained model is obtained as a pre-training model.

[0056] It should be noted that the training unit can use historical government document data (content + outline) to train the model.

[0057] For example: If you input "Notice on Environmental Protection Policy", the model can automatically match the framework of similar historical documents (such as "Policy Background → Implementation Requirements → Division of Responsibilities").

[0058] In addition, this application can also perform pre-processing such as cleaning and formatting of historical government documents, reduce noise interference (such as typos and redundant information), and improve model training efficiency and the standardization of generated text.

[0059] Example: Convert documents in different formats (Word, PDF) into structured text and remove irrelevant tables or comments.

[0060] By using a low learning rate (2e-5), overfitting can be avoided, ensuring the generalization ability of the model for government affairs domain data. By using few training rounds (3 rounds), the computational cost is reduced while ensuring that the model fully learns government affairs features. By using a moderate batch size (32), the training speed and memory occupancy can be balanced.

[0061] For example, if the number of training rounds is too large (such as 10 rounds), it may cause the model to over-memorize the training data, resulting in inflexible generated text.

[0062] Through the joint training of content and outline, the training unit can synchronously learn the details of official document content and the overall framework logic, and take into account paragraph connection and structural rationality when generating text.

[0063] For example, when generating a "work summary", the model can automatically follow the outline logic of "work review → existing problems → improvement plan".

[0064] The determination module 220 is used to input the obtained prompt words into the trained pre-trained model, output the official document outline, and extract the keywords corresponding to the official document outline; wherein, the official document outline at least includes: a title and sub-titles at a preset level; determine whether there is semantic consistency between the keywords and the prompt words, and when there is semantic consistency between the keywords and the prompt words, determine that the obtained official document outline is qualified; combine the paragraph summary and the context constraint to expand the qualified official document outline into an initial official document text.

[0065] It should be noted that the official document outline at least includes: a title and sub-titles at a preset level. The determination module 220 obtains the official document outline by inputting the prompt words, extracts the keywords for semantic consistency judgment, ensuring the qualification of the official document outline. Combining the paragraph summary and the context constraint to expand the outline into an initial official document text improves the coherence and logic of the official document content. Specifically, in this step, the demand data input by the user is obtained through a preset interface, and after extracting the prompt words, they are input into the trained pre-trained model to generate the official document outline. The paragraph summary of each paragraph is generated using natural language processing algorithms and used as reference data for the next paragraph, and the context consistency check is performed to ensure the coherence of the content.

[0066] The determination module 220 includes a verification unit, which is used to automatically generate the paragraph summary of the current paragraph content when the natural language processing algorithm generates each paragraph content based on the official document outline; use the paragraph summary of the current paragraph content as the reference data for generating the next paragraph content; after generating a paragraph of content, perform a context consistency check through the natural language processing algorithm. When the check is successful, continue to generate the next paragraph content; when the check fails, the natural language processing algorithm regenerates the next paragraph content based on the paragraph summary of the previous paragraph content as the reference data until the check is successful.

[0067] It should be noted that after the verification unit generates the current paragraph, it can automatically extract the paragraph summary (such as "Policy background: Reasons for the introduction of the environmental protection policy in a certain city"). The summary is used as a reference for generating the next paragraph (for example, when generating "Implementation requirements", the summary of "Policy background" needs to be associated). This can avoid logical jumps between paragraphs (such as directly jumping from "Background" to "Summary").

[0068] Example: When generating the "Environmental protection policy notice", if the current paragraph is "Policy background", the next paragraph will automatically be associated and generated as "Implementation scope and objectives".

[0069] Here, when generating the subsequent paragraph, its logical relevance to the previous text is verified (such as whether the keywords of the previous paragraph are mentioned). If the verification fails, it is regenerated based on the summary of the previous paragraph (such as deleting redundant descriptions that deviate from the theme).

[0070] The number of manual corrections is reduced, and the automation rate is improved.

[0071] Example: If the "Responsibility division" paragraph does not mention the key departments in the "Policy background", the system will generate them automatically.

[0072] This application forces each paragraph to inherit the core information of the summary of the previous paragraph (such as "Implementation requirements" need to include the policy name in the "Background"). This achieves a natural transition between paragraphs (such as "Background → Requirements → Division of labor" forming a logical chain).

[0073] Example: When generating the "Work summary", the "Existing problems" paragraph will automatically echo the deficiencies in the "Work review".

[0074] The official document acquisition module 230 is used to perform a preset normative review on the initial official document text, modify the content that does not meet the preset norms to content that meets the preset norms, and obtain the final official document; among them, the preset language norms at least include vocabulary usage, grammar structure, and format requirements.

[0075] The system further includes a calling module, which is used to obtain the running program corresponding to the preset normative review through a preset back-end interface, establish a triggering relationship between the running program and the preset normative review task; when the preset normative review task is triggered, the running program is called.

[0076] In addition, the embodiment of this application also provides a non-volatile computer storage medium, on which executable instructions are stored, and when the executable instructions are executed, the above-mentioned official document writing method based on the government affairs industry large model is implemented.

[0077] So far, the technical solutions of the present disclosure have been described in connection with multiple embodiments of the foregoing. However, those skilled in the art can easily understand that the protection scope of the present disclosure is not limited to these specific embodiments. Without departing from the technical principles of the present disclosure, those skilled in the art can split and combine the technical solutions in the above-mentioned various embodiments, and can also make equivalent changes or substitutions to the relevant technical features. Any changes, equivalent substitutions, improvements, etc. made within the technical concept and / or technical principles of the present disclosure will fall within the protection scope of the present disclosure.

Claims

1. A document writing method based on a government affairs industry big model, characterized in that: The method comprises: Obtain historical government document data, input the historical government document data into a pre-training module, and obtain a trained pre-training model; Input the acquired prompt words into the trained pre-trained model, output the official document outline, and extract the keywords corresponding to the official document outline; wherein the official document outline includes at least: a title and a preset subtitle; determine whether there is semantic consistency between the keywords and the prompt words, and when there is semantic consistency between the keywords and the prompt words, determine that the official document outline is qualified; combine the paragraph summary with the context constraint to expand the qualified official document outline into the initial official document text; The initial official document text is reviewed for preset normativeness, and the content that does not conform to the preset normativeness is modified to content that conforms to the preset normativeness to obtain the final official document; among which, the preset language normativeness at least includes vocabulary usage, grammatical structure and format requirements.

2. The document writing method based on the government affairs industry big model according to claim 1 is characterized in that: Obtain historical government document data, input the historical government document data into the pre-training module, and obtain a trained pre-training model, specifically including: Obtain a preset amount of historical government document data, and input the data into a pre-training module for training after data pre-processing; wherein the historical government document data includes a historical government document content part and a historical government document outline part; During the training process, the learning rate is set to 2e-5, the number of training rounds is 3, and the batch size is 32, and then a trained model is obtained as a pre-training model.

3. The document writing method based on the government affairs industry big model according to claim 1 is characterized in that: Input the obtained prompt words into the trained pre-trained model, including: Obtain the required data input by the user through the preset interface; Through semantic analysis algorithm, prompt words are extracted from demand data.

4. The document writing method based on the government affairs industry big model according to claim 1 is characterized in that: By combining paragraph summaries with contextual constraints, the qualified document outline is expanded into the initial document text, including: When the natural language processing algorithm generates each paragraph of content based on the official document outline, it automatically generates a paragraph summary of the current paragraph; Use the paragraph summary of the current paragraph as reference data for the next paragraph to generate the next paragraph; After a piece of content is generated, a context consistency check is performed through a natural language processing algorithm. When the check succeeds, the next piece of content is generated. When the check fails, the natural language processing algorithm generates the next piece of content again based on the paragraph summary of the previous piece of content as reference data until the check succeeds.

5. The document writing method based on the government affairs industry big model according to claim 1 is characterized in that: Before the initial official document is reviewed for preset normativeness and the contents that do not conform to the preset normativeness are modified to contents that conform to the preset normativeness and the final official document is obtained, the method further includes: Through the preset backend interface, obtain the running program corresponding to the preset normative review, and establish a trigger relationship between the running program and the preset normative review task; When the preset regulatory review task is triggered, the running program is called.

6. A document writing system based on a large model of government affairs industry, characterized in that: The system comprises: The model acquisition module is used to obtain historical government document data, input the historical government document data into the pre-training module, and obtain a trained pre-training model; The determination module is used to input the acquired prompt words into the trained pre-trained model, output the official document outline, and extract the keywords corresponding to the official document outline; wherein the official document outline includes at least: a title and a preset subtitle; determine whether there is semantic consistency between the keywords and the prompt words, and when there is semantic consistency between the keywords and the prompt words, determine that the official document outline is qualified; combine the paragraph summary with the context constraint to expand the qualified official document outline into the initial official document text; The official document acquisition module is used to conduct a preset normative review on the initial official document text, modify the content that does not conform to the preset normative content to content that conforms to the preset normative content, and obtain the final official document; wherein the preset language normative content at least includes vocabulary usage, grammatical structure and format requirements.

7. The document writing system based on the government affairs industry big model according to claim 6 is characterized in that: The model acquisition module includes a training unit, Used to obtain a preset amount of historical government document data, and input it into the pre-training module for training after data pre-processing; wherein the historical government document data includes the content part of the historical government document and the outline part of the historical government document; During the training process, the learning rate is set to 2e-5, the number of training rounds is 3, and the batch size is 32, and then a trained model is obtained as a pre-training model.

8. The document writing system based on the government affairs industry big model according to claim 6 is characterized in that: The determination module includes a verification unit, Used to automatically generate a paragraph summary of the current paragraph when the natural language processing algorithm generates each paragraph based on the official document outline; Use the paragraph summary of the current paragraph as reference data for the next paragraph to generate the next paragraph; After a piece of content is generated, a context consistency check is performed through a natural language processing algorithm. When the check succeeds, the next piece of content is generated. When the check fails, the natural language processing algorithm generates the next piece of content again based on the paragraph summary of the previous piece of content as reference data until the check succeeds.

9. The document writing system based on the government affairs industry big model according to claim 6 is characterized in that: The system also includes a calling module, Used to obtain the running program corresponding to the preset normative review through the preset backend interface, and establish a trigger relationship between the running program and the preset normative review task; When the preset regulatory review task is triggered, the running program is called.

10. A non-volatile computer storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed, they implement a method for writing official documents based on a large model of the government affairs industry as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Customized text generation method and device, computer equipment and storage medium

    CN121706740A

  • Official document abstract generation quality evaluation method and system based on large language model

    CN121919348A