Official document content generation and compliance verification system based on large language model

The document content generation and compliance verification system based on a large language model solves the problems of low efficiency and difficulty in ensuring quality in traditional document writing, and realizes efficient and standardized document generation and compliance verification, thereby improving the quality and consistency of documents.

CN120873172AActive Publication Date: 2025-10-31CHONGQING BORA INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511384733.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Traditional official document writing is inefficient, difficult to guarantee quality, and lacks unified standards and norms, resulting in inconsistent document quality. In particular, in emergency situations, problems such as incomplete content and unclear logic are prone to occur.

Method used

The document content generation and compliance verification system based on a large language model includes a document extraction module, a content generation module, a compliance verification module, and a writing assistance module. By analyzing writing style, constructing a knowledge graph, and verifying format and logical consistency in real time, it provides real-time modification suggestions to ensure that documents comply with regulations.

Benefits of technology

Significantly improve the efficiency and quality of official document writing, ensure the compliance of document format and content, reduce basic errors, enhance the authority and readability of official documents, and achieve seamless knowledge transfer and personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873172A_ABST
    Figure CN120873172A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, in particular to an official document content generation and compliance verification system based on a large language model. Comprising an official document extraction module used for analyzing a model essay by using a large language model, extracting a writing style and a fixed sentence pattern, constructing a knowledge graph, storing format specifications, common expressions and logic structures, mining and constructing a core intention library, and storing typical intention prototype vectors; the content generation module is used for generating paragraphs through context sensing, expanding and writing keywords, analyzing behavior data, predicting writing intention vectors and performing condition control; the compliance verification module is used for warning improper words in real time, retrieving policy verification logic and verification formats, and performing pre-verification based on intention vectors; and the writing auxiliary module is used for providing real-time modification suggestions and forward-looking prompts based on writing intention vectors. According to the technical scheme, document writing efficiency and quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, specifically to a system for generating and verifying official document content based on large language models. Background Technology

[0002] Traditional official document writing relies primarily on manual labor. From conceptualizing the framework and collecting data to writing word by word, the entire process is time-consuming and labor-intensive. This is especially true for large organizations or departments that frequently handle official documents, where staff typically spend a significant amount of time on repetitive writing tasks such as drafting routine notices and reports. Furthermore, when faced with urgent official documents, manual writing may compromise quality due to time constraints, easily resulting in incomplete content and unclear logic.

[0003] The varying writing skills, knowledge base, and experience among different staff members make it difficult to maintain consistent document quality. Novice or non-professional writers may lack a proper grasp of document formatting standards, language requirements, and logical structure, easily leading to formatting errors, inappropriate word choice, and awkward phrasing. Even experienced writers may become fatigued after long hours of intense work, resulting in oversights and affecting document quality. Furthermore, differences in writing styles and requirements across departments and scenarios, coupled with a lack of unified standards and guidelines, further exacerbate the inconsistency in document quality. Summary of the Invention

[0004] The purpose of this invention is to propose a document content generation and compliance verification system based on a large language model, which can improve the efficiency and quality of document writing.

[0005] To achieve the above objectives, this invention provides a document content generation and compliance verification system based on a large language model, comprising: The document extraction module is used to analyze uploaded sample documents through a large language model, extract the writing style and fixed sentence patterns of different document types, construct a knowledge graph of document writing, and store the format specifications, common expressions and logical structures of various documents. The document extraction module mines and constructs a core intent library for different document types, which is used to store typical intent prototype vectors in the field of document writing, with each intent corresponding to a vector representation. The content generation module is used to generate subsequent paragraphs based on the context of the input content; expand the text based on keywords; and analyze behavioral data, using a lightweight neural network model to predict the writing intent vector, which serves as a conditional control parameter. The compliance verification module is used to provide real-time alerts for inappropriate wording; it uses RAG technology to retrieve relevant policies and verify the logical consistency of the content; it automatically verifies whether the format of each element of the official document conforms to the standards; and it performs pre-verification based on the writing intent vector. The writing assistance module provides real-time editing suggestions, displayed as annotations in the document sidebar; it also offers forward-looking prompts based on writing intent vectors.

[0006] The basic solution has the advantage of building a knowledge graph and classification template library for document extraction, which can accurately match templates by institution and scenario. Users do not need to learn the format specifications and common sentence patterns of different documents from scratch. They can directly fill in the content based on the template, which greatly shortens the time for format learning and framework building.

[0007] The content generation module is context-based and can automatically match the logical structure of the document type to generate subsequent paragraphs based on the user's input of the beginning. The keyword expansion module allows users to input core information and automatically expand it into a complete expression that conforms to the style of official documents, thus lowering the threshold for non-professional writers to write official documents.

[0008] The writing assistance module provides real-time suggestions in the form of sidebar annotations and highlights questionable content to avoid users having to repeatedly read and revise after completing the first draft, thus improving writing fluency.

[0009] This technical solution enables timely compliance verification, avoids basic errors and policy risks, identifies prohibited / non-standard expressions in official documents in real time to prevent inappropriate wording from affecting the seriousness of official documents, retrieves the latest policy documents through RAG technology to verify whether the content of official documents is consistent with current policies, and avoids the risk of policy disconnection, and automatically checks whether the essential elements of official documents are complete and whether the format meets the standards, so as to avoid the invalidity of official documents due to format omissions.

[0010] The document extraction module's writing style library can distinguish the expression characteristics of documents from different institutions / scenarios, and accumulate fixed sentence patterns and logical structures to ensure that the documents written by users are not only compliant, but also in line with professional expression habits, thereby enhancing the authority and readability of the documents.

[0011] The document extraction module transforms scattered sample document experiences into a structured knowledge graph, clearly marking document types, essential elements, common sentence structures, and the relationship between policy basis and applicable scenarios. By querying the knowledge graph, users can quickly grasp the key points of writing without relying on personal experience.

[0012] Templates are categorized and managed by institution / scenario, and user-defined templates are also supported. This ensures both the universality of knowledge and meets the personalized needs of different organizations. At the same time, the template library can be iterated in real time as policies are updated and institutional needs change, ensuring the timeliness of knowledge.

[0013] By importing high-quality internal official document templates into the system, the document extraction module automatically analyzes and supplements them into the knowledge graph and template library, forming an internal exclusive knowledge base. New employees can quickly master the organization's official document writing standards through the system without long-term training, achieving seamless knowledge transfer.

[0014] By combining the core intent library built from the document extraction module, the compliance verification module can perform pre-verification based on the writing intent vector to determine whether the current writing direction aligns with the core purpose of the document type. This prevents documents from becoming off-topic due to deviating intent, ensuring compliance from a logical writing perspective. Based on the core intent library of the document extraction module, the writing assistance module can predict the writing intent vector based on the user's input and provide forward-looking prompts. This helps users plan their writing structure in advance, avoiding the discovery of deviations in thought later in the writing process and ensuring logical coherence and completeness of elements in the document.

[0015] As a feasible and preferred approach, the analysis of behavioral data includes the following: The formula for extracting text input features is as follows:

[0016] in, For word frequency, For document frequency, Total number of documents; Cursor behavior encoding, using Fourier transform to reduce the dimensionality of the trajectory sequence:

[0017] in, and Indicates the first n The cursor's two-dimensional coordinates in the document interface when sampling points are reached. n It is the summation index, representing the sequence number of the trajectory sampling point; This indicates the total number of trajectory sampling points within the sliding window; Indicates the frequency component index; Indicates the first Complex values ​​of each frequency component; i It is the imaginary unit.

[0018] As a feasible and preferred solution, the document extraction module inputs the uploaded sample documents into a pre-trained large language model; identifies the document type through a text classification algorithm; analyzes the writing style of the documents, including sentence structure, word usage habits, and paragraph layout, and extracts fixed sentence patterns and style features; and stores the extracted document type and writing style features in a database for subsequent template construction.

[0019] As a feasible and preferred solution, the document extraction module integrates document templates and format specifications from multiple channels; it uses natural language processing technology to parse the semantic relationships and structural information in the templates; based on the Neo4j graph database, it stores the parsed knowledge as nodes and edges to form a document writing knowledge graph, where nodes represent document elements and edges represent the relationships between elements; and it regularly updates the knowledge graph to incorporate the latest document formats and expressions.

[0020] As a feasible and preferred solution, the content generation module is used to input the input content into a large language model for encoding and analysis. The model understands the current writing context by analyzing the semantics and structure of the input content. Based on the contextual understanding, the model generates relevant suggestions for subsequent paragraphs. The predicted paragraphs are displayed to the user in a list format, and the user can choose to accept or reject the suggestions.

[0021] As a feasible and preferred solution, the content generation module is also used to extract keywords from the input; retrieve information related to the keywords from knowledge graphs and external policy databases; integrate the retrieved information into supplementary content by combining the text generation capabilities of the large language model; and display the supplementary content to the user for reference and editing.

[0022] As a feasible and preferred solution, the content generation module controls the style of the generated text by adjusting the generation parameters of the large language model; based on the adjusted parameters, the model generates multiple versions of the official document content; and the multiple versions of the official document content are displayed side by side.

[0023] As a feasible and preferred solution, the document extraction module includes a variational autoencoder model, which is used to decouple the semantic content and writing style of the sample document in the latent space to obtain separate content encoding vectors and style encoding vectors. The content generation module generates a content-neutral draft based on the content encoding vector; it also includes a generative adversarial network, which is used to fuse the draft with the style encoding vector to generate target style text, and a discriminator to determine whether the generated text truly conforms to the target style.

[0024] As a feasible and preferred solution, the compliance verification module stores the sensitive word library in the database; performs real-time detection of official document content through string matching algorithms; adopts a multi-pattern matching algorithm to scan the official document text word by word, identifies precisely matched sensitive words, and expands the scope of sensitive word detection based on the synonym relationships in the knowledge graph; converts the official document text and sensitive words into semantic vectors respectively, calculates similarity, and identifies semantically similar sensitive expressions; when official document content is input, it detects and highlights sensitive words in real time, and pops up an alert box in the sidebar to display the sensitivity level, sensitive word definition, and provides differentiated suggestions for different levels of sensitive words.

[0025] As a feasible and preferred writing assistance module, this feature identifies questionable content and marks and highlights it by setting rules and thresholds. Attached Figure Description Figure 1 This is a schematic diagram of the system architecture for document content generation and compliance verification based on a large language model.

[0026] Figure 2 This is a schematic diagram of the electronic device structure according to an embodiment of the present invention.

[0027] The attached figures indicate the following: electronic device 500, processor 501, communication interface 502, memory 503, and bus 504. Detailed Implementation

[0028] To make the technical solution and advantages of this application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only some embodiments of the present invention, and are only used to explain this application, not to limit it. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated; they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the accompanying drawings of the following embodiments represent the same features or components, and can be applied to different embodiments.

[0029] Furthermore, unless otherwise defined, the technical or scientific terms used in this invention description shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains.

[0030] The present invention will now be described in further detail with reference to the accompanying drawings.

[0031] Example 1 Reference Figure 1 This invention provides a document content generation and compliance verification system based on a large language model, including a document extraction module, a content generation module, and... The document extraction module includes a sample document analysis submodule, a knowledge graph construction submodule, and a template management submodule.

[0032] The sample document analysis submodule is used to analyze user-uploaded sample documents using a large language model, automatically extracting writing styles and fixed sentence structures for different document types (such as requests, reports, and notices). Specifically, the user-uploaded sample documents are input into a pre-trained large language model; the model identifies the document type, such as requests, reports, and notices, using a built-in text classification algorithm. Further analysis of the document's writing style, including but not limited to sentence structure, word choice habits, and paragraph layout, is performed to extract fixed sentence structures and style features. The extracted document type and writing style features are stored in a database for subsequent template construction.

[0033] The knowledge graph construction submodule is used to build a knowledge graph for official document writing, storing the format specifications, common expressions, and logical structures of various official documents. Specifically, it integrates official document templates and format specifications from multiple channels; uses natural language processing technology to parse the semantic relationships and structural information in the templates; and stores the parsed knowledge as nodes and edges based on graph databases such as Neo4j, forming the official document writing knowledge graph. Nodes represent official document elements (such as titles, body text, attachments, etc.), and edges represent the relationships between elements (such as inclusion, order, etc.). The knowledge graph is updated regularly to incorporate the latest official document formats and expressions.

[0034] The template management submodule is used for multi-level template management, differentiating writing requirements for different organizations and scenarios. Specifically, it categorizes and stores templates according to organization and scenario, such as government department templates, enterprise templates, and emergency notice templates; it provides a template search function, allowing users to filter suitable templates based on their organization and specific scenario; it supports users in customizing templates according to their own needs and storing customized templates in the database; and it regularly checks and updates the template library to ensure the timeliness and applicability of the templates.

[0035] The content generation module includes the associative filling submodule, the keyword expansion submodule, and the multi-version generation submodule.

[0036] The autocomplete submodule provides context-aware autocomplete functionality, predicting subsequent paragraphs based on the input. The user's input is fed into a large language model for encoding and analysis; the model understands the current writing context by analyzing the semantics and structure of the input; based on this contextual understanding, the model generates relevant suggestions for subsequent paragraphs; the predicted paragraphs are then displayed to the user in a list format, allowing the user to accept or reject the suggestions.

[0037] The keyword expansion submodule supports keyword-based document expansion, automatically supplementing relevant background information and policy basis. Specifically, it extracts keywords from user input; retrieves information related to the keywords from knowledge graphs and external policy databases; integrates the retrieved information into supplementary content using the text generation capabilities of a large language model; and displays the supplementary content to the user for reference and editing.

[0038] The multi-version generation submodule is used to generate multiple versions of the document, providing users with different writing styles to choose from. Specifically, the style of the generated text is controlled by adjusting the generation parameters of the large language model (such as intonation, maximum length, etc.). Based on the adjusted parameters, the model generates multiple versions of the official document content. These multiple versions are then displayed side-by-side to the user, who can choose the appropriate version according to their needs.

[0039] The compliance verification module includes a sensitive word detection submodule, a policy conflict detection submodule, and a format specification check submodule.

[0040] The sensitive word detection submodule has a built-in hierarchical sensitive word library (ordinary / confidential / top secret) for real-time alerts on inappropriate word usage. The sensitive word library is stored in a database. Ordinary sensitive words are derived from common errors in official document review and a list of negative expressions in government information disclosure; confidential / top secret sensitive words are derived from relevant confidentiality regulations. The system uses a string matching algorithm to perform real-time detection on the content of official documents.

[0041] Specifically, a multi-pattern matching algorithm (AC automaton algorithm) is used to scan the official document text word by word to identify precisely matched sensitive words. Based on the synonym relationships in the knowledge graph, the scope of sensitive word detection is expanded. The official document text and sensitive words are converted into semantic vectors respectively (using the Sentence-BERT model), and similarity is calculated to identify sensitive expressions with similar semantics.

[0042] When entering official document content, the system detects sensitive words in real time and immediately highlights them in the text (ordinary sensitive words in yellow, confidential sensitive words in orange, and top secret sensitive words in red), and pops up an alert box in the sidebar displaying the sensitivity level and definition of the sensitive word; it also provides differentiated suggestions for different levels of sensitive words—ordinary sensitive words are given modification suggestions (such as "'non-compliant with regulations' is suggested to be changed to 'needs further standardization'"); confidential / top secret sensitive words are prompted that "they need to be deleted or replaced with non-sensitive expressions and undergo confidentiality review", and the detection log is recorded at the same time.

[0043] The policy conflict detection submodule uses RAG technology to retrieve relevant policies and verify the logical consistency of the content. RAG technology searches for relevant policies in knowledge graphs and external policy databases to verify the logical consistency of official document content, ensuring that the content meets policy requirements. When a policy conflict is detected, the system prompts the user and makes corrections.

[0044] The format specification check submodule is used to automatically verify whether the format of each element of an official document conforms to the standard. Based on the official document format standard, the format specification check module defines the format requirements for elements such as the title, body, and attachments; it performs format verification on these elements; and when a format error is detected, the system prompts the user and makes corrections.

[0045] The writing assistance module includes a real-time suggestion submodule, an intelligent optimization submodule, and a content tagging submodule.

[0046] The real-time suggestion submodule provides real-time modification suggestions, displayed as annotations in the document's sidebar. It generates modification suggestions by analyzing the document content, combining knowledge graphs and large language models, and displays these suggestions as annotations in the document's sidebar for user reference.

[0047] The intelligent optimization submodule automatically adjusts sentence expression and paragraph structure to improve the readability of official documents. It optimizes document content using natural language processing algorithms, including checking sentence fluency and adjusting paragraph structure.

[0048] The content tagging submodule is used to highlight questionable content, making it easier for users to review it. Rules and thresholds are set to identify, tag, and highlight questionable content for user review.

[0049] Example 2 When a user enters text in the writing interface, the system captures this input in real time and passes it to the associative completion submodule. The associative completion submodule first preprocesses the input text, including word segmentation and stop word removal, to reduce noise interference. Then, it uses a pre-trained large language model to encode the processed text, converting it into a high-dimensional vector representation. This vector representation captures the semantic information and structural features of the text, providing a foundation for subsequent analysis.

[0050] After obtaining the vector representation of the text, the associative filling submodule further analyzes the inherent relationships between the vectors to understand the semantics and structure of the text. Through techniques such as clustering and topic modeling, the module can identify key themes, sentiment tendencies, and logical relationships between paragraphs in the text. For example, if the user inputs a report about project progress, the module can identify elements such as time points, key achievements, and existing problems in the report, and construct a semantic framework for the text accordingly.

[0051] Based on the results of semantic analysis, the module gains a deep understanding of the current writing context. This includes identifying the type of document the user is writing (such as a request, report, or notice), the topic of the paragraphs, and the user's possible writing intentions. By building a context model, the module can dynamically adjust its prediction strategy to more accurately match the user's writing needs.

[0052] Based on an understanding of the context, the module leverages the generative capabilities of a large language model to predict subsequent paragraphs the user might write. In this process, the module considers not only the semantic coherence of the text but also the language style and formatting conventions of official documents. For example, when predicting a paragraph about a project application, the module references common expressions and structures found in similar official documents to generate a predictive text that conforms to the standards and is semantically fluent.

[0053] The generated predicted paragraphs are displayed in a list next to the user's writing interface, each accompanied by a brief description or tag to help users quickly understand its content. The system also provides a confidence score for each predicted paragraph, reflecting the model's assessment of its accuracy.

[0054] Users can choose to accept or reject predicted paragraphs based on their needs. If a user chooses to accept a predicted paragraph, that paragraph will be automatically inserted into the user's writing text; if the user chooses to reject it, the context prediction submodule will continue to provide other prediction options until the user finds a satisfactory paragraph.

[0055] Example 3 The intelligent optimization submodule uses word segmentation algorithms (such as rule-based word segmentation, conditional random field (CRF) word segmentation, or deep learning word segmentation models) to segment each sentence in the official document and label the part of speech of each word (such as noun, verb, adjective, etc.), providing basic data for subsequent grammatical and semantic analysis.

[0056] The intelligent optimization submodule uses methods such as dependency parsing or phrase structure tree analysis to parse the grammatical structure of sentences. By identifying core components such as subject, verb, and object and their modifying relationships, the system can determine whether the sentence conforms to grammatical rules and detect potential grammatical errors, such as subject-verb disagreement and verb tense errors.

[0057] The intelligent optimization submodule uses techniques such as semantic role labeling or semantic similarity calculation to understand the meaning of sentences. By identifying entities, events, and their relationships within sentences, it determines whether the sentences are clearly expressed and logically sound, and identifies potential semantic ambiguities or unclear expressions.

[0058] For the identified grammatical and semantic issues, the system generates specific optimization suggestions. For example, for subject-verb disagreement, the system suggests adjusting the form of the subject or verb; for semantic ambiguity, the system suggests replacing vague words or adjusting the sentence structure. These suggestions will be displayed as annotations in the document sidebar for reference and selection.

[0059] This application also provides a method for generating and verifying the content of official documents based on a large language model, which utilizes the aforementioned system for generating and verifying the content of official documents based on a large language model.

[0060] Example 4 An intent-aware model for dynamic writing is constructed. A core intent library for different document types is built through a document extraction module. This library stores typical intent prototype vectors in the document writing domain, with each intent corresponding to a vector representation. Core intents include requesting approval, reporting feedback, and informing execution.

[0061] When a user is writing, the content generation module analyzes the user's input, cursor position, deletion and modification sequences in real time, and uses a lightweight neural network model (such as a miniaturized version of LSTM or Transformer) to dynamically predict the user's most likely writing intent. In this embodiment, the user's behavioral sequence is used as the key feature for intent judgment, rather than judging solely from the text.

[0062] Specifically, the content generation module captures the user's text in real time to capture the user's behavioral data, including text input sequences, cursor movement trajectories, editing operations, and time interval data; and converts the behavioral data into multi-dimensional feature vectors; and uses a lightweight temporal neural network model to encode the behavioral sequences and classify intents.

[0063] In this embodiment, the lightweight temporal neural network model adopts an improved two-layer LSTM structure with a hidden layer dimension ≤128; or a miniaturized Transformer structure with a head number ≤4 and a layer number ≤2.

[0064] The formula for extracting text input features is as follows:

[0065] in, For word frequency, For document frequency, This represents the total number of documents.

[0066] The cursor behavior encoding, using Fourier transform on the trajectory sequence, will be:

[0067] in, and This represents the two-dimensional coordinates of the cursor in the document interface at the nth sampling point; n is the summation index, representing the sequence number of the trajectory sampling point, used to traverse all cursor trajectory sampling points within the sliding window; This indicates the total number of trajectory sampling points within the sliding window; Indicates the frequency component index; Indicates the first The complex values ​​of each frequency component; i is the imaginary unit, satisfying i 2 =-1.

[0068] The process by which the content generation module predicts the writing intent vector is as follows: Set a sliding time window to collect behavior sequences In this embodiment, the sliding time window is set to 200ms.

[0069] Generate context representation through encoder :

[0070] in, Indicates model parameters.

[0071] Calculate the distribution of intent similarity Determine the dominant writing intention for the current writing task:

[0072] The softmax transformation converts the similarity scores of multiple intentions into mutually exclusive probability distributions. For the prototype matrix of the intention, T Indicates temperature parameter, ,when As the probability distribution approaches 0, it tends towards the maximum difference. Approaching When the probability distribution tends to become uniform, the probability distribution becomes uniform.

[0073] The content generation module uses the predicted writing intent vector as a conditional control parameter to ensure that the generated subsequent paragraphs or expanded content serve this core intent in both style and content. For example, when the intent is "requesting approval," the generated text will automatically emphasize necessity, compliance, and suggested solutions.

[0074] The compliance verification module is also used to receive the predicted writing intent vector and perform pre-verification based on it. For example, for documents requesting approval, the system will prioritize and rigorously verify whether the cited policy provisions are up-to-date, whether the system has the authority to approve them, and whether the attached materials are typically required. The writing assistance module is also used to receive the predicted writing intent vector and provide forward-looking prompts in the sidebar based on the writing intent vector (such as "You are writing a request, which usually requires the attachment of 'XX document'. Please confirm whether it is prepared"), thus realizing the transition from error warnings to writing guidance.

[0075] This technical solution upgrades the system from a passive text completion tool to an active intelligent writing navigation, solving the problem that the generated content may be out of touch with the user's actual goals.

[0076] Example 5 The current method of adjusting the generation parameters to control the style is rather crude. It may only adjust the formality of the vocabulary, making it difficult to generate texts that truly reflect the deep differences in writing styles between different departments.

[0077] In this embodiment, the document extraction module includes a variational autoencoder (VAE) model, which decouples the semantic content and writing style of the sample document in the latent space, obtaining separate content encoding vectors and style encoding vectors. Through training with a large amount of document data from different sources, the model is ensured to learn to separate the core content of the document from the solemn style of higher-level authorities or the friendly style of grassroots units.

[0078] Variational autoencoder (VAE) models consist of an encoder, a latent space, and a decoder. The encoder maps the input document text to the latent space, generating content-encoded vectors and style-encoded vectors; the decoder regenerates the text based on the vectors in the latent space.

[0079] During training, the VAE model learns from a large amount of official document data and acquires the ability to decouple the semantic content and writing style of official document texts in the latent space.

[0080] The encoder outputs two vectors: a content encoding vector, representing the core content of the document; and a style encoding vector, representing the writing style of the document.

[0081] In the content generation module, after receiving the content encoding vector, a content-neutral draft is first generated. The user selects the target style (e.g., the style of a community neighborhood committee or the style of the municipal party committee office), and the content generation module calls the corresponding style encoding vector from the knowledge base.

[0082] The content generation module includes a Generative Adversarial Network (GAN). The generator fuses the draft text with the style encoding vectors of the target style to generate text in the target style. The discriminator determines whether the generated text truly conforms to the target style. Through adversarial training, the generator can ultimately produce text that maintains content accuracy while also possessing different styles.

[0083] In addition, users can input or import sample documents, and the system will extract the style vector of the sample document in real time and apply it to the current writing content, achieving one-click style transfer. This technical solution meets users' deep needs for refined and personalized control over official document style.

[0084] This application embodiment also provides an electronic device 500 that utilizes the aforementioned method for generating and verifying official document content based on a large language model. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the aforementioned method for generating and verifying official document content based on a large language model. In this application embodiment, the processor is the control center of the computer method and can be a physical machine processor or a virtual machine processor.

[0085] Reference Figure 2The electronic device 500 includes at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one bus 504. The bus 504 is used for communication between these components, the communication interface 502 is used for signaling or data communication with other node devices, and the memory 503 stores machine-readable instructions executable by the processor 501. When the electronic device 500 is running, the processor 501 communicates with the memory 503 via the bus 504. When the machine-readable instructions are invoked by the processor 501, the steps of the document content generation and compliance verification method based on the large language model described above are executed.

[0086] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can implement the steps of the above-described method and apparatus for generating and verifying official document content based on a large language model.

[0087] Those skilled in the art will understand that implementing all or part of the processes in the document content generation and compliance verification method apparatus based on a large language model can be accomplished by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of various embodiments of the document content generation and compliance verification method apparatus based on a large language model. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0088] The above content is merely an embodiment of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can improve and implement this solution based on the guidance provided in this application and their own capabilities. Some typical well-known structures or systems should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A document content generation and compliance verification system based on a large language model, characterized in that, include: The document extraction module is used to analyze uploaded sample documents through a large language model, extract the writing style and fixed sentence patterns of different document types, construct a knowledge graph of document writing, and store the format specifications, common expressions and logical structures of various documents. The document extraction module mines and constructs a core intent library for different document types, which is used to store typical intent prototype vectors in the field of document writing, with each intent corresponding to a vector representation. The content generation module is used to generate subsequent paragraphs based on the context of the input content; expand the text based on keywords; and analyze behavioral data, using a lightweight neural network model to predict the writing intent vector, which serves as a conditional control parameter. The compliance verification module is used to provide real-time alerts for inappropriate wording; it also uses RAG technology to retrieve relevant policies and verify the logical consistency of the content. Automatically verify whether the format of each element of the official document conforms to the standard; perform pre-verification based on the writing intent vector; The writing assistance module provides real-time editing suggestions, which are displayed as annotations in the document's sidebar. Proactive suggestions based on writing intent vectors.

2. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, Analyze behavioral data, including the following: The formula for extracting text input features is as follows: in, For word frequency, For document frequency, Total number of documents; Cursor behavior encoding, using Fourier transform to reduce the dimensionality of the trajectory sequence: in, and Indicates the first n The cursor's two-dimensional coordinates in the document interface when sampling points are reached. n It is the summation index, representing the sequence number of the trajectory sampling point; This indicates the total number of trajectory sampling points within the sliding window; Indicates the frequency component index; Indicates the first Complex values ​​of each frequency component; i It is the imaginary unit.

3. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The document extraction module inputs the uploaded sample documents into a pre-trained large language model; identifies the document type through a text classification algorithm; analyzes the writing style of the documents, including sentence structure, word usage habits, and paragraph layout, and extracts fixed sentence patterns and style features; and stores the extracted document type and writing style features in a database for subsequent template construction.

4. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The document extraction module integrates document templates and format specifications from multiple channels; Natural language processing technology is used to analyze the semantic relationships and structural information in the sample documents; based on the Neo4j graph database, the analyzed knowledge is stored as nodes and edges to form a knowledge graph for official document writing, where nodes represent document elements and edges represent the relationships between elements; the knowledge graph is updated regularly to incorporate the latest document formats and expressions.

5. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The content generation module is used to input the already entered content into the large language model for encoding and analysis; the model understands the current writing context by analyzing the semantics and structure of the already entered content. Based on contextual understanding, the model generates relevant suggestions for subsequent paragraphs; the predicted paragraphs are then presented to the user in a list format, and the user can choose to accept or reject the suggestions.

6. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The content generation module is also used to extract keywords from the input; retrieve information related to the keywords from knowledge graphs and external policy databases; integrate the retrieved information into supplementary content by combining the text generation capabilities of the large language model; and display the supplementary content to users for reference and editing.

7. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The content generation module controls the style of the generated text by adjusting the generation parameters of the large language model; based on the adjusted parameters, the model generates multiple versions of the official document content; and the multiple versions of the official document content are displayed side by side.

8. The document content generation and compliance verification system based on a large language model according to claim 7, characterized in that, The document extraction module includes a variational autoencoder model, which decouples the semantic content and writing style of the sample document in the latent space to obtain separate content encoding vectors and style encoding vectors. The content generation module generates a content-neutral draft based on the content encoding vector; it also includes a generative adversarial network, which is used to fuse the draft with the style encoding vector to generate target style text, and a discriminator to determine whether the generated text truly conforms to the target style.

9. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The compliance verification module stores the sensitive word library in the database; performs real-time detection of official document content using string matching algorithms; employs a multi-pattern matching algorithm to scan the official document text character by character, identifying precisely matched sensitive words, and expands the scope of sensitive word detection based on synonym relationships in a knowledge graph; converts the official document text and sensitive words into semantic vectors respectively, calculates similarity, and identifies semantically similar sensitive expressions; when official document content is input, it detects and highlights sensitive words in real time, and pops up an alert box in the sidebar displaying the sensitivity level, definition of the sensitive word, and provides differentiated suggestions for different levels of sensitive words.

10. The document content generation and compliance verification system based on a large language model according to claim 1, characterized in that, The writing assistance module identifies questionable content by setting rules and thresholds, and then marks and highlights it.

Citation Information

Patent Citations

  • Manuscript writing auxiliary method and system, terminal and storage medium

    CN116432611A

  • Automatic news generation system based on intelligent writing

    CN117094291A

  • Intelligent power work summary generation method and device based on big language model retrieval enhancement generation and storage medium

    CN118551738A

  • Personalized complex report generation method based on multi-agent system

    CN118569237A

  • Intelligent creation method based on NLP intention recognition

    CN119047468A

Cited By

  • Editing enhancement system and method for document structure recognition and content suggestion generation

    CN121303068A

  • UI interface automatic generation method and system based on natural language description

    CN121579007A

  • Multi-engine collaborative official document checking method and system and related equipment

    CN122174828A