Structured document generation method and device, electronic equipment and storage medium
By leveraging semantic understanding and document generation models based on deep learning technology, structured documents can be generated automatically, solving the problems of low efficiency and unstable quality in existing technologies, and achieving efficient and accurate generation of structured documents.
Patent Information
- Application Number
- CN202510870794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies suffer from low efficiency and unstable quality in generating structured documents, and their rule engines lack flexibility, making it difficult to convert unstructured documents to structured ones.
Employing deep learning technology, this system automatically generates structured documents through semantic understanding and document generation models, including preprocessing, semantic parsing, and document completion. The model is optimized using a multi-task learning and distributed training framework.
It significantly improves the efficiency and accuracy of document writing, solves the problems of low efficiency and poor consistency of traditional manual methods, and achieves the generation of structured documents with good structural standardization and content accuracy.
Smart Images

Figure CN120850995A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, electronic device, and storage medium for generating structured documents. Background Technology
[0002] In existing technologies, the generation of structured documents (such as technical reports, patent texts, product specifications, etc.) mainly relies on manual writing or filling in simple templates. This approach has significant efficiency bottlenecks and quality problems.
[0003] Highly reliant on human experience and time: Writers need a deep understanding of the technical content and manually organize it into a structured document that conforms to a specific framework (such as chapters, entries, and formatting requirements). This process is tedious and time-consuming, especially for complex content or large documents, resulting in low efficiency.
[0004] Low standardization and prone to errors: Manual writing makes it difficult to ensure a high degree of consistency in structure, terminology and expression between different documents or between different parts of the same document, which can easily introduce subjectivity and oversights, resulting in unstable document quality.
[0005] Limitations of rule engines: Some automation tools employ template-based mechanisms using keywords or fixed rules. These methods have strict requirements on the format and expression of the input text and lack deep semantic understanding capabilities. When the technical description is complex, diverse in expression, or contains implicit logic, rule engines struggle to accurately extract key information and adapt it to the correct document structure, resulting in poor generation quality and insufficient flexibility.
[0006] The transformation from unstructured to structured text is challenging: Automatically, accurately, and efficiently converting raw, unstructured technical descriptions into the structured content required by the target framework (such as extracting specific technical features, relationships, and functional effects) is a common problem faced by existing technologies. Current methods typically require significant manual intervention and post-processing proofreading. Summary of the Invention
[0007] The main objective of this invention is to provide a method, apparatus, electronic device, and storage medium for generating structured documents, in order to solve at least one problem in the prior art. This invention can efficiently generate structured documents.
[0008] To achieve the above objectives, one aspect of this invention proposes a method for generating structured documents, the method comprising:
[0009] The technical description content is obtained, and the technical description content is preprocessed to obtain the target text;
[0010] The target text is semantically parsed based on a semantic understanding model to obtain technical features; the semantic understanding model is built based on a pre-trained language model.
[0011] Based on technical characteristics, a document generation model is used to generate structured text content; the document generation model is trained based on a deep learning model.
[0012] The structured text content is populated into the target document frame to obtain the target structured document.
[0013] In some embodiments, the technical description content is preprocessed, including the following steps:
[0014] The technical description content is preprocessed based on preset natural language processing technology;
[0015] Natural language processing techniques include text cleaning, text segmentation, stop word filtering, lemmatization, and feature standardization.
[0016] In some embodiments, the method further includes the following steps:
[0017] Multi-task learning is used to pre-train a pre-defined language model, thereby constructing a semantic understanding model;
[0018] Among them, the tasks learned in multi-task learning include entity recognition, relation extraction, and technology classification.
[0019] In some embodiments, the method further includes the following steps:
[0020] By using structured documents with pre-extracted core content as training data, a distributed training framework is used to train a pre-defined deep learning model, thus constructing a document generation model.
[0021] In some embodiments, when the structured document is a patent document, using a structured document with pre-extracted core content as training data includes the following steps:
[0022] Retrieve multiple patent documents from patent databases;
[0023] Collect the central ideas of each section of the patent document; the core content includes the central ideas of each section.
[0024] Sample pairs are constructed as training data based on the central idea and its corresponding sections.
[0025] The central idea serves as the input data, while the section content serves as the annotation content.
[0026] In some embodiments, a distributed training framework is used to train a pre-defined deep learning model to construct a document generation model, including the following steps:
[0027] Configure deep learning models based on attention-based sequence-to-sequence models;
[0028] The core content is used as input data to a deep learning model, which is then processed to generate training documents.
[0029] A similarity loss function is constructed based on the structured documents corresponding to the core content generated during training, and the model parameters of the deep learning model are optimized and adjusted based on the similarity loss function.
[0030] Increment the training round by 1, return to execute the step of inputting the core content as input data into the deep learning model until the preset training conditions are met, and use the deep learning model after the last optimization and adjustment as the document generation model.
[0031] The initial number of training rounds is 0; the preset training conditions include the number of training rounds meeting the preset number of iteration rounds or the similarity loss function meeting the preset accuracy threshold.
[0032] In some embodiments, the method further includes the following steps:
[0033] In response to the document specification structure input from the target object, the target document framework is constructed.
[0034] The document specification structure includes at least one type of structured document structure or at least one region of structured document structure.
[0035] To achieve the above objectives, another aspect of the present invention provides a structured document generation apparatus, the apparatus comprising:
[0036] The first module is used to acquire technical description content, preprocess the technical description content, and obtain the target text;
[0037] The second module is used to perform semantic parsing on the target text based on the semantic understanding model to obtain technical features; the semantic understanding model is built based on a pre-trained language model.
[0038] The third module is used to generate structured text content based on technical features using a document generation model; the document generation model is trained based on a deep learning model.
[0039] The fourth module is used to populate the target document frame with structured text content to obtain the target structured document.
[0040] In some embodiments, the apparatus further includes:
[0041] The fifth module is used to pre-train a pre-defined language model using multi-task learning, thereby constructing a semantic understanding model;
[0042] Among them, the tasks learned in multi-task learning include entity recognition, relation extraction, and technology classification.
[0043] In some embodiments, the apparatus further includes:
[0044] The sixth module is used to train a pre-defined deep learning model using structured documents with pre-extracted core content as training data, and to build a document generation model using a distributed training framework.
[0045] In some embodiments, the apparatus further includes:
[0046] The seventh module is used to construct the target document framework in response to the document specification structure input by the target object;
[0047] The document specification structure includes at least one type of structured document structure or at least one region of structured document structure.
[0048] To achieve the above objectives, another aspect of the present invention provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0049] To achieve the above objectives, another aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0050] The embodiments of this invention include at least the following beneficial effects: This invention provides a method, apparatus, electronic device, and storage medium for generating structured documents. This solution obtains technical description content, preprocesses the technical description content to obtain target text, performs semantic parsing on the target text based on a semantic understanding model to obtain technical features; wherein, the semantic understanding model is constructed based on a pre-trained language model; based on the technical features, a document generation model is used to generate structured text content; wherein, the document generation model is trained based on a deep learning model; and the structured text content is filled into the target document framework to obtain the target structured document. This invention, by introducing deep learning technology, achieves intelligent and automated generation of structured documents. Specifically, it can significantly improve document writing efficiency, effectively solve the problems of low efficiency and poor consistency in traditional manual methods, as well as the problems of insufficient flexibility and weak semantic understanding ability of rule engine methods. It has a high degree of automation, and the generated documents have good structural standardization and content accuracy. Attached Figure Description
[0051] Figure 1 This is a flowchart of a structured document generation method provided in an embodiment of the present invention;
[0052] Figure 2 This is a simplified architectural diagram of an intelligent structured document generation system based on a structured document generation method provided in an embodiment of the present invention.
[0053] Figure 3 A schematic diagram of the model training process provided in this embodiment of the invention;
[0054] Figure 4 This is a schematic diagram of the document generation sequence provided in an embodiment of the present invention;
[0055] Figure 5 A schematic diagram illustrating the application process of the structured document generation method provided in this embodiment of the invention;
[0056] Figure 6 A schematic diagram of the structured document generation device provided in this embodiment of the invention;
[0057] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this invention; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this invention as detailed in the appended claims.
[0059] It is understood that the terms “first,” “second,” etc., used in this invention may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. For example, first information may also be referred to as second information without departing from the scope of embodiments of the invention, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to determination” as used herein may be interpreted as “when…” or “when…” or “in response to determination.”
[0060] The terms “at least one,” “multiple,” “each,” “any,” etc., used in this invention, “at least one” includes one, two, or more than two; “multiple” includes two or more than two; “each” refers to each of the corresponding multiple; and “any” refers to any one of the multiple.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0062] The structured document generation method provided in this invention relates to the field of data processing technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the structured document generation method, but is not limited to the above forms.
[0063] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0064] Figure 1 This is an optional flowchart of the structured document generation method provided in the embodiments of the present invention. Figure 1 The method may include but is not limited to steps S100 to S400.
[0065] S100. Obtain the technical description content, preprocess the technical description content, and obtain the target text;
[0066] It should be noted that, in some embodiments, preprocessing the technical description content may include the following steps: preprocessing the technical description content based on preset natural language processing techniques; wherein, natural language processing techniques include text cleaning, text segmentation, stop word filtering, lemmatization, and feature standardization.
[0067] For example, in some specific implementations, the technical description text input by the user is first received and preprocessed (e.g., word segmentation, noise reduction, entity recognition). Specifically, preprocessing is a commonly used method in the field of NLP (Natural Language Processing), and can be performed through the following operations: 1. Text cleaning (e.g., noise removal, format standardization); 2. Text segmentation (e.g., sentence segmentation, word segmentation); 3. Stop word filtering (removing high-frequency but low-information words); 4. Lexical reconstruction (restoring words of different forms to their basic forms); 5. Feature standardization (e.g., vectorization, normalization, etc.).
[0068] S200. Based on the semantic understanding model, perform semantic parsing on the target text to obtain technical features;
[0069] The semantic understanding model is constructed based on a pre-trained language model.
[0070] It should be noted that in some embodiments, the method may further include the following steps: pre-training a preset language model using multi-task learning, thereby constructing a semantic understanding model; wherein, the tasks learned in multi-task learning include entity recognition tasks, relation extraction tasks, and technology classification tasks.
[0071] For example, in some specific implementations, the semantic understanding model can adopt a pre-trained language model with multi-task learning to simultaneously complete entity recognition, relation extraction, and technology classification tasks; specifically, it can perform semantic parsing on the input text based on a pre-trained language model (such as BERT, GPT) to extract technical features and key information.
[0072] S300: Based on technical features, generate structured text content using a document generation model;
[0073] The document generation model is trained based on a deep learning model.
[0074] It should be noted that in some embodiments, the method may further include the following steps: using structured documents with pre-extracted core content as training data, training a preset deep learning model using a distributed training framework, and constructing a document generation model.
[0075] For example, in some specific implementations, massive amounts of structured documents are collected from publicly available document databases as training data, and key sections (referring to various sections within the structured documents, such as the claims, description, and abstract sections in a patent document, or the project overview, objectives, background, scope, schedule, budget, and risk assessment sections in a project document) are manually annotated. Furthermore, a distributed training framework (such as TensorFlow or PyTorch) can be used to jointly train the document generation model.
[0076] It should be noted that the technical solution of the present invention is applicable to structured documents composed of multiple versions, including but not limited to the aforementioned patent documents and project proposals; specifically, a document generation model can be trained based on structured documents with pre-extracted core content, and a target document framework can be pre-constructed based on the composition structure of the sections of the structured document.
[0077] In some embodiments, when the structured document is a patent document, using the structured document with pre-extracted core content as training data may include the following steps: obtaining multiple patent documents from a patent database; collecting the central idea content of each section in the patent document; wherein the core content includes the central idea content of each section; constructing sample pairs as training data based on the central idea content and its corresponding section; wherein the central idea content is used as input data, and the section is used as labeled content.
[0078] For example, in some specific implementations, the central idea of each section in a structured document can be extracted through semantic parsing using natural language processing technology, or it can be obtained in response to the input content of the target object (i.e., manual annotation). Specifically, taking a patent document as an example, the manually annotated content can be as follows: After reading a complete patent, the central idea is extracted and input. This central idea and the annotated content (i.e., the section where the central idea is extracted, as claimed) constitute a sample pair.
[0079] In some embodiments, a distributed training framework is used to train a pre-defined deep learning model to construct a document generation model. This may include the following steps: configuring the deep learning model based on an attention-based sequence-to-sequence model; inputting core content as input data into the deep learning model and processing it to generate a trained document; constructing a similarity loss function based on the structured document corresponding to the core content and the trained document; optimizing and adjusting the model parameters of the deep learning model based on the similarity loss function; incrementing the training round by 1 and returning to the step of inputting the core content as input data into the deep learning model until the pre-defined training conditions are met; and using the last optimized and adjusted deep learning model as the document generation model. The initial training round is 0. The pre-defined training conditions include the training rounds meeting a pre-defined iteration round or the similarity loss function meeting a pre-defined accuracy threshold.
[0080] Specifically, in addition to sequence-to-sequence models, deep learning models can also employ other encoder-decoder structure sequence generation models.
[0081] For example, in some specific implementations, the training of the document generation model can be achieved as follows:
[0082] 1. Model initialization:
[0083] Choose and configure an attention-based sequence-to-sequence (Seq2Seq) model as the infrastructure (e.g., Transformer or a variant thereof). This model includes encoder and decoder structures.
[0084] The model parameters are randomly initialized or pre-trained weights are loaded.
[0085] The training round counter is initialized to 0.
[0086] 2. Data preparation and distributed loading:
[0087] Prepare a large number of training sample pairs: each sample contains the "core content" of each section of a structured document (input data X) and its corresponding, well-written "complete structured document" (target data Y, including the content of each section).
[0088] By utilizing distributed training frameworks (such as PyTorch DDP, TensorFlow MirroredStrategy, Horovod, etc.), the training dataset is split and distributed to multiple computing nodes (such as GPUs or TPUs) to achieve parallel loading and processing of data.
[0089] 3. Model training iteration:
[0090] 3.1 Forward Propagation: In the current training round, each computing node processes the batches of data assigned to it in parallel.
[0091] A batch of "core content" (X) is input into the encoder of the model.
[0092] The encoder encodes the input sequence into a sequence of context vectors containing semantic information.
[0093] The decoder (combined with an attention mechanism) generates a “training-generated document” (Y_pred) step by step based on the encoder’s output and its own state, which is usually a sequence of words.
[0094] 3.2 Loss Calculation:
[0095] The document Y_pred generated by the model is compared with the content of the corresponding section of the real structured document Y.
[0096] Construct a similarity loss function (L): This function quantifies the difference between Y_pred and Y. It can be:
[0097] Loss forms based on word overlap metrics (such as BLEU, ROUGE).
[0098] Semantic similarity loss based on the model’s internal representation (e.g., using another pre-trained model to compute the cosine similarity loss of sentence / paragraph embeddings).
[0099] Standard sequence generation loss (such as cross-entropy loss) can also be regarded as a "similarity" measure (the goal is to make the predicted sequence as consistent as possible with the real sequence).
[0100] 3.3 Backpropagation and Parameter Optimization (Distributed Synchronization):
[0101] At each computation node, the gradient of the loss L with respect to the model parameters is calculated.
[0102] Key distributed step: The distributed framework aggregates (AllReduce) the gradients across all computing nodes (e.g., calculates the average gradient).
[0103] Optimizers (such as Adam) use aggregated global gradients to synchronously update model parameters across all computing nodes, ensuring model consistency. This is the core advantage of distributed training, accelerating the training of large-scale models.
[0104] 3.4 Round Update and Judgment:
[0105] Increment the training round counter by 1.
[0106] Determine if the preset training conditions are met:
[0107] Condition 1: The number of training rounds reaches the preset maximum number of iterations (e.g., 100 rounds).
[0108] Condition 2: The calculated similarity loss L value is lower than the preset accuracy threshold (e.g., loss < 0.01).
[0109] If any of the following conditions are met: stop training and save the currently optimized model as the final document generation model.
[0110] If not satisfied: Return to step 3 to begin the next training round.
[0111] 4. Model Application (Generation Phase):
[0112] Input the "core content" of new, unseen structured documents (i.e., the technical features obtained by semantic parsing based on the semantic understanding model) into the trained document generation model.
[0113] The model automatically generates semantically relevant structured documents through an encoder-decoder structure (utilizing an attention mechanism).
[0114] S400. Fill the target document frame with the structured text content to obtain the target structured document;
[0115] It should be noted that, in some embodiments, the method may further include the following steps: in response to the document specification structure input by the target object, constructing a target document framework; wherein, the document specification structure includes at least one type of structured document structure or at least one region of structured document structure.
[0116] For example, in some implementations, a document framework can be dynamically generated based on the standard structure of a structured document (e.g., the claims, description, and abstract of a patent document; in some optional implementations, it can be further subdivided into smaller sections based on larger sections, such as the background art, invention content, and specific implementation methods in a patent document; specifically, during the model training phase, the core content is extracted for training based on the differentiated sections). Specifically, in some optional implementations, when there are significant regional differences in the structured document, the target document framework can be constructed based on the regional structured document structure as the standard structure. Taking a patent document as an example, it can adapt to the structured document structure of different countries, regions, or specific requirements by responding to the user-preset document standard structure. Finally, by filling the document framework with the structured text content generated by the document generation model, a complete and coherent structured document is generated.
[0117] In addition, some alternative implementations may also use rule engines and deep learning models to perform logical verification, terminology correction, and format optimization on the generated documents.
[0118] To explain in detail the principle of the technical solution of the present invention, the overall process of the present invention will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0119] First, it's important to note that traditional structured document writing relies heavily on manual labor, which is not only inefficient but also prone to inconsistent quality due to technical terminology or formatting issues. Existing automated tools are mostly based on template matching or simple rules, which cannot meet the needs of generating complex structured documents. Furthermore, the following technical problems in the traditional structured document writing process urgently need to be addressed:
[0120] 1. Low efficiency of manual drafting: Structured documents involve a large amount of professional knowledge, such as technical details, format specifications and legal norms; for example, in drafting patent documents, the drafters need to spend a lot of time on technical analysis, claim construction and language organization, resulting in low overall efficiency and difficulty in meeting the rapidly growing demand for patent applications.
[0121] 2. Insufficient accuracy in technical terminology and legal expression: Structured documents with high technical requirements have extremely high requirements for the accuracy of terminology and / or legal compliance. Manual writing is prone to misuse of terminology, ambiguity in expression, or legal loopholes due to insufficient professional knowledge or negligence.
[0122] 3. Difficulty in standardizing formats and specifications: Different countries or regions may have different regulations on document formats. Manual proofreading and adjustment are prone to format errors or omissions, increasing the cost of correction.
[0123] 4. Limited ability to integrate cross-disciplinary technologies: Modern, highly specialized structured documents often involve interdisciplinary technologies (such as the combination of artificial intelligence and biomedicine). Traditional methods rely on experts in a single field, making it difficult to efficiently integrate complex technical content and generate logically coherent documents.
[0124] 5. Low utilization rate of historical patent data: The existing system lacks the ability to deeply mine massive amounts of historical document data and cannot automatically associate similar technical solutions to optimize document layout and content.
[0125] The development of deep learning technology has provided new possibilities for the intelligent generation of structured documents. Therefore, this invention aims to construct an intelligent structured document generation system using deep learning technology to solve the aforementioned problems in an automated manner, thereby improving the efficiency, accuracy, and compliance of patent drafting. The following uses patent documents as an example to illustrate the application principles of the technical solutions in this invention from multiple aspects:
[0126] I. System Architecture:
[0127] The embodiments of the present invention can be implemented through a system architecture including the following modules:
[0128] 1. Input Processing Module: Receives technical description text input by the user and performs preprocessing (such as word segmentation, noise reduction, and entity recognition). In addition to text, it can also extract technical features from unstructured data such as charts and formulas and integrate them into the document.
[0129] 2. Semantic Understanding Module: Based on pre-trained language models (such as BERT, GPT), semantic parsing is performed on the input text to extract technical features and key information.
[0130] 3. Document Structure Generation Module: Dynamically generates a document framework based on the standard structure of patent documents (claims, description, and abstract). Specifically, embodiments of the present invention can automatically adapt different patent document templates according to the technical field, supporting multiple patent formats (such as USPTO and EPO).
[0131] 4. Content Filling Module: This module utilizes sequence generation models (such as Transformer) to fill extracted technical features into the document framework, generating coherent text content. Specifically, the cross-language model based on the Transformer architecture can automatically convert technical content into bilingual or multilingual languages, ensuring professionalism while meeting the requirements of international patent applications, thus solving the problem of poor adaptability of traditional translation tools in the patent field.
[0132] 5. Optimization and Verification Module: The system utilizes a rule engine and deep learning models to perform logical verification, terminology correction, and format optimization on the generated documents. Furthermore, the system incorporates a patent law knowledge base and an industry terminology database, combined with natural language processing technology, to ensure that the generated documents are logically rigorous, use accurate terminology, and comply with the format requirements of patent offices worldwide, effectively avoiding common formatting errors or descriptive ambiguities encountered in manual drafting.
[0133] In some specific implementations, the system can also collect user modification data and newly promulgated patent specifications in real time through a feedback mechanism, continuously iterate and train the model, so that the generated results continuously conform to the latest examination standards, forming a self-evolving intelligent closed loop.
[0134] In some specific application scenarios, embodiments of the present invention can achieve the following: Figure 2 The intelligent structured document generation system architecture shown includes a three-layer structure: an input layer, a deep learning processing layer, and an output layer.
[0135] II. Core Algorithms and Models:
[0136] 1. Semantic understanding model: Employs a pre-trained language model with multi-task learning to simultaneously perform entity recognition, relation extraction, and technical classification tasks.
[0137] 2. Document generation model: A sequence-to-sequence (Seq2Seq) model based on attention mechanism, fine-tuned with patent corpus to ensure that the generated text conforms to the patent writing style.
[0138] 3. Algorithm optimization: Reinforcement learning is used to iteratively optimize the generated content to improve the accuracy of technical descriptions and legal compliance.
[0139] III. Data Processing Flow:
[0140] 1. Data collection and annotation: Collect a large number of patent documents from publicly available patent databases as training data, and manually annotate the key parts (as claimed).
[0141] 2. Model Training: The document generation model is jointly trained using a distributed training framework (such as TensorFlow or PyTorch). In some specific application scenarios, such as... Figure 3 As shown, the model training process of the deep learning model (i.e., document generation model) in this embodiment of the invention includes three stages: data preprocessing, model training, and validation.
[0142] 3. Online Reasoning: After the user inputs a technical description, the system uses a pre-trained model to generate an initial draft of the patent document in real time, and provides modification suggestions through an interactive interface. For example... Figure 4 As shown, this illustrates the complete time-series process from user input to document generation in some practical application scenarios. Specifically, the server integrates the models applied in the embodiments of this invention (such as semantic understanding models and document generation models).
[0143] In some specific application scenarios, the process of this invention embodiment is applied as follows: Figure 5 As shown, based on the technical description input by the user, it is first preprocessed using a word segmenter and BERT (word segmentation and key information extraction), then a pre-trained language model (i.e., a semantic understanding model) is used to generate technical features, and finally, a document generation model is used to output the patent document; specifically, Figure 5 The relevant structures within the Chinese box are pre-configured and can be applied directly, while the document generation model can be optimized through distributed training iterations before application.
[0144] In summary, this invention automatically analyzes technical features and generates compliant patent documents using a deep learning model, significantly reducing the time spent on manual drafting and revision. This shortens the traditionally lengthy patent drafting cycle to minutes, dramatically improving overall work efficiency. Non-professionals can quickly generate high-quality patent application drafts using this system, reducing reliance on professional agents, which is particularly beneficial for SMEs and individual inventors in lowering intellectual property protection costs. Specifically, through deep learning's ability to analyze massive amounts of authorized patents, the system can intelligently identify the innovative points of technical solutions and propose suggestions for claim optimization, helping users build a more defensive patent protection scope. The synergistic effect of these technologies enables this invention to achieve a paradigm shift from experience-driven to AI-driven approaches in the field of intellectual property services.
[0145] Compared to existing technologies, this invention offers at least the following advantages: 1. Efficiency: It reduces the time required for patent drafting (traditionally several hours) to minutes. 2. Accuracy: Deep learning reduces human error and improves the rigor of technical descriptions. 3. Scalability: The system can adapt to the patent drafting needs of different technical fields and jurisdictions.
[0146] See also Figure 6 This invention also provides a structured document generation apparatus 900, which can implement the above-described structured document generation method. The apparatus may include:
[0147] The first module 901 is used to acquire technical description content, preprocess the technical description content, and obtain target text;
[0148] The second module 902 is used to perform semantic parsing on the target text based on the semantic understanding model to obtain technical features; wherein, the semantic understanding model is constructed based on a pre-trained language model;
[0149] The third module 903 is used to generate structured text content based on technical features using a document generation model; wherein, the document generation model is trained based on a deep learning model;
[0150] Module 4, 904, is used to populate the target document frame with structured text content to obtain the target structured document.
[0151] In some embodiments, the apparatus may further include:
[0152] The fifth module is used to pre-train a pre-defined language model using multi-task learning, thereby constructing a semantic understanding model;
[0153] Among them, the tasks learned in multi-task learning include entity recognition, relation extraction, and technology classification.
[0154] In some embodiments, the apparatus may further include:
[0155] The sixth module is used to train a pre-defined deep learning model using structured documents with pre-extracted core content as training data, and to build a document generation model using a distributed training framework.
[0156] In some embodiments, the apparatus may further include:
[0157] The seventh module is used to construct the target document framework in response to the document specification structure input by the target object;
[0158] The document specification structure includes at least one type of structured document structure or at least one region of structured document structure.
[0159] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0160] This invention also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described structured document generation method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0161] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0162] Please see Figure 7 , Figure 7 The hardware structure of an electronic device 1000 according to another embodiment is illustrated. The electronic device 1000 includes:
[0163] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0164] The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the structured document generation method of the embodiments of this invention.
[0165] Input / output interface 1003 is used to implement information input and output;
[0166] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0167] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0168] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0169] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described structured document generation method.
[0170] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0171] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0172] The structured document generation method, apparatus, electronic device, and storage medium provided in this invention involve: acquiring technical description content; preprocessing the technical description content to obtain target text; performing semantic parsing on the target text based on a semantic understanding model to obtain technical features; wherein the semantic understanding model is constructed based on a pre-trained language model; generating structured text content using a document generation model based on the technical features; wherein the document generation model is trained based on a deep learning model; and filling the structured text content into the target document framework to obtain the target structured document. This invention utilizes deep learning technology for intelligent structured document generation, achieving automated generation of structured documents and effectively improving the efficiency of patent drafting.
[0173] The embodiments described in this invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of this invention, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.
[0174] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0176] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0177] The terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0178] It should be understood that in this invention, "at least one (item)" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0179] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0180] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0181] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0182] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0183] The preferred embodiments of the present invention have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of the present invention should be within the scope of the claims of the present invention.
Claims
1. A method for generating structured documents, characterized in that, The method comprises the following steps: The technical description content is obtained, and the technical description content is preprocessed to obtain the target text; The target text is semantically parsed based on a semantic understanding model to obtain technical features; wherein, the semantic understanding model is constructed based on a pre-trained language model. Based on the aforementioned technical features, a document generation model is used to generate structured text content; wherein, the document generation model is trained based on a deep learning model. The structured text content is then filled into the target document frame to obtain the target structured document.
2. The method according to claim 1, characterized in that, The preprocessing of the technical description content includes the following steps: The technical description content is preprocessed based on preset natural language processing techniques. The natural language processing techniques mentioned include text cleaning, text segmentation, stop word filtering, lemmatization, and feature standardization.
3. The method according to claim 1, characterized in that, The method further includes the following steps: The semantic understanding model is constructed by pre-training a pre-defined language model using multi-task learning. The tasks learned in the multi-task learning include entity recognition, relation extraction, and technology classification.
4. The method according to claim 1, characterized in that, The method further includes the following steps: The document generation model is constructed by using structured documents with pre-extracted core content as training data and training a pre-defined deep learning model using a distributed training framework.
5. The method according to claim 4, characterized in that, When the structured document is a patent document, the step of using the structured document with pre-extracted core content as training data includes the following steps: Retrieve multiple patent documents from patent databases; Collect the core ideas of each section in the patent document; wherein, the core content includes the core ideas of each section. Sample pairs are constructed based on the central idea and its corresponding section as the training data; The central idea is used as input data, and the section is used as annotation content.
6. The method according to claim 4, characterized in that, The process of training a pre-defined deep learning model using a distributed training framework to construct the document generation model includes the following steps: The deep learning model is configured according to an attention-based sequence-to-sequence model; The core content is used as input data and fed into the deep learning model to obtain trained and generated documents. A similarity loss function is constructed based on the structured document corresponding to the core content generated during training, and the model parameters of the deep learning model are optimized and adjusted based on the similarity loss function. Increment the training round by 1, return to the step of inputting the core content as input data into the deep learning model, until the preset training conditions are met, and use the deep learning model after the last optimization and adjustment as the document generation model. The initial number of training rounds is 0; the preset training conditions include the number of training rounds satisfying a preset number of iteration rounds or the similarity loss function satisfying a preset accuracy threshold.
7. The method according to claim 1, characterized in that, The method further includes the following steps: In response to the document specification structure input by the target object, the target document framework is constructed. The document specification structure includes at least one type of structured document structure or at least one region's structured document structure.
8. A structured document generation device, characterized in that, The device includes: The first module is used to acquire technical description content, preprocess the technical description content, and obtain target text; The second module is used to perform semantic parsing on the target text based on a semantic understanding model to obtain technical features; wherein, the semantic understanding model is constructed based on a pre-trained language model; The third module is used to generate structured text content using a document generation model based on the aforementioned technical features; wherein the document generation model is trained based on a deep learning model. The fourth module is used to fill the structured text content into the target document frame to obtain the target structured document.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.