A document generation method and device, electronic equipment and storage medium

By acquiring and parsing document requirement information, determining text writing attributes, and generating document information, the problem of low efficiency in research report generation is solved, and the automatic generation and efficiency improvement of research reports are achieved.

CN121480457BActive Publication Date: 2026-03-24LIANREN HEALTHCARE BIG DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

The low efficiency of research report generation in existing technologies is mainly due to the time-consuming data collection and writing process.

Method used

By acquiring the document requirements information of the target document, analyzing the research type, paragraph topics and input source information, determining the text writing attributes of the sample document, and generating document information based on these attributes, a research report is finally automatically generated.

Benefits of technology

It enables efficient generation of research reports, reduces manual writing time, and improves document generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121480457B_ABST
    Figure CN121480457B_ABST
Patent Text Reader

Abstract

A document generation method and device, electronic equipment and storage medium are disclosed. The method comprises: obtaining document requirement information of a target document; obtaining an analysis result of analyzing the document requirement information, the analysis result comprising a research type of the target document, a paragraph theme corresponding to at least one document paragraph, document requirements and input source information, any document paragraph corresponding to a paragraph type, different paragraph types corresponding to different input source information, and the paragraph type comprising an overview type and an analysis type; determining at least one sample document based on the research type, determining a target text writing attribute of the target document based on a text writing attribute of the at least one sample document; determining document information corresponding to each document paragraph based on the target text writing attribute, the paragraph theme corresponding to each document paragraph, the document requirements and the input source information; and determining the target document based on the target text writing attribute and the document information corresponding to each document paragraph, thereby realizing automatic generation of the document.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular, to a document generation method and device, electronic equipment and storage medium. BACKGROUND

[0002] In a project research project, a research report is used to show the research results of researchers. In the prior art, researchers manually organize data and write research reports. In the process of generating a research report, the data organization and writing process requires a lot of time, which greatly increases the generation time of the research report and reduces the efficiency of generating the research report. SUMMARY

[0003] The present application provides a document generation method, device, electronic equipment and storage medium to automatically generate a document and improve the efficiency of generating the document.

[0004] According to an aspect of the present application, a document generation method is provided, which includes:

[0005] obtaining document requirement information of a target document;

[0006] obtaining an analysis result obtained by analyzing the document requirement information, the analysis result including a research type of the target document, a paragraph theme corresponding to each document paragraph, document requirements and input source information, the research type being used to represent an implementation paradigm of a research activity corresponding to the target document, any document paragraph corresponding to a paragraph type, different paragraph types corresponding to different input source information, and the paragraph type including an overview type and an analysis type;

[0007] determining at least one sample document based on the research type, and determining a target text writing attribute of the target document based on a text writing attribute of the at least one sample document;

[0008] determining document information corresponding to each document paragraph based on the target text writing attribute, the paragraph theme corresponding to each document paragraph, the document requirements and the input source information;

[0009] determining the target document based on the target text writing attribute and the document information corresponding to each document paragraph.

[0010] According to another aspect of the present application, a document generation device is provided, which includes:

[0011] a document requirement information obtaining module configured to obtain document requirement information of a target document;

[0012] The analysis result acquisition module is configured to acquire an analysis result obtained by analyzing the document requirement information, the analysis result comprising a research type of the target document, at least one document paragraph corresponding to a paragraph theme, document requirements and input source information, the research type being used to represent an implementation paradigm of a research activity corresponding to the target document, any document paragraph corresponding to a paragraph type, different paragraph types corresponding to different input source information, and the paragraph type comprising an overview type and an analysis type;

[0013] The target text writing attribute determination module is configured to determine at least one sample document based on the research type, and determine the target text writing attribute of the target document based on a text writing attribute of the at least one sample document.

[0014] The document information generation module is configured to determine document information corresponding to each document paragraph based on the target text writing attribute, the paragraph theme corresponding to each document paragraph, the document requirements and the input source information.

[0015] The target document determination module is configured to determine the target document based on the target text writing attribute and the document information corresponding to each document paragraph.

[0016] According to another aspect of the present application, an electronic device is provided, which comprises:

[0017] at least one processor; and

[0018] a memory in communication connection with the document generation at least one processor; wherein,

[0019] The document generation memory stores a computer program executable by the document generation at least one processor, and the document generation computer program is executed by the document generation at least one processor to enable the document generation at least one processor to execute the document generation method provided by any one of the embodiments of the present application.

[0020] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions, and the document generation computer instructions are used to enable the processor to implement the document generation method provided by any one of the embodiments of the present application when executed.

[0021] The technical solution of this invention provides comprehensive data support for subsequent analysis and processing by acquiring document requirement information of the target document, ensuring the efficient and accurate completion of subsequent tasks; it acquires the parsing results obtained by parsing the document generation document requirement information, including the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The document generation research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document generation paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type, achieving accurate determination of the parsing results and providing accurate data support for the subsequent generation of the target document; based on the document generation research type, at least one sample document is determined, and based on the document generation... By determining the target text writing attributes of at least one sample document, the text writing attributes of the target document to be generated are accurately determined, providing standardized data support for the subsequent generation of the target document. Based on the target text writing attributes of the document generation, the paragraph topics corresponding to each paragraph of the generated document, the document requirements, and the input source information, the document information corresponding to each paragraph of the generated document is determined, achieving accurate generation of document information and providing accurate data support for the generation of the target document. Based on the target text writing attributes of the document generation and the document information corresponding to each paragraph of the generated document, the target document is determined, realizing the automatic generation of the target document. This solves the problems of long generation time and low generation efficiency of manual document generation in the prior art, which is conducive to reducing the time of manual document writing and improving the document generation efficiency.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a document generation method provided in Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of a document generation method provided in Embodiment 2 of the present invention;

[0026] Figure 3This is a flowchart of a document generation method provided in an embodiment of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a document generation device provided in Embodiment 3 of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart of a document generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to the automatic generation of documents. The method can be executed by a document generation device, which can be implemented in hardware and / or software. This document generation device can be configured in the electronic device provided in this embodiment of the present invention. The electronic device can be a server, computer, or mobile terminal, such as a mobile phone or tablet computer. Figure 1 As shown, the method specifically includes the following steps:

[0033] S110. Obtain the document requirement information for the target document.

[0034] This invention is applicable to the automatic generation of documents, including but not limited to research reports and academic documents. This invention uses the generation of research reports in the field of medical research as an example. The target document is the document to be generated. Document requirement information is information used to guide document generation. Optionally, the document requirement information includes at least one of the following: the research object, research activities, research methods, and research purpose of the target document. Document requirement information can be obtained through a document requirement information database. For example, matching the unique identifier of the target document in the document requirement information database yields the document requirement information for the target document. The document requirement information database can store document requirement information corresponding to different target documents.

[0035] Specifically, the unique identifier of the target document is matched against the document requirement information database to obtain the document requirement information of the target document, providing comprehensive data support for subsequent analysis and processing, and ensuring that subsequent tasks are completed efficiently and accurately.

[0036] S120. Obtain the parsing results obtained by parsing the document requirement information. The parsing results include the research type of the target document, the paragraph topics corresponding to at least one document paragraph, the document requirements and input source information. The research type is used to characterize the implementation paradigm of the research activities corresponding to the target document. Each document paragraph corresponds to a paragraph type. Different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type.

[0037] The parsing results are the information obtained by parsing the document requirement information. The parsing results can be determined based on the document requirement information. For example, the document requirement information can be input into a trained parsing model for processing to obtain the parsing results. The parsing model includes, but is not limited to, neural network models; for example, the parsing model could be a large language model. The parsing results include the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The research type refers to the implementation paradigm of the research activity in the document requirement information. Research types include, but are not limited to, prospective research types and retrospective research types. The paragraph topic represents the core content of the document paragraph. The paragraph topic is a high-level summary of the paragraph content. Different document paragraphs correspond to different paragraph topics. Document requirements are the content requirements that the document paragraphs must meet. Input source information refers to the retrieval rules and retrieval locations of the data source used to populate the paragraph content of the document paragraphs. Different paragraph types correspond to different input source information; paragraph types include overview types and analysis types.

[0038] Specifically, the document requirement information is input into the trained parsing model for processing to obtain the parsing results, thus achieving accurate determination of the parsing results and providing accurate data support for the subsequent generation of target documents.

[0039] S130. Determine at least one sample document based on the research type, and determine the target text writing attributes of the target document based on the text writing attributes of the at least one sample document.

[0040] The sample documents are documents that match the research type of the target document. The number of sample documents can be one, three, or five, depending on the requirements. Sample documents can be determined based on the research type of the target document. For example, matching sample documents based on the research type of the target document can be performed. Sample documents can store documents of different research types. The semantic similarity between the research type of the target document and the research types of documents in the document database is calculated. Based on the semantic similarity, the five documents most similar to the research type of the target document are determined, and these five documents are used as sample documents. Text writing attributes are data representing the writing paradigm of documents matching the research type of the target document. Text writing attributes include, but are not limited to, writing style and analysis logic. Different sample documents can correspond to different text writing attributes. Target text writing attributes are data representing the writing paradigm of the target document. Target text writing attributes include, but are not limited to, writing style and analysis logic. Target text writing attributes can be determined based on the text writing attributes of the sample documents. For example, the text writing attributes of five sample documents are input into a trained text writing attribute determination model for processing. The common writing attributes among the text writing attributes of the five sample documents are extracted, and the common writing attributes are used as the target text writing attributes. The text writing attribute determination model includes, but is not limited to, a neural network model.

[0041] Specifically, the model matches the research type of the target document with sample documents. Sample documents can store documents of different research types. The semantic similarity between the research type of the target document and the research type of the documents in the document database is calculated. Based on the semantic similarity, the five documents most similar to the research type of the target document are determined. These five documents are used as sample documents. The text writing attributes of the five sample documents are input into a trained text writing attribute determination model for processing. The common writing attributes among the text writing attributes of the five sample documents are extracted and used as the target text writing attributes. This achieves accurate determination of the target text writing attributes and provides standardized data support for the subsequent generation of target documents.

[0042] S140. Based on the target text writing attributes, the paragraph topics corresponding to each document paragraph, the document requirements, and the input source information, determine the document information corresponding to each document paragraph.

[0043] The document information refers to the document content corresponding to the document paragraph. Document information can be determined based on the target text's writing attributes, paragraph topic, document requirements, and input source information. Taking a single document paragraph as an example, the target text's writing attributes, the paragraph topic, document requirements, and input source information are input into a trained document information generation model for processing to obtain the document information corresponding to that paragraph. The document information generation model includes, but is not limited to, neural network models.

[0044] Specifically, the target text writing attributes, the paragraph topics corresponding to each document paragraph, document requirements, and input source information are input into a trained document information generation model for processing, thereby obtaining the document information corresponding to each document paragraph. This achieves accurate generation of document information and provides accurate data support for the generation of the target document.

[0045] Optionally, the parsing result may also include the research objectives of the target document, which characterize the expected outcomes of the research activities corresponding to the target document. The research objectives are information that characterizes the expected outcomes of the research activities corresponding to the target document. The research objectives can be determined based on the document's requirement information. For example, the document's requirement information can be input into a trained parsing model for processing to obtain the research objectives of the target document. The parsing model may include, but is not limited to, a neural network model; for example, the parsing model could be a large language model.

[0046] Optionally, the paragraph type includes the analysis type; the input source information corresponding to the document paragraph of the analysis type includes the data to be analyzed that is adapted to the document paragraph of the analysis type; the document information corresponding to each document paragraph is determined based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, the document requirements, and the input source information, including: for any document paragraph of the analysis type, determining the data to be analyzed based on the paragraph topic and the document requirements; determining the data description information corresponding to the data to be analyzed based on the data to be analyzed and the target text writing attributes; determining the visualization data based on the data to be analyzed; determining the paragraph conclusion corresponding to the document paragraph based on the research objective, the data to be analyzed, and the target text writing attributes; and determining the second document information corresponding to the document paragraph of the analysis type based on the data description information, the visualization data, and the paragraph conclusion.

[0047] The document paragraphs of the analysis type can be used for data interpretation, argument analysis, and conclusion derivation. The data to be analyzed is the research data used to support the document paragraphs of that analysis type. Different document paragraphs of different analysis types correspond to different data to be analyzed. The data to be analyzed can be determined based on the paragraph's theme and document requirements. For example, for any document paragraph of any analysis type, the paragraph's theme and document requirements are input into a trained data determination model for processing to obtain the data to be analyzed corresponding to that analysis type of document paragraph. The data determination model includes, but is not limited to, neural network models. Data description information is the textual description of the research data used to support the document paragraphs of the analysis type. Data description information can be determined based on the data to be analyzed and the target text's writing attributes. For example, the data to be analyzed and the target text's writing attributes are input into a trained data description information determination model for processing to obtain the data description information corresponding to the data to be analyzed. The data description information determination model includes, but is not limited to, neural network models. Visualized data is the data used to visualize the research data used to support the document paragraphs of the analysis type. Visualized data can be determined based on the data to be analyzed. For example, the data to be analyzed is input into a trained visualization processing model for processing to obtain the visualized data corresponding to the data to be analyzed. The visualization processing model includes, but is not limited to, neural network models. The types of visualized data include, but are not limited to, image types and table types. The type of visualized data can be set according to requirements, and this invention does not impose any limitations. The paragraph conclusion is the textual content representing the viewpoint of a document paragraph of the analysis type. Different analysis types correspond to different paragraph conclusions. The paragraph conclusion can be determined based on the research objective, the data to be analyzed, and the target text writing attributes. For example, the research objective, the data to be analyzed, and the target text writing attributes can be input into a trained paragraph conclusion generation model for processing to obtain the paragraph conclusion corresponding to the document paragraph of that analysis type. Second document information is information representing the complete paragraph content of a document paragraph of the analysis type. Second document information can be determined based on data description information, visualized data, and paragraph conclusion. For example, data description information, visualized data, and paragraph conclusion can be input into a trained second document information generation model for processing to obtain the second document information corresponding to the document paragraph of that analysis type. The second document information generation model includes, but is not limited to, neural network models. Another example is that data description information, visualized data, and paragraph conclusion can be concatenated according to a preset structural order to obtain the second document information.

[0048] Specifically, for any document paragraph of any analysis type, the paragraph topic and document requirements of that analysis type are input into a trained data determination model for processing to obtain the data to be analyzed corresponding to that analysis type of document paragraph. The data to be analyzed and the target text writing attributes are input into a trained data description information determination model for processing to obtain the data description information corresponding to the data to be analyzed. The data to be analyzed is input into a trained visualization processing model for processing to obtain the visualization data corresponding to the data to be analyzed. The research objective, the data to be analyzed, and the target text writing attributes are input into a trained paragraph conclusion generation model for processing to obtain the paragraph conclusion corresponding to that analysis type of document paragraph. The data description information, visualization data, and paragraph conclusion are concatenated according to a preset first concatenation order to obtain the second document information. This achieves accurate and automatic generation of the second document information, providing accurate data support for the generation of the target document, which helps to reduce the generation time of the target document and improve the generation efficiency of the target document.

[0049] S150. Determine the target document based on the target text writing attributes and the document information corresponding to each document paragraph.

[0050] The target document can be determined based on the target text writing attributes and the document information corresponding to each document paragraph. For example, the target text writing attributes and the document information corresponding to each document paragraph can be input into a trained target document generation model for processing to obtain the target document. The target document generation model includes, but is not limited to, neural network models, such as the Dify agent framework.

[0051] Specifically, the target text writing attributes and the document information corresponding to each document paragraph are input into the trained target document generation model for processing to obtain the target document. This achieves automatic generation of target documents, which helps reduce the time spent on manual document writing and improves document generation efficiency.

[0052] Optionally, the target document is determined based on the target text writing attributes and the document information corresponding to each document paragraph, including: determining the target conclusion of the target document based on the research objective, the paragraph conclusions in the second document information, and the target text writing attributes; and determining the target document based on the document information corresponding to each document paragraph and the target conclusion.

[0053] The target conclusion refers to the textual content representing the conclusion of the target document. The target conclusion of the target document can be determined based on the research objective, the paragraph conclusions in the second document information, and the target text writing attributes. For example, the research objective, the paragraph conclusions in the second document information, and the target text writing attributes can be input into a trained target conclusion generation model for processing to obtain the target conclusion of the target document. The target conclusion generation model includes, but is not limited to, neural network models. The target document can also be determined based on the document information and target conclusions corresponding to each document paragraph. For example, the document information and target conclusions corresponding to each document paragraph can be concatenated according to a pre-set second concatenation order to obtain the target document.

[0054] Specifically, the research objective, paragraph conclusions from the second document information, and target text writing attributes are input into a trained target conclusion generation model for processing to obtain the target conclusion of the target document. The document information and target conclusions corresponding to each document paragraph are then concatenated according to a pre-set second concatenation order to obtain the target document. This achieves automatic generation of the target document, which helps reduce the time spent on manual document writing and improves document generation efficiency.

[0055] Optionally, the parsing result may also include the paragraph order of at least one document paragraph; determining the target document based on the document information and target conclusion corresponding to each document paragraph includes: splicing the document information and target conclusion corresponding to each document paragraph based on the paragraph order to obtain the target document.

[0056] The paragraph order refers to the arrangement of the document information corresponding to each paragraph. The target document can also be determined based on the paragraph order, the document information corresponding to each paragraph, and the target conclusion. For example, the target document can be obtained by concatenating the document information corresponding to each paragraph and the target conclusion according to the paragraph order.

[0057] Specifically, the document information and target conclusions corresponding to each paragraph are spliced ​​together according to the paragraph order to obtain the target document, thus realizing the automatic generation of the target document. This helps to reduce the time spent on manual report writing and improve the efficiency of target document generation.

[0058] Based on the above embodiments, the document generation method further includes: performing format verification processing on the target document, wherein the format verification processing includes at least one of the following: font adjustment, line spacing adjustment, page number setting, and visual data numbering setting.

[0059] The technical solution of this embodiment, by acquiring the document requirement information of the target document, provides comprehensive data support for subsequent analysis and processing, ensuring the efficient and accurate completion of subsequent tasks; it acquires the parsing results obtained by parsing the document generation document requirement information. The document generation parsing results include the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The document generation research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document generation paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type, achieving accurate determination of the parsing results and providing accurate data support for the subsequent generation of the target document; based on the determination of the document generation research type... At least one sample document is used to determine the target text writing attributes of the generated target document based on the text writing attributes of the sample document. This ensures accurate determination of the target text writing attributes and provides standardized data support for the subsequent generation of the target document. Furthermore, based on the target text writing attributes, the paragraph topics corresponding to each paragraph in the generated document, document requirements, and input source information, the document information corresponding to each paragraph in the generated document is determined. This ensures accurate generation of document information and provides accurate data support for the generation of the target document. Finally, based on the target text writing attributes and the document information corresponding to each paragraph in the generated document, the target document is determined, enabling automatic generation of the target document. This helps reduce the time spent on manual document writing and improves document generation efficiency.

[0060] Example 2

[0061] Figure 2 This is a flowchart of a document generation method provided in Embodiment 2 of the present invention. This embodiment is a refinement of the above embodiments. Based on the foregoing embodiments, it provides a detailed explanation of how to generate document information corresponding to each document segment based on the target text writing attributes, the paragraph topic corresponding to each document segment, document requirements, and input source information. For specific implementation methods, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here. Figure 2 As shown, the method specifically includes the following steps:

[0062] S210. Obtain the document requirement information for the target document.

[0063] S220. Obtain the parsing results obtained by parsing the document requirement information. The parsing results include the research type of the target document, the paragraph topics corresponding to at least one document paragraph, the document requirements and input source information. The research type is used to characterize the implementation paradigm of the research activities corresponding to the target document. Each document paragraph corresponds to a paragraph type. Different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type.

[0064] S230. Determine at least one sample document based on the research type, and determine the target text writing attributes of the target document based on the text writing attributes of the at least one sample document.

[0065] Optionally, the paragraph type includes an overview type; the input source information for the overview type document paragraph includes information retrieval sources adapted to the overview type document paragraph.

[0066] Overview-type document paragraphs can be used for background introductions and current situation analyses. Information retrieval sources are the sources of information used to obtain the content of overview-type document paragraphs. Information retrieval sources include, but are not limited to, document databases and the internet.

[0067] S240. For any document paragraph of the overview type, search the information retrieval sources based on the paragraph topic to determine multi-source information documents; determine the summary information of the document paragraph based on the multi-source information documents; determine the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, document requirements and summary information.

[0068] In this invention, multi-source information documents are documents that support the paragraph content of document paragraphs of the overview type. The number of multi-source information documents can be one or more, depending on the requirements; this invention does not impose any limitation on the number. Multi-source information documents can be determined through retrieval. For example, a search can be conducted in information retrieval sources based on the paragraph topic of the document paragraph of the overview type to obtain multiple multi-source information documents. Summary information is information obtained by extracting the content of the document supporting the paragraph content of the document paragraph of the overview type. Summary information can be determined based on multi-source information documents. For example, multi-source information documents can be input into a trained summary generation model for processing to obtain summary information for the document paragraph of the overview type. The summary generation model includes, but is not limited to, a neural network model. First document information is information representing the complete paragraph content of the document paragraph of the overview type. First document information can be determined based on the target text writing attributes, document requirements, and summary information. For example, the target text writing attributes, document requirements, and summary information can be input into a trained first document information generation model for processing to obtain the first document information corresponding to the document paragraph of the overview type. The first document information generation model includes, but is not limited to, a neural network model.

[0069] Specifically, based on the paragraph topic of the overview-type document paragraph, a search is conducted in the information retrieval source to obtain multi-source information documents. These multi-source information documents are then input into a trained summary generation model for processing to obtain summary information for the overview-type document paragraph. The target text writing attributes, document requirements, and summary information are then input into a trained first document information generation model for processing to obtain the first document information corresponding to the overview-type document paragraph. This achieves automatic generation of the first document information corresponding to the overview-type document paragraph, thereby improving the generation efficiency of the target document.

[0070] Optionally, after determining the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, document requirements, and summary information, the document generation method further includes: obtaining the document identification information of the multi-source information document, the document identification information being used to characterize the publication channel of the multi-source information document; and annotating the first document information based on the document identification information.

[0071] Document identification information refers to information used to characterize the publication channels of multi-source information documents. Document identification information includes, but is not limited to, title information, publisher, publishing journal, and Digital Object Identifier (DOI). Document identification information for multi-source information documents can be obtained from the multi-source information documents themselves; it can be automatically retrieved when searching for such documents. Based on the document identification information, the first document information can be annotated, with the document identification information being annotated in the summary information corresponding to the multi-source information documents within the first document information.

[0072] Specifically, when retrieving multi-source information documents, the document identification information of the multi-source information documents is automatically obtained. After determining the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, document requirements and summary information, the document identification information is marked at the summary information of the multi-source information document in the first document information, thus realizing the annotation of the first document information and facilitating the management of the first document information and the target document.

[0073] Optionally, after annotating the first document information based on the document identification information, the document generation method further includes: verifying the document identification information and the first document information to obtain a verification result, wherein the verification result indicates whether the document identification information and the first document information are consistent.

[0074] The verification result indicates whether the document identifier information and the first document information are consistent. For example, the verification result can be that the document identifier information and the first document information are consistent. Another example is that the verification result can be that the document identifier information and the first document information are inconsistent. The process of determining the verification result is as follows: The corresponding multi-source information document is determined based on the document identifier information; the semantic similarity between the multi-source information document and the first document information is calculated; when the semantic similarity between the multi-source information document and the first document information is greater than or equal to a preset similarity threshold, the verification result is that the document identifier information and the first document information are consistent; when the semantic similarity between the multi-source information document and the first document information is less than the preset similarity threshold, the verification result is that the document identifier information and the first document information are inconsistent, and the corresponding document identifier information can be removed.

[0075] Specifically, after annotating the first document information based on document identification information, the corresponding multi-source information document is determined according to the document identification information; the semantic similarity between the multi-source information document and the first document information is calculated; when the semantic similarity between the multi-source information document and the first document information is greater than or equal to a preset similarity threshold, the verification result is that the document identification information and the first document information are consistent; when the semantic similarity between the multi-source information document and the first document information is less than the preset similarity threshold, the verification result is that the document identification information and the first document information are not consistent, and the corresponding document identification information can be removed. This realizes the verification of document identification information and the first document information, which helps to improve the accuracy of document identification information annotation.

[0076] S250. Determine the target document based on the target text writing attributes and the document information corresponding to each document paragraph.

[0077] For example, see Figure 3 , Figure 3 This is a flowchart of a document generation method provided in an embodiment of the present invention.

[0078] The technical solution of this embodiment obtains the document requirement information of the target document, providing comprehensive data support for subsequent analysis and processing, ensuring the efficient and accurate completion of subsequent tasks; it obtains the parsing results obtained by parsing the document requirement information, including the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type, achieving accurate determination of the parsing results and providing accurate data support for the subsequent generation of the target document; it determines at least one sample document based on the research type, and determines the text writing attributes of at least one sample document... The system defines the target text writing attributes of the target document, enabling precise determination of these attributes and providing standardized data support for subsequent target document generation. For any overview-type document paragraph, a search is conducted in the information retrieval sources based on the paragraph topic to identify multi-source information documents. The summary information of the document paragraph is then determined based on these multi-source information documents. The first document information corresponding to the overview-type document paragraph is determined based on the target text writing attributes, document requirements, and summary information, enabling the automatic generation of the first document information for overview-type document paragraphs, thus improving the efficiency of target document generation. Finally, the target document is determined based on the target text writing attributes and the document information corresponding to each document paragraph, achieving automatic target document generation and reducing the time spent on manual document writing while improving document generation efficiency.

[0079] Example 3

[0080] Figure 4 This is a schematic diagram of a document generation device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes a document requirement information acquisition module 310, a parsing result acquisition module 320, a target text writing attribute determination module 330, a document information generation module 340, and a target document determination module 350.

[0081] The document requirement information acquisition module 310 is used to acquire the document requirement information of the target document; the parsing result acquisition module 320 is used to acquire the parsing result obtained by parsing the document requirement information. The parsing result includes the research type of the target document, the paragraph topic corresponding to at least one document paragraph, the document requirements and the input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. The paragraph types include overview type and analysis type; the target text writing attribute determination module 330 is used to determine at least one sample document based on the research type and to determine the target text writing attributes of the target document based on the text writing attributes of at least one sample document; the document information generation module 340 is used to determine the document information corresponding to each document paragraph based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, the document requirements and the input source information; and the target document determination module 350 is used to determine the target document based on the target text writing attributes and the document information corresponding to each document paragraph.

[0082] The technical solution of this embodiment obtains the document requirement information of the target document through a document requirement information acquisition module, providing comprehensive data support for subsequent analysis and processing, and ensuring the efficient and accurate completion of subsequent tasks. Through a parsing result acquisition module, it obtains the parsing results obtained from parsing the document requirement information. The parsing results include the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. Paragraph types include overview type and analysis type, achieving accurate determination of the parsing results and providing accurate data support for the subsequent generation of the target document. The module determines the target text writing attributes. The system employs several modules: a module that determines at least one sample document based on the research type, and a module that determines the target text writing attributes of the target document based on the text writing attributes of the sample document. This precise determination of the target text writing attributes provides standardized data support for the subsequent generation of the target document. A document information generation module determines the document information corresponding to each document segment based on the target text writing attributes, the paragraph topics corresponding to each document segment, document requirements, and input source information. This accurate generation of document information provides precise data support for the generation of the target document. Finally, a target document determination module determines the target document based on the target text writing attributes and the document information corresponding to each document segment. This automatic generation of the target document helps reduce the time spent on manual document writing and improves document generation efficiency.

[0083] Based on the above embodiments, optionally, the paragraph type includes an overview type.

[0084] Optionally, the input source information corresponding to the overview type document paragraph includes information retrieval sources adapted to the overview type document paragraph.

[0085] Optionally, the document information generation module 340 is further configured to: for any document paragraph of the overview type, search the information retrieval source based on the paragraph topic to determine a multi-source information document; determine the summary information of the document paragraph based on the multi-source information document; and determine the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, the document requirements, and the summary information.

[0086] Optionally, the document information generation module 340 is further configured to: after determining the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, the document requirements, and the summary information, obtain the document identification information of the multi-source information document, wherein the document identification information is used to characterize the publication channel of the multi-source information document; and annotate the first document information based on the document identification information.

[0087] Optionally, the document information generation module 340 is further configured to: after annotating the first document information based on the document identification information, verify the document identification information and the first document information to obtain a verification result, wherein the verification result indicates whether the document identification information and the first document information are consistent.

[0088] Optionally, the parsing results may also include the research objectives of the target document, which are used to characterize the expected outcomes of the research activities corresponding to the target document.

[0089] Optionally, the paragraph type includes an analysis type.

[0090] Optionally, the input source information corresponding to the document paragraph of the analysis type includes the data to be analyzed that is adapted to the document paragraph of the analysis type.

[0091] Optionally, the document information generation module 340 is further configured to: for any document paragraph of the analysis type, determine the data to be analyzed based on the paragraph topic and the document requirements; determine the data description information corresponding to the data to be analyzed based on the data to be analyzed and the target text writing attributes; determine the visualization data based on the data to be analyzed; determine the paragraph conclusion corresponding to the document paragraph based on the research objective, the data to be analyzed, and the target text writing attributes; and determine the second document information corresponding to the document paragraph of the analysis type based on the data description information, the visualization data, and the paragraph conclusion.

[0092] Optionally, the target document determination module 350 is further configured to: determine the target conclusion of the target document based on the research objective, the paragraph conclusions in the second document information, and the target text writing attributes; and determine the target document based on the document information corresponding to each of the document paragraphs and the target conclusions.

[0093] Optionally, the parsing result may also include the paragraph order of at least one of the document paragraphs.

[0094] Optionally, the target document determination module 350 is further configured to: concatenate the document information corresponding to each of the document paragraphs and the target conclusion based on the paragraph order to obtain the target document.

[0095] The document generation apparatus provided in this embodiment of the invention can execute a document generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0096] Example 4

[0097] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0098] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0099] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0100] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a document generation method.

[0101] In some embodiments, a document generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of a document generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a document generation method by any other suitable means (e.g., by means of firmware).

[0102] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] A computer program for implementing a document generation method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0104] Example 5

[0105] Embodiment 5 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a document generation method, the method comprising:

[0106] Obtain the document requirements information of the target document; obtain the parsing results obtained from parsing the document requirements information, including the research type of the target document, the paragraph topics corresponding to at least one document paragraph, document requirements, and input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type, and different paragraph types correspond to different input source information. Paragraph types include overview type and analysis type; determine at least one sample document based on the research type; determine the target text writing attributes of the target document based on the text writing attributes of at least one sample document; determine the document information corresponding to each document paragraph based on the target text writing attributes, the paragraph topics corresponding to each document paragraph, document requirements, and input source information; determine the target document based on the target text writing attributes and the document information corresponding to each document paragraph.

[0107] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0108] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0109] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0110] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0111] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0112] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A document generation method, characterized in that, include: Obtain the document requirements information for the target document; Obtain the parsing result obtained by parsing the document requirement information. The parsing result includes the research type of the target document, the paragraph topic corresponding to at least one document paragraph, the document requirements and input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type. Different paragraph types correspond to different input source information. The paragraph type includes overview type and analysis type. Based on the research type, at least one sample document is determined, and based on the text writing attributes of the at least one sample document, the target text writing attributes of the target document are determined; Based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, the document requirements, and the input source information, determine the document information corresponding to each document paragraph. The target document is determined based on the target text writing attributes and the document information corresponding to each of the document paragraphs. The paragraph type includes an overview type; the input source information corresponding to the document paragraph of the overview type includes an information retrieval source adapted to the document paragraph of the overview type; the generation of document information corresponding to each document paragraph based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, document requirements, and input source information includes: For any document paragraph of the aforementioned overview type, a search is conducted in the information retrieval source based on the paragraph topic to determine multi-source information documents; Based on the multi-source information document, the summary information of the document paragraph is determined; based on the target text writing attributes, the document requirements, and the summary information, the first document information corresponding to the document paragraph of the overview type is determined.

2. The method according to claim 1, characterized in that, After determining the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, the document requirements, and the summary information, the method further includes: Obtain the document identification information of the multi-source information document, wherein the document identification information is used to characterize the publication channel of the multi-source information document; The first document information is labeled based on the document identification information.

3. The method according to claim 2, characterized in that, After annotating the first document information based on the document identification information, the method further includes: The document identification information and the first document information are verified to obtain a verification result, which indicates whether the document identification information and the first document information are consistent.

4. The method according to claim 1, characterized in that, The parsing results also include the research objectives of the target document, which are used to characterize the expected outcomes of the research activities corresponding to the target document; the paragraph type includes the analysis type; the input source information corresponding to the document paragraph of the analysis type includes the data to be analyzed that is adapted to the document paragraph of the analysis type; The process of generating document information corresponding to each document paragraph based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, document requirements, and input source information includes: For any document paragraph of the analysis type, the data to be analyzed is determined based on the paragraph topic and the document requirements; the data description information corresponding to the data to be analyzed is determined based on the data to be analyzed and the target text writing attributes; the visualization data is determined based on the data to be analyzed; the paragraph conclusion corresponding to the document paragraph is determined based on the research objective, the data to be analyzed, and the target text writing attributes; and the second document information corresponding to the document paragraph of the analysis type is determined based on the data description information, the visualization data, and the paragraph conclusion.

5. The method according to claim 4, characterized in that, Determining the target document based on the target text writing attributes and the document information corresponding to each of the document paragraphs includes: The target conclusion of the target document is determined based on the research objective, the paragraph conclusions in the second document information, and the target text writing attributes. The target document is determined based on the document information corresponding to each of the document paragraphs and the target conclusion.

6. The method according to claim 5, characterized in that, The parsing result also includes the paragraph order of at least one paragraph of the document; The step of determining the target document based on the document information corresponding to each of the document paragraphs and the target conclusion includes: Based on the paragraph order, the document information corresponding to each of the document paragraphs and the target conclusion are concatenated to obtain the target document.

7. A document generation device, characterized in that, include: The document requirement information acquisition module is used to acquire the document requirement information of the target document; The parsing result acquisition module is used to acquire the parsing result obtained by parsing the document requirement information. The parsing result includes the research type of the target document, the paragraph topic corresponding to at least one document paragraph, the document requirements and the input source information. The research type is used to characterize the implementation paradigm of the research activity corresponding to the target document. Each document paragraph corresponds to a paragraph type. Different paragraph types correspond to different input source information. The paragraph type includes overview type and analysis type. The target text writing attribute determination module is used to determine at least one sample document based on the research type, and to determine the target text writing attribute of the target document based on the text writing attribute of the at least one sample document; The document information generation module is used to determine the document information corresponding to each document paragraph based on the target text writing attributes, the paragraph topic corresponding to each document paragraph, the document requirements and the input source information. The target document determination module is used to determine the target document based on the target text writing attributes and the document information corresponding to each of the document paragraphs. The paragraph type includes an overview type; the input source information corresponding to the document paragraph of the overview type includes an information retrieval source adapted to the document paragraph of the overview type; the document information generation module is further configured to: for any document paragraph of the overview type, search the information retrieval source based on the paragraph topic to determine a multi-source information document; determine the summary information of the document paragraph based on the multi-source information document; and determine the first document information corresponding to the document paragraph of the overview type based on the target text writing attributes, the document requirements, and the summary information.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the document generation method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the document generation method of any one of claims 1-6.

Citation Information

Patent Citations

  • Document generation method and device, equipment and medium

    CN117725895A

  • Method, device and equipment for generating document based on generative large model and medium

    CN117992569A

  • Document generation method and device, equipment, storage medium and product

    CN119474362A