A Remote Sensing Long Document Generation Method Combining Multi-Granularity Recall and Multi-Step Reflection

By constructing a long document knowledge structure structure set and expert collaborative model, combining multi-grained RAG technology and remote sensing field base model, the problems of limited generation length, poor quality and logical inconsistent generation in the long document generation in the remote sensing field are solved, and efficient, professional and coherent document generation is achieved.

CN119990075BActive Publication Date: 2025-07-18ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510012366.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-07-18
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the long document generation of existing large language models, there are problems such as limited generation length, poor generation quality, incoherence of logic, and poor consistency of outline and overall content in the field of remote sensing.

Method used

Using a multi-grained recall and multi-step reflection method, we generate and check the consistency of document content by constructing a long document knowledge structure set and expert collaborative model, including Planer, Executor, and Master, combining multi-grained RAG technology and remote sensing field base model.

Benefits of technology

It realizes efficient, professional and coherent generation of long documents in the field of remote sensing, improves the speed and quality of document generation, and ensures the professionalism and logical consistency of the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990075B_ABST
    Figure CN119990075B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating remote sensing long documents by combining multi-granularity recall and multi-step reflection, comprising the following steps: S1, constructing a long document knowledge structure set; S2, constructing an expert collaboration model, including three agents, namely Planer, Executor, and Master; S3, starting the execution plan for the overall writing task by the Planer of the CoE model; S4, parsing the input materials, and if there is a model essay input, parsing the input model essay at the same time; S5, generating a document outline; S6, using the multi-granularity RAG technology to generate the content of each chapter and paragraph; S7, merging the paragraph content to form a complete document; S8, having the Master check the content coherence of the whole text. If incoherent content is found, trigger the modification of the relevant paragraph content until the inspection is qualified and then officially output the final result. The present invention improves the semantic consistency and content coherence in the process of generating ultra-long texts, and is superior to the effect of direct single-time generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing large language models, and particularly to a method for generating remote sensing long documents by combining multi-granularity recall and multi-step reflection. Background Art

[0002] The application of generative AI based on large language models (LLMs) in office automation scenarios has become increasingly mature, and automatic long document generation is one of the typical applications. LLMs can efficiently write long texts of tens of thousands of words based on a large amount of reference knowledge, demonstrating performance beyond traditional methods and providing great convenience to users. However, in the generation of long documents in specific fields, especially in the remote sensing field with strong professionalism and knowledge intensity, existing large models still face many challenges.

[0003] Firstly, existing large models are often limited by the generation length when generating long documents. On the one hand, this is related to the architecture design of LLM series models. Although the Transformer architecture has advantages in processing long sequence data, its computational complexity increases quadratically with the sequence length, making it difficult to support the efficient generation of ultra-long texts of more than ten thousand words in practical applications. On the other hand, the "memory decay" phenomenon is likely to occur during the long text generation process, that is, the model's memory of early input information gradually weakens, affecting the coherence and logic of the generated content. In addition, in order to ensure the generation efficiency, a maximum generation length threshold is set in many application scenarios, further restricting the length of the document.

[0004] Secondly, even within an acceptable length range, there is still room for improvement in the text quality generated by large models. Specifically, it is manifested in the following aspects:

[0005] 1) Factual errors: In professional fields that require high accuracy, such as the remote sensing field where there are higher requirements for the professionalism of text content, the text generated by the model may contain professional knowledge errors, affecting the authority and credibility of the document.

[0006] 2) Logical incoherence: Long documents usually contain multiple chapters or paragraphs, requiring good connection between each part of the content. However, in the actual generation process, the logical relationship between different parts may not be close enough, or even contradictory, reducing the reading experience.

[0007] 3) How to generate full-text content that conforms to the outline consistency according to the document outline structure specified by the user is another problem that needs to be solved urgently. Ideally, the user can first provide a detailed outline, including information such as the topics and key points of each chapter, and then the model automatically generates a complete document according to the outline. However, it is still difficult to fully achieve this goal at the current technical level. On the one hand, the model needs to have a strong understanding ability to accurately parse the instruction information in the outline; on the other hand, it also needs to have excellent planning ability to ensure that the generated content strictly follows the outline framework while maintaining the richness and creativity of the content. In fact, although many generated results generally meet the outline requirements, the details are often unsatisfactory, such as deviating from the theme and omitting important information.

[0008] In summary, although the technology of generating long domain documents based on large models has made certain progress, there are still problems such as limited generation length, poor generation quality, and poor consistency between the outline and the overall content. These problems not only restrict the application scope of the technology but also point out the direction for further research. The improvement direction focuses on optimizing the model architecture, improving the generation quality, and enhancing the ability to understand and execute the consistency of complex long document generation tasks, so as to promote the application in this field to develop towards a more mature and stable direction. Summary of the Invention

[0009] To solve the existing problems, the present invention provides a method for generating remote sensing long documents by combining multi-granularity recall and multi-step reflection. The specific solution is as follows:

[0010] A method for generating remote sensing long documents by combining multi-granularity recall and multi-step reflection includes the following steps:

[0011] S1, constructing a long document knowledge structure set;

[0012] S2, constructing an expert collaboration model, that is, a CoE model, including three agents, namely Planer, Executor, and Master. Among them, Planer is responsible for overall planning, Executor is responsible for execution, and Master, the supervisor agent, is responsible for result summary and verification;

[0013] S3, after the user inputs a writing request, Planer of the CoE model starts to execute the overall writing task planning;

[0014] S4, parsing the input materials. If a sample essay is input, the input sample essay is also parsed;

[0015] S5, generating a document outline based on the parsed materials and sample essays;

[0016] S6. Based on the outline, the system adopts multi-granularity RAG technology to extract relevant information from the knowledge base and the ultra-long context, and generate the content of each chapter and paragraph.

[0017] S7. Merge the generated content of each chapter and paragraph to form a complete document.

[0018] S8. After generating the complete document content, the Master in the CoE model checks the content coherence of the whole text. If incoherent content is found, it triggers the modification of the relevant paragraph content until the check is qualified and then officially outputs the final result.

[0019] Preferably, the construction steps of step S1 are as follows:

[0020] S11. Learn from the long document model essays.

[0021] S12. Split the structure of the template input by the user, extract the outline structure and the key information of each paragraph, and form a long document writing template, where the key information of each paragraph includes the main body and keywords.

[0022] Preferably, the specific steps of learning in step S11 include:

[0023] S111. Conduct a structural analysis of the long document model essays, extract the corresponding titles, outlines, chapter structures, key information of each paragraph, and the full text of each paragraph. Among them, the key information of each paragraph includes the paragraph theme and keywords, and form writing knowledge data.

[0024] S112. Build an index structure. Specifically, for the characteristics of long documents, design multiple different granularity segmentation, keyword extraction, and different indexing methods. Among them, different granularities include different granularities of paragraphs, chapters, and the full text, and different indexing methods include keyword indexing, title vector indexing, and chapter indexing.

[0025] Preferably, the base model for generating content in steps S5 and S6 is a domain base model trained with remote sensing knowledge.

[0026] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. After the computer program runs, it executes the method described in any one of the above.

[0027] The present invention also discloses a computer system, including a processor and a storage medium. The storage medium stores a computer program, and the processor reads and runs the computer program from the storage medium to execute the method described in any one of the above.

[0028] The beneficial effects of the present invention are as follows:

[0029] The effect of the present invention is to achieve the automated and efficient generation of long documents in the field of high-quality remote sensing. Through the pre-constructed domain knowledge base and database, the system can understand the theme input by the user and automatically perform steps such as outline formulation and chapter writing, greatly improving the speed and efficiency of document generation. The following are the main functional features of the present invention:

[0030] (1) Automated process: The user only needs to provide the writing theme, and the present invention can automatically execute the entire process from outline formulation to chapter writing.

[0031] (2) Imitation mechanism based on reference models: The user can specify the structure of the reference model, and the present invention will imitate according to these structures, improving the personalization and professionalism of the document.

[0032] (3) Domain base model: The present invention adopts a domain base model trained with remote sensing knowledge, ensuring the professionalism and accuracy of the generated content.

[0033] (4) Multi-granularity RAG technology: In the stage of chapter and paragraph generation, the system adopts multi-granularity RAG technology to effectively extract relevant information from the knowledge base and the ultra-long context environment, enhancing the coherence of the content.

[0034] (5) CoE multi-expert collaborative model: The present invention simulates the multi-expert collaboration mode and realizes the precise processing of tasks through the collaborative work of three intelligent agents: Planner, Executor, and Master. The Master intelligent agent is responsible for checking the coherence of the document and triggering content modification when necessary to ensure the document quality.

[0035] (6) High professionalism: Since the present invention uses a domain base model in the field of remote sensing in the S5 and S6 stages, it is superior to general large language models in terms of the professionalism of remote sensing content generation.

[0036] In summary, the present invention realizes the efficient, professional, and coherent generation of long documents in the field of remote sensing through an automated process, an imitation mechanism, a domain professional knowledge base, multi-granularity RAG technology, and a multi-expert collaborative model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is the flowchart of the method of the present invention;

[0039] Figure 2 This is an example diagram for generating long documents in the field of remote sensing in an embodiment of the present invention. Detailed implementation manners

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0041] The present invention is a remote sensing field long document generation model solution that combines multi-granularity RAG retrieval and multi-step reflection. The multi-granularity RAG retrieval is used for semantic understanding and retrieval of long contexts to improve semantic consistency during the generation of ultra-long texts. The multi-step reflection is applied during the long text outline planning and the generation of each chapter content. The content coherence of each chapter and the overall outline can be confirmed through multi-step reflection, which is better than the effect of direct single-generation.

[0042] Such as Figure 1 , a remote sensing long document generation method combining multi-granularity retrieval and multi-step reflection, includes the following steps:

[0043] S1. Construct a long document knowledge structure set. The long document knowledge structure set in the field of remote sensing is a domain long document chapter structure knowledge base and database pre-constructed to support the generation of long documents in the field of remote sensing. Usually, the user only needs to provide a writing topic, and the system can automatically start a series of overall processes such as outline formulation and chapter writing. In particular, the present invention adds a reference model writing mechanism, and the user can directly specify the structure of the model and require the large model to write by imitation.

[0044] The specific steps are as follows:

[0045] S11. Learn from the long document model essays.

[0046] The specific steps of learning include:

[0047] S111. Conduct a structural analysis of the long document model essays, extract the corresponding titles, outlines, chapter results, key information of each paragraph, and full text of each paragraph. Among them, the key information of each paragraph includes the paragraph theme and keywords, to form writing knowledge data.

[0048] S112. Construct an index structure. Specifically, for the characteristics of long documents, design various different granularity segmentation, keyword extraction, and different index methods. Among them, different granularities include different granularities of paragraphs, chapters, and full texts, and different index methods include keyword index, title vector index, and chapter index.

[0049] S12. Split the structure of the template input by the user, extract the outline structure and the key information of each paragraph, and form a long document writing template, where the key information of each paragraph includes the main body and keywords.

[0050] S2. Build an expert collaboration model, namely the CoE (Collaboration-of-Experts) model, which refers to a multi-expert collaboration model. Multiple agents play the roles of multiple experts in the writing process. In the CoE expert collaboration model of the large model, its core concept is to achieve precise processing of tasks through the collaborative work of multiple expert models. The CoE model of the present invention includes three agents, namely Planer, Executor, and Master. Among them, Planer is responsible for overall planning, Executor is responsible for execution, and the Master supervisor agent is responsible for result summary and verification.

[0051] S3. After the user inputs a writing request, Planer of the CoE model starts to execute the overall writing task planning, determines the general framework of the document and the steps to be followed.

[0052] S4. Analyze the input materials. If a sample essay is input, analyze the input sample essay at the same time. This step is the key to ensuring that the document content is consistent with the user's needs and the given materials.

[0053] S5. Generate a document outline based on the analyzed materials and sample essays. This step is the basis of document writing, ensuring the logicality and integrity of the document structure.

[0054] S6. Based on the outline, the system adopts multi-grain RAG (multi-grain Retrieval-augmented Generation) technology to extract relevant information from the knowledge base and the ultra-long upstream context, and generate the content of each chapter and paragraph. This step ensures the professionalism and coherence of the generated content. The multi-grain RAG technology can effectively extract relevant information from the knowledge base and the ultra-long upstream context, further ensuring the professionalism and coherence of the generated content. This technology can better understand and utilize the context information by extracting information from different granularities (such as words, phrases, sentences, etc.), thereby improving the quality and relevance of the generated text.

[0055] Among them, the base model for generating the content in steps S5 and S6 is the domain base model GisRS-LongWriter trained with remote sensing knowledge.

[0056] S7. Merge the generated content of each chapter and paragraph to form a complete document. This step involves adjusting and optimizing the logical relationships between paragraphs to ensure the fluency and consistency of the document.

[0057] After S8 generates the complete document content, the intelligent agent Master in the CoE model checks the content coherence of the whole content. If incoherent content is found, the relevant paragraph content will be modified until the inspection is qualified and then the final result will be officially output.

[0058] The present invention realizes the automatic and efficient generation of long documents in the field of high-quality remote sensing. Through the pre-constructed domain knowledge base and database, the system can understand the theme input by the user and automatically perform steps such as outline formulation and chapter writing, greatly improving the speed and efficiency of document generation.

[0059] The main functional features of the present invention are as follows:

[0060] (1) Automated process: The user only needs to provide the writing theme, and the present invention can automatically execute the entire process from outline formulation to chapter writing.

[0061] (2) Imitation mechanism of reference model essays: The user can specify the structure of the model essay, and the present invention will imitate according to these structures, improving the personalization and professionalism of the document.

[0062] (3) Domain base model: The present invention adopts a domain base model trained with remote sensing knowledge, trains a base model fine-tuned for long domain documents, GisRS-LongWriter, extends the single-input and output length of the LLM to more than 10,000 words, ensures the professionalism and accuracy of the generated content, and improves the quality of long domain text generation.

[0063] (4) Multi-granularity RAG technology: Introduces the recall of relevant content of multi-granularity RAG (multi-grain Retrieval-augmented Generation), improves the recall ability of key information retrieval in long texts, improves the semantic understanding and retrieval ability of long contexts, and improves the semantic consistency during the output of long texts.

[0064] (5) CoE multi-expert collaborative model: Introduces the CoE (Collaboration-of-Experts) thinking chain of multi-Agent writing, simulates the multi-expert collaboration mode. When planning the long text outline and generating the content of each chapter, it can go through multiple steps of reflection to confirm the content coordination of each chapter and the overall outline. After the overall writing is completed, the main control Agent conducts feedback inspection to ensure the content coherence of the overall output result. Specifically, through the collaborative work of three intelligent agents, Planner, Executor, and Master, the accurate processing of tasks is realized. The Master intelligent agent is responsible for checking the coherence of the document and triggering content modification when necessary to ensure the document quality.

[0065] (6) High professionalism: Since the base model in the field of remote sensing is used in the S5 and S6 stages of the present invention, the professionalism of remote sensing content generation is superior to that of general large language models.

[0066] In summary, through the automated process, imitation mechanism, domain-specific knowledge base, multi-granularity RAG technology, and multi-expert collaborative model, the present invention realizes the efficient, professional, and coherent generation of long documents in the field of remote sensing.

[0067] As can be seen from the embodiments of the present invention Figure 2 The data example shown in Table 1. The examples listed in Table 1 represent general document writing. Usually, the user only provides the writing topic, and the system automatically starts a series of overall processes such as outline formulation and chapter writing. The example 2 in Table 1 is the reference model imitation mechanism designed by the system. The user can directly specify the structure of the model and require the large language model to imitate. Any ultra-long text (>10,000 words) actually written by the user can be used as a template or example for long text generation.

[0068] The present invention also discloses a computer-readable storage medium and a computer system. A computer program is stored on a computer-readable storage medium. After the computer program runs, it executes the method described in any one of the above. A computer system includes a processor and a storage medium. A computer program is stored on the storage medium. The processor reads and runs the computer program to execute the method described in any one of the above.

[0069] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functional form. Whether such functionality is implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in different ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0070] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0071] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0072] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0073] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0074] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A remote sensing long document generation method combining multi-granularity recall and multi-step reflection, characterized in that It includes the following steps: S1. Construct a long document knowledge structure set; The specific steps include: S11. Learn from long document model essays; S12. Split the structure of the template input by the user, extract the outline structure and key information of each paragraph, and form a long document writing template, where the key information of each paragraph includes the main body and keywords; S2. Construct an expert collaboration model, namely the CoE model, which includes three agents, namely Planer, Executor, and Master. Among them, Planer is responsible for overall planning, Executor is responsible for execution, and the Master supervisor agent is responsible for result summary and verification; S3. After the user inputs a writing request, Planer of the CoE model starts to execute the overall writing task planning, determines the document framework and steps to follow; S4. Parse the input materials. If a model essay is input, parse the input model essay at the same time; S5. Generate a document outline based on the parsed materials and model essays; S6. Based on the outline, through multi-granularity RAG, extract information from the knowledge base and the ultra-long context environment at different granularities to generate the content of each chapter and paragraph; among them, the base model for content generation is a domain base model trained with remote sensing knowledge; S7. Merge the generated content of each chapter and paragraph to form a complete document; S8. After generating the complete document content, Master in the CoE model checks the content coherence of the whole text. If incoherent content is found, trigger the modification of the relevant paragraph content until the inspection is qualified and then output the final result.

2. The method according to claim 1, wherein: The specific steps of learning in step S11 include: S111. Conduct a structural analysis of the long document model essay, extract the corresponding title, outline, chapter structure, key information of each paragraph, and full text of each paragraph of the model essay. Among them, the key information of each paragraph includes the paragraph theme and keywords, and form writing knowledge data; S112. Construct an index structure. Specifically, for the characteristics of long documents, design various different granularity segmentation, keyword extraction, and different indexing methods. Among them, different granularities include different granularities of paragraphs, chapters, and full texts, and different indexing methods include keyword indexing, title vector indexing, and chapter indexing.

3. A computer-readable storage medium, characterized in that: A computer program is stored on a medium. After the computer program runs, it executes the method described in any one of claims 1 to 2.

4. A computer system, characterized in that: It includes a processor and a storage medium. A computer program is stored on the storage medium. The processor reads and runs the computer program from the storage medium to execute the method described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Large language model retrieval enhancement generation method based on hierarchical information expansion

    CN118779425A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1