Document deep tracing system based on multi-agent cooperation

The multi-agent collaborative literature tracing system solves the problem of insufficient literature survey coverage, achieves efficient and comprehensive literature tracing, provides in-depth analysis reports, and improves research efficiency.

CN121071126BActive Publication Date: 2026-05-12HEFEI ZHONGKE LEINAO INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI ZHONGKE LEINAO INTELLIGENCE TECH CO LTD
Filing Date
2025-11-06
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing literature reviews suffer from insufficient coverage and may miss important literature.

Method used

A multi-agent collaborative deep literature tracing system is adopted, including a central control agent, an academic search agent, a literature analysis agent, a correlation analysis agent, and a report generation agent, which achieves high-coverage literature tracing through collaborative work.

Benefits of technology

It achieves high-coverage literature tracing, avoids the omission of important literature, provides a comprehensive and three-dimensional research perspective, improves the efficiency of scientific research, and discovers hidden related literature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071126B_ABST
    Figure CN121071126B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-agent cooperation's literature depth tracing system, it is related to literature tracing technical field.System includes general control agent and with general control agent communication connection academic search agent, literature analysis agent, correlation analysis agent, report generation agent;Among them, academic search agent is used to obtain tracing literature according to preset tracing starting point and reference literature search;Literature analysis agent is used to analyze tracing literature, and obtain reference list;Correlation analysis agent is used to analyze tracing literature, and obtain the correlation information between tracing literature;Report generation agent is used to generate literature tracing report according to tracing literature and correlation information;General control agent is used to obtain reference according to reference list, and send tracing literature and correlation information to report generation agent.Thereby, high coverage literature tracing can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of document tracing, and particularly relates to a document deep tracing system based on multi-agent cooperation. BACKGROUND

[0002] In the field of scientific research, document research is a crucial first step for scientific researchers to carry out research on a topic and technological innovation. The main purpose is to comprehensively understand the origin, development context, current progress and potential future trends of a certain technology or research direction.

[0003] However, the document research in the related art usually adopts manual retrieval and analysis, or uses a visual analysis tool to achieve it. However, whether it is manual retrieval and analysis or the visual analysis tool in the related art, there is a problem that the coverage is not enough and important documents may be missed. SUMMARY

[0004] The present application aims to at least solve one of the technical problems in the related art. To this end, the purpose of the present application is to provide a document deep tracing system based on multi-agent cooperation to achieve high-coverage document tracing.

[0005] To achieve the above-mentioned purpose, the embodiment of the present application provides a document deep tracing system based on multi-agent cooperation, characterized in that the system comprises a total control agent and an academic search agent, a document analysis agent, an association analysis agent, and a report generation agent which are in communication connection with the total control agent; wherein the academic search agent is configured to search for a first document according to a preset tracing starting point, obtain a reference document, search for a second document according to the reference document, and take the first document and the second document as tracing documents; the document analysis agent is configured to obtain the tracing documents and analyze the tracing documents to obtain a reference document list; the association analysis agent is configured to obtain at least two of the tracing documents and analyze the tracing documents to obtain association information between the tracing documents; the report generation agent is configured to generate a document tracing report according to the tracing documents and the association information; and the total control agent is configured to obtain the reference documents according to the reference document list, obtain the tracing documents and the association information, and send the tracing documents and the association information to the report generation agent.

[0006] In addition, the document deep tracing system based on multi-agent cooperation according to the embodiment of the present application can also have the following additional technical features:

[0007] According to one embodiment of the present invention, the overall control agent is further configured to: acquire an initial research task, generate the preset tracing starting point according to the initial research task, and send the preset tracing starting point to the academic search agent, wherein the initial research task includes at least one of the following: a link to a document to be traced, a document title, and a research topic.

[0008] According to one embodiment of the present invention, the association information includes citation relationships and similarity relationships. The association analysis agent is specifically used to: analyze the at least two source documents, obtain the metadata of each source document, obtain the citation relationship between any two source documents, and obtain the semantic similarity between any two metadata. For each semantic similarity, the similarity relationship between the corresponding two source documents is obtained based on the semantic similarity.

[0009] According to one embodiment of the present invention, the overall control agent is further configured to: add the references to a preset queue of documents to be processed; the academic search agent is specifically configured to: obtain the references in the queue of documents to be processed, so as to perform a search based on the references.

[0010] According to one embodiment of the present invention, the academic search agent is further configured to search in a preset academic database based on the preset source starting point to obtain the first document, and search in the preset academic database based on the references to obtain the second document. The system further includes: a network search agent, specifically configured to receive the preset source starting point sent by the central control agent, obtain the references from the preset pending document queue, and perform an internet search based on the preset source starting point or the references, using the search results as supplementary information; the report generation agent is specifically configured to: generate a document source tree based on the source document, the citation relationship, and the similarity relationship, generate a deep analysis report based on the source document and the supplementary information, and generate the document source report based on the document source tree and the deep analysis report.

[0011] According to one embodiment of the present invention, the report generating agent is further configured to: generate a knowledge graph with the source document as a node and the citation relationship and the similarity relationship as edges; and traverse the knowledge graph starting from the first document to obtain the document source tree.

[0012] According to one embodiment of the present invention, the document analysis agent is further configured to associate the source documents with the supplementary information, and the report generation agent is further configured to: obtain the target document from the at least two source documents, obtain the target portion of the target document, and obtain the in-depth analysis report based on the target document, the target portion, and the supplementary information associated with the target document; wherein, the target document includes at least one of source documents whose citation frequency is higher than a preset frequency threshold and source documents with similar documents, the similar documents are source documents from the at least two source documents whose similarity to the target document is greater than a preset similarity threshold, and the target portion includes the abstract of the target document.

[0013] According to one embodiment of the present invention, the academic search agent is further configured to: obtain the metadata and access link of the source document, and send the metadata and access link as preliminary results to the overall control agent; the overall control agent is further configured to: add the preliminary results to a preset processing queue; the document analysis agent is further configured to: obtain the preliminary results from the preset processing queue, obtain the source document based on the preliminary results, and analyze the source document.

[0014] According to one embodiment of the present invention, the overall control agent is further configured to: send the associated information to the report generating agent when the system encounters a preset search termination condition.

[0015] According to one embodiment of the present invention, the preset search termination condition includes at least one of the following: the number of times the academic search agent performs searches in the preset academic database reaches a preset number; the total time the academic search agent spends searching in the preset academic database reaches a preset time; the number of source documents reaches a preset number; and the duration for which the preset pending queue is an empty queue reaches a preset time threshold.

[0016] The multi-agent collaborative deep literature tracing system according to an embodiment of the present invention includes a central control agent and an academic search agent, a literature analysis agent, a correlation analysis agent, and a report generation agent, all communicatively connected to the central control agent. The academic search agent searches for a first document based on a preset tracing starting point, obtains references, searches for a second document based on the references, and uses the first and second documents as tracing sources. The literature analysis agent obtains tracing sources and analyzes them to obtain a reference list. The correlation analysis agent obtains at least two tracing sources and analyzes them to obtain correlation information between them. The report generation agent generates a literature tracing report based on the tracing sources and correlation information. The central control agent obtains references from the reference list, obtains tracing sources and correlation information, and sends the tracing sources and correlation information to the report generation agent. This enables highly comprehensive literature tracing.

[0017] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0018] Figure 1 This is a structural block diagram of a document deep tracing system based on multi-agent collaboration according to an embodiment of the present invention;

[0019] Figure 2 This is a flowchart of a document deep tracing system based on multi-agent collaboration, according to an embodiment of the present invention.

[0020] Figure 3 This is a schematic diagram of the structure of a document deep tracing system based on multi-agent collaboration according to an embodiment of the present invention. Detailed Implementation

[0021] The following description, with reference to the accompanying drawings, describes a document deep tracing system based on multi-agent cooperation according to embodiments of the present invention, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described with reference to the accompanying drawings are exemplary and should not be construed as limiting the present invention.

[0022] Figure 1 This is a structural block diagram of a document deep tracing system based on multi-agent collaboration according to an embodiment of the present invention.

[0023] like Figure 1As shown, the document deep tracing system 100 based on multi-agent collaboration includes a central control agent 101 and academic search agents 102, document analysis agents 103, correlation analysis agents 104, and report generation agents 105, all connected to the central control agent 101. The academic search agent 102 searches for a first document based on a preset tracing starting point, obtains references, searches for a second document based on the references, and uses the first and second documents as tracing sources. The document analysis agent 103 obtains tracing sources and analyzes them to obtain a reference list. The correlation analysis agent 104 obtains at least two tracing sources and analyzes them to obtain correlation information between them. The report generation agent 105 generates a document tracing report based on the tracing sources and correlation information. The central control agent 101 obtains references from the reference list, obtains tracing sources and correlation information, and sends the tracing sources and correlation information to the report generation agent 105.

[0024] Specifically, the aforementioned document deep tracing system 100 based on multi-agent collaboration includes an academic search agent 102, a document analysis agent 103, a correlation analysis agent 104, and a report generation agent 105.

[0025] Each of the aforementioned intelligent agents can be constructed based on a large language model to implement its corresponding function.

[0026] First, a preset tracing starting point is obtained. This preset tracing starting point can be generated in advance by the user or generated by the user according to their own tracing needs when tracing is required.

[0027] After obtaining the preset source starting point, the academic search agent 102 searches according to the preset source starting point, obtains the first document, and uses the first document as the source document.

[0028] After obtaining the source documents, the document analysis agent 103 analyzes the source documents and obtains a list of references.

[0029] After obtaining the list of references, the overall control agent 101 retrieves the references based on the list.

[0030] After the master control agent 101 obtains the references, the academic search agent 102 searches based on the references to obtain the second references, and uses the second references as the source references so that the literature analysis agent 103 can analyze the source references to obtain a reference list. The master control agent 101 then obtains references based on the reference list, and the academic search agent 102 searches based on the references to obtain new second references.

[0031] By repeating the above process, a reference list can be continuously obtained based on the source literature, and new source literature can be obtained based on the reference list, thereby achieving in-depth source tracing, improving the coverage of literature source tracing, and avoiding the omission of important literature.

[0032] Furthermore, upon obtaining the source documents, the association analysis agent 104 needs to analyze them. Specifically, after obtaining at least two source documents, the association analysis agent 104 analyzes the association information between them. The central control agent 101 then sends the source documents and association information to the report generation agent 105, which generates a document tracing report based on the source documents and association information.

[0033] This allows for comprehensive literature tracing.

[0034] Optionally, the aforementioned academic search agent 102 can be a dedicated ArXiv (a website that collects preprints of papers in physics, mathematics, computer science, biology, and mathematical economics) search agent, or it can be designed as a more general academic database agent that integrates API (Application Programming Interface) calls or web page information crawling capabilities for multiple academic databases such as Google Scholar, PubMed (a medical literature database), Scopus (a literature and citation database), and Web of Science (a science citation index database), thereby expanding the coverage of literature search.

[0035] The reference list obtained from the source literature can rely on the full-text understanding of large language models, or it can be combined with rule-based or traditional machine learning bibliometric analysis tools. These tools are specifically optimized for the structure of academic papers and have higher accuracy in extracting structured information such as references, authors, and institutions.

[0036] While tracing back the references (past), new literature (future) can be tracked in parallel. This new literature refers to the literature that cites the source literature, thus simultaneously constructing the source and development path of the technology. Moreover, the priority of exploration can be dynamically adjusted by a reinforcement learning model. This model learns from historically effective paths to determine which branches to explore in a limited way to find key literature more quickly. That is, when multiple source literatures are obtained, the reinforcement learning model can be used to determine which source literature's reference list to prioritize for further tracing.

[0037] In some embodiments of the present invention, the overall control agent 101 is further configured to: obtain an initial research task, generate a preset tracing starting point based on the initial research task, and send the preset tracing starting point to the academic search agent 102, wherein the initial research task includes at least one of the following: a link to a document to be traced, a document title, and a research topic.

[0038] Specifically, users need to first submit an initial research task, which can be a specific literature link, literature title, or a descriptive research topic, such as researching the Transformer model of user time series.

[0039] After the user submits the initial research task, the overall control agent 101 analyzes the initial research task and obtains the above-mentioned preset source tracing starting point.

[0040] In some embodiments of the present invention, the association information includes citation relationships and similarity relationships. The association analysis agent 104 is specifically used to: analyze at least two source documents, obtain the metadata of each source document, obtain the citation relationship between any two source documents, and obtain the semantic similarity between any two metadata. For each semantic similarity, the similarity relationship between the corresponding two source documents is obtained based on the semantic similarity.

[0041] Specifically, the aforementioned association analysis agent 104 is used to implement methodological similarity analysis and technology evolution path identification.

[0042] To implement methodological similarity analysis, the association analysis agent 104 obtains metadata for each source document, including its abstract or technical solution. After obtaining the metadata, it converts it into vector embeddings and then calculates the semantic similarity between any two metadata sets, thereby determining the similarity relationship between the two source documents based on the semantic similarity.

[0043] The aforementioned metadata includes at least one of the following: title, author, abstract, publication date, and publishing institution of the source document.

[0044] This allows for the discovery of hidden related documents that do not use the same terminology but address type issues or employ similar technical approaches.

[0045] In order to identify the path of technological evolution, we analyzed the various source documents to obtain the citation relationships of each source document.

[0046] Optionally, while obtaining the citation relationships of each source document, the publication time of each source document can also be obtained, and then a timeline of technological evolution can be constructed based on the publication time and citation relationships of each source document.

[0047] In some embodiments of the present invention, the association analysis agent 104 is also used to obtain the correlation between the source literature and the initial research task.

[0048] Specifically, to obtain relevance, metadata can be obtained for each source document. After obtaining the metadata, prompt words can be constructed based on the metadata, and these prompt words can be abstracted into a question-and-answer question. The large language model provides the answer based on the question-and-answer question. Furthermore, the relevance between the source document and the initial research task is obtained based on the answer, and source documents with relevance below a set threshold are removed. That is, references are not obtained based on these source documents, and they are not considered when generating the literature tracing report.

[0049] The following example illustrates the prompt words constructed above. In this example, the source literature is a research paper, and the initial research task is the research topic.

[0050] I. Role: You are a professional academic research analyst who specializes in evaluating the relevance of papers to research topics, providing high-quality initial screening for in-depth academic searches.

[0051] II. Task: To evaluate the relevance of the papers obtained in the pre-search phase, determine which papers are worthy of entering the in-depth analysis phase, and provide detailed relevance analysis for each paper.

[0052] III. Input Data: Research topic, original user query, metadata.

[0053] IV. Evaluation Dimensions: 1. Topic Relevance (40%)

[0054] The degree of direct relevance between the paper's topic and the user's original query.

[0055] Matching degree of key concepts and terms.

[0056] The degree of overlap in research fields.

[0057] 2. Methodological relevance (25%):

[0058] Does the research method used in this paper relate to the user's original query?

[0059] The relevance of technical approaches and solutions.

[0060] Experimental design and evaluation of the suitability of the methods.

[0061] 3. Academic value (20%):

[0062] The paper's innovation and originality.

[0063] The importance of research contributions.

[0064] Influence in related fields.

[0065] 4. Timeliness (10%):

[0066] The recentity of the paper's publication date.

[0067] The timeliness of the research content.

[0068] Relevance of technological development.

[0069] 5. Citation potential (5%):

[0070] The likelihood of the paper being cited.

[0071] Its value as a starting point for further research.

[0072] The potential for network expansion is cited.

[0073] V. Based on the above input, please extract and generate the following information:

[0074] 1. A floating-point number between `0.0` and `1.0`.

[0075] `1.0` indicates that the paper is completely focused on the search topic.

[0076] `0.5` indicates that the paper's content is partially relevant to the search topic (e.g., the topic is a subfield or application scenario).

[0077] `0.0` represents completely irrelevant.

[0078] 2. In a short paragraph, explain in detail your reasoning for providing this floating-point number. Please specify which aspects of your paper (such as research question, methodology, dataset, etc.) match or do not match the user's original query.

[0079] 3. A short paragraph in Chinese summarizing the paper's abstract, allowing users to quickly understand the paper's core ideas and main content.

[0080] 4. A paragraph that clearly describes the main research methods or technical approaches used in the paper.

[0081] 5. A paragraph summarizing the main contributions, findings, or key results of the paper.

[0082] 6. An English array containing 3-5 of the most core technical keywords. These keywords should concisely summarize the technical points of the paper.

[0083] 6. Please strictly adhere to the following JSON structure when returning the correlation analysis, without adding any additional explanations or comments. Ensure that all string fields are concise, coherent paragraphs, not lists.

[0084] json

[0085] {

[0086] "summary": "Please provide a Chinese summary of the paper's abstract here..."

[0087] "methodology": "Please describe the main research methods used in this paper..."

[0088] "Contributions": "Please summarize the main contributions or findings of the paper here..."

[0089] "technical_keywords":["keyword1","keyword2","keyword3"],

[0090] "relevance_reasoning": "Please explain the reasons for the rating in detail here..."

[0091] "relevance_score": 0.0

[0092] }

[0093] The above "technical_keywords":["keyword1","keyword2","keyword3"] refers to the keywords displayed in the paper.

[0094] The above "relevance_score":0.0 refers to displaying a floating-point number, which in this example is 0.0.

[0095] The aforementioned original user query refers to the original information entered by the user, from which the research topic is derived.

[0096] In some embodiments of the present invention, the overall control agent 101 is further configured to: add references to a preset queue of documents to be processed; the academic search agent 102 is specifically configured to: obtain references from the queue of documents to be processed, so as to perform a search based on the references.

[0097] Specifically, after obtaining the references from the reference list, the overall control agent 101 adds the references to the preset queue of documents to be processed.

[0098] The academic search agent 102 is configured to retrieve documents from the queue of documents to be processed, so that the academic search agent 102 can obtain the references obtained by the overall control agent 101.

[0099] In some embodiments of the present invention, the academic search agent 102 is further configured to search in a preset academic database based on a preset source tracing starting point to obtain a first document, and search in the preset academic database based on references to obtain a second document. The document deep tracing system 100 based on multi-agent collaboration also includes: a network search agent, specifically configured to receive a preset source tracing starting point sent by the central control agent 101, obtain references from a preset queue of documents to be processed, and perform an internet search based on the preset source tracing starting point or references, using the search results as supplementary information; and a report generation agent 105, specifically configured to: generate a document tracing tree based on the source documents, citation relationships, and similarity relationships, generate a deep analysis report based on the source documents and supplementary information, and generate a document tracing report based on the document tracing tree and the deep analysis report.

[0100] In some embodiments of the present invention, the report generation agent 105 is also used to: generate a knowledge graph with source documents as nodes and citation relationships and similarity relationships as edges; and traverse the knowledge graph starting from the first document to obtain a document source tree.

[0101] In some embodiments of the present invention, the document analysis agent 103 is further configured to associate source documents with supplementary information, and the report generation agent 105 is further configured to: obtain target documents from at least two source documents, obtain the target portion of the target documents, and obtain a deep analysis report based on the target documents, the target portion and the supplementary information associated with the target documents; wherein, the target documents include at least one of source documents whose citation frequency is higher than a preset frequency threshold and source documents with similar documents, and the similar documents are source documents from at least two source documents whose similarity to the target documents is greater than a preset similarity threshold, and the target portion includes the abstract of the target documents.

[0102] The above-mentioned method of determining target documents based on similar documents can be described as follows: for each source document obtained from the source documents, the similarity between the source document and the target document is obtained; based on the similarity and a preset similarity threshold, source documents with similar documents are found from all the source documents obtained, and the found source documents are used as target documents; and the similarity can be semantic similarity, textual similarity, etc.

[0103] In some embodiments of the present invention, the academic search agent 102 is further configured to: obtain the metadata and access links of the source documents, and send the metadata and access links as preliminary results to the central control agent 101; the central control agent 101 is further configured to: add the preliminary results to a preset processing queue; the document analysis agent 103 is further configured to: obtain the preliminary results from the processing queue, obtain the source documents based on the preliminary results, and analyze the source documents.

[0104] In some embodiments of the present invention, the overall control agent 101 is further configured to: send the associated information to the report generation agent 105 when a preset search termination condition occurs in the document deep tracing system 100 based on multi-agent collaboration.

[0105] In some embodiments of the present invention, a preset search termination condition is included, which is at least one of the following: the number of times the academic search agent 102 searches in the preset academic database reaches a preset number; the total time the academic search agent 102 spends searching in the preset academic database reaches a preset time; the number of source documents reaches a preset number; and the duration of the preset pending queue being an empty queue reaches a preset time threshold.

[0106] It should be noted that, regarding the aforementioned preset search termination condition: the preset queue to be processed is an empty queue, the overall control agent 101 can be configured to include a clock. When it detects that the preset queue to be processed is empty, it starts timing. When the timing result has reached the preset time threshold, if the preset queue to be processed is still empty, it can be considered that the academic search agent 102 is unable to search for more source documents based on the currently obtained preset source starting point or references, thus confirming the occurrence of the preset search termination condition.

[0107] In other words, the above-mentioned overall control agent 101 determines that the duration of the preset queue to be processed being an empty queue has reached a preset time threshold. This determination is made after the overall control agent 101 generates a preset tracing starting point and starts document tracing. Moreover, after the document tracing ends, the determination that the duration of the preset queue to be processed being an empty queue has reached the preset time threshold is terminated.

[0108] The following description uses a specific example.

[0109] In this specific embodiment, the aforementioned document is a paper.

[0110] See Figure 2 and Figure 3 It includes a report analyzer, search tree builder, paper processor, citation extractor, extended searcher, depth controller, knowledge graph builder, and report generator. Figure 3 The agent collaboration graph shown can be constructed using the langgraph framework (a large model agent development framework), or other mainstream agent frameworks such as AutoGen (a multi-agent application development framework) or CrewAI (a multi-agent framework).

[0111] Specifically, the report analyzer and the search tree builder are implemented by the overall control agent 101, the paper processor is implemented by the academic search agent 102, the citation extractor is implemented by the literature analysis agent 103 and the association analysis agent 104, the extended searcher is implemented by the web search agent, the depth controller is implemented by the overall control agent 101, the knowledge graph builder is implemented by the report generation agent 105, and the report generator is implemented by the report generation agent 105.

[0112] The aforementioned document deep tracing system 100 based on multi-agent collaboration works in the following manner.

[0113] First, task initialization.

[0114] Specifically, the user submits an initial research task to the multi-agent collaborative literature deep tracing system 100. This initial research task can be a specific literature link, a literature title, or a descriptive research topic, such as researching Transformer models for time series prediction. This process is called... Figure 2 User input in the middle.

[0115] Second, the central control agent 101 receives and distributes tasks.

[0116] Specifically, after receiving an initial research task, the multi-agent collaborative literature deep tracing system 100 first parses the task using a central control agent 101 to obtain a preset tracing starting point. For example, after receiving the initial research task, the central control agent 101 can determine whether the context of the initial research task is sufficient. If it is sufficient, it constructs a search tree and generates the preset tracing starting point.

[0117] The central control agent 101 is the central router of the multi-agent collaborative document deep tracing system 100. It is responsible for maintaining a global state, which includes core data such as the document queue to be processed, the set of processed papers, and the paper relationship graph. Based on the current stage of the task, the central control agent 101 decides which one or more downstream specialized agents to call to execute the specific task.

[0118] Third, parallel information acquisition.

[0119] The master control agent 101 distributes the preset source tracing starting point as the initial task to two parallel search agents: academic search agent 102 and network search agent.

[0120] The academic search agent 102 is specifically responsible for interacting with the API of the preset academic database. Upon receiving the initial task, the academic search agent 102 obtains task keywords based on the initial task, and upon obtaining references, obtains the paper titles based on the references. Then, based on the task keywords or paper titles, it accurately retrieves relevant academic papers, which are the aforementioned source documents. It then obtains the metadata (title, author, abstract) and access links of the source documents, using this metadata and access links as preliminary results.

[0121] The web search agent invokes a general search engine to conduct a broader internet search. Its purpose is to find information related to the research topic, such as technical blogs, news reports, open-source code repositories, and community discussions, as supplementary information to academic papers. The aforementioned research topic is considered the preset source starting point when the web search agent receives it, and the references when the web search agent retrieves references from the preset queue of documents to be processed.

[0122] Fourth, in-depth analysis of the paper.

[0123] After the academic search agent 102 returns preliminary results, the central control agent 101 adds the preliminary results to a preset processing queue. Next, the central control agent 101 calls the literature analysis agent 103 to retrieve the preliminary results from the preset processing queue, obtain the source documents based on the preliminary results, and perform in-depth analysis on each source document obtained.

[0124] The main responsibilities of document analysis agent 103 are:

[0125] 1. Content extraction: Download the original PDF of the paper and parse it into Markdown (Lightweight Markup Language) format.

[0126] 2. Key Information Extraction: Since the literature analysis agent 103 is based on a large language model, it can utilize the understanding capabilities of the large language model to extract core information from the full text of the paper, including: the problem solved, the core methods used, experimental results, and a list of references.

[0127] 3. Code Association: Combine the supplementary information obtained by the network search agent to associate each source document with its corresponding supplementary information.

[0128] Fifth, recursive tracing and expansion.

[0129] The reference list extracted by the literature analysis agent 103 is crucial for achieving in-depth source tracing. The central control agent 101 obtains references from this list, which are then considered new nodes to be explored. The central control agent 101 adds these references to a pre-defined queue of documents to be processed, enabling the academic search agent 102 and the web search agent to re-search based on the references in this queue. This forms a recursive exploration loop, continuously digging deeper into earlier foundational papers and tracing back to the latest cited literature.

[0130] Sixth, multi-dimensional correlation analysis.

[0131] The master control agent 101 calls the association analysis agent 104 to perform in-depth association mining on all the source documents that have been acquired, including the above-mentioned methodological similarity analysis and the above-mentioned technology evolution path identification.

[0132] Seventh, the literature source tree and in-depth analysis report are generated.

[0133] Specifically, it determines whether the exploration has reached the preset depth and whether there are no more nodes. When the exploration reaches the preset depth or there are no more nodes, it determines that the preset search termination condition has been met and the recursive process terminates.

[0134] The above-mentioned determination of whether the exploration has reached the preset depth can be made using the following three conditions. If any one of the three conditions is met, it is determined that the exploration has reached the preset depth.

[0135] Condition 1: The number of times the academic search agent 102 performs searches in the preset academic database reaches the preset number.

[0136] Condition 2: The total time spent by the academic search agent 102 searching in the preset academic database reaches the preset time.

[0137] Condition 3: The number of source documents reaches the preset number.

[0138] The above determination of whether there are no more nodes can be obtained by obtaining the time when the preset queue to be processed is empty. When the preset time when the queue to be processed is empty reaches the preset time threshold, it is determined that the academic search agent 102 can no longer search for more source documents, and thus it is determined that there are no more nodes.

[0139] After the recursive process terminates, the master control agent 101 invokes the report generation agent 105. The report generation agent 105 is responsible for receiving the information sent by the master control agent 101 and then performing the following tasks:

[0140] 1. Construct a knowledge graph: Starting with source literature, construct a structured knowledge graph with citation relationships and similarity relationships as edges.

[0141] 2. Generate a literature source tree: Starting from the first document, traverse the knowledge graph to generate a well-defined and visualized literature source tree to show the evolution of the technology.

[0142] 3. Generate an in-depth analysis report: Obtain the target document from all source documents. This target document can be pioneering, highly cited, or highly similar; it represents the core node in the aforementioned knowledge graph. Then, a detailed interpretation of the target document is performed, including:

[0143] Abstract Refinement and Core Idea Interpretation: This section extracts and interprets the target section of the target literature to help users quickly understand its contributions. The target section includes the abstract of the target literature and may also include its core methodologies.

[0144] Precision recommendations and guidance: Based on the position and role of the target literature in the technical context, suggestions are given on precision or general reading.

[0145] Related resource integration: Organize and list the supplementary information associated with each target document to form a resource list that can be directly used.

[0146] Based on the above, an in-depth analysis report is generated.

[0147] Optionally, when generating an in-depth analysis report, the aforementioned correlations can also be included in the in-depth analysis report.

[0148] 4. Generate a summary overview: Summarize the development of the entire technology field, highlighting pioneering documents and key milestone works.

[0149] Eighth, produce a comprehensive research report.

[0150] Specifically, a literature tracing report is generated based on the paper source tree and in-depth analysis report, and presented to the user in a structured, multimedia format. The user receives not just a list of documents, but a complete, interactive, comprehensive research report containing in-depth insights and practical data, thus completing the entire in-depth tracing and analysis task.

[0151] In summary, the multi-agent collaborative document tracing system of this invention includes a central control agent and several agents connected to it: an academic search agent, a document analysis agent, a correlation analysis agent, and a report generation agent. The academic search agent searches for a first document based on a preset tracing starting point, obtains references, searches for a second document based on the references, and uses the first and second documents as tracing sources. The document analysis agent obtains and analyzes the tracing sources to obtain a reference list. The correlation analysis agent obtains at least two tracing sources and analyzes them to obtain correlation information between them. The report generation agent generates a document tracing report based on the tracing sources and correlation information. The central control agent obtains references from the reference list, obtains the tracing sources and correlation information, and sends the tracing sources and correlation information to the report generation agent. This allows for highly comprehensive document tracing. Furthermore, by generating a literature source tree, complex literature networks can be automatically sorted and integrated, clearly displaying the evolution, pioneering work, key turning points, and major technical branches of any technology in a structured and visual manner. This reduces the cognitive load on researchers and helps them quickly establish a structured understanding of a field. Moreover, by setting up a network search agent, rigorous academic papers are organically integrated and linked with active open-source code, technical blogs, community discussions, and other information. This design not only verifies the practical feasibility of the theory but also provides researchers with an unprecedented comprehensive and three-dimensional research perspective encompassing theory, implementation, and community feedback, greatly enriching the breadth and depth of literature research. Furthermore, by setting up a central control agent, which autonomously plans and decomposes tasks based on a high-order initial research task and dynamically coordinates agents with different functions to perform multi-step reasoning and recursive exploration, this working mode, similar to that of a human research assistant, enables the system not only to perform retrospective source tracing but also, based on a deep understanding of the technological development trajectory, to assist researchers in making forward-looking trend predictions and discovering innovative opportunities. Moreover, automated intelligent agent workflows reduce literature review time from days or even weeks to minutes. This frees researchers from tedious and repetitive searching and filtering, allowing them to focus on more creative analysis and thinking, significantly improving the overall efficiency of scientific research. Furthermore, it introduces an intelligent agent based on deep semantic understanding for association analysis. It not only analyzes direct citations but also uncovers "hidden" literature—documents without direct citations but sharing similar ideas or offering mutual references—through vectorized analysis of paper methodologies and technical approaches. This multi-dimensional association analysis capability breaks through the limitations of traditional methods, constructing a more comprehensive and profound technical picture, effectively avoiding the limitations imposed by keywords or personal experience biases.

[0152] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein can be considered as a ordered list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0153] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0154] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0155] In the description of this specification, the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and should not be construed as limiting the present invention.

[0156] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0157] In this specification, unless otherwise stated, the terms "installation," "connection," "joining," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly defined. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0158] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0159] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A document deep tracing system based on multi-agent collaboration, characterized in that, The system includes a central control agent and several agents connected to the central control agent, including an academic search agent, a literature analysis agent, a correlation analysis agent, and a report generation agent; wherein... The academic search agent is used to search for a first document based on a preset source starting point, obtain references, search for a second document based on the references, and use the first document and the second document as source documents. The document analysis agent is used to obtain the source documents and analyze them to obtain a list of references. The association analysis agent is used to obtain at least two of the source documents and analyze them to obtain the association information between the source documents. The report generation agent is used to generate a document tracing report based on the source documents and the associated information; The overall control agent is used to obtain the references according to the reference list, and to acquire the source references and the related information, and to send the source references and the related information to the report generation agent; The association information includes citation relationships and similarity relationships, and the association analysis agent is specifically used for: The at least two source documents are analyzed to obtain the metadata of each source document, the citation relationship between any two source documents, and the semantic similarity between any two metadata. For each semantic similarity, the similarity relationship between the corresponding two source documents is obtained based on the semantic similarity. The academic search agent is further configured to search in a preset academic database based on the preset source starting point to obtain the first document, and to search in the preset academic database based on the references to obtain the second document. The system also includes: The network search agent is specifically used to receive the preset source tracing starting point sent by the central control agent, obtain the references from the preset pending document queue, and perform Internet search based on the preset source tracing starting point or the references, using the search results as supplementary information. The report-generating agent is specifically used for: A literature tracing tree is generated based on the source literature, the citation relationship, and the similarity relationship. An in-depth analysis report is generated based on the source literature and the supplementary information. The literature tracing report is also generated based on the literature tracing tree and the in-depth analysis report.

2. The document deep tracing system based on multi-agent collaboration according to claim 1, characterized in that, The overall control intelligent agent is also used for: The process involves obtaining an initial research task, generating a preset source tracing starting point based on the initial research task, and sending the preset source tracing starting point to the academic search agent. The initial research task includes at least one of the following: a link to a document to be traced, a document title, or a research topic.

3. The document deep tracing system based on multi-agent collaboration according to claim 1, characterized in that, The overall control intelligent agent is also used for: Add the references to the preset queue of documents to be processed; The academic search agent is specifically used for: Retrieve the references from the queue of documents to be processed, and perform a search based on the references.

4. The document deep tracing system based on multi-agent collaboration according to claim 1, characterized in that, The report-generating agent is also used for: A knowledge graph is generated using the source documents as nodes and the citation relationships and similarity relationships as edges; Starting from the first document, the knowledge graph is traversed to obtain the document source tree.

5. The document deep tracing system based on multi-agent collaboration according to claim 1, characterized in that, The document analysis agent is further configured to associate the source documents with the supplementary information, and the report generation agent is further configured to: The process involves obtaining the target document from the at least two source documents, obtaining the target portion of the target document, and generating the in-depth analysis report based on the target document, the target portion, and the supplementary information associated with the target document. The target document includes at least one of the following: source documents with a citation frequency higher than a preset frequency threshold, or source documents with similar documents. The similar documents are source documents from the at least two source documents whose similarity to the target document is greater than a preset similarity threshold. The target portion includes the abstract of the target document.

6. The document deep tracing system based on multi-agent collaboration according to claim 1, characterized in that, The academic search agent is also used for: Obtain the metadata and access link of the source document, and send the metadata and access link as preliminary results to the central control agent; The overall control intelligent agent is also used for: Add the preliminary results to a preset queue to be processed; The document analysis agent is also used for: The preliminary results are obtained from the preset queue to be processed, and the source documents are obtained based on the preliminary results for analysis.

7. The document deep tracing system based on multi-agent collaboration according to claim 6, characterized in that, The overall control intelligent agent is also used for: When the system encounters a preset search termination condition, the associated information is sent to the report generating agent.

8. The document deep tracing system based on multi-agent collaboration according to claim 7, characterized in that, The preset search termination conditions include at least one of the following: the number of searches performed by the academic search agent in the preset academic database reaches a preset number; the total search time of the academic search agent in the preset academic database reaches a preset time; the number of source documents reaches a preset number; and the duration of the preset pending queue being an empty queue reaches a preset time threshold.