Information deep mining method and system based on multi-agent cooperation framework

The information deep mining method using a multi-agent collaborative framework addresses the shortcomings of traditional information mining methods in processing multi-source heterogeneous information, achieving comprehensive and in-depth information acquisition and generating accurate and complete information deep mining results.

CN120804297BActive Publication Date: 2025-11-21JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511277499.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-21
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional information mining methods have limited search scope when faced with multi-source and heterogeneous information, making it difficult to achieve comprehensive coverage. Furthermore, they are insufficient in information parsing, association extraction, and deep reasoning, resulting in incomplete information acquisition, susceptibility to subjective factors, low efficiency, and difficulty in uncovering the complex relationships and potential value behind the information.

Method used

A deep information mining method based on a multi-agent collaborative framework is adopted. The planning expert agent performs task decomposition and collaborative scheduling to generate a collaborative task scheduling scheme. The search expert agent is called to perform multi-source information collaborative search. The reading expert agent performs deep analysis and correlation extraction. The analysis expert agent performs multi-level deep reasoning. Finally, the integration expert agent performs multi-dimensional verification and fusion to generate deep mining results that meet the information requirements.

Benefits of technology

It expands the scope and depth of information acquisition, uncovers implicit relationships between information, ensures the accuracy and completeness of results, and automates, simplifies, and increases the efficiency of the information mining process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804297B_ABST
    Figure CN120804297B_ABST
Patent Text Reader

Abstract

The application provides an information deep mining method and system based on a multi-agent cooperation framework. First, a to-be-mined information requirement is received, task decomposition and cooperation scheduling are performed by a planning expert agent, a cooperation task scheduling scheme is generated, then, a search expert agent performs multi-source information cooperative search according to the scheme, obtains multi-source original information, a reading expert agent deeply analyzes and correlates the multi-source original information, generates structured correlated information, an analysis expert agent performs multi-level deep reasoning on the structured correlated information, obtains a deep mining intermediate result, finally, an integration expert agent performs multi-dimensional verification and fusion on the intermediate result, generates a final information deep mining result meeting the requirement, so that the automation, intelligence and high efficiency of information mining are realized, and the potential value of the mined information is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to an information deep mining method and system based on a multi-agent collaboration framework. BACKGROUND

[0002] In the traditional field of information mining, a single processing mode is usually adopted to complete the information mining task. For example, a single search tool is relied on to collect information from limited databases or network resources, and then the collected information is preliminarily sorted and analyzed by manual or simple algorithms. However, with the rapid development of information technology and the explosive growth of data volume, the form and source of information become extremely complex and diverse, and the single processing mode gradually exposes many problems. On the one hand, the search range of the single search tool is limited, and it is difficult to cover multi-source and heterogeneous information, resulting in that the obtained information is not comprehensive and cannot meet the needs of deep mining. On the other hand, the traditional method is insufficient in information analysis, correlation extraction and deep reasoning, and it is difficult to mine the complex relationships and potential value hidden behind the information. In addition, the manual processing method is low in efficiency and is easily affected by subjective factors, and it is difficult to ensure the accuracy and consistency of information mining. SUMMARY

[0003] Therefore, the purpose of the present application is to provide an information deep mining method and system based on a multi-agent collaboration framework.

[0004] According to a first aspect of the present application, an information deep mining method based on a multi-agent collaboration framework is provided, which comprises:

[0005] receiving an information mining demand, and performing task decomposition and collaboration scheduling processing on the information mining demand by a planning expert agent, to generate a collaboration task scheduling scheme containing task hierarchical relationship and agent collaboration path;

[0006] based on the collaboration task scheduling scheme, calling a search expert agent to perform multi-source information collaborative search operation, to obtain multi-source original information containing multi-dimensional information sources;

[0007] performing deep analysis and correlation extraction processing on the multi-source original information by a reading expert agent, to generate structured correlation information with entity correlation network;

[0008] calling an analysis expert agent to perform multi-level deep reasoning processing on the structured correlation information, to obtain deep mining intermediate results containing information implicit relationships;

[0009] performing multi-dimensional verification and fusion processing on the deep mining intermediate results by an integration expert agent, to generate final information deep mining results meeting the information mining demand.

[0010] According to a second aspect of the present application, a multi-agent collaboration framework-based information deep mining system is provided, which comprises a machine readable storage medium and a processor, the machine readable storage medium stores machine executable instructions, and the processor, when executing the machine executable instructions, implements the multi-agent collaboration framework-based information deep mining method described above.

[0011] According to a third aspect of the present application, a computer readable storage medium is provided, which stores computer executable instructions, and when the computer executable instructions are executed, the multi-agent collaboration framework-based information deep mining method described above is implemented.

[0012] According to any one of the above aspects, the technical effects of the present application are as follows:

[0013] By means of the task decomposition and collaboration scheduling of the planned expert agent on the information demand to be mined, a collaboration task scheduling scheme is generated, the tasks and collaboration paths of each agent are clarified, the search expert agent performs multi-source information collaborative search according to the collaboration task scheduling scheme, and a wide range of multi-source raw information from different dimensions and sources can be obtained, which greatly expands the range and depth of information acquisition. The reading expert agent performs deep analysis and correlation extraction on the multi-source raw information to generate structured correlated information with entity correlation network, the analysis expert agent performs multi-level deep reasoning on the structured correlated information, which can mine the implicit relationships between information and reveal the potential value of information. Finally, the integration expert agent performs multi-dimensional verification and fusion on the intermediate results of deep mining to generate the final information deep mining result meeting the information demand to be mined, which ensures the accuracy and integrity of the result. Thus, the automation, intelligence and efficiency of the information mining process are realized, and valuable knowledge can be mined from complex information. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 A flowchart of the multi-agent collaboration framework-based information deep mining method provided by the embodiments of the present application is shown;

[0015] Figure 2 A component structure diagram of the multi-agent collaboration framework-based information deep mining system provided by the embodiments of the present application is shown. DETAILED DESCRIPTION

[0016] Embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not limit the technical solutions of the embodiments of the present application.

[0017] Figure 1 A flowchart of the information deep mining method and system based on the multi-agent collaboration framework provided by the embodiments of the present application is shown. It should be understood that in other embodiments, the order of some steps of the information deep mining method based on the multi-agent collaboration framework of the present embodiment can be shared according to actual needs, or some steps can be omitted or maintained. The detailed steps of the information deep mining method based on the multi-agent collaboration framework include:

[0018] The present embodiment provides an information deep mining method based on a multi-agent collaboration framework. The application scenario is the field of basic scientific literature mining. Specifically, taking the information requirement of “mechanism of action of related molecules of a new type of immunotherapy drug and progress of basic research” as an example, the specific implementation process of the method is described in detail.

[0019] Step S110: receiving an information requirement to be mined, and processing task decomposition and collaboration scheduling of the information requirement to be mined by a planning expert agent, to generate a collaboration task scheduling scheme containing a task hierarchical relationship and an agent collaboration path.

[0020] In the basic scientific literature mining scenario, first, the information requirement to be mined “mechanism of action of related molecules of a new type of immunotherapy drug and progress of basic research” is received. This requirement aims to obtain related basic research information support comprehensively and in depth through literature mining and public database data. The planning expert agent, as the task planning core in the multi-agent collaboration framework, starts to process this requirement. The planning expert agent internally contains a task analysis module, a hierarchical splitting module, an agent matching module, and a path planning module. These modules work together to complete the task decomposition and collaboration scheduling of the requirement.

[0021] Step S111: performing theme element extraction processing on the information requirement to be mined by the planning expert agent, identifying core theme words and associated theme words in the information requirement to be mined, and generating a set of requirement theme words.

[0022] The task analysis module of the planning expert agent first performs theme element extraction on the information requirement to be mined “mechanism of action of related molecules of a new type of immunotherapy drug and progress of basic research”. In the theme element extraction process, the task analysis module performs word segmentation processing on the requirement text, splitting the text into multiple independent word units. Then, through semantic understanding technology, the importance and semantic association of each word unit in the requirement are analyzed. After analysis, the core theme words are identified as “a new type of immunotherapy drug”, “mechanism of action of molecules”, “progress of basic research”, and “cell function impact”. These core theme words are the key elements that constitute the requirement and are directly related to the direction and focus of information mining.

[0023] Next, the task analysis module further mines the associated theme words related to the core theme words. Around "a new type of immunotherapy drug", the associated theme words include the molecular structure of the drug, the action target, the signal pathway regulation, the cell model application, etc.; around "molecular mechanism of action", the associated theme words include the binding mode of the drug and the cell receptor, the signal transduction pathway, the gene expression regulation, etc.; around "basic research progress", the associated theme words include the latest molecular biology research results, functional analysis experimental data, academic conference reports, etc.; around "cell function influence", the associated theme words include cell proliferation, apoptosis, immune cell activation, etc. The core theme words and the associated theme words are integrated together to generate a demand theme word set, which can be expressed as { "a new type of immunotherapy drug", "molecular mechanism of action", "basic research progress", "cell function influence", "molecular structure of a new type of immunotherapy drug", "action target of a new type of immunotherapy drug", "signal pathway regulation", "cell model application", "binding mode of the drug and the cell receptor", "signal transduction pathway", "gene expression regulation", "latest molecular biology research results", "functional analysis experimental data", "academic conference reports", "cell proliferation", "apoptosis", "immune cell activation"}.

[0024] Step S112: Based on the demand theme word set, a task level splitting operation is performed to split the to-be-mined information demand into a plurality of sub-tasks with hierarchical dependency relationship, each of the sub-tasks containing a task target description and an input-output interface definition.

[0025] Based on the generated demand theme word set, the hierarchical splitting module of the planning expert agent starts to perform the task level splitting operation. The hierarchical splitting module will gradually refine the overall demand into executable sub-tasks according to the semantic association and logical hierarchy between the theme words.

[0026] Step S1121: The core theme words in the demand theme word set are taken as root nodes by the planning expert agent to construct a theme word hierarchical tree, the root nodes representing the overall target of the to-be-mined information demand, and the child nodes of the root nodes being the associated theme words.

[0027] The hierarchical splitting module takes the core topic words "a new type of immunotherapy drug", "molecular mechanism", "basic research progress", and "cell function influence" in the demand topic word set as root nodes to construct a topic word hierarchical tree. The child nodes of the root node "a new type of immunotherapy drug" are its associated topic words, such as "molecular structure of a new type of immunotherapy drug", "action target of a new type of immunotherapy drug", and the like; the child nodes of the root node "molecular mechanism" are "binding mode of drug and cell receptor" and "signal transduction pathway"; the child nodes of the root node "basic research progress" are "latest molecular biology research results" and "functional analysis experimental data"; and the child nodes of the root node "cell function influence" are "cell proliferation", "cell apoptosis", and "immune cell activation". The topic word hierarchical tree clearly shows the hierarchical relationship between the topic words and provides a structured basis for subsequent task splitting.

[0028] Step S1122: Based on the hierarchical structure of the topic word hierarchical tree, the total target corresponding to the root node is split into a plurality of first-level subtasks, each first-level subtask corresponding to an associated topic word of a child node, and the target of the first-level subtask is to obtain the basic information related to the associated topic word.

[0029] According to the hierarchical structure of the topic word hierarchical tree, the hierarchical splitting module splits the total target "molecular mechanism and basic research progress related to a new type of immunotherapy drug" into a plurality of first-level subtasks. Each first-level subtask corresponds to an associated topic word of a child node under the root node, for example, the first-level subtask corresponding to "molecular structure of a new type of immunotherapy drug" is "obtaining basic information related to the molecular structure of a new type of immunotherapy drug"; the first-level subtask corresponding to "binding mode of drug and cell receptor" is "obtaining basic information related to the binding mode of drug and cell receptor"; the first-level subtask corresponding to "latest molecular biology research results" is "obtaining basic information related to the latest molecular biology research results"; and the first-level subtask corresponding to "cell proliferation" is "obtaining basic research data on the influence of the drug on cell proliferation". Each first-level subtask clearly indicates that the target is to obtain the basic information of the corresponding associated topic word, laying a foundation for subsequent more in-depth information mining.

[0030] Step S1123: Each first-level subtask corresponding to an associated topic word is split again, and the subdivided topic elements of the associated topic word are extracted, each subdivided topic element being taken as the target of a second-level subtask, and the target of the second-level subtask being to obtain detailed information of the subdivided topic element.

[0031] For each associated subject keyword corresponding to each first-level subtask, the hierarchical splitting module performs secondary splitting to extract the subdivided subject elements. Taking the associated subject keyword "molecular structure of a new immunotherapy drug" corresponding to the first-level subtask "obtain basic information related to the molecular structure of a new immunotherapy drug" as an example, the subdivided subject elements extracted after secondary splitting include "molecular conformation", "active group distribution", "chemical modification mode", etc., and the corresponding second-level subtasks are "obtain detailed information of the molecular conformation of a new immunotherapy drug", "obtain detailed information of the active group distribution", "obtain detailed information of the chemical modification mode", etc. Taking the associated subject keyword "cell proliferation" corresponding to the first-level subtask "obtain basic research data on the effect of a drug on cell proliferation" as an example, the subdivided subject elements after secondary splitting include "cell proliferation rate change", "cell cycle regulation mechanism", "drug dose dependence", etc., and the corresponding second-level subtasks are "obtain detailed information of the effect of a drug on the cell proliferation rate change", "obtain detailed information of the cell cycle regulation mechanism", "obtain detailed information of the effect of drug dose dependence on cell proliferation", etc. The goal of each second-level subtask is to obtain detailed information of the subdivided subject elements, making information mining more specific and in-depth.

[0032] Step S1124: Check the dependency relationship between the first-level subtasks and the second-level subtasks, and when the execution of a second-level subtask needs to depend on the output results of other first-level subtasks, mark the dependency relationship and record the dependency conditions, which include the type of required information and the integrity requirement.

[0033] The hierarchical splitting module checks the dependency relationship between the first-level subtasks and the second-level subtasks. For example, the execution of the second-level subtask "obtain detailed information of the effect of drug dose dependence on cell proliferation" needs to depend on the output results of the first-level subtask "obtain basic information related to the molecular structure of a new immunotherapy drug", because the molecular structure information may affect the dose effect analysis. At this time, the dependency relationship is marked, and the dependency conditions are recorded, the type of required information is "molecular structure information of a new immunotherapy drug", and the integrity requirement is "detailed description covering molecular conformation and active group distribution". For another example, the second-level subtask "obtain detailed information of the signal transduction pathway" may depend on the output results of the first-level subtask "obtain basic information related to the binding mode of a drug and a cell receptor", and the dependency condition is "include molecular interaction details of the binding mode". By explicitly defining the dependency relationship and the dependency conditions, the execution order of the subtasks is ensured to be reasonable, and the task execution is prevented from being blocked due to missing information.

[0034] Step S1125: Generate a subtask list according to the hierarchical order of the subject keyword hierarchical tree and the dependency relationship, each subtask in the subtask list includes task identification, parent task identification, task target description, and dependency conditions, and the subtask list is taken as the subtask splitting result with hierarchical dependency relationship.

[0035] According to the hierarchical order of the topic keyword hierarchical tree and the determined dependency relationship, the hierarchical splitting module generates a subtask list. Each subtask has a unique task identification in the list, such as T1, T2, T3, etc., and contains a parent task identification to indicate its position in the hierarchical structure, for example, the parent task identification of a two-level subtask is the corresponding one-level subtask identification. The task target description specifies the information content that the subtask needs to obtain, and the dependency condition describes the prerequisite information required for the execution of the subtask. For example, the task identification of subtask T1-1 (the first two-level subtask under a one-level subtask) is T1-1, the parent task identification is T1, the task target description is "obtain detailed information of the molecular conformation of a new type of immunotherapy drug", and the dependency condition is "none"; the task identification of subtask T2-2 (the second two-level subtask under another one-level subtask) is T2-2, the parent task identification is T2, the task target description is "obtain detailed information of the effect of drug dosage dependence on cell proliferation", and the dependency condition is "obtain molecular structure related information of a new type of immunotherapy drug, and include detailed description of molecular conformation and active group distribution" and the like. The subtask list clearly presents the hierarchical relationship and execution conditions of all subtasks.

[0036] Step S113: Obtain the capability description information of each functional intelligent agent, which includes the task type that the agent is good at processing, the information processing dimension, and the historical cooperation efficiency parameter.

[0037] The agent matching module of the planning expert agent acquires the capability description information of each functional agent. In this embodiment, the functional agents include a search expert agent, a reading expert agent, an analysis expert agent, and an integration expert agent. The capability description information of the search expert agent is that it is good at processing information search tasks in the field of basic scientific research literature, the information processing dimensions include public academic journal databases, basic research literature databases, molecular biology databases, and the like, and the historical collaboration efficiency parameter is that in the same type of task, the average completion time is within a reasonable range, and the accuracy and coverage of the search results are high. The capability description information of the reading expert agent is that it is good at processing deep analysis and entity association extraction tasks of basic research literature, the information processing dimensions include molecular information, functional experimental data, research conclusions in the literature, and the like, and the historical collaboration efficiency parameter is that the analysis speed of the literature is fast, and the accuracy of entity recognition and relationship extraction meets the requirements of basic scientific research. The capability description information of the analysis expert agent is that it is good at processing deep reasoning tasks of structured associated information, the information processing dimensions include implicit relationships between molecules, biological mechanism trend analysis, and the like, and the historical collaboration efficiency parameter is that the rationality and reliability of the reasoning result are high, and the deep meaning behind the information can be effectively mined. The capability description information of the integration expert agent is that it is good at processing verification and fusion tasks of multi-dimensional information, the information processing dimensions include consistency of associated relationships, integrity of information, and the like, and the historical collaboration efficiency parameter is that the accuracy of the fused information is high, and meets the needs of basic scientific research literature mining.

[0038] Step S114: Calculate the matching degree of the task target description of each subtask and the capability description information of each functional agent, generate a target matching matrix, and the elements in the target matching matrix represent the adaptation degree of the target.

[0039] The agent matching module calculates the matching degree between the task target description of each subtask and the ability description information of each functional agent. The matching degree calculation mainly considers three aspects: the consistency of task type, the coverage of information processing dimension, and the adaptability of historical cooperation efficiency. The consistency of task type is used to measure the consistency between the type of subtask and the type of task that the agent is good at handling; the coverage of information processing dimension is used to measure whether the information processing dimension of the agent can cover the information dimension required by the subtask; the adaptability of historical cooperation efficiency is used to refer to the matching between the historical performance of the agent in similar tasks and the current subtask. By comprehensively considering these three factors, the matching degree between each subtask and each functional agent is calculated, and a target matching matrix is generated. For example, the subtask "obtain basic information related to the molecular structure of a new immunotherapy drug" has a high matching degree with the search expert agent, because this task is mainly of the information search type, which is consistent with the type of task that the search expert agent is good at handling, and the information processing dimension of the search expert agent can cover the information sources required by the task; while the matching degree between this subtask and the analysis expert agent is relatively low, because the analysis expert agent is better at deep reasoning rather than basic information acquisition. The elements in the target matching matrix represent the adaptability degree in a specific symbol, such as high adaptability, medium adaptability, and low adaptability, which clearly shows the matching between subtasks and agents.

[0040] Step S115: Based on the target matching matrix and the hierarchical dependency relationship between subtasks, a cooperation path planning model is constructed, and the cooperation path planning model takes minimizing the total execution time of the task as the objective function.

[0041] The path planning module of the planning expert agent constructs a cooperation path planning model based on the target matching matrix and the hierarchical dependency relationship between subtasks. The objective function of the cooperation path planning model is to minimize the total execution time of the task, because in the field of basic scientific research, timely acquisition of required information is of great significance for researchers to carry out research and promote projects. The constraint conditions of the model include the hierarchical dependency relationship between subtasks, i.e., a subtask can only be executed after the completion of the parent task or the satisfaction of the dependency condition; and the load balancing of agents to avoid a certain agent from undertaking too many tasks and affecting the overall efficiency. The model also needs to consider the cooperation connection time between agents to ensure that an agent can smoothly take over the next related task after completing the current task. Through these constraint conditions and objective functions, a reasonable cooperation path planning model is constructed.

[0042] Step S116: Solve the cooperation path planning model to determine the execution agent corresponding to each subtask and the cooperation order between agents, and generate a cooperation task scheduling scheme containing the task hierarchical relationship and the cooperation path of agents.

[0043] The path planning module solves the cooperative path planning model by using a step-by-step iteration method. First, according to the hierarchical dependency relationship between sub-tasks, the general execution order framework of the sub-tasks is determined, that is, the parent task is executed first, and then the sub-task. Then, in combination with the target matching matrix, the function intelligent agent with the highest matching degree is assigned to each sub-task, while considering the load of the intelligent agent. If a certain intelligent agent is assigned too many tasks, the intelligent agent assignment of part of the sub-tasks is appropriately adjusted to ensure load balancing. After determining the correspondence between the sub-tasks and the intelligent agents, the cooperation order between the intelligent agents is further refined to determine which intelligent agent executes first and which intelligent agent executes later, as well as the nodes and ways of information transmission between the intelligent agents. Through continuous optimization and adjustment, the execution intelligent agent corresponding to each sub-task and the cooperation order between the intelligent agents are finally determined, and a cooperative task scheduling scheme is generated. The scheme lists the hierarchical relationship of the tasks, that is, which are parent tasks and which are sub-tasks, as well as their dependency relationship; at the same time, the cooperation path of the intelligent agent is clear, that is, which sub-tasks each intelligent agent is responsible for, and how the intelligent agents cooperate to complete the entire information mining task. For example, the cooperative task scheduling scheme will stipulate that the search expert intelligent agent first executes the search sub-task of basic information, and then transmits the result to the reading expert intelligent agent to execute the parsing and association extraction sub-task, and then the reading expert intelligent agent transmits the result to the analysis expert intelligent agent to execute the deep reasoning sub-task, and finally the integration expert intelligent agent executes the verification and fusion sub-task.

[0044] Step S120: Based on the cooperative task scheduling scheme, the search expert intelligent agent is called to execute the multi-source information cooperative search operation to obtain multi-source original information containing multi-dimensional information sources.

[0045] According to the generated cooperative task scheduling scheme, the system calls the search expert intelligent agent to start executing the multi-source information cooperative search operation. In the basic scientific research literature mining scene, multi-source information is crucial to a comprehensive understanding of "the mechanism of related molecules of a new type of immunotherapy drug and the progress of basic research", because different information sources may provide information from different angles and depths, which can more fully meet the needs of researchers. The search expert intelligent agent will acquire relevant information from multiple information sources according to the requirements in the cooperative task scheduling scheme for each search sub-task.

[0046] Step S121: Analyze the intelligent agent cooperation path in the cooperative task scheduling scheme to determine the search expert intelligent agent responsible for information search and the corresponding search sub-task, each search sub-task containing search topics and information source type parameters.

[0047] The search expert agent first parses the agent cooperation path in the cooperation task scheduling scheme, and determines its role and responsible search subtasks in the cooperation process. Through parsing, it is found that the search expert agent needs to be responsible for all subtasks related to basic information acquisition, which are search subtasks. Each search subtask contains a clear search topic and information source type parameter. For example, the search topic of a search subtask is "the binding mode of a new immunotherapy drug and cell receptors", and the information source type parameter is "public academic journal database, basic research literature database"; the search topic of another search subtask is "functional analysis data of the effect of a drug on cell proliferation", and the information source type parameter is "molecular biology database, public experimental data set"; the search topic of another search subtask is "the latest molecular biology research results related to a new immunotherapy drug", and the information source type parameter is "international academic conference proceedings, recently published basic research journals", etc.

[0048] Step S122: calling the search expert agent, according to the search topic and information source type parameter of the search subtask, accessing the preset search strategy library, generating a search rule set for each information source type, the search rule set including keyword combination method, retrieval field limitation and result filtering condition.

[0049] After receiving the search topic and information source type parameter of the search subtask, the search expert agent will automatically access the preset search strategy library. The search strategy library stores general search strategy templates for different information source types in the field of basic scientific research. These templates are based on a large number of basic scientific research information search practices. For example, the search strategy template for public academic journal databases includes keyword expansion methods suitable for this database, field retrieval priority, etc.; the template for molecular biology databases focuses on molecular structure, gene expression, signal pathway, etc.

[0050] The search strategy generation module will adjust the template according to the search topic. Taking the search topic "the binding mode of a new immunotherapy drug and cell receptors" and the information source type parameter "public academic journal database, basic research literature database" as an example, the module first extracts the core keywords "a new immunotherapy drug", "cell receptors", and "binding mode" from the search topic. Then, based on the strategy template for public academic journal databases, generate keyword combination methods, such as using "a new immunotherapy drug AND cell receptors AND binding mode" as the basic combination, and expand "a new immunotherapy drug OR its generic name AND cell receptors OR receptor proteins AND binding mechanism OR interaction" as alternative combinations to improve the comprehensiveness of the search.

[0051] For the search field definition aspect, for the public academic journal database, the search will be limited to the title, abstract, keyword, and experimental method fields, as these fields usually contain the core content and methodological information of the research, which can more accurately locate the literature related to "binding modes". For the basic research literature database, in addition to the above-mentioned fields, the "results and analysis" field will also be added to the search, as this field in basic research literature may describe in detail the process of drug and cell interaction.

[0052] The setting of the result filtering condition takes into account the rigor of the basic research field, for example, limiting the publication time of the literature to recent years to ensure that the information obtained is timely; limiting the language of the literature to Chinese or English, as these two languages have high recognition and reference value in international basic research literature; at the same time, filtering out review literature that only mentions but does not deeply explore the binding mode, and retaining research literature that contains experimental data or mechanism analysis. Through the above processing, a search rule set for this search subtask is generated.

[0053] Step S123: Based on the search rule set, perform parallel search on multi-dimensional information sources, real-time monitor the execution state of each search thread, record the search response time and the number of preliminary results.

[0054] The search execution module of the search expert agent will start the parallel search mechanism after obtaining the search rule set of each information source type, and simultaneously search the multi-dimensional information sources.

[0055] Step S1231: According to the number of search rule sets and the access characteristics of each information source, allocate an independent search thread for each search rule set, and set the maximum execution time and resource occupation threshold of the search thread.

[0056] The search execution module first counts the number of search rule sets, each search rule set corresponding to an information source type or a group of related information sources. At the same time, analyze the access characteristics of each information source, such as access speed, server load, data update frequency, etc. For example, academic journal full-text databases usually have large access volume, and server response speed may fluctuate due to time period; while basic biomedical literature databases have relatively small access volume and stable response.

[0057] According to these situations, an independent search thread is allocated for each search rule set. Each search thread has a unique identifier to facilitate subsequent monitoring and management. At the same time, the maximum execution time of the search thread is set, which is determined according to the historical response of the information source and the urgency of the search task, to ensure that the search is completed within a reasonable time and to avoid long-term idle occupation of resources. The resource occupation threshold includes the memory space, network bandwidth, etc. that the thread can use, to prevent a single thread from occupying too many resources and affecting the normal operation of other threads.

[0058] Step S1232: Start all search threads, and each search thread generates a search request according to the corresponding search rule set, which includes keyword combinations, search fields, and filtering conditions, and sends it to the corresponding information source platform through a preset interface protocol.

[0059] The search execution module starts all search threads in the order of the allocated threads. After each search thread is started, it generates a specific search request according to the keyword combination method, search field limitation, and result filtering condition in the corresponding search rule set. The format of the search request meets the interface requirements of the corresponding information source platform. For example, for an academic journal database that uses RESTful API interface, the search request is built in the form of HTTP request, including request header, request parameter, etc., and the request parameter is the processed keyword combination, search field, and filtering condition.

[0060] The search request is sent to the corresponding information source platform through a preset interface protocol, such as HTTPS protocol. During the sending process, the thread records the time point of the request sending, so as to calculate the response time later. For example, the thread responsible for searching the academic journal full-text database generates a request containing the keyword combination "a new type of immunotherapy drug AND immune cells AND binding method" and related field limitations, and sends it out through the interface protocol of the database.

[0061] Step S1233: Real-time collection of execution state parameters of each search thread, including current search progress, number of returned results, network connection state, and response delay time.

[0062] The thread monitoring module of the search expert intelligent agent collects the execution state of each search thread in real time. The current search progress is reflected by the ratio of the completed search steps to the total steps. For example, for information sources that require batch retrieval of results, the progress will be updated as the batches are completed. The number of returned results refers to the number of literature entries that meet the preliminary filtering conditions returned by the information source platform, which is periodically counted and updated by the thread.

[0063] The network connection state is determined by detecting whether the connection between the thread and the information source platform is stable. If the connection is interrupted, it will be marked as an abnormal state in time. The response delay time refers to the time interval from sending a retrieval request to first receiving a returned result. The thread continuously monitors and records this time to evaluate the response speed of the information source platform. These execution state parameters are transmitted in real time to the state storage unit of the thread monitoring module.

[0064] Step S1234: When it is detected that the response delay time of any one search thread exceeds the preset threshold, a priority promotion instruction is sent to the search thread, and the resource allocation weight of the search thread is adjusted to improve its data receiving priority.

[0065] The thread monitoring module compares the response delay time of each search thread with the preset threshold in real time. When it is found that the response delay time of a certain search thread exceeds the threshold, it means that the information source platform corresponding to the thread may have high load or network congestion, etc. At this time, the monitoring module sends a priority promotion instruction to the search thread.

[0066] After receiving the instruction, the search execution module adjusts the resource allocation weight of the search thread, increases its available network bandwidth and memory resources, so that it can obtain resource support in the data receiving stage and speed up the receiving speed of the result. For example, the network bandwidth proportion originally allocated to the thread is a certain proportion, and after adjustment, this proportion is appropriately increased to ensure that it can receive the returned search result faster.

[0067] Step S1235: When it is detected that the network connection state of any one search thread is abnormal, the thread retry mechanism is triggered to re-establish the network connection and continue to execute the retrieval request, and the retry number and recovery time are recorded.

[0068] If the thread monitoring module detects that the network connection state of a certain search thread is abnormal, such as connection interruption, timeout, etc., the thread retry mechanism will be triggered immediately. First, the search execution module will terminate the current connection process of the thread and release part of the resources occupied by it. Then, according to the preset retry interval time, it reinitiates the network connection request and tries to re-establish the connection with the information source platform.

[0069] Each retry operation will be recorded in the retry log, including the time point of retry, the number of retries, etc. When the connection is successfully restored, the thread will continue to execute the retrieval request from the breakpoint to ensure the continuity of the search task. At the same time, the connection recovery time is recorded to analyze the influence range and duration of network anomalies in the future. For example, if the thread responsible for accessing the basic research literature database has a connection interruption, it will try to reconnect at regular intervals after triggering the retry mechanism until the connection is restored and the search result is continued to be received.

[0070] Step S1236: Real-time update the execution state parameters of each search thread to the search state dashboard, which is used to display the real-time progress and abnormal conditions of all threads, and support dynamic adjustment of thread execution strategy.

[0071] The thread monitoring module will real-time collect the execution state parameters of each search thread, such as current search progress, number of returned results, network connection state, response delay time, etc., and real-time update them to the search state dashboard. The search state dashboard is a visual monitoring interface that displays the running conditions of all threads in a combination of charts and text. For example, use a progress bar to represent the search progress of each thread, use different colors to mark the network connection state (green for normal, red for abnormal), and use numbers to display the number of returned results and response delay time.

[0072] The operator of the search expert agent or the automatic adjustment module of the system can real-time master the search progress through the search state dashboard. When multiple threads appear response delay or abnormality, the cause can be analyzed in time, and the thread execution strategy can be dynamically adjusted, such as temporarily increasing the resource allocation of part of the threads, adjusting the keyword combination in the search rule set, etc., to optimize the overall search efficiency.

[0073] Step S124: When potential duplication is detected in the information returned by different search threads, dynamically adjust the search path through the search expert agent, and modify the keyword combination method or search field limit of the conflict path.

[0074] The duplicate information detection module of the search expert agent will real-time compare the preliminary results returned by each search thread. The basis for comparison includes the title, author, and core content of the abstract of the literature. When it is found that there are two or more articles in the information returned by different search threads that are highly similar in these core elements, it is determined that there is potential duplication.

[0075] For example, the thread responsible for searching the full-text database of academic journals and the thread responsible for searching the basic biomedical conference paper set may have returned the same research team's literature on "a new type of immune therapy drug and immune cell combination method", but the carrier of publication is different, but the core content is consistent, which is marked as potential duplication.

[0076] In view of the above, the path adjustment module of the search expert agent analyzes the conflicting search paths to determine the cause of the repetition. If the cause is that the keyword combination method is too broad, the keyword combination method of the conflicting path will be modified to add more specific qualifiers, such as "in vitro experiment" or "molecular level", to narrow the search range. If the cause is that the search field is not precise enough, the search field will be adjusted, such as searching only in the title and keyword fields, to reduce the repetitive information retrieved from the abstract field. Through the above dynamic adjustment, the information repetition rate is reduced and the effectiveness of the search results is improved.

[0077] Step S125: Collect the initial search results returned by each search thread, perform source labeling processing on the initial search results, record the source type and acquisition timestamp of each information segment, and integrate the labeled information segments into multi-source original information containing multiple dimensions of information sources.

[0078] The result collection module of the search expert agent continuously receives the initial search results returned by each search thread, which exist in the form of literature entries, abstracts, full-text segments, etc. The result collection module performs source labeling processing on each information segment, adds a source type label to the information segment according to the information source type parameter associated with the search thread, such as "academic journal full-text database", "basic research literature database", "biomedical conference proceedings", etc.

[0079] At the same time, the acquisition timestamp of each information segment is recorded, accurate to the second, to trace the acquisition order and timeliness of the information subsequently. After labeling is completed, the result integration module will preliminarily classify all labeled information segments according to the search topic, and information segments under the same search topic will be stored together. For example, all information segments related to "the combination method of a new type of immunotherapy drug and immune cells" are integrated together, and information segments related to "the molecular mechanism and functional impact of a new type of immunotherapy drug" are integrated together. Finally, multi-source original information containing multiple dimensions of information sources is formed.

[0080] Step S130: The reading expert agent performs deep analysis and correlation extraction processing on the multi-source original information to generate structured correlation information with entity correlation networks.

[0081] In the basic scientific research literature mining scenario, the multi-source original information contains a large amount of scattered and unstructured content, such as literature abstracts and experimental data descriptions, which need to be processed by the reading expert agent to extract key information and establish correlations, so as to provide valuable structured content for researchers. The reading expert agent has natural language processing and entity relationship recognition capabilities, and can deeply analyze text information and mine the internal relationships between entities.

[0082] Step S131: Input the multi-source raw information into the reading expert agent, and perform similarity calculation on the multi-source raw information to identify and remove completely duplicated information segments.

[0083] The reading expert agent converts each information segment into a feature vector that can be used for similarity calculation. The generation of the feature vector is based on the distribution of words in the information segment, semantic features, etc., and is achieved through processing steps such as tokenization, removal of stop words (such as "of", "in", etc. words without actual meaning), extraction of keywords, etc.

[0084] Then, the cosine similarity calculation method is used to calculate the similarity of the feature vectors of any two information segments. When the similarity of two information segments reaches or exceeds the pre-set threshold, it is determined to be completely duplicated. For example, two information segments both quote the abstract of the same article completely and have no differences, and their similarity will be high, which will be identified as duplicate information.

[0085] The deduplication processing module retains one of the information segments and removes the other duplicate segments. The removal process is recorded in the deduplication log, including the source type, acquisition timestamp, etc. of the removed segment, for subsequent tracing. Through this step, information redundancy is reduced and the efficiency of subsequent processing is improved.

[0086] Step S132: Perform structured parsing on the deduplicated information segments to identify core entities in the information segments, including researcher entities, institution entities, and concept entities, and record the basic attribute information of the core entities.

[0087] The entity recognition module of the reading expert agent performs structured parsing on the deduplicated information segments. This module uses an entity recognition model based on a pre-trained language model, which has been trained on a large amount of basic scientific research text corpus and can accurately identify various entities in the text.

[0088] For researcher entities, the identified content includes the name of the researcher, the institution to which they belong, and their research field, etc. For example, in the author introduction section of a paper, "Zhang, a molecular biology institute, mainly researches immune regulation mechanism" is identified, and these basic attribute information is recorded.

[0089] The identification of institution entities includes the name of the research institution, its location, and its research expertise, etc. For example, from the information segment, "a life science research institute, located in a city, has deep accumulation in the research of immune therapy molecular mechanism" is identified, and the relevant attributes are recorded.

[0090] Conceptual entities are the core of basic scientific research, including drug entities (such as "a new immunotherapy drug"), cell model entities (such as "T cell lines"), molecular mechanism entities (such as "activation of signal pathways"), and experimental technology entities (such as "flow cytometry"). For drug entities, record attributes such as generic name, research and development company, and target point; for cell model entities, record attributes such as source, type, and characteristics; for molecular mechanism entities, record attributes such as regulatory pathways and influencing factors; for experimental technology entities, record attributes such as principles and application scope.

[0091] The entity recognition module stores the identified core entities and their basic attribute information in the entity attribute database, laying the foundation for subsequent relationship extraction.

[0092] Step S133: Analyze the semantic relationships between the core entities in the information segment, extract the association type and association strength parameters between the entities, and the association type represents the action relationship between the entities, and the association strength parameter represents the degree of explicitness of the relationship.

[0093] The relationship extraction module of the reading expert agent is responsible for analyzing the semantic relationships between the core entities in the information segment. This module is also based on a pre-trained language model and combines a knowledge graph of the basic scientific research field to identify and extract relationships between entities.

[0094] Step S1331: Perform syntactic analysis on the information segment to identify the subject-predicate-object structure and the adverbial structure in the information segment, and determine the grammatical role of the core entity in the sentence.

[0095] The relationship extraction module first performs syntactic analysis on the information segment, using dependency syntactic analysis techniques to identify grammatical components such as subject-predicate-object structure and adverbial structure in the sentence. For example, in the sentence "a new immunotherapy drug affects cell function by activating the molecular signal pathway of T cells", the subject-predicate-object structure is "a new immunotherapy drug (subject) affects (predicate) cell function (object)", and the adverbial structure is "by activating the molecular signal pathway of T cells (adverb) modifying the predicate 'affects'".

[0096] Through syntactic analysis, the grammatical role of the core entity in the sentence is determined, such as "a new immunotherapy drug" being the subject, "cell function" being the object, and "the molecular signal pathway of T cells" being the core word in the adverbial structure. These grammatical role information helps to determine the potential relationship between entities.

[0097] Step S1332: Based on the grammatical role of the core entity, a pre-trained relationship classification model is called to preliminarily classify the semantic relationship between the entity pair, obtaining a candidate association type set, which contains multiple possible association types and corresponding initial confidence.

[0098] The relationship extraction module constructs entity pairs according to the syntactic roles of the core entities, such as the entity pair of "a new type of immunotherapy drug" and "cell function", and the entity pair of "a new type of immunotherapy drug" and "T cell signaling pathway".

[0099] Then, a pre-trained relationship classification model is called, which takes the entity pair and its context information in the sentence as input and outputs a set of candidate association types. The relationship classification model uses a large number of basic scientific texts annotated with entity relationships in the training process and can identify a variety of common association types, such as "activation", "inhibition", "combination", "regulation", and "associated with".

[0100] For example, for the entity pair "a new type of immunotherapy drug" and "cell function", the model may output a set of candidate association types { "enhancement, high initial confidence", "inhibition, medium initial confidence"}; for the entity pair "a new type of immunotherapy drug" and "T cell signaling pathway", the output is { "activation, high initial confidence", "combination, high initial confidence"}. The initial confidence reflects the reliability of the model's judgment of the association type, which is determined based on the classification accuracy of the model on the training data.

[0101] Step S1333: Verify the set of candidate association types in combination with the context of the information segment, check whether the association type is consistent with the context semantics, and eliminate candidate association types that are inconsistent with the context.

[0102] The relationship extraction module will verify each association type in the set of candidate association types in combination with the context of the information segment. The context includes the context before and after the sentence, the paragraph theme, the research purpose of the literature, and the like.

[0103] For example, for the candidate association type "inhibition" of the entity pair "a new type of immunotherapy drug" and "cell function", in combination with the context "the drug enhances immune cell function and specifically regulates signaling pathways", it can be determined that "inhibition" is inconsistent with the context semantics and will be eliminated.

[0104] For example, for the candidate association type of the entity pair "a new type of immunotherapy drug" and "T cell signaling pathway", "activation", "inhibition", and "no direct effect", combined with the context "the drug binds to a specific receptor on the surface of T cells, promoting the activation of the signaling pathway", it can be determined that "activation" is consistent with the context semantics; while "inhibition" and "no direct effect" are contradictory to the context description and will be eliminated. If the context is "research finds that the drug temporarily blocks the signal transduction pathway at high concentrations", then "inhibition" is consistent with the context semantics in this context, and "activation" and "no direct effect" will be eliminated due to contradiction with the context. Through the above verification method combined with the context, it is ensured that the retained candidate association type is consistent with the overall semantics of the information segment.

[0105] Step S1334: Frequency statistics are performed on the remaining candidate association types, the occurrence frequency of each association type in the information segment is calculated, and the association type with the highest occurrence frequency is taken as the final association type between entities.

[0106] After eliminating the candidate association types that are contradictory to the context, the relationship extraction module performs frequency statistics on the remaining candidate association types. The statistical process covers all sentences in the current information segment that involve the entity pair, and records the number of occurrences of each remaining candidate association type. For example, for the entity pair "a new type of immunotherapy drug" and "cell function", after context verification, the remaining candidate association types are "enhancement" and "no significant effect". In the information segment, "enhancement" occurs more frequently than "no significant effect", so "enhancement" is taken as the final association type of the entity pair. By frequency statistics to determine the final association type, the general acceptance of the association type in the information segment can be reflected, and the reliability of the association relationship is enhanced.

[0107] Step S1335: Based on the initial confidence of the final association type and the context verification result, the association strength parameter is calculated, the association type and the association strength parameter between entities are generated, and the value of the association strength parameter is positively related to the initial confidence and negatively related to the number of semantic conflicts found in the context verification process.

[0108] After determining the final association type, the relationship extraction module calculates the association strength parameter. In the calculation, the initial confidence degree of the final association type is taken as the basis, which is an index output by the relationship classification model, indicating the possibility of the association type being the real relationship of the entity pair. At the same time, the number of semantic conflicts in the context verification result is combined, which refers to the number of sentences preliminarily considered to be in conflict with the association type in the context verification process. The value of the association strength parameter increases with the increase of the initial confidence degree and decreases with the increase of the number of semantic conflicts. For example, for the entity pair "a new type of immunotherapy drug" and "T cell function", the final association type is "enhancing effect", the initial confidence degree is high, and no semantic conflict is found in the context verification process, so the value of the calculated association strength parameter is relatively large; if the final association type of another entity pair "a new type of immunotherapy drug" and "tumor cell volume" is "shrinking effect", the initial confidence degree is moderate, and a small amount of semantic conflicts are found in the context verification process, so the value of the association strength parameter will be smaller than that of the former. The association strength parameter calculated in the above manner can comprehensively reflect the reliability degree of the association type.

[0109] Step S134: based on the extracted core entities and the association types between entities, an entity association network is constructed, the entity association network taking the core entities as nodes, the association types as edge labels, and the association strength parameters as edge weights.

[0110] After extracting the core entities, the association types between entities and the association strength parameters, the network construction module starts to construct the entity association network. First, all the core entities are taken as independent nodes in the network, and each node contains the basic attribute information of the core entity, such as the node of "a new type of immunotherapy drug" containing its drug category, target point, etc. The node of "cell model" contains the model type, culture condition, etc. Then, according to the association types between entities, a connection edge is established between the corresponding two nodes, and each edge takes the association type as a label, indicating the relationship property between the two entities, for example, the edge label between "a new type of immunotherapy drug" and "T cell function" is "enhancing effect", and the edge label between "a new type of immunotherapy drug" and "cytotoxicity reaction" is "inducing effect". Finally, the association strength parameter is taken as the weight of the corresponding edge, and the size of the weight reflects the reliability and strength of the association relationship. Through the above construction method, an entity association network is formed, which takes the core entities as nodes, the association types as edge labels, and the association strength parameters as edge weights, and can directly show the relationship and strength between the core entities.

[0111] Step S135: Redundant edge simplification processing is performed on the entity association network. When there are multiple association edges of the same type between two core entities, the edge with the highest association strength parameter is retained as the main association edge, and the association strength parameters of the other edges are accumulated to the main association edge to generate the structured association information with the entity association network.

[0112] The network optimization module performs redundant edge simplification processing on the entity association network to improve the simplicity and effectiveness of the network. When it is detected that there are multiple association edges of the same type between two core entities, i.e., the edge labels of these edges are the same, indicating that they express the same type of association relationship, simplification processing is required. The processing method is to compare the association strength parameters of these association edges of the same type, retain the edge with the highest association strength parameter as the main association edge, and then accumulate the association strength parameters of the other edges to the main association edge. For example, between "a new type of immunotherapy drug" and "cell proliferation in cell models", there are three association edges with the same edge label "promoting effect", and their association strength parameters are different. After comparison, the one with the highest association strength parameter is retained as the main association edge, and the association strength parameters of the other two are accumulated to the main association edge, so that the association strength parameter of the main association edge is updated. Through the above redundant edge simplification processing, the same association relationship is avoided from being represented repeatedly, and the strength information of multiple association edges is integrated, making the entity association network more refined and the information more concentrated. The entity association network after simplification processing is the structured association information with the entity association network, which clearly presents the association relationships and strengths between core entities.

[0113] Step S140: Calling the analysis expert agent to perform multi-level deep reasoning processing on the structured association information to obtain deep mining intermediate results containing implicit relationships.

[0114] In the basic scientific research literature mining scenario, although the structured association information has presented the explicit association relationships between core entities, in order to deeply mine the molecular mechanisms and cell functions, the implicit relationships therein also need to be mined. The analysis expert agent has deep reasoning ability and can perform multi-level deep reasoning based on the entity association network in the structured association information, thereby discovering those association relationships that are not directly expressed but objectively exist, and generating deep mining intermediate results.

[0115] Step S141: Calling the analysis expert agent to perform hierarchical division on the entity association network in the structured association information, and dividing the core entities into multiple theme levels according to theme relevance, each theme level containing multiple entity subsets with close association.

[0116] The hierarchical division module of the analysis expert agent first divides the entity association network in the structured association information into hierarchical levels. The division is based on the theme relevance between core entities, which is embodied by the closeness of association between entities. The closer the association, the higher the theme relevance. In the hierarchical division process, a core theme is first determined as the starting point of the highest level, and then multiple theme levels are gradually divided according to the association degree of entities with the core theme and the mutual association between entities. Each theme level contains multiple entity subsets, and the entities within these entity subsets have close associations and certain theme relevance with other entity subsets in the same theme level. For example, taking "molecular mechanism related to a new type of immunotherapy drug" as the core theme, the highest level can be divided into "drug molecular structure related entities", "molecular target point related entities", and "cell function response related entities". Among them, the "drug molecular structure related entities" subset includes "a new type of immunotherapy drug", "drug molecular structure", "drug chemical properties", etc.; the "molecular target point related entities" subset includes "target point protein", "signal pathway molecule", "receptor", etc.; and the "cell function response related entities" subset includes "cell proliferation rate", "cell apoptosis", "immune cell activity", etc. Through the above hierarchical division, the entity association network presents a clear theme structure.

[0117] Step S142: Based on the theme level division result, cross-level association analysis is performed on entities between different theme levels to identify indirect association paths between entity subsets and calculate the cumulative association strength parameter of the indirect association paths.

[0118] The association path analysis module performs cross-level association analysis on entities between different theme levels based on the theme level division result. The analysis process first selects entity subsets from different theme levels, and then searches for indirect association paths connecting these different level entity subsets in the entity association network, i.e. paths connecting entities of different levels through one or more intermediate entities. For example, between "a new type of immunotherapy drug" in the "drug molecular structure related entities" level and "cell proliferation rate" in the "cell function response related entities" level, there may be an indirect association path "a new type of immunotherapy drug-activates signal pathway-regulates transcription factor-controls cell cycle-cell proliferation rate changes", which connects two entities of different levels through multiple intermediate entities.

[0119] After the indirect association path is identified, the cumulative association strength parameter of the path is calculated by multiplying the association strength parameters of the association edges in the path in sequence, and the result is the cumulative association strength parameter of the indirect association path, which reflects the overall association strength of the entire indirect association path. Through the above cross-level association analysis, the potential connection between different theme level entities can be found, which expands the depth and breadth of information mining.

[0120] For example, step S1421: selecting two adjacent theme levels from the theme level division result as the current analysis level, the upper theme level contains a high-level entity subset, and the lower theme level contains a low-level entity subset.

[0121] The association path analysis module selects two adjacent theme levels from the theme level division result as the current analysis level. Adjacent theme levels refer to two levels that have a direct superior-inferior relationship in the hierarchical structure, wherein the upper theme level contains a relatively macro or core high-level entity subset, and the lower theme level contains a relatively specific or subordinate low-level entity subset. For example, the "cell function response related entity" is selected as the upper theme level, which contains a high-level entity subset of "cell proliferation rate", "cell apoptosis rate", etc.; the "molecular action target point related entity" is selected as the lower theme level, which contains a low-level entity subset of "signal pathway molecule", "receptor protein", "transcription factor", etc. The two levels are closely related in theme, and the entity state of the upper theme level is directly or indirectly affected by the entity of the lower theme level. Selecting them as the current analysis level facilitates in-depth analysis of the indirect association path between the two.

[0122] The two levels are closely related in theme, and the entity state of the upper theme level is directly or indirectly affected by the entity of the lower theme level. Selecting them as the current analysis level facilitates in-depth analysis of the indirect association path between the two.

[0123] Step S1422: taking each entity in the upper theme level as a starting point and each entity in the lower theme level as an end point, performing path search in the entity association network to find all possible paths connecting the starting point and the end point, and the intermediate nodes in the searched path can be entities of the same level or other levels.

[0124] After determining the upper and lower topic hierarchy of the current analysis, the association path analysis module performs path search in the entity association network, taking each entity in the upper topic hierarchy as the starting point and each entity in the lower topic hierarchy as the ending point. The search process adopts a step-by-step expansion manner, starting from the starting point entity, and according to the association edges in the entity association network, sequentially finding entities directly associated with the starting point entity as first-level intermediate nodes, and then starting from the first-level intermediate nodes, finding entities directly associated with them as second-level intermediate nodes, and so on, until a path that can connect to the ending point entity is found. The intermediate nodes in the searched path can be entities in the upper and lower topic hierarchy of the current analysis, or entities in other topic hierarchies. For example, the starting point entity in the upper topic hierarchy is "cell proliferation rate", and the ending point entity in the lower topic hierarchy is "the molecular structure of a new type of immunotherapy drug". In the search process, the path that can be found includes "cell proliferation rate-regulatory protein-signal pathway-the molecular structure of a new type of immunotherapy drug", where "regulatory protein" belongs to the "molecular action target related entity" hierarchy, and "signal pathway" belongs to the "molecular action target related entity" hierarchy, which connects the starting point and the ending point as intermediate nodes. Through the above path search, all possible paths connecting the upper and lower topic hierarchy entities can be comprehensively found.

[0125] Step S1423: Length filtering is performed on all the searched paths, and paths with a path length less than a preset maximum value are retained as candidate indirect association paths, the path length being the number of edges contained in the path.

[0126] The path screening module performs length filtering on all the searched paths to exclude those paths that are too long and may not have much association significance. The path length is measured by the number of edges contained in the path, and each edge represents an association relationship. The longer the path length, the more indirect the association between entities. The preset maximum value is set according to the information association characteristics and reasoning requirements in the underlying scientific research literature, and is used to define the upper limit of reasonable path length. The searched paths are checked for length, and paths with a path length less than the preset maximum value are retained as candidate indirect association paths, and paths with a path length greater than or equal to the preset maximum value are excluded. For example, if the preset maximum value is 4, then the path "cell function response-regulatory signal-molecular target-drug molecular structure" with a path length of 3 will be retained as a candidate indirect association path, and the path with a path length of 5 will be excluded. Through length filtering, the number of paths for subsequent processing is reduced, and at the same time it is ensured that the retained candidate indirect association paths have certain actual association significance.

[0127] Step S1424: calculating a cumulative correlation strength parameter of each candidate indirect correlation path, the cumulative correlation strength parameter being a product of the correlation strength parameters of all edges in the candidate indirect correlation path, and the larger the product result is, the higher the overall correlation strength of the candidate indirect correlation path is.

[0128] The strength calculation module calculates the cumulative correlation strength parameter of each candidate indirect correlation path. The calculation manner is to multiply the correlation strength parameters of all edges contained in the candidate indirect correlation path in sequence, and the product result obtained is the cumulative correlation strength parameter of the candidate indirect correlation path. Since the correlation strength parameter reflects the strength of the correlation relationship of a single edge, the cumulative correlation strength parameter can comprehensively reflect the overall correlation strength of the entire path. The larger the product result is, the higher the overall correlation strength of the candidate indirect correlation path is, that is, the more significant the indirect correlation between the start entity and the end entity through the path is. For example, in the candidate indirect correlation path "a new type of immunotherapy drug-activates signal pathway-regulates transcription factor-controls cell cycle-cell proliferation rate", the correlation strength parameters of each edge are a certain value. The result obtained by multiplying them is the cumulative correlation strength parameter of the path. If the result is large, it indicates that there is a strong indirect correlation between the new type of immunotherapy drug and the cell proliferation rate through this path.

[0129] Step S1425: performing deduplication processing on the multiple candidate indirect correlation paths of the same pair of start-end entities to generate an indirect correlation path set between the entity subsets.

[0130] The path deduplication module performs deduplication processing on the multiple candidate indirect correlation paths of the same pair of start-end entities. When there are multiple candidate indirect correlation paths between the same pair of start-end entities, and these paths express the same or extremely similar correlation relationship in semantics, deduplication needs to be performed to avoid information redundancy. The deduplication processing manner is to compare the semantic connotations of these paths, identify the paths expressing the same correlation relationship, retain the path with the highest cumulative correlation strength parameter, and eliminate other paths. For example, for the start entity "a new type of immunotherapy drug" and the end entity "cell proliferation rate", there are two candidate indirect correlation paths, one is "a new type of immunotherapy drug-activates signal pathway-A regulatory protein-cell proliferation rate increases", and the other is "a new type of immunotherapy drug-regulates transcription factor-B expression-cell proliferation rate increases". If it is judged through semantic analysis that the two paths express the same correlation relationship that the drug affects cell proliferation through the regulation of molecular signal pathways, the path with the higher cumulative correlation strength parameter is retained. After deduplication processing, all the candidate indirect correlation paths obtained form an indirect correlation path set between the entity subsets. The indirect correlation path set contains indirect correlation paths of different semantics between the same pair of start-end entities.

[0131] Step S143: performing implicit relationship extraction on the indirect association path whose cumulative association strength parameter exceeds the preset threshold, to determine the implicit association type between the two entities at the ends of the indirect association path, the implicit association type representing a relationship that is not directly expressed in the original information but can be obtained through reasoning.

[0132] The implicit relationship extraction module performs implicit relationship extraction on the indirect association path in the set of indirect association paths whose cumulative association strength parameter exceeds the preset threshold. The preset threshold is set according to the requirement of the basic scientific research field for the significance of the association relationship. The indirect association path whose cumulative association strength parameter exceeds the threshold indicates that there is an indirect association between the two entities with a certain significance. For these paths, the implicit association type between the two entities at the ends of the path is determined through analysis of the logical chain and semantic connotation of each association relationship in the path. For example, for the indirect association path "a new type of immunotherapy drug - regulates signal pathway - activates transcription factor - inhibits cell proliferation", the cumulative association strength parameter exceeds the preset threshold, and through analysis, it is known that the path indicates that the new type of immunotherapy drug ultimately has an inhibitory effect on cell proliferation through a series of intermediate processes, and therefore the implicit association type between the two entities "a new type of immunotherapy drug" and "cell proliferation" is determined to be "indirect inhibitory effect". For another example, the path "a new type of immunotherapy drug - enhances immune cell activity - regulates apoptosis", the cumulative association strength parameter exceeds the preset threshold, and the implicit association type between "a new type of immunotherapy drug" and "apoptosis" is "indirect promotion". These implicit association types are not directly expressed in the original information, but through reasoning, it can be clear that they exist, enriching the association relationship between entities.

[0133] Step S144: adding the extracted implicit association type to the entity association network to generate an enhanced entity association network containing explicit association relationships and implicit association types, and taking the enhanced entity association network as an intermediate result of deep mining of information implicit relationships.

[0134] The network expansion module adds the extracted implicit association type to the original entity association network to generate an enhanced entity association network. The addition is as follows: a new association edge is established between the entities at both ends of the indirect association path, the edge label of the edge is the extracted implicit association type, the association strength parameter of the edge is referenced from the cumulative association strength parameter of the indirect association path, and is adjusted appropriately in combination with the semantic reliability of the path. For example, the implicit association type "indirect inhibitory effect" between "a new immunotherapy drug" and "cell proliferation" is added to the network, a new association edge is established between the two entities, the edge label is "indirect inhibitory effect", and the association strength parameter is determined according to the cumulative association strength parameter of the indirect association path, and is fine-tuned in combination with the number of times each association relationship in the path is mentioned in the basic biomedical literature. If the cumulative association strength parameter of the path is high, and each association relationship is mentioned in multiple authoritative documents, the association strength parameter of the "indirect inhibitory effect" association edge will be set relatively high.

[0135] After the addition of all implicit association types is completed, the original entity association network is expanded to form an enhanced entity association network containing explicit association relationships and implicit association types. The enhanced entity association network can more comprehensively reflect the relationships between entities, not only containing explicit associations directly extracted from the original information, but also covering implicit associations obtained through reasoning. The enhanced entity association network is used as an intermediate result of deep mining containing information implicit relationships.

[0136] Step S150: The deep mining intermediate result is subjected to multi-dimensional checking and fusion processing by integrating the expert agent to generate a final information deep mining result meeting the information demand to be mined.

[0137] In the basic scientific literature mining scene, the deep mining intermediate result may still have problems such as contradictory association relationships, information redundancy, and incomplete theme coverage, although it contains rich entity association information. The integration expert agent, as the last link of information processing, will check and fuse the deep mining intermediate result in multiple dimensions to generate a final information meeting the basic research demand.

[0138] Step S151: The enhanced entity association network in the deep mining intermediate result is loaded by the integration expert agent to verify the explicit association relationships and implicit association types in the enhanced entity association network, and identify mutually contradictory association relationship pairs.

[0139] The verification module of the integration expert agent first loads the enhanced entity association network, and verifies the explicit association relationship and the implicit association type one by one. The verification process will combine factors such as the authority of the information source and the degree of evidence support in the literature. For example, check whether the explicit association relationship "inhibition" and the implicit association type "promotion" between "a new type of immunotherapy drug" and "cell proliferation" exist at the same time, if they exist, it constitutes a mutually contradictory association relationship pair. For another example, check whether there is literature to support that the structure of "a new type of immunotherapy drug" is related to the activity of the signal pathway, and at the same time, there is literature to show that the structure has no significant association with the activity, if there is, it is identified as a contradictory association relationship pair.

[0140] Step S152: For the mutually contradictory association relationship pair, the credibility score of each association relationship is calculated based on the reliability parameter of the information source and the association strength parameter, and the association relationship with higher credibility score is retained.

[0141] For the identified mutually contradictory association relationship pair, the scoring module of the integration expert agent will calculate the credibility score of each association relationship. The reliability parameter of the information source is determined according to the type of the information source, such as the reliability parameter of the research published in international top basic biomedical journals is higher than that of ordinary conference papers; the association strength parameter refers to the association strength parameter of the association relationship in the enhanced entity association network. When calculating the credibility score, the reliability parameter of the information source and the association strength parameter are combined according to certain logic, for example, first determine the weight of the two in the credibility score, then combine the corresponding values of the two according to the weight to obtain the final credibility score.

[0142] Taking the contradictory association relationship pair "inhibition" and "promotion" between "a new type of immunotherapy drug" and "cell proliferation" as an example, if the association relationship "inhibition" comes from multiple basic biomedical journals with high impact factor, the reliability parameter of the information source is high, and the association strength parameter is also high, the calculated credibility score is high; while the association relationship "promotion" comes from an experimental report with small sample size, the reliability parameter of the information source and the association strength parameter are both low, the credibility score is low, then the association relationship "inhibition" is retained.

[0143] Step S153: The redundant information in the enhanced entity association network is processed, when multiple entity association paths express the same semantic relationship, they are combined into one comprehensive association path, and the association strength parameter of the comprehensive association path is the weighted average of the association strength parameters of each path.

[0144] The fusion module of the integrated expert agent processes redundant information in the enhanced entity association network. For example, the two association paths "a new immunotherapy drug-activates signal pathways-regulates cell proliferation" and "a new immunotherapy drug-enhances transcription factor activity-inhibits cell proliferation" both express the same semantic relationship "a new immunotherapy drug regulates cell proliferation through molecular signal pathways", and at this time the two paths are combined into one comprehensive association path "a new immunotherapy drug-through molecular signal pathways-regulates cell proliferation". The calculation method of the association strength parameter of the comprehensive association path is to determine the weight of the association strength parameters of the two original paths according to the frequency of their appearance in the open academic journal database, and then calculate the weighted average value as the association strength parameter of the comprehensive association path.

[0145] Step S154: Match the processed enhanced entity association network with the demand keyword set of the information demand to be mined, check whether all core keywords and associated keywords are covered, and supplement the missing association relationships related to the keywords.

[0146] The matching module of the integrated expert agent matches the processed enhanced entity association network with the demand keyword set. The core keywords in the demand keyword set are "a new immunotherapy drug", "molecular mechanisms of advanced lung cancer", "cell function impact", and "basic research progress", and the associated keywords include "mechanism of action", "cell proliferation", and "immune regulation". In the matching process, it is checked whether the enhanced entity association network contains association relationships related to each keyword. If it is found that the association relationship related to the associated keyword "immune regulation" is missing, the relevant basic research literature information is searched through the search expert agent, and the association relationship is extracted and added to the network, such as "a new immunotherapy drug-enhances T cell activity-immune regulation effect".

[0147] Step S155: Based on the matching result, convert the enhanced entity association network into structured text, and take the structured text as the final information deep mining result that meets the information demand to be mined, which includes topic hierarchical titles, entity relationship descriptions, and implicit relationship explanations.

[0148] According to the matching result, the conversion module of the expert agent integrates the enhanced entity association network into structured text. The structured text is organized according to the topic hierarchy, and the topic hierarchy titles such as "Overview of a new immunotherapy drug", "Research on the molecular mechanism of advanced lung cancer", "Analysis of cell function impact", and "Summary of basic research progress" are set. Under each title, the relationship between entities is described in detail, such as "A new immunotherapy drug activates T cells by binding to T cell surface receptors, enhances their function, and promotes the recognition and killing of tumor cells"; and the implicit relationship is explained, such as "From the data of drug regulation of signal pathways and cell apoptosis, it can be inferred that the drug may have an indirect inhibitory effect on tumor cell proliferation".

[0149] The structured text comprehensively and systematically presents the content related to the information demand to be mined, covering multiple aspects such as the molecular mechanism of the drug, the impact of cell function, and the progress of basic research. After multiple rounds of verification and fusion, the information is accurate and the logic is clear, which can provide information support for basic researchers to conduct literature mining and molecular mechanism research. This structured text is the final information depth mining result that meets the information demand to be mined.

[0150] Figure 2 An information depth mining system 100 based on a multi-agent collaboration framework is shown, which includes a processor 1001, a memory 1003, and program code stored in the memory 1003. The processor 1001 executes the above-mentioned program code to implement the steps of the information depth mining method based on the multi-agent collaboration framework.

[0151] Figure 2 The information depth mining system 100 based on the multi-agent collaboration framework includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the information depth mining system 100 based on the multi-agent collaboration framework can also include a transceiver 1004, which can be used for data interaction between the information depth mining system based on the multi-agent collaboration framework and other information depth mining systems based on the multi-agent collaboration framework, such as data transmission and / or data reception. It should be noted that the transceiver 1004 is not limited to one in actual scheduling, and the structure of the information depth mining system 100 based on the multi-agent collaboration framework does not constitute a limitation on the embodiments of the present application.

[0152] The memory 1003 is used to store the program code for executing the embodiments of the present application, and is controlled by the processor 1001 to execute. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.

[0153] The embodiment of the application provides a computer readable storage medium, which stores program codes, and the program codes are executed by a processor to implement the steps and corresponding contents of the foregoing method embodiments.

[0154] The above is only optional implementation of part of the implementation scenarios of the application. It should be pointed out that, for ordinary skilled persons in the technical field, other similar implementation methods according to the technical idea of the application without departing from the technical concept of the application also belong to the protection scope of the embodiments of the application.

Claims

1. A method for deep information mining based on a multi-agent collaborative framework, characterized in that, The method includes: The system receives requests for information to be mined, and through a planning expert agent, it performs task decomposition and collaborative scheduling to generate a collaborative task scheduling scheme that includes task hierarchy and agent collaboration paths. Based on the collaborative task scheduling scheme, the search expert agent is invoked to perform multi-source information collaborative search operations to obtain multi-source raw information containing multi-dimensional information sources. By reading expert intelligent agents, deep analysis and correlation extraction of multi-source raw information are performed to generate structured relational information with entity relational networks; The analysis expert intelligent agent is invoked to perform multi-level deep reasoning processing on the structured relational information to obtain deep mining intermediate results containing implicit relationships of information; By integrating expert intelligent agents to perform multi-dimensional verification and fusion processing on the intermediate results of deep mining, the final information deep mining results that meet the needs of the information to be mined are generated. The process involves decomposing and collaboratively scheduling the information to be mined using a planning expert agent to generate a collaborative task scheduling scheme that includes task hierarchy relationships and agent collaboration paths. The planning expert intelligent agent performs thematic element extraction processing on the information demand to be mined, identifies the core thematic words and related thematic words in the information demand to be mined, and generates a set of demand thematic words. Based on the set of demand keywords, a task hierarchical splitting operation is performed to break down the information demand to be mined into multiple sub-tasks with hierarchical dependencies. Each sub-task includes a task objective description and input / output interface definition. Obtain capability description information for each functional agent, including the task types that the agent is good at handling, information processing dimensions, and historical collaboration efficiency parameters. The matching degree of the task objective description of each subtask and the capability description information of each functional agent is calculated to generate a target matching matrix, wherein the elements in the target matching matrix represent the degree of fit of the target; Based on the target matching matrix and the hierarchical dependency relationship between subtasks, a collaborative path planning model is constructed, with the objective function being to minimize the total execution time of the task. Solve the collaborative path planning model to determine the executing agent corresponding to each subtask and the collaborative order between agents, and generate a collaborative task scheduling scheme that includes task hierarchy and agent collaborative paths.

2. The information deep mining method based on a multi-agent cooperative framework according to claim 1, characterized in that, Based on the set of demand keywords, the task hierarchy is broken down by a planning expert agent, which performs a task decomposition operation to divide the information demand to be mined into multiple sub-tasks with hierarchical dependencies, including: By using a planning expert intelligent agent, the core keywords in the demand keyword set are taken as the root node to construct a keyword hierarchy tree. The root node represents the overall goal of the information demand to be mined, and the child nodes of the root node are related keywords. Based on the hierarchical structure of the keyword hierarchy tree, the overall goal corresponding to the root node is divided into multiple first-level sub-tasks. Each first-level sub-task corresponds to a related keyword of a sub-node. The goal of the first-level sub-task is to obtain basic information related to the related keyword. The associated keywords corresponding to each first-level sub-task are further split, and the sub-topic elements of the associated keywords are extracted. Each sub-topic element is taken as the target of the second-level sub-task, and the target of the second-level sub-task is to obtain the detailed information of the sub-topic elements. Check the dependencies between first-level subtasks and second-level subtasks. When the execution of a second-level subtask depends on the output of other first-level subtasks, mark the dependency and record the dependency conditions. The dependency conditions include the type and completeness requirements of the required information. A subtask list is generated according to the hierarchical order and dependencies of the keyword hierarchy tree. Each subtask in the subtask list includes a task identifier, a parent task identifier, a task objective description, and dependency conditions. The subtask list is used as the result of subtask splitting with hierarchical dependencies.

3. The information deep mining method based on a multi-agent cooperative framework according to claim 1, characterized in that, The collaborative task scheduling scheme invokes a search expert agent to perform a multi-source information collaborative search operation, obtaining multi-source raw information containing multi-dimensional information sources, including: The agent collaboration path in the collaborative task scheduling scheme is analyzed to determine the search expert agent responsible for information search and the corresponding search sub-task. Each search sub-task includes search topic and information source type parameters. The search expert agent is invoked to access a preset search strategy library based on the search topic and information source type parameters of the search subtask, and to generate a set of search rules for each information source type. The set of search rules includes keyword combination methods, search field restrictions, and result filtering conditions. Parallel searches are performed on multi-dimensional information sources based on the set of search rules, and the execution status of each search thread is monitored in real time, with the search response time and the number of preliminary results recorded. When potential duplication of information returned by different search threads is detected, the search path is dynamically adjusted by the search expert agent, modifying the keyword combination method or search field limitation of the conflicting path. Collect the initial search results returned by each search thread, perform source marking processing on the initial search results, record the source type and acquisition timestamp of each information fragment, and integrate the marked information fragments into multi-source original information containing multi-dimensional information sources.

4. The information deep mining method based on a multi-agent cooperative framework according to claim 3, characterized in that, The process involves performing parallel searches on multi-dimensional information sources based on the set of search rules, monitoring the execution status of each search thread in real time, and recording the search response time and the number of preliminary results, including: Based on the number of search rule sets and the access characteristics of each information source, an independent search thread is allocated to each search rule set, and the maximum execution time and resource consumption threshold of the search thread are set. All search threads are started, and each search thread generates a search request based on the corresponding set of search rules. The search request includes keyword combinations, search fields and filtering conditions, and is sent to the corresponding information source platform through a preset interface protocol. The execution status parameters of each search thread are collected in real time. These parameters include the current search progress, the number of results returned, network connection status, and response latency. When the response delay time of any search thread exceeds a preset threshold, a priority increase instruction is sent to that search thread to adjust the resource allocation weight of the search thread and increase its data reception priority. When an abnormal network connection status is detected in any search thread, the thread retry mechanism is triggered to re-establish the network connection and continue to execute the search request, recording the number of retries and the recovery time. The execution status parameters of each search thread are updated to the search status dashboard in real time. The search status dashboard is used to display the real-time progress and abnormal situations of all threads and supports dynamic adjustment of thread execution strategies.

5. The information deep mining method based on a multi-agent cooperative framework according to claim 1, characterized in that, The process of deep analysis and association extraction of multi-source raw information through a reading expert intelligent agent to generate structured association information with an entity association network includes: The reading expert agent inputs multi-source raw information, performs similarity calculation on the multi-source raw information, and identifies and removes completely duplicate information segments. The deduplicated information fragments are subjected to structured parsing to identify the core entities in the information fragments. The core entities include person entities, organization entities, and concept entities. The basic attribute information of the core entities is recorded. Analyze the semantic relationships between core entities in an information fragment, and extract the association type and association strength parameters between entities. The association type represents the interaction relationship between entities, and the association strength parameter represents the explicitness of the relationship. Based on the extracted core entities and the association types and association strength parameters between entities, an entity association network is constructed. The entity association network uses core entities as nodes, association types as edge labels, and association strength parameters as edge weights. The entity association network is simplified by performing redundant edge processing. When there are multiple association edges of the same type between two core entities, the edge with the highest association strength parameter is retained as the main association edge, and the association strength parameters of other edges are added to the main association edge to generate structured association information with entity association network.

6. The information deep mining method based on a multi-agent cooperative framework according to claim 5, characterized in that, The semantic relationships between core entities in the analyzed information fragment are analyzed, and the association type and association strength parameters between entities are extracted, including: Syntactic analysis is performed on the information fragment to identify the subject-verb-object structure and adverbial-head structure in the information fragment, and to determine the grammatical role of the core entity in the sentence; Based on the grammatical roles of core entities, a pre-trained relation classification model is invoked to perform preliminary classification of the semantic relationships between entity pairs, resulting in a candidate association type set, which contains multiple possible association types and their corresponding initial confidence levels. The candidate association type set is verified by combining the context of the information fragment, checking whether the association type is consistent with the context semantics, and eliminating candidate association types that contradict the context. Frequency statistics are performed on the remaining candidate association types, and the frequency of each association type in the information fragment is calculated. The association type with the highest frequency is taken as the final association type between entities. Based on the initial confidence level of the final association type and the context verification results, the association strength parameter is calculated to generate the association type and association strength parameter between entities. The value of the association strength parameter is positively correlated with the initial confidence level and negatively correlated with the number of semantic conflicts found during the context verification process.

7. The information deep mining method based on a multi-agent cooperative framework according to claim 1, characterized in that, The expert intelligence agent is invoked to perform multi-level deep reasoning processing on the structured relational information, obtaining deep mining intermediate results containing implicit relationships, including: The analysis expert intelligent agent is invoked to perform hierarchical division of the entity association network in the structured association information, and the core entities are divided into multiple topic levels according to topic relevance. Each topic level contains multiple closely related entity subsets. Based on the topic hierarchy division results, cross-level association analysis is performed on entities at different topic levels to identify indirect association paths between entity subsets and calculate the cumulative association strength parameter of indirect association paths. Implicit relationships are extracted from indirect association paths where the cumulative association strength parameter exceeds a preset threshold. The implicit association type between the entities at both ends of the indirect association path is determined. The implicit association type represents a relationship that is not directly expressed in the original information but can be obtained through reasoning. The extracted implicit association types are added to the entity association network to generate an enhanced entity association network containing explicit association relationships and implicit association types. The enhanced entity association network is used as an intermediate result of deep mining containing implicit relationships.

8. The information deep mining method based on a multi-agent cooperative framework according to claim 1, characterized in that, The process involves integrating expert intelligence agents to perform multi-dimensional verification and fusion processing on intermediate deep mining results, generating final deep mining results that meet the requirements of the information to be mined, including: The integrated expert agent loads the enhanced entity association network from the deep mining intermediate results, verifies the explicit and implicit association types in the enhanced entity association network, and identifies contradictory association pairs. For the contradictory pairs of relationships, a credibility score is calculated for each relationship based on the reliability parameter of the information source and the relationship strength parameter, and the relationship with the higher credibility score is retained. Redundant information in the enhanced entity association network is processed. When multiple entity association paths express the same semantic relationship, they are merged into a comprehensive association path. The association strength parameter of the comprehensive association path is the weighted average of the association strength parameters of each path. The processed enhanced entity association network is matched with the set of keywords for the information to be mined, and it is checked whether all core keywords and related keywords are covered, and the associations related to missing keywords are supplemented. Based on the matching results, the enhanced entity association network is converted into structured text. The structured text is used as the final information deep mining result that meets the information mining requirements. The structured text includes topic-level titles, entity relationship descriptions, and implicit relationship explanations.

9. An information deep mining system based on a multi-agent collaborative framework, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the information deep mining method based on a multi-agent cooperative framework as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Literature retrieval and reading method based on multi-agent

    CN119760154A