Medical data analysis method based on intelligent agent

By constructing an iterative reasoning of agents with dynamic knowledge graphs and process reward models, the problem of untraceability and insufficient optimization of decision-making processes in pancreatic state assessment is solved, and accurate analysis and decision-making optimization of pancreatic state are achieved, which improves the accuracy and traceability of clinical decision-making.

CN120373426APending Publication Date: 2025-07-25NANJING UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510452238.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology lacks real-time monitoring and evaluation of decision-making processes in pancreatic status assessment, which cannot meet the traceability and risk assessment needs of precision medicine, and lacks a dynamic optimization mechanism for generating decisions, resulting in insufficient output solutions in terms of execution processes and risk control.

Method used

Adopting the medical data analysis method based on agents, by constructing a dynamically updated knowledge graph and process reward model, multiple rounds of iterative reasoning and decision optimization are realized, the knowledge graph is used to analyze multi-dimensional heterogeneous data, and combining the cross entropy loss function optimization process reward model to generate high-quality auxiliary decision-making solutions.

Benefits of technology

Accurate analysis and decision-making optimization of pancreatic state are achieved, the accuracy and objectivity of clinical decision-making are improved, the entire traceability of the decision-making process and the timeliness of information are ensured, and the agent's reasoning path is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373426A_ABST
    Figure CN120373426A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data analysis method based on an intelligent agent, and the method comprises the following steps: S1, obtaining a multi-dimensional heterogeneous data source which comprises an unstructured log, a structured attribute tag, time sequence sensor data and an iconography data report; s2, performing structured processing on the dynamic rule base document; s3, performing multiple rounds of iterative reasoning by using the data analysis agent to generate a preliminary analysis decision scheme; s4, constructing a data set for training a process reward model PRM; s5, training a process reward model as a verifier of a data analysis agent reasoning process, and optimizing the process reward model by using a cross entropy loss function; and S6, optimizing a decision scheme generation process of the data analysis agent based on a process reward model, sampling M candidate states for each decision step by the agent in a step-by-step reasoning process, scoring each candidate state by using a verifier, and selecting the state with the highest score as a final decision of the step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for analyzing medical data based on an agent. Background Art

[0002] The pancreas plays an important role in human metabolic regulation and digestive function, and its related research involves the processing and analysis of various biomedical data. Due to the complex structure of the pancreas, its status assessment usually depends on multi-source heterogeneous data, such as imaging features, pathological classifications, biochemical indicators and other multi-dimensional data, the real-time adaptation of a dynamic rule base, such as international standardized clinical guidelines, etc., and the construction of a multi-module collaborative decision-making mechanism. Based on the objective analysis of pancreas-related data, it is helpful to provide decision-making suggestions.

[0003] Currently, it is necessary to integrate multi-modal inputs including time-series sensor data (such as device status features), structured attribute tags (such as component performance parameters), and unstructured logs (such as operation records), etc. At the same time, it is necessary to dynamically match a continuously updated industry standard library (such as international standard protocols) and an empirical verification data set (such as a historical operation case library). This process places extremely high requirements on the system's knowledge graph construction ability, real-time data processing efficiency, and rule iteration adaptability.

[0004] There are still two deficiencies in the prior art:

[0005] 1) The decision-making process is not traceable: There is a lack of real-time monitoring and evaluation of the decision-making process, which cannot meet the requirements of precision medicine for program traceability and risk assessment;

[0006] 2) There is a lack of a mechanism for dynamically optimizing and continuously improving the generated decisions, resulting in deficiencies in aspects such as the execution process and risk control of the output solutions. Summary of the Invention

[0007] Aiming at the deficiencies of the above prior art, the purpose of the present invention is to provide a method for analyzing medical data based on an agent, which uses a knowledge graph, agent interactive exploration, and a process reward model PRM to optimize the search path, realizes accurate analysis and decision optimization of various heterogeneous data, and is applicable to modern intelligent medicine and an auxiliary clinical decision-making system.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0009] A method for analyzing medical data based on an agent of the present invention includes the following steps:

[0010] S1, obtaining a multi-dimensional heterogeneous data source, where the multi-dimensional heterogeneous data source includes unstructured logs, structured attribute tags, time-series sensor data, and imaging data reports;

[0011] S2. Structurally process the dynamic rule base document:

[0012] a) Establish a table of contents index by chapter to support retrieving guidelines based on the table of contents structure;

[0013] b) Construct a knowledge graph: Organize the medical entities and relationships extracted from the dynamic rule base into a graph structure to construct a knowledge graph that can be dynamically updated;

[0014] S3. Use the data analysis agent to perform multi-round iterative reasoning to generate a preliminary analysis and decision-making plan;

[0015] S4. Construct a data set for training the process reward model PRM;

[0016] S5. Train the process reward model as a verifier Verifier for the reasoning process of the data analysis agent, and optimize the process reward model PRM using the cross-entropy loss function;

[0017] S6. Optimize the decision-making plan generation process of the data analysis agent based on the process reward model. During the step-by-step reasoning process, the agent samples M candidate states for each decision step, uses the verifier to score each candidate state, and selects the state with the highest score as the final decision for this step.

[0018] Furthermore, the process of constructing the knowledge graph in step S2b) specifically includes:

[0019] b1) Extract medical entity information from the dynamic rule base using a named entity recognition algorithm;

[0020] b2) Identify the relationships between entities through dependency syntactic analysis and a relationship extraction model;

[0021] b3) Establish a knowledge graph update mechanism to regularly extract information from the latest network dynamic rule base for automatic update.

[0022] Furthermore, the knowledge graph in step S2b) includes:

[0023] The medical entity information includes: diseases, clinical stages, symptoms, examination items, drugs, treatment methods, complications, risk factors;

[0024] The types of relationships between entities include: symptom manifestations, examination means, treatment plans, prognosis, diagnostic basis, staging associations, drug associations, drug taboos, side effects, risk factors, evidence levels.

[0025] Furthermore, step S3 specifically includes:

[0026] S31. Call the guideline index tool to retrieve the original text of relevant clauses;

[0027] S32 Information extraction and entity mapping: The intelligent analysis module extracts key clinical features from the multi-dimensional heterogeneous data sources and maps them to the corresponding entities in the medical knowledge graph to construct structured data;

[0028] S33 Knowledge graph retrieval: Invoke the knowledge graph retrieval tool to analyze potential association paths based on the structured data. The potential association paths include the relationships between key clinical features and medical knowledge, as well as the logical associations between different medical concepts;

[0029] S34 Iterative optimization: According to the retrieval results and the preliminary reasoning results, identify the parts with insufficient information, generate structured information requests, and repeat any of the above steps to obtain supplementary information until the model sends an end signal;

[0030] S35 Auxiliary decision-making plan generation: Based on multi-round information analysis and combined with the cumulative data association chain, generate an auxiliary decision-making plan including medical reference information and applicable conditions to support the optimization of the clinical workflow.

[0031] Furthermore, the step S4 specifically includes:

[0032] a) Extract valid data from the official unstructured logs, where the valid data contains several cases;

[0033] b) For each case, map the information in its decision-making plan to the corresponding entities and relationships in the knowledge graph to form a subgraph where is the set of entities involved, and is the relationship involved;

[0034] c) For each training case data, sample M times of complete agent inference-generated data, that is, let the agent independently execute M times of the complete auxiliary decision-making plan generation process for the same case data and record all its interaction steps, including inference output, tool call, tool feedback, and the final decision-making plan; For a sampled dialogue τ, τ = [o1, a1, o2, a2, …, o K , a K , where K is the total number of inference steps of the sampled dialogue data τ, o i is the observed information, and a i is the output of the agent at the i-th step; At the same time, define the state s i of the agent at the i-th step as:

[0035] s i = [o1, a1, o2, a2, …, a i ;

[0036] For each sample, an automatic estimation method is used to evaluate the quality of the state s of the agent at the i-th step, i and this evaluation is used as the reward score supervision signal for this step.

[0037] Furthermore, the automatic estimation method specifically includes:

[0038] First, calculate the vector cosine similarity between the inference output I of the agent i and the decision-making plan :

[0039] By calculating the shortest hop count d(c, e) between the candidate entity c explored by the agent and the target entity e involved in the decision-making plan on the knowledge graph, the distance between nodes is obtained. The closer the distance, the higher the inference quality;

[0040] The distance between the set of candidate entities C explored in the current inference step and the set of target entities involved in the decision-making plan is:

[0041]

[0042] The score y of the automatic estimation method for the state of the agent at the i-th step si is calculated as:

[0043]

[0044] where w1 and w2 are weight coefficients, and w1 + w2 = 1, both optimized through the validation set.

[0045] Furthermore, the specific steps of step S5 include:

[0046] Use the process reward model as the agent state evaluator. Its input is the state s of the agent at the i-th step i , and the output is a scalar r, r ∈ [0, 1]; Select an open-source medical large language model for training, so that the large model predicts a positive or negative conclusion after inputting the state s i , and obtain the logarithmic probabilities l + and l - of the model for the two symbols + and -, and use the following calculation formula as the score of the process reward model for this state:

[0047] r = exp(l + ) / (exp(l + ) + exp(l - ))

[0048] Then use the following cross-entropy loss function to train the process reward model PRM:

[0049]

[0050] wherein is the state s of the agent at the i-th step i is the label of the reward score is the reward score assigned by the PRM for the state s i

[0051] Furthermore, the step S6 specifically includes:

[0052] In the test phase, for the initially collected data, the agent performs reasoning step by step; in each reasoning step, the agent samples and generates M candidate actions a ij j = 1, 2,..., M, to form multiple candidate states s ij , j = 1, 2,..., M; subsequently, a validator is used to score all the candidate states s ij and the one with the highest score is selected as the final state s of this step i* , that is

[0053] Furthermore, the decision-making scheme includes:

[0054] The execution roadmap of the decision-making scheme and the alternative decision-making scheme tree;

[0055] The visual traceability graph of the decision-making process;

[0056] The multi-dimensional risk assessment report.

[0057] Advantages of the present invention:

[0058] The present invention can make full use of historical medical data and the information in the latest international dynamic rule base, and through constructing a dynamically updated medical knowledge graph and iterative reasoning of an agent based on a reward model, achieve accurate analysis and decision optimization of various heterogeneous data, and is applicable to modern intelligent healthcare and auxiliary clinical decision-making systems. This method has:

[0059] Improve the accuracy and objectivity of clinical decision-making;

[0060] Realize the full traceability of the decision-making process, which is convenient for subsequent evaluation and review;

[0061] Dynamically update the knowledge graph to ensure the timeliness of medical information;

[0062] Optimize the reasoning path of the agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a schematic diagram of the implementation scenario of the method for medical data analysis based on an agent according to the present invention;​

[0064] Figure 2 This is a schematic diagram of the scenario for optimizing the search process through a verifier in the method for implementing an agent-based medical data analysis of the present invention;

[0065] Figure 3 This is a schematic flow diagram of the method for implementing an agent-based medical data analysis of the present invention;

[0066] Figure 4 This is a schematic structural diagram of the method for implementing an agent-based medical data analysis of the present invention. Detailed implementation manners

[0067] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with embodiments and the accompanying drawings. The content mentioned in the embodiments does not limit the present invention. The embodiments of this specification propose an agent-based medical data analysis method, which automatically generates a high-quality auxiliary assessment plan for pancreatic status by constructing a dynamically updated medical knowledge graph and iterative reasoning of an agent based on a reward model.

[0068] An agent-based medical data analysis method of the present invention includes the following steps:

[0069] S1. Obtain multi-dimensional heterogeneous data sources, where the multi-dimensional heterogeneous data sources include unstructured logs, structured attribute tags, time-series sensor data, and imaging data reports;

[0070] S2. Perform structured processing on the dynamic rule base document:

[0071] a) Establish a table of contents index by chapter to support retrieval of guidelines based on the table of contents structure;

[0072] b) Construct a knowledge graph: Organize the medical entities and relationships extracted from the dynamic rule base into a graph structure to construct a knowledge graph that can be dynamically updated;

[0073] S3. Use a data analysis agent to perform multiple rounds of iterative reasoning to generate a preliminary analysis and decision-making plan;

[0074] S4. Construct a data set for training the process reward model PRM;

[0075] S5. Train the process reward model as a verifier Verifier for the reasoning process of the data analysis agent, and optimize the process reward model PRM using the cross-entropy loss function;

[0076] S6. Optimize the decision - making plan generation process of the data - analysis intelligent agent based on the process reward model. During the step - by - step reasoning process of the intelligent agent, sample M candidate states for each decision step, use the validator to score each candidate state, and select the state with the highest score as the final decision for this step.

[0077] Further, the process of constructing the knowledge graph in step S2b) specifically includes:

[0078] b1) Extract medical entity information from the dynamic rule base using the named - entity recognition algorithm;

[0079] b2) Identify the relationships between entities through dependency syntactic analysis and relationship extraction models;

[0080] b3) Establish a knowledge - graph update mechanism to regularly extract information from the latest network dynamic rule base for automatic update.

[0081] Further, the knowledge graph in step S2b) includes:

[0082] The medical entity information includes: diseases, clinical stages, symptoms, examination items, drugs, treatment methods, complications, risk factors;

[0083] The relationship types between entities include: symptom manifestation, examination means, treatment plan, prognosis, diagnostic basis, staging association, drug association, drug contraindication, side effect, risk factor, evidence level.

[0084] Further, step S4 specifically includes:

[0085] a) Extract valid data from the official unstructured log, where the valid data contains several cases;

[0086] b) For each case, map the information in its decision - making plan to the corresponding entities and relationships in the knowledge graph to form a sub - graph where is the set of entities involved, and is the relationship involved;

[0087] c) For each training - case data, sample M times of complete intelligent - agent inference - generated data, that is, let the intelligent agent independently execute M times of the complete auxiliary decision - making plan generation process for the same case data, and record all its interaction steps, including inference output, tool call, tool feedback, and the final decision - making plan; for a sampled dialogue τ, τ = [o1,a1,o2,a2,…,o K ,a K , where K is the total number of inference steps of the sampled dialogue data τ, o i is the observation information, ai is the output of the agent at the i-th step; meanwhile, the state s of the agent at the i-th step is defined i It is expressed as:

[0088] s i = [o1, a1, o2, a2, …, a i ;

[0089] For each sample, an automatic estimation method is used to perform a quality score on the state s of the agent at the i-th step i and use it as the reward score supervision signal for this step.

[0090] Furthermore, the automatic estimation method specifically includes:

[0091] First, calculate the vector cosine similarity between the agent's inference output I i and the decision-making plan :

[0092] By calculating the shortest hop count d(c, e) between the candidate entity c explored by the agent and the target entity e involved in the decision-making plan on the knowledge graph, the distance between nodes is obtained. The closer the distance, the higher the inference quality;

[0093] The distance between the set C of candidate entities explored in the current inference step and the set of target entities involved in the decision-making plan is:

[0094]

[0095] The score of the automatic estimation method for the i-th step state of the agent The calculation method is:

[0096]

[0097] where w1 and w2 are weight coefficients, and w1 + w2 = 1, both of which are optimized through the validation set.

[0098] Furthermore, the specific steps of step S5 include:

[0099] Use the process reward model as the agent state evaluator, whose input is the state s of the agent at the i-th step i , and the output is a scalar r, r ∈ [0, 1]; Select an open-source medical large language model for training, so that the large model predicts a positive or negative conclusion after inputting the state s i and obtain the logarithmic probabilities l + and l -, and use the following calculation formula as the score of the process reward model for this state:

[0100] r = exp(l + ) / (exp(l + ) + exp(l - ))

[0101] Then use the following cross-entropy loss function to train the process reward model PRM:

[0102]

[0103] where is the reward score label of the state s i at the i-th step of the agent, is the reward score given by the PRM for the state s i .

[0104] Furthermore, the step S6 specifically includes:

[0105] In the test stage, for the initially collected data, the agent gradually performs reasoning; in each reasoning step, the agent samples and generates M candidate actions a ij j = 1, 2,..., M, to form multiple candidate states s ij , j = 1, 2,..., M; then use the validator to score all candidate states s ij , and select the one with the highest score as the final state s i* of this step, that is

[0106] Furthermore, the decision-making scheme includes:

[0107] The decision-making scheme execution roadmap and alternative decision-making scheme trees;

[0108] The visual traceability map of the decision-making process;

[0109] The multi-dimensional risk assessment report.

[0110] By generating a complete diagnosis and treatment plan through automated and multi-round interactive reasoning, and real-time quality assessment of the decision-making path, this method realizes the transparency and traceability of the decision-making process, thus greatly improving the accuracy of pancreatic state assessment and the clinical application effect. Based on this, it is particularly urgent to construct a new intelligent diagnosis and treatment framework that integrates dynamic knowledge evolution and process verification. This framework realizes the organic integration of multi-source medical knowledge by establishing an interpretable iterative reasoning mechanism, and uses the process reward model to optimize each decision in real time, and finally forms an auxiliary optimization treatment strategy that not only conforms to the evidence-based medicine standard but also takes into account the individual characteristics of patients.

[0111] In the embodiment, the system receives the medical data of patients through the hospital information system, including medical record data such as medical history, family history, and previous treatment records, pathology reports including histological type and degree of differentiation, biochemical indicators including CA19-9, CEA, and liver function indicators, and medical imaging reports: for example, imaging data reports such as CT and MRI.

[0112] In terms of guideline document processing, the system first performs structured processing on the international pancreatic cancer dynamic rule library, mainly including the following two links:

[0113] Construction of directory index:

[0114] Use the directory structure to divide the guideline chapters and generate a tree-like index structure; support queries by chapter name or number and semantic retrieval. When the query statement q is input, by calculating the semantic similarity between the query statement and the content of each sub-chapter, select the sub-chapter with the highest similarity and return the original content.

[0115] Construction of medical knowledge graph:

[0116] Use named entity recognition algorithms (such as BiLSTM-CRF) to automatically extract medical entities, such as disease names, symptoms, examination methods, drugs, and treatment methods. Subsequently, based on dependency syntactic analysis and relationship extraction models (such as BERT+Softmax), identify the internal associations between entities, form relationships such as "symptom manifestation", "examination means", and "treatment plan", and store them in the form of triples, for example, (pancreatic cancer, treatment plan, gemcitabine). To ensure the timeliness of knowledge, the system designs a periodic update mechanism (such as automatically extracting information from the latest medical literature every quarter) to update and supplement the knowledge graph. The medical knowledge graph includes:

[0117] Entity types: diseases, clinical stages, symptoms, examination items, drugs, treatment methods, complications, risk factors;

[0118] Relationship types: symptom manifestation, examination means, treatment plan, prognosis, diagnostic basis, staging association, drug association, drug contraindication, side effects, risk factors, evidence level (edge weight);

[0119] The entities and relationships defined in the graph help to establish a complete logical chain from symptoms to diagnostic strategy evaluation and then to treatment plans.

[0120] After receiving the patient data, the intelligent agent first performs information extraction and entity mapping, mapping the extracted clinical information to the corresponding entities in the constructed medical knowledge graph. Subsequently, the intelligent agent further retrieves relevant medical knowledge by calling external tools.

[0121] To enhance the accuracy of information acquisition and solution generation, based on the directory index of the dynamic rule library and the knowledge graph, the following three core tools are designed (see Table 1 for details):

[0122] Table 1 Tool List

[0123]

[0124] viewGuideDirectory: This tool parses the directory structure of the guide document. After calling this function, the system will display the information of all chapters and sub-chapters, helping to quickly browse the overall structure and content of the guide.

[0125] searchGuideDirectory: It is applicable to scenarios where specific content needs to be searched by chapter, subsection, or keyword. It provides support for subsequent reference and citation, and semantic search is supported when searching by keyword.

[0126] queryKnowledgeGraph: This tool utilizes the constructed medical knowledge graph. According to the input medical entities, relationships, and the specified number of hops, it traverses the relevant nodes within a certain range starting from the target entity and returns the information that meets the conditions. It supports complex graph query operations, facilitating users to obtain detailed data and associated information from the knowledge graph. Among them, the parameter query can be a custom query statement (supporting the SPARQL format) for performing more complex graph queries.

[0127] This method is a single-stage system. When the agent sends an end signal, the process terminates. The entire process only requires one interaction between the user and the agent, so it is easy to deploy and highly automated.

[0128] We use chain-of-thought (CoT) prompting to guide the agent to gradually solve the code location problem. In one embodiment, the prompt words can be as follows:

[0129] "You are an agent focused on medical data analysis and decision support. Your task is to generate a medical data-assisted analysis report based on a step-by-step, verifiable data reasoning process, and finally output the following content:

[0130] a) Roadmap for medical information analysis and suggestions for optimization solutions;

[0131] b) Visual association graph of the decision-making process;

[0132] c) Multi-dimensional risk assessment and response suggestions. The specific steps are as follows:

[0133] Clinical Information Analysis: Extract key medical information based on the patient's case data (including medical record text, pathology reports, biochemical indicators, and medical imaging reports);

[0134] Information Mapping and Retrieval: Map the extracted medical information to the corresponding entities in a pre-constructed medical knowledge graph and retrieve relevant entity relationships and guideline clauses;

[0135] Integration of Medical Reference Information: Analyze potential pathological features and relevant medical evidence using the above information and retrieval results;

[0136] Initial Plan Generation: Generate preliminary medical strategy suggestions that comply with the guidelines based on the medical reference information;

[0137] Plan Optimization and Risk Check: Optimize the preliminary strategy, focusing on evaluating complication avoidance and risk control to ensure the scientificity and rationality of the plan;

[0138] Requirement: Citation Requirement in the Decision-making Process: Each decision-making step must cite relevant entity relationships or guideline clauses in the knowledge graph to ensure that the reasoning process is transparent, traceable, and scientifically based.

[0139] The following is the de-identified data of the patient: {Case Data}”

[0140] The multi-round dialogue iteration process of the agent takes the patient's medical data as the initial input, and each round of iteration consists of three core links: reasoning and action generation, tool execution and feedback, and state splicing and transmission.

[0141] In the initial stage, the agent receives patient data containing information such as medical records, imaging data reports, and biochemical indicators. The agent first performs information extraction and entity mapping, mapping the key clinical information to the constructed knowledge graph.

[0142] In the reasoning and action link, the agent generates the next action based on the current state. The actions are divided into two categories: tool call requests and reasoning outputs. If external tools are needed to verify hypotheses or supplement information, the agent generates a structured tool call request, the format of which follows a predefined JSON specification, including the tool name, query parameters, and constraints. For example, when verifying the tumor stage, the "guideline index tool" may be called to retrieve the clauses on TNM staging in international guidelines, or the "knowledge graph tool" may be called to query the adjuvant treatment plan decision corresponding to this stage. If no tool intervention is required, the intermediate reasoning conclusion is directly output (such as "A significant increase in CA19-9 suggests a high possibility of pancreatic cancer").

[0143] In the tool execution and feedback phase, the external tool performs operations according to the call request and returns the results. The returned results need to meet the structured requirements. For example, the guideline retrieval tool returns the original text fragments of relevant clauses and the evidence level, and the knowledge graph tool returns the entity relationship path and its weight. The tool feedback and the agent actions together constitute a new round of observation data for updating the dialogue state.

[0144] In the state concatenation and transfer phase, the current action and the tool feedback are appended to the historical state sequence to form the input for the next iteration.

[0145] Repeat any of the above steps to obtain supplementary information until the agent sends an end signal; after multiple rounds of dialogue, generate a pancreatic health assessment, an adjuvant treatment plan decision, and an expected effect based on the cumulative evidence chain.

[0146] To collect the dataset for training the reward model, select case data with complete diagnosis and treatment process records and clear efficacy evaluations from the hospital's historical pancreatic cancer case database; for each case, map the information in its diagnosis and treatment plan to the corresponding entities and relationships in the medical knowledge graph to form a subgraph where are the involved entities, are the involved relationships; for each training case data, sample M times of complete agent inference-generated data, that is, let the agent independently execute the complete diagnosis and treatment plan generation process M times for the same case data, and record all its interaction steps, including inference outputs, tool calls, tool feedbacks, and the final diagnosis and treatment plan; for a sampled dialogue τ, τ = [o1, a1, o2, a2, …, o K , a K , where K is the total number of inference steps of the sampled dialogue data τ, o i is the observation information (including patient de-identified data or tool call feedback), and a i is the output of the agent at the i-th step (including the inference output I i and the tool call instruction T i ); at the same time, define the state s i of the agent at the i-th step (composed of the previous i - 1 steps of the dialogue and the i-th step action) can be expressed as:

[0147] s i = [o1, a1, o2, a2, …, a i ;

[0148] For each sampled dialogue, use an automatic estimation method to perform a quality score on the state s i of the agent at the i-th step, and use it as the reward score supervision signal for this step. Specifically, first calculate the agent inference output Ii and the real diagnosis and treatment plan The cosine similarity of vectors between

[0149] Meanwhile, by calculating the shortest hop count d(c, e) between the candidate entity c explored by the agent and the target entity e involved in the real diagnosis and treatment plan on the knowledge graph, the distance between nodes can be obtained. The closer the distance (the smaller the value of d), the higher the reasoning quality;

[0150] The set C of candidate entities explored in the current reasoning step and the set of target entities involved in the real diagnosis and treatment plan The distance between them is:

[0151]

[0152] Next, the automatic estimation method scores the i-th step of the agent The calculation method is:

[0153]

[0154] where w1 and w2 are weight coefficients, and w1 + w2 = 1, both optimized through the validation set.

[0155] Using the process reward model as the agent state evaluator, its input is the state s i (interleaved dialogue data), and the output is a scalar r (r ∈ [0, 1]). Specifically, an open-source medical large model (such as PMC-LLaMA) is selected for training to serve as the process reward model; the interleaved dialogue data is input into the model, and then the logarithmic probabilities l + and l - at the positions of the two symbols '+' and '-' in the output of the model are obtained, and the following calculation formula is used as the score of the process reward model for this state:

[0156] r = exp(l + ) / (exp(l + ) + exp(l - ))

[0157] The above score is used to measure the quality of the state: the closer the score value is to 1, the higher the quality of the candidate state, and the greater the probability that the agent chooses to adopt this candidate state; conversely, the lower the score value, the lower the quality of the state, and the smaller the probability of being adopted by the agent.

[0158] Then, the following cross-entropy loss function is used to train the process reward model (PRM):

[0159]

[0160] where is the state s at the i-th step i which is the reward score label is the reward score assigned by PRM for the state s i which is the reward score assigned by PRM for the state s

[0161] In the test phase, for new patient data, the agent performs reasoning step by step; in each reasoning step, the agent samples and generates M candidate actions a ij (j = 1, 2,..., M), forming multiple candidate states s ij (j = 1, 2,..., M); then the validator is used to score all candidate states s ij and the one with the highest score is selected as the final state s of this step i* , that is When the agent sends an end signal, the whole process terminates. The final output covers the execution roadmap of the evaluation plan dynamically generated based on patient data and the alternative plan tree, synchronously provides a visual traceability map (integrating guideline terms and evidence chains of tool calls) for assisting in diagnosis and treatment decision-making, and generates a multi-dimensional risk assessment report to quantify core indicators such as complication probability, drug adverse reactions, and prognosis deviation.

[0162] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An agent-based medical data analysis method, characterized in that It includes the following steps: S1. Obtain a multi-dimensional heterogeneous data source, where the multi-dimensional heterogeneous data source includes unstructured logs, structured attribute tags, time-series sensor data, and imaging data reports; S2. Perform structured processing on the dynamic rule library document: a) Establish a table of contents index by chapter to support retrieving guidelines based on the table of contents structure; b) Construct a knowledge graph: Organize the medical entities and relationships extracted from the dynamic rule library into a graph structure to construct a knowledge graph that can be dynamically updated; S3. Use a data analysis intelligent agent to perform multiple rounds of iterative reasoning to generate a preliminary analysis and decision-making plan; S4. Construct a dataset for training the process reward model PRM; S5. Train the process reward model as a verifier Verifier for the reasoning process of the data analysis intelligent agent, and optimize the process reward model PRM using the cross-entropy loss function; S6. Optimize the decision-making plan generation process of the data analysis intelligent agent based on the process reward model. During the step-by-step reasoning process of the intelligent agent, sample M candidate states for each decision step, use the verifier to score each candidate state, and select the state with the highest score as the final decision for this step.

2. The method according to claim 1, wherein The process of constructing the knowledge graph in step S2b) specifically includes: b1) Use a named entity recognition algorithm to extract medical entity information from the dynamic rule library; b2) Identify the relationships between entities through dependency syntax analysis and relationship extraction models; b3) Establish a knowledge graph update mechanism to regularly extract information from the latest network dynamic rule library for automatic update.

3. The method according to claim 2, characterized in that, The knowledge graph in step S2b) includes: The medical entity information includes: diseases, clinical stages, symptoms, examination items, drugs, treatment methods, complications, risk factors; The types of relationships between entities include: symptom manifestations, examination methods, treatment plans, prognosis, diagnostic basis, staging associations, drug associations, drug taboos, side effects, risk factors, evidence levels.

4. The method according to claim 1, wherein Step S3 specifically includes: S31. Call the guideline index tool to retrieve the original text of relevant clauses; S32. Information extraction and entity mapping: The intelligent analysis module extracts key clinical features from the multi-dimensional heterogeneous data source and maps them to the corresponding entities in the medical knowledge graph to construct structured data; S33. Knowledge graph retrieval: Call the knowledge graph retrieval tool to analyze potential association paths based on the structured data. The potential association paths include the relationships between key clinical features and medical knowledge, as well as the logical associations between different medical concepts; S34. Iterative optimization: According to the retrieval results and preliminary reasoning results, identify the parts with insufficient information, generate a structured information request, and repeat any of the above steps to obtain supplementary information until the model sends an end signal; S35. Scheme-assisted generation: Based on multiple rounds of information analysis, combine the cumulative data association chain to generate an auxiliary decision-making plan including medical reference information and applicable conditions to support the optimization of the clinical workflow.

5. The method according to claim 1, characterized in that, Step S4 specifically includes: a) Extract valid data from the official unstructured logs, and several cases are included in the valid data; b) For each case, map the information in its decision-making scheme to the corresponding entities and relationships in the knowledge graph to form a subgraph where is the set of entities involved, is the relationship involved; c) For each training case data, sample the complete agent inference-generated data M times, that is, let the agent independently execute the complete process of generating an auxiliary decision-making plan M times for the same case data, and record all its interaction steps, including inference outputs, tool calls, tool feedbacks, and the final decision-making plan; for a sampled dialogue τ, τ = [o1,a1,o2,a2,…,o K ,a K , where K is the total number of inference steps of the sampled dialogue data τ, o i is the observation information, a i is the output of the agent at the i-th step; at the same time, define the state s i of the agent at the i-th step as: s i = [o1,a1,o2,a2,…,a i ; For each sample, the state s of the agent at the i-th step is quality-scored using an automatic estimation method, and this is used as the reward score supervision signal for that step. i ​ 6. The method according to claim 5, wherein The specific automatic estimation method includes: First, calculate the cosine similarity of the vector between the inference output I of the agent i and the decision-making plan : By calculating the shortest hop count d(c, e) between the candidate entity c explored by the computing agent and the target entity e involved in the decision-making solution on the knowledge graph, the distance between nodes is obtained. The closer the distance, the higher the reasoning quality; The distance between the candidate entity set C explored in the current inference step and the target entity set involved in the decision-making scheme is as follows: The score of the automatic estimation method for the state of the agent at the i-th step The calculation method is as follows: Among them, w1 and w2 are weight coefficients, and w1 + w2 = 1, both of which are optimized through the validation set.

7. The method according to claim 1, wherein The specific steps of step S5 include: using the process reward model as the agent state evaluator, whose input is the state s of the agent at the i-th step i , and the output is a scalar r, where r ∈ [0, 1]; selecting an open-source medical large language model for training, so that the large model predicts a positive or negative conclusion after inputting the state s i , and obtaining the logarithmic probabilities l + and l - of the model for the two symbols + and -, and using the following calculation formula as the score of the process reward model for this state: r = exp(l + ) / (exp(l + ) + exp(l - )) Then use the following cross-entropy loss function to train the process reward model PRM: Among them is the state s of the agent at the i-th step i is the reward score label is the reward score assigned by the PRM for the state s i is the reward score assigned by the PRM for the state s 8. The method according to claim 1, wherein Among them, the specific steps of step S6 include: In the testing phase, for the data collected for the first time, the agent gradually conducts reasoning; in each reasoning step, the agent samples and generates M candidate actions a ij j = 1, 2,..., M, to form multiple candidate states s ij , j = 1, 2,..., M; subsequently, the validator is used to score all candidate states s ij , and the one with the highest score is selected as the final state s i* for this step, that is 9. The method according to claim 1, wherein The decision-making solution includes: The decision-making solution execution roadmap and the alternative decision-making solution tree; The visual traceability map of the decision-making process; The multi-dimensional risk assessment report.

Citation Information

Cited By

  • Intelligent agent evolution method based on GRPO and multi-stage verification

    CN120975134A

  • Bone soft tissue tumor repair prosthesis printing method based on multi-agent decision-making system

    CN121549962A

  • Medical data processing system based on dynamic knowledge graph and multi-agent cooperation

    CN121812046A