Large language model enhancement method and system for medical action chain optimization
By decomposing the structured action chain and optimizing multi-category heterogeneous retrieval of large language models, the problems of untrue answers, high response delays, and high interaction costs in intelligent question-answering systems are solved, and an efficient and accurate question-answering system is realized, which is suitable for complex reasoning and multimodal information processing.
Patent Information
- Application Number
- CN202511277263.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing large language models in intelligent question-answering systems have problems such as lack of authenticity assurance of answer content, difficulty in deep reasoning, large system response delays, high interaction costs, and rough external information call strategies. In particular, they cannot effectively support the construction of complex reasoning chains in multi-step, multi-source information fusion scenarios.
By breaking down user questions into structured action chains, utilizing multi-category heterogeneous retrieval action modules and a credibility verification scoring mechanism, the answer acquisition method for each sub-question is evaluated. Combined with the multi-reference and multi-source fact alignment index, the accuracy and consistency of the answers are ensured, unnecessary external retrieval and LLM calls are reduced, and the response efficiency and accuracy of the question-answering system are optimized.
It improves the reliability and efficiency of the question-answering system in real-world application scenarios. It is particularly suitable for intelligent question-answering tasks that have high requirements for answer accuracy and data support capabilities. It has high reliability, low interaction overhead, and multimodal adaptability, significantly reducing token usage and latency.
Smart Images

Figure CN120764699A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical big prediction model training, in particular to a medical action chain optimization large language model enhancement method and system. BACKGROUND
[0002] Current large language model (LLM) based question answering systems are widely used in multi-domain intelligent question answering scenarios. Representative methods include Chain-of-Thought, ReAct, and Self-Ask.
[0003] Chain-of-Thought is a method that generates intermediate reasoning steps to solve problems. It encourages the model to go through a series of step-by-step thinking processes before giving the final answer, emphasizing logical reasoning and cause-and-effect relationships, and is suitable for mathematical problems, logical reasoning, common sense reasoning, etc. It usually uses "Let's think step by step" as a guiding phrase, but if the initial step is wrong, subsequent reasoning may deviate from the correct direction.
[0004] ReAct is a method that combines reasoning and action, especially suitable for tasks that require interaction with the external environment, such as using tools or APIs. It gradually approaches the target by alternating reasoning and performing operations. It combines reasoning and action steps, and is suitable for tasks that require calling external tools, database queries, searches, etc. It usually follows a cycle of "think → action → observe → think again", but depends on available tools and interfaces. Each call to a tool may increase latency or cost.
[0005] Self-Ask is a method that guides the model to think deeply by asking itself questions. It encourages the model to actively ask related sub-questions when answering a question, and answer these sub-questions one by one to build a complete answer. It emphasizes self-reflection and problem decomposition capabilities. It is suitable for open-ended questions and complex reasoning tasks. It usually uses "I should ask myself…" as a guiding phrase, but may introduce unnecessary sub-questions, leading to longer reasoning paths. It requires strong context management capabilities to avoid getting lost in multiple sub-questions.
[0006] Therefore, the existing technology has the following defects:
[0007] 1. The answer content lacks authenticity guarantee, and the generated text is easy to deviate from the facts, forming "hallucinations".
[0008] 2. It is difficult to conduct in-depth reasoning, especially in multi-step, multi-source information fusion scenarios, which cannot effectively support the construction of complex reasoning chains.
[0009] 3、System response delay, high interaction cost, frequent call model leads to token consumption and calculation cost increases significantly.
[0010] 4、External information calling strategy is extensive, lacks systematic design, and cannot dynamically determine when to search and use which data source. SUMMARY
[0011] The application provides a medical action chain optimization large language model enhancement method and system, which overcomes the shortcomings of traditional large language models in complex reasoning, multi-modal information processing and answer reliability, improves the reliability, efficiency and universality of the question and answer system in real application scenarios, and is especially suitable for intelligent question and answer tasks with high requirements for answer accuracy and data support capability.
[0012] The application provides the following technical solution: a medical action chain optimization large language model enhancement method, comprising:
[0013] S1, calling a large language model LLM to decompose the user input question into several sub-questions, taking any sub-question as a node to generate a structured action chain, and setting an initial processing action scheme for each node sub-question;
[0014] S2, evaluating whether each sub-question can obtain an answer directly from the large language model LLM, if yes, obtaining a direct initial answer, if not, marking the missing;
[0015] S3, performing information retrieval on the several sub-questions based on the initial processing action scheme, and verifying the answer combined with the answer obtained in S2, if verified, performing content detection, if not verified, revising the answer based on the information retrieval result and repeating the verification until verified;
[0016] S4, performing content detection and completion on the verified answer to obtain the final answer of each node;
[0017] S5, integrating the final answer of each node to obtain the final answer of the user input question.
[0018] The method provides a structured, modular and scalable question and answer generation method, which overcomes the shortcomings of traditional large language models in complex reasoning, multi-modal information processing and answer reliability, improves the reliability, efficiency and universality of the question and answer system in real application scenarios, and is especially suitable for intelligent question and answer tasks with high requirements for answer accuracy and data support capability.
[0019] Preferably, the node contains the content of the sub-question, content analysis, initial processing action plan and behavior type. When evaluating whether each sub-question can be directly answered by the large language model LLM, content analysis is performed based on the content of the sub-question, and missing information is marked. If it is marked, it is a sub-question that cannot be directly answered.
[0020] Preferably, the initial processing action scheme is based on multiple types of pluggable information acquisition of different data types, including: web page query action, knowledge embedding action and structured data analysis action. The web page query action performs keyword query on the sub-question by calling the search engine API, and extracts the web page title and summary, calculates the semantic similarity with the original question through the embedding vector model, screens out the Top-k most relevant web page content, and uses it as factual material for subsequent answer correction. The knowledge embedding action segments the local domain knowledge and embeds it into the vector database. During retrieval, the sub-question is embedded and then compared with the knowledge base for vector comparison, and the text with the highest similarity is selected as the external support for the answer to the sub-question. The structured data analysis action automatically generates SQL statements or Python analysis code for sub-questions that require numerical information, and obtains structured results by connecting to predefined APIs or data tables.
[0021] Preferably, answer verification includes: constructing a joint semantic representation of the sub-question and its answer, performing similarity matching with the information retrieval results to obtain a multi-reference multi-source fact alignment index; if the multi-reference multi-source fact alignment index score is lower than a set threshold, replacing the original answer with high-similarity retrieval content; if the answer is missing, automatically filling in the answer from the information retrieval.
[0022] Preferably, the scoring basis of the multi-reference multi-source fact alignment index includes precision, recall and average word length, and the total score is calculated in a weighted manner.
[0023] Preferably, after content detection and completion, it is determined whether there are any missing nodes. If so, the missing nodes are returned to S1 for further processing.
[0024] This application also proposes a large language model enhancement system for optimizing a medical action chain, which is used to implement the large language model enhancement method for optimizing a medical action chain as described above, including:
[0025] Question parsing and action chain generation module: This module uses the large language model (LLM) to decompose the user's input question into several sub-questions, generates a structured action chain using any sub-question as a node, and sets an initial processing action plan for each node's sub-question.
[0026] Initial answer acquisition module: evaluates whether each sub-question can be directly answered by the large language model (LLM). If so, the initial answer is obtained directly. If not, the answer is marked as missing.
[0027] Answer verification module: performs information retrieval on several sub-questions based on the initial processing action plan, and verifies the answer based on the answer obtained in S2. If the verification passes, content detection is performed. If the verification fails, the answer is revised based on the information retrieval results and verification is repeated until the verification passes.
[0028] Answer completion module: performs content detection and completion on the verified answers to obtain the final answer for each node;
[0029] Final answer generation module: Integrate the final answers of each node to obtain the final answer to the question entered by the user.
[0030] The present invention has the following beneficial effects:
[0031] 1. This method for enhancing large language models for medical action chain optimization systematically improves the accuracy, response speed, and multimodal adaptability of traditional large language models (LLMs) in question-answering tasks by constructing a structured question decomposition mechanism, combining multiple heterogeneous retrieval action modules, and introducing a credibility verification scoring mechanism.
[0032] 2. This large language model enhancement method for optimizing medical action chains is highly reliable and significantly reduces hallucination content by evaluating the multi-reference and multi-source fact alignment index, ensuring that answers are consistent with the data source;
[0033] 3. This large language model enhancement method for optimizing medical action chains has strong multimodal support and can simultaneously process text, structured data, and vector knowledge to meet complex question-answering needs;
[0034] 4. This large language model enhancement method for optimizing medical action chains achieves low interaction overhead, adopts a knowledge boundary identification strategy to reduce unnecessary external searches and LLM calls, and significantly reduces token usage;
[0035] 5. This method for optimizing large language models for medical action chain enhancement has high response efficiency, demonstrating low interaction times and low latency in multiple question-answering benchmarks (such as WQA, FEVER, and ASQA).
[0036] 6. This large language model enhancement method for optimizing the medical action chain has good practicality and scalability. It can be deployed in the Web3 question-answering system, supports the access and call of multi-source heterogeneous data, and can ensure low latency and low call cost while significantly improving the accuracy and consistency of answers, thus meeting user needs to a large extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1is a step flow chart of a medical action chain optimization large language model enhancement method provided by embodiment 1 of the present application.
[0038] Figure 2 is an action schematic diagram of an action chain of a medical action chain optimization large language model enhancement method provided by embodiment 1 of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0040] Embodiment 1
[0041] As shown in Figure 1 and Figure 2 , a medical action chain optimization large language model enhancement method comprises the following steps:
[0042] Taking the user input question as the starting point, a large language model LLM automatically generates a structured action chain. Each question is divided into several sub-questions, and each sub-question constitutes an action node. The action node includes: sub-question content, missing flag (whether to retrieve to fill in the blank content), initial guess answer and recommended initial processing action scheme.
[0043] As shown in Figure 2 , the following is an action chain template:
[0044] ACQ = (Action1, Sub1, MF1, A1), → (Action2, Sub2, MF2, A2),...,
[0045] → (Actionn, Subn, MFn, An).
[0046] Wherein, ACQ is the original question input by the user, which is the starting point of the entire action chain reasoning; Sub is the sub-question obtained by being divided, A is the direct initial answer obtained by the large language model LLM for the sub-question; MF is the missing information flag, which marks whether the current sub-question cannot be answered by the LLM itself; Action is an initial processing action scheme allocated by the LLM.
[0047] This structure supports chain execution and can also be scheduled in parallel to improve overall response efficiency.
[0048] To improve the ability to acquire and utilize external information, a variety of pluggable information acquisition actions are designed for different data types, such as the following three actions:
[0049] 1. Web-querying action:
[0050] By calling the search engine API, the sub-problem is queried by keywords, and the web page title and abstract are extracted. The semantic similarity with the original problem is calculated by embedding the vector model, and the top-k most relevant web page content is selected as factual material for subsequent answer revision.
[0051] 2. Knowledge-encoding action:
[0052] After segmenting and cutting the local domain knowledge (such as white paper, blog, product document, etc.), it is embedded into the vector database. When retrieving, the sub-problem is embedded and compared with the knowledge base by vector, and the text with the highest similarity is selected as the external support for the answer to the sub-problem.
[0053] 3. Data-analyzing action:
[0054] For sub-problems that require numerical information, the system can automatically generate SQL statements or Python analysis code to connect pre-defined APIs or data tables to obtain structured results such as market data and historical records. This module is suitable for highly data-dependent scenarios such as finance and medicine.
[0055] After completing the action allocation, the action execution and retrieval operation is performed to meet the retrieval requirements. Each action will complete three steps: retrieving relevant information, verifying the conflict between the initial answer and the retrieved information, and inferring the missing answer using the retrieved information when necessary.
[0056] To ensure the authenticity of the answer content, the system introduces the "Evidence-Aligned Consistency Index" (EACI) to verify the consistency between the direct initial answer and the external retrieval data. The scoring basis includes precision (P), recall (R), and average word length (AWL), and the total score is calculated using a weighted method:
[0057] where α, β, γ are weight coefficients corresponding to the importance of precision (Precision), recall (Recall), and average word length (Average Word Length) of the candidate abstract. Their values can be adjusted according to specific requirements, but they must satisfy the normalization condition: α+β+γ=1.
[0058] Among them, precision is used to determine the proportion of relevant tags in the search results, and the calculation formula is: .
[0059] The recall rate is the proportion of actually relevant tags that are retrieved, and the calculation formula is: .
[0060] The average word length is the arithmetic mean of the word lengths in the candidate summaries, and is calculated as follows: .
[0061] When the score is lower than the preset threshold, the system will automatically use the search results to replace the original answer, thereby reducing the hallucination generation rate and improving the credibility of the answer.
[0062] After executing all action chain nodes, the system organizes the corrected answers to each sub-question into a unified context and, using a specific prompt template, guides the LLM in generating the final answer. The generated content possesses logical integrity, contextual consistency, and external information support, making it suitable for demanding multi-step question-answering tasks.
[0063] Example 2
[0064] In this example, the input question is "Should I get the HPV vaccine?" The system first receives this natural language question and invokes the Large Language Model (LLM) to generate an action chain based on a pre-set template. At this stage, the question is broken down into several sub-questions, such as "What is HPV?", "What are the effects of the HPV vaccine?", "Is my age within the recommended vaccination range?", "Are there any side effects from the vaccine?", "What are the current vaccination policies or guidelines?", and "What do medical experts think of the HPV vaccine?"
[0065] For each sub-question, the system will evaluate whether it can be answered directly by the model. If the knowledge within the model is insufficient or there is uncertainty, it will be marked as missing and the appropriate action module will be selected: for example, "Is my age suitable for vaccination" will be assigned to the structured data analysis action to obtain the recommended age range for vaccination in the country / region from the structured medical database; "Side effect information" and "Expert attitude" will use web query actions to capture clinical research, case feedback, expert interviews and other content from the Internet; and "Vaccination policy" and "Guideline interpretation" will use knowledge embedding actions to extract structured recommendation information from documents issued by authoritative medical institutions (such as CDC and WHO guidelines).
[0066] The system then performs information retrieval for each sub-question, constructs a joint semantic representation of the question and the preliminary answer, and performs similarity matching with the search results. Based on this, the system calculates a multi-reference, multi-source fact alignment index to determine the credibility of the preliminary answer generated by the model. If the multi-reference, multi-source fact alignment index score falls below a set threshold, the system replaces the original answer with a highly similar search result. If the preliminary answer to a sub-question is missing, the system automatically fills in the search information.
[0067] After all sub-questions have been processed, the system will reorganize the verified and completed answers and generate a logically coherent, credible, and data-supported final answer through a language model. For example, the final answer might include: "Based on the current national immunization program, World Health Organization guidelines, and statistical results from clinical studies, HPV vaccines are effective in preventing high-risk viral infections and diseases such as cervical cancer. For women aged 18 to 26 who have not completed vaccination, early vaccination is still recommended. Aside from the common minor redness and swelling at the injection site, the vaccine is generally safe, and experts generally support vaccination. Therefore, considering age, health status, and vaccination guidelines, completing the HPV vaccination at this time is a scientific and recommended decision."
[0068] Example 3
[0069] A large language model enhancement system for optimizing a medical action chain, used to implement the large language model enhancement method for optimizing a medical action chain of embodiments 1 and 2, comprising:
[0070] Question parsing and action chain generation module: This module uses the large language model (LLM) to decompose the user's input question into several sub-questions, generates a structured action chain using any sub-question as a node, and sets an initial processing action plan for each node's sub-question.
[0071] Initial answer acquisition module: evaluates whether each sub-question can be directly answered by the large language model (LLM). If so, the initial answer is obtained directly. If not, the answer is marked as missing.
[0072] Answer verification module: performs information retrieval on several sub-questions based on the initial processing action plan, and verifies the answer based on the answer obtained in S2. If the verification passes, content detection is performed. If the verification fails, the answer is revised based on the information retrieval results and verification is repeated until the verification passes.
[0073] Answer completion module: performs content detection and completion on the verified answers to obtain the final answer for each node;
[0074] Final answer generation module: Integrate the final answers of each node to obtain the final answer to the question entered by the user.
[0075] It should be noted that although the detailed description above mentions several modules or units for voice collection and processing, voiceprint recognition, and voice content recognition of the device used for action execution, this division is not mandatory. In fact, the features and functions of two or more modules or units described above can be embodied in a single module or unit. Conversely, the features and functions of a single module or unit described above can be further divided and embodied by multiple modules or units.
[0076] Those skilled in the art will appreciate that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system." Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in this disclosure. The description and embodiments are to be regarded as exemplary only, and the true scope and spirit of the invention are indicated by the claims.
[0077] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings and that various modifications and variations can be made without departing from the scope thereof, which is limited only by the appended claims.
[0078] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A large language model enhancement method for optimizing medical action chains, characterized in that: include: S1. Call the Large Language Model (LLM) to decompose the user's input question into several sub-questions, generate a structured action chain with any sub-question as a node, and set an initial processing action plan for the sub-question of each node; S2. Evaluate whether each sub-question can be answered directly by the large language model (LLM). If so, obtain the initial answer directly. If not, mark it as missing. S3: Perform information retrieval on several sub-questions based on the initial processing action plan, and verify the answers based on the answers obtained in S2. If the verification passes, perform content detection. If the verification fails, modify the initial answer based on the information retrieval results and repeat the verification until the verification passes. S4. Perform content detection and completion on the verified answers to obtain the final answer for each node; S5. Integrate the final answers of each node to obtain the final answer to the question input by the user.
2. The method for enhancing a large language model for optimizing a medical action chain according to claim 1, characterized in that: The node contains the content of the sub-question, content analysis, initial processing action plan, and behavior type. When evaluating whether each sub-question can be directly answered by the large language model (LLM), content analysis is performed based on the content of the sub-question and missing information is marked. If it is marked, it is a sub-question that cannot be directly answered.
3. The method for enhancing a large language model for optimizing a medical action chain according to claim 1 or 2, characterized in that: The initial processing action scheme is based on multiple types of pluggable information acquisition for different data types, including: web query actions, knowledge embedding actions, and structured data analysis actions.
4. The method for enhancing a large language model for optimizing a medical action chain according to claim 3, characterized in that: The web page query action calls the search engine API to perform keyword queries on sub-questions, extract web page titles and summaries, calculate the semantic similarity with the original question through the embedding vector model, screen out the top-k most relevant web page content, and use it as factual material for subsequent answer corrections.
5. The method for enhancing a large language model for optimizing a medical action chain according to claim 3, characterized in that: The knowledge embedding action segments the local domain knowledge into segments and embeds it into the vector database. During retrieval, the sub-questions are embedded and then compared with the vectors in the knowledge base. The text with the highest similarity is selected as the external support for the answer to the sub-question.
6. The method for enhancing a large language model for optimizing a medical action chain according to claim 3, characterized in that: Structured data analysis actions automatically generate SQL statements or Python analysis codes for sub-problems that require numerical information, and obtain structured results by connecting to predefined APIs or data tables.
7. The method for enhancing a large language model for optimizing a medical action chain according to claim 1, characterized in that: Answer verification includes: constructing a joint semantic representation of the sub-question and its answer, performing similarity matching with the information retrieval results to obtain a multi-reference multi-source fact alignment index. If the multi-reference multi-source fact alignment index score is lower than the set threshold, the original answer is replaced by the high-similarity retrieval content; if the answer is missing, the answer from the information retrieval is automatically filled in.
8. The method for enhancing a large language model for optimizing a medical action chain according to claim 7, characterized in that: The scoring basis of the multi-reference multi-source fact alignment index includes precision, recall and average word length, and the total score is calculated using a weighted method.
9. The method for enhancing a large language model for optimizing a medical action chain according to claim 1, characterized in that: After content detection and completion, it is determined whether there are any missing nodes. If so, the missing nodes are returned to S1 for further processing.
10. A system for enhancing a large language model for optimizing a medical action chain, for implementing the method for enhancing a large language model for optimizing a medical action chain as claimed in any one of claims 1 to 9, characterized in that: include: Question parsing and action chain generation module: This module uses the large language model (LLM) to decompose the user's input question into several sub-questions, generates a structured action chain using any sub-question as a node, and sets an initial processing action plan for each node's sub-question. Initial answer acquisition module: evaluates whether each sub-question can be directly answered by the large language model (LLM). If so, the initial answer is obtained directly. If not, the answer is marked as missing. Answer verification module: performs information retrieval on several sub-questions based on the initial processing action plan, and verifies the answer based on the answer obtained in S2. If the verification passes, content detection is performed. If the verification fails, the answer is revised based on the information retrieval results and verification is repeated until the verification passes. Answer completion module: performs content detection and completion on the verified answers to obtain the final answer for each node; Final answer generation module: Integrate the final answers of each node to obtain the final answer to the question entered by the user.
Citation Information
Patent Citations
Knowledge intensive question reasoning and generating method based on LLM
CN118798367A
Metadata-based question and answer method and related device
CN119669412A
Knowledge graph question and answer method combining large model and graph neural network
CN119848190A
Large language model answer reasoning method and system fusing memory and iterative optimization
CN119918654A
Direct current equipment operation inspection intelligent agent, evaluation method, system and related equipment
CN120163491A