Medical large model analysis method and system based on multi-strategy deep slow thinking and storage medium
By applying multi-strategy deep slow thinking medical model analysis method in speech recognition and medical knowledge retrieval, the semantic errors of the speech recognition system, insufficient rigor of medical knowledge retrieval and low information density of knowledge recall are solved, and higher system reliability, knowledge retrieval accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202510625266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the prior art, the high error rate of speech recognition systems leads to semantic errors, affecting information transmission and user decision-making; large language models lack rigor and correctness in knowledge retrieval in the medical field, making it difficult to meet the high accuracy requirements of clinical applications; traditional knowledge recall methods ignore context continuity and semantic coherence, resulting in low information density and insufficient correlation.
The medical big model analysis method based on multi-strategy deep slow thinking is adopted to generate complete search terms through contextual context and semantic analysis technology, combine private medical knowledge base and network search to perform error correction and knowledge recall, and conduct in-depth reasoning and self-dialectical thinking through natural language processing models to ensure the accuracy and credibility of the answers.
It significantly improves the reliability and user experience of the voice interaction system, improves the rigor and correctness of knowledge retrieval in the medical field, enhances the relevance and information density of knowledge recalls, and ensures the quality and practicality of generated content.
Smart Images

Figure CN120144729A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of health medical consultation, and particularly to a medical large model analysis method, system and storage medium based on multi-strategy deep slow thinking. Background Art
[0002] Semantic errors caused by ASR (Automatic Speech Recognition) lead to incorrect model generation: With the wide application of Automatic Speech Recognition (ASR) technology, especially in the medical field and various speech-based interaction platforms, the accuracy of ASR is crucial for user experience and system performance. However, current ASR systems still face many challenges in practical applications, resulting in a high recognition error rate. These errors mainly include homophone confusion, inaccurate recognition of accents and dialects, misrecognition caused by fast speech rate or unclear pronunciation, and background noise interference in noisy environments.
[0003] These ASR errors not only stay at the surface text inaccuracy, but also cause a series of problems at the semantic level. Specifically, when the ASR system incorrectly recognizes certain words or sentence structures, downstream natural language processing (NLP) models that rely on these transcribed texts, such as large language models (LLMs), may generate inaccurate, misleading or context-inconsistent responses based on the wrong information. This semantic error not only affects the correct transmission of information, but also may lead to user misunderstandings, wrong decisions, and even serious consequences in some critical application scenarios. In professional fields such as medicine, semantic errors caused by ASR may cause serious communication barriers and safety hazards.
[0004] In addition, existing error correction methods mostly rely on post-processing steps or manual rules, lacking automated and intelligent error correction capabilities, and it is difficult to correct ASR errors in real time and accurately. Therefore, there is an urgent need for a technology that can automatically detect and correct ASR errors, especially to improve accuracy at the semantic level, to make up for the deficiencies of existing systems and improve the reliability and user experience of the overall speech interaction system.
[0005] The generated retrieval terms have no context, resulting in the disappearance of the context of the current statement: In the service of large language models (LLMs), information retrieval is a key step to improve the accuracy and relevance of answers. Traditional retrieval methods usually generate retrieval terms based on the user's current query (query), and then search for relevant materials through the network. However, in multi-turn dialogue scenarios, relying only on the current query to generate retrieval terms often ignores the context and context information of the previous and subsequent conversations. This approach may lead to the generated retrieval terms not matching the theme of the entire conversation or the user's intention, thus causing deviations or errors in the retrieval results.
[0006] Lack of rigor and correctness in medical knowledge In the current medical field, with the rapid development of artificial intelligence technology, medical models based on large-scale data are gradually applied to various aspects such as clinical diagnosis, treatment plan recommendation, and health management. However, the rigor and correctness of medical knowledge are crucial for the reliability and safety of these models. The sources of medical knowledge in the open world are extensive, and the information quality is uneven, making it difficult to distinguish between true and false, and lacking unified standards and scientific basis. This uncertainty and complexity of knowledge can easily lead to misjudgments when the model processes information, thereby having a negative impact on medical decision-making. In addition, the privacy and sensitivity of medical data further limit the acquisition and application of high-quality and accurate medical knowledge. Existing medical large models still have deficiencies in knowledge verification, information screening, and accuracy guarantee, and it is difficult to fully meet the high requirements for absolute correctness in clinical applications. Especially when facing emerging diseases, complex cases, and multi-source heterogeneous data, the performance of the model is often unsatisfactory, affecting its application effect and trust in actual medical treatment. Therefore, there is an urgent need for a new technology that can effectively distinguish and screen true and scientific medical knowledge and ensure that the information on which medical large models are based has a high degree of rigor and correctness. This will not only help improve the reliability and clinical application value of medical artificial intelligence systems but also promote the development of intelligent medicine and provide more accurate and safe medical services for patients.
[0007] According to text block recall, most are useless or irrelevant knowledge:
[0008] In the service application of large language models (LLMs), common knowledge acquisition methods mainly include web search and recall from private medical knowledge bases. To effectively process massive text data, existing Retrieval-Augmented Generation (RAG) methods usually adopt text chunking technology to split large-scale documents into smaller text blocks for processing. Although this chunking processing method can improve retrieval efficiency, it also brings a series of problems.
[0009] First of all, text chunking often ignores the continuity of context and the coherence of overall semantics, resulting in the system retrieving a large number of knowledge fragments that are irrelevant or have low relevance to the user's query during the recall process. This not only increases the complexity of subsequent processing but also may introduce noise information, affecting the quality and accuracy of the generated content.Secondly, since the information density of the text fragments after chunking is relatively low and the amount of effective information contained in each fragment is limited, it becomes difficult for the system to comprehensively understand and apply relevant knowledge when integrating multiple fragments. This problem of information fragmentation further reduces the effectiveness of knowledge recall, making it difficult for the generated answers to meet the needs of users in terms of both depth and breadth.
[0010] In addition, when existing RAG methods handle specific domains or professional knowledge, it is often difficult to accurately capture subtle semantic differences and the meanings of professional terms, further affecting the accuracy and practicality of knowledge recall. Therefore, how to improve the relevance and information density of the recalled knowledge while maintaining efficient retrieval has become a technical problem that urgently needs to be solved in current large language model services. Summary of the Invention
[0011] The purpose of this application is to provide a medical large model analysis method, system, and storage medium based on multi-strategy deep slow thinking to solve one or more technical problems existing in the prior art, and at least provide a beneficial choice or create conditions.
[0012] The present application adopts the following technical solutions to achieve the above-mentioned invention purpose: The present application provides a medical large model analysis method based on multi-strategy deep slow thinking, including: Obtain the query input by the user, and deeply analyze the query through context and semantic analysis techniques to generate complete context-related retrieval terms; Recall relevant content from the knowledge base according to the generated retrieval terms and determine whether error correction is required: When the recalled content has errors or is ambiguous, use a lightweight language model to enter error correction, generate a new query after error correction, regenerate new retrieval terms based on the new query after error correction, perform web search and knowledge base recall according to the newly generated retrieval terms, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to the natural language processing large model for processing, and return the final answer to the user; When the recalled content has no errors or is not ambiguous, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to the natural language processing large model for processing, and return the final answer to the user.
[0013] Furthermore, the method of deeply analyzing the query through context and semantic analysis techniques to generate complete context-related retrieval terms includes: Through the automatic speech recognition error correction mechanism, combined with context and private medical knowledge base recall, to ensure the accuracy of the input information; Combined with the multi-round dialogue history queue, a dialogue summarizer based on a pure decoder architecture is used to extract key information and generate complete context-related retrieval terms.
[0014] Furthermore, the method for network search and knowledge base recall includes: Efficient retrieval of the private medical knowledge base through the vector database ChromaDB, which integrates screened medical resources; Conduct targeted retrieval through a crawler program within the scope of limited professional medical websites; The retrieved results are sorted and filtered for relevance through a lightweight language model based on a pure decoder architecture; The formula for the relevance score of the lightweight language model is as follows: ; In the formula, Score(q, d) represents the relevance score between the query q and the document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the retrieval algorithm based on term frequency-inverse document frequency, and LLM cls represents the relevance score output by the classifier using the lightweight language model; The formula for the classification layer is as follows: ; In the formula, p is the final relevance probability distribution, h is the hidden state of the last layer of the lightweight language model, W is the parameter matrix of the classification layer, b is the bias term of the classification layer, and represents the probability vector of the query-document relevance.
[0015] Furthermore, the method for organizing the recalled knowledge according to the current query includes: Based on the knowledge summary and refinement mechanism of the current query and context, a secretary model is used to organize various recalled knowledge; The retrieved knowledge is processed for density enhancement through a sliding window and a key information extraction algorithm, and the formula is as follows: ; In the formula, Density(T) represents the information density of the text segment, T represents the text segment, K i represents the i-th key information item, W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items.
[0016] Furthermore, the method for sending the organized knowledge and the query to the natural language processing large model for processing and returning the final answer to the user includes: Send the organized knowledge and the query to the natural language processing large model. The natural language processing large model adopts a deep thought chain mechanism and explores the reasoning path through the Monte Carlo tree search algorithm. Its node selection strategy uses the upper confidence bound formula: ; In the formula, UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is the exploration constant that controls the balance between exploration and exploitation; Continuously evaluate and verify the reasoning process through the self-dialectical thinking mechanism. Its credibility score uses the following formula: ; In the formula, Confidence represents the credibility score of the final answer, λ 1 , λ 2 , λ 3 are weight coefficients that respectively measure knowledge support, logical consistency, and context relevance. Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context; Add a classification linear layer to the last layer of the natural language processing large model to score the credibility of the final answer and return the final answer to the user.
[0017] This application provides a medical large model analysis system based on multi-strategy deep slow thinking, including: An acquisition unit for acquiring the query input by the user, deeply parsing the query through context and semantic analysis technologies, and generating complete context-related retrieval terms; An error correction unit for recalling relevant content from the knowledge base according to the generated retrieval terms and judging whether error correction is required; A first judgment and generation unit for, when the recalled content is incorrect or ambiguous, using a lightweight language model to enter error correction, generating a new query after error correction, regenerating new retrieval terms according to the new query after error correction, performing web search and knowledge base recall according to the newly generated retrieval terms, organizing the recalled knowledge according to the current query, and sending the organized knowledge and the query to the natural language processing large model for processing, and returning the final answer to the user; A second judgment generation unit, configured to directly perform network search and knowledge base recall when the recalled content has no errors or ambiguities, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to a natural language processing large model for processing, and return the final answer to the user.
[0018] The present application provides an electronic device, including a memory and a processor; The memory is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the above method.
[0019] The present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0020] The beneficial effects of the present application are as follows: The present application not only improves the relevance and information density of knowledge recall, but also effectively avoids the generation of errors and hallucinations through a self-evaluation mechanism, significantly improving the quality and practicality of the content generated by the natural language processing large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is an overall process block diagram of a medical large model analysis method based on multi-strategy deep slow thinking according to an embodiment of the present application; Figure 2 It is a schematic diagram of an AI interface in the first medical large model analysis method based on multi-strategy deep slow thinking according to an embodiment of the present application; Figure 3 It is a schematic diagram of an AI interface in the second medical large model analysis method based on multi-strategy deep slow thinking according to an embodiment of the present application; Figure 4 It is a schematic diagram of an AI interface in the third medical large model analysis method based on multi-strategy deep slow thinking according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present application realizes the accurate analysis and output of medical knowledge by constructing multiple function modules that work together. Its core is a natural language processing large model, which realizes complex reasoning capabilities through a deep thinking chain mechanism. In the specific implementation process, the system first processes the input voice information through an ASR error automatic correction module, and adopts methods such as context retrieval and private medical knowledge base recall to ensure the accuracy of the input information.
[0023] In the knowledge processing stage, the system adopts a dual knowledge base architecture: on the one hand, a strictly controlled private medical knowledge base is established to integrate high-quality resources such as authoritative medical books, research papers, and guidelines; on the other hand, the search scope of professional medical websites is restricted, and a reranker model is used for knowledge screening to ensure the reliability of external knowledge. To improve the efficiency of knowledge utilization, the system also designs a knowledge summarization mechanism based on the current query and context.
[0024] In the core natural language processing large model, a two-stage training strategy is adopted: first, the model is fine-tuned through supervised learning, and then reinforcement learning is used to deepen the training. The model realizes the automatic synthesis of thought chain data samples through the Monte Carlo Tree Search (MCTS) algorithm and the Process Reward Model (PRM) algorithm, and introduces the rejection sampling technique to ensure the quality of the samples. During the multi-round dialogue process, the system generates context-related retrieval terms through the historical dialogue summarization mechanism to maintain the coherence and accuracy of the dialogue.
[0025] Finally, the system introduces a self-dialectical thinking mechanism to autonomously evaluate the content generated by the model, and through a text integration method with high information density, ensures the accuracy and reliability of the output results. The entire processing flow forms a complete intelligent analysis closed-loop, realizing the full-process intelligent processing from input processing, knowledge acquisition, in-depth reasoning to result output.
[0026] Referring to Figures 1-4 , this application provides a medical large model analysis method based on multi-strategy deep slow thinking, including the following steps: S100, Obtain the query input by the user, and deeply analyze the query through context and semantic analysis technologies to generate complete context-related retrieval terms, specifically including: S101, The process starts with the user inputting a query. After the system receives the user input, it enters the requirement analysis stage. The goal of this stage is to accurately understand the user's intention and provide a basis for subsequent search and knowledge processing; S102, After receiving the question input by the user, the system deeply analyzes the question through context and semantic analysis technologies. The input processing module combines the multi-round dialogue history queue and uses a dialogue summarizer based on a pure decoder architecture to extract key information and generate complete context-related retrieval terms. These retrieval terms will be used for subsequent knowledge base recall and web search to ensure that the system can accurately understand the user's needs and provide targeted services. In addition, the input processing module also effectively solves problems such as homophone confusion and inaccurate accent and dialect recognition through the ASR error automatic correction mechanism, combined with context and private medical knowledge base recall.
[0027] S200. Recall content by combining the generated search terms with the knowledge base and correct errors using the capabilities of the lightweight language model. When the recalled content contains errors or is ambiguous, enter the error correction process to generate a new corrected query. The new corrected query is used to generate new search terms. Based on the newly generated search terms, conduct web searches and knowledge base recalls. Organize the recalled knowledge according to the current query and send the organized knowledge and the query to the large natural language processing model for processing, and return the final answer to the user, specifically including: S201. After accurately analyzing the requirements, the knowledge acquisition module is used to combine the content returned by the knowledge base and correct errors using the capabilities of the LLM to generate a new corrected question. The system recalls relevant information from the knowledge base based on the generated search terms and determines whether this information needs to be corrected. If the recalled content contains errors or is ambiguous, enter the error correction process; if no correction is required, directly conduct web searches and knowledge base recalls.
[0028] S202. In the error correction process, the system combines the content returned by the knowledge base and uses the capabilities of the LLM to perform intelligent error correction. The error correction process generates a new question that more accurately reflects the user's needs. The corrected question will be used again to generate new search terms to support subsequent web searches and knowledge base recalls.
[0029] S203. The system simultaneously conducts web searches and knowledge base recalls based on the newly generated search terms. The knowledge base recall part adopts a dual guarantee mechanism: on the one hand, the vector database ChromaDB is used to achieve efficient retrieval of the private medical knowledge base, which integrates strictly screened medical resources; on the other hand, a crawler program conducts targeted searches within the scope of restricted professional medical websites. All retrieval results are sorted and filtered based on the relevance by a lightweight LLM model based on a pure decoder architecture.
[0030] The formula for calculating the relevance score of the lightweight LLM model is:
[0031] In the formula, Score(q, d) represents the relevance score between the query q and the document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the classical retrieval algorithm based on term frequency - inverse document frequency, and LLM cls represents the relevance score output by the classifier using the lightweight language model. The calculation process of the classification layer is:
[0032] Where p is the final relevance probability distribution, h is the hidden state of the last layer of the LLM, W is the parameter matrix of the classification layer, and b is the bias term of the classification layer, representing the probability vector of the query-document relevance.
[0033] S204. The retrieved knowledge is processed by a sliding window and a key information extraction algorithm to enhance its density, ensuring the efficient utilization of knowledge. The knowledge density is calculated using the following formula:
[0034] Where Density(T) represents the information density of the text segment, T represents the text segment, K i represents the i-th key information item, W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items. In addition, the system has a specially designed knowledge summarization and refinement mechanism based on the current question and context, using a secretary model (another lightweight decoder-only LLM) to organize various recalled knowledge, improving the relevance and usability of the knowledge. This step ensures that the subsequent reasoning process can conduct in-depth analysis based on high-quality knowledge.
[0035] S205. The sorted high-quality knowledge and the question are input into the core natural language processing large model. This model also adopts a pure decoder architecture but maintains a large parameter scale (about 7B parameters) to ensure reasoning ability. The model uses an innovative deep thought chain mechanism to explore the reasoning path through the MCTS algorithm, and its node selection strategy uses the following formula:
[0036] Where UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is the exploration constant, controlling the balance between exploration and exploitation. During actual operation, the model continuously evaluates and validates the reasoning process through a self-dialectical thinking mechanism, and its credibility score uses:
[0037] Where Confidence represents the credibility score of the final answer, λ 1 、λ 2 、λ 3is the weight coefficient, which measures the knowledge support degree, logical consistency, and context relevance respectively. Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context. The system also adds a classification linear layer at the last layer of the natural language processing large model to score the credibility of the final answer and ensure that accurate and reliable answers are returned to users.
[0038] The entire above-mentioned step process realizes asynchronous communication between modules through a message queue, optimizes performance using a multi-level caching mechanism, and sets up a perfect error retry mechanism to ensure system stability. All core configuration parameters can be flexibly adjusted through a configuration file, supporting customized deployment according to the actual application scenario. In addition, the system also implements a complete logging and monitoring mechanism, which can track the system operation status and performance metrics in real time, providing a basis for system optimization.
[0039] This application has the following technical effects: Solve the semantic errors brought by ASR (Automatic Speech Recognition): The technical solution of this application has made significant improvements to the high error rate and its resulting semantic-level problems existing in the existing automatic speech recognition technology (ASR) in practical applications. Existing ASR systems often result in inaccurate recognized text due to reasons such as homophone confusion, inaccurate accent and dialect recognition, too fast speech rate, or environmental noise interference. These errors not only affect the user experience but also cause information misguidance in downstream natural language processing (NLP) models that rely on the transcribed text, and even cause serious consequences in critical fields such as medicine.
[0040] This application breaks through the limitations of traditional reliance on post-processing steps or manual rules by constructing a complete system for automatically correcting ASR errors, and has a high degree of automation and intelligent error correction capabilities. Specifically, the system covers multiple steps such as the generation of retrieval words for context semantic, recall from a private medical knowledge base, web search, collation of recalled text blocks to improve information density, and in-depth thinking and self-dialectics of the model. This multi-level and multi-step processing flow can significantly improve the detection and correction effect of ASR errors in terms of real-time performance and accuracy.
[0041] In addition, this application enhances the comprehensiveness and accuracy of information by integrating a private medical knowledge base and network resources, making the error correction process not only rely on preset rules but dynamically adapt to different contexts and requirements. The in-depth thinking and self-dialectics ability of the model further ensure the semantic consistency and context relevance of the error correction results.
[0042] In summary, the present application effectively makes up for the deficiencies of the prior art in terms of automation and intelligent error correction capabilities, significantly improving the reliability of the voice interaction system and the user experience.
[0043] Solving the problem of the disappearance of the retrieval word context in multi-turn conversations: Existing retrieval word generation methods usually rely only on the user's current query, ignoring the context and contextual information of the previous and subsequent conversations. This single generation method results in the retrieval words being unable to accurately reflect the theme of the entire conversation and the user's true intention, thereby causing deviations or errors in the retrieval results and affecting the quality of the answers and the user experience.
[0044] In multi-turn conversations, each question from the user often relies on the previous communication content to form a coherent context. Traditional methods lack comprehensive analysis of historical conversations and are difficult to capture the evolution of the user's intention. Especially when dealing with vague or polysemous queries, it is easy to lead to inaccuracies in the retrieval results and information omission.
[0045] To address the above problems, the present application introduces a historical conversation summarization mechanism to generate retrieval words by combining the recent few rounds of questions and answers, rather than relying only on the current query. This method can comprehensively integrate the context information of multi-turn conversations, ensuring that the generated retrieval words are more in line with the actual needs of the user and the conversation context. In addition, by designing high-quality examples (shots) for the model to refer to, the quality and consistency of retrieval word generation are further improved.
[0046] By comprehensively considering the context and contextual information of the conversation, the present application significantly improves the limitations of existing retrieval word generation methods, enhances the accuracy of information retrieval and the intelligent level of the system, thereby providing more reliable and efficient technical support for multi-turn conversation systems and greatly optimizing the user experience.
[0047] Solving the rigor, correctness, and interpretability of knowledge in the medical field: The present application significantly improves the problems of insufficient knowledge rigor and correctness existing in existing medical artificial intelligence technologies by establishing a private database with strictly controlled data sources. Currently, the knowledge sources in the medical field are extensive and of uneven quality, resulting in medical models based on large-scale data being prone to misjudgment in clinical applications, affecting the reliability and safety of medical decisions. In addition, the privacy and sensitivity of medical data further limit the acquisition of high-quality knowledge, and existing models have obvious deficiencies in knowledge verification and information screening, making it difficult to meet the high requirements for absolute accuracy in clinical practice.
[0048] This application ensures that all data sources are highly scientific and accurate by integrating highly recognized medical books, influential paper journals, authoritative cases, internationally recognized guideline consensuses, strictly screened disease libraries and phenotype libraries, as well as nationally recognized drug libraries and internationally authoritative diagnosis and treatment libraries. This high-standard data management not only effectively differentiates and screens real and scientific medical knowledge, but also greatly improves the performance of medical large models in dealing with emerging diseases, complex cases and multi-source heterogeneous data, enhancing the reliability and clinical application value of the models.
[0049] By adopting the technical solution of this application, the knowledge base of the medical artificial intelligence system is more solid and the decision-making process is more credible, greatly improving the development level of intelligent medicine, providing more accurate and safe medical services for patients, and solving the key defects of inaccurate knowledge and insufficient model reliability in the prior art.
[0050] Solve the problem of the confusion of medical knowledge on the network: This application aims to significantly improve the reliability and accuracy of the application of large language models in the medical field, and effectively improves the problems of mixed network knowledge sources and uneven information quality existing in the prior art. Currently, large language models rely on extensive network searches to obtain knowledge. However, network information, especially medical-related content, is often affected by media reports, lacks scientific argumentation, and easily leads to the "hallucination" phenomenon of the model, thus misleading users and bringing health risks and legal liabilities.
[0051] After receiving the search term and user query, this application limits the search scope to multiple professional medical query websites, blocks non-medical and media websites, strictly controls the knowledge sources, and ensures that the content input into the model is authoritative and scientifically based. In addition, the reranker model is used to sort and screen the retrieved knowledge to further improve the knowledge quality and ensure that the information referred to by the model is the highest-quality professional materials. This dual control mechanism effectively eliminates the interference of unreliable information and significantly reduces the risk of generating incorrect or misleading answers.
[0052] The advantage of this technical solution is that it enhances the credibility and security of large language models in the medical field, ensures that the information obtained by users is strictly screened and verified, and avoids potential health and legal problems caused by false information. At the same time, it improves the user's trust in the model.
[0053] Solve the problem that the retrieved knowledge has low information density for the current problem: The technical solution of this application aims to improve the low relevance and insufficient information density problems of existing large language models during the knowledge recall process. The current Retrieval-Augmented Generation (RAG) method mainly relies on text chunking technology, which divides large-scale documents into small pieces for retrieval. However, this approach often ignores the continuity of context and the coherence of overall semantics, resulting in a large amount of useless or less relevant knowledge being recalled. This not only increases the complexity of model processing but also may introduce noise, affecting the quality and accuracy of the generated content.
[0054] To address the above deficiencies, before the model generates an answer, this application adds a mechanism to summarize and refine the recalled knowledge based on the current query and context. Through this step, the system can filter out the most relevant and high-information-density knowledge fragments, thus effectively improving the relevance and effectiveness of knowledge recall. This improvement not only reduces the pressure on the model to understand and process irrelevant information but also increases the depth and breadth of the generated answer, ensuring that the answer content is more accurate and valuable.
[0055] In addition, when dealing with specific domain or professional knowledge, this application can more accurately capture subtle semantic differences and the meanings of professional terms, significantly improving the accuracy and practicality of knowledge recall. Overall, this application enhances the performance of large language models in complex application scenarios by optimizing the knowledge processing flow.
[0056] Solving the problems of recalling incorrect knowledge, model hallucinations, or random answers: This application aims to improve several key deficiencies existing in the existing large language models in the Retrieval-Augmented Generation (RAG) method. In the prior art, although text chunking improves the retrieval efficiency, it often ignores the continuity of context and the coherence of semantics, resulting in a large amount of useless or irrelevant knowledge being recalled, with serious information fragmentation, thus affecting the quality and accuracy of the generated content. In addition, when dealing with specific domain or professional knowledge, the existing RAG methods are difficult to capture subtle semantic differences and professional terms, further reducing the accuracy and practicality of knowledge recall.
[0057] This application enhances the model's autonomous evaluation ability when using recalled knowledge by introducing a self-dialectical thinking mechanism. When referring to the recalled content, the model can automatically judge its relevance and correctness to the current query, effectively filtering out irrelevant or incorrect knowledge fragments and reducing the interference of noise information. At the same time, a text integration method with high information density is adopted to ensure that when integrating multiple knowledge fragments, the model can comprehensively understand and apply relevant knowledge, thereby increasing the depth and breadth of the generated answer. During the training phase, accurate medical data and long-thinking data are introduced, and the dialectical thinking ability is strengthened, making the model more accurate and reliable when dealing with professional domain knowledge.
[0058] In summary, this application not only improves the relevance and information density of knowledge recall, but also effectively avoids the generation of errors and hallucinations through a self-assessment mechanism, significantly improving the quality and practicality of content generated by large language models.
[0059] Addressing the shortcomings of general generative models: This application aims to significantly improve the application effect of general generative models in the field of health and medicine, and proposes innovative improvements to the shortcomings of existing technologies in logical analysis, reasoning ability and application of professional knowledge. The specific technical solution includes introducing a large amount of medical pre-training corpus in the pre-training stage to ensure that the model has a deep understanding of the medical field at the basic knowledge level. At the same time, in the supervised fine-tuning (SFT) stage, a large amount of medical long-term thinking data is combined with general SFT corpus, such as mathematical reasoning and code reasoning, to comprehensively enhance the thinking and reasoning ability of the model.
[0060] Through these optimizations, the natural language processing large model is not only significantly superior to traditional large language models in logical analysis and reasoning, but also has excellent capabilities in processing complex medical queries. Its unique deep thinking mechanism can generate complete and visual reasoning links, improving the credibility and transparency of output results. In addition, the reflective learning mechanism enables the model to have the ability to self-examine and continuously optimize, significantly reducing misjudgments and hallucinations, especially in professional fields such as medical diagnosis, showing higher accuracy.
[0061] In addition, the large natural language processing model has a strong ability to understand information. Even when faced with brief or ambiguous medical input, it can supplement the context through internal reasoning, accurately grasp the user's intention, and provide professional and complete answers. These improvements not only improve the application effect of the model in the field of health and medicine, but also enhance its adaptability and performance in other complex reasoning tasks.
[0062] The computer program of this application demonstrates powerful medical knowledge processing and analysis capabilities in practical applications. The system supports intelligent processing of multiple types of medical data and achieves accurate knowledge output in different application scenarios. The specific application features and process steps of the system are as follows: 1. Input processing capability The system has established a complete intelligent error correction mechanism to address the high error rate of speech recognition (ASR) technology in practical applications. For voice input queries, the system can effectively handle issues such as confusion of homophones, inaccurate recognition of accents and dialects, too fast speech speed, or interference from environmental noise. For example, in medical query scenarios, the system can accurately identify dialect terms and automatically understand the Sichuan dialect "马巴路杀尾片是干什子" as the standard term "马巴路杀尾片是什么"; for confusion of drug names caused by similar pronunciations, such as "五子延中丸" will be automatically corrected to the standard name "五子延宗丸".
[0063] By integrating context semantic analysis and private medical knowledge base recall functions, the system breaks through the limitations of traditional methods that rely on post - processing steps or manual rules. When processing professional queries related to diagnosis and treatment plans and medical knowledge, the system combines the query context, and through entity recognition and semantic understanding, accurately grasps the user's intention. At the same time, the system maintains a multi - turn dialogue history queue, and uses a dialogue summarizer based on the Transformer architecture to extract key information, achieving a precise understanding of the query intention.
[0064] To ensure the reliability of the error - correction results, the system adopts a multi - level verification mechanism: first, it performs basic error - correction through a professional medical dictionary, then conducts semantic - level verification in combination with the context, and finally verifies through the knowledge base to ensure the professional accuracy of the error - correction results. This deep semantic understanding and knowledge support enable the system to significantly improve the detection and correction effect of ASR errors while maintaining real - time performance.
[0065] II. Retrieval Term Generation Mechanism The system innovatively introduces a historical dialogue summarization mechanism, breaking through the limitations of traditional retrieval term generation methods that only rely on the current query. When processing user queries, the system maintains a dynamic dialogue history queue, and uses a dialogue summarizer based on the Transformer architecture to extract key information, achieving a precise grasp of the multi - turn dialogue theme and user intention.
[0066] In view of the particularity of the medical field, the system adopts a Chinese - English bilingual retrieval strategy. Based on the comprehensive analysis of the dialogue history, the system can not only automatically generate standardized combinations of Chinese retrieval terms, but also simultaneously generate corresponding standard English medical terms. This bilingual retrieval mechanism ensures the comprehensiveness of knowledge coverage and can obtain the latest medical research results at home and abroad.
[0067] The system further improves the quality and consistency of retrieval term generation by designing high - quality examples (shots) as model references. During the generation process, the system considers the following key factors: The professional relevance of medical concepts; The context coherence of multi - turn dialogues; The evolution trajectory of user intentions; The integrity of the query topic; This retrieval term generation mechanism based on historical dialogue summarization significantly improves the accuracy of the system in processing vague or polysemous queries, effectively avoids information omission, and ensures that the retrieval results can accurately reflect the user's true needs and dialogue context. At the same time, the system's dynamic adaptation ability enables it to continuously optimize the retrieval strategy according to the development of the dialogue, providing users with more accurate medical knowledge services.
[0068] III. Knowledge Base Recall Architecture The system has established a strictly controlled private medical knowledge base, integrating core resources such as highly recognized medical professional books, academic journal papers with high influence, internationally recognized clinical guidelines and expert consensus, strictly screened disease databases and phenotype databases, nationally recognized drug databases, and internationally authoritative diagnosis and treatment standard databases, etc., to ensure that all data sources are highly scientific and accurate.
[0069] The API of the knowledge base module returns highly structured data, including call status flags, error codes, and the core data part. The core data covers key information such as strictly indexed reference content, authoritative document source URLs (such as national drug standards, clinical guidelines, etc.), resource type annotations (such as national standards, clinical guidelines, etc.), knowledge credibility scores, and data update timestamps, realizing the rigor and traceability of knowledge.
[0070] To ensure the absolute accuracy of medical knowledge, the system has implemented a comprehensive multi-verification mechanism. By controlling the source, it ensures that only medical knowledge released by authoritative institutions is adopted. Regularly updating the knowledge base guarantees the timeliness of information. Using cross-verification of multi-source data ensures the consistency of knowledge, and establishing a professional review mechanism ensures the rigor of content. This strict knowledge management system significantly improves the reliability of the system in dealing with emerging diseases, complex cases, and multi-source heterogeneous data, providing more accurate and safe knowledge support for clinical decision-making.
[0071] IV. Network Search and Recall System In response to the problems of mixed knowledge sources and uneven information quality in the medical field, the system has established a strict knowledge source access and screening mechanism. After receiving the search term and user query, the system strictly limits the search scope to certified professional medical data sources, effectively shielding the interference of non-professional and media websites, and ensuring the authority and science of knowledge sources.
[0072] The authoritative data sources that the system preferentially connects to include: internationally open-access medical literature databases such as PubMed, the official database of the National Medical Products Administration, the national standard drug instruction database, the traditional Chinese medicine database certified by the National Administration of Traditional Chinese Medicine, and the publicly available medical knowledge base of top-three hospitals after strict screening. Through this strict access mechanism, it is ensured that all content input into the model is highly professional and reliable.
[0073] To further improve the knowledge quality, the system uses a reranker model to intelligently sort and screen the retrieved content, preferentially retaining the most authoritative and timely professional materials. This dual control mechanism effectively reduces the risk of generating incorrect or misleading answers, significantly improving the credibility and safety of the system in the medical field, and providing users with strictly verified high-quality medical information.
[0074] V. Knowledge Organization Process To address the problems of insufficient information density and low relevance in traditional Retrieval-Augmented Generation (RAG) methods, the system innovatively introduces an intelligent knowledge processing mechanism based on the secretary model. This mechanism not only considers the literal content of the text but also fully focuses on the continuity of the context and the coherence of the overall semantics, ensuring the scientificity and effectiveness of the knowledge organization process.
[0075] In the specific implementation process, the system first performs multi-dimensional processing on the retrieved reference documents, including basic tasks such as information deduplication, content classification, and credibility assessment. Subsequently, based on the current query and conversation context, the system conducts in-depth knowledge fusion and key extraction, focusing on screening out knowledge fragments that are highly relevant to the user's needs and have a high information density. This refined processing method effectively reduces the interference of irrelevant information and improves the knowledge utilization efficiency.
[0076] Especially when dealing with professional medical knowledge, the system can accurately grasp subtle semantic differences and the meanings of professional terms, ensuring that the organized content is both professional and rigorous and easy to understand through intelligent knowledge extraction. This optimized knowledge processing flow significantly improves the quality and depth of the system's response, providing a more accurate and valuable knowledge basis for subsequent in-depth analysis.
[0077] VI. Inference Output Mechanism of the Natural Language Processing Large Model The natural language processing large model is designed based on the decoder of the Transformer architecture and uses an innovative dual-model collaborative mechanism for medical reasoning. First, the preliminary medical text understanding and analysis are completed through the basic natural language processing large model, which uses word vectorization processing and RMSNorm normalization technology to ensure the standardization of input data and computational stability.
[0078] In the core processing stage, the model achieves in-depth understanding of medical texts through multiple layers of attention mechanisms. Each layer contains a residual connection structure, effectively preventing the problem of gradient disappearance. At the same time, global context information is captured through attention weight calculation to achieve accurate understanding and analysis of complex medical texts. To improve the reasoning quality, the system introduces MCTS (Monte Carlo Tree Search) and PRM planning algorithms for sample synthesis and optimization, and screens the optimal reasoning path through a rejection sampling mechanism.
[0079] On this basis, the RAG-enhanced natural language processing large model will conduct in-depth analysis and content supplementation based on the retrieved professional medical knowledge. The system also establishes a continuous optimization mechanism, scores and screens the generated results through the Reward Model, and uses the DPO (Direct Preference Optimization) method for model alignment to improve the comprehensibility of the output while maintaining professionalism. This iterative optimization ensures that the system can continuously improve its reasoning ability and output quality.
[0080] This multi-level reasoning mechanism not only ensures the accuracy of basic medical reasoning, but also effectively supplements professional details through knowledge enhancement and continuous optimization, ultimately generating medical answers that are both professional and rigorous and easy to understand, providing reliable intelligent support for clinical decision-making.
[0081] This application provides a medical large model analysis system based on multi-strategy deep slow thinking, including: An acquisition unit is used to acquire the query input by the user, deeply analyze the query through context and semantic analysis technology, and generate complete context-related search terms; An error correction unit, used to recall relevant content from the knowledge base based on the generated search terms and determine whether error correction is required; The first judgment generation unit is used to use a lightweight language model to enter error correction when there are errors or ambiguities in the recalled content, generate a new query after error correction, regenerate new search terms after error correction, perform network search and knowledge base recall based on the newly generated search terms, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user; The second judgment generation unit is used to directly perform network search and knowledge base recall when there is no error or ambiguity in the recalled content, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user.
[0082] The electronic device provided in the present application may also include: a memory and a processor; the memory is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the aforementioned medical large model analysis method based on multi-strategy deep slow thinking.
[0083] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned medical large model analysis method based on multi-strategy deep slow thinking.
[0084] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0085] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0088] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.
Claims
1. A medical large model analysis method based on multi-strategy deep slow thinking, characterized in that: include: Get the query entered by the user, analyze the query in depth through context and semantic analysis technology, and generate complete context-related search terms; Recall relevant content from the knowledge base based on the generated search terms and determine whether correction is needed: When the recalled content is wrong or ambiguous, a lightweight language model is used to correct the error and generate a new query after the error correction. The new query after the error correction regenerates new search terms. Based on the newly generated search terms, network search and knowledge base recall are performed. The recalled knowledge is sorted according to the current query, and the sorted knowledge and query are sent to the large natural language processing model for processing, and the final answer is returned to the user. When the recalled content does not contain errors or ambiguities, a direct web search and knowledge base recall is performed, the recalled knowledge is sorted according to the current query, and the sorted knowledge and query are sent to the natural language processing large model for processing, and the final answer is returned to the user.
2. According to claim 1, a medical large model analysis method based on multi-strategy deep slow thinking is characterized in that: The method of deeply analyzing the query through context and semantic analysis technology to generate complete context-related search terms includes: The accuracy of input information is ensured through automatic correction of speech recognition errors, combined with context and private medical knowledge base recall; Combined with the multi-round dialogue history queue, the dialogue summarizer based on the pure decoder architecture is used to extract key information and generate complete context-related search terms.
3. According to claim 1, a medical large model analysis method based on multi-strategy deep slow thinking is characterized in that: Methods for conducting web searches and knowledge base recall, including: Search the private medical knowledge base through the vector database ChromaDB, which integrates the selected medical resources; Conduct targeted searches within a limited range of professional medical websites through crawlers; The search results are sorted and filtered by relevance using a lightweight language model based on a pure decoder architecture; The formula for the relevance score of the lightweight language model is as follows: ; In the formula, Score(q, d) represents the relevance score between query q and document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the retrieval algorithm based on term frequency-inverse document frequency, and LLM cls represents the relevance score of the classifier output using a lightweight language model; The formula for the classification layer is as follows: ; Where p is the final relevance probability distribution, h is the hidden state of the last layer of the lightweight language model, W is the parameter matrix of the classification layer, and b is the bias term of the classification layer, which represents the probability vector of the relevance between the query and the document.
4. According to claim 1, a medical large model analysis method based on multi-strategy deep slow thinking is characterized in that: Methods for organizing the recalled knowledge according to the current query include: Based on the knowledge summary and refinement mechanism of the current query and context, a lightweight language model is used to organize various types of recalled knowledge; The retrieved knowledge is processed through sliding window and key information extraction algorithm to increase density. The formula is as follows: ; In the formula, Density(T) represents the information density of the text segment, T represents the text segment, and K i represents the i-th key information item, W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items.
5. According to claim 1, a medical large model analysis method based on multi-strategy deep slow thinking is characterized in that: The method of sending the organized knowledge and query to the natural language processing large model for processing and returning the final answer to the user includes: The organized knowledge and query are sent to the natural language processing model. The natural language processing model adopts a deep thinking chain mechanism and explores the reasoning path through the Monte Carlo tree search algorithm. Its node selection strategy adopts the upper confidence bound formula: ; Where UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is the exploration constant, which controls the balance between exploration and exploitation. The reasoning process is continuously evaluated and verified through the self-dialectical thinking mechanism, and its credibility score is based on the following formula: ; In the formula, Confidence represents the credibility score of the final answer, λ1, λ2, and λ3 are weight coefficients, which measure knowledge support, logical consistency, and context relevance respectively. Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context. A classification linear layer is added to the last layer of the natural language processing model to score the credibility of the final answer and return the final answer to the user.
6. A medical large model analysis system based on multi-strategy deep slow thinking, characterized by: include: The acquisition unit is used to acquire the query input by the user, deeply analyze the query through context and semantic analysis technology, and generate complete context-related search terms; An error correction unit, used to recall relevant content from the knowledge base based on the generated search terms and determine whether error correction is required; The first judgment generation unit is used to use a lightweight language model to enter error correction when there are errors or ambiguities in the recalled content, generate a new query after error correction, regenerate new search terms after error correction, perform network search and knowledge base recall based on the newly generated search terms, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user; The second judgment generation unit is used to directly perform network search and knowledge base recall when there is no error or ambiguity in the recalled content, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user.
7. An electronic device, characterized in that: including memory and processor; The memory is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Interactive knowledge interaction system based on knowledge base and large language model
CN118013001A
Medical text error correction system and medical query prompt text display method and device
CN118093789A
Text error correction language model training method, text error correction method and related products
CN119398035A
Multimodal reasoning method based on Monte Carlo tree and dynamic retrieval
CN119808941A
System and method to extract software development requirements from natural language
US20210200515A1
Cited By
Technology development situation awareness system and method
CN120429414A
A technology development situation awareness system and method
CN120429414B
Multi-strategy retrieval and query adaptive combined enhanced generation method and system
CN121455985A
A multi-strategy retrieval and query adaptive combination enhanced generation method and system
CN121455985B
Text interaction method, system and equipment based on large language model and medium
CN121561031A