A medical large model analysis method, system, and storage medium based on multi-strategy deep slow thinking
Through a multi-strategy deep slow thinking method, combining context and private medical knowledge base, ASR errors are automatically corrected, complete search terms are generated, and authoritative medical resources are integrated to solve the problems of ASR errors, knowledge retrieval bias and low information density, and the reliability and accuracy of voice interaction systems and medical big models are improved.
Patent Information
- Application Number
- CN202510625266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing speech recognition technology (ASR) has a high recognition error rate in the medical field, resulting in semantic errors, affecting information transmission and user experience; large language models ignore the context and context in knowledge retrieval, resulting in bias in search results; existing medical big models have poor knowledge sources, affecting decision-making reliability; low information density and insufficient correlation during the knowledge recall process.
Through a multi-strategy deep slow thinking method, combining context and private medical knowledge base, ASR errors are automatically corrected, complete search terms are generated, authoritative medical resources are integrated, lightweight language models are used to correct errors, and efficient knowledge base and network search are used to improve information density and self-dialectical thinking to ensure the relevance and accuracy of knowledge.
It significantly improves the reliability and user experience of the voice interaction system, improves the relevance and information density of knowledge recalls, ensures the quality and practicality of generated content, and enhances the credibility and security of medical big models.
Smart Images

Figure CN120144729B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of health medical consultation, and particularly to a medical large model analysis method, system and storage medium based on multi-strategy deep slow thinking. Background Art
[0002] Semantic errors generated by ASR (Automatic Speech Recognition) lead to model generation errors:
[0003] With the wide application of Automatic Speech Recognition (ASR) technology, especially in the medical field and various speech-based interaction platforms, the accuracy of ASR is crucial for user experience and system performance. However, current ASR systems still face many challenges in practical applications, resulting in a high recognition error rate. These errors mainly include homophone confusion, inaccurate recognition of accents and dialects, misrecognition caused by fast speech rate or unclear pronunciation, and background noise interference in noisy environments.
[0004] These ASR errors not only stay at the surface level of text inaccuracy, but also cause a series of problems at the semantic level. Specifically, when the ASR system misrecognizes certain words or sentence structures, downstream Natural Language Processing (NLP) models that rely on these transcribed texts, such as Large Language Models (LLMs), may generate inaccurate, misleading or context-inconsistent responses based on the wrong information. Such semantic errors not only affect the correct transmission of information, but may also lead to user misunderstandings, wrong decisions, and even serious consequences in some critical application scenarios. In professional fields such as medicine, semantic errors caused by ASR may cause serious communication barriers and safety hazards.
[0005] In addition, existing error correction methods mostly rely on post-processing steps or manual rules, lacking automated and intelligent error correction capabilities, and it is difficult to correct ASR errors in real time and accurately. Therefore, there is an urgent need for a technology that can automatically detect and correct ASR errors, especially to improve accuracy at the semantic level, to make up for the deficiencies of existing systems and improve the reliability and user experience of the overall speech interaction system.
[0006] The generation of retrieval terms has no context, resulting in the disappearance of the context of the current sentence:
[0007] In large language model (LLM) services, information retrieval is a crucial step in improving the accuracy and relevance of answers. Traditional retrieval methods usually generate retrieval terms based on the user's current query and then search for relevant materials online. However, in multi-turn conversation scenarios, relying solely on the current query to generate retrieval terms often ignores the context and contextual information of the previous and subsequent conversations. This approach may lead to retrieval terms that do not match the overall theme of the conversation or the user's intention, resulting in biased or incorrect retrieval results.
[0008] Lack of rigor and correctness in medical knowledge:
[0009] In the current medical field, with the rapid development of artificial intelligence technology, medical models based on large-scale data are gradually applied to multiple aspects such as clinical diagnosis, treatment plan recommendation, and health management. However, the rigor and correctness of medical knowledge are crucial for the reliability and safety of these models. The sources of medical knowledge in the open world are extensive, and the information quality is uneven, making it difficult to distinguish between true and false, and lacking unified standards and scientific basis. This uncertainty and complexity of knowledge can easily lead to misjudgments when the model processes information, which in turn has a negative impact on medical decision-making. In addition, the privacy and sensitivity of medical data further limit the acquisition and application of high-quality and accurate medical knowledge. Existing medical large models still have deficiencies in knowledge verification, information screening, and accuracy guarantee, and it is difficult to fully meet the high requirements of clinical applications for absolute correctness. Especially when facing emerging diseases, complex cases, and multi-source heterogeneous data, the performance of the model is often unsatisfactory, affecting its application effect and trust in actual medical treatment. Therefore, there is an urgent need for a new technology that can effectively distinguish and screen true and scientific medical knowledge and ensure that the information on which medical large models are based is highly rigorous and correct. This will not only help improve the reliability and clinical application value of medical artificial intelligence systems but also promote the development of intelligent medicine and provide more accurate and safe medical services for patients.
[0010] When retrieving according to text chunks, most of the retrieved knowledge is useless or irrelevant:
[0011] In the service application of large language models (LLMs), common knowledge acquisition methods mainly include online search and retrieval from private medical knowledge bases. To effectively process massive amounts of text data, existing Retrieval-Augmented Generation (RAG) methods usually adopt text chunking technology to split large-scale documents into smaller text chunks for processing. Although this chunking method can improve retrieval efficiency, it also brings a series of problems.
[0012] First of all, text chunking often ignores the continuity of context and the coherence of overall semantics, resulting in the system potentially retrieving a large number of knowledge fragments that are irrelevant or have low relevance to the user's query during the recall process. This not only increases the complexity of subsequent processing but also may introduce noisy information, affecting the quality and accuracy of the generated content.
[0013] Secondly, due to the low information density of the text fragments after chunking, the amount of effective information contained in each fragment is limited, making it difficult for the system to comprehensively understand and apply relevant knowledge when integrating multiple fragments. This problem of information fragmentation further reduces the effectiveness of knowledge recall, making the generated answers difficult to meet the user's needs in terms of depth and breadth.
[0014] In addition, existing RAG methods often struggle to accurately capture subtle semantic differences and the meanings of specialized terms when dealing with specific domains or professional knowledge, further affecting the precision and practicality of knowledge recall. Therefore, how to improve the relevance and information density of the recalled knowledge while maintaining efficient retrieval has become a technical problem that urgently needs to be solved in current large language model services. Summary of the Invention
[0015] The purpose of this application is to provide a medical large model analysis method, system, and storage medium based on multi-strategy deep slow thinking to solve one or more technical problems in the prior art and at least provide a beneficial choice or create conditions.
[0016] The present application adopts the following technical solutions to achieve the above-mentioned invention purpose:
[0017] The present application provides a medical large model analysis method based on multi-strategy deep slow thinking, including:
[0018] Obtain the query input by the user, deeply analyze the query through context and semantic analysis techniques, and generate complete context-related retrieval terms;
[0019] Recall relevant content from the knowledge base according to the generated retrieval terms and determine whether error correction is required:
[0020] When the recalled content has errors or is ambiguous, use a lightweight language model to enter error correction, generate a new query after error correction, regenerate new retrieval terms based on the new query after error correction, conduct web search and knowledge base recall according to the newly generated retrieval terms, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to a natural language processing large model for processing, and return the final answer to the user;
[0021] When the recalled content has no errors or ambiguities, the recalled knowledge is organized according to the current query, and the organized knowledge and the query are sent to the natural language processing large model for processing, and the final answer is returned to the user.
[0022] Furthermore, a method for deeply parsing a query through context and semantic analysis technology to generate complete context-related retrieval terms includes:
[0023] Through the automatic speech recognition error correction mechanism, combined with context and private medical knowledge base recall, to ensure the accuracy of the input information;
[0024] Combined with the multi-round dialogue history queue, use a dialogue summarizer based on a pure decoder architecture to extract key information and generate complete context-related retrieval terms.
[0025] Furthermore, a method for web search and knowledge base recall includes:
[0026] Efficiently retrieve the private medical knowledge base through the vector database ChromaDB, which integrates screened medical resources;
[0027] Conduct targeted retrieval through a crawler program within the scope of restricted professional medical websites;
[0028] Rank and filter the retrieved results through a lightweight language model based on a pure decoder architecture;
[0029] The formula for the relevance score of the lightweight language model is as follows:
[0030] ;
[0031] In the formula, Score(q, d) represents the relevance score between query q and document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the retrieval algorithm based on term frequency-inverse document frequency, and LLM cls represents the relevance score output by the classifier of the lightweight language model;
[0032] The formula for the classification layer is as follows:
[0033] ;
[0034] In the formula, p is the final relevance probability distribution, h is the hidden state of the last layer of the lightweight language model, W is the parameter matrix of the classification layer, b is the bias term of the classification layer, and represents the probability vector of the relevance between the query and the document.
[0035] Further, a method for organizing the recalled knowledge according to the current query includes:
[0036] A knowledge summarization and refinement mechanism based on the current query and context, using a secretary model to organize various types of recalled knowledge;
[0037] The retrieved knowledge is processed by a sliding window and a key information extraction algorithm to enhance its density. The formula is as follows:
[0038] ;
[0039] In the formula, Density(T) represents the information density of the text segment, T represents the text segment, K i represents the i-th key information item, W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items.
[0040] Further, a method for sending the organized knowledge and the query to a natural language processing large model for processing and returning the final answer to the user includes:
[0041] Send the organized knowledge and the query to a natural language processing large model. The natural language processing large model uses a deep thinking chain mechanism to explore the reasoning path through the Monte Carlo tree search algorithm. Its node selection strategy uses the upper confidence bound formula:
[0042] ;
[0043] In the formula, UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is an exploration constant that controls the balance between exploration and exploitation;
[0044] Continuously evaluate and verify the reasoning process through a self-dialectical thinking mechanism. Its credibility score uses the following formula:
[0045] ;
[0046] In the formula, Confidence represents the credibility score of the final answer, λ1, λ2, and λ3 are weight coefficients that measure knowledge support, logical consistency, and context relevance respectively. Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context;
[0047] Add a classification linear layer to the last layer of the natural language processing large model to score the credibility of the final answer and return the final answer to the user.
[0048] This application provides a medical large model analysis system based on multi-strategy deep slow thinking, including:
[0049] An acquisition unit for acquiring the query input by the user, deeply parsing the query through context and semantic analysis technologies, and generating complete context-related retrieval terms;
[0050] An error correction unit for recalling relevant content from the knowledge base according to the generated retrieval terms and judging whether error correction is required;
[0051] A first judgment and generation unit, when there are errors or ambiguities in the recalled content, using a lightweight language model to enter error correction, generating a new corrected query, generating new retrieval terms for the new corrected query, performing web search and knowledge base recall according to the newly generated retrieval terms, organizing the recalled knowledge according to the current query, and sending the organized knowledge and the query to the natural language processing large model for processing, and returning the final answer to the user;
[0052] A second judgment and generation unit, when there are no errors or ambiguities in the recalled content, directly performing web search and knowledge base recall, organizing the recalled knowledge according to the current query, and sending the organized knowledge and the query to the natural language processing large model for processing, and returning the final answer to the user.
[0053] This application provides an electronic device, including a memory and a processor;
[0054] The memory is used to store instructions;
[0055] The processor is used to operate according to the instructions to execute the steps of the above method.
[0056] This application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0057] The beneficial effects of this application are as follows:
[0058] This application not only improves the relevance and information density of knowledge recall, but also effectively avoids the generation of errors and hallucinations through a self-assessment mechanism, significantly improving the quality and practicality of the content generated by the natural language processing large model. Description of the Drawings
[0059] Figure 1 It is the overall process block diagram of a medical large model analysis method provided according to an embodiment of this application;
[0060] Figure 2 Schematic diagram of the AI interface in the first medical large model analysis method based on multi-strategy deep slow thinking provided according to an embodiment of the present application;
[0061] Figure 3 Schematic diagram of the AI interface in the second medical large model analysis method based on multi-strategy deep slow thinking provided according to an embodiment of the present application;
[0062] Figure 4 Schematic diagram of the AI interface in the third medical large model analysis method based on multi-strategy deep slow thinking provided according to an embodiment of the present application. Detailed implementation manners
[0063] The present application realizes the accurate analysis and output of medical knowledge by constructing multiple functional modules that work together. Its core is a natural language processing large model, which realizes complex reasoning capabilities through a deep thinking chain mechanism. In the specific implementation process, the system first processes the input voice information through an ASR error automatic correction module, and adopts methods such as context retrieval and private medical knowledge base recall to ensure the accuracy of the input information.
[0064] In the knowledge processing link, the system adopts a dual knowledge base architecture: on the one hand, a strictly controlled medical private medical knowledge base is established to integrate high-quality resources such as authoritative medical books, paper journals, and guideline consensuses; on the other hand, by limiting the search scope of professional medical websites and using a reranker model for knowledge screening, the reliability of external knowledge is ensured. To improve the knowledge utilization efficiency, the system also designs a knowledge summary and refinement mechanism based on the current query and context.
[0065] In the core natural language processing large model, a two-stage training strategy is adopted: first, the model is fine-tuned through supervised learning, and then deep training is carried out using reinforcement learning. The model realizes the automatic synthesis of thinking chain data samples through the Monte Carlo tree search (MCTS) algorithm and the process reward model (PRM) algorithm, and introduces rejection sampling technology to ensure the sample quality. During the multi-round dialogue process, the system generates context-related retrieval terms through a historical dialogue summary mechanism to maintain the coherence and accuracy of the dialogue.
[0066] Finally, the system introduces a self-dialectical thinking mechanism to autonomously evaluate the content generated by the model, and through a high-information-density text integration method, ensures the accuracy and reliability of the output results. The entire processing flow forms a complete intelligent analysis closed-loop, realizing the full-process intelligent processing from input processing, knowledge acquisition, deep reasoning to result output.
[0067] Refer to Figures 1 - 4, this application provides a medical large model analysis method based on multi-strategy deep slow thinking, including the following steps:
[0068] S100, Obtain the query input by the user, and deeply analyze the query through context and semantic analysis technologies to generate complete context-related retrieval terms, specifically including:
[0069] S101, The process starts with the user entering a query. After the system receives the user's input, it enters the requirement analysis stage. The goal of this stage is to accurately understand the user's intention and provide a basis for subsequent search and knowledge processing;
[0070] S102, After receiving the question input by the user, the system deeply analyzes the question through context and semantic analysis technologies. The input processing module combines the multi-round dialogue history queue and uses a dialogue summarizer based on a pure decoder architecture to extract key information and generate complete context-related retrieval terms. These retrieval terms will be used for subsequent knowledge base recall and web search to ensure that the system can accurately understand the user's needs and provide targeted services. In addition, the input processing module also effectively solves problems such as homophone confusion and inaccurate accent and dialect recognition through the ASR error automatic correction mechanism, combined with context and private medical knowledge base recall.
[0071] S200, Combine the generated retrieval terms with the knowledge base recall content and use the capabilities of a lightweight language model to correct errors. When the recalled content is incorrect or ambiguous, enter the error correction link to generate a new query after error correction. The new query after error correction regenerates new retrieval terms. According to the newly generated retrieval terms, perform web search and knowledge base recall, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to the natural language processing large model for processing, and return the final answer to the user, specifically including:
[0072] S201, After accurately analyzing the requirements, the knowledge acquisition module is used to combine the content returned by the knowledge base and use the capabilities of the LLM to correct errors and generate a new question after error correction. The system recalls relevant information from the knowledge base according to the generated retrieval terms and judges whether this information needs to be corrected. If the recalled content is incorrect or ambiguous, enter the error correction link; if no correction is required, directly perform web search and knowledge base recall.
[0073] S202, In the error correction link, the system combines the content returned by the knowledge base and uses the capabilities of the LLM to perform intelligent error correction. The error correction process generates a new question that more accurately reflects the user's needs. The question after error correction will be reused to generate new retrieval terms to support subsequent web search and knowledge base recall.
[0074] In S203, the system simultaneously conducts web searches and retrieves information from the knowledge base according to the newly generated search terms. The knowledge base retrieval part adopts a dual safeguard mechanism: on the one hand, it realizes the efficient retrieval of the private medical knowledge base through the vector database ChromaDB, which integrates strictly screened medical resources; on the other hand, it conducts targeted searches within a limited range of professional medical websites through a crawler program. All retrieval results are sorted and filtered for relevance through a lightweight LLM model based on a pure decoder architecture.
[0075] The formula for calculating the relevance score of the lightweight LLM model is as follows:
[0076]
[0077] In the formula, Score(q, d) represents the relevance score between the query q and the document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the classical retrieval algorithm based on term frequency - inverse document frequency, and LLM cls represents the relevance score output by the classifier using the lightweight language model. The calculation process of the classification layer is as follows:
[0078]
[0079] In the formula, p is the final relevance probability distribution, h is the hidden state of the last layer of the LLM, W is the parameter matrix of the classification layer, b is the bias term of the classification layer, and it represents the probability vector of the relevance between the query and the document.
[0080] In S204, the retrieved knowledge is processed through a sliding window and a key information extraction algorithm to enhance the density, ensuring the efficient utilization of knowledge. The knowledge density is calculated using the following formula:
[0081]
[0082] In the formula, Density(T) represents the information density of the text segment, T represents the text segment, K i represents the i-th key information item, and W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items. In addition, the system has specially designed a knowledge summary and refinement mechanism based on the current problem and context, and uses a secretary model (another lightweight decoder-only LLM) to organize various retrieved knowledge, enhancing the relevance and usability of the knowledge. This step ensures that the subsequent reasoning process can conduct in-depth analysis based on high-quality knowledge.
[0083] S205, The high-quality knowledge after collation, together with the questions, is input into the core natural language processing large model. This model also adopts a pure decoder architecture but maintains a large parameter scale (about 7B parameters) to ensure inference ability. The model uses an innovative deep thinking chain mechanism to explore the inference path through the MCTS algorithm, and its node selection strategy uses the following formula:
[0084]
[0085] In the formula, UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is the exploration constant that controls the balance between exploration and exploitation. During actual operation, the model continuously evaluates and validates the inference process through a self-dialectical thinking mechanism, and its credibility score uses:
[0086]
[0087] In the formula, Confidence represents the credibility score of the final answer, λ1, λ2, and λ3 are weight coefficients that measure knowledge support, logical consistency, and context relevance respectively, Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the inference process, and Context_Relevance represents the relevance of the answer to the query context. The system also adds a classification linear layer at the last layer of the natural language processing large model to score the credibility of the final answer and ensure that accurate and reliable answers are returned to users.
[0088] The above entire step process realizes asynchronous communication between modules through a message queue, optimizes performance using a multi-level caching mechanism, and sets up a perfect error retry mechanism to ensure system stability. All core configuration parameters can be flexibly adjusted through a configuration file, supporting customized deployment according to the actual application scenario. In addition, the system also implements a complete logging and monitoring mechanism, which can track the system operation status and performance metrics in real time, providing a basis for system optimization.
[0089] This application has the following technical effects:
[0090] Solve the semantic errors brought by ASR (Automatic Speech Recognition):
[0091] The technical solution of this application has significantly improved the high error rate and the resulting semantic-level problems in the existing Automatic Speech Recognition (ASR) technology in practical applications. Existing ASR systems often result in inaccurate recognized text due to reasons such as homophone confusion, inaccurate recognition of accents and dialects, too fast speech rate, or environmental noise interference. These errors not only affect the user experience but also cause information misguidance in downstream Natural Language Processing (NLP) models that rely on the transcribed text, and even cause serious consequences in critical fields such as healthcare.
[0092] This application has broken through the limitations of traditional methods that rely on post-processing steps or manual rules by constructing a complete system for automatically correcting ASR errors, and has a high degree of automation and intelligent error correction capabilities. Specifically, the system covers multiple steps such as generating retrieval words based on the semantic context, recalling from a private medical knowledge base, web search, organizing the recalled text blocks to improve information density, and the model's in-depth thinking and self-dialectics. This multi-level and multi-step processing flow can significantly improve the detection and correction effects of ASR errors in terms of real-time performance and accuracy.
[0093] In addition, this application enhances the comprehensiveness and accuracy of information by integrating a private medical knowledge base and network resources, enabling the error correction process to not only rely on preset rules but also dynamically adapt to different contexts and requirements. The model's in-depth thinking and self-dialectics capabilities further ensure the semantic consistency and context relevance of the error correction results.
[0094] In summary, this application effectively makes up for the deficiencies of the existing technology in terms of automation and intelligent error correction capabilities, and greatly improves the reliability and user experience of the voice interaction system.
[0095] Solving the problem of the disappearance of the retrieval word context in multi-turn conversations:
[0096] Existing methods for generating retrieval words usually only rely on the user's current query, ignoring the context and contextual information of the previous and subsequent conversations. This single generation method results in retrieval words being unable to accurately reflect the theme of the entire conversation and the user's true intention, thereby causing deviations or errors in the retrieval results and affecting the quality of the answers and the user experience.
[0097] In multi-turn conversations, each user's question often depends on the previous communication content, forming a coherent context. Traditional methods lack a comprehensive analysis of historical conversations and are difficult to capture the evolution of the user's intention. Especially when dealing with vague or polysemous queries, it is easy to lead to inaccuracies and information omissions in the retrieval results.
[0098] To address the above problems, this application introduces a historical dialogue summarization mechanism, which generates retrieval terms by combining the Q&A from the recent few rounds instead of relying solely on the current query. This method can comprehensively integrate the context information of multi-round conversations, ensuring that the generated retrieval terms are more in line with the actual needs of users and the conversation context. In addition, by designing high-quality examples (shots) for the model to reference, the quality and consistency of retrieval term generation are further improved.
[0099] By comprehensively considering the context and context information of the conversation, this application significantly improves the limitations of existing retrieval term generation methods, enhances the accuracy of information retrieval and the intelligence level of the system, thus providing more reliable and efficient technical support for multi-round conversation systems and greatly optimizing the user experience.
[0100] Solving the rigor, correctness, and interpretability of knowledge in the medical field:
[0101] This application significantly improves the problems of insufficient knowledge rigor and correctness in existing medical artificial intelligence technologies by establishing a private database with strictly controlled data sources. Currently, the knowledge sources in the medical field are extensive and of uneven quality, resulting in misjudgments in the clinical application of medical models based on large-scale data, affecting the reliability and safety of medical decisions. In addition, the privacy and sensitivity of medical data further limit the acquisition of high-quality knowledge, and existing models have obvious deficiencies in knowledge verification and information screening, making it difficult to meet the high requirements for absolute accuracy in clinical practice.
[0102] This application ensures that all data sources are highly scientific and accurate by integrating highly recognized medical books, influential paper journals, authoritative cases, internationally recognized guideline consensuses, strictly screened disease libraries and phenotype libraries, as well as nationally recognized drug libraries and internationally authoritative diagnosis and treatment libraries. This high-standard data management not only effectively distinguishes and screens real and scientific medical knowledge but also significantly improves the performance of medical large models in dealing with emerging diseases, complex cases, and multi-source heterogeneous data, enhancing the reliability and clinical application value of the models.
[0103] By adopting the technical solution of this application, the knowledge base of the medical artificial intelligence system is more solid, and the decision-making process is more credible, greatly enhancing the development level of intelligent medicine, providing more accurate and safe medical services for patients, and solving the key defects of insufficient knowledge rigor and low model reliability in the existing technology.
[0104] Solving the confusion of medical knowledge on the network:
[0105] This application aims to significantly improve the reliability and accuracy of large language models in the medical field, and effectively addresses the problems of mixed sources of network knowledge and uneven information quality in the existing technology. Currently, large language models rely on extensive web searches to obtain knowledge. However, web information, especially medical-related content, is often influenced by media reports and lacks scientific verification, which easily leads to the "hallucination" phenomenon in model generation, misleading users and bringing health risks and legal liabilities.
[0106] After receiving the search term and user query, this application limits the search scope to multiple professional medical query websites, blocks non-medical and media websites, strictly controls the knowledge sources, and ensures that the content input into the model is authoritative and scientifically based. In addition, a reranker model is used to sort and filter the retrieved knowledge, further improving the knowledge quality and ensuring that the information referred to by the model is the highest-quality professional materials. This dual control mechanism effectively eliminates the interference of unreliable information and significantly reduces the risk of generating incorrect or misleading answers.
[0107] The advantages of this technical solution are that it enhances the credibility and security of large language models in the medical field, ensures that the information obtained by users is strictly screened and verified, and avoids potential health and legal problems caused by false information. At the same time, it improves the user's trust in the model.
[0108] To address the retrieved knowledge and the problem of low information density for the current issue:
[0109] The technical solution of this application aims to improve the low relevance and insufficient information density in the knowledge retrieval process of existing large language models. The current Retrieval-Augmented Generation (RAG) method mainly relies on text chunking technology to split large-scale documents into small pieces for retrieval. However, this approach often ignores the continuity of context and the coherence of overall semantics, resulting in a large amount of useless or less relevant knowledge being retrieved. This not only increases the complexity of model processing but also may introduce noise, affecting the quality and accuracy of the generated content.
[0110] To address the above defects, this application adds a mechanism for summarizing and refining the retrieved knowledge based on the current query and context before the model generates an answer. Through this step, the system can screen out the most relevant and high-information-density knowledge fragments, thereby effectively improving the relevance and effectiveness of knowledge retrieval. This improvement not only reduces the pressure on the model to understand and process irrelevant information but also increases the depth and breadth of the generated answer, ensuring that the answer content is more accurate and valuable.
[0111] In addition, when dealing with specific fields or professional knowledge, this application can more accurately capture subtle semantic differences and the meanings of professional terms, significantly improving the accuracy and practicality of knowledge recall. Overall, this application enhances the performance of large language models in complex application scenarios by optimizing the knowledge processing flow.
[0112] Solve the problems of recalling incorrect knowledge, model hallucinations, or random answers:
[0113] This application aims to improve several key defects in the existing large language models in the Retrieval-Augmented Generation (RAG) method. In the prior art, although text chunking improves retrieval efficiency, it often ignores the continuity of context and semantic coherence, resulting in a large amount of useless or irrelevant knowledge being recalled, with serious information fragmentation, which in turn affects the quality and accuracy of the generated content. In addition, when dealing with specific fields or professional knowledge, the existing RAG methods are difficult to capture subtle semantic differences and professional terms, further reducing the accuracy and practicality of knowledge recall.
[0114] This application enhances the model's autonomous evaluation ability when using recalled knowledge by introducing a self-dialectical thinking mechanism. When referring to the recalled content, the model can automatically judge its relevance and correctness to the current query, effectively filtering out irrelevant or incorrect knowledge fragments and reducing the interference of noise information. At the same time, a text integration method with high information density is adopted to ensure that when integrating multiple knowledge fragments, the model can comprehensively understand and apply relevant knowledge, thereby enhancing the depth and breadth of the generated answers. In the training stage, accurate medical data and long-thinking data are introduced, and the dialectical thinking ability is strengthened, making the model more accurate and reliable when dealing with professional domain knowledge.
[0115] In summary, this application not only improves the relevance and information density of knowledge recall but also effectively avoids the generation of errors and hallucinations through a self-evaluation mechanism, significantly improving the quality and practicality of the content generated by large language models.
[0116] Solve the defects of general generation models:
[0117] This application aims to significantly improve the application effect of general generation models in the field of health medicine, and proposes innovative improvements for the deficiencies in logical analysis, reasoning ability, and professional knowledge application in the prior art. The specific technical solutions include introducing a large amount of medical pre-training corpus in the pre-training stage to ensure that the model has a profound understanding of the medical field at the basic knowledge level. At the same time, in the Supervised Fine-Tuning (SFT) stage, a large amount of medical long-thinking data is combined with general SFT corpus, such as mathematical reasoning and code reasoning, to comprehensively strengthen the thinking and reasoning ability of the model.
[0118] Through these optimizations, the natural language processing large model is not only significantly superior to traditional large language models in logical analysis and reasoning, but also has excellent capabilities in processing complex medical queries. Its unique deep thinking mechanism can generate complete and visual reasoning links, improving the credibility and transparency of output results. In addition, the reflective learning mechanism enables the model to have the ability to self-examine and continuously optimize, significantly reducing misjudgments and hallucinations, especially in professional fields such as medical diagnosis, showing higher accuracy.
[0119] In addition, the large natural language processing model has a strong ability to understand information. Even when faced with brief or ambiguous medical input, it can supplement the context through internal reasoning, accurately grasp the user's intention, and provide professional and complete answers. These improvements not only improve the application effect of the model in the field of health and medicine, but also enhance its adaptability and performance in other complex reasoning tasks.
[0120] The computer program of this application demonstrates powerful medical knowledge processing and analysis capabilities in practical applications. The system supports intelligent processing of multiple types of medical data and achieves accurate knowledge output in different application scenarios. The specific application features and process steps of the system are as follows:
[0121] 1. Input processing capability
[0122] The system has established a complete intelligent error correction mechanism to address the high error rate of speech recognition (ASR) technology in practical applications. For voice input queries, the system can effectively handle issues such as confusion of homophones, inaccurate recognition of accents and dialects, too fast speech speed, or interference from environmental noise. For example, in medical query scenarios, the system can accurately identify dialect terms and automatically understand the Sichuan dialect "马巴路杀尾片是干什子" as the standard term "马巴路杀尾片是什么"; for confusion of drug names caused by similar pronunciations, such as "五子延中丸" will be automatically corrected to the standard name "五子延宗丸".
[0123] The system breaks through the limitations of traditional reliance on post-processing steps or manual rules by integrating contextual semantic analysis and private medical knowledge base recall functions. When processing professional queries on diagnosis and treatment plans and medical knowledge, the system will combine the query context and ensure accurate grasp of user intent through entity recognition and semantic understanding. At the same time, the system maintains a multi-round dialogue history queue, extracts key information through the dialogue summarizer of the transformer architecture, and achieves accurate understanding of query intent.
[0124] To ensure the reliability of the error correction results, the system adopts a multi-level verification mechanism: First, it performs basic error correction through a professional medical dictionary, then conducts semantic verification in combination with the context, and finally verifies through the knowledge base to ensure the professional accuracy of the error correction results. This deep semantic understanding and knowledge support enable the system to significantly improve the detection and correction effect of ASR errors while maintaining real-time performance.
[0125] II. Retrieval Term Generation Mechanism
[0126] The system innovatively introduces a historical dialogue summary mechanism, breaking through the limitation of traditional retrieval term generation methods that only rely on the current query. When processing user queries, the system maintains a dynamic dialogue history queue, and extracts key information through a dialogue summarizer based on the transformer architecture to accurately grasp the multi-round dialogue theme and user intent.
[0127] In view of the particularity of the medical field, the system adopts a Chinese-English bilingual retrieval strategy. Based on the comprehensive analysis of the dialogue history, the system can not only automatically generate standardized Chinese retrieval term combinations, but also simultaneously generate corresponding standard English medical terms. This bilingual retrieval mechanism ensures the comprehensiveness of knowledge coverage and can obtain the latest medical research results at home and abroad.
[0128] The system further improves the quality and consistency of retrieval term generation by designing high-quality examples (shots) as model references. During the generation process, the system considers the following key factors:
[0129] The professional relevance of medical concepts;
[0130] The context coherence of multi-round dialogues;
[0131] The evolution trajectory of user intent;
[0132] The integrity of the query topic;
[0133] This retrieval term generation mechanism based on historical dialogue summary significantly improves the accuracy of the system in processing ambiguous or polysemous queries, effectively avoids information omission, and ensures that the retrieval results can accurately reflect the user's true needs and dialogue context. At the same time, the system's dynamic adaptation ability enables it to continuously optimize the retrieval strategy according to the development of the dialogue and provide users with more accurate medical knowledge services.
[0134] III. Knowledge Base Recall Architecture
[0135] The system has established a strictly controlled private medical knowledge base, integrating core resources such as highly recognized medical professional books, high-impact academic journal papers, internationally recognized clinical guidelines and expert consensus, strictly screened disease databases and phenotype databases, nationally recognized drug databases, and internationally authoritative diagnosis and treatment standard databases, etc., to ensure that all data sources are highly scientific and accurate.
[0136] The API of the knowledge base module returns highly structured data, including call status flags, error codes, and the core data part. The core data covers key information such as strictly indexed reference content, authoritative document source URLs (such as national drug standards, clinical guidelines, etc.), resource type annotations (such as national standards, clinical guidelines, etc.), knowledge credibility scores, and data update timestamps, realizing the rigor and traceability of knowledge.
[0137] To ensure the absolute accuracy of medical knowledge, the system has implemented a comprehensive multi-verification mechanism. By controlling the source, it ensures that only medical knowledge published by authoritative institutions is adopted; regularly updating the knowledge base guarantees the timeliness of information; using cross-verification of multi-source data ensures the consistency of knowledge; and establishing a professional review mechanism ensures the rigor of content. This strict knowledge management system significantly improves the reliability of the system in dealing with emerging diseases, complex cases, and multi-source heterogeneous data, providing more accurate and secure knowledge support for clinical decision-making.
[0138] IV. Network Search and Recall System
[0139] In response to the problems of mixed knowledge sources and uneven information quality in the medical field, the system has established a strict knowledge source access and screening mechanism. After receiving the search term and user query, the system strictly limits the search scope to certified professional medical data sources, effectively shielding the interference of non-professional and media websites, and ensuring the authority and science of knowledge sources.
[0140] The authoritative data sources that the system preferentially connects to include: internationally open access medical literature databases such as PubMed, the official database of the National Medical Products Administration, the national standard drug instruction database, the traditional Chinese medicine database certified by the State Administration of Traditional Chinese Medicine, and the publicly available medical knowledge base of top-three hospitals that have been strictly screened. Through this strict access mechanism, it ensures that all content input into the model is highly professional and reliable.
[0141] To further improve the knowledge quality, the system uses a reranker model to intelligently sort and screen the retrieved content, preferentially retaining the most authoritative and up-to-date professional materials. This dual control mechanism effectively reduces the risk of generating incorrect or misleading answers, significantly improves the credibility and security of the system in the medical field, and provides users with strictly verified high-quality medical information.
[0142] V. Knowledge Organization Process
[0143] To address the issues of insufficient information density and low relevance in traditional Retrieval-Augmented Generation (RAG) methods, the system innovatively introduces an intelligent knowledge processing mechanism based on the secretary model. This mechanism not only considers the literal content of the text but also fully focuses on the continuity of the context and the coherence of the overall semantics, ensuring the scientificity and effectiveness of the knowledge organization process.
[0144] In the specific implementation process, the system first performs multi-dimensional processing on the retrieved reference documents, including basic tasks such as information deduplication, content classification, and credibility assessment. Subsequently, based on the current query and conversation context, the system conducts in-depth knowledge fusion and key extraction, focusing on screening out knowledge fragments that are highly relevant to the user's needs and have a high information density. This refined processing method effectively reduces the interference of irrelevant information and improves the knowledge utilization efficiency.
[0145] Especially when dealing with professional medical knowledge, the system can accurately grasp subtle semantic differences and the meanings of professional terms, ensuring that the organized content is both professionally rigorous and easy to understand through intelligent knowledge refinement. This optimized knowledge processing flow significantly improves the quality and depth of the system's response, providing a more accurate and valuable knowledge foundation for subsequent in-depth analysis.
[0146] VI. Inference Output Mechanism of Natural Language Processing Large Model
[0147] The natural language processing large model is designed based on the decoder of the Transformer architecture and uses an innovative dual-model collaboration mechanism for medical inference. First, the preliminary medical text understanding and analysis are completed through the basic natural language processing large model, which uses word vectorization processing and RMSNorm normalization technology to ensure the standardization of input data and computational stability.
[0148] In the core processing stage, the model achieves in-depth understanding of medical texts through multiple layers of attention mechanisms. Each layer contains a residual connection structure, effectively preventing the problem of gradient disappearance. At the same time, global context information is captured through attention weight calculation to achieve accurate understanding and analysis of complex medical texts. To improve the inference quality, the system introduces the Monte Carlo Tree Search (MCTS) and the PRM planning algorithm for sample synthesis and optimization, and screens the optimal inference path through the rejection sampling mechanism.
[0149] On this basis, the RAG-enhanced natural language processing large model will conduct in-depth analysis and content supplement based on the retrieved professional medical knowledge. The system also establishes a continuous optimization mechanism, which scores and screens the generated results through the Reward Model, and uses the DPO (direct preference optimization) method to align the model, while maintaining professionalism and improving the comprehensibility of the output. This iterative optimization ensures that the system can continuously improve its reasoning ability and output quality.
[0150] This multi-level reasoning mechanism not only ensures the accuracy of basic medical reasoning, but also effectively supplements professional details through knowledge enhancement and continuous optimization, ultimately generating medical answers that are both professional and rigorous and easy to understand, providing reliable intelligent support for clinical decision-making.
[0151] This application provides a medical large model analysis system based on multi-strategy deep slow thinking, including:
[0152] An acquisition unit is used to acquire the query input by the user, deeply analyze the query through context and semantic analysis technology, and generate complete context-related search terms;
[0153] An error correction unit, used to recall relevant content from the knowledge base based on the generated search terms and determine whether error correction is required;
[0154] The first judgment generation unit is used to use a lightweight language model to enter error correction when there are errors or ambiguities in the recalled content, generate a new query after error correction, regenerate new search terms after error correction, perform network search and knowledge base recall based on the newly generated search terms, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user;
[0155] The second judgment generation unit is used to directly perform network search and knowledge base recall when there is no error or ambiguity in the recalled content, organize the recalled knowledge according to the current query, and send the organized knowledge and query to the natural language processing large model for processing, and return the final answer to the user.
[0156] The electronic device provided in the present application may also include: a memory and a processor; the memory is used to store instructions;
[0157] The processor is used to operate according to the instructions to execute the steps of the aforementioned medical large model analysis method based on multi-strategy deep slow thinking.
[0158] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned medical large model analysis method based on multi-strategy deep slow thinking.
[0159] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0160] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0163] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.
Claims
1. A medical large model analysis method based on multi-strategy deep slow thinking, characterized in that Including: Obtain the query input by the user, deeply analyze the query through context and semantic analysis techniques, and generate complete context-related retrieval terms; Recall relevant content from the knowledge base according to the generated retrieval terms and determine whether error correction is required: When the recalled content has errors or is ambiguous, use a lightweight language model to perform error correction, generate a new query after error correction, regenerate new retrieval terms based on the new query after error correction, perform web search and knowledge base recall according to the newly generated retrieval terms, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to a large natural language processing model for processing, and return the final answer to the user; When the recalled content has no errors or is not ambiguous, directly perform web search and knowledge base recall, organize the recalled knowledge according to the current query, and send the organized knowledge and the query to a large natural language processing model for processing, and return the final answer to the user, specifically including: Send the organized knowledge and the query to a large natural language processing model. The large natural language processing model adopts a deep thought chain mechanism, explores the reasoning path through the Monte Carlo tree search algorithm, and its node selection strategy adopts the upper confidence bound formula: In the formula, UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is an exploration constant that controls the balance between exploration and exploitation; Continuously evaluate and verify the reasoning process through a self-dialectical thinking mechanism, and its credibility score adopts the following formula: Confidence = λ1·Knowledge_Support + λ2·Logic_Consistency + λ3·Context_Relevance In the formula, Confidence represents the credibility score of the final answer, λ1, λ2, and λ3 are weight coefficients that respectively measure knowledge support, logical consistency, and context relevance, Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context; Add a classification linear layer to the last layer of the large natural language processing model to score the credibility of the final answer and return the final answer to the user.
2. The medical large model analysis method based on multi-strategy deep slow thinking according to claim 1, wherein The method for deeply analyzing the query through context and semantic analysis techniques to generate complete context-related retrieval terms includes: Through a speech recognition error automatic correction mechanism, combined with context and private medical knowledge base recall to ensure the accuracy of the input information; Combine the multi-round dialogue history queue, and use a dialogue summarizer based on a pure decoder architecture to extract key information and generate complete context-related retrieval terms.
3. The medical large model analysis method based on multi-strategy deep slow thinking according to claim 1, characterized in that The methods for performing web search and knowledge base recall include: Retrieving a private medical knowledge base through the vector database ChromaDB, which integrates screened medical resources; Performing targeted retrieval through a crawler program within a limited range of professional medical websites; Sorting and filtering the retrieved results through a lightweight language model based on a pure decoder architecture; The formula for the relevance score of the lightweight language model is as follows: Score(q,d) = α·BM25(q,d) + β·LLM cls (q,d) Where Score(q, d) represents the relevance score between query q and document d, q represents the query text, d represents the candidate document, α and β are weight coefficients, BM25 represents the relevance score between the query and the document calculated by the retrieval algorithm based on term frequency - inverse document frequency, and LLM cls represents the relevance score output by the classifier using the lightweight language model; The formula for the classification layer is as follows: p = Softmax(W·h + b) In the formula, p is the final relevance probability distribution, h is the hidden state of the last layer of the lightweight language model, W is the parameter matrix of the classification layer, and b is the bias term of the classification layer, representing the probability vector of the relevance between the query and the document.
4. A method for analyzing a large medical model based on multi-strategy deep slow thinking according to claim 1, characterized in that, A method for organizing the recalled knowledge according to the current query, including: A knowledge summary and refinement mechanism based on the current query and context, using a secretary model to organize various recalled knowledge; Performing density enhancement processing on the retrieved knowledge through a sliding window and a key information extraction algorithm, and the formula is as follows: Wherein, Density(T) represents the information density of the text segment, T represents the text segment, and K i represents the i-th key information item, and W i represents the weight corresponding to the i-th key information item, L represents the text length, and n represents the total number of key information items.
5. A medical large model analysis system based on multi-strategy deep slow thinking, characterized in that, Including: An acquisition unit for acquiring the query input by the user, deeply analyzing the query through context and semantic analysis techniques, and generating complete context-related retrieval terms; An error correction unit for recalling relevant content from the knowledge base according to the generated retrieval terms and determining whether error correction is required; A first judgment and generation unit for, when the recalled content is incorrect or ambiguous, using the lightweight language model to enter error correction, generating a new corrected query, generating new retrieval terms based on the corrected new query, performing web search and knowledge base recall according to the newly generated retrieval terms, organizing the recalled knowledge according to the current query, and sending the organized knowledge and the query to a large natural language processing model for processing, and returning the final answer to the user; A second judgment and generation unit for, when the recalled content is not incorrect or ambiguous, directly performing web search and knowledge base recall, organizing the recalled knowledge according to the current query, and sending the organized knowledge and the query to a large natural language processing model for processing, and returning the final answer to the user, specifically including: Sending the organized knowledge and the query to a large natural language processing model, and the large natural language processing model adopts a deep thinking chain mechanism, exploring the reasoning path through the Monte Carlo tree search algorithm, and its node selection strategy adopts the upper confidence bound formula: In the formula, UCT represents the selection score of the node, Q(s, a) represents the average return of taking action a in state s, N(s) represents the number of times state s is visited, N(s, a) represents the number of times action a is taken in state s, and c is an exploration constant that controls the balance between exploration and exploitation; Continuously evaluating and verifying the reasoning process through a self-dialectical thinking mechanism, and its credibility score adopts the following formula: Confidence = λ1·Knowledge_Support + λ2·Logic_Consistency + λ3·Context_Relevance Where Confidence represents the confidence score of the final answer, λ1, λ2, and λ3 are weight coefficients that measure knowledge support, logical consistency, and context relevance respectively, Knowledge_Support represents the degree to which the answer is supported by the knowledge base, Logic_Consistency represents the logical consistency of the reasoning process, and Context_Relevance represents the relevance of the answer to the query context; A classification linear layer is added to the last layer of the large natural language processing model to score the confidence of the final answer and return the final answer to the user.
6. An electronic device, characterized in that, It includes a memory and a processor; The memory is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-4 are implemented.
Citation Information
Patent Citations
Interactive knowledge interaction system based on knowledge base and large language model
CN118013001A
Medical text error correction system and medical query prompt text display method and device
CN118093789A