AI-driven health question-answering method and system oriented to individual health management

By building a chronic disease knowledge database and adopting the LoRA training process and Agent intelligent body design, the problems of insufficient data and limited interaction capabilities of medical AI in the field of chronic diseases have been solved, and more efficient and professional medical Q&A services have been achieved, improving user experience and data quality.

CN120632028APending Publication Date: 2025-09-12SHAANXI OPTO DIGITAL MEDICAL CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510715478.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing medical AI in the field of chronic diseases has problems such as insufficient data, imperfect model training process and limited user interaction capabilities, resulting in insufficient accuracy and flexibility in answers, and unable to meet users' professional and comprehensive medical information needs.

Method used

By building a chronic disease knowledge database, adopting the LoRA method to carry out incremental pre-training, supervised fine-tuning and direct preference optimization training process, designing Agent intelligent bodies, and using the ollama tool to integrate models, it can realize automatic calling and reasoning answers to user questions, and correct answers by combining predefined templates and input prompts.

Benefits of technology

It has improved the quality and utilization efficiency of medical data, optimized the model training process and performance, enhanced the model interaction capabilities and user experience, expanded the application scenarios and functions of medical AI, and provided more efficient, convenient and professional medical and health services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632028A_ABST
    Figure CN120632028A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-driven health question-answering method and system oriented to individual health management, relates to the field of medical AI, is applied to a health question-answering system, and constructs a chronic disease knowledge database according to medical data obtained through screening and crawling; adopting a first training process to obtain a trained base model; designing an Agent agent; obtaining an input question of a user; the Agent agent automatically calls the chronic disease knowledge database according to an input question, and performs networking search and reasoning answering to generate a first answer; and correcting the first answer through the guidance of the predefined template and the input cue word, and outputting the corrected first answer, thereby achieving the technical effects of improving the medical data quality and utilization efficiency, optimizing the model training process and performance, and enhancing the model interaction capability and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical AI technology, and in particular relates to an AI-driven health question-and-answer method and system for individual health management. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, the application of large language models in the medical field is becoming increasingly widespread, driving the intelligent transformation of the medical industry. Models such as SenseTime's "Big Doctor," iFlytek's "iFlytek Spark Medical Big Model," and JD Health's "Jingyi Qianxun" have provided innovative solutions to numerous medical problems, significantly improving diagnostic accuracy and personalized treatment, while also significantly enhancing the overall efficiency of medical services. Aidoc's AIDiagnostic Platform can rapidly identify life-threatening conditions such as cerebral hemorrhage and pulmonary embolism in medical images, thereby improving emergency department response speed and treatment outcomes, and reducing the risk of misdiagnosis and missed diagnoses. Tempus's TempusAI model, launched by Tempus, focuses on in-depth analysis of genomic and clinical data to provide personalized treatment plans for cancer patients. Based on a patient's genomic information, medical history, and other relevant data, it provides doctors with precise treatment recommendations, optimizing the patient's treatment path and improving treatment outcomes. These models have shown great potential in practical applications. For example, when dealing with complex medical problems, they can have natural and fluent conversations with patients and provide professional and personalized health guidance, thereby greatly improving the convenience and accessibility of medical services and enabling patients to obtain effective medical advice more promptly.

[0003] However, despite the progress made in current medical AI technology, it still faces some urgent problems and challenges. First, the lack of medical data on chronic diseases is a prominent issue. The existing amount of data is difficult to meet the needs of model training, which greatly limits the further improvement of model performance. Secondly, the model training process is still imperfect, making it difficult to effectively inject domain knowledge into the model and precisely control its behavior. This may lead to suboptimal performance of the model in certain specific medical scenarios. In addition, the interaction ability between the model and the user needs to be strengthened. The current model still has shortcomings in understanding user intent and generating answers that meet user needs, which affects the user experience.

[0004] Traditional medical Q&A technology suffers from significant limitations in several key areas, including data sources, processing capabilities, model training and optimization, application effectiveness, and user experience. These limitations, like numerous shackles, severely constrain the performance and intelligent development of medical Q&A systems, significantly limiting their application value and potential for widespread adoption in real-world healthcare scenarios. To overcome these bottlenecks, technological innovation and optimization are urgently needed to enhance the performance and intelligence of medical Q&A systems across multiple dimensions, meet the growing demand for high-quality medical Q&A, and provide stronger support for the development of medical informatization and public health services.

[0005] At present, the existing technology still has the following shortcomings:

[0006] Limitations of Data Sources: Traditional medical question-and-answer technology has significant shortcomings in data acquisition, relying primarily on the construction of domain-specific question-and-answer databases. The construction of such databases is typically limited to a relatively fixed corpus and relevant knowledge within that domain. For example, in the medical field, the data the system relies on may only cover basic information such as symptom descriptions of common diseases and conventional treatment plans. Such a single data source sets narrow boundaries for the model's cognitive field of view. When faced with complex and ever-changing medical issues, especially those involving rare conditions, multi-system cross-cutting diseases, or personalized treatment plan consultations, the model struggles to provide comprehensive and accurate answers due to its inability to obtain sufficiently rich and diverse information. This limitation is reflected not only in the breadth of information but also in its depth, limiting the model's ability to understand and answer medical questions. This often leaves the system struggling to cope with complex medical scenarios and unable to meet users' urgent needs for professional and comprehensive medical information.

[0007] Insufficient data processing capabilities: Delving deeper into the data processing phase, traditional methods reveal their shortcomings in processing power. When processing data, they often only perform simple semantic analysis and keyword extraction on the text. This approach is like digging only the tip of the iceberg, failing to delve deeper into the underlying relationships and knowledge structures within the data. The medical field contains a vast and complex network of knowledge, with intricate connections between diseases and symptoms, and between test results and treatment plans. Traditional processing methods struggle to gain insight into these deep relationships, making it impossible for models to accurately grasp the core points of complex medical problems. This, in turn, significantly reduces the accuracy and completeness of answers, making it difficult to accurately provide users with practical and actionable medical advice.

[0008] Deficiencies in model training and optimization: Problems remain prominent in the dimensions of model training and optimization. In terms of the training process, traditional model training typically adopts a relatively fixed model, or only performs pre-training or simple fine-tuning, lacking refined control and in-depth optimization of the model training process. This single training process is like using a unified template to shape models with different characteristics. It is difficult to fully tap the potential of the model in specific medical scenarios, resulting in the model's performance in the professional medical field being difficult to achieve ideal conditions. In terms of model adaptability, traditional models are often only optimized for specific tasks during the training phase and lack the ability to adapt to diverse scenarios and tasks. Once encountering new problem types or changes in application scenarios, the model's performance will be like losing precise navigation, getting lost in unfamiliar areas, and its accuracy and reliability will be significantly reduced, making it impossible to quickly adjust to meet new challenges.

[0009] Deficiencies in application effectiveness and user experience: From the perspective of application effectiveness and user experience, the shortcomings of traditional medical question-and-answer technology cannot be ignored. In terms of answer quality, due to dual limitations of data and models, the answers generated by the system are often overly simplistic and general, barely scratching the surface and failing to meet users' urgent need for detailed, professional medical information. Furthermore, the accuracy and reliability of the answers lack firm guarantees, which is particularly critical in the medical field, where erroneous or misleading information can directly and adversely affect patients' health. In terms of interactive capabilities, traditional medical question-and-answer systems are quite rigid. They typically only respond based on preset templates and fixed rules, and are unable to dynamically adjust and reason based on real-time user feedback and contextual information. This lack of flexibility and intelligence in the interactive model makes it difficult for the system to establish in-depth and effective communication with users, failing to truly understand their evolving needs and, consequently, unable to provide precise services that meet their actual needs, significantly impacting the user experience and system practicality. Summary of the Invention

[0010] The embodiments of the present application provide an AI-driven health question-and-answer method and system for individual health management, thereby solving the technical problems in the prior art of medical AI in the field of chronic diseases, such as insufficient data, imperfect model training process, and limited user interaction capabilities. This achieves the technical effects of improving the quality and utilization efficiency of medical data, optimizing the model training process and performance, enhancing model interaction capabilities and user experience, and expanding the application scenarios and functions of medical AI.

[0011] In a first aspect, an embodiment of the present invention provides an AI-driven health question-and-answer method for individual health management, which is applied to a health question-and-answer system, and the method includes: step 1: constructing a chronic disease knowledge database based on medical data obtained by screening and crawling; step 2: adopting a first training process to obtain a trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the open source model in sequence based on the LoRA method; step 3: integrating the trained base model into a local multi-workflow according to the ollama tool, and designing an Agent intelligent body; step 4: obtaining the user's input question; step 5: the Agent intelligent body automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning and answering to generate a first answer; step 6: correcting the first answer through the guidance of a predefined template and input prompt words, and outputting the corrected first answer.

[0012] In a second aspect, an embodiment of the present invention further provides an AI-driven health question-and-answer system for individual health management, the system comprising: a first construction unit configured to construct a chronic disease knowledge database based on medical data obtained through screening and crawling;

[0013] A first execution unit, configured to obtain a trained base model using a first training process, wherein the first training process sequentially performs incremental pre-training, supervised fine-tuning, and direct preference optimization on the open source model based on the LoRA method;

[0014] A first design unit, configured to integrate the trained base model into a local multi-workflow according to an ollama tool and design an Agent intelligent body;

[0015] a first obtaining unit, configured to obtain a question input by a user;

[0016] A second execution unit, wherein the second execution unit is used for the Agent intelligent body to automatically call the chronic disease knowledge database according to the input question, search online and perform reasoning to answer, and generate a first answer;

[0017] The third execution unit is used to correct the first answer through the guidance of a predefined template and an input prompt word, and output the corrected first answer.

[0018] In a third aspect, an embodiment of the present invention further provides an AI-driven health question-and-answer system for individual health management, comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor, characterized in that when the processor executes the program, the steps of the method of the first aspect are implemented.

[0019] The above one or more technical solutions in the embodiments of the present invention have at least one or more of the following technical effects:

[0020] An AI-driven health question-and-answer method and system for individual health management provided in an embodiment of the present invention is applied to a health question-and-answer system, through step 1: building a chronic disease knowledge database based on medical data obtained by screening and crawling; step 2: adopting a first training process to obtain a trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the open source model in sequence based on the LoRA method; step 3: integrating the trained base model into a local multi-workflow according to the ollama tool, and designing an Agent intelligent body; step 4: obtaining the user's input question; step 5: the Agent intelligent body automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning and answering to generate a first answer; step 6: correcting the first answer through the guidance of a predefined template and input prompt words, and converting the corrected first answer into the corrected second answer. An answer is output. Therefore, the present invention constructs a chronic disease knowledge database by screening and crawling a large amount of medical data to provide rich data support for model training; innovates the model training process, adopts PT+SFT+DPO training process, and combines the lora method to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the base model to achieve domain knowledge injection and precise control of model behavior; designs an Agent intelligent body to automatically call the database, search online and perform reasoning answers according to user questions, and guides the model to generate output that meets the requirements through predefined templates and input prompt words, thereby solving the technical problems of insufficient data, imperfect model training process and limited user interaction ability in the field of chronic diseases in the existing medical AI technology, and achieving the technical effects of improving the quality and utilization efficiency of medical data, optimizing the model training process and performance, enhancing the model interaction ability and user experience, and expanding the application scenarios and functions of medical AI.

[0021] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of a flow chart of an AI-driven health question-and-answer method for individual health management in an embodiment of the present invention;

[0023] Figure 2 This is a LoRA principle diagram of an AI-driven health question-and-answer method for individual health management in an embodiment of the present invention;

[0024] Figure 3 This is a Transformer schematic diagram of an AI-driven health question-answering method for individual health management in an embodiment of the present invention;

[0025] Figure 4 Schematic diagram of a query workflow based on a chronic disease knowledge base for an AI-driven health question-and-answer method for individual health management in an embodiment of the present invention;

[0026] Figure 5 This is a schematic diagram of a user physical examination report analysis workflow of an AI-driven health question-and-answer method for individual health management in an embodiment of the present invention;

[0027] Figure 6 Schematic diagram of an online query and answer workflow of an AI-driven health question-and-answer method for individual health management in an embodiment of the present invention;

[0028] Figure 7 This is a schematic diagram of the structure of an AI-driven health question-and-answer system for individual health management in an embodiment of the present invention;

[0029] Figure 8 FIG. 4 is a schematic structural diagram of another exemplary electronic device according to an embodiment of the present invention.

[0030] Description of reference numerals: receiver 301 , processor 302 , transmitter 303 , memory 304 , bus interface 305 . DETAILED DESCRIPTION

[0031] The embodiments of the present application provide an AI-driven health question-and-answer method and system for individual health management to address the technical problems in the prior art of medical AI in the field of chronic diseases, such as insufficient data, imperfect model training process, and limited user interaction capabilities.

[0032] The technical solution in the embodiment of the present invention has the following general ideas:

[0033] An embodiment of the present invention provides an AI-driven health question-and-answer method and system for individual health management, which is applied to a health question-and-answer system. The method includes: step 1: constructing a chronic disease knowledge database based on medical data obtained by screening and crawling; step 2: adopting a first training process to obtain a trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the open source model in sequence based on the LoRA method; step 3: integrating the trained base model into a local multi-workflow according to the ollama tool, and designing an Agent intelligent body; step 4: obtaining a user's input question; step 5: the Agent intelligent body automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning and answering, and generates a first answer; step 6: correcting the first answer through the guidance of a predefined template and input prompt words, and outputting the corrected first answer, thereby achieving the technical effects of improving the quality and utilization efficiency of medical data, optimizing the model training process and performance, enhancing the model interaction capability and user experience, and expanding the application scenarios and functions of medical AI.

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0035] Example 1

[0036] This embodiment provides an AI-driven health question-and-answer method for individual health management, which is applied to a health question-and-answer system. The method specifically includes:

[0037] Step 1: Build a chronic disease knowledge database based on the medical data obtained through screening and crawling.

[0038] Specifically, the present invention constructs a chronic disease knowledge database by screening and crawling a large amount of medical data to achieve the purpose of providing rich data support for model training. For example, 330,000 question and answer data covering multiple departments can be screened to build a chronic disease knowledge base, and 2.4 million Chinese medical data sets can be crawled to provide rich and comprehensive data support for model training, ensuring that the question and answer system has a deep knowledge reserve in the field of chronic diseases and can provide users with accurate and professional chronic disease medical advice. Therefore, in response to the problem of insufficient medical data for chronic diseases, the present invention screened 330,000 question and answer medical data covering surgery, internal medicine and pediatrics, constructed the Chronic-RAG chronic disease knowledge database, and crawled and downloaded 2.4 million Chinese medical data sets, including pre-training, instruction fine-tuning and reward data sets.

[0039] Furthermore, the AI-driven health Q&A method for individual health management in this embodiment is applied to a health Q&A system. The system architecture design includes: the Q&A system adopts a layered architecture, divided into a user interaction layer, a business logic layer, and a data processing layer. The user interaction layer is responsible for receiving questions input by users and presenting answers in a natural and friendly manner; the business logic layer is responsible for coordinating the work of various modules of the intelligent agent, and rationally calling relevant tools and resources based on the type and complexity of the user's questions; the data processing layer is responsible for preprocessing user questions, extracting key information, and formatting and optimizing the answers generated by the intelligent agent to ensure the accuracy and readability of the answers.

[0040] Step 2: Use the first training process to obtain the trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the open source model based on the LoRA method.

[0041] Furthermore, it specifically includes: adopting a continuous pre-training method to perform incremental pre-training on domain document data; adopting a supervised fine-tuning method to use labeled data to accurately adjust the open source model, construct an instruction fine-tuning dataset, perform instruction fine-tuning on the open source model, and in the fine-tuning process, integrate the knowledge of the first domain into the open source model; adopting a direct preference optimization method to adjust the behavior of the open source model according to human preferences.

[0042] Furthermore, in step 2, the training process parameter optimization object of the first training process is a Transformer network structure, and the Transformer network structure is composed of multiple encoders and decoders.

[0043] Specifically, this embodiment uses a first training process to obtain a trained base model. This first training process sequentially performs incremental pre-training, supervised fine-tuning, and direct preference optimization on the open source model based on the LoRA method, thereby obtaining the trained base model. Specifically, by constructing a PT+SFT+DPO training process, the open source model is sequentially subjected to incremental pre-training, supervised fine-tuning, and direct preference optimization based on the LoRA method. This process effectively infuses domain knowledge, directly optimizes the language model, precisely controls its behavior, and effectively learns human preferences, making the model's responses more accurate and more aligned with actual user needs.

[0044] The LoRA (Low-Rank Adaptation) technology is proposed based on the conclusion that the weight update of the neural network usually has a low intrinsic dimension. Its principle is as follows Figure 2 As shown. For the pre-trained weight matrix, by The latter is represented by a low-rank decomposition W0+ΔW=W0+BA to constrain its update, where and rank r << min(d,k). During training, W0 is frozen and receives no gradient updates, while A and B contain trainable parameters. Note that both W0 and ΔW = BA are multiplied by the same input, and their respective output vectors are summed in the coordinate direction. For h = W0x, the modified forward pass yields the following.

[0045] h=W0x+ΔWx=W0x+BAx.

[0046] LoRA freezes the pre-trained model weights and injects a trainable rank factorization matrix into each layer of the Transformer architecture, significantly reducing the number of trainable parameters for downstream tasks. Compared to GPT-3175B fine-tuned with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and GPU memory requirements by 3 times. LoRA performs comparable or better than fine-tuning in terms of model quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, despite having fewer trainable parameters, higher training throughput, and no additional inference latency unlike the adapter.

[0047] The PT+SFT+DPO training process specifically includes: PT (Continuous Pretraining) performs incremental pretraining on massive domain document data to adapt the AI ​​model to the domain data distribution. After the LLM is expanded with Transformer blocks, only the newly added blocks are trained during the incremental pretraining process, effectively injecting model knowledge and greatly avoiding catastrophic forgetting. SFT (Supervised Fine-Tuning) uses annotated data to precisely adjust the model to specific medical tasks. An instruction fine-tuning dataset is constructed, and instruction fine-tuning is performed based on the pretrained model to align instruction intent. During the fine-tuning process, knowledge from the first domain (specific domain) is incorporated into the model, enabling the model to better adapt to the application needs of specific industries and enhancing its professionalism and practicality in that field. DPO (Direct Preference Optimization) achieves precise control of the language model's behavior by directly optimizing it, without the need for complex reinforcement learning. It can also effectively learn human preferences. Compared with RLHF, DPO is easier to implement and train, and has better results. The parameter optimization object of the PT+SFT+DPO training process is mainly the Transformer network structure. The Transformer network structure is as follows Figure 3 The Transformer network structure consists of multiple encoders and decoders. The input text is encoded according to the word2vec algorithm and then added to the positional encoding to form a sentence vector.

[0048]

[0049] Among them, pos is the position of the word, d model is the dimension of the position vector, which is equal to the dimension of the word encoding. model ] represents the i-th dimension of the position vector. The vector at the pos-th position can be obtained using the above formula. The encoder primarily consists of a multi-head attention layer, residual connections, and layer normalization. The sentence vectors are processed through the multi-head attention layer to capture interrelated information within the text. Residual connections and layer normalization alleviate the problem of model gradients becoming increasingly small and vanishing during propagation as the network deepens, leading to slow model parameter updates.

[0050] MultiHead(Q,K,V)=Concat(head1,...,head h )W O

[0051]

[0052] Where projection is the parameter matrix The encoder typically has multiple layers. The first encoder input is a text sequence, and the final encoder output is a set of sequence vectors. This set of sequence vectors serves as the decoder's K and V inputs, where K, V, and the decoder's output sequence vectors are equal. These attention vectors are input to the encoder-decoder attention layer of each decoder, helping the decoder focus on the appropriate locations in the input sequence.

[0053] Step 3: Integrate the trained base model into the local multi-workflow according to the ollama tool and design the Agent intelligent body;

[0054] Furthermore, the agent includes a context memory module, which stores historical conversation records and determines whether similar historical conversations have occurred before responding; a tool list module, which uses the tool list module to set a personal information interface and, in conjunction with the agent, analyzes the user's situation and extracts key information; and an execution module, which sets priorities based on key knowledge of the problem and coordinates the invocation of different tool interfaces to optimize and resolve the problem.

[0055] Furthermore, in step 3, the trained base model is integrated into the local multi-workflow according to the ollama tool, and the Agent intelligent body is designed, including: the trained base model is in the model.safetensors format, and the model.safetensors format is converted into the gguf format based on the llama.cpp tool to complete the format conversion; the converted base model is imported into the ollama model integration platform; according to the business scenario, the large model is introduced through ollama to optimize the corresponding business process.

[0056] Specifically, further, based on the ollama tool, the model is integrated into the local multi-workflow and the Agent intelligent body is designed, that is, the trained base model is integrated into the local multi-workflow according to the ollama tool, and the Agent intelligent body is designed.

[0057] Furthermore, an agent was developed that can automatically call databases, conduct online searches, and perform reasoning and answering based on user questions. Guided by predefined templates and input prompts, the model can generate outputs that better meet user needs. This interactive method not only improves the communication efficiency between users and the model, but also enables patients to obtain medical information more conveniently, enhancing the user experience. The developed agent includes the following modules: a context memory module, which is used to store and remember historical conversation records and determine whether similar historical conversations have occurred before answering; a tool list module, which is used to set personal information interfaces to help the agent quickly analyze the user's specific situation and extract key information; and an execution module, which is used to set priorities based on the key knowledge of the problem and is responsible for coordinating the calls of different tool interfaces to optimize and process the problem.

[0058] The specific design process is as follows: the trained model is in the model.safetensors format, and the format conversion is completed based on the llama.cpp tool, and the model.safetensors format is converted into the gguf format; the converted AI model is imported into the ollama model integration platform; according to the business scenario, the large model is introduced through ollama to optimize the corresponding business process. The present invention integrates the trained domain AI model into three business workflows, namely, the chronic disease knowledge base query workflow, the user physical examination report analysis workflow, and the online query answer workflow. Figure 4 、 Figure 5 、 Figure 6 shown.

[0059] Step 4: Get the user's input question;

[0060] Furthermore, in step 4, after obtaining the user's input question, it also includes: performing text preprocessing on the input question, wherein the preprocessing includes removing stop words and extracting stems; using natural language processing technology to perform semantic analysis on the input question to obtain the user's needs, wherein the health question and answer system interacts with the user and provides three forms of interaction including voice, pictures, and text; based on the semantic information of the input question, the health question and answer system determines whether the question belongs to the medical field. If so, it continues to determine the medical sub-field of the input question, and selects the medical sub-field intelligent agent corresponding to the input question, and sends the input question to the Agent intelligent agent for processing.

[0061] Specifically, after the user enters a question, the system will first perform text preprocessing on the input question, including removing stop words, stemming and other operations to simplify the question and highlight key information; then, it will use natural language processing technology to perform semantic analysis on the question to understand the user's true intentions and needs; then, based on the semantic information of the question, the health question-answering system will determine whether the question belongs to the medical field. If so, it will further determine the specific medical sub-field involved in the question, such as chronic diseases, interpretation of physical examination reports, etc., and then pass the question to the corresponding intelligent agent module for processing.

[0062] Step 5: The Agent automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning to answer, and generates a first answer.

[0063] Specifically, after receiving the user's input question, the Agent can automatically call the database, search online and infer the answer based on the user's question, and then generate the first answer.

[0064] Step 6: Correct the first answer by following the guidance of the predefined template and input prompt words, and output the corrected first answer.

[0065] Furthermore, in step 6, the first answer is corrected and the corrected answer is output through the guidance of a predefined template and input prompt words, which also includes: the health question and answer system optimizes the first answer, wherein the optimization process includes checking and correcting the logic, accuracy and completeness of the first answer to ensure that the first answer can answer the user's question, wherein the health question and answer system personalizes the first answer based on the user's historical question records and feedback information.

[0066] Specifically, predefined templates and reasonably accurate input prompts are used to enable the agent to replace the variable part in the template, and by designing reasonable and accurate input prompts, the model is guided to generate outputs that meet the needs, making the answers more flexible and accurate, and improving the intelligence and adaptability of the system. That is, through the guidance of predefined templates and input prompts, the first answer is corrected, and the corrected first answer is output. The specific answer generation and optimization process includes: after the agent generates a preliminary answer based on the received question and related tools, the question and answer system will optimize the answer. The optimization process includes checking and correcting the logic, accuracy and completeness of the answer to ensure that the answer can clearly and accurately answer the user's question. At the same time, the health question and answer system will also personalize the answer based on the user's historical question records and feedback information to better meet the user's specific needs.

[0067] Furthermore, in this embodiment, the system integration and testing also includes: integrating the intelligent agent with databases and search engines to ensure seamless collaboration and smooth data flow between the various parts. After the integration is completed, a comprehensive system test is carried out, including functional testing, performance testing, and user experience testing, to discover and fix potential problems and ensure the stability and reliability of the question-answering system. By simulating different types of user questions and usage scenarios, verifying the performance of the system in various situations, and continuously optimizing the system's response speed, accuracy, and interactive experience to achieve optimal user satisfaction, a question-answering system has been developed that supports similar question recommendations, question guidance, voice interaction, and other functions, thereby improving the interactive experience and efficiency between users and models.

[0068] Therefore, this invention effectively solves many problems faced by existing medical AI in the field of chronic diseases through innovations in data processing, model training, interactive design and function development, providing patients with more efficient, convenient and professional medical and health services, and has significant technical and social value.

[0069] Traditional medical Q&A assistants have obvious deficiencies in their technical architecture. First, they rely only on a limited Q&A database with a small amount of data and focus on common diseases. There is a lack of data on chronic diseases, making it difficult to meet the precise Q&A needs of patients with chronic diseases. Second, the training process is single, relying only on pre-training and fine-tuning, lacking direct optimization for the medical field, insufficient injection of domain knowledge, and unable to effectively learn human preferences, resulting in poor accuracy and pertinence of answers. Third, the interactive experience is poor, relying only on keyword matching and templated answers, unable to flexibly handle complex and changeable user questions, lacking functions such as reasoning and online search, and a single interactive method that cannot meet the diverse needs of users. In response to these deficiencies in the existing technologies, the present invention aims to overcome the technical problems of current medical AI in the field of chronic diseases, such as insufficient data, imperfect model training process, and limited user interaction capabilities, and to provide a more efficient, accurate, and convenient solution for the field of medical AI to better meet the needs of chronic disease patients and medical practitioners, and promote the further development and application of medical AI technology.

[0070] This invention builds a chronic disease knowledge database by screening and crawling a large amount of medical data, providing rich data support for model training; innovates the model training process, adopts the PT+SFT+DPO training process, and combines the lora method to perform incremental pre-training, supervised fine-tuning and direct preference optimization of the base model, realizing domain knowledge injection and precise control of model behavior; designs an Agent intelligent body, which automatically calls the database, searches the Internet and performs reasoning answers according to user questions, and guides the model to generate output that meets the requirements through predefined templates and input prompt words, thereby effectively improving the application effect and user experience of medical AI in the field of chronic diseases.

[0071] The beneficial effects brought about by the present invention are as follows:

[0072] First, to improve the quality and utilization efficiency of medical data: This invention constructs the Chronic-RAG chronic disease knowledge database by screening 330,000 medical question-and-answer data covering surgery, internal medicine, and pediatrics, and crawling 2.4 million Chinese medical data sets, providing rich and comprehensive data support for model training. This not only addresses the problem of insufficient chronic disease medical data, but also ensures data diversity and accuracy, enabling the model to learn more valuable information during training, thereby improving its performance and reliability.

[0073] Secondly, to optimize the model training process and performance: the constructed PT+SFT+DPO training process, combined with the lora method, performs incremental pre-training, supervised fine-tuning, and direct preference optimization on the base model Qwen2.5. This effectively injects domain knowledge and precisely controls the model's behavior through direct optimization of the language model. This innovative training process enables the model to better learn human preferences, improves the model's diagnostic accuracy and answer quality in the field of chronic diseases, and provides patients with more professional and precise medical advice.

[0074] Next, to enhance model interaction and user experience, we integrated the model into multiple local workflows using the Ollama tool and designed an agent that automatically accesses databases, conducts online searches, and generates reasoned responses based on user questions. Guided by predefined templates and input prompts, the model generates outputs that better meet user needs. This interactive approach not only improves communication efficiency between users and the model but also enables patients to more conveniently access medical information, enhancing the user experience.

[0075] Furthermore, the company is expanding the application scenarios and functionality of medical AI. The developed question-and-answer system supports features such as similar question recommendations, question guidance, and voice interaction, enriching the application scenarios of medical AI. These features can help patients express their questions more accurately and obtain more comprehensive medical advice. They also provide patients with more convenient and personalized medical services, especially for chronic disease management, better meeting their daily health management and disease diagnosis and treatment needs.

[0076] Finally, in terms of promoting the development and innovation of medical AI technology: the invention's innovations in chronic disease medical data processing, model training process optimization, interactive capability enhancement, and application scenario expansion provide new technical ideas and solutions for the field of medical AI. These innovations not only help improve the performance and application effects of current medical AI systems, but may also have a positive driving effect on the development of the entire medical AI industry, promoting the exploration and application of more related technologies. Finally, the developed question-and-answer system supports similar question recommendations, question guidance, voice interaction and other functions, which greatly improves the interactive experience and efficiency between users and models, allowing patients to obtain medical information more conveniently, and enhancing the practicality and user-friendliness of the system.

[0077] The following are alternatives to product shapes, configurations, and method execution steps of the present invention that may be substituted by other means.

[0078] First: replacement of product shape and structure

[0079] (1) Hardware replacement of health question-answering system

[0080] The health Q&A system can be integrated into hospital service platforms. It accesses users' historical medical records through the hospital's backend database, analyzes their historical data, and provides timely health status alerts and recommendations. The system can answer users' questions about chronic diseases based on a backend knowledge base or via the internet. The system can be loaded onto hospital service kiosks, allowing patients to communicate with the system through voice interaction or touchscreens, directly receiving health management advice.

[0081] (2) System structure replacement

[0082] The knowledge base storage method of the health question-and-answer system is replaced by a distributed database instead of a single centralized database. Using cloud computing or edge computing architecture can better achieve real-time data synchronization and expansion, while also improving the system's fault tolerance and availability.

[0083] Second: device replacement

[0084] (1) AI model adjustment strategy

[0085] When selecting AI models, fine-tuning can be performed on a variety of pre-trained models, such as natural language processing models like GPT4, BERT, and T5, which have demonstrated outstanding performance in dialogue generation and semantic understanding. For the sake of user data security, the selected AI model should be deployed locally whenever possible. Therefore, when making the selection, performance requirements and computing resources should be comprehensively considered and appropriate adjustments made. Alternatively, integrating a traditional rules engine with machine learning is another option. Rules engines are effective for addressing common, fixed-pattern problems, while machine learning is suitable for handling complex, dynamic problems. The synergy between the two helps improve system response speed and accuracy.

[0086] (2) Knowledge base expansion method

[0087] To ensure the timeliness and advancement of the knowledge base, text mining techniques can be used to continuously extract disease-related updates from medical literature, health information websites, and patient feedback. At the same time, a "crowdsourcing" model can be used to encourage users and experts to participate in supplementing and updating answers to frequently asked questions. Leveraging the collective wisdom of the community, this will accelerate the expansion and improvement of the knowledge base, particularly in the areas of special and rare diseases.

[0088] Third: Replacement of method steps

[0089] (1) Alternatives to data collection and preprocessing

[0090] During the data collection phase, non-traditional health tracking devices can be used. Smart wearable devices (bracelets, smart watches, etc.) can be used to directly monitor the user's physiological data (such as heart rate, blood sugar, blood pressure, etc.) in real time. These devices can be synchronized with the health assistant system in real time to further improve the accuracy and timeliness of the data. On the other hand, speech recognition technology can be used to convert the patient's voice description into text, and then combined with natural language processing technology for analysis. This method is particularly suitable for the elderly or people with limited mobility, which can reduce their input burden and improve the accessibility of the system.

[0091] (2) Model update

[0092] After signing a data service agreement with a user, the system collects user questions, processes them into a new dataset, and uses incremental learning to update the model's domain data. This allows the system to continuously learn and optimize the model's answering capabilities as users interact with the health assistant, eliminating the need for retraining each time. Incremental learning effectively responds to new situations and changes in personalized needs.

[0093] Example 2

[0094] Based on the same inventive concept as the AI-driven health question-answering method for individual health management in the aforementioned embodiment, the present invention also provides an AI-driven health question-answering system for individual health management, such as Figure 7 As shown, wherein the system includes:

[0095] A first construction unit 11 is used to construct a chronic disease knowledge database based on the medical data obtained by screening and crawling;

[0096] A first execution unit 12 is configured to obtain a trained base model by adopting a first training process, wherein the first training process is to sequentially perform incremental pre-training, supervised fine-tuning, and direct preference optimization on the open source model based on the LoRA method;

[0097] A first design unit 13 is used to integrate the trained base model into a local multi-workflow according to the ollama tool and design an Agent intelligent body;

[0098] A first obtaining unit 14, the first obtaining unit 14 is used to obtain a question input by a user;

[0099] The second execution unit 15 is used for the Agent to automatically call the chronic disease knowledge database according to the input question, search online and perform reasoning to answer, and generate a first answer;

[0100] The third execution unit 16 is configured to correct the first answer by following the guidance of a predefined template and an input prompt word, and output the corrected first answer.

[0101] Furthermore, the first execution unit 12 further includes:

[0102] A first training unit, configured to perform incremental pre-training on domain document data using a continuous pre-training approach;

[0103] a second training unit configured to employ a supervised fine-tuning approach to precisely adjust the open source model using the labeled data, construct an instruction fine-tuning dataset, perform instruction fine-tuning on the open source model, and incorporate knowledge from the first domain into the open source model during the fine-tuning process;

[0104] The third training unit is used to adjust the behavior of the open source model according to human preferences by adopting a direct preference optimization method.

[0105] Furthermore, the first execution unit 12 also includes that the training process parameter optimization object of the first training process is a Transformer network structure, and the Transformer network structure is composed of multiple encoders and decoders.

[0106] Furthermore, the Agent of the first design unit 13 includes:

[0107] A context memory module stores historical conversation records and determines whether similar historical conversations occur before answering based on the historical conversation records;

[0108] A tool list module, which sets a personal information interface through the tool list module, analyzes the user situation in combination with the Agent intelligent body and extracts key information;

[0109] The execution module sets priorities based on key knowledge of the problem and is responsible for coordinating the calls of different tool interfaces to optimize and process the problem.

[0110] Furthermore, the first design unit 13 further includes:

[0111] A first conversion unit, wherein the first conversion unit is used to convert the trained base model into a model.safetensors format based on the llama.cpp tool to complete the format conversion;

[0112] A first importing unit, configured to import the converted base model into an Ollama model integration platform;

[0113] The first introduction unit is used to introduce a large model through ollama according to the business scenario to optimize the corresponding business process.

[0114] Furthermore, the first obtaining unit 14 further includes:

[0115] A first preprocessing unit, configured to perform text preprocessing on the input question, wherein the preprocessing includes removing stop words and extracting stems;

[0116] a second obtaining unit, configured to perform semantic analysis on the input question using natural language processing technology to obtain the user's needs, wherein the health question-answering system interacts with the user and provides three forms of interaction including voice, picture, and text;

[0117] The fourth execution unit is used to determine whether the health question and answer system belongs to the medical field based on the semantic information of the input question. If so, the health question and answer system continues to determine the medical sub-field of the input question, selects the medical sub-field agent corresponding to the input question, and sends the input question to the Agent agent for processing.

[0118] Furthermore, the third execution unit 16 further includes:

[0119] A fifth execution unit, wherein the fifth execution unit is used by the health question and answer system to optimize the first answer, wherein the optimization process includes checking and correcting the logic, accuracy and completeness of the first answer to ensure that the first answer can answer the user's question, wherein the health question and answer system personalizes the first answer based on the user's historical question records and feedback information.

[0120] Furthermore, the system includes:

[0121] The sixth execution unit is used to integrate the Agent intelligent body, the chronic disease knowledge database and the search engine into the health question-and-answer system, and perform system testing after the integration is completed, wherein the system testing includes functional testing, performance testing and user experience testing.

[0122] The foregoing Figure 1 The various variations and specific examples of the AI-driven health question-and-answer method for individual health management in Example 1 are also applicable to the AI-driven health question-and-answer system for individual health management in this embodiment. Through the above detailed description of the AI-driven health question-and-answer method for individual health management, those skilled in the art can clearly understand the implementation method of the AI-driven health question-and-answer system for individual health management in this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.

[0123] Example 3

[0124] Based on the same inventive concept as the AI-driven health question-answering method for individual health management in the aforementioned embodiment, the present invention also provides an exemplary electronic device, such as Figure 8 As shown, it includes a memory 304, a processor 302, and a computer program stored in the memory 304 and executable on the processor 302. When the processor 302 executes the program, it implements the steps of any of the aforementioned AI-driven health question-and-answer methods for individual health management.

[0125] Among them, Figure 8In the embodiment of the present invention, a bus architecture (represented by bus 300) is shown. Bus 300 may include any number of interconnected buses and bridges, and bus 300 links together various circuits including one or more processors represented by processor 302 and memory represented by memory 304. Bus 300 may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same component, namely a transceiver, which provides a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 may be used to store data used by processor 302 when performing operations.

[0126] The above one or more technical solutions in the embodiments of the present invention have at least one or more of the following technical effects:

[0127] An AI-driven health question-and-answer method and system for individual health management provided in an embodiment of the present invention is applied to a health question-and-answer system, through step 1: building a chronic disease knowledge database based on medical data obtained by screening and crawling; step 2: adopting a first training process to obtain a trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the open source model in sequence based on the LoRA method; step 3: integrating the trained base model into a local multi-workflow according to the ollama tool, and designing an Agent intelligent body; step 4: obtaining the user's input question; step 5: the Agent intelligent body automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning and answering to generate a first answer; step 6: correcting the first answer through the guidance of a predefined template and input prompt words, and converting the corrected first answer into the corrected second answer. An answer is output. Therefore, the present invention constructs a chronic disease knowledge database by screening and crawling a large amount of medical data to provide rich data support for model training; innovates the model training process, adopts PT+SFT+DPO training process, and combines the lora method to perform incremental pre-training, supervised fine-tuning and direct preference optimization on the base model to achieve domain knowledge injection and precise control of model behavior; designs an Agent intelligent body to automatically call the database, search online and perform reasoning answers according to user questions, and guides the model to generate output that meets the requirements through predefined templates and input prompt words, thereby solving the technical problems of insufficient data, imperfect model training process and limited user interaction ability in the field of chronic diseases in the existing medical AI technology, and achieving the technical effects of improving the quality and utilization efficiency of medical data, optimizing the model training process and performance, enhancing the model interaction ability and user experience, and expanding the application scenarios and functions of medical AI.

[0128] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0129] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The present invention is a block or multiple blocks that specify the functions of the steps. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention is intended to include such modifications and variations.

Claims

1. An AI-driven health question-answering method for individual health management, applied to a health question-answering system, characterized in that: The method comprises: Step 1: Build a chronic disease knowledge database based on the medical data obtained through screening and crawling; Step 2: Use the first training process to obtain a trained base model, wherein the first training process is to perform incremental pre-training, supervised fine-tuning, and direct preference optimization on the open source model based on the LoRA method; Step 3: Integrate the trained base model into the local multi-workflow according to the ollama tool and design the Agent intelligent body; Step 4: Get the user's input question; Step 5: The Agent automatically calls the chronic disease knowledge database according to the input question, searches online and performs reasoning to answer, and generates a first answer; Step 6: Correct the first answer by following the guidance of the predefined template and input prompt words, and output the corrected first answer.

2. The AI-driven health question-answering method for individual health management according to claim 1, characterized in that: In step 2, the first training process is adopted to obtain the trained base model, which specifically includes: Use continuous pre-training to perform incremental pre-training on domain document data; Using a supervised fine-tuning approach, the open source model is precisely adjusted using labeled data, an instruction fine-tuning dataset is constructed, and instruction fine-tuning is performed on the open source model. During the fine-tuning process, knowledge from the first domain is incorporated into the open source model. A direct preference optimization approach is used to adjust the behavior of the open source model according to human preferences.

3. The AI-driven health question-answering method for individual health management according to claim 2, characterized in that: In step 2, the training process parameter optimization object of the first training process is a Transformer network structure, and the Transformer network structure is composed of multiple encoders and decoders.

4. The AI-driven health question-answering method for individual health management according to claim 1, characterized in that: The Agent intelligent body includes: A context memory module stores historical conversation records and determines whether similar historical conversations occur before answering based on the historical conversation records; A tool list module, which sets a personal information interface through the tool list module, analyzes the user situation in combination with the Agent intelligent body and extracts key information; The execution module sets priorities based on key knowledge of the problem and is responsible for coordinating the calls of different tool interfaces to optimize and process the problem.

5. The AI-driven health question-answering method for individual health management according to claim 1, characterized in that: In step 3, the trained base model is integrated into the local multi-workflow according to the Ollama tool, and the Agent intelligent body is designed, including: The base model after training is in the model.safetensors format. The model.safetensors format is converted to the gguf format based on the llama.cpp tool to complete the format conversion; Importing the converted base model into the Ollama model integration platform; Based on the business scenario, large models are introduced through Ollama to optimize the corresponding business processes.

6. The AI-driven health question-answering method for individual health management according to claim 1, characterized in that: In step 4, after obtaining the user's input question, the following steps are further included: Performing text preprocessing on the input question, wherein the preprocessing includes removing stop words and extracting stems; Using natural language processing technology to perform semantic analysis on the input question to obtain the user's needs, wherein the health question-answering system interacts with the user and provides three forms of interaction including voice, picture, and text; Based on the semantic information of the input question, the health question-answering system determines whether the question belongs to the medical field. If so, it continues to determine the medical sub-field of the input question, selects the medical sub-field agent corresponding to the input question, and sends the input question to the Agent agent for processing.

7. The AI-driven health question-answering method for individual health management according to claim 1, characterized in that: In step 6, the first answer is corrected by following the guidance of the predefined template and the input prompt word, and the corrected first answer is output. The step also includes: The health question and answer system optimizes the first answer, wherein the optimization process includes checking and correcting the logic, accuracy and completeness of the first answer to ensure that the first answer can answer the user's question. The health question and answer system personalizes the first answer based on the user's historical question records and feedback information.

8. The AI-driven health question-answering method for individual health management according to claim 1, wherein: Also includes: The Agent intelligent body, the chronic disease knowledge database and the search engine are integrated into the health question-and-answer system, and after the integration is completed, a system test is performed, wherein the system test includes a functional test, a performance test and a user experience test.

9. An AI-driven health question-answering system for individual health management, characterized by: The system comprises: A first construction unit, configured to construct a chronic disease knowledge database based on the medical data obtained by screening and crawling; A first execution unit, configured to obtain a trained base model using a first training process, wherein the first training process sequentially performs incremental pre-training, supervised fine-tuning, and direct preference optimization on the open source model based on the LoRA method; A first design unit, configured to integrate the trained base model into a local multi-workflow according to an ollama tool and design an Agent intelligent body; a first obtaining unit, configured to obtain a question input by a user; A second execution unit, wherein the second execution unit is used for the Agent intelligent body to automatically call the chronic disease knowledge database according to the input question, search online and perform reasoning to answer, and generate a first answer; The third execution unit is used to correct the first answer through the guidance of a predefined template and an input prompt word, and output the corrected first answer.

10. An AI-driven health question-answering system for individual health management, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Health management method and system based on multi-source data

    CN121215277A

  • A health management method and system based on multi-source data

    CN121215277B

  • AI agent construction method based on continuous and automatic optimization of user marks

    CN121327082A

  • Intelligent patrol auxiliary method and system based on large language model fine tuning

    CN121436166A