Intelligent question and answer method and device, computer equipment and storage medium
By building a comprehensive patient health portrait and introducing a reinforced learning reward mechanism, combined with search enhanced generation technology, the problem of inaccurate and personalized multimodal data integration of consultation chat robots in medical scenarios is solved, and more efficient diagnosis and consultation accuracy is achieved, and full-process intelligent medical services are supported.
Patent Information
- Application Number
- CN202510349093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The consultation chat robots in existing medical scenarios have problems such as inaccurate multimodal data integration, inaccurate image text alignment accuracy, inadequate content generated is not personalized enough, and cannot adapt to dynamic real-time information.
By obtaining the type of input data, the corresponding feature extraction model is used to generate patient health portraits, combined with the context information of the consultation dialogue, the medical knowledge base is used to retrieve consultation information and generate medical suggestions, and the reinforcement learning reward mechanism and search enhancement generation technology are introduced to optimize the medical relevance and language fluency of the consultation content.
It achieves more efficient diagnosis and consultation accuracy, provides intelligent services throughout the process covering pre-diagnosis, during-diagnosis and post-diagnosis, and improves the degree of perfection of the intelligent medical platform.
Smart Images

Figure CN120299752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and medical and health technology, and particularly relates to an intelligent question-answering method, device, computer device, and storage medium. Background Art
[0002] With the popularization of artificial intelligence, intelligent chatbots are widely used in various scenarios. For example, inquiry chatbots in medical scenarios or intelligent customer service in financial scenarios can help users quickly obtain the information they want, and also relieve the manual pressure on hospitals or financial institutions to a certain extent.
[0003] For the medical and health scenario, most current inquiry chatbots conduct multi-round conversations with patients through pictures or text information based on multi-modal vision-language models. However, most of them have main problems such as inaccurate integration of multi-modal data, inaccurate alignment accuracy of image and text, hallucinations in the generated content, lack of personalization, and inability to adapt to dynamic real-time information. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above technical deficiencies and provide an intelligent question-answering method, device, computer device, and storage medium to solve the technical problem of being unable to adapt to dynamic real-time information in the prior art.
[0005] To achieve the above technical purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides an intelligent question-answering method, including the following steps:
[0007] Obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health profile;
[0008] Obtain the context information of the inquiry conversation, and based on the context information and the patient health profile, generate new context information. When the user ends the inquiry, use the final context information as the multi-round conversation result;
[0009] Based on the multi-round conversation result and the patient health profile, retrieve corresponding inquiry information in the medical knowledge base, and generate medical advice based on the inquiry information.
[0010] In some embodiments, the obtaining input data, and based on the type of the input data, using a corresponding feature extraction model to extract features from the input data to generate a patient health profile includes:
[0011] Obtain input data and judge the type of the input data, where the type of the input data is one of image data, text data, and structured health data;
[0012] When the input data is image data, use a pre-trained residual network model to extract features from the input data to generate a patient health portrait;
[0013] When the input data is text data, use a pre-trained biomedical pre-trained language model to extract features from the input data to generate a patient health portrait;
[0014] When the input data is structured health data, use an entity embedding model to extract features from the input data to generate a patient health portrait.
[0015] In some embodiments, the step of when the input data is structured health data, using an entity embedding model to extract features from the input data to generate a patient health portrait includes:
[0016] When the input data is structured health data, perform a splitting process on the structured health data;
[0017] Respectively use the hot encoding mapping sub-model, biomedical pre-trained language sub-model, and residual network sub-model of the entity embedding model to extract features from the corresponding split structured health data;
[0018] Concatenate the feature extraction structures of the hot encoding mapping sub-model, biomedical pre-trained language sub-model, and residual network sub-model to generate a patient health portrait.
[0019] In some embodiments, the step of obtaining the context information of the consultation dialogue, generating new context information based on the context information and the patient health portrait, and using the final context information as the multi-round dialogue result when the user ends the consultation includes:
[0020] Obtain the context information of the consultation dialogue and perform an embedding process on the context information to obtain context word embeddings;
[0021] Based on the context word embeddings and the patient health portrait, use a preset MLP model to generate a context alignment vector;
[0022] Overlay the context alignment vector with the patient health portrait to generate new context information, and use the final context information as the multi-round dialogue result when the user ends the consultation.
[0023] In some embodiments, retrieving corresponding medical interview information from a medical knowledge base based on the multi-round conversation results and the patient's health profile, and generating a medical advice based on the medical interview information, includes:
[0024] Searching in a preset medical knowledge base for the medical interview information with the highest similarity to the multi-round conversation results and the patient's health profile based on the multi-round conversation results and the patient's health profile;
[0025] Generating a medical advice based on the medical interview information.
[0026] In some embodiments, the searching in a preset medical knowledge base for the medical interview information with the highest similarity to the multi-round conversation results and the patient's health profile based on the multi-round conversation results and the patient's health profile includes:
[0027] Calculating the similarity scores between each medical interview information in the medical knowledge base and the multi-round conversation results and the patient's health profile;
[0028] Obtaining the medical interview information most relevant to the multi-round conversation results and the patient's health profile based on the calculated similarity scores.
[0029] In some embodiments, the generating a medical advice based on the medical interview information includes:
[0030] Extracting keywords from the medical interview information based on a preset medical advice template;
[0031] Generating a medical advice based on the keywords in the medical interview information.
[0032] In a second aspect, the present invention further provides an intelligent question-answering device, including:
[0033] A health profile generation module, configured to obtain input data, and perform feature extraction on the input data using a corresponding feature extraction model based on the type of the input data to generate a patient's health profile;
[0034] A context information update module, configured to obtain the context information of a medical interview, generate new context information based on the context information and the patient's health profile, and use the final context information as the multi-round conversation results when the user ends the medical interview;
[0035] A medical advice generation module, configured to retrieve corresponding medical interview information from a medical knowledge base based on the multi-round conversation results and the patient's health profile, and generate a medical advice based on the medical interview information.
[0036] In a third aspect, the present invention further provides a computer device, including a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the intelligent question-and-answer method described above are implemented.
[0037] In a fourth aspect, the present invention further provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the intelligent question-and-answer method described above are implemented.
[0038] Compared with the prior art, the intelligent question-and-answer method, device, computer device, and storage medium provided by the present invention first obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health portrait. Then, obtain the context information of the consultation dialogue, and based on the context information and the patient health portrait, generate new context information until the user ends the consultation, and use the final context information as the multi-round dialogue result. Finally, based on the multi-round dialogue result and the patient health portrait, retrieve corresponding consultation information in the medical knowledge base, and generate medical suggestions based on the consultation information. Compared with the previous intelligent question-and-answer models, the present invention constructs a comprehensive patient health portrait by deeply integrating vision, language, and structured data, introduces a reinforcement learning reward mechanism to optimize the medical relevance and language fluency of the consultation content, and combines retrieval-augmented generation (RAG) technology to achieve precise recommendations for doctors, drugs, and examinations. It can more efficiently improve the accuracy of diagnosis and consultation, help the intelligent health assistant provide full-process intelligent services covering pre-consultation, in-consultation, and post-consultation, and further improve the intelligent medical platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the solutions in the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 is an exemplary system architecture diagram to which the present invention can be applied;
[0041] Figure 2 is a flowchart of an embodiment of the intelligent question-and-answer method according to the present invention;
[0042] Figure 3 is Figure 2 a flowchart of a specific embodiment of step S100 shown;
[0043] Figure 4 isFigure 2 Flow chart of a specific embodiment of step S200 shown
[0044] Figure 5 is Figure 2 Flow chart of a specific embodiment of step S300 shown
[0045] Figure 6 Schematic structural diagram of an embodiment of an intelligent question - answering device according to the present invention
[0046] Figure 7 Schematic structural diagram of an embodiment of a computer device according to the present invention
[0047] Figure 8 Schematic structural diagram of another embodiment of a computer device according to the present invention Specific embodiments
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Referring to "embodiments" herein means that specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0050] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0051] The intelligent question - answering method based on artificial intelligence provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can obtain input data through the client, and based on the type of the input data, adopt a corresponding feature extraction model to extract features from the input data to generate a patient health profile; obtain the context information of the consultation dialogue, and based on the context information and the patient health profile, generate new context information until the user ends the consultation, and use the final context information as the multi-round dialogue result; based on the multi-round dialogue result and the patient health profile, retrieve the corresponding consultation information in the medical knowledge base, and generate medical advice based on the consultation information and feedback it to the client. In the present invention, for the intelligent robot in the medical scenario, by deeply integrating vision, language, and structured data, a comprehensive patient health profile is constructed, and a reinforcement learning reward mechanism is introduced to optimize the medical relevance and language fluency of the consultation content. Combining the retrieval-augmented generation (RAG) technology, accurate recommendations for doctors, drugs, and examinations are realized. It can more efficiently improve the accuracy of diagnosis and consultation, help the intelligent health assistant provide full-process intelligent services covering pre-diagnosis, in-diagnosis, and post-diagnosis for patients, and further improve the intelligent medical platform. Among them, the client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0052] Please refer to Figure 2 , Figure 2 shows a flowchart of an embodiment of the intelligent question-answering method according to the present invention. The intelligent question-answering method is applicable to a machine customer service in a financial scenario or an intelligent consultation robot in a medical scenario, and includes steps S100 to S300.
[0053] S100. Obtain input data, and based on the type of the input data, adopt a corresponding feature extraction model to extract features from the input data to generate a patient health profile.
[0054] In this embodiment, the patient inputs data into the model, and the model will automatically extract the features of the patient's input data, and then generate a patient health portrait based on these features. Among them, the data input by the patient can be of types such as medical lesion images, descriptive texts, or structured health data. According to different data types input by the patient, the embodiments of the present invention adopt different feature extraction models for feature extraction, and then accurately construct the patient's health portrait. Exemplarily, the medical lesion image can be a gray-scale lung CT image, the descriptive text can be "Ping An Doctor, what should I do if I have been having some difficulty breathing recently?", and the structured health data includes information such as the user's basic personal information, lifestyle, past medical history, family medical history, past diagnosis records, past medical examination results, and allergens. For example, "{Personal Information}: {Gender, 32 years old, BMI = 24, etc.}, {Lifestyle}: {Smoking two to three times a day}, {Past Illness}: {Pneumonia}, {Family Medical History}: {None}, {Past Diagnosis Record}: {A certain lung CT}, {Past Medical Examination Results}: {None}, Allergens: {None}".
[0055] S200. Obtain the context information of the consultation dialogue, generate new context information based on the context information and the patient health portrait, and when the user ends the consultation, use the final context information as the multi-round dialogue result.
[0056] In this embodiment, after obtaining the patient's health portrait H, combine it with the context C of the previous generated consultation dialogue result n ={q1,a1,…,q n ,a n}(q i ,a i is the content of the i-th round of question and answer), conduct multi-round dialogues with the fine-tuned open-source LLaMa3-OpenBioLLM-8 / 30B (hereinafter referred to as OpenBioLLM) model using the health portrait and the context (initially without context), and then add the obtained results to the context, and loop this process until the user / patient ends the consultation process.
[0057] S300. Retrieve the corresponding consultation information in the medical knowledge base based on the multi-round dialogue result and the patient health portrait, and generate medical suggestions based on the consultation information.
[0058] In this embodiment, the goal is to provide accurate and personalized recommendation results for the patient's consultation according to the patient / user's health portrait H and the result C of the multi-round dialogue n . For example, "Based on your current lifestyle, physical condition and the results of the lung CT examination, Ping An Assistant recommends that you gradually quit the habit of smoking, appropriately try to take xx cough medicine to relieve cough symptoms, and promptly go to xx hospital for treatment. It is recommended to register for an expert appointment with xx expert to obtain subsequent treatment methods."
[0059] In an embodiment of the present invention, first, input data is obtained. Based on the type of the input data, a corresponding feature extraction model is used to extract features from the input data to generate a patient health portrait. Then, the context information of the consultation dialogue is obtained, and based on the context information and the patient health portrait, new context information is generated. When the user ends the consultation, the final context information is used as the result of the multi-round dialogue. Finally, based on the multi-round dialogue result and the patient health portrait, corresponding consultation information is retrieved from the medical knowledge base, and medical advice is generated based on the consultation information. Compared with the previous intelligent question-answering models, the present invention constructs a comprehensive patient health portrait by deeply integrating vision, language, and structured data, introduces a reinforcement learning reward mechanism to optimize the medical relevance and language fluency of the consultation content, and combines the retrieval-augmented generation (RAG) technology to achieve accurate recommendations for doctors, drugs, and examinations. It can more efficiently improve the accuracy of diagnosis and consultation, help the intelligent health assistant provide full-process intelligent services covering pre-consultation, in-consultation, and post-consultation, and further improve the intelligent medical platform.
[0060] In some embodiments, refer to Figure 3 , the step S100 specifically includes:
[0061] S110. Obtain input data and judge the type of the input data, where the type of the input data is one of image data, text data, and structured health data;
[0062] S120. When the input data is image data, use a pre-trained residual network model to extract features from the input data to generate a patient health portrait;
[0063] S130. When the input data is text data, use a pre-trained biomedical pre-trained language model to extract features from the input data to generate a patient health portrait;
[0064] S140. When the input data is structured health data, use an entity embedding model to extract features from the input data to generate a patient health portrait.
[0065] In this embodiment, the input data includes medical lesion images (H, W, and C are the dimensions of the image. For example, a lung CT grayscale image with a dimension of 224x224x1, and all that follow represent the set of real numbers), descriptive text (L is the length of the input text, such as "Ping An Doctor, what should I do if I have some difficulty breathing recently?"), the structured health data S of the patient, including user's basic personal information, living habits, past medical history, family medical history, past diagnosis records, past medical examination results, allergens, etc., such as "{Personal Information}: {Gender, 32 years old, BMI = 24, etc.}, {Living Habits}: {Smoking two to three times a day}, {Past Illness}: {Pneumonia}, {Family Medical History}: {None}, {Past Diagnosis Records}: {A certain lung CT}, {Past Medical Examination Results}: {None}, Allergens: {None}".
[0066] When performing feature extraction on the input content, for the medical lesion image I, use the pre-trained ResNet50 (ResNet50 is an image model for feature extraction of images based on the deep residual network, the input is an image of 224x224x3, and the output is a vector of (batch_size, num_classes)) to obtain the graph embedding d I is the dimension of the image embedding, and the dimension of the graph feature vector embedding obtained by ResNet50 is d I = 1×2048 (or batch data volume dimension × 2048); for the description text T, use the pre-trained biomedical pre-trained language model BioBERT (BioBERT is a pre-trained language model for biomedical texts, which is a model pre-trained on medical datasets based on BERT) to extract the word embeddings of the medical text d T is the dimension of the word embedding, and the dimension of the text feature embedding extracted by the BioBERT Base model is d T = 1×768, (or batch data volume dimension × 768), specifically as follows:
[0067]
[0068] For more complex structured health data S, the way of Entity embedding (Entity represents entity embedding, maps entity categories, such as the structured health data in the text, to low-dimensional vectors, and maps discrete entities to a continuous vector space) can be used to map to the embedding space d S is the dimension of the word embedding, specifically as follows:
[0069]
[0070] Furthermore, in the user portrait construction stage, in this stage, the model accurately captures the features related to the patient's health through multi-modal input content and constructs the user's health portrait H, that is, the contrastive learning loss:
[0071]
[0072] Among them, E I and E T represent data embeddings from different modalities (images and texts), such that modality pairs with the same semantics are as close as possible in the vector space, while E I and E T′ represent modality pairs with different semantics and are as far apart as possible; similarly, E I and E S are representations of images and structured information. The margin α ensures that the distance between negative sample pairs is greater than that between positive sample pairs, preventing the model from learning invalid embeddings.
[0073] For example, if E I is a lung CT scan and E T is the text description: "The patient has mild pneumonia", and E' T is the text description: "The patient has mild hepatitis", then E I and E' T have dissimilar semantics.
[0074] In some embodiments, the entity embedding model includes a one-hot encoding mapping sub-model, a biomedical pre-trained language sub-model, and a residual network sub-model. The step S140 specifically includes:
[0075] When the input data is structured health data, perform splitting processing on the structured health data;
[0076] Respectively use the one-hot encoding mapping sub-model, the biomedical pre-trained language sub-model, and the residual network sub-model to perform feature extraction on the corresponding split structured health data;
[0077] Concatenate the feature extraction structures of the one-hot encoding mapping sub-model, the biomedical pre-trained language sub-model, and the residual network sub-model to generate a patient health portrait.
[0078] In this embodiment, for those that can be simply mapped using one-hot encoding such as gender, age, BMI, etc., for medical image records, ResNet50 can be used to extract features as above, and for medical treatment text descriptions, BioBERT can also be used to extract features. Finally, the results are concatenated to obtain a complete embedding representation. Since the embeddings of different data in the structured health data are all one-dimensional (or batch data volume-dimensional) vectors, they can be directly concatenated into a high-dimensional embedding vector with a dimension of 1×d s .
[0079] Project the embeddings of the multi-modal input into a unified mapping space through contrastive learning. For example,
[0080] E I ′ = Normalize(W I ·E I ),
[0081] E T ′ = Normalize(W T ·E T ),
[0082] E S ′ = Normalize(W S ·E s ),
[0083] where Normalize is the normalization process, using L2 normalization; are different mapping matrices, d is the dimension of the shared space, and d * is the corresponding embedding dimension. Then, add them item by item for fusion to generate the patient's health portrait H containing medical history, symptoms, images, and other context information, as follows:
[0084]
[0085] In some embodiments, referring to Figure 4 , step S200 specifically includes:
[0086] S210. Obtain the context information of the consultation dialogue, and perform embedding processing on the context information to obtain context word embeddings;
[0087] S220. Based on the context word embeddings and the patient's health portrait, use a preset MLP model to generate context alignment vectors;
[0088] S230. Superimpose the context alignment vectors and the patient's health portrait to generate new context information, and when the user ends the consultation, use the final context information as the result of the multi-round dialogue.
[0089] In this embodiment, in the initial state, the health portrait H has an empty context. For example, "Ping An Doctor, I've been having some difficulty breathing lately. What's wrong?"
[0090] Dialogue process:
[0091] After obtaining H and context C0, OpenBioLLM generates the next-round result word by word through the autoregressive mechanism
[0092] (such as "Hello, do you have any other accompanying symptoms? Do you smoke a lot recently?") output. The patient / user makes an answer a n+1, such as "smoking two or three times a day recently, accompanied by symptoms of coughing and wheezing. This is the lung CT scan I had taken at the hospital").
[0093] In the nth round of conversation, after embedding the previous context with text similar to the input, the context word embedding is obtained
[0094]
[0095] Align the context vector with the healthy portrait vector through MLP
[0096]
[0097] Then add it to the healthy portrait H item by item as the input of OpenBioLLM
[0098]
[0099] The LLM then outputs a new round of conversation or question (e.g., "Based on your symptoms, lifestyle habits, and from the lung CT image you provided, your current lung health condition is quite concerning. It is recommended that you first reduce the number of cigarettes you smoke per day..."). And update the context, adding the new context to the previous context to obtain the new context C n+1 :
[0100] C n+1 = C n ∪ (Q n+1 , A n+1}
[0101] This process is then looped until the consultation ends.
[0102] Furthermore, in the multi-round conversation result generation stage, it is mainly for the training of the multi-round conversation of the text generation task of the language model. Through the autoregressive mechanism, the next round of consultation questions or diagnostic suggestions are generated. The model needs to learn how to generate subsequent questions based on the historical conversation and adjust according to the user's feedback during the conversation.
[0103] The cross-entropy function is used as the loss function in this stage
[0104]
[0105] where w t is the word generated at the t-th step, and w<t is the word generated previously. In addition, a reinforcement learning reward mechanism is introduced to optimize the quality of the consultation content. The reward function can be designed based on aspects such as the relevance of the generated questions and the accuracy of the medical advice.
[0106] R(Q n+1 , H, Cn ) = α Related(Q n+1 , H) + β Fluent(Q n+1 )
[0107]
[0108] where Related(Q n+1 , H) is the similarity between the new output of the LLM and the patient / user health profile H(E H ), represented by the cosine similarity of their embedding vectors, and the larger the value, the more relevant; Fluent(Q n+1 ) is the fluency of the LLM in the new output, evaluated by perplexity, and the higher the score, the more fluent the generated content; α and β are the weights for balancing relevance and fluency, and here α = 0.7 and β = 0.3 are set to highlight the priority of medical relevance.
[0109] In some embodiments, referring to Figure 5 , step S300 specifically includes:
[0110] S310. Based on the multi-round conversation results and the patient health profile, search for the medical consultation information with the highest similarity to the multi-round conversation results and the patient health profile in a preset medical knowledge base;
[0111] S320. Generate medical advice based on the medical consultation information.
[0112] In this embodiment, first, a private medical database is integrated, including preparing diseases, treatment methods, treatment drugs, relevant high-quality hospitals, departments, and physicians, etc., and integrating the user's health profile H and the results C n . Similar to the model input embedding representation, the data content is represented in an embedded form, and these embedding vectors are stored in the medical knowledge base. Then, the most similar medical consultation information is searched in the database according to the patient health profile and the multi-round conversation results, and medical advice is generated based on the medical consultation information.
[0113] In this stage, the model retrieves information related to medical consultation from the medical knowledge base constructed by RAG and generates corresponding personalized answers. In the retrieval stage, according to the user health profile and the context of the previous conversation, relevant content is found through vector retrieval. Then, the found relevant content is combined with the conversation context and used as the input of the LLM to generate corresponding personalized results.
[0114] In the retrieval stage, the loss is used to optimize the retrieval module so that the information retrieved from the database is highly relevant to the user's needs:
[0115]
[0116] Among them, relevant refers to the relevant content in the knowledge base; in the generation stage, the loss is similar to that in multi-turn conversations, and the cross-entropy loss can be used to train the model:
[0117]
[0118] Among them, R is the retrieved relevant information.
[0119] In some embodiments, step S310 specifically includes:
[0120] Calculate the similarity scores of each inquiry information in the medical knowledge base with the multi-turn conversation result and the patient health portrait;
[0121] Based on the calculated similarity scores, obtain the inquiry information most relevant to the multi-turn conversation result and the patient health portrait.
[0122] In this embodiment, according to the patient / user's health portrait H and the result C of the multi-round conversation n , by calculating the average score of the highest cosine similarity and Euclidean distance among the health portrait, the result of the multi-round conversation, and the content in the database, retrieve the most relevant content according to the level of the average score, and design medical advice based on this.
[0123] In some embodiments, step S320 specifically includes:
[0124] Extract keywords from the inquiry information based on a preset medical advice template;
[0125] Generate medical advice based on the keywords in the inquiry information.
[0126] Exemplarily, the key information of the patient's health portrait, such as "smoking history, mild lung nodules shown in lung CT, cough, wheezing", and the conversation context is "The patient indicates recent shortness of breath, cough with phlegm, recent increase in smoking volume, and the lung CT result shows mild pneumonia". According to the above patient information combined with the retrieved relevant medical data: the retrieved content, such as "Lung nodules are common in long-term smokers. Recommend Dr. Li, an expert in the Department of Respiratory Medicine, and cough medicine XX". After that, a personalized medical advice can be generated based on the retrieved information, including a preliminary diagnosis and possible causes, recommended departments, doctors, examination items, medication advice, and health management advice, and then have a conversation with OpenBioLLM and return the result.
[0127] The technical solution provided by the present invention is as follows: First, input data is obtained, and based on the type of the input data, a corresponding feature extraction model is used to extract features from the input data to generate a patient health profile. Then, the context information of the consultation conversation is obtained, and based on the context information and the patient health profile, new context information is generated. When the user ends the consultation, the final context information is used as the result of the multi-round conversation. Finally, based on the multi-round conversation result and the patient health profile, corresponding consultation information is retrieved from the medical knowledge base, and medical advice is generated based on the consultation information. Compared with the previous intelligent question-answering models, the present invention constructs a comprehensive patient health profile by deeply integrating vision, language, and structured data, introduces a reinforcement learning reward mechanism to optimize the medical relevance and language fluency of the consultation content, and combines the retrieval-augmented generation (RAG) technology to achieve accurate recommendations for doctors, drugs, and examinations. It can more efficiently improve the accuracy of diagnosis and consultation, help the intelligent health assistant provide full-process intelligent services covering pre-diagnosis, in-diagnosis, and post-diagnosis for patients, and further improve the intelligent medical platform.
[0128] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0129] Another embodiment of the present invention provides an intelligent question-answering device, which corresponds to the intelligent question-answering method in the above embodiment one by one. Please refer to Figure 6 , the intelligent question-answering device includes a health profile generation module 11, a context information update module 12, and a medical advice generation module 13. The detailed description of each functional module is as follows:
[0130] The health profile generation module 11 is used to obtain input data, and based on the type of the input data, a corresponding feature extraction model is used to extract features from the input data to generate a patient health profile.
[0131] The context information update module 12 is used to obtain the context information of the consultation conversation, and based on the context information and the patient health profile, generate new context information. When the user ends the consultation, the final context information is used as the result of the multi-round conversation.
[0132] The medical advice generation module 13 is used to retrieve corresponding consultation information from the medical knowledge base based on the multi-round conversation result and the patient health profile, and generate medical advice based on the consultation information.
[0133] In some embodiments, the health profile generation module 11 specifically includes:
[0134] A type judgment unit for obtaining input data and judging the type of the input data, where the type of the input data is one of image data, text data, and structured health data;
[0135] An image data processing unit for, when the input data is image data, using a pre-trained residual network model to extract features from the input data to generate a patient health portrait;
[0136] A text data processing unit for, when the input data is text data, using a pre-trained biomedical pre-trained language model to extract features from the input data to generate a patient health portrait;
[0137] A structured health data processing unit for, when the input data is structured health data, using an entity embedding model to extract features from the input data to generate a patient health portrait.
[0138] In some embodiments, the structured health data processing unit is specifically configured to:
[0139] When the input data is structured health data, perform splitting processing on the structured health data;
[0140] Respectively use the hot encoding mapping sub-model, the biomedical pre-trained language sub-model, and the residual network sub-model to extract features from the corresponding split structured health data;
[0141] Concatenate the feature extraction structures of the hot encoding mapping sub-model, the biomedical pre-trained language sub-model, and the residual network sub-model to generate a patient health portrait.
[0142] In some embodiments, the entity embedding model includes a hot encoding mapping sub-model, a biomedical pre-trained language sub-model, and a residual network sub-model, and the context information update module 12 specifically includes:
[0143] A context word embedding calculation unit for obtaining the context information of the interrogation dialogue and performing embedding processing on the context information to obtain context word embeddings;
[0144] An alignment vector calculation unit for generating a context alignment vector based on the context word embeddings and the patient health portrait by using a preset MLP model;
[0145] A dialogue result generation unit for superimposing the context alignment vector and the patient health portrait to generate new context information, and when the user ends the interrogation, taking the final context information as the multi-round dialogue result.
[0146] In some embodiments, the medical advice generation module 13 specifically includes:
[0147] An inquiry information search unit, configured to search for the inquiry information with the highest similarity to the multi-round conversation result and the patient health portrait in a preset medical knowledge base based on the multi-round conversation result and the patient health portrait;
[0148] A medical advice generation unit, configured to generate medical advice based on the inquiry information.
[0149] In some embodiments, the inquiry information search unit is specifically configured to:
[0150] Calculate the similarity scores of each inquiry information in the medical knowledge base with the multi-round conversation result and the patient health portrait;
[0151] Based on the calculated similarity scores, obtain the inquiry information most relevant to the multi-round conversation result and the patient health portrait.
[0152] In some embodiments, the medical advice generation unit is specifically configured to:
[0153] Extract keywords from the inquiry information based on a preset medical advice template;
[0154] Generate medical advice based on the keywords in the inquiry information.
[0155] In the embodiments of the present invention, first, input data is obtained, and based on the type of the input data, a corresponding feature extraction model is used to extract features from the input data to generate a patient health portrait; then, the context information of the inquiry conversation is obtained, and based on the context information and the patient health portrait, new context information is generated until the user ends the inquiry, and the final context information is used as the multi-round conversation result; finally, based on the multi-round conversation result and the patient health portrait, corresponding inquiry information is retrieved in the medical knowledge base, and medical advice is generated based on the inquiry information. Compared with the previous intelligent question-and-answer models, the present invention constructs a comprehensive patient health portrait by deeply integrating vision, language, and structured data, introduces a reinforcement learning reward mechanism to optimize the medical relevance and language fluency of the inquiry content, and combines the retrieval-augmented generation (RAG) technology to achieve accurate recommendations for doctors, drugs, and examinations. It can more efficiently improve the accuracy of diagnosis and inquiry, help the intelligent health assistant provide full-process intelligent services covering pre-diagnosis, in-diagnosis, and post-diagnosis for patients, and further improve the intelligent medical platform.
[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0157] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly restricted by order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0158] For the specific limitations of the intelligent question-answering device, reference can be made to the limitations of the intelligent question-answering method in the above text, which will not be elaborated here. Each module in the above intelligent question-answering device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0159] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an intelligent question-answering method based on artificial intelligence.
[0160] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be asFigure 8 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an intelligent question-and-answer method based on artificial intelligence
[0161] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0162] Obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health portrait;
[0163] Obtain the context information of the interrogation dialogue, and based on the context information and the patient health portrait, generate new context information. Until the user ends the interrogation, the final context information is used as the multi-round dialogue result;
[0164] Based on the multi-round dialogue result and the patient health portrait, retrieve corresponding interrogation information in the medical knowledge base, and generate medical advice based on the interrogation information.
[0165] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are realized:
[0166] Obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health portrait;
[0167] Obtain the context information of the interrogation dialogue, and based on the context information and the patient health portrait, generate new context information. Until the user ends the interrogation, the final context information is used as the multi-round dialogue result;
[0168] Based on the multi-round dialogue result and the patient health portrait, retrieve corresponding interrogation information in the medical knowledge base, and generate medical advice based on the interrogation information.
[0169] It should be noted that for the functions or steps that can be implemented by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0170] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0171] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0172] In summary, for the intelligent question-answering method, device, computer device, and storage medium provided by the present invention, input data is first obtained, and based on the type of the input data, a corresponding feature extraction model is used to extract features from the input data to generate a patient health profile. Then, the context information of the consultation conversation is obtained, and based on the context information and the patient health profile, new context information is generated until the user ends the consultation, and the final context information is used as the result of the multi-round conversation. Finally, based on the multi-round conversation result and the patient health profile, corresponding consultation information is retrieved from the medical knowledge base, and medical suggestions are generated based on the consultation information. Compared with previous intelligent question-answering models, the present invention constructs a comprehensive patient health profile by deeply integrating visual, language, and structured data, introduces a reinforcement learning reward mechanism to optimize the medical relevance and language fluency of the consultation content, and combines the retrieval-augmented generation (RAG) technology to achieve accurate recommendations for doctors, drugs, and examinations. It can more efficiently improve the accuracy of diagnosis and consultation, help the intelligent health assistant provide full-process intelligent services covering pre-diagnosis, in-diagnosis, and post-diagnosis for patients, and further improve the intelligent medical platform.
[0173] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use.
[0174] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. An intelligent question answering method, characterized in that, It includes the following steps: Obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health portrait; Obtain the context information of the consultation dialogue, and based on the context information and the patient health portrait, generate new context information. When the user ends the consultation, use the final context information as the multi-round dialogue result; Based on the multi-round dialogue result and the patient health portrait, retrieve corresponding consultation information in the medical knowledge base, and generate medical advice based on the consultation information.
2. The intelligent question-answering method according to claim 1, characterized in that, The step of obtaining input data, and based on the type of the input data, using a corresponding feature extraction model to extract features from the input data to generate a patient health portrait includes: Obtain input data and judge the type of the input data, where the type of the input data is one of image data, text data, and structured health data; When the input data is image data, use a pre-trained residual network model to extract features from the input data to generate a patient health portrait; When the input data is text data, use a pre-trained biomedical pre-trained language model to extract features from the input data to generate a patient health portrait; When the input data is structured health data, use an entity embedding model to extract features from the input data to generate a patient health portrait.
3. The intelligent question and answer method according to claim 2, wherein The step of, when the input data is structured health data, using an entity embedding model to extract features from the input data to generate a patient health portrait includes: When the input data is structured health data, perform splitting processing on the structured health data; Respectively use the hot encoding mapping sub-model, biomedical pre-trained language sub-model, and residual network sub-model of the entity embedding model to extract features from the corresponding split structured health data; Concatenate the feature extraction structures of the hot encoding mapping sub-model, biomedical pre-trained language sub-model, and residual network sub-model to generate a patient health portrait.
4. The intelligent question-answering method according to claim 1, characterized in that, The step of obtaining the context information of the consultation dialogue, and based on the context information and the patient health portrait, generating new context information. When the user ends the consultation, using the final context information as the multi-round dialogue result includes: Obtain the context information of the consultation dialogue and perform embedding processing on the context information to obtain context word embeddings; Based on the context word embeddings and the patient health portrait, use a preset MLP model to generate context alignment vectors; Overlay the context alignment vectors with the patient health portrait to generate new context information. When the user ends the consultation, use the final context information as the multi-round dialogue result.
5. The intelligent question-answering method according to claim 1, characterized in that The step of, based on the multi-round dialogue result and the patient health portrait, retrieving corresponding consultation information in the medical knowledge base and generating medical advice based on the consultation information includes: Based on the multi-round dialogue result and the patient health portrait, find the consultation information with the highest similarity to the multi-round dialogue result and the patient health portrait in a preset medical knowledge base; Generate medical advice based on the medical interview information.
6. The intelligent question-answering method according to claim 5, wherein The method of finding the medical interview information with the highest similarity to the multi-round conversation result and the patient health profile in a preset medical knowledge base based on the multi-round conversation result and the patient health profile includes: Calculate the similarity scores of each medical interview information in the medical knowledge base with the multi-round conversation result and the patient health profile; Based on the calculated similarity scores, obtain the medical interview information with the highest relevance to the multi-round conversation result and the patient health profile.
7. The intelligent question-answering method according to claim 5, wherein The method of generating medical advice based on the medical interview information includes: Extract keywords from the medical interview information based on a preset medical advice template; Generate medical advice based on the keywords in the medical interview information.
8. An intelligent question-answering device, characterized in that, It includes: A health profile generation module, configured to obtain input data, and based on the type of the input data, use a corresponding feature extraction model to extract features from the input data to generate a patient health profile; A context information update module, configured to obtain the context information of a medical interview, and based on the context information and the patient health profile, generate new context information, and when the user ends the medical interview, use the final context information as the multi-round conversation result; A medical advice generation module, configured to retrieve corresponding medical interview information in a medical knowledge base based on the multi-round conversation result and the patient health profile, and generate medical advice based on the medical interview information.
9. A computer device, characterized in that, It includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the intelligent question-and-answer method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on a computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the intelligent question-and-answer method according to any one of claims 1 to 7 are implemented.