Dialogue generation method and system for medical artificial intelligence system
The integration of general and specialized medical knowledge with targeted entity mining and context-aware encoding in medical chat systems addresses the limitations of existing systems, providing accurate and adaptive post-operative care guidance.
Patent Information
- Application Number
- CN202510382417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing medical artificial intelligence systems lack the mastery of postoperative terminology in postoperative follow-up, making it difficult to provide personalized and coherent rehabilitation guidance, and the traditional knowledge base cannot be updated quickly, resulting in answering non-questions or incorrect guidance.
Build a multi-layer medical knowledge base, identify medical terms through target entity mining algorithms, generate personalized dialogue replies based on context fusion processing, and introduce a security control mechanism.
It improves the depth and accuracy of understanding of medical entities, can quickly adapt to new symptoms and rehabilitation methods, generate coherent and clinical logic responses, reduce the burden on medical staff, provide personalized rehabilitation guidance, and ensure safety.
Smart Images

Figure CN120316218A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent medicine, and particularly relates to a dialogue generation method and system for a medical artificial intelligence system. Background Art
[0002] With the intensification of the aging population trend, osteoporotic vertebral compression fractures are in a high incidence among the elderly population. Clinically, after treatment by means such as percutaneous vertebroplasty and posterior internal fixation, patients often need continuous rehabilitation training, pain management, and complication monitoring. However, the existing postoperative follow-up methods often rely on regular outpatient clinics or telephone follow-up. The workload of medical staff is large, and patient consultations are fragmented, making it difficult to ensure personalized and refined rehabilitation advice in multi-round exchanges.
[0003] Although medical dialogue systems based on artificial intelligence have been tried in the fields of medical Q&A, health consultation, etc., there are still the following deficiencies for professional dialogues: general dialogue models lack knowledge of specialized postoperative terms and sub-domains such as rehabilitation training, and are prone to giving irrelevant answers or incorrect guidance; when patients present new symptoms or use new postoperative rehabilitation means, traditional knowledge bases cannot quickly capture and update, resulting in delayed answers or insufficient coverage; ordinary Q&A systems have limited memory and association capabilities for context information, and it is difficult to generate coherent and clinically logical multi-round responses. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the invention is to provide a dialogue generation method and system for a medical artificial intelligence system.
[0005] The present invention provides a dialogue generation method for a medical artificial intelligence system, including:
[0006] S1: Fuse general concepts with specialized knowledge to construct a multi-layer medical knowledge base;
[0007] S2: Perform masking prediction on medical terms in the multi-round dialogue corpus of patients through a target entity mining algorithm to obtain a medical entity prediction result;
[0008] S3: Perform context fusion processing on the historical interaction information and current input information of patients through an encoder to obtain a comprehensive semantic vector;
[0009] S4: Based on the multi-layer medical knowledge base and the medical entity prediction result, generate a reply for the comprehensive semantic vector through a decoder to obtain a personalized medical dialogue reply.
[0010] According to the dialogue generation method for a medical artificial intelligence system provided by the present invention, step S1 further includes:
[0011] S11: Classify and extract basic concepts according to the medical term library and establish an index to obtain a general medical dictionary database;
[0012] S12: Extract information from the information of the specialized nursing scenario to obtain a local knowledge base of specialized knowledge;
[0013] S13: Store the relationships of the local knowledge base of specialized knowledge through the triple coding method to obtain a knowledge representation;
[0014] S14: Integrate the medical dictionary database and the knowledge representation to obtain a multi-layer medical knowledge base.
[0015] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, step S2 further includes:
[0016] S21: Scan and screen the dialogue anticipation to obtain candidate word phrases representing medical meanings;
[0017] S22: Replace the candidate word phrases through masking marks to obtain a text sequence containing masking positions;
[0018] S23: Semantically encode the text sequence through a pre-trained language encoder to obtain a context hidden layer vector group;
[0019] S24: Based on the context hidden layer vector group, calculate the probability distribution of the hidden layer vector corresponding to each masking mark to obtain the prediction distribution of candidate medical entities;
[0020] S25: Output the medical entity prediction result with the highest probability according to the prediction distribution.
[0021] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, step S23 further includes:
[0022] S231: Inject position information into each token in the input text through a position encoding algorithm to obtain a first token representation containing position information;
[0023] S232: Perform context correlation calculation on the first token representation through a multi-head self-attention mechanism to obtain a second token representation considering global information;
[0024] S233: Perform a non-linear transformation on the second token representation through a feed-forward neural network to obtain a final context hidden layer vector group.
[0025] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, the expression of the prediction distribution in step S24 is:
[0026] P(e|hm ; Θ) = Softmax(W · h m + b);
[0027] where P(e|h m ; Θ) is the output probability distribution, e is the predicted entity, h m is the m-th hidden layer vector to be predicted, Θ is the set of model parameters, Softmax is the softmax function, W is the first learnable parameter, and b is the second learnable parameter.
[0028] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, after step S2, it further includes:
[0029] When the medical entity prediction result does not exist in the local knowledge base of specialized knowledge, the local knowledge base of specialized knowledge is dynamically updated.
[0030] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, step S3 further includes:
[0031] S31: Perform time series storage processing on the historical questions of the patient and the system responses to obtain a complete dialogue history record;
[0032] S32: Calculate the correlation degree of the dialogue history record to obtain a historical context segment related to the current question;
[0033] S33: Combine and represent the historical context segment and the current input text to obtain preliminarily fused context information;
[0034] S34: Based on the context information, perform global correlation modeling to obtain a semantic representation that captures long-term dependence relationships;
[0035] S35: Integrate the semantic representation and the medical entity prediction result to obtain a comprehensive semantic vector.
[0036] According to a dialogue generation method for a medical artificial intelligence system provided by the present invention, step S4 further includes:
[0037] S41: Based on the medical entity prediction result, use a decoder to generate an answer word by word for the comprehensive semantic vector to obtain a response sequence;
[0038] S42: Calculate the prediction probability of each word in the response sequence, and based on the prediction probability, generate a reply based on a multi-layer medical knowledge base to obtain an initial medical dialogue reply;
[0039] S43: Perform security control on the initial medical dialogue reply to obtain a personalized medical dialogue reply.
[0040] A dialogue generation method for a medical artificial intelligence system provided by the present invention, step S43 specifically includes:
[0041] For the content related to drug dosage and postoperative activities in the initial medical dialogue reply, call the local specialty knowledge base for matching;
[0042] When the initial medical dialogue reply contains risk information, prompt it in the output personalized medical dialogue reply.
[0043] The present invention also provides a dialogue generation system for a medical artificial intelligence system, which is used to execute a dialogue generation method for a medical artificial intelligence system as described in any one of the above, including:
[0044] A construction module: used to fuse general concepts with specialty information to construct a multi-layer medical knowledge base;
[0045] A prediction module: used to perform masking prediction on medical terms in the patient's multi-round dialogue corpus through a target entity mining algorithm to obtain a medical entity prediction result;
[0046] A fusion module, the fusion module is configured as an encoder, used to perform context fusion processing on the patient's historical interaction information and the current input information to obtain a comprehensive semantic vector;
[0047] A reply module, the reply module is configured as a decoder, used to generate a reply for the comprehensive semantic vector based on the multi-layer medical knowledge base and the medical entity prediction result to obtain a personalized medical dialogue reply.
[0048] The beneficial effects of the present invention are as follows:
[0049] A dialogue generation method and system for a medical artificial intelligence system provided by the present invention first fuse a general medical dictionary database with a local knowledge base after orthopedic surgery by constructing a multi-layer medical knowledge base, forming a medical knowledge system with distinct levels and comprehensive coverage. The multi-level knowledge structure design enables the system to master both general medical terms and professional knowledge of specialty rehabilitation at the same time, greatly improving the system's understanding depth and accuracy of medical entities. Compared with traditional dialogue systems, the present invention can accurately identify and interpret professional terms after orthopedic surgery, effectively avoiding the problems of answering irrelevant questions and giving wrong guidance.
[0050] Secondly, the present invention introduces the Target Entity Mining (TEM) algorithm. Through the masking and prediction mechanisms, it can automatically identify key medical terms in the conversation, not only improving the system's sensitivity to medical terms, but more importantly, being able to discover and learn newly emerging medical concepts. When the medical entities predicted by the system do not exist in the knowledge base, they will be supplemented to the knowledge base through the dynamic update mechanism, realizing the continuous accumulation and update of knowledge, solving the problem of difficult static update of traditional knowledge bases, and enabling the system to quickly adapt to the emergence of new symptoms and new postoperative rehabilitation methods.
[0051] The multi-round dialogue context fusion processing technology adopted by the present invention realizes the understanding and memory of long-term dialogue context through the comprehensive analysis of historical interaction information and the current input, and can remember information such as symptoms and rehabilitation conditions mentioned by the patient in previous rounds, and associate this information with the current problem to generate coherent and clinically logical responses, overcoming the defect that ordinary question-and-answer systems have limited ability to remember and associate context information.
[0052] In terms of security, the present invention designs a special security control mechanism to conduct professional verification on sensitive content such as drug dosage and postoperative activities, and can identify high-risk symptoms, providing timely medical advice in the response, effectively preventing potential harm caused by incorrect medical guidance to patients, and ensuring the security of system application.
[0053] The implementation of the present invention can significantly reduce the workload of medical staff in postoperative follow-up. The system can automatically handle patients' routine consultations and provide personalized rehabilitation guidance for patients, enabling medical staff to concentrate their limited energy on high-risk cases that require professional intervention, improving the utilization efficiency of medical resources, being able to provide rehabilitation suggestions in line with the individual conditions of patients, realizing the personalized and refined management of postoperative care, and effectively improving the rehabilitation effect and satisfaction of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0055] Figure 1 It is a schematic flowchart of a dialogue generation method for a medical artificial intelligence system provided by an embodiment of the present invention;
[0056] Figure 2 It is a schematic structural diagram of a dialogue generation system for a medical artificial intelligence system provided by an embodiment of the present invention.
[0057] Reference Signs:
[0058] 100, Building block; 200, Prediction module; 300, Fusion module; 400, Reply module. Detailed implementation manners
[0059] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts disclosed in the present invention.
[0061] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. The terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0062] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present invention. On the contrary, they are merely examples of methods and systems consistent with some aspects of the present invention as detailed in the appended claims.
[0063] The embodiments of the present invention will be described below with reference to the drawings.
[0064] As Figure 1As shown, the present invention provides a dialogue generation method for a medical artificial intelligence system, including:
[0065] S1: Integrate general concepts with specialist knowledge to construct and obtain a multi-layer medical knowledge base.
[0066] Furthermore, step S1 aims to create a multi-level medical knowledge structure system. Specifically, the specialty in the specific specialist knowledge is orthopedics. By constructing a multi-layer medical knowledge base, the system can simultaneously master general medical concepts and the proprietary knowledge of postoperative rehabilitation of osteoporotic vertebral compression fractures (OVCF) of the spine, providing a comprehensive semantic background and professional support for subsequent intelligent conversations.
[0067] Among them, step S1 further includes:
[0068] S11: Classify and extract basic concepts according to a medical term library and establish an index to obtain a general medical dictionary database.
[0069] Specifically, step S11 extracts medical information from existing publicly available medical term libraries or professional medical literature to obtain basic concepts and general semantic relationships corresponding to common diseases, drugs, diagnostic and treatment methods, etc. Subsequently, these extracted information needs to be organized in a hierarchical classification or graph structure form, aiming to identify conventional medical terms (such as "osteoporosis", "analgesic drugs", etc.) during subsequent dialogue processing.
[0070] Step S11 essentially establishes a basic framework of general medical knowledge for the dialogue method of the present invention, ensuring the ability to understand a wide range of medical terms and concepts, laying a foundation for in-depth understanding of specialist knowledge. Through this method of classification extraction and index establishment, the system can quickly locate and retrieve medical concepts, improving the efficiency and accuracy of subsequent dialogue processing.
[0071] S12: Extract information from the information of the specialist nursing scenario to obtain a local knowledge base of specialist knowledge.
[0072] In a specific embodiment, the specialty in the present invention is orthopedics. For orthopedics, step S12 mainly focuses on the postoperative care scenario of OVCF. Key information such as hospital rehabilitation manuals, electronic medical records, follow-up records, and expert interviews are extracted. The purpose is to collect and organize professional knowledge directly related to the postoperative rehabilitation of spinal osteoporotic vertebral compression fractures. This kind of professional knowledge usually does not appear in general medical dictionaries, but is crucial for providing accurate postoperative care guidance. By extracting information in the specialty field, the present invention can obtain more refined, professional, and targeted knowledge, such as postoperative precautions for specific surgical methods (percutaneous vertebroplasty, posterior internal fixation, etc.), treatment methods for common complications (incision infection, bone cement leakage, nerve compression, etc.), and professional content such as the applicable timing and methods of rehabilitation exercises (lumbar and back muscle exercises, lower limb muscle strength training, etc.), thereby establishing a local knowledge base of specialty knowledge.
[0073] S13: Store the relationships of the local knowledge base of specialty knowledge through triple encoding to obtain a knowledge representation.
[0074] In step S13 of the present invention, triple or equivalent relationship structures are used to store orthopedic proprietary knowledge. The triple structure is usually expressed in the form of "head entity, relationship, tail entity", such as "percutaneous vertebroplasty, postoperative complication, incision infection" and "lumbar and back muscle exercise, applicable stage, 2 - 4 weeks after surgery".
[0075] The triple encoding method adopted by the present invention has highly structured characteristics, can clearly express the semantic relationships between medical concepts, and is convenient for the system to perform reasoning and querying. Through triple encoding, complex medical knowledge is transformed into a form that can be efficiently processed by a computer, laying a structured foundation for subsequent knowledge retrieval, relationship reasoning, and dynamic update, and is particularly suitable for expressing various relationships between entities in the medical field, such as multi-dimensional associations like indications, contraindications, pathogenesis, treatment methods, etc.
[0076] S14: Integrate the medical dictionary database and the knowledge representation to obtain a multi-layer medical knowledge base.
[0077] In step S14, through the fusion of the above-mentioned two-layer or multi-layer knowledge bases, the present invention can take into account both general medical concepts and proprietary information on spinal osteoporosis postoperative rehabilitation, laying a foundation for medical entity annotation, relationship reasoning, and dynamic update in multi-round conversations. The integration process organically combines two knowledge bases at different levels to form a medical knowledge system with clear levels and comprehensive coverage.
[0078] S2: Use the target entity mining algorithm to perform masking prediction on medical terms in the patient's multi-round conversation corpus to obtain medical entity prediction results.
[0079] In step S2 of the present invention, the Targeted Entity Mining (TEM) algorithm is adopted to solve the problem of insufficient understanding of medical entities. In step S2, by identifying and predicting medical terms in multiple rounds of conversations between patients and the system, the medical entity information contained in the conversations can be automatically discovered. Based on the TEM technology, not only can known medical terms be identified, but more importantly, newly emerging medical concepts can be learned and identified, realizing the discovery and understanding of unknown medical entities. The masked prediction method is a self-supervised learning method, enabling the system to continuously learn and adapt to new medical terms in practical applications, especially the professional expressions in the field of postoperative rehabilitation of OVCF.
[0080] Among them, step S2 further includes:
[0081] S21: Scan and screen the conversation corpus to obtain candidate word phrases representing medical meanings.
[0082] In the multi-round conversation corpus between the patient and the system, in step S21, words or phrases that may represent medical meanings (such as pain score, soreness and swelling in the waist and back, etc.) are detected first. The system scans the conversation content through word analysis technology to identify terms or expressions that may have medical meanings.
[0083] For the postoperative rehabilitation scenario of OVCF, these candidate word phrases may include symptoms described by the patient (such as soreness, numbness), rehabilitation activities (such as lumbar back muscle exercise), drug use conditions, or complication manifestations. Effective scanning and screening can ensure that the system does not miss important medical information, while avoiding processing too much irrelevant content, improving the accuracy and efficiency of subsequent predictions.
[0084] S22: Replace the candidate word phrases with masking markers to obtain a text sequence containing masking positions.
[0085] Furthermore, in step S22, mainly the symbols "[MASK]" or "[hidden]" are used for replacement to form a text sequence with several masking positions. For example, when the patient inputs "I have been suffering from severe soreness and swelling in my waist and back recently. Do I need to stay in bed and rest?", it will be converted to "I have been suffering from [MASK] recently. Do I need to stay in bed and rest?". The masking marker method is the basis of self-supervised learning. By artificially masking, the model can learn how to predict the masked content according to the context subsequently.
[0086] S23: Semantically encode the text sequence through a pre-trained language encoder to obtain a group of context hidden layer vectors.
[0087] The present invention introduces a medical pre-trained language encoder that maps a masked text sequence to a set of hidden layer vectors, where each hidden layer vector corresponds to a word (or sub-word) position in the input sequence and is used to represent its contextual semantics.
[0088] The pre-trained language encoder in the medical field can better understand medical terms and expressions, capture the semantic features of medical texts. By converting the text into hidden layer vector representations, the system can process semantic information in a high-dimensional space, providing a rich semantic representation basis for subsequent entity prediction.
[0089] Among them, step S23 further includes:
[0090] S231: Inject position information into each token in the input text through a position encoding algorithm to obtain a first token representation containing position information.
[0091] Furthermore, the purpose of the position encoding algorithm is to add position information to each token in the input sequence, enabling the model to recognize the position of words in a sentence and thus understand the impact of word order on semantics.
[0092] In the applicable scenario of the present invention, namely medical conversations, symptom descriptions, time expressions, and causal relationships highly depend on the position and order of words. For example, "pain relieved two days after surgery" and "pain after surgery relieved two days later" express different time states. Through position information injection, a preliminary token representation containing position semantics can be obtained, providing sequence structure information for subsequent context understanding.
[0093] S232: Perform context correlation calculation on the first token representation through a multi-head self-attention mechanism to obtain a second token representation considering global information.
[0094] The multi-head self-attention mechanism can automatically learn the correlation relationships between words in the text, enabling the representation of each token to consider the global information of the entire sequence. In medical conversations, since medical symptom descriptions are often scattered in multiple sentences and long-distance semantic associations need to be established. For example, a patient may describe "low back pain" at the beginning of a conversation and mention "aggravated when standing" in a subsequent sentence. The system needs to associate these scattered pieces of information. Through the multi-head self-attention mechanism, the system can capture such complex context dependencies and generate token representations containing global semantic information, greatly improving the ability to understand medical conversation content.
[0095] S233: Perform a non-linear transformation on the second token representation through a feed-forward neural network to obtain a final set of context hidden layer vectors.
[0096] In step S233, the masked text sequence is mapped to a set of hidden layer vectors to generate a final set of context hidden layer vectors. The hidden layer vectors contain rich semantic information, integrating not only the meaning of the words themselves, but also the position information and global context information.
[0097] S24: Based on the set of context hidden layer vectors, calculate the probability distribution for the hidden layer vector corresponding to each masked token to obtain the prediction distribution of candidate medical entities.
[0098] In step S24, the present invention calculates the probability distribution of each possible medical entity based on the context hidden layer vectors, especially the vectors corresponding to the masked positions. The prediction mechanism can determine what the most likely masked medical term is according to the context. For example, in the context of "I have been having [MASK] recently", the system may predict the probabilities of symptom description words such as "low back pain", "swelling and soreness", or "numbness". This context-based probability prediction is the core of target entity mining, enabling the system to identify and understand key medical information in the dialogue, including symptoms, drugs, treatment methods, etc.
[0099] Among them, the expression of the prediction distribution in step S24 is:
[0100] P(e|h m ; Θ) = Softmax(W·h m +b);
[0101] Among them, P(e|h m ; Θ) is the output probability distribution, e is the predicted entity, h m is the m-th hidden layer vector to be predicted, Θ is the set of model parameters, Softmax is the softmax function, W is the first learnable parameter, and b is the second learnable parameter.
[0102] Furthermore, the expression of the objective function for the above prediction part is:
[0103]
[0104] Among them, L TEM is the loss function in the prediction stage, M is the total number of hidden layer vectors, and e * is the true label of the m-th hidden layer vector.
[0105] S25: Output the medical entity prediction result with the highest probability according to the prediction distribution.
[0106] In step S25, the entity with the highest probability is output, representing the most likely specific medical entity at that position. For the dialogue scenario of patients after OVCF surgery, accurate medical entity recognition is crucial for understanding patients' symptoms and providing precise rehabilitation guidance, and it is the basic guarantee for the present invention to provide professional medical dialogue services.
[0107] Wherein, after step S2, it further includes:
[0108] When the medical entity prediction result does not exist in the local knowledge base of specialist knowledge, the local knowledge base of specialist knowledge is dynamically updated.
[0109] Furthermore, if the predicted entity does not exist or is not perfectly associated in the local knowledge base of orthopedics, the present invention can decide whether to add or correct the entry according to the confidence threshold and the manual review rule, thereby maintaining the sensitivity of the knowledge base to newly emerging or rare terms in postoperative care and forming an entity collection mechanism in the form of "online learning" or "semi-online learning". However, in the actual implementation process, medical staff need to regularly check the system marks and confirm them to maintain the accuracy of the knowledge base.
[0110] S3: The encoder performs context fusion processing on the patient's historical interaction information and the current input information to obtain a comprehensive semantic vector.
[0111] Step S3 needs to save the interaction history of the patient in multiple rounds of conversations, including the questions raised by the patient, the current symptom description, and the key entities predicted in the past, and encode the multi-round historical information together with the current round of input into a context vector. In the OVCF postoperative rehabilitation scenario, patients often need to interact multiple times to describe symptom changes and rehabilitation progress. Since it is necessary to remember the previous conversation content to provide accurate continuous guidance, the present invention performs context fusion in step S3, which is the basis for generating coherent and clinically logical multi-round responses.
[0112] Wherein, step S3 further includes:
[0113] S31: Perform time series storage processing on the patient's historical questions and system responses to obtain a complete conversation history record.
[0114] Through time series storage, the system can track important medical information such as the development trajectory of the patient's symptoms, changes in drug use, and progress of rehabilitation activities. The complete conversation history record forms the patient's longitudinal health data, providing comprehensive background information for the system and enabling it to provide suggestions based on the patient's overall situation rather than a single conversation.
[0115] S32: Calculate the correlation degree of the conversation history record to obtain a historical context segment related to the current question.
[0116] Specifically, since not all historical information is relevant to the current question, in step S32, it is necessary to intelligently screen out the most relevant historical segments. In the context of OVCF postoperative rehabilitation, the selective attention mechanism in step S32 is particularly important for accurately locating the historical information most relevant to the current rehabilitation problem from the long conversation history, improving the pertinence and accuracy of the response.
[0117] S33: Combine and represent the historical context segment and the current input text to obtain preliminarily fused context information.
[0118] The combined representation in step S33 is a fusion that needs to preserve the semantic structure and importance weights of each part of the information. Through fusion, a more comprehensive understanding of the context can be obtained. For example, when the patient simply asks, "Is the pain normal?" the system needs to understand by combining information such as the type of surgery, the number of days after surgery, and the pain description mentioned by the patient before.
[0119] S34: Based on the context information, perform global correlation modeling to obtain a semantic representation that captures long-term dependencies.
[0120] Global correlation modeling can capture the deep semantic connections between the conversation contents at different time points, understand the causal relationships and medical development laws with a long time span. During the OVCF postoperative rehabilitation process, the symptom changes, rehabilitation progress, and complication risks of patients often have their inherent medical laws and correlations, and the system needs to have the ability to understand long-term dependencies. For example, the system needs to correlate the obvious pain in the early postoperative period with the emergence of new discomfort after activities a few weeks later to determine whether it is a normal rehabilitation process or a potential complication. Based on the global correlation modeling in step S34, the system can generate a coherent response that conforms to clinical logic, rather than a simple response based on fragmentary information.
[0121] S35: Integrate the semantic representation and the medical entity prediction result to obtain a comprehensive semantic vector.
[0122] In step S35, by integrating the semantic representation and the medical entity prediction result, the system combines text understanding with medical knowledge to form a comprehensive semantic vector that contains general language understanding and professional medical information.
[0123] S4: Based on the multi-layer medical knowledge base and the medical entity prediction result, use a decoder to generate a reply for the comprehensive semantic vector to obtain a personalized medical dialogue reply.
[0124] In step S4, the present invention combines a medical pre-trained language encoder and a decoder structure to perform generative processing on the multi-turn dialogue context, output a response more targeted to the medical scenario, and the obtained personalized medical dialogue response integrates the multi-layer medical knowledge base, the identified medical entities, and the comprehensive semantic vector constructed in the previous steps to generate the final dialogue response.
[0125] Among them, step S4 further includes:
[0126] S41: Based on the medical entity prediction result, use the decoder to generate an answer word by word for the comprehensive semantic vector to obtain a response sequence.
[0127] In step S41, the decoder generates an answer sequence word by word (or sub-word by sub-word) based on the medical entity prediction result and the entity information output by target entity mining (such as mild pain, suggesting lumbar and back muscle exercise, etc.). The process adopts an autoregressive generation method, and based on the generated words and input information, gradually generates a complete answer.
[0128] S42: Calculate the prediction probability of each word in the response sequence, and based on the prediction probability, generate a response based on the multi-layer medical knowledge base to obtain an initial medical dialogue response.
[0129] Furthermore, step S42 selects the most appropriate words to form a response by calculating the generation probability of each candidate word. The calculation process of step S2 not only follows the general rules of the language model but also incorporates the professional information of the multi-layer medical knowledge base. By querying the triple relationship in the knowledge base, accurate medical information can be introduced when generating a response, such as professional knowledge like "2 - 4 weeks after percutaneous vertebroplasty is the appropriate period for lumbar and back muscle function training". The knowledge base fusion mechanism not only ensures that the initial medical dialogue response output by the system is both fluent and natural but also contains professional and accurate medical information.
[0130] S43: Perform safety control on the initial medical dialogue response to obtain a personalized medical dialogue response.
[0131] Furthermore, the safety control mechanism in step S43 ensures the guarantee of not providing wrong suggestions that may endanger the patient's health, greatly improving the safety and reliability of the system in the actual medical scenario, and ensuring that patients after OVCF surgery can obtain professional and safe rehabilitation guidance.
[0132] Among them, step S43 specifically includes:
[0133] For the content related to drug dosage and postoperative activities in the initial medical dialogue response, call the local specialty knowledge base for matching;
[0134] When the initial medical dialogue reply contains risk information, a prompt is given in the output personalized medical dialogue reply.
[0135] Furthermore, if the generated content involves highly sensitive information such as medication dosage and postoperative exercise intensity, the system will call the local orthopedic knowledge base for precise matching or rule restriction; when it is found that the patient has multiple complications, a long duration of symptoms, or other high-risk situations, the system will prompt in the output answer for timely reexamination or further medical treatment; and this situation can be marked for medical staff for manual follow-up and dynamic maintenance of the knowledge base.
[0136] In addition, when the patient reports the same symptom for a long time, the symptom has not improved, or there are multiple complications in multiple rounds of dialogue, the system will prompt the patient for timely reexamination or further medical treatment in the answer; at the same time, such high-risk cases are marked for key attention by medical staff and subsequent update of the knowledge base.
[0137] As Figure 2 shown, the present invention also provides a dialogue generation system for a medical artificial intelligence system, including:
[0138] A construction module 100: used to fuse general concepts with specialty information to construct a multi-layer medical knowledge base;
[0139] A prediction module 200: used to perform masked prediction on medical terms in the patient's multi-round dialogue corpus through a target entity mining algorithm to obtain a medical entity prediction result;
[0140] A fusion module 300, where the fusion module 300 is configured as an encoder for context fusion processing of the patient's historical interaction information and the current input information to obtain a comprehensive semantic vector;
[0141] A reply module 400, where the reply module 400 is configured as a decoder for generating a reply for the comprehensive semantic vector based on the multi-layer medical knowledge base and the medical entity prediction result to obtain a personalized medical dialogue reply.
[0142] Furthermore, the implementation of the present invention will be introduced below in combination with specific configuration environments and application environments.
[0143] Hardware environment: The server configuration includes a multi-core CPU and an appropriate amount of GPUs (such as NVIDIA A100) or other devices supporting deep learning acceleration for training and inference on large-scale dialogue data; Data storage: A relational database or a graph database is used to store the local orthopedic knowledge base for rapid retrieval, update, and visualization of triple information. To adapt to the dynamic update requirements of manual review, version management and backtracking of newly added entity or relationship entries need to be supported.
[0144] Data Interface: The follow-up system obtains basic information, surgical type, diagnosis results, postoperative examination records, etc. through the patient-facing APP or web terminal; Dialogue Input: Patients can input postoperative symptoms or consultation questions on the platform in text or voice-to-text form.
[0145] Taking a 67-year-old female patient as an example, one week after OVCF surgery, she began to report symptoms of "aching and soreness in the lower back, occasional numbness". Based on the method and system of the present invention, the specific processing flow is as follows.
[0146] S101: Patient input (in text form): "Doctor, I've been feeling a bit [redacted] in my lower back lately, especially after sitting for a long time, and it gets [redacted]. What's going on?"
[0147] S102: Masking Marking and Encoding: Through word segmentation and preliminary entity detection, the suspected postoperative symptom description part is identified and marked with "[redacted]"; the whole paragraph semantics is embedded using a medical pre-trained language encoder to obtain a context representation vector.
[0148] S103: Entity Prediction calculates the most likely entity words, such as "aching and soreness", "numbness", etc.; based on the patient's existing osteoporosis and postoperative wound healing information, symptom descriptions related to "postoperative muscle strain" or "nerve stimulation" are matched from the knowledge base.
[0149] S104: Knowledge Base Update: If "aching and soreness in the lower back" already exists in the orthopedics local knowledge base, relevant information is directly obtained, such as "take appropriate rest, avoid sitting for a long time, and perform lower back muscle function training"; if a new symptom combination form (such as "nerve numbness combined with pain on different sides") is encountered, the system will list it in the pending review list. After regular review by medical staff, if it is confirmed that the entry is a new postoperative manifestation, the corresponding triple entry will be updated or supplemented in the knowledge base.
[0150] S105: Dialogue Generation Combining the patient's current situation, the system outputs personalized suggestions, such as "You may have muscle tension in your lower back caused by long-term sitting posture. You can perform 3 - 5 times of mild stretching training for the lower back muscles every day, and at the same time keep your daily walking distance around 3000 steps; if the numbness persists or the pain intensifies, you need to seek medical attention in time."
[0151] The multi-round dialogue interaction of the method and system based on the present invention is introduced below in combination with specific application scenarios.
[0152] The First Round of Dialogue:
[0153] Patient: "My pain index in the waist is probably 4 points. Can I do some simple activities?"
[0154] System-generated answer: "Based on your situation, you can try mild lumbar and back muscle function training, twice a day, 10 minutes each time, and pay attention to rest."
[0155] Second round of conversation (2 days later):
[0156] Patient: "I feel sore after the training now. What should I do?" The TEM module identifies "soreness" as a key symptom and retrieves that "Increasing lumbar and back muscle endurance training may cause slight soreness", which is a normal phenomenon. If the degree is within the tolerable range, it can continue;
[0157] System answer: "Slight soreness is a common reaction. You can apply hot compress or massage after training, 5 - 10 minutes each time. If the pain increases significantly, seek medical attention in time."
[0158] Third round of conversation (two weeks after surgery):
[0159] Patient: "Can I try to get out of bed and move around for a longer time?"
[0160] The system queries the activity suggestions for the two - week period after surgery in the local orthopedic knowledge base and combines the patient's previous symptom feedback to output the answer: "You can gradually extend the walking or standing time, not exceeding 30 minutes each time, and record the lumbar and back pain level in time. If there is acute pain or numbness, be vigilant about excessive spinal load and seek medical attention in time."
[0161] A dialogue generation method and system for a medical artificial intelligence system provided by the present invention combines a multi - source medical dictionary database and a local knowledge base for after - orthopedic - surgery, can provide richer and more reliable medical background information, and reduce the occurrence of incorrect answers; secondly, through the TEM module, unknown or new terms appearing in the dialogue can be learned and incorporated in real time, and combined with the method of regular manual review to ensure that the knowledge base is synchronized with clinical practice; in addition, the medical pre - trained language encoder is trained on orthopedic corpus to ensure context concatenation, key information extraction and dialogue coherence; furthermore, during the regular follow - up stage, the automatically generated postoperative care suggestions can reduce the pressure of repeated answers from medical staff and at the same time provide convenient and personalized rehabilitation guidance for patients.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A dialogue generation method for a medical artificial intelligence system, characterized in that, Including: S1: Integrate general concepts with specialist knowledge to construct a multi-layer medical knowledge base; S2: Use the target entity mining algorithm to perform masking prediction on medical terms in the multi-round conversation corpus of patients to obtain medical entity prediction results; S3: Use an encoder to perform context fusion processing on the historical interaction information and current input information of the patient to obtain a comprehensive semantic vector; S4: Based on the multi-layer medical knowledge base and the medical entity prediction results, use a decoder to generate a reply for the comprehensive semantic vector to obtain a personalized medical dialogue response.
2. The dialogue generation method for a medical artificial intelligence system according to claim 1, characterized in that, Step S1 further includes: S11: Classify and extract basic concepts according to the medical term library and establish an index to obtain a general medical dictionary database; S12: Extract information from the specialist nursing scenario information to obtain a local knowledge base of specialist knowledge; S13: Store the relationships of the local knowledge base of specialist knowledge through triple coding to obtain a knowledge representation; S14: Integrate the medical dictionary database and the knowledge representation to obtain a multi-layer medical knowledge base.
3. The dialogue generation method for a medical artificial intelligence system according to claim 1, wherein, Step S2 further includes: S21: Scan and screen the dialogue corpus to obtain candidate word phrases representing medical meanings; S22: Replace the candidate word phrases with masking tokens to obtain a text sequence containing masking positions; S23: Semantically encode the text sequence through a pre-trained language encoder to obtain a context hidden layer vector group; S24: Based on the context hidden layer vector group, calculate the probability distribution of the hidden layer vector corresponding to each masking token to obtain the prediction distribution of candidate medical entities; S25: Output the medical entity prediction result with the highest probability according to the prediction distribution.
4. A dialogue generation method for a medical artificial intelligence system according to claim 3, characterized in that, Step S23 further includes: S231: Inject position information into each token in the input text through a position encoding algorithm to obtain a first token representation containing position information; S232: Perform context correlation calculation on the first token representation through a multi-head self-attention mechanism to obtain a second token representation considering global information; S233: Perform non-linear transformation on the second token representation through a feed-forward neural network to obtain the final context hidden layer vector group.
5. A dialogue generation method for a medical artificial intelligence system according to claim 3, characterized in that, The expression of the prediction distribution in step S24 is: P(e|h m ; Θ) = Softmax(W·h m + b); Among them, P(e|h m ; Θ) is the output probability distribution, e is the predicted entity, h m is the m-th hidden layer vector to be predicted, Θ is the set of model parameters, Softmax is the softmax function, W is the first learnable parameter, and b is the second learnable parameter.
6. The dialogue generation method for a medical artificial intelligence system according to claim 1, characterized in that After step S2, it further includes: When the medical entity prediction result does not exist in the local knowledge base of specialist knowledge, dynamically update the local knowledge base of specialist knowledge.
7. A dialogue generation method for a medical artificial intelligence system according to claim 1, characterized in that, Step S3 further includes: S31: Perform time series storage processing on the historical questions and system replies of the patient to obtain a complete dialogue history record; S32: Calculate the correlation degree of the dialogue history record to obtain a historical context fragment related to the current question; S33: Combine and represent the historical context fragment and the current input text to obtain preliminarily fused context information; S34: Based on the context information, perform global correlation modeling to obtain a semantic representation capturing long-term dependence relationships; S35: Integrate the semantic representation and the medical entity prediction result to obtain a comprehensive semantic vector.
8. A dialogue generation method for a medical artificial intelligence system according to claim 1, characterized in that, Step S4 further includes: S41: Based on the medical entity prediction result, use a decoder to generate answers word by word for the comprehensive semantic vector to obtain a response sequence; S42: Calculate the prediction probability of each word in the response sequence, and based on the prediction probability, generate a reply based on a multi-layer medical knowledge base to obtain an initial medical dialogue reply; S43: Perform security control on the initial medical dialogue reply to obtain a personalized medical dialogue reply.
9. A method for generating a dialogue for a medical artificial intelligence system according to claim 8, characterized in that, Step S43 specifically includes: For the content related to drug dosage and postoperative activities in the initial medical dialogue reply, call the local specialty knowledge base for matching; When the initial medical dialogue reply contains risk information, give a prompt in the output personalized medical dialogue reply.
10. A dialogue generation system for a medical artificial intelligence system, which is used to execute a dialogue generation method for a medical artificial intelligence system according to any one of claims 1 to 9, characterized in that, It includes: A construction module: used to fuse general concepts with specialty information to construct a multi-layer medical knowledge base; A prediction module: used to perform masked prediction on medical terms in the patient's multi-round dialogue corpus through a target entity mining algorithm to obtain a medical entity prediction result; A fusion module, which is configured as an encoder, used to perform context fusion processing on the patient's historical interaction information and the current input information to obtain a comprehensive semantic vector; A reply module, which is configured as a decoder, used to generate a reply for the comprehensive semantic vector based on the multi-layer medical knowledge base and the medical entity prediction result to obtain a personalized medical dialogue reply.
Citation Information
Patent Citations
Method for constructing stroke medical knowledge graph
CN112420212A
Dialogue generation method and device, electronic equipment and storage medium
CN116522873A
Data knowledge dual-drive intelligent medical dialogue system and method based on knowledge graph
CN117077786A
Large language model question and answer generation method based on knowledge graph enhancement
CN118227769A
Dialog system answering method based on sentence paraphrase recognition
US20230069935A1