A medical data processing method and system based on a large AI model
The disease trend chart and language model constructed through AI large-scale models automate patient return visits, solving the high efficiency and personalized problems of medical return visits, reducing the burden on medical staff, and improving the efficiency of return visits.
Patent Information
- Application Number
- CN202510056796.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The existing medical follow-up work requires a lot of manpower and material resources, and it is difficult to target the individual follow-up problem of different patients, and it is difficult to build an effective mechanism.
The medical data processing method based on AI big model is adopted, and the disease description information is converted into query vectors. The back-access answer template is constructed using the disease trend chart and language big model, and the return visit is automated and the disease development trend is analyzed.
Automatic follow-up visits and condition analysis for different patients are realized, which reduces the labor force of medical staff, improves the efficiency of follow-up visits, and provides rapid condition trend analysis assistance.
Smart Images

Figure CN119938852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and specifically to a medical data processing method and system based on an AI large model. Background Art
[0002] The hospital's follow-up visits to patients are an important part of the medical service system, which has a positive significance for improving the quality of medical services and the health outcomes of patients. Doctors can monitor the recovery of patients through follow-up visits, provide necessary rehabilitation guidance, help patients better understand their own conditions, and reduce the possibility of disease recurrence. At the same time, follow-up visits are also an important way to collect information on the development of diseases and evaluate the treatment effects, and these data are very valuable for medical research and the formulation of clinical guidelines.
[0003] However, effective follow-up work requires a large amount of manpower, material resources and time. In addition, the injuries and illnesses of different patients are different, and different follow-up questions need to be determined for each patient. Therefore, it is very difficult to build an effective follow-up mechanism. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a medical data processing method and system based on an AI large model to solve the problems in the background art.
[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A medical data processing method based on an AI large model of the present invention includes the steps of:
[0007] Obtain the condition description information of the current patient, wherein the condition description information is extracted from the medical record or the physical examination report;
[0008] Convert the condition description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes various condition description labels, possible symptoms of various condition description labels, and the condition development trends corresponding to the symptoms;
[0009] Construct a follow-up question and answer template based on the possible symptoms of the current patient, and conduct a follow-up visit to the current patient based on the follow-up question and answer template and a called language large model to obtain follow-up data;
[0010] Extract symptom feedback information from the follow-up data, and analyze the condition development trend of the current patient based on the symptom feedback information and the condition trend map to obtain a condition development analysis result.
[0011] In an embodiment of the present application, the construction process of the condition trend map includes:
[0012] Obtain the medical records or physical examination reports of multiple patients, and obtain the review reports of multiple patients;
[0013] Extract the disease description information from the medical records or physical examination reports of the multiple patients; and extract the symptom descriptions from the review reports; and annotate the review reports to obtain a review conclusion, wherein the review conclusion includes improvement, deterioration, and to be observed;
[0014] Classify multiple patients based on the disease description information to obtain multiple patient categories, the symptom descriptions of the multiple patient categories, and the review conclusions; and extract the review conclusions corresponding to each symptom description of each patient category;
[0015] Calculate the probability of each review conclusion of the symptom description , , wherein, is the total number of review conclusions of the symptom description, is the th number of review conclusions of the symptom description;
[0016] Convert the symptom descriptions of multiple patient categories into text vectors;
[0017] Based on multiple patient categories, the text vectors of the symptom descriptions of multiple patient categories, and the probability of each review conclusion of the symptom description Construct a disease trend graph.
[0018] In an embodiment of the present application, the method for extracting the disease description information includes:
[0019] Obtain the medical record text or physical examination report text of the patient;
[0020] Clean and segment the medical record text or physical examination report text to obtain multiple words;
[0021] Perform entity extraction on the multiple words to obtain a first target entity, a second target entity, and a third target entity, wherein the first target entity is a body part, the second target entity is a type of injury or illness, and the third target entity is the severity;
[0022] Take the first target entity, the second target entity, and the third target entity with a distance less than a preset distance threshold as co-occurring entities, and construct structured disease description information based on the co-occurring entities.
[0023] In an embodiment of the present application, construct a follow-up questionnaire template based on the possible symptoms of the current patient, including:
[0024] Classify all possible symptoms based on the disease development trend to obtain the possible symptoms for each disease development trend;
[0025] Rank the possible symptoms for each disease development trend based on probability to obtain the order of the possible symptoms for each disease development trend;
[0026] Construct a question template and an answer template for each possible symptom, and determine the dialogue turn of each question template for each disease development trend based on the order of the possible symptoms for each disease development trend;
[0027] Construct a follow-up Q&A template based on the question template, answer template, and dialogue turn of each question template for each disease development trend.
[0028] In an embodiment of the present application, conduct a follow-up visit to the current patient based on the follow-up Q&A template and the invoked language large model to obtain follow-up data, including:
[0029] S1. Determine the current dialogue turn of the current disease development trend, and send the question template of the current dialogue turn in the follow-up Q&A template to the current patient;
[0030] S2. When receiving the text feedback from the current patient, perform entity extraction on the text feedback, where the text feedback is generated based on the question template of the current dialogue turn;
[0031] S3. Match the extracted entity with the answer template of the current dialogue turn in the follow-up Q&A template;
[0032] S4. When the entity extracted in the current turn matches the answer template of the current dialogue turn in the follow-up Q&A template, and the extracted entity is an affirmative word or a symptom description entity indicating affirmation, add the entity extracted in the current turn to the feedback data, enter the next dialogue turn, and return to step S1 until the dialogue turn ends, where the answer template includes one or more combinations of an affirmative word, a negative word, and a symptom description entity;
[0033] S5. When the entity extracted in the current turn matches the answer template of the current dialogue turn in the follow-up Q&A template, and the extracted entity is a negative word or a symptom description entity indicating negation, accumulate the probability corresponding to the current dialogue turn into the total probability. When the total probability is less than the preset switching probability threshold, enter the next dialogue turn and return to step S1; when the total probability is greater than or equal to the preset switching probability threshold, switch to the dialogue of the next disease development trend;
[0034] S6. When the entity extracted in the current round does not match the answer template for the current conversation round in the callback answer template, and the extracted entity is a symptom description entity, match the symptom description entity with all the answer templates in the callback answer template. If there is a match, skip the corresponding question template in the subsequent conversation rounds, add the entity extracted in the previous round to the feedback data, repeat the current conversation round, and return to step S1. When the number of repetitions exceeds the preset number of times, proceed to the next conversation round and return to step S1. If there is no match, record the symptom description entity, repeat the current conversation round, and return to step S1. When the number of repetitions exceeds the preset number of times, proceed to the next conversation round and return to step S1.
[0035] S7. When the entity extracted in the current round does not match the answer template for the current conversation round in the callback answer template and does not include a symptom description entity, forward the text feedback to the fine-tuned large language model, and guide the current patient based on the fine-tuned large language model until an entity that matches the answer template is extracted from the text feedback of the current patient. Add the extracted entity to the feedback data, proceed to the next conversation round, and return to step S1 until the conversation round ends.
[0036] In one embodiment of the present application, the method for fine-tuning the large language model includes:
[0037] Obtain a guiding conversation template;
[0038] Train the pre-trained large language model based on the guiding conversation template to obtain a fine-tuned large language model.
[0039] In one embodiment of the present application, analyzing the disease development trend of the current patient based on the symptom feedback information and the disease trend graph, and obtaining a disease development analysis result, including:
[0040] Convert each symptom in the symptom feedback information into a second query vector respectively;
[0041] Match the second query vector with the disease trend graph to obtain the probability of the disease development trend of each second query vector;
[0042] Sum the probabilities of each disease development trend, and use the disease development trend with the highest probability as the disease development analysis result.
[0043] In one embodiment of the present application, converting the disease description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector respectively, includes:
[0044] Convert the entities in the disease description information or the entities in the symptom feedback information into word vectors based on a look-up table, and extract the position encodings of the multiple entities based on the exponential function;
[0045] Multiply the word vectors with a pre-constructed first parameter matrix to obtain a word vector matrix; and multiply the position encodings with a pre-constructed second parameter matrix to obtain a position matrix;
[0046] Fuse the word vector matrix and the position matrix to obtain a fusion matrix;
[0047] Multiply the fusion matrix with a pre-constructed third parameter matrix to obtain an encoding result;
[0048] Convert the encoding result into a first query vector or a second query vector through a look-up table method.
[0049] In an embodiment of the present application, it further includes:
[0050] Send the symptom feedback information and the disease development analysis result to the doctor;
[0051] After receiving the confirmation information from the doctor, send the disease development analysis result to the current patient; after receiving the modification information from the doctor, send the modified disease development analysis result to the current patient.
[0052] The present application also provides a medical data processing system based on an AI large model, including:
[0053] An acquisition module, configured to acquire the disease description information of the current patient, wherein the disease description information is extracted from the medical record or the physical examination report;
[0054] A query module, configured to convert the disease description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed disease trend map, wherein the disease trend map includes multiple disease description labels, possible symptoms of multiple disease description labels, and the disease development trends corresponding to the symptoms;
[0055] A follow-up visit module, configured to construct a follow-up visit answer template based on the possible symptoms of the current patient, and conduct a follow-up visit to the current patient based on the follow-up visit answer template and a called language large model to obtain follow-up visit data;
[0056] A data processing module, configured to extract symptom feedback information from the follow-up visit data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0057] The beneficial effects of the present invention are as follows: A medical data processing method and system based on an AI large model of the present invention extract the real-time condition of a patient, find the possible symptoms corresponding to the current condition description information in the condition trend graph, and construct a return visit question-and-answer template based on the possible symptoms to automatically return visit the patient. When making a return visit, considering the complexity of the conversation, a language large model is introduced for assistance to ensure that the symptom feedback information of the patient can be obtained. Based on the symptom feedback information of the patient and the condition trend graph, it can be roughly analyzed whether the current condition trend of the patient is improving, so as to provide rapid analysis assistance for doctors. This application provides an automatic return visit mechanism based on the conversation template and the language large model, which can automatically return visit patients with different conditions and automatically analyze, thereby reducing the return visit workload of medical staff and improving the return visit efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described below in conjunction with the drawings and embodiments:
[0059] Figure 1 is a flowchart of a medical data processing method based on an AI large model shown in an embodiment of the present application;
[0060] Figure 2 is a schematic structural diagram of a condition trend graph in an embodiment of the present application;
[0061] Figure 3 is a schematic diagram of a return visit process in an embodiment of the present application;
[0062] Figure 4 is a structural diagram of a medical data processing system based on an AI large model shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0064] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the layers related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the layers in actual implementation. The type, quantity, and ratio of each layer in actual implementation can be arbitrarily changed, and the layer layout type may also be more complex.
[0065] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details.
[0066] Figure 1 is a flowchart of a medical data processing method based on an AI large model shown in an embodiment of the present application. As Figure 1 shown, a medical data processing method based on an AI large model in this embodiment may include the steps:
[0067] S110, obtain the condition description information of the current patient, where the condition description information is extracted from the medical record or physical examination report;
[0068] This application mainly conducts follow-up visits for injured and sick patients. Since the conditions of different injured and sick patients are different, it is necessary to extract the condition description from the medical record or physical examination report of the most recent time.
[0069] There are significant differences in the condition descriptions of different hospitals and different doctors. If the semantic understanding method is used to summarize the condition, it is relatively difficult. Therefore, this application uses the entity extraction method to extract key entities and uses the key entities to construct a structured condition description for subsequent processing and matching.
[0070] Specifically, the process of obtaining the condition description information of the current patient through entity extraction is as follows:
[0071] S111, obtain the medical record text or physical examination report text of the patient;
[0072] In this embodiment, the medical record and physical examination report of the patient can be in electronic format or paper files. If it is a paper file, it is necessary to use OCR (Optical Character Recognition) to scan the paper file and extract the text information to obtain the medical record text or physical examination report text therein.
[0073] S112, clean and segment the medical record text or physical examination report text to obtain multiple words;
[0074] The main purpose of text cleaning is to remove irrelevant characters such as special symbols, punctuation marks, and line breaks in the text.
[0075] The purpose of word segmentation is to divide the text into multiple words, and then match or recognize the multiple words to verify whether the multiple words are the entities required by this application.
[0076] Among them, the word segmentation process adopts existing word segmentation algorithms, such as the forward maximum matching method, reverse maximum matching method, and bidirectional maximum matching method based on rules; the hidden Markov model and conditional random field based on statistical methods; the support vector machine, neural network model, etc. based on machine learning. This application does not make any restrictions here.
[0077] S113. Extract entities from the multiple words to obtain a first target entity, a second target entity, and a third target entity, where the first target entity is a body part, the second target entity is a type of injury or illness, and the third target entity is the severity level.
[0078] The disease description in this application is a fixed structure, and the structure is: body part + type of injury or illness + severity level. Therefore, the three extracted entities are the body part entity, the type of injury or illness entity, and the severity description entity.
[0079] For example, "sinus + inflammation / chronic inflammation + mild", "left tibia + fracture + mild", etc.
[0080] For entity extraction, a rule-based matching scheme, statistical model algorithms, machine learning methods, etc. can also be used. In this application, when applied in the medical field, the required word entities are relatively few. Therefore, a method of pre-building an entity database and performing matching is used to extract entities.
[0081] The entity templates in the entity database of this application are stored in the form of vectors. When matching, the extracted entities are converted into query vectors, and then matching is performed based on the cosine similarity.
[0082] S114. Use the first target entity, the second target entity, and the third target entity whose distances are less than a preset distance threshold as contemporaneous entities, and construct structured disease description information based on the contemporaneous entities.
[0083] In the medical record, there may be entities in different periods that meet the structural requirements of the disease description information. To avoid extracting entities in different periods and constructing false disease description information, this application screens contemporaneous entities based on the character distance.
[0084] Finally, extract the entities that meet the requirements and construct structured disease description information in the fixed format of "body part + type of injury or illness + severity level".
[0085] Since different diseases may present different symptoms during the recovery period. For example, rib fractures may present symptoms such as pain, difficulty breathing, and fever during the recovery period. If the pain subsides and breathing becomes easier, it indicates improvement. If the pain intensifies or even symptoms such as fever and cough appear, it indicates the deterioration of the condition.
[0086] Based on the above corresponding relationships among the disease conditions, symptoms, and disease development trends, this application constructs a disease trend map in advance based on big data to clarify the relationships among various disease conditions, various symptoms of various disease conditions, and the disease development trends represented by various symptoms, so as to replace doctors in making simple trend judgments.
[0087] Specifically, the construction process of the disease trend map is as follows:
[0088] (1) Obtain the medical records or physical examination reports of multiple patients, and obtain the review reports of multiple patients;
[0089] In this application, since it is necessary to extract the symptoms and review conclusions of patients during the recovery period, the medical records, physical examination reports, and review reports of patients with review reports need to be used as basic data.
[0090] Since there are numerical changes in various physical indicators in the physical examination report and usually few symptom descriptions, symptoms can be extracted from the medical records, or experienced medical staff can annotate the physical examination report to obtain the possible symptoms corresponding to the physical examination values.
[0091] In addition, if the patient is hospitalized, symptom description information can also be extracted from the nursing records and inpatient medical record reports.
[0092] (2) Extract disease description information from the medical records or physical examination reports of the multiple patients; extract symptom descriptions from the review reports; and annotate the review reports to obtain review conclusions, where the review conclusions include improvement, deterioration, and to be observed;
[0093] The extraction of disease description information is as described above and will not be elaborated here.
[0094] The extraction of symptom descriptions is also based on target entities. Generally speaking, symptom descriptions are also composed of a body part entity + a symptom entity. Therefore, the following process can be adopted for extraction:
[0095] (2-1) Clean and segment the text in the medical record or physical examination report to obtain multiple words;
[0096] (2-2) Convert the multiple words into query vectors;
[0097] (2-3) Use the entity template library to match the query vectors. The matching method is cosine similarity, and the matched entities are used as target entities;
[0098] (2-4) Construct symptom descriptions based on the target entities, such as "abdomen + pain".
[0099] (3) Classify multiple patients based on the described illness information to obtain multiple patient categories, symptom descriptions of multiple patient categories, and reexamination conclusions; and extract the reexamination conclusions corresponding to each symptom description of each patient category.
[0100] After obtaining the basic data, first classify the basic data based on the described illness. Thus, multiple patients with the same described illness are grouped in one data unit.
[0101] Then, reclassify the data unit, and group the basic data with the same symptom description in one data sub-unit, so as to find out which reexamination conclusions exist for the same symptom description corresponding to the same described illness. Since even for the same described illness, different people may have different symptoms due to different physiques, corresponding to different reexamination conclusions. Therefore, in the constructed data sub-units, there will be multiple reexamination conclusions for the same symptom description.
[0102] (4) Calculate the probability of each reexamination conclusion of the symptom description , , where is the total number of reexamination conclusions of the symptom description, is the number of the -th reexamination conclusion of the symptom description;
[0103] For different reexamination conclusions, the present application calculates their probabilities. For example, for rib - fracture - severe, the corresponding symptom is dyspnea, the probability of improvement is 10%, the probability of deterioration is 70%, and the probability of pending observation is 20%.
[0104] (5) Convert the symptom descriptions of multiple patient categories into text vectors;
[0105] (6) Based on multiple patient categories, text vectors of symptom descriptions of multiple patient categories, and the probability of each reexamination conclusion of the symptom description construct a disease trend graph.
[0106] Finally, for the convenience of subsequent matching with symptom description information, the present application converts the symptom descriptions of multiple patient categories in the database into text vectors. Then, construct a disease trend graph based on a tree structure.
[0107] Figure 2 is a schematic structural diagram of the disease trend graph in an embodiment of the present application, and the constructed disease trend graph is as Figure 2 shown.
[0108] S120. Convert the disease description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed disease trend map, where the disease trend map includes multiple disease description tags, possible symptoms of multiple disease description tags, and the disease development trends corresponding to the symptoms;
[0109] After constructing the disease trend map, it can be used to determine the disease description information and possible symptoms of the current patient. Here, the symptoms refer to various symptoms that may occur during improvement, deterioration, and maintenance during the recovery period.
[0110] S130. Construct a follow-up questionnaire template based on the possible symptoms of the current patient, and conduct a follow-up visit to the current patient based on the follow-up questionnaire template and the called large language model to obtain follow-up data;
[0111] In this application, a follow-up questionnaire template is established based on the possible symptoms of the current patient, aiming to obtain whether the current patient has the symptoms in the question through inquiry.
[0112] The follow-up questionnaire template is constructed according to the principle of the highest efficiency, that is, the follow-up visit is completed in the shortest time to avoid the antipathy of the patient. The process is as follows:
[0113] S131. Classify all possible symptoms based on the disease development trend to obtain the possible symptoms of each disease development trend;
[0114] For example, if the disease description of the current patient is "sinus + inflammation / chronic inflammation + mild", the corresponding improvement symptoms are: "nose + normal breathing", "upper respiratory tract + reduced secretions", etc., and the corresponding worsening symptoms are: "head - pain", "upper respiratory tract + bloody secretions", etc.;
[0115] S132. Sort the possible symptoms of each disease development trend based on probability to obtain the order of the possible symptoms of each disease development trend;
[0116] Since the probabilities of different symptoms occurring in each disease development trend are different, in order to quickly obtain the positive / negative feedback of the patient when asking questions, the symptoms with higher probabilities are placed in the front and the symptoms with lower probabilities are placed in the back.
[0117] S133. Construct a question template and an answer template for each possible symptom, and determine the dialogue turns of each question template for each disease development trend based on the order of the possible symptoms of each disease development trend;
[0118] After completing the sorting, the corresponding question template and answer template can be constructed.
[0119] Both the question template and the answer template are constructed based on symptom information. For example, if the symptoms in sequence A are "nose + normal breathing", the constructed question template is "Has your nose become breathing smoothly recently?", and the corresponding answer templates are "Yes", "No", "Increased nasal discharge", etc., to correspond to various answers from the other party.
[0120] Multiple question templates with the same disease development trend are sorted according to probability, so as to conduct follow-up visits with users according to the probability size.
[0121] S134. Construct a follow-up question and answer template based on the question template, answer template for each disease development trend, and the number of dialogue turns for each question template.
[0122] Finally, based on the question templates, answer templates for three independent disease development trends, and the number of dialogue turns for each question template, three groups of question and answer templates are respectively constructed.
[0123] In this application, automatic follow-up visits are generally performed when the patient is in the recovery period. For example, an online follow-up visit process is automatically triggered 10 days after the patient goes home.
[0124] During the follow-up visit, first, an opening statement will be made according to the personal information address of the patient. For example:
[0125] "XX, hello, I am the follow-up robot. Now I am conducting a follow-up visit to you. Please touch X on the screen to start the follow-up visit."
[0126] When receiving the feedback generated by touching the screen, the follow-up visit will start. Figure 3 This is a schematic diagram of the follow-up visit process in an embodiment of this application, as Figure 3 shown. The specific process is as follows:
[0127] S1. Determine the current dialogue turn of the current disease development trend, and send the question template of the current dialogue turn in the follow-up question and answer template to the current patient.
[0128] In this application, the default trend is the improvement of the disease. Therefore, the question and answer template sequence corresponding to the improvement trend is selected to perform the follow-up visit. At the beginning, the question template of the first round is sent to the user.
[0129] If it is not the first round, the question template of the corresponding round is sent to the user.
[0130] For example: "Has the pain under your ribs improved?"
[0131] S2. When receiving the text feedback from the current patient, perform entity extraction on the text feedback, where the text feedback is generated based on the question template of the current dialogue turn.
[0132] After sending the question template, wait for text feedback from the current patient. If there is text feedback, extract the entities in the text feedback. If there is no feedback after waiting for more than T, end the follow-up visit and mark it as unresponsive.
[0133] The purpose of entity extraction is to find the content related to the answer template, so as to judge the feedback of the current patient.
[0134] For example, the feedback text of the current patient is "No, the pain under the ribs has worsened recently";
[0135] At this time, extract the target entities "No", "under the ribs", and "pain has worsened".
[0136] S3, match the extracted entities with the answer template of the current dialogue turn in the follow-up visit answer template;
[0137] In this process, the extracted entities are matched by the matching method. The entity vectors in the template library are used to match the feedback text to extract the target entities. The matching criterion is still the cosine similarity.
[0138] S4, when the entities extracted in the current turn match the answer template of the current dialogue turn in the follow-up visit answer template, and the extracted entities are affirmative words or symptom description entities indicating affirmation, add the entities extracted in the current turn to the feedback data, enter the next dialogue turn and return to step S1 until the dialogue turn ends, where the answer template includes one or more combinations of affirmative words, negative words, and symptom description entities;
[0139] If the extracted entities match the answer template of the current turn and indicate affirmation. It means that the intention of this round of dialogue is confirmed as positive, and the affirmative entities and the symptoms of the current turn are added to the feedback data. As the feedback information of the current turn.
[0140] At this time, directly enter the next round of dialogue of the current disease development trend until the follow-up visit of the current disease development trend is completed.
[0141] Specifically, the symptom description entity indicating affirmation is matched with the symptom description entity vector in the answer template of the current dialogue turn and does not contain negative expressions. For example, if the symptom description entity vector in the answer template of the current turn contains "under the ribs" and "pain", and the feedback text contains the entities "under the ribs" and "pain has worsened", it means affirmation.
[0142] S5. When the entity extracted in the current round matches the answer template for the current conversation round in the follow-up visit answer template, and the extracted entity is a negative word or a symptom description entity indicating negation, accumulate the probability corresponding to the current conversation round into the total probability. When the total probability is less than the preset switching probability threshold, proceed to the next conversation round and return to step S1; when the total probability is greater than or equal to the preset switching probability threshold, switch to the conversation for the next disease development trend;
[0143] If the extracted entity matches the answer template for the current round and indicates negation, it means that the intention of this round of conversation is confirmed as negative. Discard the entity for the current round. At the same time, accumulate the symptom probability for the current round into the total probability in.
[0144] When the total probability is less than the preset probability threshold, directly proceed to the next round of conversation for the current disease development trend. If the current patient has denied multiple times, resulting in the total probability being greater than or equal to the preset probability threshold, then it is basically possible to abandon the current disease development trend. Therefore, jump to the next disease development trend, such as the disease deterioration trend.
[0145] In addition, if during the follow-up visit for the disease deterioration trend, the total probability is also greater than or equal to the preset probability threshold. Then directly mark the disease development trend of the current patient as "to be observed".
[0146] Since this application sorts according to the probability size, the questions earlier in the order correspond to greater probabilities. Generally, when the patient denies after answering 2 - 3 questions, it will automatically switch to another disease development trend to avoid the patient's negative emotions.
[0147] S6. When the entity extracted in the current round does not match the answer template for the current conversation round in the follow-up visit answer template, and the extracted entity is a symptom description entity, match the symptom description entity with all the answer templates in the follow-up visit answer template. If there is a match, skip the corresponding question template in the subsequent conversation rounds, add the entity extracted in the previous round to the feedback data, and repeat the current conversation round and return to step S1 until the number of repetitions exceeds the preset number of times, then proceed to the next conversation round and return to step S1; if there is no match, record the symptom description entity, and repeat the current conversation round and return to step S1 until the number of repetitions exceeds the preset number of times, then proceed to the next conversation round and return to step S1;
[0148] If the text feedback by the current patient mentions a new symptom description, perform a global match.
[0149] For example, the question template is "Has the pain under your ribs improved?";
[0150] If the feedback text is: "My breathing has not been very smooth recently", then through word segmentation and template matching, the entities "breathing", "not", and "smooth" are extracted.
[0151] This feedback does not correspond to the current question, but may appear in subsequent conversations. At this time, match the question templates of other rounds with the extracted entities. If a match is found, record the feedback entities corresponding to the questions in the background for the corresponding rounds. At the same time, by means of re-prompting, guide the patient back to the question of the current round. After multiple guidances, when the patient still does not answer the answer related to the current round, it means that the patient deliberately avoids this question. At this time, skip this round and enter the next round of conversation.
[0152] S7. When the entity extracted in the current round does not match the answer template of the current conversation round in the return visit answer template and does not contain symptom description entities, forward the text feedback to the fine-tuned large language model, and guide the current patient based on the fine-tuned large language model until an entity matching the answer template is extracted from the text feedback of the current patient, add the extracted entity to the feedback data, enter the next conversation round and return to step S1 until the conversation round ends.
[0153] If no text related to the current round of conversation or symptom description is extracted from the text feedback by the patient. Then, in the case of not being able to understand the user's intention, call the large language model to have a conversation with the user. Guide the user back to the conversation of the current round.
[0154] The large language model adopted in this application is a fine-tuned large language model, and relevant corpora are used to adjust the pre-trained large model. So that the large model can understand the user's intention and positively answer the user's question, and then prompt the user to return to the conversation of the current round. Avoid situations where it is not flexible enough and the user experience is poor.
[0155] The fine-tuning method of the large language model includes:
[0156] Obtain a guiding conversation template;
[0157] Train the pre-trained large language model based on the guiding conversation template to obtain a fine-tuned large language model.
[0158] In the guiding dialogue template of this application, the question template can be set arbitrarily because after entity extraction in the previous text, it has been possible to confirm that the patient's feedback has nothing to do with the current round of dialogue. Therefore, only a fixed rule needs to be set so that after the large model answers the user's question positively, the guiding process can be executed. That is, ensure that there are examples in the training dataset marked with the correct round and question template, which can help the model learn how to identify and return to the correct dialogue path.
[0159] S140, extract the symptom feedback information from the follow-up data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0160] Finally, after obtaining all the feedback entities with positive feedback, combine them with the disease trend map for analysis to obtain a disease development analysis result, including:
[0161] Convert each symptom in the symptom feedback information into a second query vector;
[0162] Match the second query vector with the disease trend map to obtain the probability of the disease development trend of each second query vector;
[0163] Sum the probabilities of each disease development trend, and take the disease development trend with the highest probability as the disease development analysis result.
[0164] Since in the dialogue process described above, feedback entities corresponding to multiple disease trends may be extracted. Therefore, this application sums the probabilities of the feedback entities corresponding to each disease trend, and takes the disease development trend with the highest probability as the disease development analysis result.
[0165] In addition, the symptom feedback information and the disease development analysis result are also sent to the doctor;
[0166] After receiving the confirmation information from the doctor, send the disease development analysis result to the current patient; after receiving the modification information from the doctor, send the modified disease development analysis result to the current patient.
[0167] Thus, it assists the doctor in performing the follow-up work, and at the same time, with the doctor's confirmation, it can ensure that the analysis of the disease development trend of the current patient is not wrong. Greatly improving the efficiency of hospital follow-up.
[0168] In an embodiment of this application, converting the disease description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector, includes:
[0169] (1) Convert the entities in the condition description information or the entities in the symptom feedback information into word vectors based on look-up tables, and extract the position encodings of the multiple entities based on the exponential function;
[0170] (2) In this embodiment, the word vectors are converted by means of look-up table mapping, such as One-Hot encoding. The position encodings are used to mark the position of each word by using the exponential function.
[0171] (3) Multiply the word vectors with a pre-constructed first parameter matrix to obtain a word vector matrix; and multiply the position encodings with a pre-constructed second parameter matrix to obtain a position matrix;
[0172] (4) Fuse the word vector matrix and the position matrix to obtain a fusion matrix;
[0173] (5) Multiply the fusion matrix with a pre-constructed third parameter matrix to obtain an encoding result;
[0174] (6) Convert the encoding result into a first query vector or a second query vector by means of look-up table.
[0175] In this application, a first parameter matrix W1 is used to fuse with the word vectors, and a second parameter matrix W2 is used to fuse with the position encodings. Finally, the obtained word vector matrix and position matrix are fused (which can be added) to obtain a fusion matrix. The fusion matrix is multiplied with a third parameter matrix W3 to obtain an encoding result. The encoding result contains both the word vectors of each word and the position information of each word. Therefore, it can retain the semantic information of the question text. Finally, through the way of look-up table mapping, the encoding result is converted into a query vector, so as to convert the semantic information into a vector that can be recognized by the system.
[0176] A medical data processing method based on an AI large model of the present invention extracts the real-time condition of a patient, finds the possible symptoms corresponding to the current condition description information in the condition trend graph, and constructs a return visit question-and-answer template based on the possible symptoms to automatically return visit the patient. During the return visit, considering the complexity of the conversation, a language large model is introduced for assistance to ensure that the symptom feedback information of the patient can be obtained. Based on the symptom feedback information of the patient and the condition trend graph, it can be roughly analyzed whether the current condition trend of the patient is getting better, so as to provide quick analysis assistance for doctors. This application provides an automatic return visit mechanism based on a conversation template and a language large model, which can automatically return visit patients with different conditions and automatically analyze, thereby reducing the return visit workload of medical staff and improving the return visit efficiency.
[0177] As Figure 4 shown, this application also provides a medical data processing system based on an AI large model, including:
[0178] An acquisition module, configured to acquire the condition description information of the current patient, wherein the condition description information is extracted from a medical record or a physical examination report;
[0179] A query module, configured to convert the condition description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes various condition description tags, possible symptoms of various condition description tags, and the condition development trends corresponding to the symptoms;
[0180] A follow-up visit module, configured to construct a follow-up visit answer template based on the possible symptoms of the current patient, and conduct a follow-up visit to the current patient based on the follow-up visit answer template and a called language large model to obtain follow-up visit data;
[0181] A data processing module, configured to extract symptom feedback information from the follow-up visit data, and analyze the condition development trend of the current patient based on the symptom feedback information and the condition trend map to obtain a condition development analysis result.
[0182] A medical data processing system based on an AI large model according to the present invention extracts the real-time condition of a patient, finds the possible symptoms corresponding to the current condition description information in a condition trend map, and constructs a follow-up visit answer template based on the possible symptoms to automatically follow up the patient. During the follow-up visit, considering the complexity of the conversation, a language large model is introduced for assistance to ensure that symptom feedback information of the patient can be obtained. Based on the symptom feedback information of the patient and the condition trend map, it can be roughly analyzed whether the current condition trend of the patient is improving, thereby providing rapid analysis assistance for doctors. This application provides an automatic follow-up mechanism based on a conversation template and a language large model, which can automatically follow up and analyze patients with different conditions, thereby reducing the follow-up workload of medical staff and improving the follow-up efficiency.
[0183] This embodiment further provides an electronic terminal, including: a processor and a memory;
[0184] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any method in this embodiment.
[0185] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0186] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used for communication, and the processor and the transceiver are used to run the computer programs to enable the electronic terminal to execute each step of the above method.
[0187] In this embodiment, the memory may include a Random Access Memory (RAM) and may also include a non-volatile memory, such as at least one disk memory.
[0188] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0189] In the above embodiments, although the present invention has been described in conjunction with specific embodiments of the present invention, according to the previous description, many substitutions, modifications, and variations of these embodiments will be obvious to those of ordinary skill in the art. The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.
[0190] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A medical data processing method based on the AI large model, characterized in that, Including the steps: Obtain the description information of the current patient's condition, where the description information of the condition is extracted from the medical record or physical examination report; Convert the described condition information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed disease trend map, where the disease trend map includes multiple disease description tags, possible symptoms of multiple disease description tags, and the disease development trends corresponding to the symptoms; the construction process of the disease trend map includes: obtaining the medical records or physical examination reports of multiple patients, and obtaining the follow-up reports of multiple patients; extracting the described condition information from the medical records or physical examination reports of the multiple patients; and extracting the symptom descriptions from the follow-up reports; and annotating the follow-up reports to obtain a follow-up conclusion, where the follow-up conclusion includes improvement, deterioration, and to be observed; classifying the multiple patients based on the described condition information to obtain multiple patient categories, the symptom descriptions of the multiple patient categories, and the follow-up conclusions; and extracting the follow-up conclusions corresponding to each symptom description of each patient category; calculating the probability of each follow-up conclusion of the symptom description , , where is the total number of follow-up conclusions of the symptom description, is the number of the th follow-up conclusion of the symptom description; convert the symptom descriptions of the multiple patient categories into text vectors; based on the multiple patient categories, the text vectors of the symptom descriptions of the multiple patient categories, and the probability of each follow-up conclusion of the symptom description construct a disease trend map; Construct a follow-up questionnaire template based on the possible symptoms of the current patient, and conduct a follow-up visit to the current patient based on the follow-up questionnaire template and the called language model to obtain follow-up data; Extract symptom feedback information from the follow-up data, and analyze the development trend of the current patient's condition based on the symptom feedback information and the condition trend graph to obtain the analysis result of the condition development.
2. The medical data processing method based on the AI large model according to claim 1, wherein The method for extracting the description information of the condition includes: Obtain the text of the patient's medical record or physical examination report; Clean and segment the text of the medical record or physical examination report to obtain multiple words; Extract entities from the multiple words to obtain a first target entity, a second target entity, and a third target entity, where the first target entity is a body part, the second target entity is a type of injury or disease, and the third target entity is the severity; Use the first target entity, the second target entity, and the third target entity with a distance less than the preset distance threshold as synchronous entities, and construct structured description information of the condition based on the synchronous entities.
3. The medical data processing method based on the AI large model according to claim 1, characterized in that, Construct a follow-up questionnaire template based on the possible symptoms of the current patient, including: Classify all possible symptoms according to the development trend of the condition to obtain the possible symptoms of each development trend of the condition; Sort the possible symptoms of each development trend of the condition based on probability to obtain the order of the possible symptoms of each development trend of the condition; Construct a question template and an answer template for each possible symptom, and determine the dialogue turn of each question template of each development trend of the condition based on the order of the possible symptoms of each development trend of the condition; Construct a follow-up questionnaire template based on the question templates, answer templates, and dialogue turns of each question template of each development trend of the condition.
4. A medical data processing method based on an AI large model according to claim 3, characterized in that, Conduct a follow-up visit to the current patient based on the follow-up questionnaire template and the called language model to obtain follow-up data, including: S1. Determine the current dialogue turn of the current development trend of the condition, and send the question template of the current dialogue turn in the follow-up questionnaire template to the current patient; S2. When receiving the text feedback from the current patient, extract entities from the text feedback, where the text feedback is generated based on the question template of the current dialogue turn; S3. Match the extracted entities with the answer template of the current dialogue turn in the follow-up questionnaire template; S4. When the entities extracted in the current turn match the answer template of the current dialogue turn in the follow-up questionnaire template, and the extracted entities are affirmative words or symptom description entities indicating affirmation, add the entities extracted in the current turn to the feedback data, enter the next dialogue turn, and return to step S1 until the dialogue turn ends, where the answer template includes one or more combinations of affirmative words, negative words, and symptom description entities. S5. When the entity extracted in the current round matches the response template of the current conversation round in the callback response template, and the extracted entity is a negative word or a symptom description entity indicating negation, accumulate the probability corresponding to the current conversation round into the total probability. When the total probability is less than the preset switching probability threshold, proceed to the next conversation round and return to step S1; when the total probability is greater than or equal to the preset switching probability threshold, switch to the conversation of the next disease development trend; S6. When the entity extracted in the current round does not match the response template of the current conversation round in the callback response template, and the extracted entity is a symptom description entity, match the symptom description entity with all response templates in the callback response template. If there is a match, skip the corresponding question template in the subsequent conversation rounds, add the entity extracted in the previous round to the feedback data, and repeat the current conversation round and return to step S1 until the number of repetitions exceeds the preset number, then proceed to the next conversation round and return to step S1; if there is no match, record the symptom description entity, and repeat the current conversation round and return to step S1 until the number of repetitions exceeds the preset number, then proceed to the next conversation round and return to step S1; S7. When the entity extracted in the current round does not match the response template of the current conversation round in the callback response template and does not include a symptom description entity, forward the text feedback to the fine-tuned large language model, and guide the current patient based on the fine-tuned large language model until an entity matching the response template is extracted from the text feedback of the current patient, add the extracted entity to the feedback data, proceed to the next conversation round and return to step S1 until the conversation round ends.
5. A medical data processing method based on an AI large model according to claim 4, characterized in that The fine-tuning method of the large language model includes: Obtain a guiding conversation template; Train the pre-trained large language model based on the guiding conversation template to obtain a fine-tuned large language model.
6. The medical data processing method based on the AI large model according to claim 1, wherein Analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend graph, and the disease development analysis result includes: Convert each symptom in the symptom feedback information into a second query vector respectively; Match the second query vector with the disease trend graph to obtain the probability of the disease development trend of each second query vector; Sum the probabilities of each disease development trend, and use the disease development trend with the highest probability as the disease development analysis result.
7. A medical data processing method based on an AI large model according to claim 6, characterized in that Converting the disease description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector respectively, includes: Convert the entity in the disease description information or the entity in the symptom feedback information into a word vector based on looking up a table, and extract the position encoding of multiple entities based on the exponential function; Multiply the word vector by a pre-constructed first parameter matrix to obtain a word vector matrix; and multiply the position encoding by a pre-constructed second parameter matrix to obtain a position matrix; Fuse the word vector matrix and the position matrix to obtain a fusion matrix; Multiply the fusion matrix with a pre-constructed third parameter matrix to obtain a coding result; Convert the coding result into a first query vector or a second query vector by means of table lookup.
8. A medical data processing method based on an AI large model according to claim 1, characterized in that, It further includes: Send the symptom feedback information and the disease development analysis result to the doctor; After receiving the confirmation information from the doctor, send the disease development analysis result to the current patient; After receiving the modification information from the doctor, send the modified disease development analysis result to the current patient.
9. A medical data processing system based on an AI large model, characterized in that, It includes: An acquisition module, configured to acquire the disease description information of the current patient, wherein the disease description information is extracted from the medical record or the physical examination report; A query module for converting the disease description information into a first query vector, and determining the possible symptoms of the current patient based on the first query vector and a pre-constructed disease trend map, where the disease trend map includes multiple disease description tags, possible symptoms of multiple disease description tags, and the disease development trends corresponding to the symptoms; the construction process of the disease trend map includes: obtaining medical records or physical examination reports of multiple patients, and obtaining follow-up reports of multiple patients; extracting disease description information from the medical records or physical examination reports of the multiple patients; and extracting symptom descriptions from the follow-up reports; and annotating the follow-up reports to obtain follow-up conclusions, where the follow-up conclusions include improvement, deterioration, and to be observed; classifying multiple patients based on the disease description information to obtain multiple patient categories, symptom descriptions of the multiple patient categories, and follow-up conclusions; and extracting the follow-up conclusions corresponding to each symptom description of each patient category; calculating the probability of each follow-up conclusion of the symptom description , , where, is the total number of follow-up conclusions of the symptom description, is the number of the th follow-up conclusion of the symptom description; converting the symptom descriptions of multiple patient categories into text vectors; constructing a disease trend map based on multiple patient categories, the text vectors of the symptom descriptions of the multiple patient categories, and the probability of each follow-up conclusion of the symptom description A follow-up visit module, configured to construct a follow-up visit answer template based on the possible symptoms of the current patient, and conduct a follow-up visit on the current patient based on the follow-up visit answer template and the called language large model to obtain follow-up visit data; A data processing module, configured to extract symptom feedback information from the follow-up visit data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend graph to obtain a disease development analysis result.
Citation Information
Patent Citations
Medical AI assistant implementation method and system based on data driving and large model
CN118098585A