Medical data processing method and system based on AI large model
Through the medical data processing method based on AI large-scale models, automated follow-up visits and analysis of the development trends of the disease are solved, and the problem of large manpower and time investment in medical follow-up work is improved, and the return visit efficiency is provided and rapid analysis assistance is provided.
Patent Information
- Application Number
- CN202510056796.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Effective medical follow-up work requires a lot of manpower, material resources and time. Since each patient has different injuries and conditions, it is necessary to determine different follow-up problems for each patient, and it is very difficult to build an effective follow-up mechanism.
The medical data processing method based on AI big model is adopted to obtain the patient's condition description information, convert it into a query vector, and determine possible symptoms based on the pre-constructed condition trend chart, build a back-visit template, and use the language big model to perform automatic return visits, extract symptom feedback information, and analyze the development trend of the disease.
Automatic follow-up visits and condition analysis for patients with different conditions have been realized, which has reduced the labor of medical staff, improved the efficiency of follow-up, and provided doctors with rapid analysis assistance.
Smart Images

Figure CN119938852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and specifically to a medical data processing method and system based on an AI big model. Background Art
[0002] Hospital follow-up visits to patients are an important part of the medical service system, and they have a positive impact on improving the quality of medical services and the health outcomes of patients. Doctors can monitor the patient's recovery through follow-up visits, provide necessary rehabilitation guidance, help patients better understand their own conditions, and reduce the possibility of disease recurrence. At the same time, follow-up visits are also an important way to collect information on disease development and evaluate treatment effects. These data are very valuable for medical research and the formulation of clinical guidelines.
[0003] However, effective follow-up work requires a lot of manpower, material resources and time. In addition, different patients have different injuries and conditions, and different follow-up questions need to be determined for each patient, so it is very difficult to build an effective follow-up mechanism. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a medical data processing method and system based on an AI big model to solve the problems in the background technology.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A medical data processing method based on an AI big model of the present invention comprises the following steps: Obtaining the current patient's condition description information, wherein the condition description information is extracted from a medical record or a physical examination report; Converting the condition description information into a first query vector, and determining possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; Constructing a questionnaire answer template based on the possible symptoms of the current patient, and performing a questionnaire answer template and the called language macro model to interview the current patient to obtain questionnaire data; Symptom feedback information is extracted from the follow-up data, and the disease development trend of the current patient is analyzed based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0006] In one embodiment of the present application, the process of constructing the disease trend map includes: Obtain medical records or physical examination reports of multiple patients, and obtain reexamination reports of multiple patients; Extracting condition description information from the medical records or physical examination reports of the multiple patients; extracting symptom descriptions from the review reports; and marking the review reports to obtain review conclusions, wherein the review conclusions include improvement, deterioration, and to be observed; Based on the condition description information, multiple patients are classified to obtain multiple patient types and symptom descriptions and review conclusions of the multiple patient types; and the review conclusion corresponding to each symptom description of each patient type is extracted; Calculate the probability of each review conclusion for the symptom description , ,in, The total number of review conclusions for symptom descriptions, The first The number of review conclusions; Convert symptom descriptions of multiple patient categories into text vectors; The probability of each review conclusion based on multiple patient categories, text vectors of symptom descriptions of multiple patient categories, and symptom descriptions Construct a disease trend map.
[0007] In one embodiment of the present application, the method for extracting the condition description information includes: Obtain the patient's medical records or physical examination report text; Cleaning and segmenting the medical record text or physical examination report text to obtain multiple words; Perform entity extraction on the multiple words to obtain a first target entity, a second target entity, and a third target entity, wherein the first target entity is a body part, the second target entity is an injury type, and the third target entity is a severity level; The first target entity, the second target entity and the third target entity whose distances are less than a preset distance threshold are taken as concurrent entities, and structured disease description information is constructed based on the concurrent entities.
[0008] In one embodiment of the present application, a questionnaire answer template is constructed based on the possible symptoms of the current patient, including: Classify all possible symptoms based on the development trend of the disease, and obtain the possible symptoms of each development trend of the disease; The possible symptoms of each disease development trend are sorted based on probability to obtain the order of the possible symptoms of each disease development trend; Constructing a question template and an answer template for each possible symptom, and determining the conversation turn of each question template for each disease development trend based on the order of possible symptoms of each disease development trend; A questionnaire answering template is constructed based on the question template, answer template and conversation rounds of each question template for each disease development trend.
[0009] In one embodiment of the present application, the current patient is visited based on the visit answer template and the called language macro model to obtain visit data, including: S1, determining the current dialogue round of the current disease development trend, and sending the question template of the current dialogue round in the response and answer template to the current patient; S2, when receiving text feedback from the current patient, performing entity extraction on the text feedback, wherein the text feedback is generated based on the question template of the current dialogue round; S3, matching the extracted entity with the answer template of the current dialogue round in the reply answer template; S4, when the entity extracted in the current round matches the answer template of the current dialogue round in the reply answer template, and the extracted entity is an affirmative word or an affirmative symptom description entity, the entity extracted in the current round is added to the feedback data, the next dialogue round is entered and the process returns to step S1 until the dialogue round ends, wherein the answer template includes one or a combination of affirmative words, negative words, and symptom description entities; S5, when the entity extracted in the current round matches the answer template of the current dialogue round in the response answer template, and the extracted entity is a negative word or a symptom description entity that represents negation, the probability corresponding to the current dialogue round is accumulated into the probability sum, and when the probability sum is less than the preset switching probability threshold, the next dialogue round is entered and the process returns to step S1; when the probability sum is greater than or equal to the preset switching probability threshold, the process switches to the next dialogue on the development trend of the disease; S6, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and the extracted entity is a symptom description entity, the symptom description entity is matched with all the answer templates in the reply answer template. If they match, the corresponding question template is skipped in the subsequent dialogue round, the entity extracted in the previous round is added to the feedback data, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; if they do not match, the symptom description entity is recorded, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; S7, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and does not contain the symptom description entity, the text feedback is forwarded to the fine-tuned language model, and the current patient is guided based on the fine-tuned language model until an entity matching the answer template is extracted from the text feedback of the current patient, the extracted entity is added to the feedback data, and the next dialogue round is entered and returns to step S1 until the dialogue round ends.
[0010] In one embodiment of the present application, the method for fine-tuning the large language model includes: Get the guided dialogue template; The pre-trained language model is trained based on the guided dialogue template to obtain a fine-tuned language model.
[0011] In one embodiment of the present application, the disease progression trend of the current patient is analyzed based on the symptom feedback information and the disease progression map to obtain a disease progression analysis result, including: Convert each symptom in the symptom feedback information into a second query vector; Based on matching the second query vector with the disease trend map, obtaining the probability of the disease development trend of each second query vector; The probability of each disease development trend is summed up, and the disease development trend with the greatest probability is taken as the disease development analysis result.
[0012] In one embodiment of the present application, converting the condition description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector respectively, includes: Converting entities in the condition description information or entities in the symptom feedback information into word vectors based on a table lookup, and extracting position codes of the multiple entities based on an exponential function; Multiplying the word vector with a pre-constructed first parameter matrix to obtain a word vector matrix; and multiplying the position code with a pre-constructed second parameter matrix to obtain a position matrix; Fusing the word vector matrix with the position matrix to obtain a fusion matrix; Multiplying the fusion matrix with a pre-constructed third parameter matrix to obtain a coding result; The encoding result is converted into a first query vector or a second query vector by table lookup.
[0013] In one embodiment of the present application, it also includes: Sending the symptom feedback information and the disease progression analysis results to a doctor; After receiving the confirmation information from the doctor, the disease progression analysis result is sent to the current patient; after receiving the modification information from the doctor, the modified disease progression analysis result is sent to the current patient.
[0014] The present application also provides a medical data processing system based on an AI big model, comprising: An acquisition module is used to acquire the condition description information of the current patient, wherein the condition description information is extracted from the medical record or physical examination report; A query module, configured to convert the condition description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; A return visit module, used for constructing a return visit answer template based on the possible symptoms of the current patient, and return visit the current patient based on the return visit answer template and the called language big model to obtain return visit data; The data processing module is used to extract symptom feedback information from the follow-up data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0015] The beneficial effects of the present invention are: a medical data processing method and system based on an AI big model of the present invention extracts the patient's real-time condition, finds the possible symptoms corresponding to the current condition description information in the condition trend map, and constructs a response template based on the possible symptoms to automatically return the patient. During the return visit, considering the complexity of the conversation, a language big model is introduced for assistance to ensure that the patient's symptom feedback information can be obtained. Based on the patient's symptom feedback information and the condition trend map, it can be roughly analyzed whether the patient's current condition trend has improved, thereby providing doctors with rapid analysis assistance. The present application provides an automatic return visit mechanism based on conversation templates and language big models, which can automatically return visits and automatically analyze patients with different conditions, thereby reducing the return visit workload of medical staff and improving return visit efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 It is a flowchart of a medical data processing method based on an AI big model shown in an embodiment of the present application; Figure 2 This is a schematic diagram of the structure of a disease trend graph in one embodiment of the present application; Figure 3 This is a schematic diagram of the return visit process in an embodiment of the present application; Figure 4 It is a structural diagram of a medical data processing system based on an AI big model shown in one embodiment of the present application. DETAILED DESCRIPTION
[0017] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0018] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show the layers related to the present invention rather than being drawn according to the number, shape and size of the layers in actual implementation. In actual implementation, the type, quantity and proportion of each layer may be changed arbitrarily, and the layer layout may also be more complicated.
[0019] In the following description, numerous details are discussed to provide a more thorough explanation of embodiments of the present invention; however, it is apparent to one skilled in the art that embodiments of the present invention may be practiced without these specific details.
[0020] Figure 1 is a flowchart of a medical data processing method based on an AI large model shown in an embodiment of the present application, such as Figure 1 As shown, a medical data processing method based on an AI big model in this embodiment may include the following steps: S110, obtaining the condition description information of the current patient, wherein the condition description information is extracted from the medical record or physical examination report; This application is mainly for follow-up visits to injured and sick patients. Since the conditions of different injured and sick patients are different, it is necessary to extract the condition description from the most recent medical records or physical examination reports.
[0021] There are large differences in the descriptions of medical conditions in different hospitals and by different doctors. It is difficult to summarize the medical conditions using semantic understanding. Therefore, this application uses entity extraction to extract key entities and use key entities to construct a structured description of the medical condition for subsequent processing and matching.
[0022] Specifically, the process of obtaining the current patient's condition description information through entity extraction is as follows: S111, obtaining the patient's medical record text or physical examination report text; In this embodiment, the patient's medical records and physical examination reports can be in electronic format or in paper format. If they are paper files, OCR (Optical Character Recognition) is required to scan the paper files and extract text information to obtain the medical record text or physical examination report text.
[0023] S112, cleaning and segmenting the medical record text or physical examination report text to obtain a plurality of words; The main purpose of text cleaning is to remove irrelevant characters such as special symbols, punctuation marks, line breaks, etc. in the text.
[0024] The purpose of word segmentation is to divide text into multiple words, and then match or identify the multiple words to verify whether the multiple words are the entities required by this application.
[0025] Among them, the word segmentation process adopts existing word segmentation algorithms, such as rule-based forward maximum matching method, reverse maximum matching method, and bidirectional maximum matching method; hidden Markov model and conditional random field based on statistical methods; support vector machine and neural network model based on machine learning, etc. This application does not make any restrictions here.
[0026] S113, performing entity extraction on the multiple words to obtain a first target entity, a second target entity, and a third target entity, wherein the first target entity is a body part, the second target entity is an injury type, and the third target entity is a severity; The disease description in this application is a fixed structure, which is: body part + injury type + severity. Therefore, the three entities extracted are the body part entity, the injury type entity and the severity description entity.
[0027] For example, "sinus+inflammation / chronic inflammation+mild", "left tibia+fracture+mild", etc.
[0028] For entity extraction, rule-based matching schemes, statistical model algorithms, and machine learning methods can also be used. In this application, it is used in the medical field and relatively few word entities are required. Therefore, an entity database is built in advance and matched to extract entities.
[0029] The entity templates in the entity database of the present application are stored in the form of vectors. During matching, the extracted entities are converted into query vectors and then matched based on cosine similarity.
[0030] S114, taking the first target entity, the second target entity and the third target entity whose distances are less than a preset distance threshold as concurrent entities, and constructing structured disease description information based on the concurrent entities.
[0031] In the medical records, there may be entities from different periods that meet the structural requirements of the condition description information. In order to avoid extracting entities from different periods and constructing false condition description information, this application screens entities from the same period based on character distance.
[0032] Finally, the entities that meet the requirements are extracted and structured disease description information is constructed according to the fixed format of "body part + injury type + severity".
[0033] Different conditions may have different symptoms during the recovery period. For example, rib fractures may cause pain, difficulty breathing, fever, etc. during the recovery period. If the pain is relieved and breathing becomes easier, it means that the condition is getting better. If the pain is aggravated, or even fever or coughing occurs, it means that the condition is getting worse.
[0034] Based on the above-mentioned correspondence between disease conditions, symptoms and disease development trends, this application constructs a disease trend map in advance based on big data to clarify the relationship between multiple diseases, multiple symptoms of multiple diseases, and disease development trends represented by multiple symptoms, so as to replace doctors in making simple trend judgments.
[0035] Specifically, the process of constructing the disease trend map is as follows: (1) Obtain medical records or physical examination reports of multiple patients, and obtain reexamination reports of multiple patients; In the present application, since it is necessary to extract the symptoms of the patient during the recovery period and the review conclusions, the medical records, physical examination reports and review reports of the patients with review reports are required as basic data.
[0036] Since there are numerical changes in various physical indicators in the physical examination report, there is usually little description of symptoms. Therefore, symptoms can be extracted from medical records, or experienced medical staff can mark the physical examination report to obtain possible symptoms corresponding to the physical examination values.
[0037] In addition, if the patient is hospitalized, symptom description information can also be extracted from nursing records and hospitalization medical records.
[0038] (2) extracting condition description information from the medical records or physical examination reports of the multiple patients; extracting symptom descriptions from the review reports; and marking the review reports to obtain review conclusions, wherein the review conclusions include improvement, deterioration, and to be observed; The extraction of the condition description information is as described above and will not be repeated here.
[0039] The extraction of symptom description is also based on the target entity. Generally speaking, symptom description is also composed of body part entity + symptom entity. Therefore, the following process can be adopted for extraction: (2-1) Clean and segment the text in the medical record or physical examination report to obtain multiple words; (2-2) Convert multiple words into query vectors; (2-3) Use the entity template library to match the query vector. The matching method is cosine similarity, and the matched entity is used as the target entity. (2-4) Construct symptom descriptions based on the target entity, such as “abdomen + pain”.
[0040] (3) Classifying multiple patients based on the disease description information to obtain multiple patient categories and symptom descriptions and review conclusions for the multiple patient categories; and extracting the review conclusion corresponding to each symptom description of each patient category; After obtaining the basic data, the basic data is first classified based on the description of the condition, so that multiple patients with the same description of the condition are classified into one data unit.
[0041] Then the data unit is divided again, and the basic data with the same symptom description is divided into a data sub-unit, so as to find out which review conclusions will exist for the same symptom description corresponding to the same condition description. Because even for the same condition description, different people will have different symptoms due to their different physical constitutions, and correspondingly different review conclusions will be obtained. Therefore, in the constructed data sub-unit, there will be multiple review conclusions for the same symptom description.
[0042] (4) Calculate the probability of each review conclusion for the symptom description , ,in, The total number of review conclusions for symptom descriptions, The first The number of review conclusions; For different review conclusions, this application calculates their probabilities. For example, rib-fracture-severe, the corresponding symptom is dyspnea, the probability of improvement is 10%, the probability of deterioration is 70%, and the probability of waiting for observation is 20%.
[0043] (5) Convert symptom descriptions of multiple patient types into text vectors; (6) Text vectors based on multiple patient types, symptom descriptions of multiple patient types, and the probability of each review conclusion based on the symptom description Construct a disease trend map.
[0044] Finally, in order to facilitate the subsequent matching of symptom description information, the present application converts the symptom descriptions of multiple patient types in the database into text vectors, and then constructs a disease trend map based on the tree structure.
[0045] Figure 2is a schematic diagram of the structure of the disease trend map in one embodiment of the present application, and the disease trend map constructed is as follows Figure 2 shown.
[0046] S120, converting the condition description information into a first query vector, and determining possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; After the disease trend map is constructed, the disease trend map can be used to determine the current patient's disease description information and possible symptoms. The symptoms here refer to various symptoms that may occur during the recovery period, when the disease gets better, when it gets worse, and when it stays the same.
[0047] S130, constructing a questionnaire answering template based on the possible symptoms of the current patient, and performing a questionnaire answering template and the called language macro model to interview the current patient to obtain questionnaire data; In this application, a questionnaire answer template is established based on the possibility of the current patient, with the purpose of obtaining, through inquiry, whether the current patient has the symptoms in the question.
[0048] The construction of the questionnaire answer template is based on the principle of maximum efficiency, that is, to complete the questionnaire in the shortest time and avoid patients' disgust. The process is as follows: S131, classifying all possible symptoms based on the development trend of the disease, and obtaining possible symptoms of each development trend of the disease; For example, the current patient's condition is described as "sinus + inflammation / chronic inflammation + mild", and the corresponding improved symptoms are: "nose + breathing normally", "upper respiratory tract + reduced secretions", etc., and the corresponding worsening symptoms are: "head-pain", "upper respiratory tract + secretions with blood", etc.; S132, sorting the possible symptoms of each disease development trend based on probability to obtain the order of the possible symptoms of each disease development trend; Since the probability of different symptoms occurring is different in each disease development trend, in order to quickly obtain the patient's affirmative / negative feedback when asking questions, the symptoms with a higher probability are placed in front and the symptoms with a lower probability are placed in the back.
[0049] S133, constructing a question template and an answer template for each possible symptom, and determining the conversation turn of each question template for each disease development trend based on the order of the possible symptoms of each disease development trend; After completing the sorting, you can build the corresponding question template and answer template.
[0050] Both the question template and the answer template are constructed based on the symptom information. For example, if the symptom of order A is "nose + normal breathing", the constructed question template is "Has your nose become easier to breathe recently?" The corresponding answer templates are "yes / yes", "no / no / no", "increased nasal discharge", etc. to correspond to the other party's various answers.
[0051] Multiple question templates for the same disease development trend are sorted according to probability, so that revisits are performed with users based on the probability.
[0052] S134, constructing a callback answer template based on the question template and answer template of each disease development trend and the dialogue round of each question template.
[0053] Finally, three sets of question and answer templates were constructed based on the three independent question templates and answer templates of disease development trends and the dialogue rounds of each question template.
[0054] In the present application, automatic follow-up visits are generally performed when the patient is in the recovery period, for example, the online follow-up visit process is automatically triggered 10 days after the patient returns home.
[0055] During the follow-up visit, we will first open the conversation based on the patient's personal information address, for example: "XX, Hello, I am a return visit robot. I will be returning your visit now. Please touch the X on the screen to start the return visit." When receiving feedback from the touch screen, the return visit begins. Figure 3 Schematic diagram of the return visit process in an embodiment of the present application, such as Figure 3 As shown, the specific process is as follows: S1, determining the current dialogue round of the current disease development trend, and sending the question template of the current dialogue round in the response and answer template to the current patient; This application takes improvement of the condition as the default trend, so the question and answer template sequence corresponding to the improvement trend is selected to perform a follow-up visit. At the beginning, the question template of the first round is sent to the user.
[0056] If it is not the initial round, the question template of the corresponding round is sent to the user.
[0057] For example: "Is the pain under your ribs getting better?" S2, when receiving text feedback from the current patient, performing entity extraction on the text feedback, wherein the text feedback is generated based on the question template of the current dialogue round; After sending the question template, wait for text feedback from the current patient. If there is text feedback, extract the entity in the text feedback. If there is no feedback after waiting for more than T, end the follow-up and mark it as unresponsive.
[0058] The purpose of entity extraction is to find content related to the answer template in order to determine the current patient's feedback.
[0059] For example, the current patient's feedback text is "No, the pain under the ribs has worsened recently"; At this time, the target entities "none", "under the ribs", and "intensified pain" are extracted.
[0060] S3, matching the extracted entity with the answer template of the current dialogue round in the reply answer template; In this process, the extracted entities are matched by using the entity vectors in the template library and the feedback text to extract the target entities. The matching criterion is still cosine similarity.
[0061] S4, when the entity extracted in the current round matches the answer template of the current dialogue round in the reply answer template, and the extracted entity is an affirmative word or an affirmative symptom description entity, the entity extracted in the current round is added to the feedback data, the next dialogue round is entered and the process returns to step S1 until the dialogue round ends, wherein the answer template includes one or a combination of affirmative words, negative words, and symptom description entities; If the extracted entity matches the answer template of the current round and indicates affirmation, it means that the intention of this round of dialogue is confirmed to be positive, and the affirmative entity and the symptoms of the current round are added to the feedback data as the feedback information of the current round.
[0062] At this point, directly enter the next round of dialogue on the current development trend of the disease until the follow-up visit on the current development trend of the disease is completed.
[0063] Specifically, the symptom description entity representing affirmation is an entity that matches the symptom description entity vector in the answer template of the current dialogue round and does not contain a negative expression. For example, if the symptom description entity vector in the answer template of the current round contains "under the ribs" and "pain", and the feedback text contains the entities "under the ribs" and "intensified pain", it means affirmation.
[0064] S5, when the entity extracted in the current round matches the answer template of the current dialogue round in the response answer template, and the extracted entity is a negative word or a symptom description entity that represents negation, the probability corresponding to the current dialogue round is accumulated into the probability sum, and when the probability sum is less than the preset switching probability threshold, the next dialogue round is entered and the process returns to step S1; when the probability sum is greater than or equal to the preset switching probability threshold, the process switches to the next dialogue on the development trend of the disease; If the extracted entity matches the answer template of the current round and indicates negation, it means that the intention of this round of dialogue is confirmed to be negative. Discard the entity of the current round. At the same time, accumulate the symptom probability of the current round to the total probability middle.
[0065] In the probability sum When the probability is less than the preset threshold, the next round of dialogue on the current disease development trend will be directly entered. When the probability is greater than or equal to the preset probability threshold, the current disease development trend can be basically abandoned, and the next disease development trend, such as the disease worsening trend, can be jumped to.
[0066] In addition, if the condition worsens during the follow-up visit, the total probability If the probability is greater than or equal to the preset probability threshold, the current patient's condition development trend is directly marked as "to be observed".
[0067] Since this application is sorted by probability, the higher the probability of the question, the higher the probability of the question. Generally, when the patient denies the answer after answering 2-3 questions, the application will automatically switch to another trend of the disease development to avoid the patient's negative emotions.
[0068] S6, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and the extracted entity is a symptom description entity, the symptom description entity is matched with all the answer templates in the reply answer template. If they match, the corresponding question template is skipped in the subsequent dialogue round, the entity extracted in the previous round is added to the feedback data, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; if they do not match, the symptom description entity is recorded, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; If a new symptom description is proposed in the text of the current patient feedback, a global match is performed.
[0069] For example, the question template was “Has the pain under your ribs gotten better?”; If the feedback text is: "My breathing has not been very smooth recently", the entities "breathe", "not", and "smooth" are extracted through word segmentation and template matching.
[0070] This feedback does not correspond to the current question, but may appear in subsequent conversations. At this time, match the question templates of other rounds with the extracted entities. If they match, the feedback entities of the corresponding round of questions are recorded in the background. At the same time, guide the patient back to the current round of questions by prompting again. If the patient still does not answer the relevant answers to the current round after multiple guidance, it means that the patient intends to avoid the question. At this time, skip this round and enter the next round of conversation.
[0071] S7, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and does not contain the symptom description entity, the text feedback is forwarded to the fine-tuned language model, and the current patient is guided based on the fine-tuned language model until an entity matching the answer template is extracted from the text feedback of the current patient, the extracted entity is added to the feedback data, and the next dialogue round is entered and returns to step S1 until the dialogue round ends.
[0072] If the patient's feedback does not contain any text related to the current conversation or the description of the symptoms, and the user's intention cannot be understood, the language model is called to have a conversation with the user, guiding the user back to the current conversation.
[0073] The language model used in this application is a fine-tuned language model, which uses relevant corpus to adjust the pre-trained model. This allows the model to understand the user's intention and answer the user's question directly before prompting the user to return to the current round of dialogue. This avoids the situation of being inflexible and having a poor user experience.
[0074] Methods for fine-tuning large language models include: Get the guided dialogue template; The pre-trained language model is trained based on the guided dialogue template to obtain a fine-tuned language model.
[0075] The guidance dialogue template and question template in this application can be set arbitrarily, because after the entity extraction in the previous text, it has been confirmed that the patient's feedback is irrelevant to the current round of dialogue. Therefore, it is only necessary to set a fixed rule to let the big model answer the user's question directly before executing the guidance process. That is, make sure that there are examples with the correct rounds and question templates marked in the training data set, which can help the model learn how to identify and return to the correct dialogue path.
[0076] S140, extracting symptom feedback information from the follow-up data, and analyzing the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0077] Finally, after obtaining all the positive feedback entities, we analyze them in combination with the disease trend map to obtain the disease development analysis results, including: Convert each symptom in the symptom feedback information into a second query vector; Based on matching the second query vector with the disease trend map, obtaining the probability of the disease development trend of each second query vector; The probability of each disease development trend is summed up, and the disease development trend with the greatest probability is taken as the disease development analysis result.
[0078] Since multiple feedback entities corresponding to disease trends may be extracted during the above-mentioned conversation process, the present application sums the probabilities of the feedback entities corresponding to each disease trend and takes the disease development trend with the highest probability as the disease development analysis result.
[0079] In addition, the symptom feedback information and the disease progression analysis results are sent to the doctor; After receiving the confirmation information from the doctor, the disease progression analysis result is sent to the current patient; after receiving the modification information from the doctor, the modified disease progression analysis result is sent to the current patient.
[0080] This can assist doctors in performing follow-up visits, and with the doctor's confirmation, it can ensure that the analysis of the current patient's condition development trend will not be wrong, greatly improving the efficiency of hospital follow-up visits.
[0081] In one embodiment of the present application, converting the condition description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector respectively, includes: (1) converting entities in the condition description information or entities in the symptom feedback information into word vectors based on a table lookup, and extracting position codes of the multiple entities based on an exponential function; (2) The word vector in this embodiment is converted by table lookup mapping, such as One-Hot encoding. Position encoding uses an exponential function to mark the position of each word.
[0082] (3) multiplying the word vector with a pre-constructed first parameter matrix to obtain a word vector matrix; and multiplying the position code with a pre-constructed second parameter matrix to obtain a position matrix; (4) fusing the word vector matrix with the position matrix to obtain a fusion matrix; (5) multiplying the fusion matrix by a pre-constructed third parameter matrix to obtain a coding result; (6) Convert the encoding result into a first query vector or a second query vector by table lookup.
[0083] In this application, the first parameter matrix W1 is used to fuse with the word vector, and the second parameter matrix W2 is used to fuse with the position code. Finally, the word vector matrix and the position matrix are fused (can be added) to obtain a fusion matrix. The fusion matrix is multiplied with the third parameter matrix W3 to obtain the encoding result. The encoding result includes both the word vector of each word and the position information of each word, so the semantic information of the question text can be retained. Finally, the encoding result is converted into a query vector by table lookup mapping, thereby converting the semantic information into a vector that can be recognized by the system.
[0084] The present invention provides a medical data processing method based on an AI big model, which extracts the patient's real-time condition, finds the possible symptoms corresponding to the current condition description information in the condition trend map, and constructs a questionnaire answer template based on the possible symptoms to automatically return the patient. During the return visit, considering the complexity of the conversation, a language big model is introduced for assistance to ensure that the patient's symptom feedback information can be obtained. Based on the patient's symptom feedback information and the condition trend map, it can be roughly analyzed whether the patient's current condition trend has improved, thereby providing doctors with rapid analysis assistance. The present application provides an automatic return visit mechanism based on conversation templates and language big models, which can automatically return visits and automatically analyze patients with different conditions, thereby reducing the return visit workload of medical staff and improving return visit efficiency.
[0085] like Figure 4 As shown, the present application also provides a medical data processing system based on an AI big model, comprising: An acquisition module is used to acquire the condition description information of the current patient, wherein the condition description information is extracted from the medical record or physical examination report; A query module, configured to convert the condition description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; A return visit module, used for constructing a return visit answer template based on the possible symptoms of the current patient, and return visit the current patient based on the return visit answer template and the called language big model to obtain return visit data; The data processing module is used to extract symptom feedback information from the follow-up data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
[0086] The present invention provides a medical data processing system based on an AI big model, which automatically returns the patient by extracting the patient's real-time condition, finding the possible symptoms corresponding to the current condition description information in the condition trend map, and constructing a return answer template based on the possible symptoms. During the return visit, considering the complexity of the conversation, a language big model is introduced for assistance to ensure that the patient's symptom feedback information can be obtained. Based on the patient's symptom feedback information and the condition trend map, it can be roughly analyzed whether the patient's current condition trend has improved, thereby providing doctors with rapid analysis assistance. The present application provides an automatic return visit mechanism based on conversation templates and language big models, which can automatically return visits and automatically analyze patients with different conditions, thereby reducing the return visit workload of medical staff and improving return visit efficiency.
[0087] This embodiment also provides an electronic terminal, including: a processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.
[0088] The computer-readable storage medium in this embodiment can be understood by ordinary technicians in this field: all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.
[0089] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes each step of the above method.
[0090] In this embodiment, the memory may include a random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0091] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0092] In the above-mentioned embodiments, although the present invention has been described in conjunction with the specific embodiments of the present invention, many replacements, modifications and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. The embodiments of the present invention are intended to cover all such replacements, modifications and variations falling within the broad scope of the appended claims.
[0093] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. A medical data processing method based on AI big model, characterized in that: Includes steps: Obtaining the current patient's condition description information, wherein the condition description information is extracted from a medical record or a physical examination report; Converting the condition description information into a first query vector, and determining possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; Constructing a questionnaire answer template based on the possible symptoms of the current patient, and performing a questionnaire answer template and the called language macro model to interview the current patient to obtain questionnaire data; Symptom feedback information is extracted from the follow-up data, and the disease development trend of the current patient is analyzed based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
2. A medical data processing method based on AI big model according to claim 1, characterized in that: The process of constructing the disease trend map includes: Obtain medical records or physical examination reports of multiple patients, and obtain reexamination reports of multiple patients; Extracting condition description information from the medical records or physical examination reports of the multiple patients; extracting symptom descriptions from the review reports; and marking the review reports to obtain review conclusions, wherein the review conclusions include improvement, deterioration, and to be observed; Based on the condition description information, multiple patients are classified to obtain multiple patient types and symptom descriptions and review conclusions of the multiple patient types; and the review conclusion corresponding to each symptom description of each patient type is extracted; Calculate the probability of each review conclusion for the symptom description , ,in, The total number of review conclusions for symptom descriptions, The first The number of review conclusions; Convert symptom descriptions of multiple patient categories into text vectors; The probability of each review conclusion based on multiple patient categories, text vectors of symptom descriptions of multiple patient categories, and symptom descriptions Construct a disease trend map.
3. A medical data processing method based on AI big model according to claim 2, characterized in that: The method for extracting the condition description information comprises: Obtain the patient's medical records or physical examination report text; Cleaning and segmenting the medical record text or physical examination report text to obtain multiple words; Perform entity extraction on the multiple words to obtain a first target entity, a second target entity, and a third target entity, wherein the first target entity is a body part, the second target entity is an injury type, and the third target entity is a severity level; The first target entity, the second target entity and the third target entity whose distances are less than a preset distance threshold are taken as concurrent entities, and structured disease description information is constructed based on the concurrent entities.
4. The medical data processing method based on AI big model according to claim 2 is characterized in that: Constructing a questionnaire answer template based on the possible symptoms of the current patient, including: Classify all possible symptoms based on the development trend of the disease, and obtain the possible symptoms of each development trend of the disease; The possible symptoms of each disease development trend are sorted based on probability to obtain the order of the possible symptoms of each disease development trend; Constructing a question template and an answer template for each possible symptom, and determining the conversation turn of each question template for each disease development trend based on the order of possible symptoms of each disease development trend; A questionnaire answering template is constructed based on the question template, answer template and conversation rounds of each question template for each disease development trend.
5. The medical data processing method based on AI big model according to claim 4 is characterized in that: The current patient is visited based on the visit answer template and the called language model to obtain visit data, including: S1, determining the current dialogue round of the current disease development trend, and sending the question template of the current dialogue round in the response and answer template to the current patient; S2, when receiving text feedback from the current patient, performing entity extraction on the text feedback, wherein the text feedback is generated based on the question template of the current dialogue round; S3, matching the extracted entity with the answer template of the current dialogue round in the reply answer template; S4, when the entity extracted in the current round matches the answer template of the current dialogue round in the reply answer template, and the extracted entity is an affirmative word or an affirmative symptom description entity, the entity extracted in the current round is added to the feedback data, the next dialogue round is entered and the process returns to step S1 until the dialogue round ends, wherein the answer template includes one or a combination of affirmative words, negative words, and symptom description entities; S5, when the entity extracted in the current round matches the answer template of the current dialogue round in the response answer template, and the extracted entity is a negative word or a symptom description entity that represents negation, the probability corresponding to the current dialogue round is accumulated into the probability sum, and when the probability sum is less than the preset switching probability threshold, the next dialogue round is entered and the process returns to step S1; when the probability sum is greater than or equal to the preset switching probability threshold, the process switches to the next dialogue on the development trend of the disease; S6, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and the extracted entity is a symptom description entity, the symptom description entity is matched with all the answer templates in the reply answer template. If they match, the corresponding question template is skipped in the subsequent dialogue round, the entity extracted in the previous round is added to the feedback data, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; if they do not match, the symptom description entity is recorded, and the current dialogue round is repeated and returns to step S1, until the number of repetitions exceeds the preset number, enter the next dialogue round and return to step S1; S7, when the entity extracted in the current round does not match the answer template of the current dialogue round in the reply answer template, and does not contain the symptom description entity, the text feedback is forwarded to the fine-tuned language model, and the current patient is guided based on the fine-tuned language model until an entity matching the answer template is extracted from the text feedback of the current patient, the extracted entity is added to the feedback data, and the next dialogue round is entered and returns to step S1 until the dialogue round ends.
6. The medical data processing method based on AI big model according to claim 5 is characterized in that: The fine-tuning method of the large language model includes: Get the guided dialogue template; The pre-trained language model is trained based on the guided dialogue template to obtain a fine-tuned language model.
7. The medical data processing method based on AI big model according to claim 1 is characterized in that: Analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result, including: Convert each symptom in the symptom feedback information into a second query vector; Based on matching the second query vector with the disease trend map, obtaining the probability of the disease development trend of each second query vector; The probability of each disease development trend is summed up, and the disease development trend with the greatest probability is taken as the disease development analysis result.
8. The medical data processing method based on AI big model according to claim 7 is characterized in that: Converting the condition description information into a first query vector, and converting each symptom in the symptom feedback information into a second query vector, including: Converting entities in the condition description information or entities in the symptom feedback information into word vectors based on a table lookup, and extracting position codes of the multiple entities based on an exponential function; Multiplying the word vector with a pre-constructed first parameter matrix to obtain a word vector matrix; and multiplying the position code with a pre-constructed second parameter matrix to obtain a position matrix; Fusing the word vector matrix with the position matrix to obtain a fusion matrix; Multiplying the fusion matrix with a pre-constructed third parameter matrix to obtain a coding result; The encoding result is converted into a first query vector or a second query vector by table lookup.
9. The medical data processing method based on AI big model according to claim 1 is characterized in that: Also includes: Sending the symptom feedback information and the disease progression analysis results to a doctor; After receiving the confirmation information from the doctor, the disease progression analysis result is sent to the current patient; After receiving the modified information from the doctor, the modified disease progression analysis result is sent to the current patient.
10. A medical data processing system based on AI big model, characterized in that: include: An acquisition module is used to acquire the condition description information of the current patient, wherein the condition description information is extracted from the medical record or physical examination report; A query module, configured to convert the condition description information into a first query vector, and determine the possible symptoms of the current patient based on the first query vector and a pre-constructed condition trend map, wherein the condition trend map includes a plurality of condition description tags, possible symptoms of the plurality of condition description tags, and condition development trends corresponding to the symptoms; A return visit module, used for constructing a return visit answer template based on the possible symptoms of the current patient, and return visit the current patient based on the return visit answer template and the called language big model to obtain return visit data; The data processing module is used to extract symptom feedback information from the follow-up data, and analyze the disease development trend of the current patient based on the symptom feedback information and the disease trend map to obtain a disease development analysis result.
Citation Information
Patent Citations
Illness state analysis method and device, electronic equipment and storage medium
CN113707307A
Disease inquiry process auxiliary information generation method based on large language model
CN117690581A
Medical question answering system based on large language model
CN117851558A
Medical large model question answering system based on medical records
CN118016324A
Medical AI assistant implementation method and system based on data driving and large model
CN118098585A