Outpatient pre-examination triage method and system based on deep learning
By extracting symptom descriptive words and emotional polarity words using deep learning technology and calculating temperature parameters to adjust the attention of the triage model, the problem of interference from patients' vague words and emotional words is solved, and the accuracy and adaptability of the triage path are improved.
Patent Information
- Application Number
- CN202511216809.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing electronic triage technology based on natural language processing is easily interfered with when faced with patients' vague or emotional words, leading to a decrease in the accuracy of the triage path.
Using a deep learning-based approach, the attention distribution of the triage path recommendation model is adjusted by extracting symptom descriptive words, keywords, and sentiment polarity words, and calculating temperature parameters. The triage path is optimized through multi-round iterative question-and-answer dialogue, and relevant questions are generated to clarify or supplement symptom information.
This improved the fit between the triage pathway and patients, increased the accuracy of triage results, and continuously optimized the accuracy of the triage model through data accumulation.
Smart Images

Figure CN120748649B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical triage, in particular to a pre-examination triage method and system for outpatients based on deep learning. BACKGROUND
[0002] Triage is a process in emergency medical care that classifies patients according to the severity of their condition and the specialist they belong to, and arranges the order of treatment. Its core is to achieve condition classification and department diversion through rapid data collection and analysis, to prioritize critical cases and improve rescue efficiency, while maintaining treatment order and optimizing patient experience.
[0003] Currently, in the prior art, electronic triage technology based on natural language processing has been gradually applied to outpatient process optimization. Such systems usually collect patient expression text through a dialogue interface and use pre-trained models to match triage paths. However, in actual hospital work scenarios, patients often do not have professional knowledge, and their expressions often contain ambiguous words, emotional words and other interference information, which can easily interfere with the model, reduce the focus on real symptom words, and affect the accuracy of the triage path. SUMMARY
[0004] To solve the technical problem of inaccurate triage results caused by confusing and ambiguous words in patient expression text, the present application provides a pre-examination triage method and system for outpatients based on deep learning, which uses the following technical solutions:
[0005] The present application proposes a pre-examination triage method for outpatients based on deep learning, which includes:
[0006] During the process of asking and answering questions with the pre-set interactive interface at the patient end, the expression text input by the patient is obtained; symptom description words, keywords related to triage, and emotional polarity words are extracted from the expression text;
[0007] According to the information entropy difference characteristics between the symptom description words and the keywords, the co-occurrence characteristics between the symptom description words and the emotional polarity words, and the semantic expression characteristics of the symptom description words, the temperature parameter for adjusting the attention distribution in the triage path recommendation model is determined;
[0008] The expression text is input into the triage path recommendation model with the set temperature parameter to obtain candidate triage categories and corresponding triage paths; based on the candidate triage categories, candidate questions are generated; the candidate questions are pushed to the patient end, and the patient's new responses to the candidate questions are received to obtain new expression text;
[0009] The steps of calculating the temperature parameter, inputting the model to obtain the candidate triage category and triage path, generating and pushing the candidate question, and obtaining the new expression text are repeatedly performed for multiple rounds of iterative question and answer until a preset termination condition is met.
[0010] The latest expression text obtained in the last round of iteration is input into the triage path recommendation model to obtain the target triage category and the target triage path.
[0011] Further, the temperature parameter determination process includes:
[0012] Based on the information entropy difference feature between the symptom description words and the keywords, a symptom description sufficiency index of the patient is determined, which quantitatively expresses the coverage degree of the key information related to triage in the text.
[0013] Based on the semantic expression characteristics of the symptom description words, a symptom description ambiguity index of the patient is determined.
[0014] Based on the co-occurrence characteristics between the symptom description words and the sentiment polarity words, an emotion association strength index of the patient is determined, which quantitatively expresses the statistical association strength of the co-occurrence of the sentiment polarity words and the symptom description words in the text.
[0015] According to a preset function relationship among the symptom description sufficiency index, the symptom description ambiguity index, and the emotion association strength index, a temperature parameter is calculated.
[0016] Further, the symptom description sufficiency index determination process includes:
[0017] The information entropy of all symptom description words is calculated to obtain a first information entropy;
[0018] The information entropy of all keywords is calculated to obtain a second information entropy;
[0019] Based on the first information entropy and the second information entropy, a symptom description sufficiency index is obtained through a preset sufficiency index calculation function.
[0020] Further, the symptom description ambiguity index determination process includes:
[0021] From the expression text, ambiguous words belonging to a preset ambiguous word dictionary are identified; and a plurality of triage category sets containing a plurality of symptom reference words are obtained;
[0022] The total number of ambiguous words in the expression text is compared with the total number of symptom description words to obtain a second ratio;
[0023] The symptom description words are semantically clustered to obtain at least one semantic cluster, each semantic cluster containing a plurality of symptom description words that are semantically similar.
[0024] For any symptom description word in each semantic cluster, determine whether there is a target symptom reference word in the set of triage categories that meets the similarity condition with the any symptom description word; the set of triage categories that has the target symptom reference word is taken as the target set of triage categories;
[0025] For each target set of triage categories, perform ratio operation on the sum of the number of target symptom reference words in the target set of triage categories and the sum of the number of all target symptom reference words to obtain a third ratio of the target set of triage categories;
[0026] Based on the second ratio and the third ratio, determine the symptom description ambiguity index through a preset ambiguity calculation function.
[0027] Further, based on the second ratio and the third ratio, determine the symptom description ambiguity index through a preset ambiguity calculation function, comprising:
[0028] Calculate the product of the second ratio and the third ratio, and take the product as the priority score of the target set of triage categories;
[0029] According to the order from large to small of the priority score, select a preset number of target sets of triage categories as core sets of triage categories; calculate the variance of the priority scores of all core sets of triage categories, normalize the inverse of the variance to obtain a concentration index;
[0030] Multiply the second ratio and the concentration index to obtain the symptom description ambiguity index.
[0031] Further, the sentiment polarity word at least includes an exclamation word and a degree adverb, and the emotion association strength index determination process comprises:
[0032] Statistically count the total number of sentiment polarity words appearing in the expression text; and respectively count the appearance frequency of the exclamation word and the appearance frequency of the degree adverb in the expression text;
[0033] Identify all symptom description words and sentiment polarity words pairs that commonly appear in a sentence or a preset context window in the expression text as effective co-occurrence pairs; statistically count the total number of effective co-occurrence pairs;
[0034] Perform ratio operation on the total number of effective co-occurrence pairs and the total number of sentiment polarity words to obtain a sentiment-symptom co-occurrence strength ratio;
[0035] According to the exclamation word frequency, the degree adverb frequency and the sentiment-symptom co-occurrence strength ratio, calculate the emotion association strength index through a preset association strength calculation function.
[0036] Further, the candidate question generation process comprises:
[0037] extracting key symptom feature words associated with the candidate triage category from the expression text;
[0038] generating a prompt template of the key symptom feature words, the candidate triage category and the preset question by inputting the key symptom feature words, the candidate triage category and the preset question into a preset natural language generation model;
[0039] obtaining a plurality of initial candidate questions output by the natural language generation model.
[0040] Further, the outpatient pre-examination triage method based on deep learning further comprises:
[0041] obtaining a historical dialogue record set, the historical dialogue record set containing a plurality of historical dialogue records, each historical dialogue record including a historical expression text sequence generated in a plurality of historical question and answer dialogue rounds, a historical triage category sequence and a corresponding historical candidate question sequence;
[0042] For each initial candidate question, a historical dialogue record subset matching the current candidate triage category is retrieved from the historical dialogue record set; in the historical dialogue record subset, a historical candidate question similar in semantics to the initial candidate question is identified as a similar question; based on the statistical characteristics of the similar question in the historical dialogue record subset, a screening score of the initial candidate question is calculated;
[0043] From the screening scores of all initial candidate questions, the initial candidate question with the highest screening score is selected as the candidate question and pushed to the patient end.
[0044] Further, there is an effective historical triage category similar in meaning to the current candidate triage category in the historical triage category sequence, and the outpatient pre-examination triage method based on deep learning further comprises:
[0045] The statistical characteristics include the co-occurrence frequency between the similar question and the effective historical triage category in its corresponding historical dialogue record, the occurrence frequency of the similar question in its corresponding historical dialogue record and the occurrence round distribution characteristics of the similar question in its corresponding historical dialogue record.
[0046] An outpatient pre-examination triage system based on deep learning, the system comprising a memory, a processor and a computer program stored in the memory and running on the processor, the processor implementing the steps of the outpatient pre-examination triage method based on deep learning when executing the computer program.
[0047] The present application has the following beneficial effects:
[0048] The application determines the temperature parameter by quantifying the key information of the symptom description words in the expression text, the semantic ambiguity degree of the expression text, and the association degree between the emotional polarity words and the symptom description words, injects the temperature parameter into the triage path recommendation model, so that the model can adaptively adjust the attention distribution according to the characteristics of the expression text, and consider the possible disordered expression or emotional situation of the patient. Next, through multiple rounds of iterative question and answer dialogues, not only the adaptation degree between the triage path and the patient can be improved, but also based on the determined candidate triage path, more appropriate candidate questions can be generated to further clarify or supplement the symptom information. Moreover, through each round of question and answer dialogue, the temperature parameter and the triage path are updated in real time, and the triage path model is continuously optimized, and the result accuracy of the patient triage matching can be further improved with data accumulation. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0050] Figure 1 A flow chart of a deep learning-based outpatient pre-examination triage method provided by an embodiment of the present application;
[0051] Figure 2 An example diagram of a temperature parameter determination process provided by an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined application purpose, the following describes the deep learning-based outpatient pre-examination triage method and system according to the present application, its specific implementation, structure, features and effects in detail, as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0054] The specific scheme of the deep learning-based outpatient pre-examination triage method and system provided by the present application is described in detail below with reference to the drawings.
[0055] Please refer to Figure 1It shows a deep learning-based outpatient pre-examination triage method flowchart provided by one embodiment of the application, and the method comprises the following steps:
[0056] S101, in the process of asking and answering dialogue between the patient end and the preset interactive interface, obtaining the expression text input by the patient; extracting the symptom description word, the keyword related to triage, and the sentiment polarity word from the expression text.
[0057] It should be noted that the preset interactive interface is a channel for information exchange between the patient end and the computer, and the specific interface arrangement of the preset interactive interface is determined according to actual needs, which is not limited in the embodiment.
[0058] The keyword is a word that can express the central content of the expression text related to triage, and the specific extraction method of the keyword is a technical means familiar to those skilled in the art, which will not be repeated in the embodiment. For example, a pre-set language network graph related to medical triage is used, and then the network graph analysis is performed on the expression text to find words or phrases with important roles on the graph, which are the keywords of the expression text.
[0059] It should be noted that the specific extraction method of the symptom description word is a technical means familiar to those skilled in the art, which will not be repeated in the embodiment. For example, an existing medical word segmentation tool is used to segment and extract the expression text, such as extracting the symptom description word by pkuseg.
[0060] The sentiment polarity word refers to the emotional tendency word carried by the emotional expression, and the specific extraction method of the sentiment polarity word is a technical means familiar to those skilled in the art, which will not be repeated in the embodiment. For example, a pre-set sentiment polarity dictionary is used to extract the sentiment polarity word, such as using the Linguistic Inquiry and Word Count (LIWC) as the pre-set sentiment polarity dictionary to extract the sentiment polarity word.
[0061] S102: According to the information entropy difference characteristics between the symptom description word and the keyword, the co-occurrence characteristics between the symptom description word and the sentiment polarity word, and the semantic expression characteristics of the symptom description word, the temperature parameter for adjusting the attention distribution in the triage path recommendation model is determined.
[0062] The temperature parameter, in the attention mechanism of the triage path recommendation model, affects the behavior of the triage path recommendation model by adjusting the output probability distribution of the softmax function.
[0063] It needs to be understood that due to different self-expression abilities of different patients, and the expression ability of the patient may also be affected by the physical condition of the patient, in order to make it more accurate to grasp the chief complaint of the patient, considering the influence of the symptom description word, the keyword and the sentiment polarity word in the expression text on the grasping of the chief complaint of the patient, the triage path recommendation model can be adaptively adjusted by adjusting the temperature parameter.
[0064] The temperature parameter determination process, as shown in Figure 2 , includes:
[0065] S102-1: Determine the symptom description sufficiency index of the patient based on the information entropy difference feature between the symptom description word and the keyword.
[0066] It needs to be pointed out that the symptom description sufficiency index quantitatively expresses the coverage degree of the key information related to triage in the text.
[0067] In this embodiment, the information entropy of all symptom description words is calculated to obtain a first information entropy; the information entropy of all keywords is calculated to obtain a second information entropy; and based on the first information entropy and the second information entropy, the symptom description sufficiency index is obtained through a preset sufficiency index calculation function.
[0068] It needs to be pointed out that the specific way of calculating the information entropy is a technology familiar to those skilled in the art, and will not be described here.
[0069] Since the larger the first information entropy is, the more uniform the probability of the appearance of the symptom description word is when the patient describes the symptom, the symptoms described are more miscellaneous, and there is no focus on describing the chief complaint, and the larger the first information entropy of the symptom description word is, and the closer it is to the second information entropy of the keyword, the more dispersed the effective symptom description word in the text is when describing the symptom, therefore, assuming that the first information entropy is less than the second information entropy , the symptom description sufficiency index can be represented by the following preset sufficiency index calculation function:
[0070]
[0071] Among them, represents the symptom description sufficiency index; represents the first information entropy; represents the second information entropy.
[0072] It needs to be pointed out that a small positive offset is added to the denominator , which mathematically ensures that the denominator Strictly greater than zero, so that the denominator of the division operation is completely eliminated the possibility of zero, to ensure the robustness and executability of the calculation formula, and Is preset to a small positive value, the selection principle of the specific value is: in Normal working condition, significantly greater than zero, The impact of the calculation results Can be ignored, for example, The specific value is determined according to the actual demand, Can be 0.01, the embodiment is not limited specifically.
[0073] Among them, it should be noted that in the actual scene, the size of the first information entropy and the second information entropy may exist following two cases, one case is Less than , another case is Greater than , however, in Greater than , if the formula To calculate the sufficiency index of symptom description, instead will appear with the first information entropy is greater, and with the first information entropy and the second information entropy between the smaller the difference, the effective symptom description word in the text distribution is more gathered, therefore, in order to avoid the division operation due to the size of the numerator and denominator between the problem, resulting in the calculated Expected expression of the meaning of ambiguity, can be in Greater than , through the formula To calculate the sufficiency index of symptom description.
[0074] S102-2: based on the semantic expression characteristics of the symptom description word, determine the ambiguity degree index of the patient's symptom description.
[0075] In the embodiment, ambiguous words belonging to a preset ambiguous word dictionary are identified from the expression text; a plurality of preset triage category sets are obtained, and the triage category sets include a plurality of symptom reference words; a second ratio is obtained by ratio calculation of a total number of the ambiguous words in the expression text and a total number of the symptom description words; semantic clustering is performed on the symptom description words to obtain at least one semantic cluster, and each semantic cluster includes a plurality of symptom description words that are semantically similar; for any symptom description word in each semantic cluster, it is determined whether there is a target symptom reference word in each triage category set that meets a similarity condition with the symptom description word; a triage category set in which the target symptom reference word exists is taken as a target triage category set; for each target triage category set, a third ratio of the target triage category set is obtained by ratio operation of a total number of the target symptom reference words in the target triage category set and a total number of all the target symptom reference words; and a symptom description ambiguity index is determined based on the second ratio and the third ratio through a preset ambiguity degree calculation function.
[0076] It should be noted that the specific content of the preset ambiguous word dictionary is determined according to actual needs, and the embodiment is not limited specifically, for example, words such as “maybe”, “probably”, “sometimes”, “not sure” and the like can be set as ambiguous words in the preset ambiguous word dictionary, and the specific construction method of the preset ambiguous word dictionary is a technical means familiar to those skilled in the art, and the embodiment will not be repeated.
[0077] The triage category is a department or department classification label predefined by a medical institution, for example, “pediatrics”, “eye emergency” and the like.
[0078] It should be noted that the plurality of preset triage category sets and the specific information of the symptom reference words can be summarized from the triage categories commonly used by most medical institutions and the diseases treated by each triage category, and the embodiment is not limited specifically.
[0079] It should be noted that the third ratio represents the degree of professionalism of the words used by the patient when describing the symptom description word. Since the plurality of symptom reference words in the preset triage category set are usually some summarized professional words, for example, “acute upper respiratory tract infection” and the like, if the number of target symptom reference words that meet the similarity condition with a certain symptom description word in a certain target triage category set is more, it reflects that the patient uses more professional words when describing the symptom description word, and the ambiguity is less.
[0080] It should be noted that the specific method of performing semantic clustering on the symptom description words is a technical means familiar to those skilled in the art, and the embodiment will not be repeated, for example, the K-Means clustering method is used to perform semantic clustering on the symptom description words.
[0081] The similarity condition includes that the similarity between any symptom description word and the symptom reference word exceeds a preset similarity threshold.
[0082] It should be noted that the specific calculation method of the similarity is a technology familiar to those skilled in the art, and the embodiment will not be repeated. For example, a first word vector of any symptom description word and a second word vector of the symptom reference word are generated, and the cosine similarity between the first word vector and the second word vector is calculated. The cosine similarity can be used as the similarity between any symptom description word and the symptom reference word.
[0083] It should be noted that the specific value of the preset similarity threshold is determined according to actual needs, and the embodiment does not make specific limitations. For example, the preset similarity threshold can be 0.5.
[0084] In order to accurately determine the symptom description ambiguity index through the preset ambiguity degree calculation function, as a possible implementation manner, the product of the second ratio and the third ratio is calculated, and the product is used as the priority score of the target triage category set. According to the order from large to small of the priority score, a preset number of target triage category sets are selected as the core triage category set. The variance of the priority scores of all core triage category sets is calculated, the inverse of the variance is normalized to obtain a concentration index. The second ratio is multiplied by the concentration index to obtain the symptom description ambiguity index.
[0085] It should be noted that the specific value of the preset number is determined according to actual needs, and the embodiment does not make specific limitations. For example, the preset number can be 5.
[0086] Since the greater the concentration index, the more concentrated the priority score distribution of the core triage category set (that is, the smaller the variance), the priority score of each core triage category set fluctuates less, which reflects that the patient's symptom description is highly matched with multiple different triage categories, and the scores of these categories are difficult to distinguish primary and secondary, resulting in the inability to determine the primary triage direction, and therefore the symptom information expressed by the patient has a high degree of ambiguity. The greater the second ratio, the more the total number of ambiguous words in the expression text, and the more ambiguous words reflect that the expression text is more ambiguous. Therefore, the symptom description ambiguity index can be represented by the following preset ambiguity degree calculation function:
[0087]
[0088] Wherein, The symptom description ambiguity index is represented by; The total number of symptom description words is represented by; The total number of ambiguous words is represented by; The concentration index is represented by.
[0089] It should be noted that in actual medical cases, the content and manner of expression of different patients are not fixed, but even if different patients, in the question and answer dialogue in the preset interactive interface, the patient corresponding to the patient end can extract the symptom description word "pain" even if the patient simply expresses his own condition, for example, "whole body pain", and the total number of symptom description words cannot be zero.
[0090] S102-3: Based on the co-occurrence characteristics between the symptom description word and the emotional polarity word, determine the emotional association strength index of the patient, which quantifies the statistical association strength of the emotional polarity word and the symptom description word appearing together in the text.
[0091] In this embodiment, the total number of emotional polarity words appearing in the expression text is counted, and the appearance frequency of exclamation words and the appearance frequency of degree adverbs in the expression text are counted respectively; all pairs of symptom description words and emotional polarity words that co-occur in a sentence or a preset context window in the expression text are identified as effective co-occurrence pairs; the total number of effective co-occurrence pairs is counted; the total number of effective co-occurrence pairs and the total number of emotional polarity words are ratio operated to obtain the emotional-symptom co-occurrence strength ratio; and the emotional association strength index is calculated according to the exclamation word frequency, the degree adverb frequency and the emotional-symptom co-occurrence strength ratio through a preset association strength calculation function.
[0092] It should be noted that the emotional polarity word at least includes exclamation words and degree adverbs.
[0093] It should be noted that the context window indicates the range considered before and after processing the expression text, for example, within the previous and subsequent 5 words, which is not specifically limited in this embodiment.
[0094] The effective co-occurrence pair is used to indicate the symptom description word and the emotional polarity word that co-occur in a sentence or a preset context window.
[0095] Since the greater the emotional-symptom co-occurrence strength ratio, the greater the degree of association between the emotional word and the symptom description word, the patient's expression of the symptom description word is associated with a stronger emotion, rather than the independent appearance of the emotional polarity word, the greater the frequency of the appearance of the exclamation word and the degree adverb, the stronger the emotional reaction expressed by the patient, and if the exclamation word and the degree adverb appear at the same time as the symptom description word, it reflects the severity of the patient's symptoms, further enhancing the strength of emotional expression, therefore, the emotional association strength index can be calculated by the following preset association strength calculation function:
[0096]
[0097] Wherein, Emotional association strength index is represented by The total number of effective co-occurrence pairs is represented by denotes the total number of sentiment polarity words; denotes a logarithmic function; denotes the base of the logarithm, and a is greater than 1; i denotes the frequency of exclamation words; denotes the frequency of degree adverbs.
[0098] It should be noted that the specific value of a can be determined according to the actual situation of the text, for example, a can take the natural base e, and the present embodiment will not be described again.
[0099] It should be noted that in actual medical situations, the content and manner of expression of different patients are not fixed, but even if different patients, in the question and answer dialogue with the pre-set interactive interface, due to the patient's own uncomfortable condition, under normal circumstances, the patient corresponding to a certain patient will use sentiment polarity words even if the patient simply expresses his own condition, for example, "I have a very headache", but it does not exclude the case that the patient does not use sentiment polarity words, then is set to 0, and is set to 1.
[0100] S102-4: According to the pre-set function relationship among the symptom description sufficiency index, the symptom description ambiguity index and the emotion correlation strength index, the temperature parameter is calculated.
[0101] It should be understood that in actual working conditions, based on the symptom description sufficiency index, the symptom description ambiguity index and the emotion correlation strength index of most patients, it can be found that the patient's expression can be summarized into the following three cases, one case is clear description type: the patient directly and clearly describes the typical symptoms based on his own cognition; another case is disordered description type: the patient provides multiple symptom information in disorder; another case is fuzzy emotion type: the patient describes the symptoms with non-professional words, the description of the symptoms is relatively general (such as feeling weak all over the body, and can't say where it hurts), and the body discomfort may cause psychological anxiety, and multiple symptoms unrelated to the main complaint are provided, and the symptoms are described with emotional language, therefore, based on the determined symptom description sufficiency index, the symptom description ambiguity index and the emotion correlation strength index, the patient's expression mode can be objectively classified.
[0102] For example, the symptom description sufficiency index, the symptom description ambiguity index and the emotion correlation strength index can be distinguished by setting a classification threshold, so as to realize the objective classification of the patient's expression mode, and the classification threshold is shown in Table 1, wherein it can be found that when the symptom description sufficiency index is greater than 0.7, the symptom description ambiguity index is less than 0.5, and the emotion correlation strength index is less than 0.3, the patient's expression mode can be classified as clear description type.
[0103] Table 1
[0104] Indicator interval Symptom description adequacy indicator interval Symptom description ambiguity indicator interval Emotion association strength indicator interval Clear description type b is greater than 0.7 f is less than 0.5 j is less than 0.3 Disordered description type b is less than 0.7 f is greater than 0.5 j is greater than 0.3 and less than or equal to 0.8 Fuzzy emotion type b is less than 0.7 f is greater than 0.5 j is greater than 0.8
[0105] wherein b represents the symptom description sufficiency index; f represents the symptom description ambiguity index; j represents the emotion association intensity index.
[0106] It should be noted that the specific value of the classification threshold can be summarized according to the historical dialogue data of the medical institution, and the embodiment is not limited specifically.
[0107] It should be understood that based on the symptom description sufficiency index, the symptom description ambiguity index and the emotion association intensity index of the current patient, the category to which the expression mode of the current patient belongs is determined, and the symptom description sufficiency index, the symptom description ambiguity index and the emotion association intensity index of the current patient are added to the symptom description sufficiency index set, the symptom description ambiguity index set and the emotion association intensity index set of the same category respectively; the mean value of all symptom description sufficiency indexes in the same category is calculated, and the mean value of the symptom description sufficiency index of the current patient is determined; the mean value of all symptom description ambiguity indexes in the same category is calculated, and the mean value of the symptom description ambiguity index of the current patient is determined; the mean value of all emotion association intensity indexes in the same category is calculated, and the mean value of the emotion association intensity index of the current patient is determined. Therefore, in order to further improve the influence effect of the determined temperature parameter on the attention distribution, the mean value of the symptom description sufficiency index can be used as the current symptom description sufficiency index of the current patient, the mean value of the symptom description ambiguity index can be used as the current symptom description ambiguity index of the current patient, and the mean value of the emotion association intensity index can be used as the current emotion association intensity index of the current patient. According to the preset function relationship between the current symptom description sufficiency index, the current symptom description ambiguity index and the current emotion association intensity index, the temperature parameter is calculated.
[0108] Since the higher the symptom description sufficiency index, the clearer the symptom description words contained in the expression text, the more likely the patient is to be a clear description type, and the focus needs to be on a small number of core symptom words; the lower the symptom description ambiguity index, the more disordered and ambiguous symptom description words contained in the expression text, the more likely the patient is to be a disordered description type, and the attention range needs to be expanded; the lower the emotion association intensity index, the more emotional expression symptom description words contained in the expression text, the more likely the patient is to be a fuzzy emotion type, and the attention range needs to be expanded. Therefore, the temperature parameter can be represented by the following preset function relationship:
[0109]
[0110] Wherein, b represents the symptom description sufficiency index; f represents the symptom description ambiguity index; j represents the emotion association strength index. represents a basic value of a temperature parameter; represents a temperature parameter.
[0111] It should be noted that, The specific value of is determined according to the actual situation, which can be set by the person skilled in the art according to the existing technical experience, for example, Can be 0.2.
[0112] S103: input the expression text into the triage path recommendation model with the set temperature parameter, obtain the candidate triage category and the corresponding triage path; generate candidate questions based on the candidate triage category; push the candidate questions to the patient end, and receive the new response of the patient to the candidate questions, and obtain the new expression text.
[0113] It should be noted that, due to different expression modes of different patients, in general, only one round of question and answer dialogue cannot understand the patient's chief complaint, therefore, according to the determined candidate triage category, candidate questions are generated and pushed to the patient end, and then the new expression text is determined to further clarify or supplement the symptom information.
[0114] In this embodiment, the key symptom feature words associated with the candidate triage category are extracted from the expression text; the key symptom feature words, the candidate triage category and the preset question generation prompt template are input into the preset natural language generation model; and a plurality of initial candidate questions output by the natural language generation model are obtained.
[0115] The key symptom feature words are the symptom description words associated with the candidate triage category among the plurality of symptom description words contained in the expression text.
[0116] For example, the candidate triage category is otology, and the respective symptom description words contained in the expression text are "ear pain", "hard to hear" and "cause dizziness", then the symptom description words associated with the candidate triage category are "ear pain" and "hard to hear", that is, the key symptom feature words are "ear pain" and "hard to hear".
[0117] It should be noted that the specific construction method of the preset natural language generation model is a well-known technical means to the person skilled in the art, for example, the GPT model can be used to construct the preset natural language generation model.
[0118] It should be noted that the specific construction method of the preset question generation prompt template is determined according to the actual demand, which is not limited in this embodiment, for example, the CO-STAR framework can be used to construct a high-quality prompt template.
[0119] It should be noted that the initial candidate question is prepared for subsequent screening of candidate questions, and is finally used to further clarify or supplement the symptom information.
[0120] In this embodiment, a historical dialogue record set is obtained, the historical dialogue record set containing a plurality of historical dialogue records, each historical dialogue record including a historical triage category sequence and a corresponding historical candidate question sequence generated in a plurality of rounds of historical question and answer dialogue; for each initial candidate question, a historical dialogue record subset matching the current candidate triage category is retrieved from the historical dialogue record set; in the historical dialogue record subset, a historical candidate question similar in semantics to the initial candidate question is identified as a similar question; based on statistical characteristics of the similar question in the historical dialogue record subset, a screening score of the initial candidate question is calculated; from the screening scores of all initial candidate questions, an initial candidate question with the highest screening score is selected as a candidate question and pushed to the patient end.
[0121] It should be noted that the historical dialogue record set comes from historical dialogue records generated in the historical question and answer dialogue process between the patient end and the preset interactive interface, and the historical triage category sequence includes historical triage categories determined at different times, and the corresponding historical candidate question sequence includes historical candidate questions generated at different times, in chronological order from early to late of each round of historical question and answer dialogue.
[0122] It can be understood that the historical dialogue record subset is used to indicate the part of the historical dialogue record matching the current candidate triage category.
[0123] It should be noted that the specific determination method of the similar question is a technology familiar to those skilled in the art, and this embodiment will not be repeated, for example, the similar question can be matched and determined based on the semantic similarity (word vector cosine similarity or sentence vector similarity) between the initial candidate question and the historical candidate question.
[0124] It should be noted that the statistical characteristics include the co-occurrence frequency between the similar question and the effective historical triage category in its corresponding historical dialogue record, the occurrence frequency of the similar question in its corresponding historical dialogue record, and the occurrence round distribution characteristics of the similar question in its corresponding historical dialogue record.
[0125] Effective historical triage category, correct triage result verified by actual medical treatment in the historical dialogue.
[0126] In this embodiment, the point mutual information between the similar question and the effective historical triage category in its corresponding historical dialogue record is calculated to obtain a correlation degree indicator; the frequency of occurrence of the effective historical triage category is identified; and the product of the frequency of occurrence of the effective historical triage category and the correlation degree indicator is calculated to obtain a first score.
[0127] It should be noted that the correlation degree index is used to measure the correlation between the similar question and the effective historical triage category in the corresponding historical dialogue record, and the point mutual information is calculated in a specific manner known to those skilled in the art, and the embodiment will not be repeated.
[0128] In the embodiment, the historical key symptom feature words related to the effective historical triage category are extracted from the historical expression text; for each historical key symptom feature word, it is judged whether there is a first target symptom reference word in the set of preset triage categories that satisfies the first similarity condition with the historical key symptom feature word; and the product of the total number of the first target symptom reference word and the occurrence frequency of the similar question in the corresponding historical dialogue record is calculated to obtain a second score.
[0129] It should be noted that the specific description of the set of preset triage categories can refer to step S102-2 of the embodiment, and the embodiment will not be repeated.
[0130] The first similarity condition includes that the similarity between the historical key symptom feature word and each symptom reference word in the set of preset triage categories exceeds a preset first similarity threshold.
[0131] It should be noted that the specific calculation method of the similarity is a technical means known to those skilled in the art, and the embodiment will not be repeated, for example, a third word vector of the historical key symptom feature word and a fourth word vector of the symptom reference word in the set of preset triage categories are generated, and the cosine similarity between the third word vector and the fourth word vector is calculated, which can be used as the similarity between the historical key symptom feature word and each symptom reference word in the set of preset triage categories.
[0132] It should be noted that the specific value of the preset first similarity threshold is determined according to actual needs, and the embodiment is not limited, for example, the preset first similarity threshold can be 0.5.
[0133] In the embodiment, the number of rounds of the historical question and answer dialogue remaining in the corresponding historical dialogue record after the similar question is pushed is obtained, and the round of the historical question and answer dialogue in which the similar question is located in the corresponding historical dialogue record is obtained; based on the number of rounds of the remaining historical question and answer dialogue and the round of the historical question and answer dialogue in which the similar question is located, a third score is obtained through a preset third score calculation function.
[0134] Since, if the number of remaining historical question and answer dialog turns is less, it means that after the similar question is pushed, the symptom information is further clarified or supplemented, the question and answer dialog turns are shortened, the similar question contributes more to ending the historical question and answer dialog process, if the turn of the historical question and answer dialog where the similar question is located is later, and the cumulative dialog time to reach the turn of the historical question and answer dialog where the similar question is located is shorter, it means that only a small number of question and answer dialog turns are needed to understand the symptom information, so the similar question contributes more to ending the historical question and answer dialog process, therefore, the third score can be represented by the following preset third score calculation function:
[0135]
[0136] wherein, represents the third score; represents the number of remaining historical question and answer dialog turns; represents the turn of the historical question and answer dialog where the similar question is located; T represents the total time consumed by the historical question and answer dialog reaching the th turn in the historical dialog record.
[0137] wherein, it is necessary to note that the th turn of the historical question and answer dialog is a historical question and answer dialog that has already occurred and has carried out question and answer dialog with the patient end, and in this process, time must be consumed, so T cannot be zero.
[0138] wherein, it is necessary to note that a very small positive offset is added to the denominator , which mathematically ensures that the denominator is strictly greater than zero, thereby completely eliminating the possibility of a zero denominator in the division operation, ensuring the robustness and executability of the calculation formula, and is preset to a small positive value, and the selection principle of the specific value is: under normal working conditions where is significantly greater than zero, the influence on the calculation result can be ignored, for example, the specific value of is determined according to actual needs, which can be 0.01, and the present embodiment does not make specific limitations.
[0139] wherein, it is necessary to note that the specific value of the first score is a normalized value, the specific value of the second score is a normalized value, and the specific value of the third score is a normalized value, and the specific way of normalization is a technical means familiar to those skilled in the art, which will not be described in detail in the present embodiment, for example, normalization can be performed by averaging.
[0140] In this embodiment, the first score, the second score and the third score are weighted and summed to obtain a screening score of the initial candidate question by a preset first weight of the first score, a preset second weight of the second score and a preset third weight of the third score.
[0141] It should be noted that the sum of the preset first weight, the preset second weight and the preset third weight is a positive integer 1, and the specific values of the preset first weight, the preset second weight and the preset third weight are determined according to actual conditions, which are not limited in this embodiment. For example, the preset first weight can be 0.3, the preset second weight can be 0.45, and the preset third weight can be 0.25.
[0142] S104: Repeat the steps of calculating the temperature parameter, inputting the model to obtain the candidate triage category and the triage path, generating and pushing the candidate question, and obtaining the new expression text, and perform multiple rounds of iterative question and answer until a preset termination condition is met.
[0143] It should be noted that the preset termination condition includes that the cumulative question and answer conversation time reaches a preset time threshold, or that an instruction to end the question and answer conversation is received, or that the number of iterations reaches a preset number threshold.
[0144] It should be noted that the specific value of the preset time threshold is determined according to actual conditions, which is not limited in this embodiment. For example, the preset time threshold can be 3 minutes.
[0145] It should be noted that the specific value of the preset number threshold is determined according to actual conditions, which is not limited in this embodiment. For example, the preset number threshold can be 5.
[0146] The triage path is a registration or treatment process guide recommended to the patient corresponding to the patient end.
[0147] It should be noted that the specific description of calculating the temperature parameter can refer to step S102 of this embodiment, and the specific description of inputting the model to obtain the candidate triage category and the triage path, generating and pushing the candidate question, and obtaining the new expression text can refer to step S103 of this embodiment, which will not be repeated here.
[0148] S105: Input the latest expression text obtained in the last round of iteration into the triage path recommendation model to obtain a target triage category and a target triage path.
[0149] It can be understood that the target triage category and the target triage path can be pushed to the patient end and only used to assist in the management of the triage process, without providing the final medical diagnosis or treatment decision.
[0150] One embodiment of the present application provides a deep learning-based outpatient pre-examination triage system, which comprises a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the deep learning-based outpatient pre-examination triage method when executing the computer program.
[0151] It should be noted that the above-mentioned embodiment sequence is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or can be advantageous.
[0152] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment mainly explains the difference from other embodiments.
Claims
1. A deep learning-based outpatient pre-examination triage method, characterized in that, The method comprises: During the process of the patient end interacting with the preset interactive interface, obtaining the expression text input by the patient; extracting the symptom description words, the keywords related to triage, and the emotional polarity words from the expression text; According to the information entropy difference characteristics between the symptom description words and the keywords, the co-occurrence characteristics between the symptom description words and the emotional polarity words, and the semantic expression characteristics of the symptom description words, the temperature parameter for adjusting the attention distribution in the triage path recommendation model is determined; The expression text is input into the triage path recommendation model with the set temperature parameter to obtain the candidate triage categories and the corresponding triage paths; based on the candidate triage categories, the candidate questions are generated; the candidate questions are pushed to the patient end, and the new responses of the patient to the candidate questions are received to obtain new expression texts; The steps of calculating the temperature parameter, inputting the model to obtain the candidate triage categories and the triage paths, generating and pushing the candidate questions, and obtaining the new expression texts are repeatedly executed for multiple rounds of iterative question and answer until the preset termination condition is met; The latest expression text obtained in the last round of iteration is input into the triage path recommendation model to obtain the target triage category and the target triage path; The temperature parameter determination process comprises: Based on the information entropy difference characteristics between the symptom description words and the keywords, the symptom description sufficiency index of the patient is determined, which quantifies the coverage degree of the key information related to triage in the expression text; Based on the semantic expression characteristics of the symptom description words, the symptom description ambiguity index of the patient is determined; Based on the co-occurrence characteristics between the symptom description words and the emotional polarity words, the emotion association strength index of the patient is determined, which quantifies the statistical association strength of the co-occurrence of the emotional polarity words and the symptom description words in the text; According to a preset function relationship among the symptom description adequacy index, the symptom description ambiguity index, and the emotion association strength index, a temperature parameter is calculated; the temperature parameter is represented by the following preset function relationship: wherein b represents the symptom description adequacy index, f represents the symptom description ambiguity index, and j represents the emotion association strength index, a basic value of the temperature parameter is represented by the following preset function relationship: the temperature parameter is represented by the following preset function relationship: The symptom description sufficiency index determination process comprises: The information entropy of all symptom description words is calculated to obtain a first information entropy; The information entropy of all keywords is calculated to obtain a second information entropy; Based on the first information entropy and the second information entropy, a symptom description sufficiency index is obtained through a preset sufficiency index calculation function; the symptom description sufficiency index is represented by the following preset sufficiency index calculation function: Wherein, The symptom description sufficiency index is represented by The first information entropy is represented by The second information entropy is represented by The minimum positive offset is represented by Less than ; or Wherein Greater than ; The symptom description ambiguity index determination process comprises: Fuzzy words in the expression text are identified from the preset fuzzy word dictionary; and a plurality of triage category sets are obtained, wherein the triage category sets contain a plurality of symptom reference words; The total number of fuzzy words in the expression text is compared with the total number of symptom description words to obtain a second ratio; The symptom description words are semantically clustered to obtain at least one semantic cluster, and each semantic cluster contains a plurality of symptom description words with similar semantics; For any symptom description word in each semantic cluster, it is determined whether there is a target symptom reference word in each triage category set that meets the similarity condition with the any symptom description word; the triage category set containing the target symptom reference word is taken as a target triage category set; For each target triage category set, the total number of target symptom reference words in the target triage category set is compared with the total number of all target symptom reference words to obtain a third ratio of the target triage category set; Based on the second ratio and the third ratio, the ambiguity degree index is determined by a preset ambiguity degree calculation function. The symptom description ambiguity index is determined based on the second ratio and the third ratio by a preset ambiguity calculation function, and the method comprises the following steps: A product of the second ratio and the third ratio is calculated, and the product is taken as a priority score of the target triage category set; A preset number of target triage category sets are selected as core triage category sets in descending order of the priority score; A variance of the priority scores of all core triage category sets is calculated, and the reciprocal of the variance is normalized to obtain a concentration index; The second ratio is multiplied by the concentration index to obtain the symptom description ambiguity index; The sentiment polarity words at least include exclamation words and degree adverbs, and the emotion association strength index determination process comprises the following steps: The total number of sentiment polarity words appearing in the expression text is counted, and the appearance frequency of exclamation words and the appearance frequency of degree adverbs in the expression text are counted respectively; All pairs of symptom description words and sentiment polarity words that co-occur in a sentence or a preset context window in the expression text are identified as effective co-occurrence pairs, and the total number of effective co-occurrence pairs is counted; According to the exclamation word frequency, the degree adverb frequency and the emotion-symptom co-occurrence intensity ratio, the emotion correlation intensity index is calculated through a preset correlation intensity calculation function; the emotion correlation intensity index is calculated through the following preset correlation intensity calculation function: wherein, the emotion correlation intensity index is represented by, the total number of valid co-occurrence pairs is represented by, the total number of emotional polarity words is represented by, the logarithmic function is represented by, the base of the logarithm is represented by, and a is greater than 1, i represents the exclamation word frequency, the degree adverb frequency is represented by. 2.The outpatient pre-examination triage method based on deep learning of claim 1, wherein, The total number of effective co-occurrence pairs is divided by the total number of sentiment polarity words to obtain a sentiment-symptom co-occurrence strength ratio; The candidate question generation process comprises the following steps: Key symptom feature words associated with the candidate triage category are extracted from the expression text; The key symptom feature words, the candidate triage category, and a preset question generation prompt template are input into a preset natural language generation model; 3.The deep learning-based outpatient pre-examination triage method according to claim 2, characterized in that, A plurality of initial candidate questions output by the natural language generation model are obtained. The method further comprises the following steps: A historical dialogue record set is obtained, and the historical dialogue record set comprises a plurality of historical dialogue records, each historical dialogue record comprising a historical triage category sequence and a corresponding historical candidate question sequence generated in a plurality of historical question and answer dialogue rounds; For each initial candidate question, a historical dialogue record subset matching the current candidate triage category is retrieved from the historical dialogue record set; in the historical dialogue record subset, a historical candidate question similar in semantics to the initial candidate question is identified as a similar question; and a screening score of the initial candidate question is calculated based on statistical characteristics of the similar question in the historical dialogue record subset; 4.The outpatient pre-examination triage method based on deep learning of claim 3, wherein, From the screening scores of all initial candidate questions, an initial candidate question with the highest screening score is selected as a candidate question and pushed to the patient terminal. There is an effective historical triage category similar in meaning to the current candidate triage category in the historical triage category sequence, and the method further comprises the following steps:
5. A deep learning-based outpatient pre-examination triage system, characterized in that, The statistical characteristics include the co-occurrence frequency between the similar question and the effective historical triage category in the corresponding historical dialogue record of the similar question, the appearance frequency of the similar question in the corresponding historical dialogue record of the similar question, and the appearance round distribution characteristics of the similar question in the corresponding historical dialogue record of the similar question. The system comprises a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the method according to any one of claims 1-4 when executing the computer program.
Citation Information
Patent Citations
Internet medical triage method and system
CN114822800A
Intelligent hospital guide method and system
CN115274086A