Intelligent follow-up method and system for traditional chinese and western medicine logical voice interaction based on emotion recognition
By capturing and analyzing patients' voice signals in real time, and combining this with a rule base for the association between traditional Chinese and Western medicine emotions and symptoms to generate personalized follow-up questions, the system has solved the problems of existing systems being unable to recognize emotions and lacking logical follow-up questions, thereby improving the completeness of symptom collection and the accuracy of diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN JUNSHAN MEDICAL CORE TECHNOLOGY CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-26
AI Technical Summary
Existing auxiliary diagnostic systems cannot identify patients' emotional states, leading to incomplete symptom collection and a lack of logical questioning that integrates traditional Chinese and Western medicine, thus affecting the accuracy of diagnosis and treatment.
By capturing patients' voice signals in real time, filtering environmental noise, extracting emotional features and inputting them into an emotion recognition model, and combining them with a rule base for the association between traditional Chinese and Western medicine emotions and symptoms to generate personalized follow-up questions, symptom information is collected in real time, and secondary matching is triggered when new symptoms are identified.
It enables accurate identification of patient emotions, improves the completeness of symptom collection and the accuracy of auxiliary diagnosis and treatment, and is suitable for auxiliary diagnosis and treatment scenarios combining traditional Chinese and Western medicine.
Smart Images

Figure CN122290644A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical auxiliary diagnosis and treatment technology, and in particular to a method and system for intelligent questioning based on emotion recognition and logical voice interaction between traditional Chinese and Western medicine. Background Technology
[0002] In the field of medical auxiliary diagnosis and treatment, voice-interactive auxiliary diagnosis and treatment systems have been widely used because they can improve consultation efficiency and reduce the workload of medical staff. However, most existing auxiliary diagnosis and treatment systems can only recognize symptom keywords in the patient's voice and cannot perceive the patient's emotional state. In clinical practice, a patient's emotions are often closely related to physical symptoms, and emotional abnormalities often induce or aggravate physical symptoms.
[0003] Meanwhile, the existing system's questioning process lacks specificity and fails to conduct in-depth questioning in conjunction with the logic of traditional Chinese and Western medicine diagnosis and treatment. It neither utilizes the correlation between emotions and physical symptoms to supplement symptom collection nor meets the needs of integrated traditional Chinese and Western medicine auxiliary diagnosis and treatment in clinical practice. This easily leads to incomplete symptom collection and omission of key symptoms, thereby affecting the accuracy of auxiliary diagnosis and treatment and failing to provide complete and effective data support for subsequent diagnosis and treatment analysis. Summary of the Invention
[0004] The purpose of this invention is to propose an intelligent questioning method for integrated traditional Chinese and Western medicine logic voice interaction based on emotion recognition, aiming to solve the technical problems of existing auxiliary diagnosis and treatment voice interaction methods being unable to recognize patients' emotions and lacking integrated traditional Chinese and Western medicine logic in questioning, resulting in incomplete symptom collection.
[0005] The present invention is implemented as follows: a method for intelligent questioning based on emotion recognition in both traditional Chinese and Western medicine logical voice interaction, comprising the following steps: The patient's voice signal is captured in real time and the voice signal is filtered for environmental noise. Extract speech features that represent emotions from the filtered speech signal; The extracted speech features are input into a preset emotion recognition model. Through feature matching and probability calculation of the model, the current emotion type of the patient and the corresponding emotion recognition confidence level are output. Based on the emotion type and the confidence level of emotion recognition, as well as the preset rule base for the association between traditional Chinese and Western medicine emotions and symptoms, corresponding follow-up questions are generated. The generated follow-up questions are read aloud, and the patient's voice responses are collected in real time. The effective symptom information in the responses is extracted. The patient's emotional type, mentioned symptoms, and supplementary symptoms in the follow-up responses are associated according to a preset format. The emotional association attributes corresponding to the symptoms are labeled to form symptom-emotion association data and stored.
[0006] Furthermore, the method also includes: The system detects whether the patient's voice response contains any new symptoms not yet recorded in the symptom database. If so, a secondary matching process is triggered. The new symptom is used as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database for matching. Depending on whether the patient has identified a clear emotion, a corresponding supplementary question is generated. The supplementary question is then played and the response is collected, and the new symptom is recorded in the symptom database. The new symptom is repeatedly verified and the secondary matching is performed until no new symptoms are found or the preset maximum number of follow-up questions is reached. Finally, all collected data is associated and stored.
[0007] Another objective of this invention is to propose an intelligent questioning system for logical voice interaction between traditional Chinese and Western medicine based on emotion recognition, the system comprising: The speech signal acquisition and noise reduction module is used to capture the patient's speech signal in real time and perform environmental noise filtering on the speech signal; The speech feature extraction module is used to extract speech features that represent emotions from the filtered speech signal; The emotion recognition and confidence determination module is used to input the extracted speech features into a preset emotion recognition model, and output the patient's current emotion type and corresponding emotion recognition confidence determination through feature matching and probability calculation of the model. The follow-up questioning script generation module generates corresponding follow-up questioning scripts based on the emotion type and the confidence level of emotion recognition, as well as the preset Chinese and Western medicine emotion-symptom association rule library. The voice interaction and data integration module is used to broadcast the generated follow-up questions aloud, collect the patient's voice responses in real time, extract the effective symptom information from the responses, associate the patient's emotion type, mentioned symptoms, and supplementary symptoms in the follow-up responses according to a preset format, label the emotion association attributes corresponding to the symptoms, form symptom-emotion association data, and store it.
[0008] Furthermore, the system also includes: A new symptom secondary matching and supplementary questioning module has been added to detect whether there are new symptoms in the patient's voice response that have not been recorded in the symptom database. If so, a secondary matching process is triggered. The new symptom is used as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database. Based on whether the patient has identified a clear emotion, a corresponding supplementary question is generated. The supplementary question is then played and the response is collected, and the new symptom is recorded in the symptom database. The new symptom is repeatedly verified and secondary matching is performed until there are no new symptoms or the preset maximum number of follow-up questions is reached. Finally, all collected data is associated and stored.
[0009] Beneficial effects of the present invention This invention proposes an intelligent questioning method and system based on emotion recognition and integrated traditional Chinese and Western medicine (TCM) logic voice interaction. The method includes: voice emotion acquisition and recognition; real-time capture and filtering of patient voice signals; extraction of emotion-related voice features, inputting them into a trained and optimized emotion recognition model; outputting emotion type and confidence level, and processing according to interval judgment strategies; retrieving associated symptoms of corresponding emotions from a TCM-symptom association rule base; combining this with patient-mentioned symptoms for screening and generating personalized in-depth questioning scripts according to TCM-Western medicine diagnostic logic; executing voice interaction to complete script playback, response collection, and data processing; if new symptoms are detected, triggering secondary matching and supplementary questioning; and finally forming standardized symptom-emotion association data storage. This invention achieves accurate recognition of patient voice emotions, combines TCM and Western medicine logic for in-depth questioning, improves the completeness of symptom collection and the accuracy of auxiliary diagnosis, and is suitable for TCM-Western medicine integrated auxiliary diagnosis scenarios. Attached Figure Description
[0010] Figure 1 This is a flowchart of a preferred embodiment of the present invention: an intelligent questioning method for logical voice interaction between traditional Chinese and Western medicine based on emotion recognition. Figure 2 This is a structural diagram of a preferred embodiment of the present invention: a logical voice interaction intelligent questioning system for traditional Chinese and Western medicine based on emotion recognition. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. For ease of explanation, only the parts related to the embodiments of this invention are shown. It should be understood that the specific embodiments described herein are merely for explaining this invention and are not intended to limit this invention.
[0012] This invention proposes an intelligent questioning method and system based on emotion recognition and integrated traditional Chinese and Western medicine (TCM) logic voice interaction. The method includes: voice emotion acquisition and recognition; real-time capture and filtering of patient voice signals; extraction of emotion-related voice features, inputting them into a trained and optimized emotion recognition model; outputting emotion type and confidence level, and processing according to interval judgment strategies; retrieving associated symptoms of corresponding emotions from a TCM-symptom association rule base; combining these with symptoms already mentioned by the patient for screening; and generating personalized in-depth questioning scripts according to TCM-Western medicine diagnostic logic; executing voice interaction to complete script playback, response collection, and data processing; triggering secondary matching and supplementary questioning if new symptoms are detected; and finally forming standardized symptom-emotion association data storage. This invention achieves accurate recognition of patient voice emotions, combines TCM and Western medicine logic for in-depth questioning, improves the completeness of symptom collection and the accuracy of auxiliary diagnosis, and is suitable for TCM-Western medicine integrated auxiliary diagnosis scenarios.
[0013] Figure 1This is a flowchart of a preferred embodiment of the present invention for an intelligent questioning method based on emotion recognition for logical voice interaction between traditional Chinese and Western medicine; it includes the following steps: S1. Capture the patient's voice signal in real time and perform environmental noise filtering on the voice signal; In this embodiment of the invention, an adaptive filtering algorithm is used to filter out environmental noise to obtain a clean speech signal; S2, extract speech features representing emotions from the filtered speech signal; In this embodiment of the invention, the speech features representing emotions include at least one of the following: sound frequency (e.g., higher frequency when anxious, lower frequency when depressed), pitch fluctuation (drastic pitch changes when emotions fluctuate greatly, stable pitch when calm), speech rate (faster speech rate when irritable, slower speech rate when depressed), tone intensity (loud tone when excited, weak tone when depressed), speech pause interval (frequent and short pauses when nervous), and abnormal breathing sounds (rapid / weak). S3, input the extracted voice features into the preset emotion recognition model, and output the patient's current emotion type and corresponding emotion recognition confidence level through feature matching and attribution probability calculation of the model. The emotional types mentioned include, but are not limited to, anxiety, depression, irritability, calmness, anger, and depressive tendencies. Specifically, the extracted speech features are input into a preset emotion recognition model. Through feature matching and attribution probability calculation, the emotion type with the highest attribution probability is determined as the current patient's emotion type. The feature matching weighted score is then fused with the highest attribution probability to obtain the emotion recognition confidence score. The emotion recognition confidence score is then determined according to a preset confidence interval.
[0014] Furthermore, the probability calculation refers to the model calculating the probability of the input speech features belonging to each preset emotion type, such as anxiety, depression, irritability, calmness, anger, and depressive tendency; and taking the highest probability as the current emotion type, while weighting and fusing the highest probability with the feature matching degree to obtain the final emotion recognition confidence.
[0015] The emotion recognition model takes multi-dimensional emotion-related speech features as input. These input features are key characteristics that can characterize emotions, such as sound frequency, pitch fluctuation, speech rate, tone intensity, and pause intervals, extracted from the filtered speech signal of the patient. The model is trained and iteratively optimized using a large number of clinical speech emotion samples (5,000 to 100,000 sets) from different age groups and accents. During training, a standard emotion template is used as a reference, and the sample speech features are matched with the standard emotion template for similarity learning. At the same time, the feature matching weights and probability calculation model parameters are optimized to enable the model to recognize emotions that are adapted to the speech characteristics of different groups of people. After training and optimization, the model outputs the patient's emotion type and the corresponding emotion recognition confidence score. The process for calculating the confidence level of emotion recognition is as follows: The extracted speech features, such as sound frequency, pitch fluctuation, speech rate, tone intensity, and pause interval, are matched one by one with standard emotion templates such as anxiety, depression, and calmness to obtain individual matching scores for each feature. The feature comprehensive score is obtained by weighting and summing the individual matching scores according to preset weights. The probability of a speech feature set belonging to each preset emotion type is calculated using an emotion recognition model. The highest attribution probability and the feature comprehensive score are mutually verified to form the final emotion recognition confidence level, and the corresponding follow-up questioning strategy is executed according to the confidence level interval.
[0016] The confidence level for emotion recognition is determined as follows: If the confidence level of emotion recognition is ≥ the first threshold (e.g., 80%), it is considered high confidence. The emotion type is then confirmed and in-depth follow-up questions from both traditional Chinese and Western medicine are initiated. If the second threshold is less than or equal to the confidence level of emotion recognition and less than the first threshold, it is considered a medium confidence level, which is then judged as a suspicious emotion and gentle guided questioning is initiated. If the confidence level of emotion recognition is less than the second threshold (e.g., 60%), it is considered low confidence, and no specific emotion is determined; only standardized symptom collection is performed.
[0017] For example, taking a patient's voice as indicating anxiety with a confidence level of 88%, the process of calculating the confidence level for emotion recognition is as follows: Feature extraction was performed on the patient's speech, which included: sound frequency of 350Hz, pitch fluctuation of 80Hz, speech rate of 180 words / minute, high tone intensity, and short and frequent pauses.
[0018] Single Feature Matching and Scoring (0-100 points): The above features are compared with the standard anxiety emotion template, and scores are assigned based on the degree of match. Sound frequency matching score: 90 points; Pitch variation matching score: 85 points; Speech rate matching accuracy: 90 points; Tone intensity matching score: 85 points; Pause interval matching accuracy: 90 points; The weighted summation yields a comprehensive feature score, which is then calculated using preset weights (the sum of the weights is 1): Comprehensive Feature Score = 90 × 0.2 + 85 × 0.2 + 90 × 0.2 + 85 × 0.2 + 90 × 0.2 = 88 points The probability calculation and confidence output model simultaneously calculate the probability of attributing the speech feature combination to each emotion: Belonging anxiety probability: 88%; Probability of feeling agitated by belonging: 10%; Probability of being classified as another emotion: ≤2%; The model verifies the highest attribution probability with the feature comprehensive score, and outputs the emotion type as anxiety, with an emotion recognition confidence level of 88%.
[0019] Confidence interval determination: Because 88% ≥ 80%, it was determined to be of high confidence, the emotion type was confirmed, and in-depth follow-up questions using both traditional Chinese and Western medicine were initiated.
[0020] S4. Based on the emotion type and the confidence level of emotion recognition, as well as the preset rule base for the association between traditional Chinese and Western medicine emotions and symptoms, generate corresponding follow-up questions. The TCM and Western medicine emotion-symptom association rule base adopts a hierarchical storage structure, constructing separate TCM emotion-symptom association sub-bases and Western medicine emotion-symptom association sub-bases. The two sub-bases are independent of each other and can be called synchronously. Among them, the TCM emotion-symptom association sub-database is constructed based on the TCM syndrome differentiation and treatment theory. It stores the corresponding associations between different emotions and TCM syndrome types and physical symptoms, and clarifies the impact of emotions on the internal organs, qi and blood, and the specific physical manifestations they cause. For example, anxiety corresponds to the liver qi stagnation syndrome, and is associated with symptoms such as chest tightness, hypochondriac pain, belching, abdominal distension, irritability and anger; depression corresponds to the heart and spleen deficiency syndrome, and is associated with symptoms such as fatigue, sallow complexion, insomnia and dreaminess, loss of appetite and other symptoms; irritability corresponds to the liver fire excess syndrome, and is associated with symptoms such as headache, red eyes, bitter taste in the mouth, insomnia and other symptoms.
[0021] The Western Medicine Emotion-Symptom Association Sub-Database is constructed based on modern medical theories. It stores the corresponding relationships between different emotions and disease tendencies and physical symptoms, clarifying the physiological changes and related symptoms caused by emotional abnormalities. For example, low mood corresponds to depressive tendencies and is associated with symptoms such as insomnia, fatigue, poor concentration, and loss of appetite; anxiety corresponds to anxiety tendencies and is associated with symptoms such as chest tightness, palpitations, dizziness, sweating, and insomnia; irritability corresponds to emotional stress response and is associated with symptoms such as blood pressure fluctuations, headaches, and palpitations.
[0022] In this embodiment of the invention, the traditional Chinese and Western medicine emotion-symptom association rule base can also support manual editing, updating and expansion, and can add or adjust the association between emotion and symptoms based on clinical diagnosis and treatment experience and medical research progress.
[0023] The step of generating corresponding follow-up questions based on the emotion type and emotion recognition confidence level, as well as a preset traditional Chinese and Western medicine emotion-symptom association rule base, includes: S41. Based on the patient's emotion type and confidence level, and combined with the physical symptoms mentioned by the patient in the voice interaction, simultaneously retrieve all associated symptoms of the corresponding emotion from the Chinese and Western medicine emotion-symptom association rule base. S42, screen the retrieved related symptoms (e.g., remove symptoms that the patient has actively mentioned, retain related symptoms that have not been mentioned, and ensure that the follow-up questions are not repetitive and are targeted), and adjust the intensity of follow-up questions based on the confidence level of the emotion assessment. In this embodiment of the invention, the higher the confidence level, the more specific the follow-up questions should be. When the confidence level is in a reasonable range (60%-79%), the follow-up questions should be mainly guiding, avoiding excessive follow-up questions.
[0024] S43, generate corresponding follow-up questions according to the corresponding diagnosis and treatment logic of traditional Chinese medicine and Western medicine; In this embodiment of the invention, the dialogue in the TCM diagnosis and treatment logic focuses on asking follow-up questions about symptoms related to the syndrome type, which is in line with the needs of TCM syndrome differentiation; for example, in response to anxiety and the mention of chest tightness, the dialogue generates questions such as "Do you experience chest tightness accompanied by rib pain? Does chest tightness worsen when you are emotionally agitated? Do you experience belching or abdominal distension?". In Western medicine's diagnostic logic, the dialogue focuses on asking follow-up questions about symptoms related to the disease tendency, which aligns with the needs of Western medicine diagnosis; for example, in response to anxiety and the mention of chest tightness, the dialogue might ask, "Besides chest tightness, do you experience palpitations, dizziness, or sweating? Are the episodes of chest tightness related to emotional fluctuations? How long does each episode last?"
[0025] Specifically, for high or medium confidence levels, all associated symptoms corresponding to the emotion type are retrieved from the pre-defined TCM-Western medicine emotion-symptom association rule base. Combined with the physical symptoms already mentioned by the patient, the mentioned symptoms are removed, and the unmentioned associated symptoms are retained. The intensity of follow-up questions is adjusted according to the confidence level of emotion recognition. Specific follow-up questions are generated for high confidence levels, and guided follow-up questions are generated for medium confidence levels. Corresponding TCM and Western medicine follow-up questions are generated according to the TCM and Western medicine diagnosis and treatment logic, respectively. For low confidence levels, only the patient's chief complaint symptoms are extracted and standardized, structured questioning is performed. For example, when the confidence level of emotion recognition is below 60%, no emotion judgment, no emotion association inference, and no in-depth TCM-Western medicine emotion-symptom follow-up questions are performed. Only standardized, structured questioning is performed based on the physical symptoms already mentioned by the patient. The specific execution logic is as follows: No emotional association is performed, and no traditional Chinese medicine or Western medicine emotion-symptom association rule bases are invoked. Only the patient's chief complaint is extracted. For example, if a patient says, "I have chest discomfort and can't sleep well," NLP technology is used to extract the content as chest tightness / chest discomfort and insomnia. This leads to a standardized symptom-based consultation process where each chief complaint is followed up with questions along fixed dimensions, without involving emotions.
[0026] S5 will broadcast the generated follow-up questions aloud, collect the patient's voice responses in real time, and extract the effective symptom information from the responses; associate the patient's emotional type, mentioned symptoms, and supplementary symptoms in the follow-up responses according to a preset format, label the emotional association attributes (TCM syndrome type / Western medicine disease tendency) corresponding to the symptoms, form symptom-emotion association data and store it. For example, in one embodiment of the present invention, the patient's voice response is collected in real time, and the voice recognition function is called simultaneously to convert the patient's voice response into text data, filter out invalid responses (such as disordered voice and irrelevant statements), and extract valid response information (such as having rib pain, not having palpitations, etc.).
[0027] Furthermore, in this embodiment of the invention, the speed of the voice broadcast can be adaptively adjusted according to the patient's speech rate, and a voice replay function is supported; Furthermore, in this embodiment of the invention, after step S5, step S6 is also included. S6, detect whether there are any new symptoms in the patient's voice response that have not been recorded in the symptom database collected this time. If so, trigger a secondary matching process: use the new symptom as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database for matching, generate supplementary follow-up questions based on whether the patient has identified a clear emotion, execute supplementary question broadcast and response collection, and record the new symptom into the collected symptom database; repeatedly verify the new symptom and perform secondary matching until there are no new symptoms or the preset maximum number of follow-up questions is reached, and finally store all collected data in association; in this embodiment of the invention, the preset maximum number of follow-up questions can be 2.
[0028] In this embodiment of the invention, the specific implementation of triggering the secondary matching process is as follows: S61. The physical symptoms in the patient's response are compared in real time with the symptoms collected in the symptom database (symptoms initially mentioned by the patient + symptoms supplemented in the first round of follow-up questions). If a symptom that has not been recorded is found (such as the patient mentioning "bitter taste in mouth" after chest tightness in the first round of follow-up questions), it is determined to be a new symptom and a second matching is triggered. The collected symptom database includes symptoms initially mentioned by the patient and symptoms supplemented during the first round of follow-up questioning. S62 uses newly added symptoms as the core keywords and retrieves the preset Chinese and Western medicine emotion-symptom association rule library for matching; If the patient has previously identified a clear emotion (such as anxiety), that is, the confidence level of emotion identification is ≥ the first preset threshold (such as confidence level ≥ 80%), then the intersection rule of new symptoms + identified emotions will be matched first (such as "bitter taste in the mouth + anxiety" corresponds to the syndrome of excessive liver fire in traditional Chinese medicine). If the patient previously had suspicious emotions, that is, the confidence level of emotion recognition is between the second preset threshold and the first preset threshold, it will also be treated as an already identified emotion, and the intersection rule of new symptoms + already identified suspicious emotions will be prioritized. If the patient does not identify a specific emotion, i.e., the confidence level of emotion identification is less than the second preset threshold (e.g., confidence level < 60%), then only the general TCM-Western medicine association rule for this new symptom is matched, without any emotion indication.
[0029] S63 generates targeted scripts based on the matching results; When there is a clear or suspected emotional indication (i.e., the confidence level of emotion recognition is ≥ the first preset threshold or the second threshold is ≤ the confidence level of emotion recognition < the first threshold), the dialogue is generated in combination with the logic of traditional Chinese and Western medicine diagnosis and treatment (e.g., "Does the bitter taste in your mouth that you mentioned worsen when you are anxious or irritable?"). When there is no emotional focus (emotion recognition confidence level < second preset threshold), only standardized symptom follow-up questions are asked (such as "How long have you had a bitter taste in your mouth? Is it accompanied by dry mouth?"), and the verification script does not repeat the questions asked in the first round.
[0030] S64: Broadcast the supplementary script and collect responses. Enter the new symptoms and response information into the collected symptom database. Check again for new symptoms. If no new symptoms are found, terminate the process. If so, repeat the matching (preset maximum 2 times to avoid excessive questioning). Finally, associate and store all collected data (initial symptoms + first round of supplements + second round of supplements + emotional information) to complete the complete collection.
[0031] Example 1 This embodiment provides a specific implementation of an intelligent questioning method based on emotion recognition for logical voice interaction between traditional Chinese and Western medicine. In this embodiment, the system has been fully deployed and pre-configured, specifically including: The system architecture adopts a "front-end interactive terminal (medical tablet) + back-end processing server (high-performance computing server) + data storage unit (encrypted database)" structure. A hybrid machine learning model based on CNN-LSTM was deployed as the emotion recognition model. After debugging, the confidence thresholds were set as follows: ≥80% high confidence, 60%-79% medium confidence, and <60% low confidence. An adaptive filtering algorithm was used to filter out environmental noise below 50dB, the speech sampling frequency was 16kHz, and the speech recognition conversion accuracy was ≥95%. A rule database linking emotions and symptoms in Traditional Chinese Medicine (TCM) and Western medicine has been established. The TCM sub-database includes: Anxiety → Liver Qi Stagnation Syndrome → Chest tightness, hypochondriac pain, belching, abdominal distension, irritability, frequent sighing; Depression → Heart and Spleen Deficiency Syndrome → Fatigue, sallow complexion, insomnia, poor appetite, forgetfulness; Irritability → Liver Fire Excess Syndrome → Headache, red eyes, bitter taste in mouth, insomnia, constipation. The Western medicine sub-database includes: Anxiety → Anxiety Tendency → Chest tightness, palpitations, dizziness, sweating, insomnia, muscle tension; Depression → Depressive Tendency → Insomnia, fatigue, poor concentration, decreased appetite, weight fluctuations; Irritability → Emotional Stress Response → Blood Pressure Fluctuations, headache, palpitations, chest tightness, nausea. The rule database is open to manual editing by medical staff. Preset voice interaction parameters: voice playback speed is 120 words / minute, supports voice command "say it again" to trigger replay, and the maximum number of follow-up questions in the second matching is 2.
[0032] In this example, the patient is in a high-confidence anxiety scenario, and the specific implementation steps are as follows: Step 1: Voice emotion collection and recognition; The patient activated the voice consultation function through the front-end medical tablet. The system read out the guiding words, "Please describe your recent physical discomfort symptoms with your voice." The patient replied in voice, "I have been feeling tightness in my chest for the past week, and I can't sleep well at night. I am very anxious."
[0033] The system captures the patient's speech signal through a microphone at a sampling frequency of 16kHz; it uses an adaptive filtering algorithm to filter out environmental noise to obtain a clean speech signal; it extracts emotion-related features from the clean signal, which are: sound frequency 350Hz (200-300Hz higher than the calm state), pitch fluctuation 80Hz (30-50Hz higher than the calm state), speech rate 180 words / minute (120-150 words / minute higher than the normal speech rate), and short and frequent pauses; the emotion determination unit inputs the above features into a CNN-LSTM hybrid model, the model outputs the emotion type as "anxiety", the emotion recognition confidence is 88% (≥80%, high confidence), and transmits the result along with the symptoms mentioned by the patient (chest tightness, insomnia).
[0034] Step 2, generating in-depth follow-up questions; Based on the emotion of "anxiety", the corresponding associated symptoms are retrieved simultaneously from the traditional Chinese medicine and Western medicine emotion-symptom association rule base: the traditional Chinese medicine sub-base retrieves chest tightness, rib pain, belching, abdominal distension, irritability, and frequent sighing; the Western medicine sub-base retrieves chest tightness, palpitations, dizziness, sweating, insomnia, and muscle tension.
[0035] Excluding symptoms already mentioned by the patient such as chest tightness and insomnia, retaining symptoms not mentioned in Traditional Chinese Medicine: hypochondriac pain, belching, abdominal distension, irritability, and frequent sighing; and symptoms not mentioned in Western Medicine: palpitations, dizziness, sweating, and muscle tension; due to the high confidence level of 88%, the follow-up question intensity was set to specific.
[0036] Generate simplified dialogue based on the logic of traditional Chinese medicine and Western medicine diagnosis and treatment: Traditional Chinese medicine dialogue: "Do you experience chest tightness accompanied by rib pain? Do you usually experience belching or abdominal bloating? Does your chest tightness worsen when you are emotionally agitated?"; Western medicine dialogue: "Besides chest tightness and insomnia, do you experience palpitations, dizziness, or sweating? Do you usually feel muscle tension or tightness?", and transmit the dialogue to the interactive execution module.
[0037] Step 3, voice interaction execution; The questioning script was read aloud to the patient at a speed of 120 words per minute, following the order of "Traditional Chinese Medicine first, then Western Medicine." The patient replied in voice, "I have bloating and pain in my rib area, occasional belching, but no palpitations, dizziness, or sweating."
[0038] The speech recognition function was used to convert the data into text. After filtering out invalid expressions, valid symptom information was extracted: hypochondriac pain, occasional belching, no palpitations, no dizziness, and no sweating. Emotional association attributes were also marked: hypochondriac pain and belching correspond to the TCM syndrome of liver qi stagnation, while no palpitations, dizziness, and no sweating correspond to the exclusion symptoms of anxiety tendency in Western medicine. The emotion type, initial symptoms, and supplementary symptoms were preliminarily associated and organized to form preliminary symptom-emotion association data.
[0039] Step 4, follow up with additional questions through secondary matching; The symptoms reported by patients are compared in real time with the collected symptom database (chest tightness, insomnia). No new symptoms that have not been entered are detected, so the secondary matching process is not triggered.
[0040] Step 5: Data storage and process termination; The final symptom-emotion correlation data were organized according to the preset format: Patient emotion - anxiety (confidence 88%); TCM correlation symptoms - chest tightness (initial), hypochondriac pain (supplementary), belching (supplementary), syndrome type - liver qi stagnation; Western medicine correlation symptoms - chest tightness (initial), insomnia (initial), excluding palpitations, dizziness, and sweating, disease tendency - anxiety tendency.
[0041] The data is encrypted and transmitted to a backend encrypted database for persistent storage. Then, a voice announcement is made saying, "Symptom collection is complete. Thank you for your cooperation," completing the voice interaction and symptom collection process. Patients can view the collected symptom information on a medical tablet, and medical staff can retrieve the data through the backend processing server for auxiliary diagnosis and analysis.
[0042] Example 2 This embodiment is a scenario where the confidence level of anxiety is high and a secondary matching is triggered. Based on the system preset configuration in embodiment 1, the patient's initial voice response is "chest tightness, poor sleep, and a feeling of unease". The emotion recognition model outputs the emotion type "anxiety" with a confidence level of 70% (60%-79%, medium confidence level).
[0043] Generate guiding questioning techniques: Traditional Chinese medicine question: "Do you usually feel discomfort in your rib area or feel like sighing?"; Western medicine question: "Do you occasionally experience palpitations or dizziness?"
[0044] The patient responded via voice that "I occasionally feel discomfort in my flanks and often have a bitter taste in my mouth." After extracting the valid symptom information, it was detected that "bitter taste in the mouth" was a new symptom that had not been entered into the symptom database and had not reached the maximum number of follow-up questions (2 times), thus triggering a secondary matching process.
[0045] Since the patient has identified a clear emotion (anxiety, confidence level 70% ≥ 60%), the intersection rule of "bitter taste in mouth + anxiety" is matched first. The corresponding syndrome of excessive liver fire in traditional Chinese medicine and the corresponding emotional stress response in Western medicine are retrieved to generate supplementary follow-up questions such as "Does your bitter taste in mouth occur occasionally or every day? Does the bitter taste in mouth worsen when you are emotionally agitated?"
[0046] The system reads the supplementary script, and the patient replies, "It happens every day, and it gets worse when I'm emotionally agitated." The system extracts this information and enters it into the symptom database. After a second check, no new symptoms are found, and the second matching process is terminated.
[0047] Finally, the symptom-emotion association data was organized and stored to complete the symptom collection. The data was then added to associate the bitter taste in the mouth with the TCM syndrome of excessive liver fire and the Western medicine emotional stress response. The symptom collection results were improved by combining the initial anxiety.
[0048] Example 3 This embodiment is a low-confidence scenario. Based on the system preset configuration in embodiment 1, the patient's initial voice response is "I feel discomfort in my chest and cannot sleep well". The voice features extracted by the emotion recognition model have a low matching degree with the standard emotion template, and the output emotion recognition confidence is 55% (<60%, low confidence).
[0049] The emotional association is disabled, and the traditional Chinese and Western medicine emotion-symptom association rule base is not accessed. Only the patient's chief symptoms, chest tightness and insomnia, are extracted using NLP technology. A standardized symptom consultation process is then initiated, generating follow-up questions for each symptom according to fixed dimensions: For chest tightness, "How long have you had chest tightness? Where is the chest tightness located? Is it persistent or intermittent?"; For insomnia, "How long have you had insomnia? How many hours of sleep do you get each day? Do you have difficulty falling asleep or wake up easily?".
[0050] The standardized questioning script is broadcast, the patient's response is collected and effective symptom information is extracted. Only the chief complaint is supplemented and no in-depth questioning related to emotions is conducted. Finally, the collected symptom data is organized and stored to complete the symptom collection.
[0051] Corresponding to the above embodiment of the intelligent questioning method for logical voice interaction between traditional Chinese and Western medicine based on emotion recognition, Figure 2 The diagram shows a structural block diagram of an intelligent questioning system based on emotion recognition for logical voice interaction between traditional Chinese and Western medicine, provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0052] The speech signal acquisition and noise reduction module is used to capture the patient's speech signal in real time and perform environmental noise filtering on the speech signal; The speech feature extraction module is used to extract speech features that represent emotions from the filtered speech signal; The emotion recognition and confidence determination module is used to input the extracted speech features into a preset emotion recognition model, and output the patient's current emotion type and corresponding emotion recognition confidence determination through feature matching and probability calculation of the model. The follow-up questioning script generation module generates corresponding follow-up questioning scripts based on the emotion type and the confidence level of emotion recognition, as well as the preset Chinese and Western medicine emotion-symptom association rule library. The voice interaction and data integration module is used to broadcast the generated follow-up questions aloud, collect the patient's voice responses in real time, extract the effective symptom information from the responses, associate the patient's emotion type, mentioned symptoms, and supplementary symptoms in the follow-up responses according to a preset format, label the emotion association attributes corresponding to the symptoms, form symptom-emotion association data, and store it.
[0053] Furthermore, the system also includes: A new symptom secondary matching and supplementary questioning module is added to detect whether there are new symptoms in the patient's voice response that have not been recorded in the symptom database. If so, a secondary matching process is triggered. The new symptom is used as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database for matching. Depending on whether the patient has identified a clear emotion, a corresponding supplementary question is generated. The supplementary question is then played and the response is collected, and the new symptom is recorded in the symptom database. The new symptom is repeatedly verified and secondary matching is performed until there are no new symptoms or the preset maximum number of follow-up questions is reached. Finally, all collected data is associated and stored.
[0054] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by program instructions and related hardware. The program can be stored in a computer-readable storage medium, such as ROM, RAM, disk, optical disk, etc.
[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A voice-interactive intelligent questioning method based on emotion recognition, characterized in that, Includes the following steps: The patient's voice signal is captured in real time and the voice signal is filtered for environmental noise. Extract speech features that represent emotions from the filtered speech signal; The extracted speech features are input into a preset emotion recognition model. Through feature matching and probability calculation of the model, the current emotion type of the patient and the corresponding emotion recognition confidence level are output. Based on the emotion type and the confidence level of emotion recognition, as well as the preset rule base for the association between traditional Chinese and Western medicine emotions and symptoms, corresponding follow-up questions are generated. The generated follow-up questions are read aloud, and the patient's voice responses are collected in real time. The effective symptom information in the responses is extracted. The patient's emotional type, mentioned symptoms, and supplementary symptoms in the follow-up responses are associated according to a preset format. The emotional association attributes corresponding to the symptoms are labeled to form symptom-emotion association data and stored.
2. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 1, characterized in that, The method further includes: The system detects whether the patient's voice response contains any new symptoms not yet recorded in the symptom database. If so, a secondary matching process is triggered. The new symptom is used as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database for matching. Depending on whether the patient has identified a clear emotion, a corresponding supplementary question is generated. The supplementary question is then played and the response is collected, and the new symptom is recorded in the symptom database. The new symptom is repeatedly verified and the secondary matching is performed until no new symptoms are found or the preset maximum number of follow-up questions is reached. Finally, all collected data is associated and stored.
3. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 1, characterized in that, The speech features that characterize emotions include at least one of the following: sound frequency, pitch variation, speech rate, tone intensity, speech pause intervals, and abnormal breathing sounds.
4. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 1, characterized in that, The process for calculating the confidence level of emotion recognition is as follows: The extracted speech features are matched one by one with the standard emotion template to obtain the individual matching score for each feature; The feature comprehensive score is obtained by weighting and summing the individual matching scores according to preset weights. The probability of a speech feature set belonging to each preset emotion type is calculated using an emotion recognition model. The highest attribution probability and the feature comprehensive score are mutually verified to form the final emotion recognition confidence level, and the corresponding follow-up questioning strategy is executed according to the confidence level interval.
5. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 4, characterized in that, The confidence level for emotion recognition is determined as follows: If the confidence level of emotion recognition is greater than or equal to the first threshold, it is considered high confidence. The emotion type is then confirmed and in-depth follow-up questions from both traditional Chinese and Western medicine are initiated. If the second threshold is less than or equal to the confidence level of emotion recognition and less than the first threshold, it is considered a medium confidence level, which is then judged as a suspicious emotion and gentle guided questioning is initiated. If the confidence level of emotion recognition is less than the second threshold, it is considered low confidence, and no specific emotion is determined; only standardized symptom collection is performed.
6. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 5, characterized in that, The traditional Chinese medicine and Western medicine emotion-symptom association rule base adopts a hierarchical storage structure, and separately constructs a traditional Chinese medicine emotion-symptom association sub-base and a Western medicine emotion-symptom association sub-base. Among them, the TCM emotion-symptom association sub-database is constructed based on the TCM syndrome differentiation and treatment theory. It stores the corresponding association between different emotions and TCM syndrome types and physical symptoms, and clarifies the influence of emotions on the internal organs, qi and blood and the specific physical manifestations they cause. The Western Medicine Emotion-Symptom Association Sub-Database is constructed based on modern medical theories. It stores the corresponding relationships between different emotions and disease tendencies and physical symptoms, and clarifies the physiological changes and related symptoms caused by emotional abnormalities.
7. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 1, characterized in that, The step of generating corresponding follow-up questions based on the emotion type and emotion recognition confidence level, as well as a preset traditional Chinese and Western medicine emotion-symptom association rule base, includes: Based on the patient's emotion type and confidence level, and combined with the physical symptoms mentioned by the patient in the voice interaction, all related symptoms of the corresponding emotion are simultaneously retrieved from the Chinese and Western medicine emotion-symptom association rule base. The retrieved related symptoms are screened, and the intensity of follow-up questions is adjusted based on the confidence level of the emotion assessment. Based on the corresponding diagnostic and treatment logic of traditional Chinese medicine and Western medicine, corresponding follow-up questions are generated respectively.
8. The intelligent questioning method for voice interaction based on emotion recognition as described in claim 2, characterized in that, The specific implementation of triggering the secondary matching process is as follows: The physical symptoms reported by patients are compared with the symptom database collected this time in real time. If a symptom that has not been recorded is found, it is determined to be a new symptom and a second matching is triggered. Using newly added symptoms as the core keywords, the system retrieves and matches them against a pre-defined database of traditional Chinese and Western medicine emotion-symptom association rules. If the confidence level of emotion recognition is greater than or equal to the first preset threshold, then the intersection rule of new symptoms + recognized emotions will be matched first. If the confidence level of emotion recognition is between the second preset threshold and the first preset threshold, it will be treated as an already recognized emotion, and the intersection rule of new symptoms + already recognized suspicious emotions will be prioritized. If the confidence level of emotion recognition is less than the second preset threshold, then only the general Chinese and Western medicine association rules for the new symptom are matched, without emotion indication; Generate targeted scripts based on the matching results: When the confidence level of emotion recognition is greater than or equal to the first preset threshold or the second threshold is less than or equal to the confidence level of emotion recognition and less than the first threshold, the script is generated by combining the logic of traditional Chinese and Western medicine diagnosis and treatment. When the confidence level of emotion recognition is less than the second preset threshold, only standardized symptom follow-up questions are asked, and the verification script does not repeat the questions asked in the first round. The system broadcasts supplementary scripts and collects responses. New symptoms and response information are entered into the symptom database. The system then checks for new symptoms. If no new symptoms are found, the process is terminated. If new symptoms are found, the matching process is repeated. Finally, all collected data is associated and stored to complete the full data collection.
9. A voice-interactive intelligent questioning system based on emotion recognition, characterized in that, The system includes: The speech signal acquisition and noise reduction module is used to capture the patient's speech signal in real time and perform environmental noise filtering on the speech signal; The speech feature extraction module is used to extract speech features that represent emotions from the filtered speech signal; The emotion recognition and confidence determination module is used to input the extracted speech features into a preset emotion recognition model, and output the patient's current emotion type and corresponding emotion recognition confidence determination through feature matching and probability calculation of the model. The follow-up questioning script generation module generates corresponding follow-up questioning scripts based on the emotion type and the confidence level of emotion recognition, as well as the preset Chinese and Western medicine emotion-symptom association rule library. The voice interaction and data integration module is used to broadcast the generated follow-up questions aloud, collect the patient's voice responses in real time, extract the effective symptom information from the responses, associate the patient's emotion type, mentioned symptoms, and supplementary symptoms in the follow-up responses according to a preset format, label the emotion association attributes corresponding to the symptoms, form symptom-emotion association data, and store it.
10. The voice-interactive intelligent questioning system based on emotion recognition as described in claim 9, characterized in that, The system also includes: A new symptom secondary matching and supplementary questioning module has been added to detect whether there are new symptoms in the patient's voice response that have not been recorded in the symptom database. If so, a secondary matching process is triggered. The new symptom is used as the core keyword to retrieve the traditional Chinese and Western medicine emotion-symptom association rule database. Based on whether the patient has identified a clear emotion, a corresponding supplementary question is generated. The supplementary question is then played and the response is collected, and the new symptom is recorded in the symptom database. The new symptom is repeatedly verified and secondary matching is performed until there are no new symptoms or the preset maximum number of follow-up questions is reached. Finally, all collected data is associated and stored.