Closed-loop medical voice interaction system and method
By integrating voice generation, acquisition, understanding and management modules into a closed-loop medical voice interaction system, and combining real-time data flow control and emotion recognition, the non-closed-loop interaction problem of existing systems is solved, realizing a natural and coherent intelligent consultation experience, and improving the efficiency and user experience of telemedicine.
Patent Information
- Application Number
- CN202510921415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing medical voice interaction systems lack closed-loop interaction capabilities, with each functional module operating independently. This results in a discontinuous and mechanical consultation process, failing to achieve the natural and smooth experience of a real doctor. Furthermore, there are response delays and data transmission delays, affecting the continuity of the conversation and the patient experience.
A closed-loop medical voice interaction system is adopted, which deeply integrates voice generation, acquisition, understanding, management and data flow control. It uses BERT encoder and CRF decoder to identify medical entities, combines Transformer and WaveGAN to generate natural speech, and adopts WebSocket protocol to realize close collaboration between modules. It can perceive the dialogue status in real time and dynamically adjust the consultation strategy, and introduces emotion recognition module and consistency verification mechanism.
It achieves a natural consultation process similar to that of a real doctor, with end-to-end latency controlled within 500 milliseconds. The system can accurately identify key information, avoid repeated questions or omissions, perceive the patient's emotions in real time and adjust the consultation rhythm, and generate a natural and smooth consultation summary, thereby improving the user experience and consultation efficiency of telemedicine.
Smart Images

Figure CN120977301A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing and natural language processing, and particularly relates to a closed-loop medical speech interaction system and method. BACKGROUND
[0002] With the rapid development of telemedicine and smart healthcare, the demand for speech interaction technology in the medical consultation field is growing. However, the biggest problem of existing medical speech interaction systems is the lack of true closed-loop interaction capability, and each functional module is independent and lacks cooperation, resulting in a discontinuous and mechanical whole consultation process, which is far from the natural and smooth experience of real doctor consultation.
[0003] The existing system usually adopts a simple process of "speech recognition-text processing-template reply", each link runs independently, and lacks a feedback mechanism. When the patient says "I have a headache for three days", the system can only recognize and record this information, but it cannot immediately ask for key information such as the specific location, degree, and accompanying symptoms of the pain, like a real doctor. More seriously, when the patient describes contradictory information in the subsequent conversation, such as "it started yesterday" and then mentions "it has been a week", the system cannot detect and actively clarify, but can only passively record these conflicting information.
[0004] This non-closed-loop interaction mode is also reflected in the system's inability to dynamically adjust the consultation strategy according to the progress of the conversation. Real doctors will flexibly adjust the inquiry direction according to the patient's answers, while existing systems often mechanically inquire according to a fixed list of questions, even if some information has been mentioned in previous answers, the system will still ask repeatedly, wasting time and affecting the patient's experience. At the same time, due to the lack of global control of the conversation state, the system cannot judge when the information collection is sufficient, often resulting in incomplete information due to premature end of the consultation, or patient fatigue due to excessive inquiry.
[0005] There is a serious delay and synchronization problem in the data transmission between the modules of the existing system. From the patient speaking to hearing the system's reply, it goes through multiple links such as speech collection, transmission, recognition, understanding, decision-making, synthesis, and playback, and the processing time of each link adds up, resulting in a whole response delay often exceeding several seconds, completely destroying the coherence of the conversation. Moreover, when the network jitters or a module processes slowly, the whole system will appear to be stuck, the speech will be interrupted, and in severe cases, the conversation will be interrupted.
[0006] Therefore, it is urgent to build a real closed-loop medical voice interaction system, which needs to realize the deep integration and real-time collaboration of voice generation, collection, understanding, management and other links, can conduct natural, coherent and intelligent inquiry dialogue like real doctors, while ensuring millisecond response speed and accurate perception of patient emotions, and truly realize the intelligent medical inquiry experience of "understanding, accurate answer, fast response and temperature". SUMMARY
[0007] To overcome the shortcomings of the prior art, the present application provides a closed-loop medical voice interaction system and method, which can flexibly select an inquiry path according to the current dialogue state. The system no longer asks questions in a fixed order, but dynamically decides the next question according to the collected information, making the inquiry process more efficient and targeted.
[0008] To achieve the above-mentioned purpose, the present application provides a closed-loop medical voice interaction system and method, which comprises a voice generation unit, a voice collection unit, a semantic understanding unit, a dialogue management unit and a data flow controller.
[0009] The voice generation unit comprises a text encoder and an acoustic model, the text encoder converts the inquiry text into a 512-dimensional semantic vector, and the acoustic model generates a 16kHz sampling rate voice signal based on the semantic vector and is divided into 320 sampling point voice frames according to 20ms time length;
[0010] The voice collection unit comprises a VAD detector and a voice recognizer, the VAD detector uses an energy threshold of-40dB and a zero-crossing rate threshold of 0.25 to identify valid voice segments, and the voice recognizer converts the voice segments into text using CTC decoding;
[0011] The semantic understanding unit comprises a BERT encoder and a CRF decoder, the BERT encoder generates a 768-dimensional context vector, and the CRF decoder performs sequence labeling based on a predefined 13-class medical entity label;
[0012] The dialogue management unit maintains a state vector S=[s1,s2,...,si,...,sN] containing symptom slot, time slot and part slot, where si∈[0,1] represents the filling degree of the i-th slot, and when si=1, it is determined that the information collection is complete; 15 i
[0013] The data flow controller establishes a bidirectional channel through the WebSocket protocol, coordinates the data transmission between each unit through the queue buffer mechanism, and ensures that the end-to-end delay is less than 500ms.
[0014] Further, the voice generation unit specifically comprises:
[0015] The inquiry content generation module constructs an inquiry decision tree with a depth of 5 layers and a node number of 127, each node containing a {question text, sub-node pointer, transition condition} triple, and selects the next node according to the state vector S:
[0016] When s1<0.5, the chief complaint collection branch is selected
[0017] When 0.5≤s1<0.8 and s2<0.5, the symptom details branch is selected
[0018] When , the summary confirmation branch is selected;
[0019] The acoustic feature generation module adopts a Transformer structure, including 6 layers of encoders and 6 layers of decoders, and maps the text sequence T=[t1,t2,...,t n ] into 80-dimensional mel spectrum F=[f1,f2,...,f m ], where m=n×k, k is a time length prediction factor; where t i represents the i-th character in the text sequence, and f i represents the i-th mel spectrum frame (80-dimensional vector);
[0020] The vocoder module adopts a WaveGAN generator, inputs the 80-dimensional mel spectrum, and generates a 16kHz audio waveform through 5 layers of up-sampling convolution layers, with convolution kernel sizes of [16, 16, 4, 4, 4] and step sizes of [8, 8, 2, 2, 2].
[0021] Further, the voice collection unit specifically includes:
[0022] The audio preprocessing module adopts spectral subtraction for noise reduction:
[0023] Y(ω)=X(ω)-α·N(ω)
[0024] Where X(ω) is the input spectrum, N(ω) is the noise estimate, and α=2.5 is the over-subtraction factor;
[0025] The human voice activity detection module calculates the energy E j and the zero-crossing rate Z j of each frame:
[0026]
[0027] Where L=320 is the frame length (corresponding to 20ms, 16kHz sampling rate);
[0028] x j [n] the n-th sampling point of the j-th frame
[0029] sgn(x) is a sign function, defined as:
[0030] When E j -40dB is marked as speech, when Z j >0.25 is confirmed as voiced;
[0031] The endpoint detection module adopts a double threshold method, sets a high threshold and a low threshold Wherein:
[0032] is the average energy of M frames;
[0033] is the energy standard deviation
[0034] The start and end point detection is realized by a state machine.
[0035] Further, the medical entity recognition of the semantic understanding unit adopts the following label system:
[0036] B-SYMPTOM / I-SYMPTOM: symptom entity, such as "headache", "fever";
[0037] B-BODY / I-BODY: body part, such as "head", "chest";
[0038] B-TIME / I-TIME: time expression, such as "three days ago", "last night";
[0039] B-DEGREE / I-DEGREE: degree description, such as "severe", "mild";
[0040] B-DRUG / I-DRUG: drug name O: non-entity label;
[0041] The Viterbi algorithm is used for optimal label sequence decoding, the transition probability matrix , and the emission probability is calculated by the last layer of BERT.
[0042] Further, the state updating mechanism of the dialogue management unit is:
[0043] Define the state vector S=[s1,s2,...,s 15 ], wherein:
[0044] s1-s5: chief complaint related slots (symptom name, occurrence time, duration, severity, frequency of onset);
[0045] s6-s 10 : accompanying symptom slots;
[0046] s 11 -s 13 : History slot (disease history, medication history, allergy history);
[0047] s 14 -s 15 : Other information slot;
[0048] Update rule: when the entity corresponding to slot i is detected, s i = min(s i + c i , 1.0);
[0049] where c i is the entity recognition confidence.
[0050] Further, it further includes a sentiment recognition module:
[0051] Extract prosodic feature vector P = [p1, p2,..., p6] from the speech signal;
[0052] p1: Mean of fundamental frequency (normalized to the range of 70-400Hz)
[0053] p2: Standard deviation of fundamental frequency
[0054] p3: Mean of energy (normalized to 0-1)
[0055] p4: Standard deviation of energy
[0056] p5: Speech rate (phonemes per second, normalized)
[0057] p6: Pause rate (pause duration / total duration)
[0058] Map feature vector P to 5 sentiment labels: calm, anxious, doubtful, eager, negative, using SVM classifier (RBF kernel, γ = 0.1, C = 1.0).
[0059] Further, the data flow controller includes:
[0060] A ring buffer with a capacity of 100 frames, managed by read and write pointers, and triggered when (write pointer - read pointer) > 80 for flow control;
[0061] Timestamp synchronization mechanism, each speech frame carries a timestamp t i = t0 + i × 20ms, where t0 is the Unix timestamp (milliseconds) when the system starts playing;
[0062] Frame loss recovery strategy, when the frame sequence number is detected to be discontinuous, linear interpolation is used to reconstruct the missing frame:
[0063] f lost= 0.5 x (f prev + f next )
[0064] A closed-loop medical voice interaction method, applicable to the closed-loop medical voice interaction system described above, comprising the following steps:
[0065] S1: initialize state vector Generate the first question text according to the root node of the diagnosis decision tree;
[0066] S2: Convert the text to speech through the acoustic model, and push it through WebSocket after cutting it into 20ms, and the client receives and plays it;
[0067] S3: Collect user voice, start recording after VAD detects the starting point, stop after detecting the ending point, and perform ASR recognition on the valid voice segment;
[0068] S4: Perform medical entity recognition on the recognized text to extract entity list ε = {(entity text, label, position, confidence)};
[0069] S5: Update the state vector S according to the entity list ε, and calculate the completeness
[0070] S6: When score < 0.8, select the next node in the decision tree according to the unfilled slot, and return to S1; when Score ≥ 0.8, generate a diagnosis summary;
[0071] Among them, the context window of each round of interaction retains the last 5 rounds of question and answer pairs, and the part exceeding is compressed and stored by using the TextRank algorithm to extract key sentences.
[0072] Further, between steps S4 and S5, a consistency checking step is also included:
[0073] Convert the entity list ε to a 768-dimensional vector representation, v current is the entity vector of the current round, v history is the average of the entity vectors of the last 5 rounds of dialogue;
[0074] Calculate the cosine similarity between the current entity vector and the historical entity vector:
[0075]
[0076] When sim < 0.6, generate a confirmation question: "You mentioned [historical symptoms] before, now you say [current symptoms], please confirm whether..." ;
[0077] When a time logic conflict (such as "yesterday started" and "last for a week" appearing at the same time) is detected, prefer to ask for clarification.
[0078] Further, the inquiry summary generation step comprises:
[0079] Convert the filled state vector S into a structured text template: "Patient complaint: [s1 corresponding entity], occurrence time: [s2 corresponding entity], duration: [s3 corresponding entity]..."
[0080] Rewrite the structured text into natural language summary using the seq2seq model (3 layers of encoder-decoder, hidden dimension 256);
[0081] Calculate the ROUGE-L score of the summary text and the original dialogue, and regenerate when the score is <0.7.
[0082] Compared with the prior art, the beneficial effects of the present application are:
[0083] 1. The present application provides a closed-loop medical voice interaction system and method, which realizes the natural inquiry process similar to real doctors through the deep integration of five units of voice generation, collection, understanding, management and data flow control, and can realize real-time perception of dialogue state and dynamic adjustment of inquiry strategy. The modules are closely coordinated through the WebSocket bidirectional channel, and the end-to-end delay is controlled within 500 milliseconds, ensuring the fluency and coherence of the dialogue.
[0084] 2. The present application provides a closed-loop medical voice interaction system and method, which adopts BERT encoder combined with CRF decoder, can accurately identify 13 types of medical entities, including symptoms, parts, time, degree and other key information. Through real-time updating and completeness calculation of state vector, the system can accurately grasp the information collection progress, avoid repeated inquiry or omission of important information. Especially the introduction of consistency checking mechanism can actively find and correct contradictions in patient description, ensuring the reliability of information collection.
[0085] 3. The present application provides a closed-loop medical voice interaction system and method, which can realize real-time perception of patient's emotional state through analysis of voice prosody characteristics, and adjust the inquiry rhythm and tone accordingly. When detecting anxiety or eagerness of the patient, the system will speed up the inquiry process or give appropriate comfort, reflecting the humanistic care of medical service. At the same time, the voice generated by the voice synthesis technology based on Transformer is natural and smooth, avoiding mechanical voice output. BRIEF DESCRIPTION OF DRAWINGS
[0086] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0087] Figure 1 is a schematic diagram of the system architecture of the present application DETAILED DESCRIPTION
[0088] The technical solutions of the present application will be described more clearly and completely by combining the drawings and the description of the preferred embodiments of the present application.
[0089] As Figure 1 shown, it is composed of six core functional units, coordinated by a central data flow controller. The system operates in a closed-loop manner, starting from the speech generation unit, which contains a text encoder, an acoustic model, and a vocoder, converting the consultation text into a speech signal sent to the user. The user's voice answer is received and processed by the speech acquisition unit, which converts the speech into text through a VAD detector, a speech recognizer, and an audio preprocessing component. The converted text is analyzed by the semantic understanding unit, in which the BERT encoder, CRF decoder, and medical entity tagging system jointly identify the key medical entities. The dialogue management unit then updates the dialogue state vector according to the identified medical entities and decides the next consultation content through the consultation decision tree and consistency check, generating new consultation text sent back to the speech generation unit, forming a complete closed loop. The system also contains an emotion recognition module that analyzes the emotional state from the user's speech through prosodic features and an SVM classifier; the state display function monitors the slot filling degree of the consultation progress and generates a consultation summary at the appropriate time. The central data flow controller establishes a bidirectional channel based on the WebSocket protocol to coordinate the data transmission between the units, ensuring that the system reacts quickly with an end-to-end delay of less than 500 milliseconds. This closed-loop design enables the system to continuously and effectively collect patient information, guide the conversation through a structured consultation process, and generate a complete consultation summary after collecting enough information.
[0090] As a specific embodiment, the closed-loop medical voice interaction system provided by the present application realizes natural and smooth voice consultation interaction between doctors and patients. The system builds a complete closed-loop interaction system through the close cooperation of the five core modules of the speech generation unit, the speech acquisition unit, the semantic understanding unit, the dialogue management unit, and the data flow controller.
[0091] In practical applications, when the patient enters the consultation interface, the system first initializes the dialogue state. The speech generation unit generates the first question based on the pre-set consultation decision tree, such as "Where do you feel unwell?" This text question is converted into a 512-dimensional semantic vector by the text encoder, and then a natural and fluent voice is generated by the acoustic model. The acoustic model uses a Transformer architecture, including 6 layers of encoder and 6 layers of decoder, which can accurately map text to mel-spectrum features. Subsequently, the WaveGAN vocoder converts these spectral features into high-quality 16kHz speech signals.
[0092] The generated speech is divided into 20ms speech frames, each containing 320 sampling points, and is pushed to the patient end in real time through the WebSocket protocol. This streaming method ensures extremely low interaction delay, with the entire end-to-end delay controlled within 500ms, so that the patient does not feel significant waiting time.
[0093] When the patient starts answering after hearing the question, the speech collection unit immediately starts working. The VAD detector accurately identifies when the patient starts speaking and when they stop by monitoring the speech energy and zero-crossing rate. The system sets the energy threshold to -40dB and the zero-crossing rate threshold to 0.25, which are optimized through a large number of medical scene tests and can accurately detect human voice in the background noise of a hospital environment. The audio preprocessing module uses spectral subtraction to effectively remove environmental noise and improve recognition accuracy.
[0094] After the collected speech is converted into text by the CTC decoder, the semantic understanding unit begins to analyze the patient's answer content. The BERT encoder converts the text into a 768-dimensional context vector, fully understanding the semantic information. The CRF decoder performs sequence labeling based on pre-defined 13 medical entity labels, accurately identifying key information such as symptom name, body part, time description, severity, etc. For example, when the patient says "I have a headache for three days, it's very severe," the system can identify "headache" as a symptom entity, "three days" as a time entity, and "very severe" as a degree description.
[0095] The dialogue management unit is the intelligent core of the entire system, which maintains a state vector containing multiple slots and tracks the consultation progress in real time. The chief complaint-related slots include symptom name, occurrence time, duration, severity, and frequency of onset; the accompanying symptoms slot records other related symptoms; the past history slot covers disease history, medication history, and allergy history, etc. important information. Whenever a new entity is identified, the filling degree of the corresponding slot will be updated according to the recognition confidence.
[0096] The system introduces a consistency checking mechanism. When detecting contradictions between the patient's current description and historical information, the system generates clarifying questions. For example, if the patient first says "headache started yesterday" and then mentions "has been lasting for a week," the system will actively ask: "You just said it started yesterday, but you also mentioned it has been lasting for a week. Could you please confirm the specific time again?" This intelligent error correction capability greatly improves the accuracy of information collection.
[0097] The system extracts prosodic features such as fundamental frequency, energy, speech rate, and pause rate from the speech signal, and judges the patient's emotional state through an SVM classifier, divided into five categories: calm, anxious, puzzled, urgent, and negative. When detecting that the patient's emotion is anxious or urgent, the system will appropriately adjust the examination strategy, such as speeding up the examination pace or giving soothing responses.
[0098] The examination decision adopts a decision tree structure with a depth of 5 layers and contains 127 nodes. Each node contains question text, child node pointer, and transition condition. The system dynamically selects the examination path according to the completeness of the current state vector: when the chief complaint information collection is insufficient, it prioritizes the chief complaint collection branch; when the chief complaint is basically complete but needs to be further understood, it enters the symptom details branch; when the information collection is basically completed, it enters the summary confirmation branch.
[0099] The ring buffer of the data flow controller can accommodate 100 frames of speech data, and accurately manages the data flow through read and write pointers. The flow control mechanism is automatically triggered when the buffer is close to full to prevent data overflow. The timestamp synchronization mechanism ensures the continuity of speech playback, even in the case of network jitter, to maintain smoothness. For occasional frame loss, the system uses a linear interpolation algorithm to reconstruct the missing speech frames, ensuring user auditory experience.
[0100] During the entire examination process, the system retains the complete context of the last 5 rounds of dialogue, and the excess part is compressed and stored by extracting key information through the TextRank algorithm. This context management strategy not only ensures the coherence of the dialogue, but also avoids the problem of excessive memory occupation.
[0101] When the system judges that the information collection completeness reaches 80% or more, it automatically generates an examination summary. First, the state vector is converted into a structured text template, and then it is rewritten into a natural and smooth summary text through a seq2seq model. The system also calculates the ROUGE-L score of the summary text and the original dialogue to ensure the accuracy and completeness of the summary.
[0102] The application realizes a truly intelligent medical diagnosis system by organically combining voice interaction, natural language processing, emotion calculation and other technologies. The whole interaction process is natural and smooth, just like a conversation with a real doctor, greatly improving the user experience and diagnosis efficiency of remote medical treatment. The modular design of the system also facilitates subsequent function expansion and performance optimization, and has a broad application prospect.
[0103] The above specific embodiments only describe the preferred embodiments of the application, and do not limit the protection scope of the application. Without departing from the design concept and spirit of the application, various modifications, substitutions and improvements of the technical solutions of the application made by those skilled in the art according to the description and drawings of the application shall belong to the protection scope of the application. The protection scope of the application is determined by the claims.
Claims
1. A closed-loop medical voice interaction system, characterized in that, It includes a speech generation unit, a speech acquisition unit, a semantic understanding unit, a dialogue management unit, and a data flow controller; The speech generation unit includes a text encoder and an acoustic model. The text encoder converts the consultation text into a 512-dimensional semantic vector. The acoustic model generates a speech signal with a sampling rate of 16kHz based on the semantic vector and divides it into speech frames with 320 sampling points in 20ms duration. The voice acquisition unit includes a VAD detector and a voice recognizer. The VAD detector uses an energy threshold of -40dB and a zero-crossing rate threshold of 0.25 to identify valid voice segments. The voice recognizer uses CTC decoding to convert the voice segments into text. The semantic understanding unit includes a BERT encoder and a CRF decoder. The BERT encoder generates a 768-dimensional context vector, and the CRF decoder performs sequence annotation based on 13 predefined medical entity labels. The dialogue management unit maintains a state vector S = [s1, s2, ..., s3] containing symptom slots, time slots, and location slots. 15 ], where s i ∈[0,1] represents the fill degree of the i-th slot, when It is determined that the information collected is complete at that time; The data flow controller establishes a bidirectional channel through the WebSocket protocol and coordinates data transmission between units using a queue caching mechanism to ensure that the end-to-end latency is less than 500ms.
2. The closed-loop medical voice interaction system according to claim 1, characterized in that, The speech generation unit specifically includes: The consultation content generation module constructs a consultation decision tree with a depth of 5 layers and 127 nodes. Each node contains a triple of {question text, child node pointer, and transition condition}. The next node is selected based on the state vector S. When s1 < 0.5, select the main complaint collection branch. When 0.5 ≤ s1 < 0.8 and s2 < 0.5, select the symptom details branch. when When selecting a summary confirmation branch; The acoustic feature generation module adopts a Transformer architecture, containing a 6-layer encoder and a 6-layer decoder, to generate text sequences T = [t1, t2, ..., t...]. n The mapping is an 80-dimensional spectrum F = [f1, f2, ..., f m ], where m = n × k, k is the duration prediction factor; where t i f represents the i-th character in the text sequence. i This represents the i-th Mel spectrum frame; The vocoder module uses a WaveGAN generator, which takes an 80-dimensional mem spectrum as input and generates a 16kHz audio waveform through 5 upsampling convolutional layers. The convolutional kernel sizes are [16,16,4,4,4] and the stride is [8,8,2,2,2].
3. The closed-loop medical voice interaction system according to claim 1, characterized in that, The voice acquisition unit specifically includes: The audio preprocessing module uses spectral subtraction for noise reduction. Y(ω)=X(ω)-α·N(ω) Where X(ω) is the input spectrum, N(ω) is the noise estimate, and α = 2.5 is the over-subtraction factor; The human voice activity detection module calculates the energy E of each frame. j and the zero-crossing rate Z j : Where L = 320 is the frame length; x j [n] The nth sampling point of the j-th frame sgn(x) is a sign function, defined as: When E j >-40dB is marked as speech, when Z j A value >0.25 indicates a voiced sound; The endpoint detection module uses a dual-threshold method, setting a high threshold. and low threshold in: The average energy of M frames; Energy standard deviation Start and end point detection is achieved through a state machine.
4. The closed-loop medical voice interaction system according to claim 1, characterized in that, The medical entity recognition of the semantic understanding unit adopts the following labeling system: B-SYMPTOM / I-SYMPTOM: Symptom entity; B-BODY / I-BODY: Body parts; B-TIME / I-TIME: Time representation; B-DEGREE / I-DEGREE: Degree description; B-DRUG / I-DRUG: Drug name; O: Non-physical label; The Viterbi algorithm is used for optimal label sequence decoding, and the transition probability matrix is obtained. The emission probability is calculated from the last layer of BERT.
5. A closed-loop medical voice interaction system according to claim 1, characterized in that, The status update mechanism of the dialogue management unit is as follows: Define the state vector S = [s1, s2, ..., s 15 ],,in: s1-s5: Slots related to the main complaint; s6-s 10 Accompanying symptoms slot; s 11 -s 13 Historical slots; s 14 -s 15 Other information slots; Update rule: When the entity corresponding to slot i is detected, s i =min(s) i +c i ,1.0); Where c i Confidence level for entity identification.
6. The closed-loop medical voice interaction system according to claim 1, characterized in that, It also includes an emotion recognition module: Extract prosodic feature vector P = [p1, p2, ..., p6] from speech signal; p1: Mean fundamental frequency; p2: Standard deviation of fundamental frequency; p3: Average energy; p4: Energy standard deviation; p5: Speech rate; p6: Pause rate; An SVM classifier (RBF kernel, γ = 0.1, C = 1.0) was used to map the feature vector P to five emotion labels: calm, anxiety, doubt, urgency, and negativity.
7. A closed-loop medical voice interaction system according to claim 1, characterized in that, The data flow controller includes: A circular buffer with a capacity of 100 frames is used, managed by read and write pointers, and flow control is triggered when the value exceeds 80. The timestamp synchronization mechanism ensures that each audio frame carries a timestamp t. i =t0 + i × 20ms The playback speed is adjusted by the playback terminal based on the timestamp; where t0 is the Unix timestamp at which the system started playback; The frame loss recovery strategy uses linear interpolation to reconstruct lost frames when a discontinuous frame sequence is detected. f lost =0.5×(f prev +f next )。 8. A closed-loop medical voice interaction method, applicable to a closed-loop medical voice interaction system according to any one of claims 1-7, characterized in that, Includes the following steps: S1: Initialize the state vector The first question text is generated based on the root node of the consultation decision tree; S2: Convert the text into speech using an acoustic model, segment it into 20ms segments, and push it via WebSocket. The client receives and plays the speech. S3: Collect user voice. Recording starts when the VAD detects the start point and stops when the end point is detected. Perform ASR recognition on valid voice segments. S4: Perform medical entity recognition on the identified text and extract the entity list ε = {(entity text, label, location, confidence)}; S5: Update the state vector S based on the entity list ε, and calculate the completeness. S6: When score < 0.8, select the next node in the decision tree based on the unfilled slots and return to S1; when score ≥ 0.8, generate a consultation summary. In this process, the context window for each round of interaction retains the question-answer pairs from the most recent 5 rounds, and the excess parts are extracted and compressed using the TextRank algorithm.
9. A closed-loop medical voice interaction method according to claim 8, characterized in that, A consistency check step is also included between steps S4 and S5: Convert the entity list ε into a 768-dimensional vector representation, v current v is the entity vector for the current round. history The average entity vectors from the past 5 rounds of dialogue; Calculate the cosine similarity between the current entity vector and historical entity vectors: When sim < 0.6, a confirmation question is generated: "You previously mentioned [historical symptoms], now you are talking about [current symptoms], please confirm whether..."; When a time logic conflict is detected, clarification should be requested first.
10. A closed-loop medical voice interaction method according to claim 8, characterized in that, The steps for generating the consultation summary include: The state vector S is converted into a structured text template: "Patient's chief complaint: [entity corresponding to s1], time of occurrence: [entity corresponding to s2], duration: [entity corresponding to s3]..." The seq2seq model (3 layers each for encoder and decoder, 256 hidden dimensions) is used to rewrite structured text into natural language summaries. Calculate the ROUGE-L score between the summary text and the original dialogue, and regenerate if the score is <0.7.
Citation Information
Cited By
Streaming audio synthesis method and device, storage medium and electronic device
CN121600905A