Medical pre-inquiry real-time interaction method and system, terminal and medium
By combining professional medical knowledge base and large language model, using double buffers to manage audio streams and adjusting voice output according to patient characteristics, the problem of insufficient professionalism and lag in the existing medical dialogue system is solved, and an efficient and accurate medical pre-diagnosis interactive experience is achieved.
Patent Information
- Application Number
- CN202510205519.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-04
AI Technical Summary
The existing medical dialogue system lacks professionalism and accuracy, has poor flexibility in knowledge updates, and is prone to lag during voice interaction and the inability to adjust the voice effect according to patient characteristics, affecting the interactive experience.
Combining professional medical knowledge base and large language model, using double buffers to manage audio streams, adjust voice output according to patient characteristics, match the medical knowledge context through a multi-stage knowledge retrieval enhancement mechanism, build targeted propts and asynchronously process speech synthesis.
It improves the professionalism and accuracy of medical pre-diagnosis, can easily respond to knowledge updates, ensure the continuity of audio playback and interactive experience effect, and dynamically adjust the voice output to adapt to patient characteristics.
Smart Images

Figure CN120255840A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a real-time interactive method, system, terminal and medium for medical pre-consultation. Background Art
[0002] Medical pre-consultation is an important part of doctor training. Traditional training methods mainly rely on standardized patients or simulated patients, which have problems such as high cost, low efficiency, and single scenarios. In recent years, with the development of large language model (LLM) technology, AI-based medical dialogue systems have gradually become a new solution. However, the existing medical dialogue systems still have the following problems: (1) Lack of professionalism and accuracy: Existing systems directly use large language models for dialogue. Due to the lack of support from professional medical knowledge, it is easy to generate inaccurate, clinically unrealistic or even completely wrong answers. To address this problem, related technologies use medical knowledge to fine-tune neural network large models, and then use the fine-tuned neural network large models for pre-consultation. However, this method has poor flexibility in knowledge update. If medical knowledge needs to be updated, the large model needs to be re-fine-tuned and trained, which is time-consuming, consumes a large amount of resources, has a high training cost, and lacks targeted interaction. It only replies based on overall medical knowledge, lacks pertinence and authenticity, and reduces accuracy.
[0003] (2) Fragmented interaction experience: Currently, when a medical dialogue system performs voice interaction, it generally stores voice data in a single buffer. Once there is a delay in voice synthesis or the processing speed slows down, the data in the buffer may be quickly exhausted, thus affecting the continuity and stability of voice output, prone to stuttering, and unable to dynamically adjust the language expression and voice characteristics of the reply according to personal characteristics such as the age and gender of the patient, affecting the coherence and realism of the dialogue. Summary of the Invention
[0004] To solve the above problems, the present invention provides a real-time interactive method, system, terminal and medium for medical pre-consultation. By combining a professional medical knowledge base with a large language model, it improves the professionalism and accuracy of medical pre-consultation, and can conveniently respond to knowledge updates. At the same time, a dual-buffer structure is used to manage the audio stream, ensuring the continuity of audio playback, effectively avoiding playback stuttering caused by waiting for synthesis, and can adjust the voice effect of voice output according to the identity characteristics of the simulated patient, improving the interactive experience effect.
[0005] In a first aspect, the technical solution of the present invention provides a real-time interactive method for medical pre-consultation, including the following steps: Recognize the input user consultation voice, and convert the recognition result of the user consultation voice into text form, denoted as consultation text; Screen for medical knowledge contexts similar to the consultation text from a pre - constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, denoted as relevant text; Based on the pre - constructed prompt template, construct the current consultation prompt through the consultation text and the relevant text; Input the current consultation prompt into the large - language model to output the consultation response content; Perform speech synthesis on the consultation response content to generate consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than the threshold, transfer the data in the secondary buffer to the main buffer; Asynchronously process speech synthesis and voice output; Among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0006] In an optional implementation, the medical knowledge in the professional medical knowledge base is chunked by patient dimension, and each knowledge chunk corresponds to the medical information of a patient; The screening of the medical knowledge context similar to the consultation text from the pre - constructed professional medical knowledge base specifically includes: Perform vector quantization encoding on each knowledge chunk in the professional medical knowledge base through the BGE - Large - zh model; Perform vector quantization encoding on the consultation text through the BGE - Large - zh model; Based on the vector quantization encoding of the consultation text and the vector quantization encoding of each knowledge chunk, screen out several knowledge chunks whose cosine similarity with the consultation text is greater than the similarity threshold to form a first candidate knowledge chunk set; Calculate the semantic matching degree between the consultation text and each knowledge chunk in the first candidate knowledge chunk set through a deep - learning model based on a cross - encoder; Screen out at least one knowledge chunk from the candidate knowledge chunk set whose semantic matching degree is greater than the matching degree threshold to form a second candidate knowledge chunk set; Construct a medical knowledge context through the second candidate knowledge chunk set.
[0007] In an optional implementation, constructing a medical knowledge context through the second candidate knowledge chunk set specifically includes: Detect whether the number of knowledge chunks in the second candidate knowledge chunk set exceeds the upper threshold ; If not, directly construct a medical knowledge context with all the knowledge chunks in the second candidate knowledge chunk set; If so, sort all the knowledge chunks in the second candidate knowledge chunk set according to the semantic matching degree; Select knowledge chunks with the highest semantic matching degree from the second set of candidate knowledge chunks to form a third set of candidate knowledge chunks; Construct a medical knowledge context for all the knowledge chunks in the third set of candidate knowledge chunks.
[0008] In an alternative embodiment, the input user speech is recognized, and the user speech recognition result is converted into text form, specifically including: Preprocess the user speech and extract the acoustic feature sequence using the short-time Fourier transform , expressed as,
[0009] where is the user speech input at time is the window function, is the time offset, is the angular frequency; Generate a number of candidate texts based on the acoustic features; Calculate the decoding probability between each candidate text and the acoustic feature sequence ; Select the candidate text with the highest decoding probability and record it as the consultation text.
[0010] In an alternative embodiment, the consultation reply content is synthesized into voice to generate consultation reply voice data, the consultation reply voice data is stored in a buffer, the reply voice data is extracted from the main buffer for voice output, and when the data in the main buffer is less than the threshold, the data in the secondary buffer is transferred to the main buffer, specifically including: Segment the consultation reply content according to the segmentation identifier, and the length of each segment is less than a preset length threshold; Generate consultation reply voice data for each segment in sequence; Initially, store the consultation reply voice data in the main buffer; When the main buffer is full, store the subsequent consultation reply voice data in the secondary buffer in sequence; When the consultation reply voice data in the secondary buffer exceeds the first threshold, extract the reply voice data from the main buffer for voice output in sequence; When the consultation reply voice data in the main buffer is less than the second threshold, transfer the consultation reply voice data in the secondary buffer to the main buffer in sequence, and asynchronously process the generation operation of the consultation reply voice data.
[0011] In an alternative embodiment, the method further includes the following steps: Perform voice activity detection in real time and calculate the short-time energy of the detected voice activity within the time window and the zero-crossing rate within the time window ,
[0012]
[0013] wherein represents the amplitude of the voice signal at the moment, is the window function; When the following conditions are met, the voice output interruption mechanism is triggered,
[0014] wherein and are the thresholds of energy and zero-crossing rate respectively; When the voice output interruption mechanism is triggered, the consultation reply voice data in the main buffer and the secondary buffer is cleared, and the voice synthesis operation is stopped.
[0015] In an optional embodiment, the voice effect of the voice output is adjusted according to the identity characteristics of the simulated patient in the consultation text, which specifically includes the following steps: Extract the identity characteristics of the simulated patient from the consultation text, including age and gender; Based on the pre-set rules, configure the voice effect parameters according to the identity characteristics, including pitch parameters, sound speed parameters, and volume parameters; When performing voice output, perform voice output according to the configured voice effect parameters.
[0016] In a second aspect, the technical solution of the present invention provides a real-time interactive system for medical pre-consultation, including a consultation text acquisition module, which is used to recognize the input user consultation voice, convert the recognition result of the user consultation voice into text form, and record it as the consultation text; a relevant text acquisition module, which is used to screen the medical knowledge context similar to the consultation text from the pre-constructed professional medical knowledge base, and the medical knowledge context is composed of the medical information of several patients, and is recorded as the relevant text; a consultation prompt generation module, which is used to construct the current consultation prompt based on the pre-constructed prompt template through the consultation text and the relevant text; a consultation reply content acquisition module, which is used to input the current consultation prompt into the large language model and output the consultation reply content; A voice output module, which is used to synthesize the consultation reply content into voice to generate consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than a threshold, transfer the data in the secondary buffer to the main buffer; asynchronously process voice synthesis and voice output; wherein, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0017] In a third aspect, the technical solution of the present invention provides a terminal, including: A memory, which is used to store the real-time interactive program for medical pre-consultation; A processor, which is used to implement the steps of the real-time interactive method for medical pre-consultation as described in any one of the above when executing the real-time interactive program for medical pre-consultation.
[0018] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a real-time interactive program for medical pre-consultation is stored. When the real-time interactive program for medical pre-consultation is executed by a processor, the steps of the real-time interactive method for medical pre-consultation as described in any one of the above are implemented.
[0019] A real-time interactive method, system, terminal and medium for medical pre-consultation provided by the present invention have the following beneficial effects compared with the prior art: First, match the medical knowledge context from a professional medical knowledge base, then construct a prompt based on the consultation text and the medical knowledge context, and use the newly constructed targeted prompt to output the reply content. By combining the professional medical knowledge base with a large language model, the professionalism and accuracy of medical pre-consultation are improved, and knowledge updates can be conveniently responded to; at the same time, a dual-buffer structure of a main buffer and a secondary buffer is used to manage the audio stream, extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than a threshold, transfer the data in the secondary buffer to the main buffer, and cooperate with asynchronous processing of voice synthesis and voice output to ensure the continuity of audio playback, effectively avoiding playback stuttering caused by waiting for synthesis; and the voice effect of the voice output can be adjusted according to the identity characteristics of the simulated patient, improving the interactive experience effect. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solution of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of a real-time interactive method for medical pre-consultation provided by an embodiment of the present invention.
[0021] Figure 2 Construct a flow diagram for the medical knowledge context.
[0022] Figure 3 This is a schematic block diagram of the structure of a medical pre-consultation real-time interaction system provided by an embodiment of the present invention.
[0023] Figure 4 This is a schematic diagram of the structure of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0024] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the specific embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0026] The following explains the key terms that appear in the present invention.
[0027] LLM: Large Language Model, a large language model, is an artificial intelligence model based on deep learning. Through training with a large amount of text data, it has powerful language understanding and generation capabilities.
[0028] prompt: Prompt word. When interacting with a language model, it is the instruction, question, or text input by the user to guide the model to generate content.
[0029] RAG: Retrieval-Augmented Generation, a retrieval-augmented generation technology that combines retrieval and generation technologies to retrieve information from an external knowledge base to assist the language model in generating answers.
[0030] TTS: Text-to-Speech, a text-to-speech streaming technology that converts text into speech.
[0031] VAD: Voice Activity Detection, a technology used to detect whether there is human speech activity in a speech signal.
[0032] BGE-Large-zh Model: The BGE large-scale Chinese model is a model used for natural language processing tasks.
[0033] SenseVoiceSmall Model of the FunASR Framework: FunASR is an open-source speech recognition framework by Alibaba, and SenseVoiceSmall is an encoder-only speech base model in it, used for fast speech understanding.
[0034] Microsoft Edge-TTS Engine: The speech synthesis engine of the Microsoft Edge browser, used to convert text into speech.
[0035] Figure 1 This is a schematic flowchart of a real-time interactive method for medical pre-consultation provided by an embodiment of the present invention. Among them, Figure 1 The execution subject can be a real-time interactive system for medical pre-consultation. The real-time interactive method for medical pre-consultation provided by an embodiment of the present invention is executed by a computer device. Correspondingly, the real-time interactive system for medical pre-consultation runs in the computer device. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0036] As Figure 1 shown, the method includes the following steps.
[0037] S1. Recognize the input user consultation voice, convert the recognition result of the user consultation voice into text form, and record it as the consultation text.
[0038] S2. Screen medical knowledge contexts similar to the consultation text from a pre-constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, and is recorded as the relevant text.
[0039] S3. Based on a pre-constructed prompt template, construct the current consultation prompt through the consultation text and the relevant text.
[0040] S4. Input the current consultation prompt into a large language model and output the consultation reply content.
[0041] S5. Synthesize the consultation reply content into consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than the threshold, transfer the data in the secondary buffer to the main buffer; asynchronously process voice synthesis and voice output. Among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0042] The medical pre-consultation real-time interaction method provided in this embodiment combines a professional medical knowledge base with a large language model to improve the professionalism and accuracy of medical pre-consultation, can conveniently respond to knowledge updates, and at the same time uses a double-buffer structure to manage the audio stream to ensure the continuity of audio playback, effectively avoiding playback jams caused by waiting for synthesis, and can adjust the voice effect of the voice output according to the identity characteristics of the simulated patient to improve the interactive experience effect.
[0043] To further understand the present invention, a specific embodiment is provided below to further elaborate on the present invention in detail. The specific embodiment includes the following steps.
[0044] SS1, construct a professional medical knowledge base.
[0045] First, a comprehensive professional medical knowledge base is constructed through multi-source data collection, integrating clinical basic information such as department classification, demographic characteristics, chief complaint and current medical history, past medical history, physical examination, preoperative diagnosis, and surgical plan. At the same time, it includes test and examination data such as blood routine (HGB, RBC, WBC, etc.), biochemical indicators (ALT, AST, TP, etc.), coagulation function (PT, APTT, etc.), electrolytes, infectious indicators, as well as medical imaging data such as electrocardiogram and X-ray and cardiopulmonary function evaluation results. The original data is standardized using the international standard medical terminology system to construct a multi-dimensional disease-symptom-treatment plan relationship map and a drug-indication-contraindication association network, and the key clinical decision points in the standard consultation process are extracted to form a systematic medical knowledge system.
[0046] The medical knowledge in the professional medical knowledge base is divided into knowledge blocks based on the patient dimension, and each knowledge block corresponds to a patient's medical information. The medical information includes the whole-process data such as basic information, chief complaint symptoms, test and examination results, and treatment plans. By structurally integrating various clinical indicators, medical records, and follow-up data of patients, the integrity and relevance of medical information are ensured. Each knowledge block contains at least multi-dimensional meta-information such as the patient's demographic characteristics, clinical manifestations, and test results, realizing knowledge retrieval based on real cases and similar case matching.
[0047] SS2, recognize the input user consultation voice, and convert the recognition result of the user consultation voice into text form, denoted as the consultation text.
[0048] In the voice recognition link, this specific embodiment samples the SenseVoiceSmall model based on the FunASR framework. The input audio data is first preprocessed, then acoustic features are extracted, and finally the consultation text is generated, including the following steps.
[0049] SS2.1, preprocess the user voice and extract the acoustic feature sequence using the short-time Fourier transform , expressed as,
[0050] Wherein, is the user voice input at a moment, is a window function, is a time offset, is an angular frequency.
[0051] SS2.2, generate a number of candidate texts according to acoustic features.
[0052] SS2.3, calculate the decoding probability between each candidate text and the acoustic feature sequence therebetween.
[0053] The SenseVoiceSmall model is based on the Transformer architecture, and its decoding probability can be expressed as:
[0054] Wherein, is the recognition sequence result, is the recognition result sequence in the th element at the th position, represents the subsequence composed of all elements before the th position. The SenseVoiceSmall model can be optimized by introducing a medical domain-specific hot word list to improve the recognition accuracy of professional medical terms.
[0055] SS2.4, screen out the candidate text with the highest decoding probability and record it as the consultation text.
[0056] SS3, screen out the medical knowledge context similar to the consultation text from the pre-constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients and is recorded as the relevant text.
[0057] Figure 2 is a schematic diagram of the medical knowledge context construction process. This specific embodiment adopts a multi-stage knowledge retrieval enhancement mechanism (RAG) to ensure the professionalism and accuracy of medical consultation responses. The first stage is the vectorized retrieval stage, and the second stage introduces a cross-encoder for re-ranking, including the following steps.
[0058] SS3.1, perform vectorized encoding on each knowledge block in the professional medical knowledge base through the BGE-Large-zh model.
[0059] In the vectorization stage, the BGE-Large-zh model trained with a large-scale Chinese corpus is used to perform semantic encoding on the knowledge blocks. For the knowledge block , and its vector representation is:
[0060] SS3.2, vectorize and encode the consultation text through the BGE-Large-zh model.
[0061] First, structurally analyze the doctor's description (such as age, gender, clinical symptoms, relevant background, etc.) to extract key feature information. Use this feature information as the query vector and perform similarity matching in the medical knowledge base through the RAG retrieval mechanism.
[0062] When receiving a user query , after processing the feature information, perform the same vectorization processing on the query:
[0063] SS3.3, based on the vectorized encoding of the consultation text and the vectorized encoding of each knowledge block, screen out several knowledge blocks with a cosine similarity greater than the similarity threshold to the consultation text, and form the first candidate knowledge block set.
[0064] In the first stage, perform a preliminary screening through vector similarity calculation, and use cosine similarity to measure the semantic relevance between the query vector and the knowledge block vector:
[0065] SS3.4, calculate the semantic matching degree between the consultation text and each knowledge block in the first candidate knowledge block set through a deep learning model based on a cross-encoder.
[0066] In the second stage, introduce a cross-encoder for precise re-ranking, and use a deep learning model to perform a finer-grained modeling of the semantic matching relationship between the query and the knowledge block. Its scoring function is:
[0067] SS3.4.1, perform data preparation.
[0068] Obtain the consultation text input by the user, clean it to remove interference information such as special characters and punctuation marks. Then perform word segmentation to split the text into individual words or sub-word units. Exemplarily, for the consultation text "What should I do about my recent headache", after word segmentation, it may be ["I", "recently", "headache", "what should I do"].
[0069] For each knowledge block in the first candidate knowledge block set, perform the same cleaning and word segmentation operations. Exemplarily, a knowledge block is "Headache may be caused by lack of sleep", and after word segmentation, it is ["headache", "may", "be", "caused", "by", "lack", "of", "sleep"].
[0070] SS3.4.2, Construct the input format.
[0071] Combine the consultation text and each knowledge block to form the input of the model. Specifically, concatenate the two together and add a special separator in the middle. Exemplarily, concatenate the above consultation text and knowledge block as "What should I do if I have a headache recently [SEP] Headache may be caused by lack of sleep".
[0072] SS3.4.3, Load the pre-trained model.
[0073] The deep learning model based on the cross-encoder is BERT of the Transformer architecture. Load the pre-trained model weights and configurations on a large-scale corpus from the model repository.
[0074] SS3.4.4, Model inference.
[0075] Convert the constructed input sequence into a format that the model can process. Specifically, map words or sub-words to corresponding IDs, and at the same time generate an attention mask to identify the valid positions in the input sequence.
[0076] Input the processed input data into the loaded cross-encoder model. The model encodes the input through the internal multi-layer Transformer structure to capture the semantic interaction information between the consultation text and the knowledge block.
[0077] The model outputs a score representing the degree of semantic matching.
[0078] SS3.4.5, Calculate the semantic matching degree.
[0079] Perform the above model inference steps for each combination of the knowledge block and the consultation text to obtain a series of scores. These scores represent the semantic matching degrees between the consultation text and each knowledge block in the first candidate knowledge block set.
[0080] SS3.5, Filter out at least one knowledge block with a semantic matching degree greater than the matching degree threshold from the candidate knowledge block set to form the second candidate knowledge block set.
[0081] SS3.6, Construct the medical knowledge context through the second candidate knowledge block set.
[0082] Based on the re-ranking score, select the most relevant knowledge block set to construct the medical context, and at the same time control the context length through the knowledge block quantity limit, expressed as:
[0083] Where Relevance threshold, used to filter low-relevance knowledge, is the upper limit of the number of returned knowledge chunks, used to control the context length.
[0084] SS3.6.1, Detect whether the number of knowledge chunks in the second set of candidate knowledge chunks exceeds the upper limit threshold .
[0085] SS3.6.2, If not, directly construct the medical knowledge context with all the knowledge chunks in the second set of candidate knowledge chunks.
[0086] SS3.6.3, If so, sort all the knowledge chunks in the second set of candidate knowledge chunks according to the semantic matching degree.
[0087] SS3.6.4, Screen out from the second set of candidate knowledge chunks knowledge chunks with the highest semantic matching degree to form the third set of candidate knowledge chunks.
[0088] SS3.6.5, Construct the medical knowledge context with all the knowledge chunks in the third set of candidate knowledge chunks.
[0089] The information of all knowledge chunks is merged according to the pre-configured logic and output in the set format to form the final medical knowledge context.
[0090] The screened medical knowledge context is combined with the user query , and input into the large language model fine-tuned with medical domain knowledge to generate professional and accurate medical consultation responses:
[0091] Through this multi-stage knowledge retrieval enhancement mechanism, combining the efficiency of vector retrieval and the accuracy of cross-encoding, relevant medical knowledge can be quickly located, and the professionalism and accuracy of the answers during the conversation can be ensured.
[0092] SS4, Based on the pre-constructed prompt template, construct the current consultation prompt through the consultation text and relevant texts; input the current consultation prompt into the large language model to output the consultation response content.
[0093] For the determined matching patient cases after the above steps, automatically use the targeted prompt template to form the current consultation prompt. Exemplary: { Basic role setting: - You are now playing the role of this patient, aged {age}, gender {gender} - The chief complaint is {chief_complaint}, lasting for {duration} - You should answer based on the information in the medical record. Principles of answering: - Only answer based on the information recorded in the medical record. - For information not recorded in the medical record, indicate uncertainty or not remembering. - Use an expression that conforms to the patient's background. } Through this method of constructing prompts based on real cases, the authenticity and professionalism of the conversation can be ensured. During the actual consultation process, answers are strictly based on the matched medical record information, and reasonable ambiguity is shown for information not recorded in the medical record. This design is more in line with the performance characteristics of real patients. The patient simulation scheme based on real case matching not only simplifies the system implementation complexity but also provides a highly realistic consultation experience, enabling doctors to conduct pre-consultation training in an environment close to the real scenario.
[0094] SS5, perform speech synthesis on the consultation reply content to generate consultation reply voice data and output the consultation reply voice.
[0095] In the speech synthesis link, EdgeTTS engine is used based on a double buffer to achieve streaming voice output for low-latency real-time speech synthesis, which specifically includes the following steps.
[0096] SS5.1, segment the consultation reply content according to the segmentation identifier, and the length of each segment is less than the preset length threshold.
[0097] SS5.2, generate consultation reply voice data for each segment in sequence.
[0098] SS5.3, initially, store the consultation reply voice data in the main buffer.
[0099] SS5.4, when the main buffer is full, store the subsequent consultation reply voice data in the secondary buffer in sequence.
[0100] SS5.5, when the consultation reply voice data in the secondary buffer exceeds the first threshold, extract the reply voice data from the main buffer in sequence for voice output.
[0101] SS5.6, when the consultation reply voice data in the main buffer is less than the second threshold, transfer the consultation reply voice data in the secondary buffer to the main buffer in sequence, and simultaneously asynchronously process the operation of generating consultation reply voice data.
[0102] Specifically, the reply text generated by the large language model is intelligently segmented, mainly at the end punctuation such as full stops and question marks, and at the same time, the maximum segmentation length threshold is set to avoid overly long single-segment text.
[0103] This specific embodiment manages the audio stream using a dual-buffer structure. The main buffer is responsible for storing the audio data being played, while the secondary buffer preloads the synthesis results of the subsequent text. When the data in the main buffer is played to a preset threshold (usually 80%), the data in the secondary buffer is automatically transferred to the main buffer, and at the same time, the synthesis task of the next text is processed asynchronously. This mechanism ensures the continuity of audio playback and effectively avoids playback stuttering caused by waiting for synthesis.
[0104] In addition, for each text segment in this specific embodiment, the speech parameters, including speech rate, pitch, volume, etc., are configured through the SSML (Speech Synthesis Markup Language) markup language to meet the requirements of different dialogue scenarios, so as to dynamically adjust the language expression and speech characteristics of the response according to personal characteristics such as the age and gender of the patient. Specifically, it includes the following steps.
[0105] Step 1: Extract the identity characteristics of the simulated patient from the consultation text, including age and gender.
[0106] Step 2: Based on the pre-set rules, configure the voice effect parameters according to the identity characteristics, including pitch parameters, speech rate parameters, and volume parameters.
[0107] Step 3: When performing voice output, perform voice output according to the configured voice effect parameters.
[0108] This specific embodiment can dynamically adjust the voice effect according to the patient characteristics by adjusting the pitch, rate, and volume parameters, which is expressed as:
[0109] It should be noted that this specific embodiment adopts a patient recognition mechanism of "one description, precise matching". When the doctor describes the basic information of the patient (such as "a 35-year-old female with abdominal pain for 3 days"), the identity characteristics can be automatically extracted, and then the voice effect parameters are configured based on the pre-set rules. It can be understood that the pre-set rules represent the voice effect parameters corresponding to different identity characteristics.
[0110] SS6, perform voice activity detection in real time.
[0111] This specific embodiment performs voice activity detection in real time to achieve the real-time interruption function, simulating the natural interaction method in the real consultation process. Specifically, the real-time interruption function is achieved through the combination judgment of the energy threshold and the zero-crossing rate.
[0112] Calculate the short-time energy of the detected voice activity within the time window through the following formula inside and zero-crossing rate ,
[0113]
[0114] wherein, represents the amplitude of the voice signal at moment, is the window function.
[0115] When the following conditions are met, the voice output interruption mechanism is triggered.
[0116] wherein, and are the thresholds of energy and zero-crossing rate respectively.
[0117] When the voice output interruption mechanism is triggered, the consultation reply voice data in the main buffer and the secondary buffer is cleared, and the voice synthesis operation is stopped, ensuring that the system can respond to the user's interruption behavior in a timely manner. Through this design, the performance index of the first-word output delay being less than 500 milliseconds is achieved, providing a smooth and natural voice interaction experience for the medical pre-consultation system.
[0118] In the above text, an embodiment of a medical pre-consultation real-time interaction method has been described in detail. Based on the medical pre-consultation real-time interaction method described in the above embodiment, an embodiment of the present invention also provides a medical pre-consultation real-time interaction system corresponding to this method.
[0119] Figure 3 FIG. is a schematic block diagram of the structure of a medical pre-consultation real-time interaction system provided by an embodiment of the present invention. In this embodiment, the medical pre-consultation real-time interaction system 300 can be divided into multiple functional modules according to the functions it performs, such as Figure 3 shown. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory.
[0120] The consultation text acquisition module 310 is used to recognize the input user consultation voice, and convert the recognition result of the user consultation voice into text form, denoted as the consultation text.
[0121] The relevant text acquisition module 320 is used to screen medical knowledge contexts similar to the consultation text from a pre-constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, denoted as the relevant text.
[0122] The consultation prompt generation module 330 is used to construct the current consultation prompt based on a pre - constructed prompt template through the consultation text and relevant text.
[0123] The consultation reply content acquisition module 340 is used to input the current consultation prompt into a large - language model and output the consultation reply content.
[0124] The voice output module 350 is used to synthesize the consultation reply content into consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than a threshold, transfer the data in the secondary buffer to the main buffer; asynchronously process voice synthesis and voice output; among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0125] The medical pre - consultation real - time interaction system in this embodiment is used to implement the foregoing medical pre - consultation real - time interaction method. Therefore, the specific implementation manners in this system can be seen in the embodiment part of the medical pre - consultation real - time interaction method in the previous text. Therefore, its specific implementation manners can refer to the descriptions of the corresponding individual part embodiments and will not be elaborated here.
[0126] In addition, since the medical pre - consultation real - time interaction system in this embodiment is used to implement the foregoing medical pre - consultation real - time interaction method, its functions correspond to those of the above - mentioned method and will not be repeated here.
[0127] Figure 4 The following is a schematic structural diagram of a terminal 400 provided by an embodiment of the present invention, including: a processor 410, a memory 420, and a communication unit 430. When the processor 410 implements the medical pre - consultation real - time interaction program stored in the memory 420, the following steps are implemented: Recognize the input user consultation voice, convert the recognition result of the user consultation voice into text form, denoted as the consultation text; Screen medical knowledge contexts similar to the consultation text from a pre - constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, denoted as relevant text; Based on a pre - constructed prompt template, construct the current consultation prompt through the consultation text and relevant text; Input the current consultation prompt into a large - language model and output the consultation reply content; Perform voice synthesis on the consultation reply content to generate consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than a threshold, transfer the data in the secondary buffer to the main buffer; Asynchronously process voice synthesis and voice output; Among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0128] The present invention also provides a computer storage medium, which can be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.
[0129] The computer storage medium stores a medical pre-consultation real-time interaction program. When the medical pre-consultation real-time interaction program is executed by a processor, the following steps are implemented: Recognize the input user consultation voice, convert the user consultation voice recognition result into text form, denoted as consultation text; Screen medical knowledge contexts similar to the consultation text from a pre-constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, denoted as relevant text; Based on a pre-constructed prompt template, construct the current consultation prompt through the consultation text and the relevant text; Input the current consultation prompt into a large language model to output the consultation reply content; Perform voice synthesis on the consultation reply content to generate consultation reply voice data, store the consultation reply voice data in a buffer. The buffer includes a main buffer and a secondary buffer. Extract the reply voice data from the main buffer for voice output, and when the data in the main buffer is less than a threshold, transfer the data in the secondary buffer to the main buffer; Asynchronously process voice synthesis and voice output; Among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
[0130] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A real-time interactive method for medical pre-consultation, characterized in that, It includes the following steps: Recognize the input user consultation voice, convert the recognition result of the user consultation voice into text form, and record it as the consultation text; Screen the medical knowledge context similar to the consultation text from the pre-constructed professional medical knowledge base. The medical knowledge context is composed of the medical information of several patients, and is recorded as the relevant text; Based on the pre-constructed prompt template, construct the current consultation prompt through the consultation text and the relevant text; Input the current consultation prompt into the large language model and output the consultation reply content; Perform speech synthesis on the consultation reply content to generate consultation reply voice data, store the consultation reply voice data in the buffer area. The buffer area includes a main buffer area and a secondary buffer area. Extract the reply voice data from the main buffer area for voice output, and when the data in the main buffer area is less than the threshold, transfer the data in the secondary buffer area to the main buffer area; Asynchronously process speech synthesis and voice output; Among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
2. The real-time interactive method for medical pre-consultation according to claim 1, wherein The medical knowledge in the professional medical knowledge base is chunked by patient dimension, and each knowledge chunk corresponds to the medical information of a patient; The screening of the medical knowledge context similar to the consultation text from the pre-constructed professional medical knowledge base specifically includes: Perform vector quantization encoding on each knowledge chunk in the professional medical knowledge base through the BGE-Large-zh model; Perform vector quantization encoding on the consultation text through the BGE-Large-zh model; Based on the vector quantization encoding of the consultation text and the vector quantization encoding of each knowledge chunk, screen out several knowledge chunks whose cosine similarity with the consultation text is greater than the similarity threshold, and form the first candidate knowledge chunk set; Calculate the semantic matching degree between the consultation text and each knowledge chunk in the first candidate knowledge chunk set through a deep learning model based on a cross encoder; Screen out at least one knowledge chunk from the candidate knowledge chunk set whose semantic matching degree is greater than the matching degree threshold, and form the second candidate knowledge chunk set; Construct the medical knowledge context through the second candidate knowledge chunk set.
3. The real-time interactive method for medical pre-consultation according to claim 2, wherein Constructing the medical knowledge context through the second candidate knowledge chunk set specifically includes: Check whether the number of knowledge chunks in the second set of candidate knowledge chunks exceeds the upper limit threshold ; If not, directly construct the medical knowledge context with all the knowledge chunks in the second candidate knowledge chunk set; If so, sort all the knowledge chunks in the second candidate knowledge chunk set according to the semantic matching degree; Select knowledge chunks with the highest semantic matching degree from the second set of candidate knowledge chunks to form the third set of candidate knowledge chunks; Construct the medical knowledge context with all the knowledge chunks in the third candidate knowledge chunk set.
4. The real-time interactive method for medical pre-consultation according to claim 3, wherein, Recognize the input user voice, and convert the recognition result of the user voice into text form, specifically including: Preprocess the user's speech and extract the acoustic feature sequence using the short-time Fourier transform , denoted as Among them, is the user speech input at a moment, is a window function, is a time offset, is an angular frequency; Generate several candidate texts according to the acoustic features; Calculate the decoding probabilities between each candidate text and the acoustic feature sequence ; Screen out the candidate text with the highest decoding probability and record it as the consultation text.
5. The real-time interactive method for medical pre-consultation according to claim 4, wherein Perform speech synthesis on the consultation reply content to generate consultation reply voice data, store the consultation reply voice data in the buffer area, extract the reply voice data from the main buffer area for voice output, and when the data in the main buffer area is less than the threshold, transfer the data in the secondary buffer area to the main buffer area, specifically including: Segment the consultation reply content according to the segment identifier, and the length of each segment is less than the preset length threshold; Generate consultation reply voice data for each segment in chronological order; Initially, store the consultation reply voice data in the main buffer; After the main buffer is full, store the subsequent consultation reply voice data in the secondary buffer in sequence; When the consultation reply voice data in the secondary buffer exceeds the first threshold, extract the reply voice data from the main buffer in sequence for voice output; When the consultation reply voice data in the main buffer is less than the second threshold, transfer the consultation reply voice data in the secondary buffer to the main buffer in sequence, and asynchronously process the operation of generating the consultation reply voice data.
6. The real-time interactive method for medical pre-consultation according to claim 5, characterized in that The method further includes the following steps: Perform voice activity detection in real time and calculate the short-time energy of the detected voice activity within the time window through the following formula and the zero-crossing rate within , Among them, represents the amplitude of the speech signal at the moment, is the window function; When the following conditions are met, trigger the voice output interruption mechanism, wherein, and are the thresholds of energy and zero-crossing rate respectively; When the voice output interruption mechanism is triggered, clear the consultation reply voice data in the main buffer and the secondary buffer, and stop the voice synthesis operation.
7. The real-time interactive method for medical pre-consultation according to claim 6, wherein Adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text, specifically including the following steps: Extract the identity characteristics of the simulated patient from the consultation text, including age and gender; Based on the pre-set rules, configure the voice effect parameters according to the identity characteristics, including pitch parameters, sound speed parameters, and volume parameters; When performing voice output, perform voice output according to the configured voice effect parameters.
8. A real-time interactive system for medical pre-consultation, characterized in that, Including, A consultation text acquisition module, used to recognize the input user consultation voice, and convert the recognition result of the user consultation voice into text form, denoted as the consultation text; A relevant text acquisition module, used to screen the medical knowledge context similar to the consultation text from the pre-constructed professional medical knowledge base, and the medical knowledge context is composed of the medical information of several patients, denoted as the relevant text; A consultation prompt generation module, used to construct the current consultation prompt based on the pre-constructed prompt template through the consultation text and the relevant text; A consultation reply content acquisition module, used to input the current consultation prompt into the large language model and output the consultation reply content; A voice output module, used to synthesize the consultation reply content into consultation reply voice data, store the consultation reply voice data in the buffer, the buffer includes a main buffer and a secondary buffer, extract the reply voice data from the main buffer for voice output, and transfer the data in the secondary buffer to the main buffer when the data in the main buffer is less than the threshold; asynchronously process voice synthesis and voice output; among them, when performing voice output, adjust the voice effect of the voice output according to the identity characteristics of the simulated patient in the consultation text.
9. A terminal, characterized in that, Including: A memory, used to store the medical pre-consultation real-time interaction program; A processor, used to implement the steps of the medical pre-consultation real-time interaction method as described in any one of claims 1-7 when executing the medical pre-consultation real-time interaction program.
10. A computer-readable storage medium, characterized in that, The medical pre-consultation real-time interaction program is stored on the readable storage medium, and when the medical pre-consultation real-time interaction program is executed by the processor, it implements the steps of the medical pre-consultation real-time interaction method as described in any one of claims 1-7.
Citation Information
Cited By
Voice generation method and electronic equipment
CN121725763A
Voice generation method and electronic device
CN121725763B