Medical intelligent voice interaction processing system
By combining directional microphone array and voiceprint biometric lock technology with the medical semantic analysis module of the AI decision-making layer, the problems of noise interference and missed diagnosis in medical scenarios are solved, efficient and accurate voice interaction processing and diagnosis support are achieved, and the quality of medical services is improved.
Patent Information
- Application Number
- CN202510978408.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-14
AI Technical Summary
In existing medical scenarios, environmental noise interference, high risk of missed diagnosis and efficiency bottlenecks seriously restrict the improvement of medical service quality and diagnosis and treatment efficiency. Traditional recording equipment cannot distinguish between the voices of doctors and patients. The speech recognition system is limited by homophone interference and lacks the ability to understand medical semantics. The electronic medical record auxiliary system cannot dynamically identify missed items.
A directional microphone array, voiceprint biometric lock and doctor ID card integrated module are used for sound collection and filtering, combined with the environmental noise reduction module and role separation engine, and the medical semantic parsing module and missed diagnosis warning module of the AI decision layer are used to extract medical entities and provide examination suggestions, which are output through the AR visualization interface and electronic medical record automatic generation module.
It achieves accurate collection and separation of the voices of doctors and patients in noisy environments, reduces noise interference, significantly improves voice recognition accuracy and diagnostic comprehensiveness, reduces missed diagnosis and over-examination rates, and improves diagnosis and treatment efficiency and privacy protection.
Smart Images

Figure CN120784010A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent medical technology, and in particular relates to a medical intelligent voice interaction processing system. Background Art
[0002] In modern medical scenarios, problems such as environmental noise interference, high risk of missed diagnosis, and efficiency bottlenecks seriously restrict the improvement of medical service quality and diagnosis and treatment efficiency.
[0003] To address the above issues, existing technical solutions mainly include traditional recording equipment, voice recognition systems, and electronic medical record assistance systems. Traditional recording equipment uses omnidirectional microphones with noise reduction algorithms such as spectral subtraction. Voice recognition systems work through a process of voice-to-text conversion, keyword matching, and template generation. Electronic medical record assistance systems rely on pre-set consultation templates. However, these existing technologies all have significant flaws: traditional recording equipment cannot distinguish between the voices of doctors, patients, and family members, and the accuracy rate of key information extraction is less than 50%; voice recognition systems are limited by homophone interference, such as misidentifying "hepatitis" as "dry eyes," and lack medical semantic understanding capabilities; electronic medical record assistance systems cannot dynamically identify missed items, making it difficult to associate symptoms with examination recommendations.
[0004] In view of the shortcomings of the existing technology, in order to effectively solve the problems of voice interaction and medical record processing in medical scenarios, the present invention proposes a medical intelligent voice interaction processing system. Summary of the Invention
[0005] The purpose of the present invention is to provide a medical intelligent voice interaction processing system, aiming to solve the problems raised in the above background technology.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] Medical intelligent voice interaction processing system, including:
[0008] The main control sound collection layer includes a directional microphone array, voiceprint biometric lock, and doctor ID card integration module, which is used to directionally collect and filter the voice signals of doctors and patients;
[0009] The data processing layer includes an environmental noise reduction module and a character separation engine for noise reduction and separation of the doctor's and patient's voices;
[0010] The AI decision-making layer includes a medical semantic parsing module, a missed diagnosis warning module, and an examination recommendation engine, which are used to extract medical entities, generate differential diagnosis trees, and recommend examination plans;
[0011] The output layer includes an AR visualization interface and an electronic medical record automatic generation module, which is used to prompt doctors and generate structured medical records;
[0012] The system realizes voice interaction processing in medical scenarios through the collaboration of voiceprint identity authentication, multimodal biometrics, intelligent scene perception, dynamic sound source positioning and real-time medical decision support.
[0013] Furthermore, the output signal y(t) is expressed as:
[0014] y(t)=w H x(t);
[0015] Among them, w H is the conjugate transpose of w; x(t) is an N×1 column vector containing the signal received by each MEMS unit.
[0016] Furthermore, in the data processing layer:
[0017] The environmental noise reduction module uses a deep noise suppression algorithm to improve the signal-to-noise ratio by ≥20dB, and maps the input noisy speech signal x(n) to an estimate of the clean speech through the neural network model f(x). Expressed as:
[0018]
[0019] Among them, θ is the parameter of the neural network;
[0020] The role separation engine uses the uPIT-BLSTM model to capture contextual information through forward hidden states and backward hidden states, separate the doctor and patient voices, and generate independent audio tracks; its forward and backward propagation formulas are:
[0021]
[0022] in, is the hidden state of the forward LSTM at time t; LSTM represents the calculation of the LSTM unit; is the hidden state of the forward LSTM at time t-1; x t is the input feature vector at time t; is the hidden state of the backward LSTM at time t; is the hidden state of the backward LSTM at time t+1.
[0023] Furthermore, in the AI decision-making layer:
[0024] The medical semantic parsing module is based on the BiLSTM-CRF model of the SNOMED CT terminology library, which extracts symptoms, signs, and medical history entity information from the speech-to-speech text. The model establishment process includes data preparation, vocabulary construction, and model architecture establishment.
[0025] The missed diagnosis warning module is driven by the knowledge graph, generates a differential diagnosis tree in real time, calculates the weights of undiscussed nodes, and triggers graded inspection recommendations based on the weights. Its algorithm includes inputting a real-time voice stream and performing text conversion, entity extraction, knowledge graph traversal, and outputting graded inspection recommendations.
[0026] The examination recommendation engine combines the QALY weight optimization model to recommend a graded examination plan based on cost-benefit analysis.
[0027] Furthermore, in the medical semantic parsing module, the architecture of the BiLSTM-CRF model includes an embedding layer, a bidirectional long short-term memory network layer, and a conditional random field layer. The embedding layer converts the input text sequence into a low-dimensional vector representation, the bidirectional long short-term memory network layer processes forward and reverse text sequence information, and the conditional random field layer considers the dependency relationship between adjacent labels to predict entity boundaries and categories.
[0028] Furthermore, in the output of the missed diagnosis warning module, the hierarchical triggering mechanism of the examination recommendations is: when the weight is greater than 0.7, a strong recommendation is generated for unexamined items with a mortality rate greater than 5%; when the weight is in the range of 0.3-0.7, a general recommendation is generated for auxiliary examination items with a misdiagnosis rate greater than 30%; when the weight is less than 0.3, an optional examination recommendation is generated; where the weight = disease mortality rate × (1-coverage rate of examined items).
[0029] Furthermore, the AR visualization interface uses eye tracking technology to project the supplementary consultation items prompted by the missed diagnosis warning module into the doctor's AR glasses field of view in red highlight form.
[0030] The medical intelligent voice interaction processing method is applied to the medical intelligent voice interaction processing system described above, and includes the following steps:
[0031] Directedly collect and filter the voice signals of doctors and patients through the main control sound collection layer;
[0032] The collected speech signals are subjected to noise reduction and character separation through the data processing layer;
[0033] The AI decision layer extracts medical entities from processed speech signals, generates differential diagnosis trees, and recommends examination plans;
[0034] AR visualization prompts and electronic medical record generation are performed through the output layer.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention deeply combines acoustic role separation with dynamic reasoning of medical knowledge graphs for the first time. Traditional speech recognition technology is limited to general processing of speech-to-text, while the present invention realizes the precise locking of the "voice of the main control role" in medical scenarios through the voiceprint biometric lock mechanism of doctors and patients (such as using MFCC features bound to UWB positioning technology). Specifically, when the doctor wears a badge with an integrated UWB tag, the system dynamically activates the beamforming algorithm based on the spatial distance threshold (1.5 meters), and combines it with voiceprint verification technology to effectively filter out irrelevant personnel (such as family members) and environmental noise interference. This "acoustic + spatial" two-dimensional verification mechanism breaks through the technical bottleneck of single voice separation and provides a new idea for the accurate collection of medical conversations.
[0037] 2. The dynamic mandatory inspection item prediction algorithm in the present invention breaks through the limitations of existing technologies. Existing auxiliary diagnosis systems mostly rely on static rule bases or simple keyword matching, while the present invention constructs a three-dimensional graph of symptoms-diseases-examinations and introduces a real-time weight calculation model. Taking the patient's description of "chest pain" but not "radiation to the left shoulder" as an example, the system dynamically calculates the associated weights of the undiscussed symptoms by traversing the knowledge graph nodes. The calculation formula is "weight = disease mortality rate × (1-coverage of inspected items)", which realizes the technological leap from static rule reasoning to dynamic knowledge graph traversal, significantly improving the comprehensiveness and accuracy of the diagnosis process.
[0038] 3. The present invention solves the special needs of medical scenarios through privacy-oriented interactive design and localized data processing architecture. Traditional systems need to upload voice data to the cloud, which poses a risk of privacy leakage and has high latency. The present invention uses a federated learning framework to complete voice recognition, entity extraction, and knowledge graph matching locally. The data is fully encrypted and does not need to be transmitted externally. At the same time, when the doctor receives a missed diagnosis prompt through AR glasses, the system only projects it to the doctor's personal field of view in a "line of sight locked" manner, and outsiders cannot peek. For example, when the system detects that the patient has not described a "history of drug allergies", the AR interface will automatically highlight the entry when the doctor looks at the electronic medical record area, while the patient's perspective remains seamless. This "unconscious" privacy protection mechanism takes into account both the efficiency of information prompts and the psychological comfort of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is the system framework diagram.
[0040] Figure 2 Schematic diagram of the ring PCB and MEMS microphone.
[0041] Figure 3 This is the flow chart of the missed diagnosis warning algorithm. DETAILED DESCRIPTION
[0042] In order to have a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention is now described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.
[0043] The specific implementation of the present invention is described in detail below with reference to specific embodiments.
[0044] The present invention provides a medical intelligent voice interaction processing system, the framework of which is shown in the following figure: Figure 1 Shown, including:
[0045] 1. Main control sound collection layer:
[0046] 1. Directional microphone array: A 4cm diameter circular PCB board is used as the carrier, on which 8 MEMS microphones with a signal-to-noise ratio greater than 70dB are integrated ( Figure 2 ), can clearly capture weak sound signals in complex medical environments. Equipped with a beamforming chip (TI TLVAIC3254), it supports dynamic sound source tracking, automatically adjusting the microphone's sound pickup direction based on the doctor and patient's position, keeping the main lobe width within a ±15° range to accurately lock onto the sound source. This greatly improves the targetedness and effectiveness of sound collection and effectively reduces external noise interference.
[0047] Assume that there are N MEMS units in a uniform linear array and the received signal vector x(t) = [x1(t), x2(t), …, x N (t)] T , x(t) is an N×1 column vector containing the signal received by each MEMS unit, where t represents time; x n (t) represents the signal received by the nth MEMS unit at time t, where n = 1, 2, ..., N; T represents the transposition. The output signal y(t) after the sound beamforming is collected is expressed as y(t) = w H x(t), where w=[w1,w2,…,w N ] T , is the beamforming weight vector, which is also an N×1 column vector; wn represents the weight of the nth MEMS unit; w H is the conjugate transpose of w.
[0048] 2. Voiceprint biometric lock:
[0049] The doctor's side pre-registers the voiceprint (MFCC+Delta-MFCC composite features), and the patient's side dynamically extracts the voiceprint and automatically filters the family members' voiceprints (blacklist matching), thereby building a voiceprint biometric lock to further improve the targeted nature of voice collection.
[0050] 3. Doctor ID card integration module;
[0051] The doctor's badge integrates a UWB positioning chip (accuracy ±5cm) to ensure effective sound reception within 1.5 meters, and an integrated bone conduction vibration unit provides privacy (sound leakage rate less than 5%). The patient's wristband is bound to an RFID tag, which works in conjunction with the doctor's badge's UWB tag to dynamically activate beamforming based on a distance threshold (1.5 meters) to optimize sound collection.
[0052] 2. Data processing layer:
[0053] 1. Environmental noise reduction module: Using the Deep Noise Suppression (DNS) algorithm, the signal-to-noise ratio is improved by ≥20dB, which can effectively reduce environmental noise interference.
[0054] In the DNS algorithm based on deep learning, the input noisy speech signal x(n) is usually processed through a series of processes. Assuming that the neural network model is f(x), it attempts to learn a mapping that maps the noisy speech to an estimate of the clean speech. It can be expressed as:
[0055]
[0056] Among them, θ is the parameter of the neural network, including weights and biases.
[0057] 2. Role Separation Engine: Using the speaker separation algorithm (uPIT-BLSTM model), the uPIT-BLSTM (Unified Permutation-Invariant Training for Bidirectional Long-Short Term Memory) model is a model for speaker separation. and backward The hidden state captures contextual information, separates the doctor's and patient's speech, and generates independent audio tracks.
[0058] The forward propagation formula is:
[0059]
[0060] in, is the hidden state of the forward LSTM at time t, which calculates the hidden state at the current moment based on the hidden state at the previous moment and the current input; x t is the input feature vector at time t, such as the spectral features of audio; LSTM represents the calculation of the LSTM unit; is the hidden state of the forward LSTM at time t-1.
[0061] The backpropagation formula is:
[0062]
[0063] in, is the hidden state of the backward LSTM at time t, which is calculated based on the hidden state at the next moment and the current input; is the hidden state of the backward LSTM at time t+1.
[0064] 3. AI Decision-Making Layer
[0065] The separated speech enters the AI decision-making layer, which includes a medical semantic parsing module, a missed diagnosis warning module, and an examination recommendation engine. The medical semantic parsing module converts speech into text and extracts key entity information based on the BiLSTM-CRF model of the SNOMED CT terminology library; the missed diagnosis warning module traverses and analyzes symptom information based on the knowledge graph (UpToDate clinical database), generates a differential diagnosis tree in real time, and determines whether there is a risk of missed diagnosis; the examination recommendation engine combines a cost-benefit analysis model to recommend a graded examination plan based on the analysis results. The details are as follows:
[0066] 1. Medical semantic parsing module: This module is based on the BiLSTM-CRF model of the SNOMED CT terminology library and can identify entity information such as symptoms, signs, and medical history in medical texts, providing strong support for medical information processing and analysis.
[0067] The model establishment, training and verification process is as follows:
[0068] 1.1 Build the model;
[0069] (1) Data preparation: First, a large amount of text data is collected from various medical data sources such as electronic medical records and medical literature. This data should cover different medical fields and topics to ensure the generalization ability of the model. Then, professional medical annotators will annotate the medical entities in the text and determine the boundaries and categories of each entity. For example, "pneumonia" is annotated as a "disease" category, and "aspirin" is annotated as a "drug" category. Finally, the annotated data is divided into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used to adjust the model's hyperparameters and monitor the model's training process, and the test set is used to evaluate the model's final performance.
[0070] (2) Vocabulary Construction: First, the training set text is traversed and all the words or characters that appear are constructed into a vocabulary. Each word or character is assigned a unique index for representation in the model. For words or characters that do not appear in the vocabulary, special tags are used. Then, terms from the SNOMED CT terminology are added to the vocabulary and assigned indexes to enable the model to better utilize the prior knowledge in the terminology.
[0071] (3) Model architecture construction: First, the embedding layer converts each word or character in the input text sequence into a low-dimensional vector representation. Pre-trained word vectors such as Word2Vec or GloVe can be used, or word vectors can be randomly initialized and learned during the training process. Character-level embedding can use models such as CNN to extract the feature vectors of characters. Then the bidirectional long short-term memory network (BiLSTM) layer receives the vector sequence output by the embedding layer. It can process forward and reverse text sequence information at the same time and better capture the long-term dependencies in the text. At each time step, BiLSTM outputs a hidden state vector containing contextual information. Finally, the conditional random field (CRF) layer takes the output of the BiLSTM layer as input. The CRF layer is used to model the entire sequence, considering the dependencies between adjacent labels, so as to more accurately predict the boundaries and categories of entities. It calculates the probability of each possible label sequence and selects the sequence with the highest probability as the final prediction result.
[0072] 1.2 Model Training: First, define the loss function. For medical entity recognition tasks based on the BiLSTM-CRF model, the negative log-likelihood loss function is commonly used. This loss function measures the difference between the model's predicted label sequence and the true label sequence. Model parameters are adjusted by minimizing the loss function. Next, select an optimizer. Common optimizers such as stochastic gradient descent (SGD), Adagrad, Adadelta, RMSProp, or Adam can be used for model training. The Adam optimizer generally performs well in practice because it can adaptively adjust the learning rate and update parameters based on the gradient history of each parameter. Next, set hyperparameters. These include the learning rate, batch size, number of training epochs, number of LSTM units, and embedding dimension. The choice of these hyperparameters affects model performance and training speed. Experimentation is usually required to adjust hyperparameters to find the optimal combination. For example, setting the learning rate to 0.001, the batch size to 32 or 64, and the number of training epochs to 50 to 100 epochs can be helpful. Finally, train the model. During training, the training set data is fed into the model in batches according to the batch size. For each batch of data, the model will forward propagate to calculate the prediction results and calculate the loss value based on the loss function. It will then use the backpropagation algorithm to calculate the gradient of each parameter and use the optimizer to update the parameters based on the gradient. This process is repeated until the preset number of training rounds is reached or the loss value converges.
[0073] 1.3 Model Validation: During the training process, the validation set is used to evaluate the model performance at regular intervals. The validation set data is input into the model to obtain the model's prediction results. Evaluation indicators such as precision, recall, and F1 value are calculated between the prediction results and the true labels.
[0074] Precision is the ratio of true positive examples to samples predicted as positive examples. The calculation formula is: Precision = TP / (FP+TP), where TP is the number of true positive examples and FP is the number of false positive examples.
[0075] Recall is the ratio of true positive examples to those predicted as positive examples. The calculation formula is: Recall = TP / (FN + TP), where FN is the number of false negative examples.
[0076] The F1 value is the harmonic mean of precision and recall, and the calculation formula is:
[0077] Adjust the model parameters based on the validation set evaluation results. If the F1 value on the validation set stops improving or overfitting occurs (the training set loss continues to decrease while the validation set loss begins to increase), consider adjusting hyperparameters, such as reducing the learning rate, increasing regularization, adjusting the batch size, or changing the model structure. Then retrain the model until optimal performance on the validation set is achieved. During training, save the model parameters that perform best on the validation set. This model is the final model used in the actual application. It exhibits optimal performance on unseen validation set data and has good generalization ability.
[0078] 2. Missed Diagnosis Warning Module: This module is driven by a knowledge graph (UpToDate clinical database) and can generate a differential diagnosis tree (DDx Tree) in real time, providing clinicians with fast and accurate diagnostic support, helping them better handle complex clinical cases and improve medical quality. The algorithm flow, knowledge graph construction, differential diagnosis tree generation, and real-time update optimization process are as follows:
[0079] 2.1 Missed Diagnosis Warning Algorithm Process ( Figure 3 ):
[0080] (1) Input: Real-time speech stream → text conversion (medical-specific ASR model with an error rate of less than 5%).
[0081] The system receives a real-time speech stream as initial data and converts it into text using a medical-specific ASR model with an incremental learning architecture. This model is optimized for vocabulary and language characteristics in the medical field and supports online updates of the hospital's localized terminology library (such as department-specific disease categories). With fewer than 100 training samples, the recognition accuracy is greater than 90%, and the overall error rate is kept below 5%. By continuously learning the specialized terminology of specific hospital scenarios, the model can accurately convert the speech of doctors and patients into text format, providing the foundational data for subsequent analysis and processing.
[0082] (2) Entity extraction;
[0083] Symptoms: fever, cough, chest pain, etc. (based on SNOMED CT codes).
[0084] Time description: "Last 3 days" → quantified to 72 hours.
[0085] Based on SNOMED CT codes, key entity information is accurately extracted from the converted text. For symptoms, it can identify common symptoms such as fever, cough, and chest pain. For time descriptions, it can quantify ambiguous expressions such as "lasting three days" into 72 hours, standardizing time information and facilitating subsequent correlation analysis. This precise entity extraction provides key information for in-depth analysis of the patient's condition.
[0086] (3) Knowledge graph traversal: symptom node → associated disease → verification of required items (e.g., if chest pain is detected but an electrocardiogram is not checked, an early warning will be triggered).
[0087] Starting with the extracted symptom nodes, the system conducts a traversal analysis within a knowledge graph built from the UpToDate clinical database. The system automatically associates related diseases and verifies mandatory checklists for each condition. For example, if a patient describes "chest pain" and the system determines an electrocardiogram (ECG) was not performed, an early warning mechanism will be triggered to identify potential missed diagnoses, ensuring comprehensive and accurate diagnosis.
[0088] (4) Output;
[0089] The system uses a dynamic mandatory inspection item prediction algorithm based on a three-dimensional symptom-disease-examination map to calculate the weight of undiscussed nodes in real time: weight = disease mortality rate × (1-coverage of examined items). Based on this weight, combined with a hierarchical triggering mechanism for examination recommendations, the system outputs different levels of examination recommendations: when the weight is greater than 0.7, a strong recommendation is generated for unexamined items with a mortality rate greater than 5%. For example, if a patient has a headache but has not undergone a CT scan, the system will prompt a brain tumor screening. When the weight is in the range of 0.3-0.7, a general recommendation is generated for auxiliary examination items with a misdiagnosis rate greater than 30%. For example, if a patient has a cough but has not been asked about his allergy history, the system will recommend an IgE test. When the weight is less than 0.3, an optional examination recommendation is generated. This hierarchical output mechanism can help doctors rationally arrange examination items, improve diagnosis and treatment efficiency, and avoid excessive examinations.
[0090] 2.2 Knowledge Graph Construction: A large amount of clinical knowledge data, including disease symptoms, signs, examination results, diagnostic criteria, and treatment methods, is collected from the UpToDate clinical database. This data forms the foundation for constructing the knowledge graph. Natural language processing techniques are used to extract information from the collected text data. For example, entities such as the disease name, related symptoms, and common causes, as well as the relationships between them, are extracted from disease descriptions. For example, from the sentence "Pneumonia is often accompanied by fever, cough, and sputum," the disease entity "pneumonia" and symptom entities such as "fever," "cough," and "sputum" are extracted, and relationships such as "pneumonia-accompanied-fever" are established. The extracted knowledge is integrated to eliminate redundancy and inconsistency in the data. For example, different descriptions of the same disease in different literature are unified, such as "myocardial infarction" and "acute myocardial infarction," into a single standardized disease entity. Furthermore, knowledge from different sources is integrated, such as integrating the relationships between symptoms and diseases and between diseases and causes into a comprehensive knowledge graph. Use appropriate knowledge representation methods, such as Resource Description Framework (RDF) or property graphs, to formally represent the entities and relationships in the knowledge graph. For example, using RDF, "pneumonia," "fever," and "accompanied by" can be represented as a triple: <pneumonia, accompanied by, fever>, making this knowledge accessible to computers.
[0091] 2.3 Generate a differential diagnosis tree: Clinicians or patients input information such as patient symptoms, signs, and related test results into the system. For example, input symptoms such as "fever, cough, and chest pain." The system queries the knowledge graph based on the input symptom information, searches for disease entities related to these symptoms and their relationships, and finds a list of diseases that may cause these symptoms. For example, the knowledge graph queries "pneumonia," "pleurisy," and "lung cancer," which may all be accompanied by symptoms such as "fever, cough, and chest pain." Construct a differential diagnosis tree, using the input symptom set as the root node of the differential diagnosis tree and possible diseases as intermediate nodes, connected to the root node through relationships such as "may cause." For example, disease nodes such as "pneumonia," "pleurisy," and "lung cancer" are connected to the root node "fever, cough, and chest pain." Further, based on the characteristics and relevant knowledge of the disease, subnodes are added under each disease node, such as common causes of the disease, different types, and related test results. For example, under the "Pneumonia" node, add subnodes such as "Bacterial Pneumonia" and "Viral Pneumonia," as well as test result subnodes such as "Elevated White Blood Cells in a Routine Blood Test" and "Chest X-ray Shows Pulmonary Inflammatory Shadows." By continuously expanding and connecting nodes, a complete differential diagnosis tree is formed. Each branch in the tree represents a possible diagnostic path. The path from the root node to the leaf nodes can help doctors gradually narrow down the diagnostic scope and ultimately determine an accurate diagnosis.
[0092] 2.4 Real-time Updates and Optimization: The UpToDate clinical database is regularly updated with information on new clinical research findings, changes in disease diagnostic criteria, and advances in treatment methods. The knowledge graph is updated accordingly to reflect these changes. For example, when a new disease associated with certain symptoms is discovered, it is added to the knowledge graph, and new nodes and relationships are added to the differential diagnosis tree accordingly. Machine learning and artificial intelligence technologies are used to optimize the knowledge graph and differential diagnosis tree. For example, by analyzing large amounts of clinical case data, new associations between certain symptoms and diseases can be discovered, or the structure of existing differential diagnosis trees can be adjusted to better reflect clinical reality, thereby improving diagnostic accuracy and efficiency. The system provides a user feedback interface. When using the differential diagnosis tree, clinicians can compare their actual diagnosis results with the system-generated differential diagnosis and provide feedback to the system. If the system's diagnosis results are found to be inconsistent with the actual situation, the system will correct and improve the knowledge graph and differential diagnosis tree based on this feedback, continuously improving system performance and reliability.
[0093] 3. Examination Recommendation Engine: This engine combines a cost-benefit analysis model (QALY weight optimization) to recommend a tiered examination plan based on the analysis results. The model establishment, training, and validation process is as follows:
[0094] 3.1 Model Building: First, identify the medical intervention or health policy to be analyzed and the specific objectives for assessing its cost-effectiveness. For example, when studying the cost-effectiveness of a new drug for a specific disease, the goal might be to determine the optimal use strategy for the drug in different populations to maximize health benefits. Next, select the model structure: Based on the research question and available data, choose an appropriate cost-effectiveness analysis model structure. Common models include decision tree models and Markov models. Decision tree models are suitable for simple decision-making scenarios, while Markov models are more suitable for addressing the dynamic processes and long-term effects of diseases. Next, define the model parameters, including cost parameters, health status parameters, and transition probability parameters. Finally, collect the various data required for the model, including cost data, health status data, and transition probability data. Data sources can include hospital medical records, medical insurance reimbursement records, clinical trial data, epidemiological studies, etc. Furthermore, ensure data quality and reliability, and appropriately address missing data through interpolation or sensitivity analysis.
[0095] 3.2 Model Training: Initial QALY weights must be set for each health state at the beginning of model training. These initial values can be based on existing literature, universal health utility scales, or subjective expert judgment. For example, based on previous research, the QALY weight for mild symptom states of a chronic disease can be initialized to 0.7, and the QALY weight for severe symptom states can be initialized to 0.4. An appropriate training algorithm, such as gradient descent or genetic algorithm, is selected to optimize the QALY weights. Algorithm parameters such as the learning rate and number of iterations are set. The learning rate determines the step size for each update of the QALY weights, while the number of iterations determines the termination criteria for training. The collected data is input into the model, and the QALY weights are continuously adjusted according to the selected training algorithm to ensure that the model predictions closely match the actual data. At each iteration, the model's loss function, such as mean squared error or logarithmic loss, is calculated to evaluate model performance. The QALY weights are then updated based on the gradient of the loss function, gradually decreasing the loss function. For example, when using the gradient descent algorithm, the partial derivative of the loss function with respect to the QALY weight is multiplied by the learning rate and then subtracted from the current QALY weight to obtain the updated QALY weight. During training, regularly monitor the model's performance indicators, such as the loss function value and prediction accuracy. You can plot the change in the loss function over the number of iterations to observe whether the model has converged. If the loss function no longer decreases significantly after a certain number of iterations, the model may have converged, and training can be stopped.
[0096] 3.3 Model Validation: First, divide the collected data into a training dataset and a validation dataset, with approximately 70-80% for training and 20-30% for validation. Ensure similar characteristics and distributions between the two datasets to ensure the validity of the validation results. Apply the trained model to the validation dataset and calculate performance metrics on the validation dataset, such as mean squared error, mean absolute error, and prediction accuracy. These metrics assess the model's generalization to new data, that is, its performance on unseen data. Perform a sensitivity analysis of key model parameters (including the QALY weight). By varying the QALY weight, observe how the model results change. For example, adjust the QALY weight for a particular health state by a certain percentage, such as 10% or 20%, and then recalculate the model's cost-effectiveness metrics, such as the incremental cost-effectiveness ratio (ICER). If the change in ICER exceeds a pre-defined threshold, this QALY weight is highly sensitive to the model's results and requires further careful processing and validation. Compare the developed cost-effectiveness analysis model with other existing similar models or benchmarks. This can be used to compare the model's performance metrics, predictions, or cost-effectiveness conclusions. If a new model demonstrates superior accuracy, reliability, or clinical utility, it demonstrates its validity and value. Medical experts, health economists, and other professionals in related fields should be invited to review the model's rationality, the plausibility of its assumptions, and the clinical significance of its results. The model can also be validated on a small scale in clinical practice to observe whether the model's predictions align with actual clinical conditions, further verifying the model's validity and applicability.
[0097] Through the above steps, a QALY weight optimization model for cost-effectiveness analysis can be established, trained, and validated, providing a scientific basis for medical decision-making. It is important to note that model establishment and validation is an iterative process, and adjustments and improvements may be required based on actual circumstances to improve the model's accuracy and reliability.
[0098] 4. Output layer:
[0099] 1. AR visualization interface: With the help of eye tracking technology, the doctor's line of sight is locked, and the supplementary medical consultation items prompted by the missed diagnosis warning module are projected in red highlight form into the doctor's field of view of AR glasses, realizing a convenient reminder function.
[0100] 2. Automatic electronic medical record generation: This function generates structured output of medical consultation records, examination recommendations, and differential diagnosis flowcharts. This feature improves the efficiency and standardization of medical record keeping, helping doctors fully understand the patient's condition, thereby improving the efficiency and accuracy of diagnosis and treatment, and optimizing medical service processes.
[0101] The specific implementation of the present invention is described in detail below with reference to specific embodiments.
[0102] Example 1: Verification of the effectiveness of the system's core functions;
[0103] 1. In a noisy outpatient environment, the system uses a directional microphone array and voiceprint biometric lock technology to accurately lock onto the voices of doctors and patients, effectively filtering out interference from family members and environmental noise. For example, under background noise levels of 60-75dB, measured data (Table 1) shows that voice recognition accuracy can reach 85%. Compared to traditional omnidirectional microphones (recognition rates are less than 50% when the signal-to-noise ratio is less than 10dB), this technology achieves a breakthrough in key information extraction capabilities. This technology not only solves the problem of sound mixing, but also dynamically tracks the spatial position of doctors and patients through UWB positioning, ensuring complete capture of conversations within 1.5 meters and avoiding blind spots caused by movement.
[0104] Table 1 Comparison of speech recognition technology effects
[0105]
[0106]
[0107] 2. In terms of missed diagnosis warning, the system's dynamic analysis capabilities based on the medical knowledge graph are particularly outstanding. By analyzing symptom correlations in real time and generating a differential diagnosis tree, the comprehensiveness of diagnosis is significantly improved. Data shows (Table 2) that the missed diagnosis rate of rare diseases has been significantly reduced from 18.7% in traditional manual consultations to 4.3%. For example, when a patient describes "joint pain" but does not mention "light sensitivity", the system will immediately trigger a lupus erythematosus screening recommendation and automatically prescribe an antinuclear antibody test, reducing the diagnosis time for such diseases from an average of 2.3 visits to 1, significantly improving clinical efficiency. In addition, the AI-driven examination recommendation engine, combined with a cost-effectiveness model, has reduced the over-examination rate from 32% to 12%. For example, if an emergency chest pain patient does not undergo an electrocardiogram in a timely manner, the system will project a red highlighted alert to the doctor through AR glasses, helping to shorten the treatment time (D2B) for myocardial infarction patients from 90 minutes to 58 minutes, directly improving the success rate of rescue.
[0108] Table 2 Effect data of missed diagnosis warning system
[0109] Indicator Traditional manual diagnosis AI-assisted system Improvement range Rare disease missed diagnosis rate 18.7% 4.3% Reduced by 77% Number of visits required for diagnosis (case of lupus erythematosus) 2.3 times 1 time Reduced by 56.5% Over-examination rate 32% 12% Reduced by 62.5% Treatment time for heart attack patients (D2B) 90 minutes 58 minutes Shortened by 35.6%
[0110] 3. In terms of efficiency and privacy, doctors receive real-time prompts through AR glasses, and missed items are only presented in a visually highlighted form in their personal field of view, which not only avoids patient anxiety caused by traditional screen displays, but also reduces information overload interference. Data shows (Table 3) that the automated generation of electronic medical records reduces paperwork time by 40%, the completeness of structured output consultation records reaches 95%, and the automatic association of examination results and diagnostic recommendations significantly reduces the burden on doctors. On the hardware level, the patient wristband is made of medical silicone material and can be worn continuously for 24 hours without irritation. The localized federated learning framework ensures that model optimization can be completed without external data transmission, fully complying with the privacy compliance requirements of HIPAA and GDPR.
[0111] Table 3 Work efficiency improvement data
[0112] Function Traditional way New technology Improvement effect Time of clerical work 100% benchmark Reduced by 40% Efficiency improved by 66.7% Completeness of diagnosis records Not applicable 95% Structured output
[0113] These innovations make the technology not only suitable for outpatient and emergency scenarios, but also can be seamlessly connected to the hospital's HIS and PACS systems, driving the comprehensive upgrade of the diagnosis and treatment process towards intelligence and precision.
[0114] Example 2: Typical scenario implementation verification (scenario-based cases);
[0115] This example uses two real-world scenarios, outpatient rare disease screening and emergency chest pain triage, to verify the practicality and effectiveness of the medical intelligent voice interaction processing system in specific diagnosis and treatment processes, and combines data comparison to demonstrate how the technology improves clinical efficiency and accuracy.
[0116] Sub-scenario 1: Diabetes complications screening in the endocrinology clinic;
[0117] Configuration environment:
[0118] Application Department: Endocrinology Department of a tertiary hospital;
[0119] System connection: hospital information system (HIS), laboratory database, electronic health record (EHR);
[0120] Hardware configuration: intelligent microphone array, doctor AR glasses, UWB positioning base station, edge computing server;
[0121] Knowledge base version: Diabetes complications knowledge graph v3.2 (covering 12 types of complications such as neuropathy, retinopathy, and nephropathy).
[0122] Operation process steps:
[0123] Table 4 Operation process steps
[0124]
[0125]
[0126] Typical operation process example:
[0127] 1. The patient complained of frequent thirst, increased urination, and occasional numbness in the toes in the past three months;
[0128] 2. Doctor added: The patient has a 7-year history of diabetes and has had blurred vision for the past two months;
[0129] 3. System analysis: Combined with the medical history, it is detected that foot examination is not mentioned → triggering diabetic foot screening;
[0130] 4. Dynamic atlas: Associated symptoms: "blurred vision" → recommend retinal examination; "lower limb numbness" → recommend nerve conduction study;
[0131] 5. Examination and prescription: Automatically generate HbA1c, fundus photography, nerve conduction, and urine microalbumin test orders;
[0132] 6. Result feedback: Abnormal nerve conduction test values are highlighted through AR glasses when doctors review patient files;
[0133] 7. Diagnostic suggestion: The system prompts "diabetic peripheral neuropathy (moderate)" and recommends a treatment plan.
[0134] Comparison of performance indicators:
[0135] Table 5 Comparison of effect indicators
[0136]
[0137]
[0138] Clinical benefit analysis: The early diagnosis rate of diabetic peripheral neuropathy increased from 42% to 89%; the screening coverage of diabetic retinopathy increased from 65% to 98%; the average time from the onset of symptoms to diagnosis was shortened from 106 days to 14 days; the number of patients seen by doctors per day increased by 40%, while the quality of diagnosis improved; through the federated learning framework, the risk of patient data leakage was reduced by 95%.
[0139] Sub-scenario 2: Emergency chest pain triage;
[0140] Configuration: Emergency department pre-examination triage desk, integrated electrocardiograph and POCT equipment.
[0141] Operational process: A patient complains of chest pain for 2 hours, and the system detects that an ECG has not been performed. An ECG is strongly recommended. The ECG shows ST-segment elevation, automatically triggering a cardiology consultation.
[0142] Effect: D2B time (from presentation to balloon dilation) was reduced from 90 minutes to 58 minutes in acute myocardial infarction patients (Table 6).
[0143] Table 6 Emergency chest pain triage
[0144]
[0145] The data of the two types of scenarios collectively show that the system can effectively make up for the shortcomings of the traditional diagnosis and treatment process and promote the intelligent and precise transformation of medical services.
[0146] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the concept of the present application, can make a number of deformation and improvement, these should be considered as the protection scope of the present application, these will not affect the effect and practicality of the patent of the present application.
Claims
1. Medical intelligent voice interaction processing system, characterized by: include: The main control sound collection layer includes a directional microphone array, a voiceprint biometric lock, and a doctor's ID card integrated module, which is used to directionally collect and filter the voice signals of doctors and patients; The data processing layer includes an environmental noise reduction module and a character separation engine for noise reduction and separation of the doctor's and patient's voices; The AI decision-making layer includes a medical semantic parsing module, a missed diagnosis warning module, and an examination recommendation engine, which are used to extract medical entities, generate differential diagnosis trees, and recommend examination plans; The output layer includes an AR visualization interface and an automatic electronic medical record generation module, which is used to prompt doctors and generate structured medical records; The system realizes voice interaction processing in medical scenarios through the collaboration of voiceprint identity authentication, multimodal biometrics, intelligent scene perception, dynamic sound source positioning and real-time medical decision support.
2. The medical intelligent voice interaction processing system according to claim 1, characterized in that: The directional microphone array adopts beamforming technology, and the output signal y(t) is expressed as: y(t)=w H x(t); Among them, w H is the conjugate transpose of w; x(t) is an N×1 column vector containing the signal received by each MEMS unit.
3. The medical intelligent voice interaction processing system according to claim 1, characterized in that: In the data processing layer: The environmental noise reduction module uses a deep noise suppression algorithm to improve the signal-to-noise ratio by ≥20dB, and maps the input noisy speech signal x(n) to an estimate of the clean speech through the neural network model f(x). Expressed as: Among them, θ is the parameter of the neural network; The role separation engine uses the uPIT-BLSTM model to capture contextual information through forward hidden states and backward hidden states, separate the doctor and patient voices, and generate independent audio tracks; its forward and backward propagation formulas are: in, is the hidden state of the forward LSTM at time t; LSTM represents the calculation of the LSTM unit; is the hidden state of the forward LSTM at time t-1; x t is the input feature vector at time t; is the hidden state of the backward LSTM at time t; is the hidden state of the backward LSTM at time t+1.
4. The medical intelligent voice interaction processing system according to claim 1, characterized in that: In the AI decision-making layer: The medical semantic parsing module is based on the BiLSTM-CRF model of the SNOMED CT terminology library, which extracts symptoms, signs, and medical history entity information from the speech-to-speech text. The model establishment process includes data preparation, vocabulary construction, and model architecture establishment. The missed diagnosis warning module is driven by the knowledge graph, generates a differential diagnosis tree in real time, calculates the weights of undiscussed nodes, and triggers graded inspection suggestions based on the weights; Its algorithm process includes inputting real-time voice stream and performing text conversion, entity extraction, knowledge graph traversal, and outputting graded inspection suggestions; The examination recommendation engine combines the QALY weight optimization model to recommend a graded examination plan based on cost-benefit analysis.
5. The medical intelligent voice interaction processing system according to claim 4, characterized in that: In the medical semantic parsing module, the architecture of the BiLSTM-CRF model includes an embedding layer, a bidirectional long short-term memory network layer, and a conditional random field layer. The embedding layer converts the input text sequence into a low-dimensional vector representation, the bidirectional long short-term memory network layer processes forward and reverse text sequence information, and the conditional random field layer considers the dependency relationship between adjacent labels to predict entity boundaries and categories.
6. The medical intelligent voice interaction processing system according to claim 4, characterized in that: In the output of the missed diagnosis warning module, the hierarchical triggering mechanism for examination recommendations is as follows: when the weight is greater than 0.7, a strong recommendation is generated for unchecked items with a mortality rate greater than 5%; when the weight is in the range of 0.3-0.7, a general recommendation is generated for auxiliary examination items with a misdiagnosis rate greater than 30%; When the weight is less than 0.3, an optional inspection recommendation is generated; where weight = disease mortality rate × (1-coverage of inspected items).
7. The medical intelligent voice interaction processing system according to claim 1, characterized in that: The AR visualization interface uses eye tracking technology to project the supplementary consultation items prompted by the missed diagnosis warning module into the doctor's AR glasses field of view in red highlight form.
8. A medical intelligent voice interaction processing method, applied to the medical intelligent voice interaction processing system according to any one of claims 1 to 7, characterized in that: The following steps are involved: Directedly collect and filter the voice signals of doctors and patients through the main control sound collection layer; The collected speech signals are subjected to noise reduction and character separation through the data processing layer; The AI decision layer extracts medical entities from processed speech signals, generates differential diagnosis trees, and recommends examination plans; AR visualization prompts and electronic medical record generation are performed through the output layer.