A speech recognition-based cross-language barrier-free auxiliary diagnosis and treatment method

By collecting speech signals and lip video images, and combining them with the AFM model and medical knowledge graph, the problem of medical terminology mistranslation in cross-language assisted diagnosis and treatment has been solved, achieving highly accurate and safe cross-language assisted diagnosis and treatment.

CN121214914BActive Publication Date: 2026-02-24BEIJING XINRUISHI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511454581.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-02-24
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Existing cross-language-assisted diagnosis and treatment methods lack a deep understanding of medical context, cannot effectively distinguish between medical terms that sound similar but have different meanings, leading to mistranslation and medical risks, and do not quantify or assess the uncertainties in the identification and translation process.

Method used

By acquiring speech signals and lip video images, acoustic features and visual lip reading features are extracted. The AFM model is used for speech recognition, and the text is corrected and translation confidence is assessed by combining medical knowledge graphs. The comprehensive uncertainty score is calculated and the risk level is classified, and corresponding interactive operations are performed.

Benefits of technology

It improves the accuracy of speech recognition and the reliability of clinical semantics in cross-language assisted diagnosis and treatment, reduces the medical risks caused by misidentification and translation bias, and realizes a reliable, safe and intelligent cross-language assisted diagnosis and treatment throughout the entire process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214914B_ABST
    Figure CN121214914B_ABST
Patent Text Reader

Abstract

The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The application discloses a cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, and relates to the technical field of language recognition. The
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language recognition technology, and in particular to a cross-language barrier-free assisted diagnosis and treatment method based on speech recognition. Background Technology

[0002] With the deep integration of artificial intelligence and medical information technology, intelligent assisted diagnosis and treatment mechanisms have shown great potential in improving the efficiency of medical services and optimizing doctor-patient communication. Among them, speech recognition methods, as the core means of human-computer interaction, have been widely used in clinical consultation, electronic medical record generation, and telemedicine. In recent years, deep learning models, especially encoder-decoder architectures based on attention mechanisms, have made progress in speech recognition tasks, effectively processing continuous speech and generating high-accuracy text transcriptions. At the same time, multimodal fusion methods, by combining audio signals and visual lip movement information, have further improved the accuracy of speech recognition in complex acoustic environments (such as noisy hospital environments). In addition, the development of natural language processing methods has enabled machine translation centers to continuously improve the translation quality in multilingual scenarios, providing technical support for cross-language communication.

[0003] Nevertheless, existing cross-language-assisted diagnosis and treatment methods still have room for improvement. First, speech recognition methods mostly rely on joint decoding of acoustic-language models, lacking a deep perception of medical context and failing to effectively distinguish between medical terms that sound similar but have different meanings. This affects the accuracy of subsequent translation and diagnostic decisions. When dealing with multilingual patient consultations, the uncertainty in the recognition and translation process is not quantified and integrated for evaluation, making it difficult to judge the reliability of the output translation and posing a potential risk of medical complications due to mistranslation. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a cross-language barrier-free assisted diagnosis and treatment method based on speech recognition to solve the problems affecting translation and diagnostic decision-making.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] This invention provides a cross-language barrier-free assisted diagnosis and treatment method based on speech recognition, comprising:

[0008] Acquire the patient's speech signals and lip video images, and extract features to obtain acoustic features and visual lip reading features;

[0009] After temporally aligning the acoustic features and visual lip reading features, the AFM model is used for speech recognition to obtain the initial recognition text and speech recognition confidence sequence.

[0010] After constructing a medical knowledge graph and matching it with the initial identified text, the medical context confidence sequence of the initial identified text is calculated.

[0011] Based on the medical context confidence sequence, a preset medical threshold is used for comparison. The initial identified text is replaced and corrected according to the comparison results to generate high-confidence clinical text.

[0012] High-confidence clinical texts are input into a machine translation center for translation, and the translation confidence sequence of the target language translation is obtained.

[0013] A comprehensive uncertainty score is calculated based on speech recognition confidence sequences and medical context confidence sequences, combined with translation confidence sequences.

[0014] Based on the comprehensive uncertainty score, risk level ranges are divided and corresponding interactive operations are performed.

[0015] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, the step of acquiring the patient's speech signal and lip video images, and performing feature extraction to obtain acoustic features and visual lip reading features specifically includes:

[0016] The speech signal is divided into short time frames. After applying a Hamming window to each short time frame, a fast Fourier transform is performed to obtain the spectrum, and acoustic features are extracted from the spectrum.

[0017] After locating the lip region using a face detection algorithm on the lip video image, a convolutional neural network is used to extract the spatiotemporal visual features of the lip region as visual lip reading features.

[0018] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, wherein: the temporal alignment of acoustic features and visual lip reading features refers to aligning the acoustic features and visual lip reading features with timestamps, and then using a linear interpolation method to resample the visual lip reading features to the frame rate of the acoustic features.

[0019] As a preferred embodiment of the cross-language barrier-free assisted diagnosis and treatment method based on speech recognition described in this invention, the step of using an AFM model for speech recognition to obtain initial recognition text and speech recognition confidence sequence specifically involves:

[0020] The temporally aligned acoustic features and visual lip-reading features are input into the AFM model. The AFM model uses an attention mechanism to dynamically calculate the weights of the acoustic features and visual lip-reading features, and performs weighted fusion to generate a fused feature vector.

[0021] After the fused feature vectors are input into the encoder-decoder structure of the AFM model for sequence mapping, the decoder generates the vocabulary probability value for each time step.

[0022] The word with the highest probability value at each time step is selected and connected in chronological order to form the initial recognition text.

[0023] The probability values ​​of the selected words are arranged in chronological order to form a speech recognition confidence sequence.

[0024] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, wherein: the construction of the medical knowledge graph specifically includes:

[0025] Collect medical entities and establish treatment, symptom, and examination relationships between them using entity linking methods to form structured data;

[0026] Using a medical ontology building tool, structured data is converted into a triplet format to obtain a medical knowledge graph.

[0027] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, wherein: the calculation of the medical context confidence sequence of the initial recognized text specifically includes:

[0028] The semantic similarity score of each word in the initial identified text is calculated by comparing it with the medical entities in the medical knowledge graph.

[0029] Based on the order in which each word appears in the initial recognition text, the semantic similarity scores of each word are arranged to form a medical context confidence sequence.

[0030] As a preferred embodiment of the cross-language barrier-free assisted diagnosis and treatment method based on speech recognition described in this invention, the method involves: comparing a medical context confidence sequence with a preset medical threshold; replacing and correcting the initial recognized text based on the comparison results to generate high-confidence clinical text, specifically:

[0031] Based on a medical context confidence sequence, a preset medical threshold is used;

[0032] Each confidence value in the medical context confidence sequence is compared with a medical threshold. For words with confidence values ​​lower than the medical threshold, medical entities with similar pronunciation to the words are searched in the medical knowledge graph and the corresponding words in the initial recognition text are replaced. For words with confidence values ​​higher than the medical threshold, the original words in the initial recognition text are retained. The initial recognition text after replacement and correction constitutes a high-confidence clinical text.

[0033] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method described in this invention, the method involves: inputting high-confidence clinical text into a machine translation center for translation, and obtaining a translation confidence sequence of the target language translation, specifically:

[0034] High-confidence clinical text is input into the machine translation center. The encoder of the machine translation center converts the word sequence of the high-confidence clinical text into a context vector representation. The decoder generates target language words and their corresponding probability values ​​step by step based on the context vector representation.

[0035] At each time step, select the target language word with the highest probability value;

[0036] The selected target language words are arranged in chronological order to form the target language translation, and the probability values ​​of the selected target language words are arranged in chronological order to form the translation confidence sequence.

[0037] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, wherein: the comprehensive uncertainty score is calculated based on the speech recognition confidence sequence and the medical context confidence sequence, combined with the translation confidence sequence, specifically as follows:

[0038] The arithmetic mean of the speech recognition confidence sequence, the medical context confidence sequence, and the translation confidence sequence are calculated respectively to obtain the average confidence of speech recognition, the average confidence of medical context, and the average confidence of translation.

[0039] The average confidence scores of speech recognition, medical context, and translation are substituted into the weighted uncertainty fusion formula to calculate the overall uncertainty score.

[0040] As a preferred embodiment of the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method of the present invention, the step of dividing risk level intervals and performing corresponding interactive operations specifically includes:

[0041] Based on the comprehensive uncertainty score, divide the risk level into Z consecutive intervals;

[0042] The interval mapping algorithm is used to map the comprehensive uncertainty score to the corresponding risk level interval, triggering and executing the corresponding interactive operation.

[0043] The beneficial effects of this invention are as follows: By using an attention-based dynamic weighted recognition (AFM model), combined with semantic confidence assessment based on medical knowledge graphs and a threshold-driven text correction mechanism, the accuracy of speech recognition and the reliability of clinical semantics in complex environments of cross-language assisted diagnosis and treatment mechanisms are improved. It not only enhances noise resistance by utilizing multimodal information, but also achieves intelligent error correction of misidentified medical terms through medical context confidence sequences, generating high-confidence clinical text. Furthermore, it combines translation confidence to conduct comprehensive uncertainty assessment and risk grading interaction, effectively reducing the medical risks caused by speech misrecognition or translation deviations. It realizes a reliable, safe, and intelligent cross-language assisted diagnosis and treatment system with a full-link reliability from "perception-recognition-semantic verification-translation-risk feedback", which is especially suitable for precision medical communication in multilingual patient scenarios. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of a cross-language barrier-free assisted diagnosis and treatment method based on speech recognition.

[0046] Figure 2 A flowchart for obtaining the initial recognition text and speech recognition confidence sequence.

[0047] Figure 3 A flowchart for obtaining the target language translation and translation confidence sequence.

[0048] Figure 4 A flowchart for obtaining a medical knowledge graph. Detailed Implementation

[0049] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0050] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0051] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0052] Reference Figures 1-4 This is one embodiment of the present invention, which provides a cross-language barrier-free assisted diagnosis and treatment method based on speech recognition, including the following steps:

[0053] S1. Collect the patient's voice signal and lip video images, and extract features to obtain acoustic features and visual lip reading features;

[0054] S1.1 It should be noted that when patients communicate about their condition, audio and video capture devices should be used for real-time recording to obtain the patients' voice signals and lip video images.

[0055] All of the above information has been agreed to by the user and is used for legitimate purposes.

[0056] S1.2 Divide the speech signal into short time frames, apply a Hamming window to each short time frame for windowing, and then perform a fast Fourier transform to obtain the spectrum.

[0057] It should be noted that a fixed-length sliding window (e.g., a window length of 25 milliseconds and a window movement step of 10 milliseconds) is used to continuously capture segments of the real-time acquired speech signal, resulting in a short-time frame sequence. A Hamming window is applied to each short-time frame in the sequence, which involves multiplying the Hamming window function with the short-time frame point by point. A Fast Fourier Transform is then performed on each windowed short-time frame to convert it into a frequency domain representation, thus obtaining the spectrum of each short-time frame. It should be noted that setting the window length to 25 milliseconds and the window movement step to 10 milliseconds represents a perfect balance between theoretical accuracy and engineering practicality in speech signal processing, ensuring that the extracted acoustic features accurately reflect the spectral characteristics of the speech while fully capturing dynamic changes.

[0058] S1.3 Extract acoustic features from the spectrum;

[0059] It should be noted that a Mel filter bank is used to map the spectrum of each short frame onto the Mel scale. Specifically, the Mel filter bank consists of a set of triangular bandpass filters. The energy value of each frequency point in the spectrum is multiplied by the weighting coefficient of the corresponding triangular bandpass filter, and then summed to obtain the energy value output by each Mel filter. All triangular bandpass filters in the Mel filter bank sequentially perform the above weighted summation calculation, thus completing the mapping.

[0060] After mapping the spectrum of each short frame onto the Mel scale, the logarithmic energy of the output of each Mel filter is calculated, and all logarithmic energies are sequentially subjected to discrete cosine transform to obtain the cepstral coefficient sequence.

[0061] The discrete cosine transform of logarithmic energy can be expressed by the following formula:

[0062] ;

[0063] in, Indicates the first cepstral coefficients, This indicates the total number of filters in the Mel filter bank. Indicates the index of the Mel filter. This represents the index of the cepstral coefficients to be determined. Represents pi;

[0064] The first 12 cepstral coefficients in the cepstral coefficient sequence are extracted as Mel frequency cepstral coefficients, which constitute the acoustic features characterizing the short-time spectral characteristics of the speech signal.

[0065] It should also be noted that the extraction of the first 12 cepstral coefficients as Mel frequency cepstral coefficients is based on the physical characteristics of acoustic features. When the number of extracted cepstral coefficients is less than 12, it will lead to the loss of consonant detail features, and when it is greater than 12, it is susceptible to environmental noise and channel distortion.

[0066] S1.4 After locating the lip region using a face detection algorithm on the lip video image, a convolutional neural network is used to extract the spatiotemporal visual features of the lip region as visual lip reading features.

[0067] Images of the lip region from historical patients were collected as training and validation sets. The training set was input into a convolutional neural network (CNN). After the CNN outputs the prediction results of visual lip reading features through forward propagation, the gradient of the loss function with respect to the CNN parameters was calculated using the backpropagation algorithm. The weight parameters of the CNN were then updated using the adaptive moment estimation algorithm. The forward propagation, loss calculation, backpropagation, and parameter update processes were repeated until the recognition accuracy of the CNN on the validation set no longer improved with increasing training epochs, reaching a stable state. This yielded the trained CNN.

[0068] Real-time video images of the patient's lips are input into a face detection algorithm to locate the coordinates of the lip region. Based on these coordinates, the lip region image is cropped from the video images. The cropped lip region image is then input into a trained convolutional neural network (CNN). The CNN extracts the spatial features of the lip region image through multiple 2D convolutional and pooling layers. The spatial features of multiple consecutive frames of the lip region image are combined along the time dimension, and then the spatiotemporal visual features of the lip region image are extracted through a 3D convolutional layer. The output layer outputs a feature vector, which is the visual lip reading feature.

[0069] S2. After temporally aligning the acoustic features and visual lip reading features, use the AFM model to perform speech recognition and obtain the initial recognition text and speech recognition confidence sequence.

[0070] S2.1 Input the temporally aligned acoustic features and visual lip-reading features into the AFM model. The AFM model uses an attention mechanism to dynamically calculate the weights of the acoustic features and visual lip-reading features, and performs weighted fusion to generate a fused feature vector.

[0071] It should be noted that after aligning the acoustic features and visual lip reading features with timestamps, a linear interpolation method is used to resample the visual lip reading features to the frame rate of the acoustic features, thereby obtaining temporally aligned acoustic features and visual lip reading features.

[0072] Acoustic features, visual lip reading features, and corresponding speech transcription text annotations aligned to historical time sequences were collected as training and validation sets. The training set is input into the AFM model (Attention Fusion Mechanism model). The AFM model dynamically calculates the weights of historical acoustic features and visual lip-reading features through an attention mechanism and generates a fused feature vector. The AFM model then uses a recurrent neural network to convert the fused feature vector sequence into a hidden state sequence. The hidden state sequence is input into the decoder of the AFM model for word prediction processing to obtain the predicted text sequence. The loss value between the predicted text sequence and the speech transcription text annotation is calculated using a connectionist temporal classification loss function. The gradient of the loss value with respect to the AFM model parameters is calculated using the backpropagation algorithm. The attention weight matrix and encoder-decoder parameters of the AFM model are updated using an adaptive moment estimation algorithm. Specifically, the adaptive moment estimation algorithm calculates the first-order moment estimate and the second-order moment estimate of the AFM model based on the gradient of the loss function of the AFM model. At the same time, the adaptive moment estimation algorithm corrects the bias of the first-order moment estimate and the second-order moment estimate, and uses the square root of the corrected first-order moment estimate and the corrected second-order moment estimate to obtain the update amount of the attention weight matrix and encoder-decoder parameters of the AFM model. The adaptive moment estimation algorithm applies the update to the attention weight matrix and the current values ​​of the encoder-decoder parameters of the AFM model to complete the parameter update. The forward propagation, loss calculation, backpropagation, and parameter update processes are repeated until the word error rate of the AFM model converges on the validation set, resulting in the trained AFM model.

[0073] The temporally aligned acoustic features and visual lip-reading features are input into the trained AFM model. The AFM model calculates the attention weights for the acoustic features and the visual lip-reading features through an attention computation layer. The calculation of the attention weights is based on the contextual relevance of the acoustic features and the visual lip-reading features, which is obtained through a query-key-value attention mechanism. The AFM model uses the calculated attention weights to perform a weighted summation of the acoustic features and the visual lip-reading features, and then concatenates the weighted summation of the acoustic features and the visual lip-reading features into a fused feature vector.

[0074] S2.2 After inputting the fused feature vector into the encoder-decoder structure of the AFM model for sequence mapping, the decoder generates the word probability value for each time step.

[0075] It should be noted that the fused feature vector is input into the encoder of the AFM model. The encoder performs forward computation from the start to the end of the fused feature vector through the forward Long Short-Term Memory (LSTM) network to obtain the hidden states of the forward LSM network. The encoder also performs reverse computation from the end to the start of the fused feature vector through the backward LSM network to obtain the hidden states of the backward LSM network. The hidden states of the forward and backward LSM networks at each time step are concatenated to form the hidden states of a multi-layer bidirectional LSM network. The hidden states of the multi-layer bidirectional LSM network at each time step are arranged in chronological order to form a hidden state sequence.

[0076] After inputting the hidden state sequence into the decoder of the AFM model, the relevant parts of the hidden state sequence are dynamically focused. Combined with the decoder's output from the previous time step, the lexical probability distribution for the current time step is generated using the Softmax function. The lexical probability distribution is a vector with the same dimension as the vocabulary size, where each element represents the probability value of the corresponding word.

[0077] The word with the highest probability value at each time step is selected and connected in chronological order to form the initial recognition text. At the same time, the probability values ​​of the selected words are arranged in chronological order to form a speech recognition confidence sequence.

[0078] S3. Construct a medical knowledge graph and, after matching it with the initial identified text, calculate the medical context confidence sequence of the initial identified text;

[0079] S3.1 It should be noted that medical entities (including disease names, clinical symptoms, drug names, examination items, and body parts) are collected from authoritative medical textbooks, clinical guidelines, and medical dictionaries. Treatment relationships, symptom relationships, and examination relationships between medical entities are established through entity linking methods to form structured data (in tabular form). The tabular structured data is then converted into a topic-verb-object triple format using a medical ontology construction tool to obtain a medical knowledge graph.

[0080] S3.2 Calculate the semantic similarity between each word in the initial identified text and the medical entities in the medical knowledge graph, and obtain the semantic similarity score for each word;

[0081] It should be noted that the initial identified text is segmented to obtain a word sequence. A pre-trained word embedding model (such as the Word2Vec model) is used to map each word in the word sequence to a low-dimensional, dense real-valued vector (i.e., a word vector). This process maps semantically similar words to nearby positions in the vector space.

[0082] Query all medical entities in the medical knowledge graph and obtain the word vectors of each medical entity. Calculate the semantic similarity score between each word vector in the word sequence and the word vector of each medical entity in the medical knowledge graph using the cosine similarity algorithm, expressed by the formula:

[0083] ;

[0084] in, The semantic similarity score represents the number of words. Words Word vectors, Represents medical entities Word vectors;

[0085] Based on the order in which each word appears in the initial recognition text, the semantic similarity scores corresponding to each word are arranged in the order of the word sequence to form the medical context confidence sequence of the initial recognition text.

[0086] S4. Based on the medical context confidence sequence, a preset medical threshold is used for comparison. The initial recognized text is replaced and corrected according to the comparison results to generate high-confidence clinical text.

[0087] S4.1 It should be noted that a preset medical threshold is used based on the medical context confidence sequence;

[0088] Each confidence value in the medical context confidence sequence is compared with a medical threshold. For words in the medical context confidence sequence with confidence values ​​lower than the medical threshold, the medical entity with the highest pronunciation similarity to the word is searched in the medical knowledge graph. The found medical entity is used to replace the corresponding word in the initial recognition text. For words in the medical context confidence sequence with confidence values ​​higher than the medical threshold, the original words in the initial recognition text are retained. The word sequence after replacement and correction constitutes a high-confidence clinical text.

[0089] It should be noted that the medical threshold ranges from 0.5 to 0.7, chosen as a balance between clinical risk tolerance and terminology recognition accuracy. A medical threshold below 0.5 results in a large number of incorrect terms being retained, posing a significant medical safety hazard; a medical threshold above 0.7 results in a large number of correct terms being incorrectly corrected, disrupting the consultation process.

[0090] S5. Input the high-confidence clinical text into the machine translation center for translation and obtain the translation confidence sequence of the target language translation;

[0091] S5.1 It should be noted that the high-confidence clinical text is input into the machine translation center. The encoder of the machine translation center sequentially reads the embedding vector of each word in the high-confidence clinical text through a recurrent neural network (RNN). The RNN transmits contextual information word by word through hidden states. The hidden state update at each time step depends on the embedding vector of the current word and the hidden state of the previous time step. After the RNN processes the last word of the high-confidence clinical text, the hidden state of the final time step serves as the context vector for the entire high-confidence clinical text. The decoder of the machine translation center takes the context vector as input at the initial time step and generates the first hidden state through the RNN. The decoder of the machine translation center calculates attention weights based on the hidden state of the current time step. These attention weights are used to weight and combine the information in the context vector. The decoder of the machine translation center concatenates the attention-weighted context vector with the hidden state of the current time step and inputs it into a fully connected layer. The fully connected layer performs a linear transformation using a weight matrix and a bias vector to generate a linear output vector. The linear output vector is input into the Softmax function. The Softmax function performs an exponential operation on each element of the linear output vector and normalizes the result so that the sum of all elements equals 1, obtaining the probability value of the target language vocabulary. At each time step, the target language vocabulary with the highest probability value is selected. The selected target language vocabulary is arranged in chronological order to form the target language translation, and the probability values ​​of the selected target language vocabulary are arranged in chronological order to form a translation confidence sequence.

[0092] S6. Calculate the comprehensive uncertainty score based on the speech recognition confidence sequence and the medical context confidence sequence, combined with the translation confidence sequence;

[0093] S6.1 It should be noted that the arithmetic mean of the speech recognition confidence sequence, the medical context confidence sequence, and the translation confidence sequence should be calculated separately to obtain the average confidence of speech recognition, the average confidence of medical context, and the average confidence of translation; the average confidence of speech recognition is expressed by the formula:

[0094] ;

[0095] Where S represents the average confidence level of speech recognition. This represents the total number of elements in the speech recognition confidence sequence. This represents the index variable of an element in the speech recognition confidence sequence. Indicates the first [number] in the speech recognition confidence sequence. The specific numerical value of each element;

[0096] It should also be noted that the formula for the arithmetic mean of the speech recognition confidence sequence, the medical context confidence sequence, and the translation confidence sequence is the same. Specifically, when When representing the element values ​​in the speech recognition confidence sequence, the result is the average confidence score of speech recognition; when... When representing element values ​​in a medical context confidence sequence, the result is the average confidence score of the medical context; when... When representing the element values ​​in the translation confidence sequence, the result is the average translation confidence.

[0097] The average confidence scores of speech recognition, medical context, and translation are substituted into the weighted uncertainty fusion formula to calculate the comprehensive uncertainty score, which is expressed as follows:

[0098] ;

[0099] in, This represents the overall uncertainty score. This represents the average confidence level of speech recognition. This represents the average confidence level in the medical context. This represents the average confidence level of the translation. The weighting coefficients represent the average confidence level of speech recognition. The weighting coefficients represent the average confidence level in the medical context. Weighting coefficients representing the average confidence level of the translation;

[0100] It should be noted that the weighting coefficients for the average confidence level of speech recognition (0.2), translation (0.3), and medical context (0.5) are obtained through objective quantification using the FMEA tool, based on the principles of clinical risk level quantitative assessment.

[0101] S7. Based on the comprehensive uncertainty score, divide the risk level range and perform the corresponding interactive operations;

[0102] S7.1. Based on the comprehensive uncertainty score, divide the risk level into Z consecutive intervals;

[0103] It should be noted that, based on the comprehensive uncertainty score, three consecutive risk level intervals are defined. The first risk level interval is [0, 0.3), which is defined as low risk; the second risk level interval is [0.3, 0.7), which is defined as medium risk; and the third risk level interval is [0.7, 1], which is defined as high risk.

[0104] It should also be noted that the risk level ranges are determined based on historical data statistics. The first risk level range is set to [0, 0.3), the second risk level range is set to [0.3, 0.7), and the third risk level range is set to [0.7, 1]. This is because 0.3 is the minimum threshold to ensure a safety baseline; below 0.3, the risk is out of control, and above 0.3, efficiency decreases unnecessarily. 0.7 is the safety red line that must be adhered to to avoid serious medical errors; below 0.7, safety is out of control, and above 0.7, availability decreases sharply.

[0105] S7.2 Use the interval mapping algorithm to map the comprehensive uncertainty score to the corresponding risk level interval, trigger and execute the corresponding interactive operation;

[0106] Furthermore, an interval mapping algorithm is used to compare the overall uncertainty score with the boundary values ​​of the low-risk interval [0, 0.3), the medium-risk interval [0.3, 0.7), and the high-risk interval [0.7, 1]. When the overall uncertainty score falls within the low-risk interval, the corresponding interactive operation is triggered and executed, meaning the target language translation is directly output to the display interface. When the overall uncertainty score falls within the medium-risk interval, the corresponding interactive operation is triggered and executed, meaning the target language translation is output to the display interface, and words in the translation confidence sequence with values ​​below the confidence threshold (e.g., 0.8) are highlighted on the display interface. When the overall uncertainty score falls within the high-risk interval, the corresponding interactive operation is triggered and executed, meaning that for words in the translation confidence sequence with values ​​below the confidence threshold, a corresponding clarification request statement is generated. This clarification request statement is played to the patient via a voice playback device, and the patient's voice feedback to the clarification request statement is re-acquired and input into the AFM model for a new round of speech recognition and subsequent processing. It should be noted that the confidence threshold of 0.8 is set based on clinical safety and operational efficiency. A confidence threshold higher than 0.8 may lead to excessive conservatism and fragmentation of the consultation process; a confidence threshold lower than 0.8 may lead to insufficient risk control and overlooking medical terminology errors.

[0107] This embodiment also provides a computer device applicable to the cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the cross-language barrier-free auxiliary diagnosis and treatment method based on speech recognition as proposed in the above embodiment.

[0108] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0109] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the speech recognition-based cross-language barrier-free assisted diagnosis and treatment method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0110] In summary, this invention improves the accuracy of speech recognition and the reliability of clinical semantics in complex environments through: dynamic weighted recognition based on attention mechanism (AFM model), semantic confidence assessment combined with medical knowledge graph, and threshold-driven text correction mechanism. It not only enhances noise resistance by utilizing multimodal information, but also achieves intelligent error correction of misidentified medical terms through medical context confidence sequences, generating high-confidence clinical text. Furthermore, it combines translation confidence for comprehensive uncertainty assessment and risk grading interaction, effectively reducing medical risks caused by speech misrecognition or translation deviation. It realizes a reliable, safe, and intelligent cross-language diagnostic assistance system from "perception-recognition-semantic verification-translation-risk feedback", which is especially suitable for precision medical communication in multilingual patient scenarios.

[0111] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A cross-language barrier-free assisted diagnosis and treatment method based on speech recognition, characterized in that: include, Acquire the patient's speech signals and lip video images, and extract features to obtain acoustic features and visual lip reading features; After temporally aligning the acoustic features and visual lip-reading features, an AFM model is used for speech recognition to obtain the initial recognition text and speech recognition confidence sequence, specifically: The temporally aligned acoustic features and visual lip-reading features are input into the AFM model. The AFM model uses an attention mechanism to dynamically calculate the weights of the acoustic features and visual lip-reading features, and performs weighted fusion to generate a fused feature vector. After the fused feature vectors are input into the encoder-decoder structure of the AFM model for sequence mapping, the decoder generates the vocabulary probability value for each time step. The word with the highest probability value at each time step is selected and connected in chronological order to form the initial recognition text. The probability values ​​of the selected words are arranged in chronological order to form a speech recognition confidence sequence; After constructing a medical knowledge graph and matching it with the initial identified text, the medical context confidence sequence of the initial identified text is calculated. Based on the medical context confidence sequence, a preset medical threshold is used for comparison. The initial identified text is replaced and corrected according to the comparison results to generate high-confidence clinical text. High-confidence clinical texts are input into a machine translation center for translation, and the translation confidence sequence of the target language translation is obtained. A comprehensive uncertainty score is calculated based on speech recognition confidence sequences and medical context confidence sequences, combined with translation confidence sequences. Based on the comprehensive uncertainty score, risk level ranges are divided and corresponding interactive operations are performed.

2. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 1, characterized in that: The process involves acquiring the patient's voice signals and lip video images, and performing feature extraction to obtain acoustic and visual lip-reading features. The speech signal is divided into short time frames. After applying a Hamming window to each short time frame, a fast Fourier transform is performed to obtain the spectrum, and acoustic features are extracted from the spectrum. After locating the lip region using a face detection algorithm on the lip video image, a convolutional neural network is used to extract the spatiotemporal visual features of the lip region as visual lip reading features.

3. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 2, characterized in that: The temporal alignment of acoustic features and visual lip reading features refers to aligning the acoustic features and visual lip reading features with timestamps, and then using a linear interpolation method to resample the visual lip reading features to the frame rate of the acoustic features.

4. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 3, characterized in that: The construction of the medical knowledge graph specifically involves: Collect medical entities and establish treatment, symptom, and examination relationships between them using entity linking methods to form structured data; Using a medical ontology building tool, structured data is converted into a triplet format to obtain a medical knowledge graph.

5. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 4, characterized in that: The calculation of the medical context confidence sequence of the initial recognized text is specifically as follows: The semantic similarity score of each word in the initial identified text is calculated by comparing it with the medical entities in the medical knowledge graph. Based on the order in which each word appears in the initial recognition text, the semantic similarity scores of each word are arranged to form a medical context confidence sequence.

6. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 5, characterized in that: Based on a medical context confidence sequence, a preset medical threshold is used for comparison. The initial identified text is then replaced and corrected according to the comparison results to generate high-confidence clinical text. Specifically: Based on a medical context confidence sequence, a preset medical threshold is used; Each confidence value in the medical context confidence sequence is compared with a medical threshold. For words with confidence values ​​lower than the medical threshold, medical entities with similar pronunciation to the words are searched in the medical knowledge graph and the corresponding words in the initial recognition text are replaced. For words with confidence values ​​higher than the medical threshold, the original words in the initial recognition text are retained. The initial recognition text after replacement and correction constitutes a high-confidence clinical text.

7. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 6, characterized in that: High-confidence clinical texts are input into a machine translation center for translation, and the translation confidence sequence of the target language translation is obtained, specifically: High-confidence clinical text is input into the machine translation center. The encoder of the machine translation center converts the word sequence of the high-confidence clinical text into a context vector representation. The decoder generates target language words and their corresponding probability values ​​step by step based on the context vector representation. At each time step, select the target language word with the highest probability value; The selected target language words are arranged in chronological order to form the target language translation, and the probability values ​​of the selected target language words are arranged in chronological order to form the translation confidence sequence.

8. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 7, characterized in that: The comprehensive uncertainty score is calculated based on the speech recognition confidence sequence, the medical context confidence sequence, and the translation confidence sequence, specifically as follows: The arithmetic mean of the speech recognition confidence sequence, the medical context confidence sequence, and the translation confidence sequence are calculated respectively to obtain the average confidence of speech recognition, the average confidence of medical context, and the average confidence of translation. The average confidence scores of speech recognition, medical context, and translation are substituted into the weighted uncertainty fusion formula to calculate the overall uncertainty score.

9. The cross-language barrier-free assisted diagnosis and treatment method based on speech recognition as described in claim 8, characterized in that: The process of dividing risk level ranges and executing corresponding interactive operations is as follows: Based on the comprehensive uncertainty score, divide the risk level into Z consecutive intervals; The interval mapping algorithm is used to map the comprehensive uncertainty score to the corresponding risk level interval, triggering and executing the corresponding interactive operation.

Citation Information

Patent Citations

  • Lip language recognition method for patient with language disorder in hospital environment

    CN116959060A

  • Power grid dispatching method and system based on knowledge graph

    CN117012185A