Virtual doctor diagnosis auxiliary system integrating visual diffusion and voice emotion recognition

By integrating visual diffusion and speech emotion recognition technology and integrating medical images and voice data, the problems of single mode and emotional deficiency in the virtual doctor diagnosis system are solved, and more efficient and accurate remote diagnosis is achieved.

CN120496812APending Publication Date: 2025-08-15ZHEJIANG YISHAN SMART MEDICAL RES CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510693248.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing virtual doctor diagnosis system relies on single modal data, making it difficult to fully capture patient condition information, lack of emotional judgment, and insufficient image fusion, resulting in high misdiagnosis rate and low diagnostic efficiency.

Method used

Using fusion visual diffusion and speech emotion recognition technology, medical image preprocessing, feature extraction, multimodal fusion, speech data acquisition and comprehensive diagnostic decision-making modules are used to integrate medical images and speech emotion characteristics, and comprehensive analysis is performed using deep learning algorithms.

Benefits of technology

It improves the accuracy and efficiency of telemedicine diagnosis, can better capture subtle changes and patient emotions in complex cases, and provides more accurate diagnostic suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496812A_ABST
    Figure CN120496812A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual doctor diagnosis auxiliary system fusing visual diffusion and voice emotion recognition. The virtual doctor diagnosis auxiliary system comprises a medical image preprocessing module, a medical image feature extraction module, a multi-modal fusion module, a medical voice data acquisition module, an emotion state feature extraction module and a comprehensive diagnosis decision module. According to the invention, a visual diffusion and speech emotion recognition fusion technology is adopted, the disease of the current patient can be more accurately judged according to the medical image and the speech of the patient, and the problems of low detection efficiency, high cost, poor accuracy, strong dependence on personnel and the like in a traditional method are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical information computer processing technology, relates to virtual doctor diagnosis assistance technology, and specifically relates to a virtual doctor diagnosis assistance system that integrates visual diffusion and voice emotion recognition. Background Art

[0002] With advances in medical technology and the growth of the internet, telemedicine services have become an integral part of the modern healthcare system. The importance of virtual doctor diagnosis systems has become increasingly prominent, especially during public health emergencies. These systems allow patients to obtain preliminary diagnostic recommendations by uploading medical images (such as X-rays, CT scans, or MRIs) and engaging in voice conversations. This is particularly beneficial for patients living in remote areas or with limited mobility. They can also improve the efficiency of daily health management and chronic disease monitoring. While existing virtual doctor diagnosis systems have improved the accessibility and efficiency of healthcare services to a certain extent, they still have some significant limitations: 1) Single-modality limitations: Existing systems primarily rely on a single data type, such as text-based question-answering or simple medical image analysis. This single-modality approach struggles to fully capture a patient's condition; in many cases, a more accurate diagnosis requires the integration of multiple data sources. 2) Lack of emotional assessment: Traditional speech processing techniques focus on understanding content. However, in clinical settings, patients' emotional expressions (such as anxiety and pain) can provide doctors with important diagnostic clues. Failing to consider emotional factors can lead to misdiagnosis or missed diagnoses. Simultaneously modeling speech and image features can provide better disease assessment. 3) Image fusion challenge: When processing multimodal image data, existing methods often fail to effectively integrate information from different imaging technologies, resulting in insufficient feature extraction and affecting the final diagnostic accuracy. Summary of the Invention

[0003] This paper addresses the shortcomings of existing virtual doctor diagnosis systems, such as low flexibility, high misdiagnosis rates, and poor interactive behavior recognition accuracy. By integrating visual diffusion models with speech emotion recognition, this approach aims to improve the accuracy and efficiency of remote medical diagnosis.

[0004] To achieve the above object, the present invention adopts the following technical solutions: The virtual doctor diagnosis assistance system, which integrates visual diffusion and speech emotion recognition, includes the following modules: Medical image preprocessing module: used to improve the quality and standardize medical image data from various sources.

[0005] Medical image feature extraction module: Build an image matching model guided by key point attention mechanism to process medical images of different modalities.

[0006] Multimodal fusion module: An encoder module guided by a keypoint attention mechanism is used to achieve multimodal medical image fusion.

[0007] Medical voice data collection module: collects various sounds of patients in different health conditions to construct a medical voice dataset for emotional state feature extraction training.

[0008] Emotional state feature extraction module: Emotional state feature extraction is performed based on the improved WavLM technology.

[0009] Comprehensive diagnosis decision module: Based on multimodal medical image feature representation and speech feature representation, it outputs possible disease classification results.

[0010] In one possible implementation, the medical image preprocessing module first applies a histogram equalization algorithm to redistribute pixel values and enhance image contrast. Next, the image is converted from the spatial domain to the frequency domain using a Fourier transform, and filters are applied to remove noise and reduce blur. Ultimately, a high-quality, standardized image is output.

[0011] In one possible implementation, the medical image feature extraction module first applies the Haralick algorithm to extract texture features from medical images of different modalities. It then generates keypoint maps corresponding to the medical images of different modalities, capturing the connections between corresponding regions across the images and ensuring that keypoints in the same region are annotated across medical images of different modalities.

[0012] In one possible implementation, the encoder module guided by the keypoint attention mechanism is as follows: The keypoint map extracted by the medical image feature extraction module guides the construction of the keypoint attention mechanism layer in the encoder module. This layer uses spatial information to enhance and focus the propagation and fusion of locally important information. Next, the Haralick algorithm is used to extract keypoint map features, which are further processed by the keypoint attention mechanism layer to generate a set of attention features. Finally, a multi-layer neural network is used to further fuse features from different modal images, integrating their attention features to output the final fused multimodal medical image feature representation.

[0013] In one possible implementation, the medical voice data acquisition module uses a high-quality microphone to collect voice data, systematically recording various voice characteristics of patients under different health conditions and assigning a unique numerical category code to each voice characteristic. The continuous voice signal is then segmented into short time frames, and normalization techniques are applied to adjust the volume level of each voice sample to ensure similar energy levels.

[0014] In a possible implementation, the medical speech data acquisition module further includes frequency domain noise reduction using a combination of frame processing, Hanning window function, and fast Fourier transform (FFT).

[0015] In a possible implementation, the improved WavLM technology is specifically as follows: The WavLM model is a binary classification task. To transform it into a multi-classification task, the number of nodes in the final layer's output is changed to the number of sound categories. To ensure that the original sound waveform is not lost, a waveform feature conversion layer is added. The output of the conversion layer is concatenated with the results of the feature extraction layer. The concatenated speech features are used as input to the fully connected prediction layer, which outputs the final emotional state features.

[0016] In one possible implementation, the comprehensive diagnosis and decision module utilizes a multimodal fusion module to extract and fuse features from multimodal medical images uploaded by the patient to generate a multimodal medical image feature representation. Next, the emotional state feature extraction module extracts emotional state and other vocal characteristics from the patient's spoken conversation, retaining the original waveform information. These features are then concatenated to form a speech feature representation. The combined multimodal medical image and speech feature representations are then input into a deep neural network (DNN) model, which learns and integrates information from these two heterogeneous data sources to output a possible disease classification result.

[0017] Through the above technical means, the present invention can effectively improve the efficiency and accuracy of virtual doctors in diagnosing patients' conditions in remote online medical diagnosis.

[0018] The beneficial effects of the present invention are as follows: To address the above issues, this paper proposes a virtual doctor diagnostic assistance system that integrates vision diffusion and speech emotion recognition (SER). This system offers the following significant advantages: 1) Deep multimodal integration of medical images: The introduction of a visual diffusion model more effectively extracts and fuses key features from diverse medical images, ensuring that subtle changes in complex cases can be captured. Furthermore, the speech emotion recognition module analyzes the patient's emotional fluctuations during speech, providing an additional dimension for diagnosis and improving diagnostic accuracy. 2) Enhanced recognition of patient emotional state: The improved WavLM technology allows for in-depth exploration of emotional information in speech, including but not limited to emotions such as tension, anger, and sadness. These emotional indicators can help doctors better understand the patient's psychological state, enabling more personalized judgments. Furthermore, emotional information can be integrated with multimodal image features as an additional auxiliary feature for comprehensive diagnosis and treatment. 3) Improved diagnostic efficiency and accuracy: By organically combining image and speech features and utilizing advanced deep learning algorithms for comprehensive analysis, more accurate diagnostic recommendations can be provided in a shorter timeframe. This approach is particularly advantageous for diseases with similar symptoms but different causes.

[0019] The system not only overcomes some key shortcomings of existing technologies but also opens up new paths for future intelligent, personalized medical services. By fully leveraging the value of multimodal medical data and voice data during interactions, it provides patients with more accurate and efficient diagnostic support and promotes the rational allocation of medical resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0021] Figure 1 This is a schematic diagram of the virtual doctor diagnosis assistance process that integrates visual diffusion and voice emotion recognition in an embodiment of the present invention.

[0022] Figure 2 Schematic diagram of multimodal medical image fusion based on key graph attention mechanism in an embodiment of the present invention.

[0023] Figure 3 Schematic diagram of the improved WavLM technology according to an embodiment of the present invention.

[0024] Figure 4 Schematic diagram of integrating heterogeneous data sources into a DNN model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0026] The purpose of this invention is to improve the accuracy of patient disease recognition during telemedicine. By integrating visual diffusion with speech emotion recognition, this technology can more accurately determine the patient's current disease based on medical images and patient speech, avoiding the problems of traditional methods such as low detection efficiency, high cost, poor accuracy, and strong dependence on personnel.

[0027] The present application also provides a method for implementing a virtual doctor diagnosis assistance system that integrates visual diffusion and speech emotion recognition, which specifically includes the following steps: Step (1): Improve the quality of medical image data from various sources and achieve standardization through medical image preprocessing modules.

[0028] A histogram equalization algorithm is applied to redistribute pixel values and enhance image contrast, thereby highlighting key features of the captured area. Next, a Fourier transform is used to convert the image from the spatial domain to the frequency domain, and filters are applied to this domain to remove noise and reduce blur, ensuring image clarity and accuracy. Step (2): Process medical images of different modalities through the medical image feature extraction module.

[0029] The medical image feature extraction module applies the Haralick algorithm to extract texture features from medical images of different modalities. Then, based on clinical needs, it labels key points of important anatomical structures in various medical images and records the (x, y) coordinates of these key points within the image plane, ensuring that key points in the same region are labeled across medical images of different modalities. Next, using this key point labeling information, it generates key point maps corresponding to medical images of different modalities, capturing the connections between corresponding regions across images. Step (3): Perform multimodal medical image fusion through the multimodal fusion module.

[0030] The multimodal fusion module uses an encoder module guided by a keypoint attention mechanism to achieve multimodal medical image fusion. The keypoint attention mechanism layer in the encoder module is constructed based on the keypoint map extracted by the medical image feature extraction module. The keypoint map features are then extracted using the Haralick algorithm. These features are further processed by the keypoint attention mechanism layer, emphasizing the information in the keypoint regions and generating a set of attention features. Finally, a multi-layer neural network is used to further fuse features from images of different modalities, integrating their attention features to output the final fused multimodal medical image feature representation. Step (4): The medical speech data collection module is used to collect various sounds of patients in different health states to construct a medical speech dataset for emotional state feature extraction training.

[0031] High-quality recording microphones were selected to systematically record the various voice characteristics of patients in different health states, and each voice characteristic was assigned a unique numerical category code. The continuous speech signal was then segmented into short time frames of 20 to 30 milliseconds. Finally, normalization techniques were applied to adjust the volume level of each speech sample. To further improve the quality of the audio signal, a combination of frame processing, Hanning window function, and fast Fourier transform was used for frequency domain noise reduction. Step (5): Extract emotional state features through the emotional state feature extraction module.

[0032] Emotional state feature extraction is performed based on the improved WavLM technology. The WavLM model is a binary classification task. To transform it into a multi-classification task, the number of nodes output by the last layer is changed to the number of sound categories. At the same time, to ensure that the original sound waveform is not lost, an original waveform feature conversion layer is added. The output of the conversion layer is spliced with the result of the feature extraction layer. The spliced speech features are input into the fully connected prediction layer to output the final emotional state features. Step (6): Through the comprehensive diagnosis decision module, based on the multimodal medical image feature representation and speech feature representation, the possible disease classification results are output.

[0033] First, the multimodal fusion module is used to extract and fuse the features of the multimodal medical images uploaded by the patient to obtain a multimodal medical image feature representation. Secondly, in terms of speech processing, the emotional state feature extraction module is used to extract the emotional state and other sound characteristics from the patient's voice conversation, while retaining the original waveform information. These features are then spliced together as a speech feature representation. Then, the multimodal medical image feature representation and the speech feature representation are combined and input into the deep neural network (DNN) model. Finally, the DNN model learns and integrates the information of these two heterogeneous data sources, and outputs possible disease classification results, thereby supporting doctors to make more accurate diagnostic decisions in long-distance situations.

[0034] In one possible implementation, the medical image preprocessing module receives and standardizes medical image data from various sources (such as MRI, CT, PET, US, etc.) and improves its quality by following the steps below: Use histogram equalization algorithm to enhance image contrast: For each input medical image , calculate the number of pixels at each gray level , where i represents the grayscale value, according to the histogram , calculate the cumulative distribution function , the calculation formula is: Afterwards Normalized to the range of [0, L-1], L is the maximum grayscale level of the image, the formula is as follows: Where M and N represent the height and width of the image. The enhanced image is obtained through the above steps. .

[0035] Furthermore, the image is subjected to Fourier transform to remove noise and reduce blur. The enhanced image is transformed into Converting from the spatial domain to the frequency domain is done as follows: in is the representation in the frequency domain. In the frequency domain, the value of the noise signal is usually large. Since the above histogram equalization algorithm step amplifies the noise signal for easy identification, a low-pass filter is selected to process the frequency domain image. The low-pass filter transfer function is defined as follows: in represents the distance from the point (u,v) to the center, is a custom cutoff frequency. The denoised image is obtained by processing with a low-pass filter. .

[0036] In one possible implementation, the medical image feature extraction module constructs an image matching model guided by a keypoint attention mechanism to process medical images of different modalities (such as MRI, CT, PET, etc.) by following the steps below: First, the Haralick algorithm is used to extract texture features from medical images of different modalities. For each input medical image , calculate its gray level co-occurrence matrix (GLCM) at a specific direction and distance , where d represents the distance between pixel pairs, is the direction. A set of contrast features is further extracted as texture features, and the extraction method is as follows: Where i and j represent different gray levels, and P(i, j) is an element in the gray-level co-occurrence matrix.

[0037] Next, a keypoint map associated with each medical image is generated to ensure that keypoints in the same region are annotated across different modalities. Based on their medical knowledge, professional physicians annotate keypoints at important anatomical structures in various medical images (such as joint centers in skeletal images or specific points on the boundaries of cardiac images) and record their (x, y) coordinates within the image plane. A binary image is created for each medical image as the keypoint map, where only pixels at keypoint locations have a value of 1, and all other pixels have a value of 0.

[0038] In one possible implementation, Figure 2 As shown in the figure, in order to realize multimodal medical image fusion based on key point attention mechanism, the following steps are performed: Introduction of positional spatial information: For each input medical image , the corresponding key point map is obtained through the medical image feature extraction module , where the keypoint map is a binary map and the locations of keypoints are marked as significant. In order to enhance and focus on local important information, a keypoint attention mechanism layer is added to the encoder module, which accepts feature maps from different modalities. and the corresponding keypoint graph .

[0039] The workflow for setting up the keypoint attention mechanism layer is as follows: First, the keypoint graph we accept Convert to feature map Tensors of the same size are used to apply element-by-element to the feature map. Next, the key point map is used to adjust the weights in the feature map to emphasize the information in the key point area. The attention feature map is calculated by the following formula : in It is a learnable parameter that controls the influence of key points on features; Represents element-by-element multiplication. After the above processing, a set of feature maps adjusted by the key point force mechanism are obtained. , these feature maps focus more on the key point areas in the image. , spliced together to form multimodal medical image features : In one possible implementation, the medical voice data acquisition module operates as follows: (1) Audio Collection: Noise-Cancelling Microphones are preferred to reduce background noise. Unidirectional microphones are also used. In medical environments, unidirectional microphones primarily pick up sounds in front of you, effectively reducing ambient noise. Maintain an appropriate distance between the microphone and the sound source (patient), typically between 15-30 cm, and collect audio data from various patients in real offline consultation environments.

[0040] Audio labeling: Assign a numeric category code to each sound feature, such as wheezing: 1, wheezing: 2, shortness of breath: 3, coughing: 4, normal: 5, sneezing: 6, hoarseness: 7.

[0041] (2) Speech signal processing: First, perform short time frame segmentation: divide the continuous speech signal s(t) into short time frames , each frame length ranges from 20 to 30 milliseconds. Segmentation is performed using overlapping Hanning windows, with the following formula: Where T is the frame shift and w(t) is the window function. In order to make all speech samples have similar energy levels, root mean square (RMS) normalization is applied. First, the RMS value of each speech frame is calculated: Where Q is the number of speech frames.

[0042] Then, adjust the volume by the following formula: In order to further improve the quality of audio signals, a combination of Hanning window function, fast Fourier transform (FFT), frequency domain filtering and noise reduction is adopted.

[0043] First, each frame after adjusting the volume Apply the Hanning window function, the formula is as follows: Where U represents the number of divided frames, and then the Hanning window is multiplied point by point with each frame of data to obtain the windowed time domain signal. : Perform fast Fourier transform on the windowed time domain signal: A low-pass filter H(k) is then applied to the frequency domain signal to filter and reduce noise in the frequency domain. A medical speech dataset is constructed based on the processed data.

[0044] In one possible implementation, the emotional state feature extraction module operations include: like Figure 3 As shown, the WavLM model architecture is adjusted: The original WavLM model has a fully connected layer at the end, which is used for binary classification tasks and has two nodes. To support multi-classification tasks, the number of output nodes in the last layer of the WavLM model needs to be modified to equal the number of sound categories. Assuming there are C emotional states, the new output layer should have C nodes. Softmax is used as the activation function to ensure that the output is a probability distribution, as shown in the following formula: in is the output score of the cth category, C is the total number of sound categories, based on the constructed medical speech dataset, the cross entropy loss function is used to train the model, and the calculation formula is: in are the elements in the true label vector.

[0045] In order to ensure that the original sound waveform is not lost, an additional original waveform feature conversion layer is introduced into the model, which can directly process the original audio data and generate feature representation . Use 1D Convolutional Layer to build the original waveform feature conversion layer. The input waveform length is , the number of channels is 1 (mono), then the output of a one-dimensional convolution layer can be expressed as: Where W is the convolution kernel weight matrix, b is the bias term, ∗ represents the convolution operation, and σ is the ReLU activation function, which ensures that the output feature dimension of the conversion layer matches the feature dimension extracted by WavLM for subsequent feature fusion.

[0046] The result of the original waveform feature conversion layer output , and the WavLM feature extraction layer results Splicing to form voice features : The “;” indicates the concatenation operation of the feature vector. The input is fed into a new fully connected layer, which is responsible for learning and predicting emotion category features. After passing through the Softmax activation function, the predicted probability distribution for each category is obtained, thereby determining the final emotion category. Through these steps, the improved WavLM model can be effectively used to extract multi-class emotion state features while preserving the original audio information, improving the model's ability to recognize complex emotions and providing real-time understanding of the patient's current emotional state during the conversation.

[0047] In one possible implementation, the comprehensive diagnosis decision module constructs a deep learning model that combines multimodal medical images and speech recognition features to assist doctors in making remote diagnosis decisions. It uses speech recognition features and multimodal medical image features fused with multimodal medical images to diagnose patient diseases, thus avoiding the problems of low prediction accuracy and low efficiency brought by a single modality. According to the above module processing, the speech features and multimodal medical image features are respectively and . Define an attention mechanism method for speech features and multimodal medical image features Fusion is performed to assign different weights to speech features and multimodal medical image features to dynamically adjust the importance of the features.

[0048] First, use the dot product to calculate the speech features and multimodal medical image features Similarity: In order to ensure that the sum of the weights between different modalities is 1, the similarity scores need to be normalized. and The weights corresponding to speech features and multimodal medical image features are calculated as follows: The original features are weighted and summed according to the calculated weights to generate the final fusion features : This approach implements the basic functionality of the attention mechanism through simple similarity calculations and normalization operations: dynamically adjusting feature weights based on their relationships. This avoids complex matrix operations and multi-head mechanisms, making the model more lightweight and easier to implement. At the same time, this simple attention mechanism still effectively captures interactive information between different modalities, thereby improving the effectiveness of multimodal data fusion.

[0049] like Figure 4 As shown, the final fusion feature The input is fed into the constructed DNN network to predict the patient's current disease category and assist doctors in making medical diagnoses. The DNN network shown in the figure includes an input layer, a hidden layer, a regularization function, and an output layer. The input layer receives data that is a fusion of speech features and multimodal medical image features. The hidden layer uses three fully connected layers. Each fully connected layer is followed by a ReLU activation function, batch normalization, and a Dropout layer to improve generalization. L1 regularization is used to better constrain model complexity. The final output layer is a fully connected network with nodes equal to the number of disease categories, C. It uses a Softmax activation function to generate a classification probability distribution and output possible disease classification results, thereby supporting doctors in making more accurate diagnostic decisions in remote situations.

[0050] The above description is a further detailed description of the present invention in conjunction with specific / preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art of the present invention may make various substitutions or modifications to the described embodiments without departing from the scope of the present invention, and such substitutions or modifications should be considered to fall within the scope of protection of the present invention.

[0051] Parts of the present invention that are not described in detail belong to the common knowledge of those skilled in the art.

Claims

1. A virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition, characterized by: Includes the following modules: Medical image preprocessing module: used to improve the quality and standardize medical image data from various sources; Medical Image Feature Extraction Module: Builds an image matching model guided by a keypoint attention mechanism to process medical images of different modalities; Multimodal fusion module: an encoder module guided by keypoint attention mechanism to achieve multimodal medical image fusion; Medical voice data collection module: collects various sounds of patients in different health states to build a medical voice dataset for emotional state feature extraction training; Emotional state feature extraction module: Emotional state feature extraction based on improved WavLM technology; Comprehensive diagnosis decision module: Based on multimodal medical image feature representation and speech feature representation, it outputs possible disease classification results.

2. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The medical image preprocessing module first applies a histogram equalization algorithm to redistribute pixel values and enhance image contrast; then, it converts the image from the spatial domain to the frequency domain through Fourier transform, and applies filters on this basis to remove noise and reduce blur; finally, it outputs a high-quality, standardized image.

3. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The medical image feature extraction module first applies the Haralick algorithm to extract texture features from medical images of different modalities; then generates key point maps corresponding to medical images of different modalities, captures the connection between corresponding areas between images, and ensures that key points of the same area are marked in medical images of different modalities.

4. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The encoder module guided by the key point attention mechanism is as follows: The key point graph extracted by the medical image feature extraction module guides the construction of the key point attention mechanism layer in the encoder module. This layer uses positional spatial information to enhance and focus on the propagation and fusion of locally important information. Then, the Haralick algorithm is used to extract key point graph features, which are further processed by the key point attention mechanism layer to generate a set of attention features. Finally, a multi-layer neural network is used to further fuse features from images of different modalities and integrate their attention features to output the final fused multimodal medical image feature representation.

5. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The medical voice data acquisition module selects a high-quality recording device microphone to collect voice data, systematically records the various voice characteristics of patients in different health states, and assigns a unique digital category code to each voice characteristic; then, it divides the continuous voice signal into short time frames and applies normalization technology to adjust the volume level of each voice sample to ensure that they have similar energy.

6. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 5 is characterized in that: A combined method including frame processing, Hanning window function and fast Fourier transform is used to perform frequency domain noise reduction.

7. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The improved WavLM technology is as follows: The WavLM model is a binary classification task. To transform it into a multi-classification task, the number of nodes in the last layer output is changed to the number of sound categories. At the same time, to ensure that the original sound waveform is not lost, an original waveform feature conversion layer is added. The output of the conversion layer is spliced with the result of the feature extraction layer. The spliced speech features are used as the input of the fully connected prediction layer, which outputs the final emotional state features.

8. The virtual doctor diagnosis assistance system integrating visual diffusion and speech emotion recognition according to claim 1 is characterized in that: The comprehensive diagnosis and decision module uses the multimodal fusion module to extract and fuse features from the multimodal medical images uploaded by the patient to obtain a multimodal medical image feature representation. Secondly, the emotional state feature extraction module extracts the emotional state and other sound characteristics from the patient's voice conversation, retaining the original waveform information, and splicing these features as a voice feature representation. Then, the multimodal medical image feature representation and speech feature representation are combined and input into the deep neural network model to learn and integrate the information of these two heterogeneous data sources and output possible disease classification results.

Citation Information

Cited By

  • Remote acquisition system based on four diagnosis methods of traditional Chinese medicine

    CN121171535A