Intelligent mental disease identification method and device based on multi-modal data and medium
By combining multimodal data fusion of self-report text of the disease, ear images and acupoint area impedance/temperature information, the deep learning model is used to achieve refined classification of mental illness, solving the problems of high misdiagnosis rate and insufficient generalization ability in the existing technology, and providing the accuracy and targeting of clinical diagnosis.
Patent Information
- Application Number
- CN202510327565.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The diagnosis of mental illness in the prior art depends on subjective scale evaluation, with high misdiagnosis rate and low efficiency, and insufficient generalization ability of existing automated analysis models and lack of joint modeling capabilities of multimodal data.
By obtaining the self-report text information of the disease, ear image data, and impedance and/or temperature information of the auricle and erectus areas, the multi-head attention mechanism and the Transformer architecture capture features, and combining deep learning models to fusion of multimodal data to achieve intelligent identification of mental illness.
It has achieved refined classification and identification of mental illnesses, can distinguish the severity of the disease, provides a targeted basis for clinical treatment, and improves diagnostic efficiency and accuracy.
Smart Images

Figure CN120496866A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and in particular to methods, devices, and media for intelligently identifying mental illnesses based on multimodal data. Background Art
[0002] In the field of medical artificial intelligence, the accurate diagnosis of mental illnesses (such as depression and cognitive impairment) is of great significance. Its clinical diagnosis relies on subjective scale evaluation and doctor experience, and there are problems such as high misdiagnosis rate and low efficiency. To solve the above problems, the existing technology mainly adopts auxiliary diagnosis methods based on single modality data (such as electroencephalogram or MRI), but the existing automated analysis model of single modality data has insufficient generalization ability and lacks the ability to jointly model multimodal data (such as physiological signals, images, and text). Although the existing multimodal pre-training models such as MedViLL have demonstrated the potential of cross-modal representation learning in the medical field, the existing technology does not provide effective multimodal data fusion analysis for mental illnesses.
[0003] It can be seen that traditional diagnostic methods have problems such as reliance on subjective judgment, high misdiagnosis rate, and low efficiency. Existing automated analysis models also face the problems of insufficient generalization ability and difficulty in joint modeling of multimodal data. Summary of the Invention
[0004] In view of the above problems, the present invention is proposed to provide a method, device and medium for intelligent identification of mental illness based on multimodal data that solves the above technical problems or at least partially solves the above technical problems.
[0005] One aspect of the present invention provides a method for intelligently identifying mental illness based on multimodal data, the method comprising:
[0006] Obtaining text information of the subject's self-reported condition, ear image data, and impedance and / or temperature information of each designated acupoint in the subject's concha and helix areas;
[0007] Perform local feature extraction on ear image data to obtain image feature vectors;
[0008] The impedance and / or temperature information of each acupoint area is converted into a vector form, and the multi-head attention mechanism in the Transformer architecture is used to capture the correlation characteristics between the impedance and / or temperature information of different acupoint areas to obtain the measurement value feature vector;
[0009] Extract the semantic feature vector of the self-reported text information of the disease;
[0010] The image feature vector, the measurement value feature vector, and the semantic feature vector are concatenated and fused to form a multimodal fusion feature vector;
[0011] The multimodal fusion feature vector is input into the pre-trained disease intelligent recognition model so that the disease intelligent recognition model can classify the input multimodal fusion feature vector according to the relationship between the multimodal fusion feature vector learned in the model training process and various categories of mental illnesses of different severities, and output the prediction result. The prediction result is a probability vector, which represents the probability of the sample belonging to various categories of mental illnesses of different severities and the normal state.
[0012] Optionally, the method further includes:
[0013] Before extracting local features from the ear image data, the ear image data is resized and pixel values are normalized;
[0014] performing Z-score normalization on the impedance and / or temperature information of each well before converting the impedance and / or temperature information of each well into a vector form; and
[0015] Before extracting the semantic feature vector of the self-described text information of the medical condition, the self-described text information of the medical condition is segmented to be truncated to a preset maximum length.
[0016] Optionally, after forming the multimodal fusion feature vector, the method further includes:
[0017] A multimodal joint modeling network composed of multiple layers of Transformer blocks arranged in series is used to perform feature transformation on the multimodal fusion feature vector for multiple iterations to obtain a multimodal fusion feature vector with deep fusion and feature enhancement;
[0018] Among them, each layer of Transformer block contains a multi-head attention mechanism and a feedforward neural network. The multi-head attention mechanism is used to perform attention calculation on the feature vector input to the current layer Transformer block, capture the correlation between different modal features, and perform feature transformation and enhancement through the feedforward neural network. The processed feature vector is passed to the next layer Transformer block. After multi-layer iterative processing, a multimodal fusion feature vector that has undergone deep fusion and feature enhancement is output.
[0019] Optionally, obtaining the impedance and / or temperature information of each designated acupoint in the concha and helix regions of the subject includes:
[0020] A pre-trained acupoint detection model is used to perform acupoint segmentation detection on the ear image data of the subject to be tested, and the acupoint distribution information of the concha and helix areas of the current subject to be tested is obtained;
[0021] generating an acupoint measurement and positioning template for the subject according to the acupoint distribution information, and transmitting the acupoint measurement and positioning template to a device for collecting ear image data, so as to display the acupoint measurement and positioning template in a video guide frame of an image acquisition interface of the acquisition device, and guiding a sensor device to collect impedance values and / or temperature values of each designated acupoint in the concha and helix regions of the subject by aligning the acupoint measurement and positioning template with the ear image of the subject;
[0022] Receive the impedance value and / or temperature value of each designated acupoint in the concha and helix areas of the subject uploaded by the sensor device.
[0023] Optionally, the impedance values and / or temperature values of each designated acupoint area are sorted according to the acquisition time and / or acupoint number.
[0024] Optionally, the acupuncture point detection model includes a convolutional neural network backbone network, a neck network, and a head network;
[0025] The convolutional neural network backbone network is used to extract features of different scales and levels from the ear image data of the subject to obtain multi-scale image feature information;
[0026] Neck network, used to fuse the multi-scale image feature information extracted by the convolutional neural network backbone network;
[0027] The head network is used to predict the acupoint area bounding box and key point coordinates of the fused features using a decoupled head structure to obtain the acupoint area distribution information of the concha and helix areas of the current subject.
[0028] Optionally, the training step of the disease intelligent identification model includes:
[0029] Obtain labeled sample data, each sample data including the patient's self-described text information of the condition, ear image data, impedance and / or temperature information of each designated acupoint in the patient's concha and helix, and the patient's disease label;
[0030] Performing size unification and pixel value normalization operations on the ear image data in each sample data to obtain a preprocessed image data sample, performing Z-score normalization on the impedance and / or temperature information in each sample data to obtain a preprocessed measurement data sample, and performing word segmentation processing on the self-description text information in each sample data to truncate it to a preset maximum length to obtain a preprocessed text data sample;
[0031] A convolutional neural network (CNN) is used to extract local features from each image data sample, and the feature dimension is reduced through a pooling layer network to obtain an image feature vector sample. The impedance and / or temperature information of each acupoint in each measurement data sample is converted into a vector form, and the multi-head attention mechanism in the Transformer architecture is used to capture the correlation features between the impedance and / or temperature information of different acupoints to obtain a measurement value feature vector sample. The BERT language model is used to encode each text data sample and extract the semantic feature vector sample of the text data sample. The image feature vector sample, measurement value feature vector sample, and semantic feature vector sample corresponding to each sample data are aligned in feature dimension, and the feature vector samples from different modalities are spliced and fused through feature cascading to form a multimodal fusion feature vector sample of each sample data.
[0032] Input the multimodal fusion feature vector samples of each sample data and the disease labels corresponding to each sample into the preset fully connected neural network classifier to train the disease intelligent recognition model;
[0033] During the model training process, the cross-entropy loss function is used to measure the loss value between the prediction result of each sample data of the model and the actual disease label, and the Adam optimizer is used to optimize the network parameters of the disease intelligent recognition model to minimize the loss function and achieve model performance optimization.
[0034] Optionally, the impedance and / or temperature information of the designated acupoints includes impedance values and / or temperature values of 18 acupoints in the concha area and 13 acupoints in the antihelix area.
[0035] Another aspect of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the method for intelligent identification of mental illness based on multimodal data as described in any one of the above items are implemented.
[0036] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for intelligent identification of mental illness based on multimodal data as described in any one of the above items.
[0037] The intelligent identification method, device and medium for mental illness based on multimodal data provided by the embodiments of the present invention provide a richer and more comprehensive information dimension for auxiliary identification of mental illness by innovatively integrating three modal data: ear image features, impedance and temperature information of each acupuncture point, and self-reported information of the patient. It also uses an advanced deep learning model architecture to fuse multimodal data. The model can automatically learn the association and complementary information between different modal data, thereby realizing refined classification and identification of mental illnesses, such as depression (mild, moderate and severe) and cognitive impairment (mild, moderate and severe). The present invention can not only realize auxiliary identification of the presence or absence of mental illness, but also distinguish the different severity of the disease, providing a more targeted basis for the formulation of clinical treatment plans, and has important clinical application value.
[0038] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Various other advantages and benefits will become apparent to those skilled in the art by reading the detailed description of the preferred embodiment below. The accompanying drawings are only for the purpose of illustrating the preferred embodiment and are not to be considered as limiting the present invention. In the accompanying drawings:
[0040] Figure 1 This is a flowchart of a method for intelligently identifying mental illness based on multimodal data according to an embodiment of the present invention;
[0041] Figure 2 Schematic diagram showing the mapping of the acupuncture point measurement and positioning template to the video guide frame in an embodiment of the present invention;
[0042] Figure 3 This is a flowchart of impedance and temperature information standardization processing in an embodiment of the present invention;
[0043] Figure 4 This is a flowchart of processing text information of a self-reported condition in an embodiment of the present invention;
[0044] Figure 5 This is a flowchart for implementing feature extraction of impedance and temperature information in an embodiment of the present invention;
[0045] Figure 6 This is a flowchart for implementing feature extraction of self-described text information of a medical condition in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0047] Those skilled in the art will understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art in the art to which the present invention pertains. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with those in the context of the prior art and, unless specifically defined, will not be interpreted in an idealized or overly formal sense.
[0048] Example 1
[0049] The embodiment of the present invention provides a method for intelligent identification of mental illness based on multimodal data, such as Figure 1 As shown, the method for intelligent identification of mental illness based on multimodal data proposed by the present invention includes the following steps:
[0050] S1. Obtain the subject's self-described text information on his / her condition, ear image data, and impedance and / or temperature information of each designated acupoint in the concha and helix areas of the subject.
[0051] There is a certain correspondence between the ear and various parts of the human body. By observing the shape, color, and luster of the ear, it can effectively assist in the diagnosis of mental illnesses, such as people with cognitive decline associated with depression and the continuous spectrum of dementia (subjective cognitive decline, mild cognitive impairment, mild to moderate dementia). The intelligent identification method for mental illness based on multimodal data proposed in the present invention, by obtaining multimodal data involving the physiological characteristics of the concha area of the subject and the patient's self-report, assists in the intelligent identification of mental illnesses such as depression and cognitive impairment, thereby effectively improving the recognition efficiency and accuracy. Among them, the multimodal data includes ear image data, impedance values and / or temperature values of 18 acupoints in the concha area and 13 acupoints in the antihelix area, and text information of the patient's self-report of the condition.
[0052] S2. Extract local features from the ear image data to obtain an image feature vector. Specifically, a convolutional neural network (CNN) can be used to extract local features from the ear image data, and a pooling layer network can be used to reduce feature dimensions to obtain the image feature vector.
[0053] This embodiment uses a feature extraction network based on a convolutional neural network (CNN) and a Transformer to extract features from multimodal data. The feature extraction network consists of three parts to extract features from different modal data. The feature extraction network requires powerful computing capabilities, such as a high-performance GPU, to accelerate the convolution and Transformer calculations.
[0054] Specifically, for image data, a convolutional neural network (CNN) is used to perform convolution operations on the blood vessels and dandruff images of the ear to extract local features of the image, such as edges and textures, and reduce the feature dimensions through pooling operations to obtain image feature vectors.
[0055] S3. Convert the impedance and / or temperature information of each acupoint area into a vector form, and use the multi-head attention mechanism in the Transformer architecture to capture the correlation features between the impedance and / or temperature information of different acupoint areas to obtain the measurement value feature vector.
[0056] Specifically, the impedance and / or temperature information of each well area is converted into a vector form as follows: each well area corresponds to an impedance value Z and a temperature value T. Z and T are standardized (Z-score and normalization), and Z and T of each well area are combined into a two-dimensional vector [Z, T}, which is then arranged into a sequence by well area order. Furthermore, after converting the impedance and / or temperature information into a vector form, the multi-head attention mechanism in the Transformer is used to capture the correlation features between the impedance and / or temperature information of different partitions by calculating the attention weights between features at different positions, thereby obtaining the corresponding feature vector.
[0057] S4. Extracting semantic feature vectors of the self-described medical condition text information. Specifically, the BERT language model can be used to encode the self-described medical condition text information to extract the semantic feature vectors of the self-described medical condition text information.
[0058] Specifically, for the text information of self-reported medical conditions, the pre-trained BERT language model based on Transformer is used to perform semantic understanding and feature extraction on the text to obtain the semantic feature vector of the text.
[0059] S5. Concatenate and fuse the image feature vector, the measurement value feature vector, and the semantic feature vector to form a multimodal fusion feature vector. Specifically, the image feature vector, the measurement value feature vector, and the semantic feature vector are aligned in terms of their feature dimensions. Features from different modalities are then concatenated and fused using feature concatenation to form a multimodal fusion feature vector.
[0060] Specifically, the image branch CNN extracts image features (output dimension 2048) and can randomly sample 180 regional features; the measurement branch fully connected layer maps impedance and temperature into a 768-dimensional embedding vector; the text branch BERT encodes the self-described text information of the disease into a 768-dimensional feature vector. By extracting features from multimodal data, representative features are extracted from data of different modalities, and the multimodal data is converted into a unified feature representation, which facilitates subsequent model learning and classification. After completing the feature extraction of all input data, the features of each modality are embedded into a unified dimensional space, which can be 768 dimensions, to achieve feature dimension alignment, and end after successfully splicing into a multimodal fusion feature vector. The multimodal fusion feature vector is generated as the input of the subsequent classifier, which effectively extracts the key features of the multimodal data, enhances the model's understanding and learning ability of the data, and provides strong support for accurate diagnosis.
[0061] S6. Input the multimodal fusion feature vector into the pre-trained disease intelligent recognition model, so that the disease intelligent recognition model can classify the input multimodal fusion feature vector according to the relationship between the multimodal fusion feature vector learned in the model training process and various categories of mental illnesses of different severity, and output the prediction result. The prediction result is a probability vector, which respectively represents the probability of the sample belonging to various categories of mental illnesses of different severity and the normal state.
[0062] The intelligent identification method for mental illness based on multimodal data provided by the embodiment of the present invention provides a richer and more comprehensive information dimension for auxiliary identification of mental illness by innovatively integrating three modal data: ear image features, impedance and temperature information of each acupuncture point, and self-reported information of the patient. It also uses an advanced deep learning model architecture to fuse multimodal data. The model can automatically learn the association and complementary information between different modal data, thereby realizing refined classification and identification of mental illnesses, such as depression (mild, moderate and severe) and cognitive impairment (mild, moderate and severe). The present invention can not only realize auxiliary identification of the presence or absence of mental illness, but also distinguish the different severity of the disease, providing a more targeted basis for the formulation of clinical treatment plans, and has important clinical application value.
[0063] In an embodiment of the present invention, the steps for obtaining the impedance and / or temperature information of each designated acupoint in the concha and helix areas of the subject include: using a pre-trained acupoint detection model to perform acupoint segmentation detection on the ear image data of the subject to be tested, and obtaining acupoint distribution information of the concha and helix areas of the current subject to be tested; generating an acupoint measurement and positioning template of the subject to be tested based on the acupoint distribution information, and sending the acupoint measurement and positioning template to an ear image data acquisition device to display the acupoint measurement and positioning template in a video guide frame of an image acquisition interface of the acquisition device, and guiding the sensor device to collect the impedance value and / or temperature value of each designated acupoint in the concha and helix areas of the subject to be tested by matching and aligning the acupoint measurement and positioning template with the ear image of the subject; and receiving the impedance value and / or temperature value of each designated acupoint in the concha and helix areas of the subject to be tested uploaded by the sensor device. The acupoint measurement and positioning template is an acupoint distribution image composed of 18 acupoints and key acupoint locations in the concha area and 13 acupoints and key acupoint locations in the antihelix area, in which the boundaries of the acupoints and the key feature points of the acupoints are marked with lines.
[0064] The acupoint detection model in this embodiment is trained using the YOLOv8 model. This model has been trained with a large amount of ear image data and can adapt to the differences in ear characteristics of different individuals. It can accurately segment the boundaries of each acupoint and the key feature points of the acupoint, providing a basis for accurate acupoint positioning with subsequent impedance and temperature value collection.
[0065] Specifically, an acquisition device equipped with an acupoint detection model is used to acquire images of the patient's ear, obtain the patient's ear image data and input it into the acupoint detection model. The acupoint detection model will analyze the acquired ear images and quickly and accurately identify the acupoint distribution information of the concha and helix areas. Ear acupoints are in two forms: acupoint points and acupoint areas. The acupoint distribution information in this application includes 18 acupoint areas and key acupoint points in the concha area, and 13 acupoint areas and key acupoint points in the antihelix area. After completing the precise segmentation and detection of the acupoint distribution information and when the acquisition device is ready, the acquisition device maps the acupoint measurement and positioning template to the video guide frame, such as Figure 2As shown, an intuitive guidance interface is generated and displayed on the screen of the acquisition device. Medical staff or patients follow the instructions of the guide box to accurately place the impedance sensor and / or temperature sensor in the corresponding acupuncture point area and key acupuncture point. The sensor is connected to the acquisition device. After it is in place, the device starts to collect the impedance value and / or temperature value of each acupuncture point in real time. The data is transmitted to the internal storage module of the device in real time. The impedance value and / or temperature value of each acupuncture point area are sorted and stored according to the acquisition time and / or acupuncture point number and displayed on the screen. The video guidance interface of the acquisition device is clear and easy to operate. The impedance sensor and temperature sensor are accurate and sensitive and have stable data transmission and storage capabilities. The present invention combines acupuncture point information with video guidance, and uses video guidance to guide the sensor to be accurately placed in the corresponding acupuncture point to accurately obtain data and realize real-time transmission, storage and display. This operation improves the accuracy of data acquisition, reduces operational errors, and real-time feedback data is convenient for adjustment, ensuring data quality. It is also simple to operate and easy to promote, providing convenience for large-scale data acquisition and clinical diagnosis. The acquisition of impedance and temperature values of all target acupoints is completed after the data verification is complete and no abnormalities are found. The real-time impedance and temperature data of 18 acupoint areas in the concha area and 13 acupoint areas in the antihelix area are obtained, realizing real-time and accurate acquisition of impedance and / or temperature, improving data validity and diagnostic accuracy.
[0066] The present invention provides a video-guided data acquisition method that maps 18 acupoint areas in the concha region, 13 acupoint areas in the antihelix region, and key acupoint points onto a video guide frame, enabling real-time acquisition of impedance and / or temperature. In traditional data acquisition processes, impedance and / or temperature acquisition is often disconnected from acupoint location, resulting in the inability to accurately map the collected data to specific acupoints, impacting data validity and diagnostic accuracy. By combining acupoint information with video guidance, the present invention provides intuitive and accurate guidance for data acquisition. During the acquisition process, medical personnel or patients simply follow the instructions in the video guide frame to accurately place the impedance and temperature sensors at the corresponding acupoint areas or points to achieve real-time data acquisition. This method not only improves data acquisition accuracy but also effectively reduces data errors caused by improper operation. Furthermore, real-time collected data can be promptly fed back to the operator, facilitating the timely detection and adjustment of abnormal data, further ensuring data quality. Furthermore, the video-guided data acquisition method offers the advantages of ease of operation and widespread adoption, making it applicable in various medical settings and facilitating large-scale data acquisition and clinical diagnosis.
[0067] In a specific embodiment, the acquisition device may be a multimodal data acquisition device that integrates a video guidance function, an impedance sensor, a temperature sensor, and the like.
[0068] In an embodiment of the present invention, the acupoint detection model includes a convolutional neural network backbone network, a neck network and a head network; the convolutional neural network backbone network is used to extract features of different scales and levels on the ear image data of the subject to be tested to obtain multi-scale image feature information; the neck network is used to perform feature fusion on the multi-scale image feature information extracted by the convolutional neural network backbone network; the head network is used to use a decoupling head structure to predict the acupoint area boundary box and key point coordinates of the fused features to obtain the acupoint distribution information of the concha and helix areas of the current subject to be tested.
[0069] Specifically, the convolutional neural network backbone of the acupoint detection model draws on the CSPDarkNet architecture. Through a series of convolution operations, it performs preliminary feature extraction on the input ear image data, providing the foundation for subsequent more refined feature processing and analysis. The convolutional neural network backbone includes the C2f module and the SPPF module. During the initial feature extraction process, components such as the C2f and SPPF modules are used to capture feature maps at various levels to obtain rich image feature information. The C2f module primarily captures feature maps at various levels and enables further processing and integration of input features. It effectively extracts feature information at different scales and levels, enabling the model to more comprehensively and deeply represent ear image features, thereby improving the model's ability to identify acupoint regions and key points. The C2f module includes the Bottleneck module. As a component of the C2f module, the Bottleneck module plays a role in optimizing feature extraction within the overall network architecture. It processes the input features through a specific convolution kernel configuration and connection scheme, minimizing computational effort while preserving important feature information as much as possible, avoiding loss of critical information during the feature extraction process, thereby improving model efficiency and accuracy. The SPPF module is responsible for feature extraction and representation enhancement. It further processes and enhances input features, effectively expanding the receptive field and rapidly aggregating contextual information. This enhances the model's ability to represent acupoint regions and keypoints in ear images, resulting in superior performance in detection and segmentation tasks. After the backbone network extracts multi-scale features, the Neck network utilizes an optimized Path Aggregation Network (PANet) to refine and fuse these features. This effective integration and transfer of features at different scales enables the model to fully utilize feature information at different levels, improving feature quality and usability and providing a more accurate feature foundation for subsequent predictions. The Head network is responsible for generating the final prediction. By replacing the coupled head with a decoupled head structure and employing specific strategies, this structural change and strategy enable the Head module to more flexibly process and analyze input features, resulting in more accurate predictions of acupoint region bounding boxes and keypoint coordinates. The decoupled head structure allows the model to perform more refined processing for different tasks, improving prediction accuracy and reliability. The patient's ear image is input and, after a series of operations such as convolution, C2f modules, upsampling, and concat, different layers output feature maps of varying sizes. Finally, the Detect layer detects the output of a specific layer to reduce system errors and detect all objects in the image. The model treats object detection as a regression problem, directly predicting the bounding boxes of acupoint regions and the coordinates of key points through a neural network. Image features are extracted using operations such as convolution, pooling, and feature fusion.This operation provides a precise acupoint location foundation for subsequent data collection, ensuring that the data corresponds to the correct acupoint area, providing a reliable basis for extracting relevant features, and improving the data quality and performance of the diagnostic model. The process ends after successfully outputting the segmentation results and key point location information for 18 acupoint areas in the concha region and 13 acupoint areas in the antihelix region. This achieves automated, high-precision acupoint segmentation and key point detection, improving data collection efficiency and accuracy, reducing manual positioning errors, and adapting to individual ear feature differences.
[0070] The present invention introduces the YOLOv8 model to segment and detect 18 acupoint areas in the concha area and 13 acupoint areas in the antihelix area. The YOLOv8 model has efficient target detection and segmentation capabilities, and can quickly and accurately identify acupoint areas and acupoint points. In the medical data collection process, accurate acupoint positioning is crucial. The traditional manual positioning method is not only inefficient, but also easily affected by subjective factors, resulting in inaccurate positioning. The application of the YOLOv8 model realizes automated, high-precision acupoint segmentation and key point detection, greatly improving the efficiency and accuracy of data collection. Through training on a large amount of ear image data, the YOLOv8 model can adapt to the differences in ear features of different individuals, accurately segment the boundaries of each acupoint, and mark key feature points for acupoints. These precise segmentation and detection results provide a reliable basis for the subsequent extraction of blood vessels, dandruff and other features of ear images, ensuring the consistency and accuracy of the data, thereby improving the performance of the entire diagnostic model.
[0071] In an embodiment of the present invention, the ear image data is resized and pixel value normalized before local feature extraction is performed on the ear image data; the impedance and / or temperature information of each acupoint area is Z-score standardized before being converted into vector form; and before extracting the semantic feature vector of the self-described text information of the medical condition, the self-described text information of the medical condition is word segmented to be truncated to a preset maximum length.
[0072] Specifically, in a specific embodiment, when new test data needs to be input into the disease intelligent recognition model for recognition, the data input interface program needs to pre-process the acquired multimodal data, including ear image data, impedance information and temperature information of 18 acupoint areas in the concha area and 13 acupoint areas in the antihelix area, and self-described text information of the disease condition. The ear image data is pre-processed according to a predetermined size and format, including adjusting the image size to a uniform size, normalizing pixel values, etc. The impedance and temperature information are numerically standardized, such as normalizing to a specific interval, such as Figure 3As shown in the figure, the obtained impedance and temperature information is previewed and checked to determine whether there are missing values or abnormal values in the data. If there are, the missing values are filled with the mean value, and the 3σ principle (three sigma principle) is used to filter out abnormal values. The characteristic mean u and standard deviation σ of the processed impedance and temperature information are calculated respectively. The standardization calculation method is: (characteristic value - characteristic mean u) / standard deviation σ, and the standardized impedance and temperature data are output and converted into fixed-dimensional impedance and temperature data vectors. The self-report text information of the disease is segmented and encoded, as shown in the figure. Figure 3 As shown, the acquired text information is preprocessed to remove special characters, convert to lowercase, etc. The word segmentation tool jieba can be used for word segmentation, and then the BERT language model is used for word vector encoding, including word embedding, position embedding, and segmentation embedding, to obtain the encoded self-description text data, and then the processed data is input into the subsequent corresponding feature extraction network. Furthermore, for ear image data, a convolutional neural network CNN can be used to extract local features of ear image data. For the impedance and temperature information converted into vector form, the multi-head attention mechanism in Transformer can be used to capture the correlation features between the impedance and temperature information of different partitions by calculating the attention weights between the features at different positions, such as Figure 5 As shown in the figure, the impedance and temperature data vectors are input into the input layer of the multi-head attention mechanism to generate the query vector Q, key vector K and value vector V. The attention weights between the features at different positions are calculated by the self-attention mechanism and combined with the value vector V to obtain the weighted value vector. The multi-head attention results are concatenated and feature fusion is performed through the fully connected layer to output the associated feature vector. For the self-report text information of the disease, such as Figure 6 As shown in FIG, the encoded self-description text data (word embedding, position embedding, segmentation embedding) is input into the BERT language model to obtain the semantic feature vector of the self-description text information of the disease.
[0073] The data input interface program must adapt to the format requirements of different modal data and have the computing resources required for data preprocessing. Based on the characteristics of different modal data, appropriate preprocessing algorithms are used to convert the data into a format suitable for model processing to ensure data consistency and compatibility. This operation preprocesses the raw multimodal data so that it can be effectively received and processed by the model, providing a foundation for subsequent feature extraction and model training. The process ends when all input data has been preprocessed and successfully transferred to the feature extraction module. The preprocessed multimodal data is obtained and output to the next module in a unified format, improving data availability, reducing the complexity of model training, and improving model training efficiency.
[0074] In an embodiment of the present invention, after forming a multimodal fusion feature vector, the method further includes: using a multimodal joint modeling network composed of multiple layers of Transformer blocks arranged in series to perform feature transformation on the multimodal fusion feature vector for multiple iterations to obtain a multimodal fusion feature vector with deep fusion and feature enhancement; wherein each layer of Transformer block includes a multi-head attention mechanism and a feedforward neural network, and the multi-head attention mechanism is used to perform attention calculation on the feature vector input into the current layer Transformer block to capture the correlation between different modal features, and the feedforward neural network is used to perform feature transformation and enhancement, and the processed feature vector is passed to the next layer of Transformer block. After multi-layer iterative processing, the multimodal fusion feature vector with deep fusion and feature enhancement is output.
[0075] Specifically, after the feature extraction network outputs the multimodal fusion feature vector, the multi-layer Transformer block begins feature enhancement processing. The multimodal fusion feature vector is input into the multi-layer Transformer block. Each layer of the Transformer block contains a multi-head attention mechanism, a feedforward neural network, and a residual connection and layer normalization network. Optionally, a 12-layer Transformer block and a 12-layer multi-head self-attention mechanism are set to allow cross-modal interaction between signals and images, and text is generated by autoregression. In the multi-head attention mechanism, multiple attention calculations are performed on the input feature vector to capture the complex relationship between different modal features. The feature is then transformed and enhanced through a feedforward neural network, residual connection, and layer normalization network. The processed feature vector is passed to the next layer of Transformer block. After multi-layer processing, a vector that has undergone deep fusion and feature enhancement is output. The multi-layer Transformer block has a large amount of computation and requires high-performance computing equipment, such as multi-GPU parallel computing to support fast calculations. The multi-head attention mechanism in the Transformer block can focus on different aspects of the input features in parallel, while the feedforward neural network further transforms and nonlinearly processes the attention output. Through multi-layer stacking, it continuously explores the deep connections between features. This operation deeply fuses and enhances the multimodal fusion features, further exploring the complex relationships between features of different modalities and improving the model's expressiveness and classification performance. The multimodal fusion feature vector is processed through all the preset Transformer layers, and the deeply processed and fused feature vector is output for subsequent classification operations. This improves the model's understanding and processing capabilities of multimodal data, enhances the model's robustness, and helps improve diagnostic accuracy.
[0076] In an embodiment of the present invention, after the multi-layer Transformer block completes feature processing and outputs a multimodal fusion feature vector that has undergone deep fusion and feature enhancement, the pre-built disease intelligent identification model, namely the fully connected neural network classifier, begins to work. The feature vector processed by the multi-layer Transformer block is input into the disease intelligent identification model, which performs nonlinear transformation on the input features through multiple hidden layers and finally outputs the prediction result through the output layer. This embodiment uses depression and cognitive impairment as examples of predicted diseases, that is, the prediction result is a probability vector, which represents the probability of the sample belonging to depression (mild), depression (moderate), depression (severe), cognitive impairment (mild), cognitive impairment (moderate), cognitive impairment (severe), and normal state respectively. The fully connected neural network performs a linear transformation on the input features through the weight matrix, then introduces nonlinearity through the activation function, continuously combines and transforms the features, and finally outputs the classification result. This operation classifies the input multimodal data based on the relationship between the multimodal features learned by the model and the disease, outputs the diagnosis result, and realizes the judgment of depression and cognitive impairment. The classification calculation of the input feature vector is completed, and the predicted probability vector is output, ending the process. The resulting probability vectors predict that the sample belongs to different categories (depression (mild), depression (moderate), depression (severe), cognitive impairment (mild), cognitive impairment (moderate), cognitive impairment (severe), and normal state). This enables classification diagnosis of multimodal data and provides a diagnostic reference for doctors. During model training, the cross-entropy loss function is used to measure the difference between the predicted results and the true labels. By minimizing this loss function, the model parameters are continuously adjusted. The Adam optimizer is selected, combining the advantages of the Adagrad and RMSProp optimizers to adaptively adjust the learning rate of each parameter. Based on the calculated cross-entropy loss, the gradient is calculated and the update step size of the model parameters is adjusted based on the first-order moment estimate and the second-order moment estimate of the gradient. The model parameters are updated in the direction of decreasing the loss function, thereby gradually optimizing the model performance.
[0077] In an embodiment of the present invention, the training steps of the disease intelligent recognition model include:
[0078] First, obtain labeled sample data. Each sample data includes the patient's self-described text information of the condition, ear image data, impedance and / or temperature information of each designated acupoint in the patient's concha and helix areas, and the patient's disease label.
[0079] Second, the ear image data in each sample data is resized and pixel value normalized to obtain preprocessed image data samples, the impedance and / or temperature information in each sample data is Z-score standardized to obtain preprocessed measurement data samples, and the self-described text information of the disease in each sample data is word segmented to be truncated to a preset maximum length to obtain preprocessed text data samples.
[0080] In this embodiment, multimodal data collection is required in advance with the patient's authorization for data use to obtain data for model training. After the patient completes the clinical examination and authorizes the use of the data, the collection system (ear collection equipment, impedance sensor, electronic medical record system) starts working. First, the information collector's self-description is converted into text content. For the text content, the annotation personnel mark the key sentences related to the disease in the information collector's self-description text based on the description of symptoms related to depression and cognitive impairment, diagnosis results and other information to obtain the self-description text information of the condition, and mark the corresponding disease type; for ear image data, the annotation personnel mark the disease-related features and corresponding disease categories of each partition based on the morphology, texture and other features in the ear image, combined with medical knowledge; for impedance and / or temperature information table data, the annotation personnel mark whether the data is abnormal and its association with depression or cognitive impairment based on the numerical range and change trend of the data, combined with clinical experience. The annotation personnel must have professional medical knowledge and use preset auxiliary annotation software to implement sample annotation. The auxiliary annotation software has a friendly annotation interface. Based on their medical knowledge and experience, annotators manually judge and label data from different modalities. Assisted annotation software provides data display, annotation recording, and storage, providing labeled training data for subsequent model training. This allows the model to learn the corresponding relationships between different modal data and diseases, a key step in accurate classification and diagnosis. Labeling is completed after all cleaned data has been completed and passed quality review. This results in labeled multimodal medical data samples, each containing data from different modalities and corresponding disease labels. This provides high-quality supervised data for model training and helps improve the diagnostic accuracy of the model for depression and cognitive impairment. Concha images of the 18 subareas are grayscale normalized to a scale of 0-255 and a uniform size of 512×512. Impedance (Ω) and temperature (°C) values are normalized using the Z-score. Patient self-report text is tokenized and truncated to a maximum length of 256 words to eliminate data heterogeneity and ensure input consistency.
[0081] Third, a convolutional neural network (CNN) is used to extract local features from each image data sample, and the feature dimension is reduced through a pooling layer network to obtain an image feature vector sample. The impedance and / or temperature information of each acupoint in each measurement data sample is converted into a vector form, and the multi-head attention mechanism in the Transformer architecture is used to capture the correlation features between the impedance and / or temperature information of different acupoints to obtain a measurement value feature vector sample. The BERT language model is used to encode each text data sample and extract the semantic feature vector sample of the text data sample. The image feature vector sample, measurement value feature vector sample, and semantic feature vector sample corresponding to each sample data are aligned in feature dimension, and the feature vector samples from different modalities are spliced and fused through feature cascade to form a multimodal fusion feature vector sample of each sample data.
[0082] Fourth, the multimodal fusion feature vector samples of each sample data and the disease labels corresponding to each sample are input into the preset fully connected neural network classifier to train the disease intelligent recognition model;
[0083] Fifth, during the model training process, the cross-entropy loss function is used to measure the loss value between the predicted results of each sample data of the model and the actual disease label, and the Adam optimizer is used to optimize the network parameters of the disease intelligent recognition model to minimize the loss function and achieve model performance optimization.
[0084] In this embodiment, after model construction is complete, the training task is initiated, and the model training program begins processing the labeled multimodal medical data. The labeled multimodal medical data is divided into a training set, a validation set, and a test set according to a specific ratio, such as 70% for training, 15% for validation, and 15% for testing. The training set data is sequentially input into each module of the model. According to the calculation process set in the model construction, forward propagation calculations are performed to obtain prediction results. The cross-entropy loss is then calculated based on the predicted results and the true labels. The Adam optimizer is used to calculate the gradient based on the loss value and update the model parameters. During the training process, the model is regularly verified using the validation set data. Metrics such as loss and accuracy on the validation set are calculated to observe the model training effect. Model training is performed on a server with powerful computing power, equipped with multiple high-performance GPUs to accelerate the training process, and with sufficient memory to store training data and model parameters. Based on the training principles of deep learning, the prediction results are calculated through forward propagation, the prediction error is measured using a loss function, and the model parameters are adjusted through an optimizer. Training is continuously iterated to enable the model to learn the relationship between multimodal data and diseases. This operation enables the model to learn effective feature representations and classification decision boundaries from a large amount of labeled data, continuously optimizing model performance and improving the model's ability to diagnose depression and cognitive impairment. When the model's loss on the validation set stops decreasing or decreases very slowly, such as when the loss decreases by less than a set threshold, such as 0.001, for multiple consecutive training rounds, or when the model's accuracy on the validation set reaches a certain stable level, such as when the accuracy fluctuations for multiple consecutive training rounds are less than a set threshold, such as 0.01, training is stopped, resulting in a trained intelligent disease recognition model, also known as the earvill model. The model has the ability to assist in the diagnosis of depression and cognitive impairment based on multimodal data input from the concha region of the ear. This model is trained to effectively process multimodal data and accurately identify diseases, providing a powerful auxiliary tool for actual clinical diagnosis.
[0085] Furthermore, after model training is complete, or if it is found during actual application that the model's diagnostic accuracy needs further improvement, the trained EarViLL model can be optimized using a model optimization algorithm. First, the model is cross-validated. The training set data is further divided into multiple subsets. A 5-fold cross-validation is performed, where the training set is divided into five subsets. Four subsets are used for training and one subset for validation, and this process is repeated multiple times to comprehensively evaluate the model's performance. Based on the cross-validation results, the model's hyperparameters are adjusted, such as the number of Transformer layers, the number of neurons in the fully connected neural network's hidden layers, and the learning rate. The model is then retrained and validated, comparing the performance of the model under different hyperparameter settings and selecting the hyperparameter combination with the best performance. Furthermore, the amount of training data can be increased by collecting more multimodal medical data from the concha region to update the sample data. The data can then be relabeled, the model trained, and validated to observe changes in model performance. The model optimization process is performed in an environment with sufficient computing power and storage space to record and analyze large amounts of experimental data. Cross-validation evaluates the model's generalization capabilities more comprehensively by splitting the training and validation sets multiple times. Hyperparameter tuning explores the optimal model configuration by changing the model's structure and training parameters. Increasing the amount of training data utilizes more sample information, enabling the model to learn more comprehensive features and relationships. Model optimization can further improve the model's classification accuracy and generalization capabilities, making it more adaptable to different application scenarios.
[0086] The intelligent identification method, device and medium for mental illness based on multimodal data provided by the embodiments of the present invention provide a richer and more comprehensive information dimension for auxiliary identification of mental illness by innovatively integrating three modal data: ear image features, impedance and temperature information of each acupuncture point, and self-reported information of the patient. It also uses an advanced deep learning model architecture to fuse multimodal data. The model can automatically learn the association and complementary information between different modal data, thereby realizing refined classification and identification of mental illnesses, such as depression (mild, moderate and severe) and cognitive impairment (mild, moderate and severe). The present invention can not only realize auxiliary identification of the presence or absence of mental illness, but also distinguish the different severity of the disease, providing a more targeted basis for the formulation of clinical treatment plans, and has important clinical application value.
[0087] For simplicity of description, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because certain steps can be performed in other orders or simultaneously according to the embodiments of the present invention. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0088] Example 2
[0089] Another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for intelligent identification of mental illness based on multimodal data. Figure 1 Steps S1-S6 are shown.
[0090] In the specific implementation process of the second embodiment, reference may be made to the first embodiment, and the corresponding technical effects are achieved.
[0091] Example 3
[0092] An embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the embodiment of the method for intelligent identification of mental illness based on multimodal data are implemented. Figure 1 Steps S1-S6 are shown.
[0093] In the specific implementation process of Example 3, reference may be made to Example 1, and the corresponding technical effects are achieved.
[0094] Furthermore, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of the present invention and to form different embodiments. For example, any of the claimed embodiments may be used in any combination.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for intelligent identification of mental illness based on multimodal data, characterized in that: The method comprises: Obtaining text information of the subject's self-reported condition, ear image data, and impedance and / or temperature information of each designated acupoint in the subject's concha and helix areas; Perform local feature extraction on ear image data to obtain image feature vectors; The impedance and / or temperature information of each acupoint area is converted into a vector form, and the multi-head attention mechanism in the Transformer architecture is used to capture the correlation characteristics between the impedance and / or temperature information of different acupoint areas to obtain the measurement value feature vector; Extract the semantic feature vector of the self-reported text information of the disease; The image feature vector, the measurement value feature vector, and the semantic feature vector are concatenated and fused to form a multimodal fusion feature vector; The multimodal fusion feature vector is input into the pre-trained disease intelligent recognition model so that the disease intelligent recognition model can classify the input multimodal fusion feature vector according to the relationship between the multimodal fusion feature vector learned in the model training process and various categories of mental illnesses of different severities, and output the prediction result. The prediction result is a probability vector, which represents the probability of the sample belonging to various categories of mental illnesses of different severities and the normal state.
2. The method according to claim 1, characterized in that The method further comprises: Before extracting local features from the ear image data, the ear image data is resized and pixel values are normalized; performing Z-score normalization on the impedance and / or temperature information of each well before converting the impedance and / or temperature information of each well into a vector form; and Before extracting the semantic feature vector of the self-described text information of the medical condition, the self-described text information of the medical condition is segmented to be truncated to a preset maximum length.
3. The method according to claim 1, characterized in that After forming the multimodal fusion feature vector, the method further includes: A multimodal joint modeling network composed of multiple layers of Transformer blocks arranged in series is used to perform feature transformation on the multimodal fusion feature vector for multiple iterations to obtain a multimodal fusion feature vector with deep fusion and feature enhancement; Among them, each layer of Transformer block contains a multi-head attention mechanism and a feedforward neural network. The multi-head attention mechanism is used to perform attention calculation on the feature vector input to the current layer Transformer block, capture the correlation between different modal features, and perform feature transformation and enhancement through the feedforward neural network. The processed feature vector is passed to the next layer Transformer block. After multi-layer iterative processing, a multimodal fusion feature vector that has undergone deep fusion and feature enhancement is output.
4. The method according to claim 1, wherein Obtaining impedance and / or temperature information of each designated acupoint in the concha and helix regions of the subject includes: A pre-trained acupoint detection model is used to perform acupoint segmentation detection on the ear image data of the subject to be tested, and the acupoint distribution information of the concha and helix areas of the current subject to be tested is obtained; generating an acupoint measurement and positioning template for the subject according to the acupoint distribution information, and transmitting the acupoint measurement and positioning template to a device for collecting ear image data, so as to display the acupoint measurement and positioning template in a video guide frame of an image acquisition interface of the acquisition device, and guiding a sensor device to collect impedance values and / or temperature values of each designated acupoint in the concha and helix regions of the subject by aligning the acupoint measurement and positioning template with the ear image of the subject; Receive the impedance value and / or temperature value of each designated acupoint in the concha and helix areas of the subject uploaded by the sensor device.
5. The method according to claim 2, characterized in that The impedance values and / or temperature values of each designated acupoint area are sorted according to the acquisition time and / or acupoint number.
6. The method according to claim 2, characterized in that The acupuncture point detection model includes a convolutional neural network backbone network, a neck network, and a head network; The convolutional neural network backbone network is used to extract features of different scales and levels from the ear image data of the subject to obtain multi-scale image feature information; Neck network, used to fuse the multi-scale image feature information extracted by the convolutional neural network backbone network; The head network is used to predict the acupoint area bounding box and key point coordinates of the fused features using a decoupled head structure to obtain the acupoint area distribution information of the concha and helix areas of the current subject.
7. The method according to claim 1, characterized in that The training steps of the disease intelligent identification model include: Obtain labeled sample data, each sample data including the patient's self-described text information of the condition, ear image data, impedance and / or temperature information of each designated acupoint in the patient's concha and helix, and the patient's disease label; Performing size unification and pixel value normalization operations on the ear image data in each sample data to obtain a preprocessed image data sample, performing Z-score normalization on the impedance and / or temperature information in each sample data to obtain a preprocessed measurement data sample, and performing word segmentation processing on the self-description text information in each sample data to truncate it to a preset maximum length to obtain a preprocessed text data sample; A convolutional neural network (CNN) is used to extract local features from each image data sample, and the feature dimension is reduced through a pooling layer network to obtain an image feature vector sample. The impedance and / or temperature information of each acupoint in each measurement data sample is converted into a vector form, and the multi-head attention mechanism in the Transformer architecture is used to capture the correlation features between the impedance and / or temperature information of different acupoints to obtain a measurement value feature vector sample. The BERT language model is used to encode each text data sample and extract the semantic feature vector sample of the text data sample. The image feature vector sample, measurement value feature vector sample, and semantic feature vector sample corresponding to each sample data are aligned in feature dimension, and the feature vector samples from different modalities are spliced and fused through feature cascading to form a multimodal fusion feature vector sample of each sample data. Input the multimodal fusion feature vector samples of each sample data and the disease labels corresponding to each sample into the preset fully connected neural network classifier to train the disease intelligent recognition model; During the model training process, the cross-entropy loss function is used to measure the loss value between the prediction result of each sample data of the model and the actual disease label, and the Adam optimizer is used to optimize the network parameters of the disease intelligent recognition model to minimize the loss function and achieve model performance optimization.
8. The method according to claim 5, characterized in that The impedance and / or temperature information of the designated acupoints includes impedance values and / or temperature values of 18 acupoints in the concha area and 13 acupoints in the antihelix area.
9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-modal information-based depression prediction method and related device
CN116110567A
Ear acupoint recognition method and system based on image processing
CN118968547A
Health state classification method, system and equipment based on multi-modal data and medium
CN119046766A
Mental disease category identification method and device, computer equipment and storage medium
CN119377822A
Auricular point real-time detection and auxiliary positioning method and system and storage medium
CN119445056A
Cited By
Otological disease prediction method based on multi-modal data fusion and confidence evaluation
CN121096599A
Ear disease prediction method based on multi-modal data fusion and confidence evaluation
CN121096599B
Mental disease identification method combining electronic medical record and electroencephalogram signal
CN122314351A