Health state classification model based on multi-modal data, training method and classification method

Through the health status classification model of multimodal data, the limitations of determining health status by a single physical examination report are solved, and the accurate identification and rapid classification of multimodal data are achieved.

CN120448905APending Publication Date: 2025-08-08HANGZHOU HUAXI INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510530829.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-23
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately determine the user's health status through a single user physical examination report, and there are limitations.

Method used

A health state classification model based on multimodal data is adopted, including a multimodal feature extraction module and a multimodal feature fusion encoder, and a variety of health data are processed through a multimodal attention module and a convolution module, and the fusion feature extraction and optimization are carried out.

Benefits of technology

It realizes accurate identification of multimodal user data, quickly determines the user's health status, and provides help for the classification of health status based on multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448905A_ABST
    Figure CN120448905A_ABST
Patent Text Reader

Abstract

The invention discloses a health state classification model based on multi-modal data, a training method and a classification method. The model comprises a multi-modal feature extraction module, a multi-modal attention module and a convolution module. The multi-modal attention module comprises a feature head; the multi-modal feature extraction module is used for extracting data features of the modal health data sets to obtain modal features matched with the modal health data sets; the multi-modal attention module is used for respectively storing each modal feature into each feature head, and respectively processing candidate modal features in the target feature head and reference modal features in other feature heads to obtain optimized target modal features; the convolution module is used for processing each target modal feature and the first feature after the jump connection of each modal feature through different convolution layers to obtain a fusion feature; according to the scheme, help can be provided for accurately and rapidly determining the health state of the user based on the multi-source data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a health status classification model based on multimodal data, and a training and classification method. Background Art

[0002] In recent years, with the continuous improvement of residents' living standards, people have paid more and more attention to their own health status; at this stage, the health status of users is mainly obtained by medical staff through interpretation of users' physical examination reports.

[0003] With the rapid development of artificial intelligence technology, large models have demonstrated remarkable processing capabilities in many fields such as natural language processing, image recognition and data analysis. More and more artificial intelligence technologies are being used for the intelligent interpretation of physical examination reports to determine the user's health status.

[0004] However, determining a user's health status solely through their physical examination report has certain limitations and is difficult to accurately determine the user's health status. How to determine a classification model that can identify multimodal user data and provide assistance in accurately and quickly determining the user's health status based on multi-source data is a key research issue in the industry. Summary of the Invention

[0005] The present invention provides a health status classification model, training and classification method based on multimodal data to solve the problem that determining a user's health status solely through the user's physical examination report has certain limitations and is difficult to accurately determine the user's health status. It can identify multimodal user data and provide assistance for accurately and quickly determining the user's health status based on multi-source data.

[0006] According to one aspect of the present invention, a health status classification model based on multimodal data is provided. The health status classification model based on multimodal data includes: a multimodal feature extraction module and a multimodal feature fusion encoder; the multimodal feature fusion encoder includes: a multimodal attention module and a convolution module; the multimodal feature extraction module is communicatively connected to the multimodal feature fusion encoder; the multimodal attention module includes at least two feature heads; wherein the number of the feature heads is the same as the number of modalities of the multimodal data;

[0007] The multimodal feature extraction module is used to extract data features of each modal health data group respectively to obtain modal features matching each modal health data group; wherein the number of modal features is the same as the number of modalities included in the modal health data group;

[0008] The multimodal attention module is used to store each of the modal features in each feature head respectively, and process the candidate modal features in the target feature head with the reference modal features in other feature heads respectively to obtain the optimized target modal features;

[0009] The convolution module is used to process each of the target modal features and the first features after jump connection of each modal feature through different convolution layers to obtain a fusion feature.

[0010] According to another aspect of the present invention, a method for training a health status classification model based on multimodal data is provided, the method comprising:

[0011] Acquire sample data; the sample data includes a modal health data group and label data corresponding to each modal health data group; the same modal health data group includes at least two model data;

[0012] Inputting the sample data into a health status classification model based on multimodal data for iterative training, and obtaining a health status classification model based on multimodal data for classifying health status when an iteration stop condition is met;

[0013] The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data described in any one of the embodiments of the present invention.

[0014] According to another aspect of the present invention, a health status classification method based on multimodal data is provided, the method comprising:

[0015] Acquire target health data of at least two modalities of a target user;

[0016] Inputting the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user;

[0017] Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data described in any embodiment of the present invention, and is trained by the training method of the health status classification model based on multimodal data described in any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a device for training a health status classification model based on multimodal data, the device comprising:

[0019] A sample data acquisition module is used to acquire sample data; the sample data includes a modal health data group and label data corresponding to each modal health data group; the same modal health data group contains at least two model data;

[0020] a training module, configured to input the sample data into a health status classification model based on multimodal data for iterative training, and obtain a health status classification model based on multimodal data for classifying health status when an iteration stopping condition is satisfied;

[0021] The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of the embodiments of the present invention.

[0022] According to another aspect of the present invention, a device for classifying health status based on multimodal data is provided, the device comprising:

[0023] A health data acquisition module, configured to acquire target health data of a target user in at least two modalities;

[0024] a health warning level determination module, configured to input the target health data into the health status classification model based on multimodal data for processing, and obtain the health warning level of the target user;

[0025] Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any embodiment of the present invention, and is trained by the training method of the health status classification model based on multimodal data as described in any embodiment of the present invention.

[0026] According to another aspect of the present invention, an electronic device is provided, comprising:

[0027] at least one processor; and

[0028] a memory communicatively connected to the at least one processor; wherein,

[0029] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the health status classification model based on multimodal data or the health status classification method based on multimodal data described in any embodiment of the present invention.

[0030] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions, and the computer instructions are used to enable a processor to implement the training method of a health status classification model based on multimodal data or the health status classification method based on multimodal data described in any embodiment of the present invention when executed.

[0031] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the training method for a health status classification model based on multimodal data, or the health status classification method based on multimodal data, as described in any embodiment of the present invention.

[0032] The technical solution of the embodiment of the present invention provides a health status classification model based on multimodal data, which may include: a multimodal feature extraction module and a multimodal feature fusion encoder; the multimodal feature fusion encoder includes: a multimodal attention module and a convolution module; the multimodal feature extraction module is communicatively connected to the multimodal feature fusion encoder; the multimodal attention module includes at least two feature heads; wherein the number of the feature heads is the same as the number of modalities of the multimodal data; the multimodal feature extraction module is used to extract data features of each modal health data group respectively, and obtain each modal feature matching each modal health data group; wherein the number of modal features is the same as the number of modal features included in the modal health data group. The number of modalities contained is the same; the multimodal attention module is used to store each of the modal features respectively in each feature head, and process the candidate modal features in the target feature head with the reference modal features in other feature heads respectively to obtain the optimized target modal features; the convolution module is used to process each of the target modal features and the first feature after jump connection of each modal feature through different convolution layers to obtain a fusion feature, which solves the problem that there are certain limitations in determining the user's health status through a single physical examination report of the user, and it is difficult to accurately determine the user's health status. It can identify multimodal user data and provide assistance for accurately and quickly determining the user's health status based on multi-source data.

[0033] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0035] Figure 1 1 is a structural diagram of a health status classification model based on multimodal data provided according to the first embodiment of the present invention;

[0036] Figure 2 1 is a schematic structural diagram of a multimodal feature extraction module provided according to the first embodiment of the present invention;

[0037] Figure 3 1 is a schematic structural diagram of a multimodal fusion encoder provided according to embodiment 1 of the present invention;

[0038] Figure 4 2 is a schematic diagram of the principle of a multimodal attention module provided according to the first embodiment of the present invention;

[0039] Figure 5 is a flowchart of a method for training a health status classification model based on multimodal data according to a second embodiment of the present invention;

[0040] Figure 6 is a flowchart of a health status classification method based on multimodal data provided according to embodiment three of the present invention;

[0041] Figure 7 2 is a schematic structural diagram of a training device for a health status classification model based on multimodal data according to a fourth embodiment of the present invention;

[0042] Figure 8 2 is a schematic structural diagram of a health status classification device based on multimodal data according to a fifth embodiment of the present invention;

[0043] Figure 9 It is a structural diagram of an electronic device for implementing a training method for a health status classification model based on multimodal data or a health status classification method based on multimodal data according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0045] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0046] Example 1

[0047] Figure 1 is a structural diagram of a health status classification model based on multimodal data provided according to the first embodiment of the present invention. The model can be used to determine the health status of a user, such as Figure 1 As shown, the device includes: a multimodal feature extraction module 110 and a multimodal feature fusion encoder 120; the multimodal feature fusion encoder 120 includes: a multimodal attention module 121 and a convolution module 122; the multimodal feature extraction module 110 is communicatively connected to the multimodal feature fusion encoder 120; the multimodal attention module 121 includes at least two feature heads ( Figure 1 (not shown in the figure); wherein the number of the feature heads is the same as the number of modalities of the multimodal data.

[0048] Among them, multimodal data may include: video data, audio data, text data, sensor data, tactile data, biometric data and social network data; among them, one type of data can be called a data modality; the multimodal data involved in this embodiment may include two or more of the above data at the same time, and this embodiment does not limit the specific number of modalities.

[0049] It should be noted that Figure 1 The description is only given by taking images, audio and text data as examples, which is not a limitation of this embodiment.

[0050] Optionally, in this embodiment, the multimodal feature extraction module 110 can be used to extract data features of each modal health data group respectively to obtain modal features that match each of the modal health data groups; wherein the number of modal features is the same as the number of modalities contained in the modal health data group.

[0051] In this embodiment, the modal health data group may include at least two of the following modal data: video data, audio data, text data, sensor data, tactile data, biometric data, and social network data.

[0052] For example, a user's health data set of each modality may include the user's physical examination report (text data), electrocardiogram (image data), 24-hour blood glucose monitoring data (biometric data), and family members' illness data (social network data).

[0053] It can be understood that in this embodiment, the features in a health data group of each modality can be the same features, that is, display data for the same target, for example, text data, audio data or image data based on the user's heart examination.

[0054] In an optional implementation of this embodiment, the multimodal feature extraction module 110 can extract the features of each modal health data group through a convolutional neural network, thereby obtaining each modal feature that matches each modal health data group; illustratively, if the modal health data group contains image data, text data, and audio data, then each modal feature can be an image feature, a text feature, and an audio feature, respectively.

[0055] Figure 2 This is a schematic diagram of the structure of a multimodal feature extraction module provided according to the first embodiment of the present invention. It should be noted that: Figure 2 The description is only based on image data, audio data and text data as examples. In this embodiment, the multimodal feature extraction module can also be used to extract features of other health data such as sensor data, tactile data, biometric data, etc. Figure 2 As can be seen from the figure, for different modal data, the multimodal feature extraction module can adopt different methods to extract features. For example, it can extract image features through convolutional neural networks, extract audio features through Mel spectrograms, and extract text features through LSTM (Long Short-Term Memory).

[0056] Optionally, in this embodiment, the multimodal attention module 121 can be used to store each of the modal features in each feature head respectively, and process the candidate modal features in the target feature head with the reference modal features in other feature heads respectively to obtain the optimized target modal features.

[0057] The target feature head is any feature head in the multimodal attention module 121 , and the other feature heads are all feature heads in the multimodal attention module 121 except the target feature head.

[0058] It can be understood that in this embodiment, after the features of each modality of health data are extracted through the multimodal feature extraction module, there is no need to splice, copy, etc. the extracted modal features. Instead, the features of different models are directly stored in different feature heads of the multimodal attention module 121, and the candidate modal features in the target feature head are processed separately with the reference modal features in other feature heads, and then each feature is optimized to obtain optimized modal features; for example, in the above example, optimized text features, image features, and audio features are obtained.

[0059] Optionally, in this embodiment, the convolution module is used to process each of the target modal features and the first feature after jump connection of each modal feature through different convolution layers to obtain a fusion feature.

[0060] In this embodiment, the convolution module may include a 1*1 convolution layer, a 3*3 depth convolution layer and a 1*1 convolution layer arranged in sequence, which can mainly fuse multiple text features, image features and audio features into a fusion feature.

[0061] Figure 3 is a structural diagram of a multimodal fusion encoder provided according to the first embodiment of the present invention, such as Figure 3 As shown, the multimodal feature fusion encoder may further include: a feedforward network and a normalization layer; the feedforward network is respectively communicated with the convolution module and the normalization layer.

[0062] Optionally, in this embodiment, the feedforward network can be used to extract the second feature after the fusion feature is jump-connected with each of the target modal features to obtain the target fusion feature; the normalization layer can be used to process the third feature after the target fusion feature is jump-connected with the fusion feature to obtain a unified feature.

[0063] like Figure 2 The health status classification model based on multimodal data shown may also include a fully connected layer; the fully connected layer is communicatively connected to the multimodal feature fusion encoder; the fully connected layer can be used to map the unified features to the health warning level.

[0064] Optionally, in this embodiment, image features, audio features and text features are used as inputs of the multimodal attention module, the output of the multimodal attention module is jump-connected with the image features, audio features and text features and then used as inputs of the convolution module, the output of the convolution module is jump-connected with the output of the multimodal attention module and then input into the feedforward network, the output of the feedforward network is jump-connected with the output of the convolution module and then input into the normalization layer, the normalization layer outputs unified features and then inputs into the fully connected layer, and the fully connected layer outputs the health warning level.

[0065] Figure 4 This is a schematic diagram of the principle of a multimodal attention module provided by Example 1 of the present invention. The features of multiple modal data are extracted. In the multimodal attention module, the features between each two modalities will independently calculate the attention weight to adjust the influence of each feature on the final prediction result in various situations. Since the features of each modality correspond to an independent, non-shared and learnable weight transformation W, each two modalities share a learnable modal feature fusion transformation W. cross In this embodiment, W cross The role of is to fuse the information between multimodal data. Therefore, through the above calculation, the information between various modalities can be fused to adaptively enhance the influence of important features on the prediction results and weaken the influence of irrelevant features on the results. Figure 4 As shown in Figure 2, the multimodal attention module adjusts the weights of different modal features and obtains fused features. For example, the fused features can be expressed as:

[0066]

[0067] in, W represents the fusion of the two modal features m1 and m2. cross(m1,m2) represents the feature fusion transformation between m1 and m2, W m1 (F m1 ) represents the modality-independent weight transformation of m1, W m2 (F m2 ) represents the m2 modality-independent weight transformation, F m1 , F m2 Represent the initial features extracted from the two modes m1 and m2 respectively.

[0068] The technical solution of this embodiment provides a health status classification model based on multimodal data, which may include: a multimodal feature extraction module and a multimodal feature fusion encoder; the multimodal feature fusion encoder includes: a multimodal attention module and a convolution module; the multimodal feature extraction module is communicatively connected to the multimodal feature fusion encoder; the multimodal attention module includes at least two feature heads; wherein the number of the feature heads is the same as the number of modalities of the multimodal data; the multimodal feature extraction module is used to extract data features of each modal health data group respectively, and obtain each modal feature matching each modal health data group; wherein the number of modal features is the same as the number of modal features contained in the modal health data group. The number of modalities is the same; the multimodal attention module is used to store each of the modal features respectively in each feature head, and process the candidate modal features in the target feature head with the reference modal features in other feature heads respectively to obtain the optimized target modal features; the convolution module is used to process each of the target modal features and the first feature after the jump connection of each modal feature through different convolution layers to obtain a fusion feature, which solves the problem that there are certain limitations in determining the user's health status through a single physical examination report of the user, and it is difficult to accurately determine the user's health status. It can identify multimodal user data and provide assistance for accurately and quickly determining the user's health status based on multi-source data.

[0069] Example 2

[0070] Figure 5 This is a flowchart of a method for training a health status classification model based on multimodal data provided in accordance with the second embodiment of the present invention. This embodiment is applicable to the case of training a health classification model based on multimodal data involved in this embodiment. The method can be executed by a training device for a health status classification model based on multimodal data. The training device for a health status classification model based on multimodal data can be implemented in the form of hardware and / or software. The training device for a health status classification model based on multimodal data can be configured in electronic devices such as computers, servers, or tablet computers. Figure 5 As shown, the method includes:

[0071] Step 510: Obtain sample data.

[0072] The sample data includes a modal health data group and label data corresponding to each modal health data group; the same modal health data group includes at least two model data;

[0073] In this embodiment, health status sample data can be obtained, including image samples, audio samples, and text samples; the health status of the health status sample data is annotated to form label data. Among them, image samples include patient facial images and patient medical images (such as CT images, electrocardiograms, etc.); audio samples include recordings of patients' self-descriptions of their health conditions and medical test recordings (such as heartbeats, breathing sounds, and other medical auscultation sounds); text samples include patients' self-written texts about their health conditions, test reports, etc. The above samples come from various tertiary-level A hospitals, physical examination centers, and their corresponding patients.

[0074] Optionally, after obtaining the above samples, professional doctors will conduct multiple rounds of labeling to obtain label data. The label contains five levels of health status, namely healthy, sub-healthy, false healthy, sick, and seriously ill. Among them, health: refers to a state in which both the body and mind are in good condition, without disease, and can effectively cope with various needs of daily life; sub-health: a state between health and disease, manifested by some non-specific symptoms, such as fatigue, insomnia, etc., but has not yet reached the diagnostic criteria for the disease; false health: although there are no obvious symptoms, there may be problems such as abnormal blood lipids, high blood sugar, and blood pressure fluctuations. The psychological state is under high pressure for a long time, which may lead to psychological problems such as anxiety and depression; sick: including symptoms (such as fever, cough, pain) and signs (such as jaundice, accelerated heart rate); seriously ill: a serious disease state that poses a greater threat to life, including serious abnormalities in body temperature, blood pressure, pulse, respiratory rate, etc., and simultaneous dysfunction of multiple important organs such as the heart, lungs, liver, and kidneys.

[0075] In an optional implementation of this embodiment, after obtaining the above sample data, the obtained health status sample data may be further normalized. For example, the normalization calculation method may be:

[0076]

[0077] Among them, X represents the original data, μ represents the mean of the data, σ represents the standard deviation of the data, and X norm Represents the normalized data.

[0078] It should be noted that, in this embodiment, the sample data is obtained only after authorization by the user, and the acquisition method is reasonable and legal.

[0079] Step 520: Input the sample data into a health status classification model based on multimodal data for iterative training. When an iteration stopping condition is met, a health status classification model based on multimodal data for classifying health status is obtained.

[0080] The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any of the above embodiments.

[0081] Among them, the iteration stopping condition can be an accuracy condition, for example, stopping training when the training accuracy reaches 99.99%; it can also be an iteration number condition, for example, stopping training when the iteration reaches 100,000 times, 200,000 times or 500,000 times, which is not limited in this embodiment.

[0082] Optionally, in this embodiment, after obtaining each sample data, each sample data can be input into the health status classification model based on multimodal data in the embodiment of the present invention for iterative training. When the iteration stopping condition is met, the trained (i.e., model parameter optimized) health status classification model based on multimodal data is output.

[0083] Optionally, in an example of this embodiment, health status sample data and label data are input into the classification network model, and image samples, audio samples, and text samples are respectively subjected to feature extraction through convolutional neural networks, Mel spectrum graphs, and long short-term memory networks (LSTMs) to obtain image features, audio features, and text features. The image features, audio features, and text features are input into the multimodal attention module for fusion to obtain fused features. The fused features are sequentially passed through the convolution module, feedforward network, and normalization layer to output unified features; the unified features are input into the fully connected layer, and the fully connected layer outputs the health warning level.

[0084] Optionally, in this embodiment, when training the classification network model, a cross entropy loss function may be used to train the model. Specifically, the loss function may be:

[0085]

[0086] Among them, x i represents the i-th sample, n represents the total number of samples, q(x i ) represents the sample x i The predicted probability, p(x i ) represents the sample x i The true label.

[0087] The solution of this embodiment obtains sample data and inputs the sample data into a health status classification model based on multimodal data for iterative training. When the iteration stopping condition is met, a health status classification model based on multimodal data for classifying health status is obtained. The model parameters of the health status classification model based on multimodal data can be adjusted to obtain a health status classification model with higher accuracy, thereby providing a basis for subsequent rapid prediction of the user's health status.

[0088] Example 3

[0089] Figure 6 This is a flowchart of a health status classification method based on multimodal data provided in accordance with the third embodiment of the present invention; this embodiment is applicable to the case where a user's health status is predicted by a trained health classification model based on multimodal data. The method can be executed by a health status classification device based on multimodal data. The health status classification device based on multimodal data can be implemented in the form of hardware and / or software. The health status classification device based on multimodal data can be configured in electronic devices such as computers, servers or tablet computers. Figure 6 As shown, the method includes:

[0090] Step 610: Obtain target health data of at least two modalities of the target user.

[0091] Optionally, the target user may be any user who does not belong to the user of the sample data obtained above.

[0092] Optionally, the target health data may include the target user's physical examination report, a doctor's diagnosis record, or monitoring data of the user by different sensors, such as the user's exercise data, diet data, or monitoring data of indicators such as blood sugar, etc., which are not limited in this embodiment.

[0093] Step 620: Input the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user.

[0094] Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data described in any embodiment of the present invention, and is trained by the training method of the health status classification model based on multimodal data described in any embodiment of the present invention.

[0095] Optionally, in this embodiment, after obtaining target health data of at least two modalities about the target user, the acquired target health data can be further input into a health status classification model based on multimodal data trained based on the above embodiment, and the target health data can be processed by the health status classification model based on multimodal data to obtain the health status classification result of the target user.

[0096] Optionally, in this embodiment, the target health data is input into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user, which may include: extracting each target health data modal feature of the target health data through the multimodal feature extraction module; storing each target health data modal feature in each feature head respectively through the multimodal attention module, and processing the candidate health data modal feature in the target feature head with the reference health data modal features in other feature heads respectively to obtain the optimized first target health data modal feature; and The accumulation layer processes the first feature after the jump connection between each first target health data modal feature and each target health data modal feature to obtain the fused health data modal feature; the feedforward network extracts the second feature after the jump connection between the fused health data modal feature and the first target health data modal feature to obtain the target health data fusion feature; the normalization layer processes the third feature after the jump connection between the target health data fusion feature and the fused health data modal feature to obtain the unified health data feature; the fully connected layer maps the unified health data feature to the health warning level, and outputs the health warning level.

[0097] In an optional implementation of this embodiment, after the target health data is input into the health status classification model based on multimodal data, each target health data modal feature of the target health data can be first extracted through the multimodal feature extraction module; further, the multimodal attention module in the multimodal fusion feature encoder can be used to store each target health data modal feature in each feature head, and the candidate health data modal features in the target feature head are processed with the reference health data modal features in other feature heads to obtain the optimized first target health data modal features; further, the different convolution layers of the convolution module can be used to process each first target health data modal feature. The first feature after the target health data modal feature and each target health data modal feature are jump-connected is processed to obtain a fused health data modal feature; further, the second feature after the fused health data modal feature and the first target health data modal feature are jump-connected can be extracted through a feedforward network to obtain a target health data fusion feature; further, the third feature after the target health data fusion feature and the fused health data modal feature are jump-connected can be processed through a normalization layer to obtain a unified health data feature; finally, the unified health data feature can be mapped to a health warning level through a fully connected layer, and the health warning level can be output.

[0098] The solution of this embodiment obtains target health data of at least two modalities of the target user; inputs the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user, which can quickly and accurately determine the health level of the target user and provide assistance to the target user in subsequent life, diet and other aspects.

[0099] Example 4

[0100] Figure 7 Schematic diagram of a training device for a health status classification model based on multimodal data according to the fourth embodiment of the present invention. Figure 7 As shown, the device includes: a sample data acquisition module 710 and a training module 720.

[0101] The sample data acquisition module 710 is configured to acquire sample data; the sample data includes a modal health data group and label data corresponding to each modal health data group; and the same modal health data group includes at least two model data.

[0102] A training module 720 is configured to input the sample data into a health status classification model based on multimodal data for iterative training, and obtain a health status classification model based on multimodal data for classifying health status when an iteration stop condition is satisfied;

[0103] The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of the embodiments of the present invention.

[0104] The solution of this embodiment obtains sample data through a sample data acquisition module; the sample data includes modal health data groups and label data corresponding to each modal health data group; the same modal health data group contains at least two model data; the sample data is input into a health status classification model based on multimodal data through a training module for iterative training, and when the iteration stop condition is met, a health status classification model based on multimodal data for classifying the health status is obtained, and the model parameters of the health status classification model based on multimodal data can be adjusted to obtain a health status classification model with higher accuracy, providing a basis for subsequent rapid prediction of the user's health status.

[0105] The training device for the health status classification model based on multimodal data provided in an embodiment of the present invention can execute the training method for the health status classification model based on multimodal data provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0106] Example 5

[0107] Figure 8 is a structural diagram of a health status classification device based on multimodal data according to a fifth embodiment of the present invention; Figure 8 As shown, the device includes: a health data acquisition module 810 and a health warning level determination module 820.

[0108] The health data acquisition module 810 is used to acquire target health data of at least two modalities of the target user;

[0109] The health warning level determination module 820 is configured to input the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user;

[0110] Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data described in any embodiment of the present invention, and is trained by the training method of the health status classification model based on multimodal data described in any embodiment of the present invention.

[0111] The solution of this embodiment obtains target health data of at least two modalities of the target user through the health data acquisition module; inputs the target health data into the health status classification model based on multimodal data for processing through the health warning level determination module to obtain the health warning level of the target user, which can quickly and accurately determine the health level of the target user and provide assistance to the target user in subsequent life, diet and other aspects.

[0112] In an optional implementation of this embodiment, the health warning level determination module is specifically configured to:

[0113] Extracting each target health data modality feature of the target health data by the multimodal feature extraction module;

[0114] The multimodal attention module stores each of the target health data modal features in each feature head, and processes the candidate health data modal features in the target feature head with the reference health data modal features in other feature heads to obtain an optimized first target health data modal feature;

[0115] Processing each first target health data modal feature and the first feature after jump connection of each target health data modal feature through different convolutional layers of the convolution module to obtain a fused health data modal feature;

[0116] Extracting the second feature after the fusion health data modality feature and the first target health data modality feature are jump-connected through a feedforward network to obtain the target health data fusion feature;

[0117] Processing the third feature after the jump connection between the target health data fusion feature and the fused health data modality feature through a normalization layer to obtain a unified health data feature;

[0118] The unified health data features are mapped to a health warning level through a fully connected layer, and the health warning level is output.

[0119] The health status classification device based on multimodal data provided in an embodiment of the present invention can execute the health status classification method based on multimodal data provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0120] In the technical solutions of the embodiments of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of multimodal data (such as physical examination report data) involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0121] Example 6

[0122] Figure 9 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0123] like Figure 9As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0124] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0125] The processor 11 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a training method for a health status classification model based on multimodal data, or a health status classification method based on multimodal data; wherein the training method for a health status classification model based on multimodal data includes: obtaining sample data; the sample data includes a modal health data group and label data corresponding to each of the modal health data groups; the same modal health data group contains at least two model data; the sample data is input into a health status classification model based on multimodal data for iterative training, and when the iteration stop condition is met, a health status classification model based on multimodal data for classifying health status is obtained. A health status classification method based on multimodal data includes: obtaining target health data of at least two modalities of a target user; inputting the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user.

[0126] In some embodiments, the training method of the health state classification model based on multimodal data, or the health state classification method based on multimodal data may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, the training method of the health state classification model based on multimodal data, or one or more steps of the health state classification method based on multimodal data described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the training method of the health state classification model based on multimodal data, or the health state classification method based on multimodal data, by any other appropriate means (for example, by means of firmware).

[0127] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0132] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0133] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0134] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

[0135] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the database detection method provided in any embodiment of the present application.

[0136] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.

[0138] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A health status classification model based on multimodal data, characterized in that: The health status classification model based on multimodal data includes: a multimodal feature extraction module and a multimodal feature fusion encoder; the multimodal feature fusion encoder includes: a multimodal attention module and a convolution module; the multimodal feature extraction module is communicatively connected to the multimodal feature fusion encoder; the multimodal attention module includes at least two feature heads; wherein the number of the feature heads is the same as the number of modalities of the multimodal data; The multimodal feature extraction module is used to extract data features of each modal health data group respectively to obtain modal features matching each modal health data group; wherein the number of modal features is the same as the number of modalities included in the modal health data group; The multimodal attention module is used to store each of the modal features in each feature head respectively, and process the candidate modal features in the target feature head with the reference modal features in other feature heads respectively to obtain the optimized target modal features; The convolution module is used to process each of the target modal features and the first features after jump connection of each modal feature through different convolution layers to obtain a fusion feature.

2. The health status classification model based on multimodal data according to claim 1, characterized in that: The multimodal feature fusion encoder further includes: a feedforward network and a normalization layer; the feedforward network is communicatively connected to the convolution module and the normalization layer respectively; The feedforward network is used to extract the second feature after the fusion feature is jump-connected with each of the target modal features to obtain the target fusion feature; The normalization layer is used to process the target fusion feature and the third feature after the fusion feature is jump-connected to obtain a unified feature.

3. The health status classification model based on multimodal data according to claim 2, characterized in that: The health status classification model based on multimodal data further includes: a fully connected layer; the fully connected layer is communicatively connected to the multimodal feature fusion encoder; The fully connected layer is used to map the unified features to the health warning level.

4. The health status classification model based on multimodal data according to claim 1, characterized in that: The modality health data set includes at least two of the following modality data: Image data, video data, audio data, text data, sensor data, tactile data, biometric data, and social network data.

5. A training method for a health status classification model based on multimodal data, characterized in that: include: Acquire sample data; the sample data includes a modal health data group and label data corresponding to each modal health data group; the same modal health data group includes at least two model data; Inputting the sample data into a health status classification model based on multimodal data for iterative training, and obtaining a health status classification model based on multimodal data for classifying health status when an iteration stop condition is met; The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of claims 1 to 4.

6. A health status classification method based on multimodal data, characterized in that: include: Acquire target health data of at least two modalities of a target user; Inputting the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user; Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of claims 1-4, and is trained by the training method of the health status classification model based on multimodal data described in claim 5.

7. The health status classification method based on multimodal data according to claim 6, characterized in that: The step of inputting the target health data into the health status classification model based on multimodal data for processing to obtain the health warning level of the target user includes: Extracting each target health data modality feature of the target health data by the multimodal feature extraction module; The multimodal attention module stores each of the target health data modal features in each feature head, and processes the candidate health data modal features in the target feature head with the reference health data modal features in other feature heads to obtain an optimized first target health data modal feature; Processing each first target health data modal feature and the first feature after jump connection of each target health data modal feature through different convolutional layers of the convolution module to obtain a fused health data modal feature; Extracting the second feature after the fusion health data modality feature and the first target health data modality feature are jump-connected through a feedforward network to obtain the target health data fusion feature; Processing the third feature after the jump connection between the target health data fusion feature and the fused health data modality feature through a normalization layer to obtain a unified health data feature; The unified health data features are mapped to a health warning level through a fully connected layer, and the health warning level is output.

8. A training device for a health status classification model based on multimodal data, characterized in that: include: A sample data acquisition module is used to acquire sample data; the sample data includes a modal health data group and label data corresponding to each modal health data group; the same modal health data group contains at least two model data; a training module, configured to input the sample data into a health status classification model based on multimodal data for iterative training, and obtain a health status classification model based on multimodal data for classifying health status when an iteration stopping condition is satisfied; The structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of claims 1 to 4.

9. A health status classification device based on multimodal data, characterized in that: include: A health data acquisition module, configured to acquire target health data of a target user in at least two modalities; a health warning level determination module, configured to input the target health data into the health status classification model based on multimodal data for processing, and obtain the health warning level of the target user; Among them, the structure of the health status classification model based on multimodal data is the health status classification model based on multimodal data as described in any one of claims 1-4, and is trained by the training method of the health status classification model based on multimodal data described in claim 5.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the health status classification model based on multimodal data as described in claim 5, or the health status classification method based on multimodal data as described in claims 6-7.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable the processor to implement the training method of the health status classification model based on multimodal data according to claim 5, or the health status classification method based on multimodal data according to claims 6-7 when executed.

12. A computer program product, comprising a computer program, which, when executed by a processor, implements the training method for a health status classification model based on multimodal data according to claim 5, or the health status classification method based on multimodal data according to claims 6-7.