Equipment invariance enhanced multi-modal deep learning model and application thereof

By jointly modeling cough audio and text information using a multimodal deep learning model, the problems of device differences and multi-label recognition are solved, achieving high accuracy and robust diagnosis of respiratory diseases, which is suitable for community and remote diagnosis scenarios.

CN120954690AActive Publication Date: 2025-11-14BOJIANG LIFE SCI (SHANGHAI) CO LTD

Patent Information

Application Number
CN202511485464.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-14
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing technologies for respiratory disease screening suffer from audio data distribution drift due to equipment differences, insufficient model stability and generalization ability, and single-modal information is insufficient to fully understand patient information. Traditional models cannot effectively handle multi-label pathological structures.

Method used

We employ a device-invariant enhanced multimodal deep learning model that jointly models cough audio and structured text information. We use the Transformer architecture to extract audio and text features and fuse them through a dual-path cross-modal multi-head attention mechanism. Combined with a device adversarial branch module and an adaptive loss function, we improve the model's generalization performance and multi-label recognition capability.

Benefits of technology

It achieves high-accuracy recognition capabilities across different devices and multi-label scenarios, enhances the robustness and generalization performance of the model, adapts to disease comorbidities in real clinical settings, and supports community and remote diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954690A_ABST
    Figure CN120954690A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment invariance enhanced multi-modal deep learning model and application thereof, and the model comprises an input module, a coding module, a modal fusion module, a classification module, an equipment adversarial branch module, an output module and a final total loss function. Extracting features through a coding module to obtain audio features and text features; after the audio features and the text features are processed by the modal fusion module to obtain final joint representation, the classification module processes the final joint representation to obtain classification results and corresponding probabilities, and the classification results and the corresponding probabilities are output by the output module; the equipment confrontation branch module enables an audio encoder to confront an equipment classification head in a training stage; and finally, introducing the total loss function in a training stage to optimize the model. The model provided by the invention has the capability of realizing high-accuracy identification on various respiratory system diseases on the premise of not depending on acquisition equipment of a specific brand, and shows excellent generalization performance and robustness in multi-equipment and multi-center data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence-assisted clinical diagnosis technology, specifically relating to a multimodal deep learning model with enhanced device invariance and its application. Background Technology

[0002] Respiratory diseases are a major public health problem contributing to rising morbidity and mortality worldwide, posing a significant health burden, especially among adults and the elderly. Conditions such as chronic obstructive pulmonary disease (COPD), interstitial lung disease, chronic bronchitis, and various infectious diseases often present with nonspecific symptoms in their early stages, including cough, sputum production, and shortness of breath. Cough is one of the most common and earliest clinical signs. Traditional respiratory disease screening relies on lung imaging, pulmonary function tests, or auscultation by physicians. These methods typically require sophisticated equipment and personnel, making widespread coverage in communities or resource-constrained areas difficult. In recent years, with the development of artificial intelligence and acoustic modeling technologies, automated cough sound recognition methods have gradually emerged as a non-invasive, low-cost, highly accessible, and remotely deployable alternative, considered an important direction for achieving early screening and intelligent assisted diagnosis of respiratory diseases.

[0003] Despite initial progress in cough sound analysis, current research faces several challenges that limit the widespread application of related models in real-world scenarios. First, audio data distribution drift caused by device differences severely impacts model stability and generalization ability. In practical applications, cough sounds may be collected from devices of various brands, models, and microphone distances. If the model lacks device invariance, performance degradation or even failure is likely. Second, most existing methods rely solely on single-modal audio information, neglecting crucial diagnostic information from patient demographics (such as age, gender, and body type) and symptom descriptions (such as cough, fever, and dyspnea), thus limiting the model's comprehensive understanding and diagnostic capabilities. Furthermore, in clinical practice, patients often have multiple comorbid respiratory diseases, but traditional models, employing single-label classification strategies, cannot effectively model complex multi-label pathological structures, resulting in insufficient adaptability to real-world clinical problems. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention provides a device invariance-enhanced multimodal deep learning model and its application, which is used to jointly model cough audio and structured text information to achieve robust recognition of respiratory diseases.

[0005] To solve the above-mentioned technical problems and achieve the above-mentioned technical effects, the present invention is implemented through the following technical solution: A device-invariant enhanced multimodal deep learning model includes: The input module is responsible for receiving the cough audio tensor, which has been preprocessed and converted into a Mel spectrum, as well as text information including symptom descriptions and demographic information, and converting it into vector representations to obtain audio vectors and text vectors. The encoding module is responsible for using two independent audio encoders and text encoders to extract features from the audio vector and the text vector respectively, thereby obtaining audio features and text features. The modality fusion module is responsible for using a dual-path cross-modal multi-head attention mechanism to sequentially map, exchange, compress, concatenate, and fuse the audio features and the text features to achieve deep multimodal interaction and obtain the final joint representation. The classification module is responsible for using the trained disease classification head to perform a multi-label respiratory disease classification task on the final joint representation, and to identify the classification results and corresponding probabilities. The device adversarial branch module is responsible for introducing a gradient inversion strategy during the training phase, enabling the audio encoder to learn device-independent representations in the process of adversarializing the device classification head, thereby reducing the interference of device information and improving the generalization performance of multi-label respiratory disease classification tasks. The output module is responsible for outputting the classification results obtained from the recognition and the corresponding probabilities; The final total loss function is introduced during the model training phase and uses a bi-branch multi-objective optimization system consisting of device invariance loss and disease joint loss to evaluate and optimize model performance.

[0006] Furthermore, the patient's cough audio was collected using acquisition devices including smartphones, remote microphones, or wearable devices.

[0007] Furthermore, the audio encoder uses a Transformer architecture to extract frame-level features from the audio vector, with an output dimension of [batch, T, d], which represents the audio features; Furthermore, the text encoder uses a multi-layer Transformer to perform contextual modeling on the text vector (word embedding vector) and outputs the semantic representation of each word, i.e., the text feature.

[0008] Furthermore, the dual-path cross-modal multi-head attention mechanism of the modality fusion module includes two linear layers, an audio-to-text multi-head attention layer, a text-to-audio multi-head attention layer, two attention pooling layers, a pooling feature connection layer, and a multimodal fusion layer; The two linear layers respectively map the audio features and text features input to the modality fusion module into audio projections and text projections; the audio projections and text projections are in the same embedding space, ensuring embedding alignment between different modalities; The audio-to-text multi-head attention layer uses the audio projection as the query vector and the text projection as the key vector and value vector. Through the synergistic effect of the three, feature transformation is performed to obtain the audio projection after multimodal interaction. The text-to-audio multi-head attention layer uses the text projection as the query vector and the audio projection as the key vector and value vector. Through the synergistic effect of the three, feature transformation is performed to obtain the text projection after multimodal interaction. The two attention pooling layers compress the audio projection and the text projection after multimodal interaction, respectively, to obtain audio pooling features and text pooling features. The pooling feature connection layer concatenates the audio pooling features and the text pooling features; The multimodal fusion layer performs multimodal fusion on the concatenated audio pooling features and the text pooling features to generate the final joint representation after multimodal interaction.

[0009] Furthermore, the disease classification head in the classification module adopts a multi-layer MLP structure, including a normalization layer, a regularization layer, a GELU activation function, and a linear layer, to adapt to multi-label respiratory system disease classification tasks and output a probability distribution with dimensions [batch, num_labels].

[0010] Furthermore, the device adversarial branch module includes a self-attention pooling layer, a gradient inversion layer, and a device classification head; wherein, During the training phase, when the audio features extracted by the audio encoder are input into the device adversarial branch module, the self-attention pooling layer first aggregates the audio features and then sends them to the gradient inversion layer. During forward propagation, the gradient inversion layer directly sends the aggregated audio features to the device classification head for the classification task of the acquisition device. At the same time, during backward propagation, the gradient inversion layer reverses the gradient sign, causing the audio encoder to learn feature representations that are opposite to the target of the device classification head during training, that is, to learn audio representations that are independent of the acquisition device. After being aggregated by the self-attention pooling layer, these features are sent to the device classification head through the gradient inversion layer for the device classification task. This allows the optimization target of the audio encoder to form an adversarial relationship with the device classification head, thereby mitigating the distribution offset between different acquisition devices and improving the model's generalization ability to different acquisition devices.

[0011] Furthermore, in designing the final total loss function, a bi-branch multi-objective optimization system composed of the device invariance loss and the disease joint loss is constructed to suppress the distribution shift caused by different acquisition devices between training and deployment scenarios, thereby enhancing the model's generalization ability to unstructured environmental changes; wherein, The device invariance loss includes an invariance risk minimization term and a device classification loss. The weighted combination of the two forms the device invariance loss, which improves the model's generalization ability to different acquisition devices. The joint disease loss includes disease classification loss, audio-text contrast loss, and modality alignment loss. These three are adaptively weighted through an introduced uncertainty weighting mechanism and combined to form the joint disease loss, which better integrates audio and text features and finds the correlation between them for respiratory disease classification tasks. The adaptive weighting is achieved through learnable log-variance parameters to dynamically integrate the importance of different subtasks. During training, a linear warmup training strategy is adopted to ensure that the disease classification task is learned first, so as to avoid strong interference from the device adversarial task in the early stage of training and prevent the audio encoder from being misled in the early stage of training. After the disease classification task is stable, the device adversarial task is gradually introduced to constrain it. The final optimization goal is a comprehensive goal of device robustness and disease recognition. The device invariance loss and the disease joint loss are weighted and summed to form the final total loss function, and the hyperparameters are optimized by grid search.

[0012] Furthermore, when designing the invariance loss of the aforementioned device, 1) The invariant risk minimization term, i.e. the IRM penalty term, is based on the Invariant Risk Minimization theory and introduces a gradient penalty term after the device environment is divided, so as to encourage the model to learn the shared feature space in different device environments. The specific approach is to divide the sub-environment according to the acquisition device, calculate the main task loss gradient in each sub-environment, and minimize the variability of the main task gradient in each sub-environment to enhance the consistency of shared features among acquisition devices. 2) The device classification loss adopts the standard cross-entropy loss, which is introduced during device recognition training. During device recognition training, the device adversarial branch module uses a gradient inversion layer to connect the audio encoder and the device classification head, prompting the audio encoder to generate indistinguishable device representations, and adversarially training the audio encoder to confuse the device recognition ability of the device classification head. 3) Apply sigmoid dynamic gradient adjustment to the invariant risk minimization term to avoid accuracy degradation of IRM due to early addition; 4) The device invariance loss is composed of the sigmoid dynamic gradient adjustment, the invariant risk minimization term, and the device classification loss through a weighted combination, thereby improving the model's generalization ability to different acquisition devices.

[0013] Furthermore, when designing the combined loss for the aforementioned diseases, 1) The disease classification loss adopts binary cross-entropy loss with Logits or cross-entropy loss. The binary cross-entropy loss with Logits corresponds to mixed training or multi-label tasks, and the cross-entropy loss corresponds to single training or single-label tasks. 2) The audio-text contrast loss is calculated using the cosine similarity between the audio embedding α and the text embedding t, and consists of two parts: positive sample pair loss and negative sample pair loss. The objective of the positive sample pair loss is to maximize the similarity of positive sample pairs; The objective of the negative sample pair loss is to minimize the similarity of negative sample pairs, and to separate positive and negative samples using margin m. The audio-text contrast loss is a weighted sum of the positive sample pair loss and the negative sample pair loss, which can improve the consistency of patient multimodal representation. 3) The modal alignment loss adopts the Euclidean distance alignment loss of the fused modal features, which includes two parts: cosine similarity loss and mean square error loss. The goal of the cosine similarity loss is to align the audio and text embeddings so that their semantic similarity is as close to 1 as possible. The mean squared error loss further optimizes the alignment, minimizing the Euclidean distance between the audio and text embeddings; The modality alignment loss is a weighted sum of the cosine similarity loss and the mean square error loss. 4) The joint disease loss is composed of the disease classification loss, the audio-text comparison loss, and the modality alignment loss through adaptive weighting. The joint disease loss is weighted by the disease classification loss, the audio-text comparison loss, and the modality alignment loss through uncertainty weight parameters σ1, σ2, and σ3, all with an initial weight of 1, and includes a regularization term to stabilize training. During training, σ1, σ2, and σ3 are optimized along with other parameters of the model using gradient descent, as follows: Step 1. Calculate the forward propagation loss: Calculate each loss term based on the current model's predictions, and weight each loss term based on the current values ​​of σ1, σ2, and σ3, and calculate the total loss. Step 2. Calculate the gradient: The backpropagation process will calculate the gradient for each σ by taking derivatives; Step 3. Update σ; Step 4. The optimizer updates the σ value of each value using gradient descent based on the current learning rate η; Step 5. Iterative optimization: As training iterates, the model dynamically adjusts the values ​​of σ1, σ2, and σ3 according to the actual data distribution to balance the various losses; By optimizing the overall joint loss for the aforementioned diseases, the model can simultaneously improve in three aspects: classification, contrastive learning, and modality alignment, and dynamically adjust the weights of each part of the loss, ultimately achieving performance improvement.

[0014] An AI-assisted diagnostic system for respiratory diseases, built using a multimodal deep learning model enhanced with the aforementioned device invariance.

[0015] The beneficial effects of this invention are as follows: In constructing a device-invariant enhanced multimodal deep learning model, this invention first jointly models cough audio and structured text information to mine complementary information contained in different data modalities. Then, by introducing a device adversarial module into the audio encoder structure, it uses a gradient inversion strategy to counter the device classification head, learns device-independent audio representations, and improves device recognition performance. Furthermore, by constructing a novel joint loss function and further introducing an invariant risk minimization strategy on this basis, it effectively suppresses the distribution shift caused by different acquisition devices between training and deployment scenarios, further enhancing the model's generalization ability to unstructured environmental changes. In addition, the constructed multi-label learning framework can simultaneously identify multiple respiratory diseases to adapt to the prevalence of disease comorbidity in real-world scenarios.

[0016] Therefore, the model of this invention has the ability to accurately identify a variety of respiratory diseases without relying on specific brand data collection devices, and exhibits excellent generalization performance and robustness in multi-device, multi-center data. The model of this invention not only has strong adaptability in clinical data, but also provides strong support for its deployment in community health service stations, remote consultation platforms, and edge intelligent devices.

[0017] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the invention and to implement it according to the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Specific embodiments of the present invention are given in detail below with reference to the accompanying drawings. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is an overall framework diagram of the multimodal deep learning model for enhancing the invariance of the device in this invention; Figure 2 This is a schematic diagram of the internal flow of the audio→text multi-head attention layer and the text→audio multi-head attention layer in the model of this invention; Figure 3This is a schematic diagram of the internal flow of the attention pooling layer in the model of this invention; Figure 4 This is a schematic diagram of the internal process of the multimodal fusion layer in the model of this invention; Figure 5 This is a schematic diagram of the internal flow of the disease classification head in the model of this invention; Figure 6 This is a schematic diagram of the internal process of the self-attention pooling layer in the model of this invention; Figure 7 This is a schematic diagram of the internal process of the equipment classification head in the model of this invention; Figure 8 This is a flowchart illustrating the calculation of the final total loss function in the model of this invention; Figure 9 This is a schematic diagram of the sigmoid dynamic gradient adjustment used when training the model in this invention; Figure 10 This is a schematic diagram illustrating the dynamic adjustment of saturation in the linear warmup training strategy used when training the model in this invention. Figure 11 The figure shows the comparative experimental results of the model of this invention and several existing models in multi-label tasks; where Figure 11 In the diagram, 'a' represents the AUROC index. Figure 11 In this diagram, 'b' represents the AUPRC index. Figure 12 The figure shows the experimental results of modal ablation of the model of this invention; Figure 13 The figure shows the results of the adversarial / non-adversarial comparison experiment of the model of this invention. Detailed Implementation

[0019] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings to provide a clearer understanding of the invention's purpose, features, and advantages. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of the invention, but are merely illustrative of the essential spirit of the invention's technical solution.

[0020] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known apparatuses, structures, and techniques associated with this application may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.

[0021] Unless the context requires otherwise, throughout the specification and claims, the word “comprising” and its variations, such as “including” and “having”, shall be understood to have an open, inclusive meaning, that is, to be interpreted as “including, but not limited to”.

[0022] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.

[0023] The singular forms “a” and “the” used in this specification and the appended claims include plural references unless otherwise expressly stated herein. It should be noted that the term “or” is generally used to mean “and / or” unless otherwise expressly stated herein.

[0024] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0025] See Figure 1 As shown, this invention proposes a device-invariant enhanced multimodal deep learning model for jointly modeling cough audio and structured text information to achieve robust identification of respiratory diseases. The model consists of six core modules: an input module, an encoding module, a modality fusion module, a classification module, an output module, and a device adversarial branch module.

[0026] See Figure 1 As shown, the input module is responsible for receiving two types of modalities: one is the cough audio tensor, which is preprocessed and converted into a Mel spectrum, and the other is text information, including symptom description and demographic information. The input module also converts the cough volume tensor and text information into vector representations to obtain audio vectors and text vectors.

[0027] In terms of audio, since the multimodal deep learning model of the present invention has the ability to adversarially model differences in audio devices, the cough audio can support data input from acquisition devices such as smartphones, remote microphones or wearable devices, and can be flexibly adapted to different acquisition environments, terminal forms and medical infrastructure.

[0028] In terms of text, in addition to the two modalities of demographic information and symptom questionnaire information, the text information may also include other medical data modalities, such as vital signs such as body temperature, heart rate, and blood oxygen saturation (SpO2).

[0029] See Figure 1 As shown, the encoding module includes an audio decoder and a text decoder. The audio decoder is responsible for extracting audio features from the audio vector, and the text decoder is responsible for extracting text features from the text vector.

[0030] Preferably, the audio encoder uses a Transformer architecture to extract frame-level features from the audio vector, with an output dimension of [batch, T, d], which represents the audio features. The audio encoder includes, but is not limited to, AudioSpectogram Transformer (AST), ATST, BEAts, CLAP, Efficient Audio Transformer, SpeechT5, Wav2Vec and its variants, and Whisper.

[0031] Preferably, the text encoder uses a multi-layer Transformer to perform contextual modeling on the text vector (word embedding vector), outputting the semantic representation of each word, i.e., the text features. The text encoder includes, but is not limited to, BERT, Mamba2, Phi-3, PhiMoE, PhoBERT, Refomer, RemBERT, RetriBERT, RoBERTa, RoCBert, Roformer, T5, Zamba, etc.

[0032] The modality fusion module is responsible for using a dual-path cross-modal multi-head attention mechanism to sequentially map, exchange, compress, concatenate, and fuse the audio features and the text features to achieve deep multimodal interaction and obtain the final joint representation.

[0033] See Figure 1 As shown, the dual-path cross-modal multi-head attention mechanism of the modality fusion module includes two linear layers, an audio-to-text multi-head attention layer, a text-to-audio multi-head attention layer, two attention pooling layers, a pooling feature connection layer, and a multimodal fusion layer.

[0034] The two linear layers respectively map the audio features and text features input to the modality fusion module to audio projection and text projection, respectively; the audio projection and the text projection are in the same embedding space, ensuring embedding alignment between different modalities.

[0035] The audio-to-text multi-head attention layer uses audio_project as the query vector, text_project as the key vector and value vector, and performs feature transformation through the synergistic effect of the three to obtain the audio_project after multimodal interaction.

[0036] The text-to-audio multi-head attention layer uses text_project as the query vector, audio_project as the key vector and value vector, and performs feature transformation through the synergistic effect of the three to obtain the text_project after multimodal interaction.

[0037] See Figure 2 As shown, both the audio-to-text multi-head attention layer and the text-to-audio multi-head attention layer are supplemented with additional regularization (Dropout) and layer normalization (LayerNorm) on the basis of their respective multi-head attention layers.

[0038] The audio-to-text interaction process is as follows: When text_project is input as the key and value and audio_project as the query into the audio-to-text multi-head attention layer, the multi-head attention layer inside performs feature transformation through the synergistic effect of the three to achieve deep multimodal interaction; the attention output is then regularized, and the regularized output is combined with the query for layer normalization, and finally the audio_project after multimodal interaction is output.

[0039] The text-to-audio interaction process is as follows: When audio_project is input as the key and value and text_project as the query into the text-to-audio multi-head attention layer, the multi-head attention layer inside performs feature transformation through the synergistic effect of the three to achieve deep multimodal interaction; the attention output is then regularized, and the regularized output is combined with the query for layer normalization, and finally the text_project after multimodal interaction is output.

[0040] The interaction flow between Key, Value, and Query is as follows: First, the attention score is calculated between the query and all keys using a dot product or scaled dot product to reflect the relevance strength. Then, the attention score is Softmax normalized to obtain the attention weight. Finally, the attention weight is used to sum the values ​​in a weighted manner to generate the final output.

[0041] The two attention pooling layers compress the audio_project and text_project after multimodal interaction, respectively, to obtain audio pooling features and text pooling features.

[0042] See Figure 3As shown, the process of the attention pooling layer compressing audio_project / text_project is as follows: when the audio_project / text_project after multimodal interaction is input into the corresponding attention pooling layer, it is first processed by the mapping of the linear layer and the attention weights are calculated. Then, after normalization by the Softmax activation function, the attention weights are converted into a probability distribution. Finally, the audio_project / text_project after multimodal interaction is weighted and summed using the probability distribution to obtain a fixed-length representation, that is, the output audio / text pooling features, which are used for cross-modal interaction compression.

[0043] The pooling feature concatenation layer concatenates the audio pooling features and the text pooling features.

[0044] The multimodal fusion layer performs multimodal fusion on the concatenated audio pooling features and the text pooling features to generate the final joint representation after multimodal interaction.

[0045] See Figure 4 As shown, the process of fusing audio pooling features and text pooling features in the multimodal fusion layer is as follows: When the concatenated audio pooling features and text pooling features are input into the multimodal fusion layer, they are first mapped to a multimodal comprehensive projection through a linear layer, then the attention weights are obtained through a multi-head attention mechanism, then regularization is performed, and the regularization output is combined with the input multimodal comprehensive features for layer normalization. Finally, the final joint representation after multimodal interaction is output.

[0046] See Figure 1 As shown, the classification module is responsible for using the trained disease classification head to perform multi-label respiratory disease classification on the final joint representation, and outputting the classification results and corresponding probabilities.

[0047] The disease classification head in the classification module adopts a multi-layer MLP structure (including LayerNorm, Dropout, GELU activation function and linear layer) to adapt to multi-label respiratory disease classification tasks and output a probability distribution with dimensions [batch, num_labels].

[0048] See Figure 5As shown, the process of the multi-label respiratory disease classification task is as follows: when the final joint representation obtained by fusion is input into the disease classification head, it first goes through the mapping process of the first linear layer, then through the feature transformation and regularization process of the first layer using LayerNorm, GELU activation function and Dropout, then through the mapping process of the second linear layer, then through the feature transformation and regularization process of the second layer using LayerNorm, GELU activation function and Dropout, and finally through the mapping process of the third linear layer, thereby mapping the processed high-dimensional features to the classification label space and outputting the final disease classification result and probability distribution.

[0049] See Figure 1 As shown, the device adversarial branch module aims to alleviate the distribution offset between different acquisition devices. It is responsible for introducing a gradient inversion strategy during the training phase, so that the audio encoder learns device-independent representations in the process of adversarializing the device classification head, thereby reducing the interference of device information and improving the generalization performance of the multi-label respiratory disease classification task. Its loss is only introduced during training.

[0050] The device adversarial branch module includes a self-attention pooling layer, a gradient reversal layer (GRL), and a device classification head. During the training phase, when the audio features extracted by the audio encoder are input into the device adversarial branch module, the self-attention pooling layer first aggregates the audio features and then sends them to the GRL. During forward propagation, the GRL directly sends the aggregated audio features to the device classification head for the task of classifying the acquisition devices. Simultaneously, during backward propagation, the GRL reverses the gradient sign, causing the audio encoder to learn feature representations that are opposite to the target of the device classification head during training, i.e., learning audio representations that are independent of the acquisition devices. These representations are then aggregated by the self-attention pooling layer and sent to the device classification head through the corresponding gradient reversal layer for the device classification task. This allows the audio encoder's optimization target to adversarially compete with the device classification head, improving the model's generalization ability to different acquisition devices.

[0051] See Figure 6 As shown, the aggregation process of the audio features is similar to the compression process of audio / text projection, both of which adopt the attention pooling mechanism: when the audio features are input into the self-attention pooling layer, they are first processed by the mapping of the linear layer and the attention weights are calculated. Then, after normalization by the Softmax activation function, the weights are summed to finally obtain a fixed-length representation for internal modality aggregation.

[0052] See Figure 7As shown, the process of the device classification task is as follows: after the aggregated audio features are sent to the device classification head via GRL, they first go through the mapping process of the first linear layer, then through the GELU activation function, and finally through the mapping process of the second linear layer, thereby mapping the processed high-dimensional features to the classification label space and outputting the final collection device classification result and probability distribution.

[0053] See Figure 1 As shown, the output module is responsible for outputting the classification results and corresponding probabilities obtained from the recognition.

[0054] The model of this invention uses a certain amount of multimodal data from past patients and corresponding disease labels and device labels to train the classification of respiratory diseases.

[0055] The multimodal data includes preprocessed cough audio tensors converted to Mel spectrum from past patients and textual information including symptom descriptions and demographic information.

[0056] The disease tags record clinical diagnostic information corresponding to the multimodal data and serve as classification supervision signals to guide the optimization of the disease classification head (classification supervision). These disease tags typically indicate a single respiratory disease type or a combination of multiple tags associated with the patient's past history. Examples include: single disease: asthma, COPD, pneumonia, bronchiectasis; multiple comorbidities: COPD + lung cancer, asthma + allergic rhinitis. This multi-tag structure is used to depict the patient's true clinical condition, overcoming the limitations of traditional "single-tag classification."

[0057] The device tags record the acquisition source information corresponding to multimodal data, used for adversarial learning / invariant representation learning to help the model eliminate domain bias between different acquisition devices (adversarial supervision). These device tags typically indicate the device category from which the audio or text samples originated, such as: mobile phone model (iPhone, Huawei, Xiaomi, etc.); microphone type (professional recording equipment vs. ordinary headphone microphone); clinical data input source (electronic medical record systems of different hospitals). These tags do not represent medical significance, but only reflect differences in data acquisition conditions.

[0058] Both the disease label and the device label are used at the supervisory signal level. That is, the disease label acts on the final joint representation output by the intermediate layer of the model, and the device label acts on the audio features output by the intermediate layer of the model.

[0059] During model training, the disease labels and the final joint representation are input into the disease classification head to directly supervise disease prediction. The role of the disease labels is to serve as a classification supervision signal, directly optimizing the classifier output. In multi-label scenarios, this helps the model learn co-occurrence patterns among diseases, improving its adaptability to complex clinicopathological situations and achieving supervised disease prediction.

[0060] During model training, the device labels and audio features are input into the device classification head, and adversarial constraints are applied to the main representation extractor through the gradient inversion layer (GRL). The device labels are used for adversarial learning or domain-invariant representation learning during training to help the model reduce the impact of different acquisition conditions on the final classification performance, achieve cross-device generalization ability, and eliminate domain bias.

[0061] When the model of this invention is trained and is in the deployment / inference phase, only the multimodal data of the patient to be identified (such as cough audio tensor, clinical text information) needs to be input. No additional device labels are required. The model automatically achieves device-invariant representation extraction through the adversarial constraints learned during the training phase and outputs respiratory disease prediction results.

[0062] During the model training phase, this invention introduces a final total loss function to evaluate and optimize model performance. See also Figure 8 As shown, in the design of the final total loss function, this invention constructs a bi-branch multi-objective optimization system composed of device invariance loss and disease joint loss, thereby effectively suppressing the distribution shift caused by different acquisition devices between training and deployment scenarios and enhancing the model's generalization ability to unstructured environmental changes.

[0063] 1. Loss due to equipment invariability.

[0064] The device invariance loss aims to improve the model's generalization ability to different acquisition devices. It includes an invariance risk minimization term and a device classification loss, which are formed by a weighted combination of the two.

[0065] 1) The Invariant Risk Minimization (IRM) term, or IRM Penalty, is based on the Invariant Risk Minimization (IRM) theory. It introduces a gradient penalty term after dividing the model into device environments to encourage the model to learn shared feature spaces across different device environments. Specifically, it divides the model into sub-environments based on the acquisition device, calculates the main task loss gradient for each sub-environment, and minimizes the variability of the main task gradient in each sub-environment, thus strengthening the consistency of shared features among acquisition devices. The formula for the IRM penalty term is as follows: (1); In equation (1), This represents a collection of training environments, each representing a subset of datasets with different distribution characteristics, such as different devices, different centers, or different patient groups. 'e' indicates a single environment index. ; This represents a feature extraction function (or feature transformation mapping) that maps the original input x to a feature space representation. ; This represents the classifier function, which receives feature representations. And output the prediction results ; The hyperparameter represents the weighting factor of the invariant risk constraint, used to balance empirical risk and environmental invariance constraint. This represents the gradient of the classifier weights with w=1.0 fixed. The smaller the norm of the gradient, the better the feature extractor performs in this environment. The resulting characteristics are more environmentally invariant; among them Let be the empirical risk in environment e, and let be the expected loss of the sample in that environment. It represents the square of the Euclidean norm, used to measure gradient strength.

[0066] 2) The Device Classification Loss uses the standard Cross-Entropy Loss (CE), introduced during device recognition training. During device recognition training, the device adversarial branch module uses GRL to connect the audio encoder and the device classification head, prompting the audio encoder to generate indistinguishable device representations, adversarially training the audio encoder to confuse the device classification head's device recognition capabilities. The formula for the standard cross-entropy loss is as follows: (2); In equation (2), This represents a device classifier function that maps audio features to logits for each device category; This represents the audio representation (embedding, feature vector) from the audio encoder ( / feature extractor). This indicates the device's true label.

[0067] 3) Since the IRM penalty is a gradient penalty term, it is not suitable to add it early on, therefore see [reference needed]. Figure 9 As shown, this invention uses sigmoid dynamic gradient adjustment for the invariant risk minimization term, thus avoiding accuracy decay caused by early inclusion of the IRM; the formula for the sigmoid dynamic gradient adjustment is as follows: (3); In equation (3), t is the current step, and T is the total number of steps, calculated as follows: ; It is the maximum weight for resisting loss; k represents the steepness, which determines the growth rate; It is the normalized training progress, ranging from [0,1].

[0068] 4) As a stability enhancement mechanism, the resistance loss coefficient is gradually adjusted, and... , The equipment invariance loss is calculated by weighted combination, and its loss term is formulated as follows: (4).

[0069] 2. Combined losses due to disease.

[0070] To better integrate audio and text features and identify their correlations for respiratory disease classification tasks, this invention designs a joint disease loss function to evaluate and optimize model performance. This joint disease loss function consists of three parts: disease classification loss, audio-text contrast loss (improving consistency in patient multimodal representations), and modality alignment loss. The contribution of each loss is dynamically adjusted during training using uncertainty weights, thereby defining the total loss value.

[0071] 1) Disease classification loss (L) classification ).

[0072] Depending on the training strategy and the nature of the task, disease classification loss can be either binary cross-entropy loss with Logits or cross-entropy loss. For mixed training or multi-label tasks, the disease classification loss uses binary cross-entropy loss with Logits; for single training or single-label tasks, the disease classification loss uses cross-entropy loss.

[0073] The formula for the binary cross-entropy loss with Logits is as follows: (5); In equation (5), xi It is the model's predicted logits; The target label (needs to be one-hot encoded and converted to floating-point type); , where i is the sigmoid function; N is the total number of samples.

[0074] The formula for the cross-entropy loss is as follows: (6); In equation (6), K is the number of categories; It is the predicted probability of the k-th class after softmax; It is the target one-hot encoding.

[0075] 2) Audio-text contrast loss (L contrastive ).

[0076] The audio-text contrast loss is calculated using the cosine similarity between the audio embedding α and the text embedding t, and consists of two parts: positive sample pair loss and negative sample pair loss. The objective of the positive sample pair loss is to maximize the similarity of positive sample pairs (audio and text semantically consistent), and the formula is as follows: (7); In equation (7), ; N is the number of sample pairs; a i This represents the audio modal input of the i-th sample; t i This represents the text modal input of the i-th sample.

[0077] The objective of the negative sample pair loss is to minimize the similarity between negative sample pairs (where audio and text semantics are inconsistent), and to separate positive and negative samples using margin m, as shown in the following formula: (8); In equation (8), ; N is the number of sample pairs; m is the similarity margin, which is the hyperparameter of the margin between negative samples; a i This represents the audio modal input of the i-th sample; t i This represents the text modal input of the i-th sample.

[0078] The total audio-text contrast loss is a weighted sum of the losses from the positive sample pairs and the negative sample pairs, as shown in the following formula: (9).

[0079] 3) Modal alignment loss (L) alignment ).

[0080] The modality alignment loss employs Euclidean distance alignment loss based on the fused modality features, which combines two parts: cosine similarity loss and mean squared error loss. The goal of the cosine similarity loss is to align the audio and text embeddings so that their semantic similarity is as close to 1 as possible, and the formula is as follows: (10); In equation (10), ; N is the number of sample pairs; a i This represents the audio modal input of the i-th sample; t i This represents the text modal input of the i-th sample.

[0081] The mean squared error loss further optimizes the alignment by minimizing the Euclidean distance between the audio and text embeddings, as shown in the following formula: (11); In equation (11), N is the number of sample pairs; a i This represents the audio modal input of the i-th sample; t i This represents the text modal input of the i-th sample.

[0082] The total modal alignment loss is a weighted sum of the cosine similarity loss and the mean square error loss, and its formula is as follows: (12).

[0083] 4) Disease-related combined loss and uncertainty weights.

[0084] To automatically adjust the weights based on the dynamic characteristics of each loss term, the disease joint loss weights the disease classification loss, audio-text contrast loss, and modality alignment loss with uncertainty weight parameters σ1, σ2, and σ3 (all with an initial weight of 1), and includes a regularization term to stabilize training.

[0085] The total joint disease loss is a weighted sum of three parts: disease classification loss, audio-text contrast loss, and modality alignment loss, as shown in the following formula: (13); In equation (13), Classification loss weights ; Audio-text contrast loss weighting term ; Modal alignment loss weight term .

[0086] During training, σ1, σ2, and σ3 are optimized along with other parameters of the model using gradient descent, as follows: Step 1. Calculate the forward propagation loss: Calculate each loss term based on the current model's predictions. Based on the current values ​​of σ1, σ2, and σ3, the individual loss terms are weighted and the total loss is calculated. ; Step 2. Calculate the gradient: The backpropagation process will differentiate and calculate the gradient for each σ. The calculation formula is as follows: (14); In equation (14), i=1,2,3 correspond to the three loss terms: disease classification loss, audio-text comparison loss, and modality alignment loss. L i This represents the individual loss value for the i-th task; The uncertainty parameter for the i-th task is a trainable scalar that reflects the noise or uncertainty of the task's output; the higher the uncertainty, the smaller the weight. This is represented as the gradient of the total loss with respect to the uncertainty parameter of the i-th task; Step 3. Update σ; Step 4. The optimizer updates the σ value of each step using gradient descent based on the current learning rate η. The calculation formula is as follows: (15); Step 5. Iterative optimization: As training iterates, the model dynamically adjusts the values ​​of σ1, σ2, and σ3 according to the actual data distribution to balance the various losses.

[0087] By optimizing the overall disease joint loss, the model can be optimized simultaneously in classification, contrastive learning, and modality alignment, and the weights of each part of the loss can be dynamically adjusted to ultimately improve performance.

[0088] 3. Final total loss function.

[0089] To ensure that the equipment adversarial task does not significantly interfere with the disease classification task during the initial training phase, see [link to training]. Figure 10 As shown, this invention employs a linear warmup training strategy during training to prevent the audio encoder from being misled in the early stages of training, ensuring that the disease classification task is learned first. Once the disease classification task is stable, adversarial tasks are gradually introduced to constrain the learning process. The linear warmup training strategy is as follows: (16); In equation (16), t is the current step and T is the total number of steps; This indicates the parameters used to adjust the equipment to its maximum resistance level; The step indicates the end of the warm-up phase.

[0090] The ultimate optimization objective is a comprehensive goal of combining device robustness and disease recognition. The device invariance loss and the joint disease loss are weighted and summed to form the final total loss function, and the hyperparameters are optimized using grid search. The formula for the final total loss function is as follows: (17).

[0091] The device-invariant enhanced multimodal deep learning model proposed in this invention achieves highly accurate identification of various respiratory diseases without relying on specific brand data acquisition devices, and demonstrates excellent generalization performance and robustness in multi-device, multi-center data. This not only fully verifies the model's adaptability to clinical data, but also provides strong support for its deployment in community health service stations, remote consultation platforms, and edge intelligent devices. The following are the results of comparative experiments in different tasks: 1. Results of the multi-label respiratory disease identification task.

[0092] In multi-label classification tasks, the research team of this invention compared and evaluated a variety of existing models and the device-invariant enhanced multimodal deep learning model proposed in this invention.

[0093] See Figure 11As shown, on a multicenter, multi-device dataset containing seven types of respiratory diseases (chronic obstructive pulmonary disease (COPD), upper respiratory tract infection (URTI), lower respiratory tract infection (LRTI), interstitial lung disease (ILD), asthma (ASTHMA), bronchiectasis (BRONCHIECTASIS), and chronic bronchitis (CB)), the model of this invention achieved the best results on both AUROC and AUPRC, significantly outperforming the baseline model, demonstrating its strong generalization ability and practical application potential in complex multimodal data.

[0094] 2. Evaluation of the effectiveness of modal ablation and equipment adversarial training.

[0095] To evaluate the role of each modality in multimodal learning, the research team of this invention conducted modal ablation experiments. First, the individual and joint contributions of audio modality, demographic information (text), and symptom description (text) in recognition were evaluated (demographics + symptoms, audio, audio + demographics, audio + symptoms, audio + demographics + symptoms, audio + demographics + symptoms + device adversarial).

[0096] See Figure 12 As shown in the experimental results, the single-modal audio model performs well on the basic classification task, while demographic features and symptom information provide complementary information in multi-label recognition. The complete multimodal model, which integrates audio, demographic information, and symptom descriptions, significantly outperforms models using only single-modal or bimodal information inputs in the AUC metrics across all six tests, demonstrating stronger generalization ability and clinical applicability, highlighting the advantages of modal collaboration in disease identification. This indicates that multi-source medical information collaborative modeling has crucial value for the intelligent identification of respiratory diseases.

[0097] Furthermore, to verify the effectiveness of adversarial training in device generalization, the research team of this invention systematically compared the performance of models using and not using adversarial mechanisms in all tasks (binary classification tasks of different modalities, binary classification tasks of different diseases, and multi-label tasks).

[0098] See Figure 13 As shown, the results consistently demonstrate that the introduction of adversarial strategies significantly improves the AUC index, indicating that the adversarial mechanism can effectively suppress device-related biases and enhance the model's stability and generalization ability in cross-device environments.

[0099] 3. The role of equipment bias robustness and invariant risk minimization strategies.

[0100] To address the distributional differences arising from multi-device audio acquisition, the research team of this invention introduced a device adversarial mechanism and an invariant risk minimization (IRM) strategy into the model. As shown in Table 1, in cross-device migration tests, the model without the adversarial strategy exhibited a significant performance decline. However, after introducing the combined device adversarial and IRM mechanism, the model's classification performance on the target device was significantly improved, demonstrating strong device invariance and deployment stability. Although the device adversarial mechanism using the IRM strategy had a significantly longer iteration time per round under the same conditions, it significantly shortened the number of iterations and enhanced the model's convergence. This strategy is expected to promote the widespread application of intelligent respiratory disease identification models in multi-center clinical settings.

[0101] Table 1. Comparison of adversarial tasks (including IRM), adversarial tasks (excluding IRM), and non-adversarial tasks.

[0102] 4. Performance and Deployment Comparison of Models with Different Parameter Scales To further explore the impact of model parameter size on performance and deployment capabilities, the research team of this invention designed two model versions with approximately twice the number of parameters for comparison: the Base-size model and the Large-size model. The Large-size model has approximately twice the number of parameters as the Base-size model, representing model settings in resource-constrained and computationally abundant scenarios, respectively. As shown in Table 2, experimental results show that the Large-size model achieves an average AUROC of 0.8571 in multi-label tasks, slightly higher than the Base-size model's 0.8405, demonstrating a certain performance gain.

[0103] Table 2 Comparison of Model Parameters

[0104] However, while maintaining near-perfect accuracy, the Base-size model offers significant advantages in terms of lightweight design, with inference latency and memory consumption significantly lower than the Large-size model. This makes it particularly suitable for mobile devices, edge devices, or resource-constrained primary healthcare scenarios. This comparison highlights the importance of model compression and efficient deployment strategies in practical medical AI applications and provides a reference for the design of subsequent edge intelligent diagnostic systems.

[0105] This invention provides an effective solution for constructing AI models of cough sounds for real-world applications, demonstrating the significant value of multimodal fusion and device-robust modeling in the digital screening of adult respiratory diseases. The findings will propel cough sound AI technology from experimental settings to practical clinical applications, particularly in primary healthcare, remote health monitoring, and chronic disease screening for the elderly, showcasing broad applicability and high clinical translational potential. Therefore, the device invariance enhancement multimodal deep learning model proposed in this invention can be applied to build or manufacture AI-assisted diagnostic systems, devices or apparatuses for respiratory diseases, thereby helping to achieve early screening and intelligent assisted diagnosis of respiratory diseases.

[0106] Furthermore, in addition to respiratory diseases, the device invariance-enhanced multimodal deep learning model proposed in this invention can also be applied to other medical audio scenarios, such as abnormal heart sounds, abnormal lung sounds, snoring analysis, early warning of acute asthma attacks, or screening for sleep-disordered breathing.

[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal deep learning model with enhanced device invariance, characterized in that, include: The input module is responsible for receiving the cough audio tensor, which has been preprocessed and converted into a Mel spectrum, as well as text information including symptom descriptions and demographic information, and converting it into vector representations to obtain audio vectors and text vectors. The encoding module is responsible for using two independent audio encoders and text encoders to extract features from the audio vector and the text vector respectively, thereby obtaining audio features and text features. The modality fusion module is responsible for using a dual-path cross-modal multi-head attention mechanism to sequentially map, exchange, compress, concatenate, and fuse the audio features and the text features to achieve deep multimodal interaction and obtain the final joint representation. The classification module is responsible for using the trained disease classification head to perform a multi-label respiratory disease classification task on the final joint representation, and to identify the classification results and corresponding probabilities. The device adversarial branch module is responsible for introducing a gradient inversion strategy during the training phase, enabling the audio encoder to learn device-independent representations in the process of adversarializing the device classification head, thereby reducing the interference of device information and improving the generalization performance of multi-label respiratory disease classification tasks. The output module is responsible for outputting the classification results obtained from the recognition and the corresponding probabilities; The final total loss function is introduced during the model training phase and uses a bi-branch multi-objective optimization system consisting of device invariance loss and disease joint loss to evaluate and optimize model performance.

2. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: The patient's cough audio was collected using acquisition devices including smartphones, remote microphones, or wearable devices.

3. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: In the encoding module The audio encoder uses a Transformer architecture to extract frame-level features from the audio vector, and the output dimension is [batch, T, d], which is the audio feature; The text encoder uses a multi-layer Transformer to perform context modeling on the text vectors and outputs the semantic representation of each text vector, i.e., the text features.

4. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: The dual-path cross-modal multi-head attention mechanism of the modality fusion module includes two linear layers, an audio-to-text multi-head attention layer, a text-to-audio multi-head attention layer, two attention pooling layers, a pooling feature connection layer, and a multimodal fusion layer; The two linear layers respectively map the audio features and text features input to the modality fusion module into audio projections and text projections; the audio projections and text projections are in the same embedding space, ensuring embedding alignment between different modalities; The audio-to-text multi-head attention layer uses the audio projection as the query vector and the text projection as the key vector and value vector. Through the synergistic effect of the three, feature transformation is performed to obtain the audio projection after multimodal interaction. The text-to-audio multi-head attention layer uses the text projection as the query vector and the audio projection as the key vector and value vector. Through the synergistic effect of the three, feature transformation is performed to obtain the text projection after multimodal interaction. The two attention pooling layers compress the audio projection and the text projection after multimodal interaction, respectively, to obtain audio pooling features and text pooling features. The pooling feature connection layer concatenates the audio pooling features and the text pooling features; The multimodal fusion layer performs multimodal fusion on the concatenated audio pooling features and the text pooling features to generate the final joint representation after multimodal interaction.

5. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: The disease classification head in the classification module adopts a multi-layer MLP structure, including a normalization layer, a regularization layer, a GELU activation function, and a linear layer, to adapt to multi-label respiratory disease classification tasks and output a probability distribution with dimensions [batch, num_labels].

6. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: The device adversarial branch module includes a self-attention pooling layer, a gradient inversion layer, and a device classification head; wherein... During the training phase, when the audio features extracted by the audio encoder are input into the device adversarial branch module, the self-attention pooling layer first aggregates the audio features and then sends them to the gradient inversion layer. During forward propagation, the gradient inversion layer directly sends the aggregated audio features to the device classification head for the classification task of the acquisition device. At the same time, during backward propagation, the gradient inversion layer reverses the gradient sign, causing the audio encoder to learn feature representations that are opposite to the target of the device classification head during training, that is, to learn audio representations that are independent of the acquisition device. After being aggregated by the self-attention pooling layer, these features are sent to the device classification head through the gradient inversion layer for the device classification task. This allows the optimization target of the audio encoder to form an adversarial relationship with the device classification head, thereby mitigating the distribution offset between different acquisition devices and improving the model's generalization ability to different acquisition devices.

7. The device invariance enhancement multimodal deep learning model according to claim 1, characterized in that: In designing the final total loss function, a bi-branch multi-objective optimization system composed of the device invariance loss and the disease joint loss is constructed to suppress the distribution shift caused by different acquisition devices between training and deployment scenarios, thereby enhancing the model's generalization ability to unstructured environmental changes; wherein, The device invariance loss includes an invariance risk minimization term and a device classification loss. The weighted combination of the two forms the device invariance loss, which improves the model's generalization ability to different acquisition devices. The joint disease loss includes disease classification loss, audio-text contrast loss, and modality alignment loss. These three are adaptively weighted through an introduced uncertainty weighting mechanism and combined to form the joint disease loss, which better integrates audio and text features and finds the correlation between them for respiratory disease classification tasks. The adaptive weighting is achieved through learnable log-variance parameters to dynamically integrate the importance of different subtasks. During training, a linear warmup training strategy is adopted to ensure that the disease classification task is learned first, so as to avoid strong interference from the device adversarial task in the early stage of training and prevent the audio encoder from being misled in the early stage of training. After the disease classification task is stable, the device adversarial task is gradually introduced to constrain it. The final optimization goal is a comprehensive goal of device robustness and disease recognition. The device invariance loss and the disease joint loss are weighted and summed to form the final total loss function, and the hyperparameters are optimized by grid search.

8. The device invariance enhancement multimodal deep learning model according to claim 7, characterized in that: In the aforementioned equipment invariance loss 1) The invariant risk minimization term, i.e. the IRM penalty term, is based on the Invariant Risk Minimization theory and introduces a gradient penalty term after the device environment is divided, so as to encourage the model to learn the shared feature space in different device environments; The specific approach is to divide the data acquisition device into sub-environments, calculate the main task loss gradient in each sub-environment, and minimize the variability of the main task gradient in each sub-environment to enhance the consistency of shared features among the data acquisition devices. 2) The device classification loss adopts the standard cross-entropy loss, which is introduced during device recognition training. During device recognition training, the device adversarial branch module uses a gradient inversion layer to connect the audio encoder and the device classification head, prompting the audio encoder to generate indistinguishable device representations, and adversarially training the audio encoder to confuse the device recognition ability of the device classification head. 3) Apply sigmoid dynamic gradient adjustment to the invariant risk minimization term to avoid accuracy degradation of IRM due to early addition; 4) The device invariance loss is composed of the sigmoid dynamic gradient adjustment, the invariant risk minimization term, and the device classification loss through a weighted combination, thereby improving the model's generalization ability to different acquisition devices.

9. The device invariance enhancement multimodal deep learning model according to claim 7, characterized in that: In the combined losses of the aforementioned diseases 1) The disease classification loss adopts binary cross-entropy loss with Logits or cross-entropy loss. The binary cross-entropy loss with Logits corresponds to mixed training or multi-label tasks, and the cross-entropy loss corresponds to single training or single-label tasks. 2) The audio-text contrast loss is calculated using the cosine similarity between the audio embedding α and the text embedding t, and consists of two parts: positive sample pair loss and negative sample pair loss. The objective of the positive sample pair loss is to maximize the similarity of positive sample pairs; The objective of the negative sample pair loss is to minimize the similarity of negative sample pairs, and to separate positive and negative samples using margin m. The audio-text contrast loss is a weighted sum of the positive sample pair loss and the negative sample pair loss, which can improve the consistency of patient multimodal representation. 3) The modal alignment loss adopts the Euclidean distance alignment loss of the fused modal features, which includes two parts: cosine similarity loss and mean square error loss. The goal of the cosine similarity loss is to align the audio and text embeddings so that their semantic similarity is as close to 1 as possible. The mean squared error loss further optimizes the alignment, minimizing the Euclidean distance between the audio and text embeddings; The modality alignment loss is a weighted sum of the cosine similarity loss and the mean square error loss. 4) The joint disease loss is composed of the disease classification loss, the audio-text comparison loss, and the modality alignment loss through adaptive weighting. The joint disease loss is weighted by the disease classification loss, the audio-text comparison loss, and the modality alignment loss through uncertainty weight parameters σ1, σ2, and σ3, all with an initial weight of 1, and includes a regularization term to stabilize training. During training, σ1, σ2, and σ3 are optimized along with other parameters of the model using gradient descent, as follows: Step 1. Calculate the forward propagation loss: Calculate each loss term based on the current model's predictions, and weight each loss term based on the current values ​​of σ1, σ2, and σ3, and calculate the total loss. Step 2. Calculate the gradient: The backpropagation process will calculate the gradient for each σ by taking derivatives; Step 3. Update σ; Step 4. The optimizer updates the σ value of each value using gradient descent based on the current learning rate η; Step 5. Iterative optimization: As training iterates, the model dynamically adjusts the values ​​of σ1, σ2, and σ3 according to the actual data distribution to balance the various losses; By optimizing the overall joint loss for the aforementioned diseases, the model can simultaneously improve in three aspects: classification, contrastive learning, and modality alignment, and dynamically adjust the weights of each part of the loss, ultimately achieving performance improvement.

10. An AI-assisted diagnostic system for respiratory diseases built using a multimodal deep learning model with enhanced device invariance as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Cardiovascular disease prediction system based on multi-modal fusion

    CN119446530A

  • Multi-modal enhanced representation collaborative learning pneumonia image recognition method

    CN120563854A

  • Intelligent sacroiliac joint lesion identification method and system based on MRI (Magnetic Resonance Imaging)

    CN120727244A

Cited By

  • Task awareness query guidance-based chronic obstructive pulmonary disease auxiliary judgment system

    CN121583513A

  • A task-aware query guidance based chronic obstructive pulmonary disease auxiliary judgment system

    CN121583513B