A respiratory disease risk early warning method based on multi-modal data fusion

By fusing multimodal data, calculating modality consistency weights, and performing collaborative interactive fusion, the problems of insufficient utilization of multimodal data and insufficient early warning in traditional methods are solved, and efficient and accurate early warning of respiratory disease risks is achieved.

CN122455334APending Publication Date: 2026-07-24FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-23
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional respiratory disease risk warning methods rely on single-modal data, which cannot fully explore the complementarity of multimodal information, ignore intermodal conflicts, and lack guidance from prior clinical knowledge, resulting in insufficient sensitivity of early warning.

Method used

By fusing multimodal data, calculating modal consistency weights, and employing a collaborative interactive fusion algorithm that combines consistency weighting terms and adaptive residual enhancement terms, modal contributions are dynamically adjusted to generate fused feature vectors for respiratory disease risk warning.

Benefits of technology

It improves the detection rate of early and occult cases, enhances the accuracy and sensitivity of early warning, is applicable to risk identification of complex cases, and improves the clinical prospectivity and robustness of early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122455334A_ABST
    Figure CN122455334A_ABST
Patent Text Reader

Abstract

The present application relates to the field of respiratory disease risk early warning, and more particularly to a respiratory disease risk early warning method based on multi-modal data fusion. First, based on the obtained multi-modal data, preprocessing is performed to obtain preprocessed multi-modal data; feature extraction is performed on the preprocessed multi-modal data to obtain a multi-modal feature vector; and the consistency weight of the mode is calculated based on the multi-modal feature vector. Then, based on the multi-modal feature vector and the consistency weight of the mode, a fusion feature vector is obtained through a collaborative interaction fusion algorithm; the high-risk probability of respiratory disease is calculated based on the fusion feature vector; and finally, the early warning result is obtained based on the high-risk probability of respiratory disease. The technical problems of insufficient clinical multi-modal data fusion, limited conflict processing capability between multi-modal data, lack of clinical prior knowledge guidance, and insufficient sensitivity of early warning of respiratory disease risk are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of respiratory disease risk early warning, and in particular to a respiratory disease risk early warning method based on multimodal data fusion. Background Technology

[0002] Traditional respiratory disease risk early warning systems typically rely on epidemiological surveys, hospital reports, and laboratory testing, which have inherent time lags and limitations. With increasing urban population density, deteriorating air quality, and frequent outbreaks of novel viruses, traditional monitoring and early warning methods face significant challenges in terms of sensitivity, real-time performance, and accuracy. Therefore, there is an urgent need to develop new, more intelligent, and efficient disease risk prediction and early warning technologies. Respiratory disease risk early warning methods based on multimodal data fusion can overcome the limitations of traditional methods in terms of timeliness, personalization, and prediction accuracy, and have broad application prospects.

[0003] However, existing methods still have problems such as insufficient fusion of clinical multimodal data, limited ability to handle conflicts between multimodalities, lack of guidance from prior clinical knowledge, and insufficient sensitivity in early warning of respiratory disease risks. Summary of the Invention

[0004] This invention provides a respiratory disease risk warning method based on multimodal data fusion to address the problems of traditional methods that mostly rely on a single modality and cannot fully explore the complementarity and synergy between multimodal information; traditional methods often ignore information conflicts between modalities, especially when signals are hidden in the early stages of the disease, leading to prediction distortion; traditional methods use data-driven mechanisms in isolation, lacking clinical experience constraints, making it difficult to ensure the rationality and credibility of decisions; and traditional methods can only identify risks when symptoms are obvious or physiological indicators are significantly abnormal, resulting in poor sensitivity to early and hidden cases.

[0005] The present invention provides a respiratory disease risk early warning method based on multimodal data fusion, comprising the following steps: S1. Acquire multimodal data and preprocess it to obtain preprocessed multimodal data; extract features from the preprocessed multimodal data to obtain multimodal feature vectors; based on the multimodal feature vectors, calculate the cosine similarity of each pair of feature vectors of different modes to obtain the mode consistency weight. S2. Based on multimodal feature vectors and consistency weights of different modalities, a collaborative interactive fusion algorithm is used to select the mode with the largest L2 norm of the multimodal feature vectors as the dominant mode. The feature vectors of the two modes in each mode pair are multiplied element-wise to generate a nonlinear interactive feature vector for each mode pair. Combined with the feature vector of the dominant mode, a gating scalar is generated. The gating scalar is multiplied element-wise with the nonlinear interactive feature vector of each mode pair to perform gating modulation. Combined with the consistency weighting term and the adaptive residual enhancement term, the fused feature vector is obtained. In the collaborative interactive fusion algorithm, the feature vectors of different modes are linearly weighted and summed based on the consistency weights of different modes to obtain the consistency weighting term. Using residuals, the part of the feature vector of the dominant mode that is negatively correlated with the consistency weight is used to obtain the adaptive residual enhancement term. The fused feature vector is mapped to a one-dimensional scalar space through a fully connected linear transformation, and a nonlinear activation function is applied to calculate the high-risk probability of respiratory diseases. The high-risk probability of respiratory diseases is compared with the set risk judgment threshold to obtain the warning result.

[0006] Preferably, S1 specifically includes: The preprocessed multimodal data includes preprocessed physiological data vectors, preprocessed chest images, preprocessed symptom text, and preprocessed medical history text. Features are extracted from the preprocessed physiological data vectors using a fully connected neural network to obtain high-dimensional semantic feature vectors. Deep convolutional feature extraction is performed on the preprocessed chest images using a pre-trained convolutional neural network to obtain image semantic feature vectors. Semantic encoding is performed on the preprocessed symptom text and preprocessed medical history text using a pre-trained language model to obtain symptom semantic feature vectors and medical history risk feature vectors. The high-dimensional semantic feature vectors, image semantic feature vectors, symptom semantic feature vectors, and medical history risk feature vectors constitute the multimodal feature vector.

[0007] Preferably, S1 specifically includes: The absolute deviation between the cosine similarity and the elements in the predefined clinical prior correlation matrix is ​​obtained by subtracting the cosine similarity from the elements in the clinical prior correlation matrix.

[0008] Preferably, S1 specifically includes: For each mode, the absolute deviation values ​​from all other modes are summarized to generate a mode consistency deviation index. The mode consistency deviation index is then negatively calculated and exponentially operated to obtain the mode exponential score.

[0009] Preferably, S1 specifically includes: The modality consistency weight is obtained by dividing the modality's exponent score by the sum of all modality exponent scores and then normalizing it.

[0010] Preferably, S2 specifically includes: In the collaborative interaction fusion algorithm, the residual enhancement coefficient is obtained by subtracting the consistency weight of the dominant mode from 1; the adaptive residual enhancement term is obtained by multiplying the residual enhancement coefficient with the feature vector of the dominant mode.

[0011] Preferably, S2 specifically includes: In the collaborative interaction fusion algorithm, based on the nonlinear interaction feature vector of each mode pair, the cosine similarity between the feature vector of the dominant mode and the feature vector of the dominant mode is calculated; based on the cosine similarity, a gated scalar is generated through the sigmoid function.

[0012] Preferably, S2 specifically includes: The gating scalar is multiplied by the nonlinear interaction eigenvector of the mode pair, then multiplied by the product of the consistency weights of the two modes in the mode pair, and summed to generate all the interaction terms that have been modulated by gating pairwise.

[0013] Preferably, S2 specifically includes: The consistency weighting term, the adaptive residual enhancement term, and all interaction terms that have undergone pairwise gating modulation are added together to obtain the fused feature vector.

[0014] The beneficial effects of the technical solution of the present invention are: 1. By quantifying the cosine similarity between multimodal feature vectors and the deviation of elements in the clinical prior correlation matrix, the contribution weight of each modality in subsequent fusion can be dynamically allocated. This can automatically identify and suppress the interference of abnormal or noisy modalities, while strengthening the contribution of modalities that exhibit typical pathological features. This enhances the ability to respond to real clinical patterns and helps improve the accuracy and robustness of fusion judgment.

[0015] 2. By combining consistency weighting terms, adaptive residual enhancement terms, and all interaction terms that have undergone pairwise gating modulation to model, the conservative decision-making logic of clinicians when faced with conflicting information can be simulated. This can effectively preserve early or hidden risk signals and avoid them being submerged by averaging during fusion, thereby improving the detection rate of complex or boundary samples, which has important clinical prospective value.

[0016] 3. By dynamically selecting the dominant mode through the L2 norm and applying consistency weights for reverse modulation of the residual enhancement mechanism, features that are inconsistent with other modes but may have diagnostic value are strengthened, improving the response to modes that deviate significantly from risk information. This is especially suitable for complex cases with weak signals and heterogeneous manifestations in the early stage of the disease, and enhances the sensitivity in atypical scenarios.

[0017] 4. The Hadamard product interaction is performed on all modal pairs, and dynamic gating amplification or suppression is applied based on the cosine similarity with the dominant modality. This significantly improves the ability to model fine-grained synergistic relationships between modalities, captures subtle interaction information between multimodalities, and helps to discover deep pathological patterns. Attached Figure Description

[0018] Figure 1 This is a flowchart of a respiratory disease risk early warning method based on multimodal data fusion as described in this invention. Detailed Implementation

[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0021] The following description, in conjunction with the accompanying drawings, details a specific scheme for a respiratory disease risk early warning method based on multimodal data fusion provided by the present invention.

[0022] See attached document Figure 1 The diagram illustrates a flowchart of a respiratory disease risk warning method based on multimodal data fusion, provided by an embodiment of the present invention. The method includes the following steps: S1. Acquire multimodal data and preprocess it to obtain preprocessed multimodal data; extract features from the preprocessed multimodal data to obtain multimodal feature vectors; based on the multimodal feature vectors, calculate the cosine similarity of each pair of feature vectors of different modes to obtain the mode consistency weight.

[0023] On the day of the patient's visit, multimodal data is obtained from the hospital information system and the image archiving and communication system. The obtained multimodal data includes: physiological data, such as body temperature, heart rate, blood oxygen saturation, respiratory rate, etc.; chest images; symptom text strings, such as descriptions of the patient's chief complaint of cough, difficulty breathing, etc.; and past medical history text strings, such as electronic medical record history.

[0024] Furthermore, the acquired multimodal data undergoes targeted preprocessing to obtain preprocessed multimodal data, including preprocessed physiological data vectors, preprocessed chest images, preprocessed symptom text, and preprocessed medical history text. Specifically, the physiological data is converted into a mathematical vector form, resulting in a physiological data vector. This vector is then subjected to Z-score standardization to eliminate the influence of differences in dimensions and numerical ranges among different indicators. The chest images are resized to a fixed resolution and grayscale normalized to ensure consistent pixel value distribution, resulting in preprocessed chest images. The symptom text strings and medical history text strings are segmented and standardized using medical terminology, resulting in preprocessed symptom text and preprocessed medical history text, ensuring that the text data is clean and retains clinical semantic information.

[0025] Feature extraction is performed on the preprocessed multimodal data to obtain multimodal feature vectors, including high-dimensional semantic feature vectors, image semantic feature vectors, symptom semantic feature vectors, and medical history risk feature vectors. Specifically, a high-dimensional semantic feature vector is obtained by extracting from the preprocessed physiological data vector through a multi-layer fully connected neural network. The multi-layer fully connected neural network adopts a typical feedforward structure, consisting of three fully connected layers, including two hidden layers and one output layer, but no input layer. The dimension of the input layer is equal to the dimension of the preprocessed physiological data vector. The first hidden layer maps the input to a higher-dimensional feature space to enhance the feature representation capability, and the number of nodes can be set to twice the input dimension. The second hidden layer compresses and reconstructs the high-dimensional features, and the number of nodes can be set to half that of the first hidden layer to extract stable semantic representations. The output layer maps the features to a high-dimensional semantic feature vector of a fixed dimension, such as 128 or 256, as the input for subsequent multimodal fusion. Both hidden layers use the ReLU function to improve non-linear expression capability and avoid gradient vanishing. The output layer does not use an activation function or uses linear activation to preserve the original distribution information of the features. Batch normalization can be added between layers to stabilize the training process, and a lightweight Dropout, such as 0.2, can be used to suppress overfitting. Preprocessed chest images are processed using a pre-trained ResNet-50 convolutional neural network for deep convolutional feature extraction, yielding image semantic feature vectors. A language model, such as ClinicalBERT, based on a Transformer encoder architecture and pre-trained on clinical medical data, is then used to perform contextual semantic encoding on the preprocessed symptom text, resulting in symptom semantic feature vectors. The ClinicalBERT model takes a sequence of tokens from the preprocessed symptom text as input, with a sequence length not exceeding 128. Each token is mapped to a 768-dimensional word vector with positional encoding. The model employs a 12-layer Transformer encoder structure, with a hidden layer dimension of 768, 12 self-attention heads, and a subspace dimension of 64 for each attention head (768 / 12). The feedforward network's intermediate layers have a dimension of 3072. The output is the 768-dimensional vector corresponding to the [CLS] position in the last layer, which serves as the symptom semantic feature vector. Finally, the preprocessed medical history text is processed using the aforementioned language model based on a Transformer encoder architecture and pre-trained on clinical medical data for long-range dependency semantic encoding, yielding a medical history risk feature vector.

[0026] Based on multimodal feature vectors, the contribution weight of each modality in subsequent fusion is dynamically calculated by quantifying the cosine similarity between multimodal feature vectors and the deviation of elements in the clinical prior correlation matrix, i.e., consistency weight, thereby realizing intelligent control of complementary and conflicting information between modalities.

[0027] For each pair of feature vectors from different modalities, a cosine similarity is calculated. Cosine similarity reflects the degree of directional consistency between the feature vectors of the two modalities in a high-dimensional semantic space, with a value ranging from -1 to 1. A larger positive value indicates that the feature vectors of the two modalities point to more similar risk patterns, while a negative value indicates that the feature vectors of the two modalities are in opposite directions, potentially indicating conflict. Furthermore, the deviation between the cosine similarity calculated for each pair of feature vectors from different modalities and the elements in the clinical prior correlation matrix is ​​quantified. The magnitude of the absolute deviation is calculated. A smaller absolute deviation indicates that the current patient's modal relationship more closely matches the typical clinical presentation pattern of respiratory diseases. A larger absolute deviation indicates abnormal modal incoordination, possibly corresponding to latent manifestations or noise interference in the early stages of the disease. The elements in the clinical prior correlation matrix are determined based on the average Pearson correlation coefficient between similar feature vectors extracted from domestic respiratory disease cohort studies or set in conjunction with the consensus of respiratory medicine experts, with a value ranging from -1 to 1.

[0028] For each mode, the absolute deviation values ​​from other modes are summarized to form a mode consistency deviation index. The lower the consistency deviation index, the higher the coordination between the mode and other modes. The consistency deviation index of each mode is negatively calculated and then exponentially operated so that the mode with the lower consistency deviation index receives an exponentially higher exponential score, while the mode with the higher consistency deviation index is exponentially suppressed.

[0029] The exponent scores of different modes are normalized so that their sum is 1, thus obtaining the consistency weights of different modes. The calculation formula is as follows: ; in, Indicates the first Consistency weights for each modality; Represents an exponential fraction; Represents an exponential function; The consistency deviation index is obtained by summing the absolute deviation values ​​of each mode from all other modes except itself. Indicates the first The first mode and the first The absolute deviation of the cosine similarity calculated from the eigenvectors of the modality from the elements in the clinical prior correlation matrix is ​​used to represent the _th _th_ modality. The absolute deviation value of each mode; Indicates the first Feature vectors of each modality; Indicates the transpose symbol; Indicates the first Feature vectors of each modality; and They represent the first The feature vector of the i-th mode and the i-th mode The L2 norm of the eigenvectors of each modality; Indicates the first The eigenvector of the i-th modality and the i-th modality The cosine similarity of the feature vectors of two modalities is used to measure the directional consistency of the feature vectors of two modalities in a high-dimensional feature space. In the clinical prior correlation matrix, the first... Line 1 Column elements are used to reflect the clinical scenarios of respiratory diseases. The modality and the first The expected correlation strength of each modality.

[0030] By introducing prior knowledge from the clinical domain as a reference benchmark, the consistency between the current modal feature vectors of patients is dynamically evaluated, thereby assigning reasonable contribution weights to each modality and achieving risk-oriented intelligent fusion decision-making.

[0031] S2. Based on multimodal feature vectors and consistency weights of different modalities, a collaborative interactive fusion algorithm selects the modality with the largest L2 norm among the multimodal feature vectors as the dominant modality. The feature vectors of the two modalities in each modality pair are multiplied element-wise to generate a nonlinear interactive feature vector for each modality pair. This nonlinear interactive feature vector is then combined with the feature vector of the dominant modality to generate a gating scalar. The gating scalar is then multiplied element-wise with the nonlinear interactive feature vector of each modality pair to perform gating modulation. Combined with a consistency weighting term and an adaptive residual enhancement term, a fused feature vector is obtained. In the collaborative interactive fusion algorithm, the feature vectors of different modalities are linearly weighted and summed based on the consistency weights of different modalities to obtain a consistency weighting term. An adaptive residual enhancement term is obtained based on the negative correlation between the feature vector of the dominant modality and the consistency weights using residuals. The fused feature vector is mapped to a one-dimensional scalar space through a fully connected linear transformation, and a nonlinear activation function is applied to calculate the high-risk probability of respiratory diseases. The high-risk probability of respiratory diseases is compared with a set risk judgment threshold to obtain a warning result.

[0032] Based on multimodal feature vectors and consistency weights of different modes, a collaborative interaction fusion algorithm is used to perform fine pairwise gating modulation on the nonlinear interactions between all pairs of modes, with the most significant risk mode, i.e. the dominant mode, as the dynamic guiding signal. At the same time, by combining consistency weighting terms and adaptive residual enhancement terms, risk-oriented high-sensitivity fusion is achieved to obtain the fused feature vector.

[0033] Using the consistency weights of different modalities as coefficients, the feature vectors of different modalities are linearly weighted and summed to obtain the consistency weighted term, which reflects the preliminary assessment of the overall consistency of evidence in clinical diagnosis.

[0034] By analyzing the L2 norm of the eigenvectors of different modalities, the modality with the largest L2 norm is selected as the dominant modality. After selecting the dominant modality, an adaptive residual enhancement term is obtained by adding only the portion of the dominant modality's eigenvector negatively correlated with its own consistency weight, using residual form. Specifically, the higher the consistency weight of the dominant modality, the more coordinated it is with other modalities, resulting in a smaller residual enhancement strength to avoid redundant emphasis; conversely, the lower the consistency weight of the dominant modality, the more significant the conflict between it and other modalities, resulting in a larger residual enhancement strength. This adaptive residual mechanism accurately simulates the conservative decision-making logic of clinicians when faced with modality conflict—preferring to err on the side of caution and prioritizing the most anomalous evidence—effectively preventing early or occult respiratory disease signals from being averaged out.

[0035] The algorithm iterates through all different modal pairs. For each modal pair, the feature vectors of the two modes are first multiplied element-wise by the Hadamard product to generate the nonlinear interaction feature vector of each modal pair. The Hadamard product can capture the synergistic or complementary effects of the two modes in the feature space in each dimension. The nonlinear interaction feature vector provides rich potential synergistic information for subsequent selective amplification. At the same time, the nonlinear interaction feature vector of each modal pair is also controlled by the product of the consistency weights of the two corresponding modes. The mode with higher consistency weights contributes more to the nonlinear interaction.

[0036] For each mode pair's nonlinear interaction feature vector, the cosine similarity between the nonlinear interaction and the dominant mode's feature vector is calculated. A positive cosine similarity close to 1 indicates that the nonlinear interaction supports the dominant risk, i.e., a supportive interaction; a negative cosine similarity close to 0 indicates that the nonlinear interaction conflicts with or is unrelated to the dominant risk, i.e., a contradictory or irrelevant interaction. Based on the cosine similarity, a gating scalar with a value between 0 and 1 is obtained through the sigmoid function. This scalar is then multiplied element-wise with the nonlinear interaction feature vector of each mode pair to achieve soft-gating modulation. When the gating scalar is close to 1, the supportive interaction is almost completely preserved and amplified; when the gating scalar is close to 0, contradictory or irrelevant interactions are significantly suppressed.

[0037] The consistency weighting term, the adaptive residual enhancement term, and all interaction terms modulated by pairwise gating are added together to obtain the fused feature vector, calculated as follows: ; in, Represents the fused feature vector; This represents the consistency weighting term, which uses the consistency weights of different modalities as coefficients to perform a linear weighted summation on the feature vectors of different modalities. This represents the adaptive residual enhancement term, which takes the form of a residual and only adds the part of the dominant mode's feature vector that is negatively correlated with its own consistency weight. Indicates the residual enhancement factor; Represents the consistency weight of the dominant mode; The feature vector representing the dominant mode; Index indicating the dominant mode, ; This represents all interaction items that have undergone pairwise gating modulation; Summing over all unordered mode pairs; Indicates the first Consistency weights for each modality; Indicates the first The first mode and the first The nonlinear interaction eigenvectors of the mode pairs of the ... The eigenvector of the i-th modality and the i-th modality The feature vectors of each modality are multiplied element by element; Represents the Hadamard product; Indicates the first The modality and the first A modal gated scalar.

[0038] Using the fused feature vector as input, a fully connected neural network classifier (i.e., a single-layer fully connected linear transformation) maps the fused feature vector to a one-dimensional scalar space. Then, a sigmoid nonlinear activation function is applied to compress it to the [0,1] interval, yielding the high-risk probability of respiratory diseases. This high-risk probability is stored in a historical sample database for subsequent calculations.

[0039] The risk assessment threshold is determined by comparing the probability of a high risk of respiratory disease with a risk determination threshold. If the probability of a high risk of respiratory disease is greater than the risk determination threshold, the patient is marked as high-risk and a high-risk status indicator appears in the patient list or details page of the hospital information system. Otherwise, the patient is marked as low-risk. The risk determination threshold is determined using a fixed quantile threshold method. Statistical analysis is performed on historical sample data in the historical sample database, and a preset high quantile of the probability distribution is selected as the risk determination threshold. The high quantile can be selected as the 90th percentile or the 95th percentile, that is, the respiratory disease high-risk probability value in the highest 10% or 5% of the historical sample data is used as the risk determination threshold.

[0040] The entire S2 step can significantly improve the detection sensitivity of early and occult cases, reduce the risk of missed diagnosis, improve the level of respiratory disease risk management, help optimize the allocation of medical resources, and reduce the incidence of adverse patient outcomes. It has important clinical application value and public health significance.

[0041] In summary, a respiratory disease risk early warning method based on multimodal data fusion has been developed.

[0042] The order of the embodiments is for illustrative purposes only and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0043] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0044] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A respiratory disease risk early warning method based on multimodal data fusion, characterized in that, Includes the following steps: S1. Acquire multimodal data and preprocess it to obtain preprocessed multimodal data; Feature extraction is performed on the preprocessed multimodal data to obtain multimodal feature vectors; based on the multimodal feature vectors, the cosine similarity of each pair of feature vectors from different modes is calculated to obtain the mode consistency weight. S2. Based on multimodal feature vectors and consistency weights of different modes, a collaborative interaction fusion algorithm is used to select the mode with the largest L2 norm of the multimodal feature vectors as the dominant mode. The feature vectors of the two modes in each mode pair are multiplied element-wise to generate a nonlinear interaction feature vector for each mode pair. Combined with the feature vector of the dominant mode, a gated scalar is generated. The gated scalar is multiplied element-wise with the nonlinear interaction feature vector of each mode pair to perform gated modulation. Combined with the consistency weighting term and the adaptive residual enhancement term, the fused feature vector is obtained. In the collaborative interaction fusion algorithm, the feature vectors of different modes are linearly weighted and summed based on the consistency weights of different modes to obtain the consistency weight term; the residual is used to obtain the adaptive residual enhancement term based on the part of the feature vector of the dominant mode that is negatively correlated with the consistency weight. The fused feature vectors are mapped to a one-dimensional scalar space through a fully connected linear transformation, and a nonlinear activation function is applied to calculate the high-risk probability of respiratory diseases. The probability of high risk of respiratory diseases is compared with the set risk assessment threshold to obtain the early warning result.

2. The respiratory disease risk early warning method based on multimodal data fusion according to claim 1, characterized in that, S1 specifically includes: The preprocessed multimodal data includes preprocessed physiological data vectors, preprocessed chest images, preprocessed symptom text, and preprocessed medical history text. Features are extracted from the preprocessed physiological data vectors using a fully connected neural network to obtain high-dimensional semantic feature vectors. Deep convolutional feature extraction is performed on the preprocessed chest images using a pre-trained convolutional neural network to obtain image semantic feature vectors. Semantic encoding is performed on the preprocessed symptom text and preprocessed medical history text using a pre-trained language model to obtain symptom semantic feature vectors and medical history risk feature vectors. The high-dimensional semantic feature vectors, image semantic feature vectors, symptom semantic feature vectors, and medical history risk feature vectors constitute the multimodal feature vector.

3. The respiratory disease risk early warning method based on multimodal data fusion according to claim 1, characterized in that, S1 specifically includes: The absolute deviation between the cosine similarity and the elements in the predefined clinical prior correlation matrix is ​​obtained by subtracting the cosine similarity from the elements in the clinical prior correlation matrix.

4. The respiratory disease risk early warning method based on multimodal data fusion according to claim 3, characterized in that, S1 specifically includes: For each mode, the absolute deviation values ​​from all other modes are summarized to generate a mode consistency deviation index. The mode consistency deviation index is then negatively calculated and exponentially operated to obtain the mode exponential score.

5. A respiratory disease risk early warning method based on multimodal data fusion according to claim 4, characterized in that, S1 specifically includes: The modality consistency weight is obtained by dividing the modality's exponent score by the sum of all modality exponent scores and then normalizing it.

6. The respiratory disease risk early warning method based on multimodal data fusion according to claim 1, characterized in that, S2 specifically includes: In the collaborative interaction fusion algorithm, the residual enhancement coefficient is obtained by subtracting the consistency weight of the dominant mode from 1; the adaptive residual enhancement term is obtained by multiplying the residual enhancement coefficient with the feature vector of the dominant mode.

7. A respiratory disease risk early warning method based on multimodal data fusion according to claim 6, characterized in that, S2 specifically includes: In the collaborative interaction fusion algorithm, based on the nonlinear interaction feature vector of each mode pair, the cosine similarity between the feature vector of the dominant mode and the feature vector of the dominant mode is calculated; based on the cosine similarity, a gated scalar is generated through the sigmoid function.

8. The respiratory disease risk early warning method based on multimodal data fusion according to claim 7, characterized in that, S2 specifically includes: The gating scalar is multiplied by the nonlinear interaction eigenvector of the mode pair, then multiplied by the product of the consistency weights of the two modes in the mode pair, and summed to generate all the interaction terms that have been modulated by gating pairwise.

9. A respiratory disease risk early warning method based on multimodal data fusion according to claim 8, characterized in that, S2 specifically includes: The consistency weighting term, the adaptive residual enhancement term, and all interaction terms that have undergone pairwise gating modulation are added together to obtain the fused feature vector.