Pulmonary infectious disease prediction system based on multimodal data fusion
Through a multimodal data fusion system, combining clinical data, medical imaging and environmental factors, the attention mechanism is used to perform feature fusion and timing analysis, which solves the problem that a single data source analysis in the existing technology cannot capture the impact and timing changes of multiple factors, and achieves more efficient prediction and management of infectious pulmonary diseases.
Patent Information
- Application Number
- CN202510067485.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-16
AI Technical Summary
In the prediction and diagnosis of infectious lung diseases, the prior art rely on the analysis of a single data source and cannot fully capture the influence of multiple factors and timing changes, resulting in inaccuracy in prediction and management efficiency.
A system based on multimodal data fusion is adopted, and clinical data, medical imaging data and environmental factor data are collected in real time through the data acquisition module. The feature extraction module extracts physiological indicators, images and environmental features. The dynamic fusion module adopts an attention mechanism for feature fusion, and the prediction modeling module performs timing analysis to predict disease probability and progress trend.
It significantly improves the prediction accuracy and management efficiency of infectious lung diseases, and can more comprehensively reflect the multidimensional factors of the disease, and dynamically monitor and manage disease progress trends.
Smart Images

Figure CN119480124B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a lung infectious disease prediction system based on multimodal data fusion. Background Art
[0002] In existing technologies, the prediction and diagnosis of pulmonary infectious diseases usually rely on the analysis of a single data source, such as clinical data or medical imaging data. These methods mainly use traditional statistical models or single-modality machine learning algorithms to model and predict patients' physical signs, blood test indicators or imaging features. In addition, some methods may combine macro data such as disease prevalence for auxiliary analysis.
[0003] The main problem with existing technologies is that single-modality analysis methods cannot fully capture the multi-factorial impact of pulmonary infectious diseases, including the complex relationship between patients' physiological characteristics, imaging manifestations, and environmental factors. In addition, these methods usually fail to dynamically analyze the temporal changes of data and cannot effectively predict the progression trend of the disease, thus limiting the dynamic monitoring and management of patients' conditions.
[0004] Therefore, it is necessary to provide a system that can integrate multimodal data features and capture temporal variation patterns to improve the prediction accuracy of pulmonary infectious diseases and the efficiency of disease management. Summary of the invention
[0005] The present application provides a lung infectious disease prediction system based on multimodal data fusion to improve the accuracy and reliability of lung infectious disease prediction.
[0006] The present application provides a pulmonary infectious disease prediction system based on multimodal data fusion, comprising:
[0007] A data acquisition module, for collecting clinical data, medical imaging data and environmental factor data of patients in real time, wherein the clinical data includes physical sign data, blood test data and symptom description data, the medical imaging data includes chest X-rays, CT scan images and ultrasound images, and the environmental factor data includes air quality index, meteorological data and regional infectious disease prevalence data;
[0008] A feature extraction module, used to extract physiological index feature vectors from the clinical data, image feature vectors from the medical imaging data, and environmental impact feature vectors from the environmental factor data;
[0009] A dynamic fusion module, used to use an attention mechanism to adaptively assign weights to the physiological indicator feature vector, the image feature vector, and the environmental impact feature vector, and fuse them into a unified multimodal feature representation;
[0010] The predictive modeling module performs time series analysis on the multimodal feature representation to obtain the time series characteristics of the multimodal feature representation; based on the time series characteristics, the probability value of the patient suffering from a lung infectious disease and the possible disease progression trend are predicted.
[0011] This application has the following beneficial technical effects:
[0012] (1) The present invention integrates clinical data, medical imaging data and environmental factor data, and overcomes the limitations of traditional single data source analysis through multimodal fusion, which can more comprehensively reflect the multidimensional factors of pulmonary infectious diseases, thereby significantly improving the accuracy and reliability of disease prediction. (2) The dynamic fusion module realizes the adaptive weight allocation of feature vectors through the attention mechanism, so that the most relevant features can be highlighted in different scenarios, improving the intelligence of data fusion and enhancing the system's ability in complex data association analysis. (3) The predictive modeling module uses time series analysis to extract the temporal dynamic change pattern of multimodal features, which can effectively predict the probability of disease occurrence and progression trend, and provide a scientific basis for personalized treatment and dynamic disease management. (4) The system design adopts a modular architecture with clear functional division, which is convenient for adding new data sources or improving existing modules. It has good scalability and can adapt to different application scenarios, providing a flexible and efficient solution for medical diagnosis and public health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of a lung infectious disease prediction system based on multimodal data fusion provided in the first embodiment of the present application. DETAILED DESCRIPTION
[0014] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0015] The first embodiment of the present application provides a lung infectious disease prediction system based on multimodal data fusion. Figure 1 , which is a schematic diagram of the first embodiment of the present application. Figure 1 The first embodiment of the present application provides a lung infectious disease prediction system based on multimodal data fusion and is described in detail.
[0016] The pulmonary infectious disease prediction system based on multimodal data fusion includes a data acquisition module 101 , a feature extraction module 102 , a dynamic fusion module 103 and a prediction modeling module 104 .
[0017] The data acquisition module 101 is used to collect the patient's clinical data, medical imaging data and environmental factor data in real time, wherein the clinical data includes physical sign data, blood test data and symptom description data, the medical imaging data includes chest X-rays, CT scan images and ultrasound images, and the environmental factor data includes air quality index, meteorological data and regional infectious disease prevalence data.
[0018] The data acquisition module 101 is a basic component of the prediction system provided in this embodiment, and its function is to collect multimodal data related to pulmonary infectious diseases in real time and efficiently. These data include clinical data, medical imaging data, and environmental factor data of patients, which comprehensively cover the key information required for disease prediction. The design of this module fully considers the diversity and complexity of data sources, and ensures the accuracy, timeliness and reliability of data collection through technical means.
[0019] For the collection of clinical data, module 101 records the patient's vital signs data, such as body temperature, pulse, blood pressure and blood oxygen saturation, in real time by integrating a variety of biosensor devices. These devices can use wireless communication technology (such as Bluetooth or ZigBee) to establish a connection with the module to achieve automatic data transmission. In addition, blood test data can be achieved by docking with the laboratory information management system, supporting automatic extraction of key indicators, including white blood cell count, C-reactive protein level and blood gas analysis results. Symptom description data is entered by the patient through the electronic consultation platform, and supplemented by natural language processing (NLP) technology to convert unstructured text into structured feature data to support subsequent analysis.
[0020] In terms of medical imaging data, module 101 uses the standardized DICOM protocol and integrates seamlessly with the medical image storage and communication system (PACS), and can receive chest X-rays, CT scan images, and ultrasound images. These data are encrypted during transmission to ensure privacy protection and data integrity. At the same time, the module supports basic preprocessing functions such as image denoising, resolution adjustment, and region segmentation to provide high-quality input for feature extraction. Through this design, the module can adapt to the data standards of different medical institutions and has high compatibility.
[0021] The collection of environmental factor data is achieved through real-time connection with external data sources, such as meteorological monitoring stations, public health databases, and air quality monitoring platforms. Module 101 uses an API interface to regularly call information such as air quality index (AQI), temperature, humidity, wind speed, and prevalence of infectious diseases, and combines geolocation technology (such as GPS) to associate environmental data with the specific location of the patient. This function makes the collected environmental data more accurate and provides an important reference for subsequent disease prediction.
[0022] The uniqueness of module 101 also lies in its data verification and optimization mechanism. To improve the reliability of data, the module has a built-in outlier detection function, which uses machine learning models to identify and correct noise and errors in the data collection process. In addition, the module also supports dynamic sampling rate adjustment of data, such as increasing the sampling frequency when the patient's symptoms change dramatically to capture more key information. In low-resource environments, the module can also achieve offline data storage through local caching function, and automatically upload after the network is restored.
[0023] Module 101 can be further extended to personalized data collection strategies. For example, based on the patient's specific situation and medical history, the module can dynamically adjust the collection parameters and focus on specific data types. For example, for patients with previous lung diseases, the frequency and resolution of image data collection can be prioritized; for patients in high-risk areas for infectious diseases, real-time monitoring of environmental factor data can be strengthened. This design improves the intelligence and adaptability of the system, making the module not only suitable for conventional disease prediction, but also able to respond to public health emergencies or the precise health monitoring needs of specific populations.
[0024] In summary, the data acquisition module 101 not only lays the foundation for the system's prediction capabilities, but also demonstrates a high degree of flexibility through efficient collection, verification and dynamic optimization of multi-dimensional data.
[0025] Furthermore, the data acquisition module includes a data verification unit for performing real-time consistency check and integrity verification on the collected clinical data, medical imaging data and environmental factor data during the data acquisition process, specifically including:
[0026] By comparing historical data trends, identifying outliers and detecting data loss, the accuracy and reliability of the collected data can be ensured. At the same time, it supports automatic generation of predicted values for missing data to supplement incomplete records, thereby improving the overall quality of the data.
[0027] The data verification unit is designed to ensure the accuracy and integrity of multimodal data, providing a reliable data foundation for subsequent prediction and analysis. The implementation of the data verification unit focuses on real-time inspection and dynamic completion, which can perform efficient and accurate processing at each stage of data collection. Its core functions include historical data trend comparison, outlier identification, data loss detection, and missing data completion.
[0028] First, the data verification unit determines the rationality of the collected data by comparing the currently collected data with the historical data trends of the patient or environment. This process involves the storage, indexing and rapid retrieval of historical data, and the system can identify whether the data conforms to the expected change pattern. For example, in clinical data, whether the changes in body temperature or blood pressure are consistent with the patient's previous health records. If there is a significant deviation, it is marked as suspicious data and enters the next step of outlier analysis.
[0029] Secondly, the data verification unit uses a built-in outlier recognition algorithm to screen outliers for each modality of data collected. For example, for clinical data, the algorithm can detect measurements that do not conform to the physiological range; for medical imaging data, it analyzes whether the image quality meets the clarity requirements, such as whether there are blurred areas or artifact interference; for environmental factor data, it determines whether the information provided by the monitoring equipment or external data source is within the normal range, such as whether the fluctuation of the air quality index is reasonable. This multi-dimensional outlier recognition mechanism can effectively improve the credibility of the data.
[0030] When detecting missing data, the data verification unit combines modality-specific acquisition rules to quickly find missing key data points. For example, for blood test data, if certain important indicators (such as white blood cell count or C-reactive protein level) are missing, the system can mark these data and list them in the verification report. This function is crucial in dealing with complex multi-modal data acquisition.
[0031] When missing data is found, the data verification unit further supports automatic generation of predicted values for missing data to supplement incomplete records. This function is based on the inherent relationship between historical data and currently collected data, and infers missing values by building a prediction model (such as interpolation or regression analysis based on time series). For example, in a time series of multiple blood tests of a patient, the system can predict missing indicator values through trend analysis; for air quality index data, interpolation calculations can be performed in combination with meteorological conditions and regional historical data. The generated predicted values are marked as completed data and accompanied by a confidence score for reference in subsequent processing.
[0032] Through the above functions, the data verification unit not only significantly improves the accuracy of collected data, but also can dynamically respond to a variety of potential problems during the data collection process, such as equipment failure, fluctuations in collection conditions, or instability of external data sources. Its modular design ensures that it can adapt to the complex needs of multimodal data while providing high-quality input for downstream feature extraction and predictive modeling.
[0033] Furthermore, the data acquisition module includes a dynamic sampling unit for adaptively adjusting the frequency of data acquisition and data type priority according to the patient's condition changes or dynamic fluctuations of environmental factors, specifically including:
[0034] When the patient's condition changes dramatically, the dynamic sampling unit prioritizes increasing the frequency of collecting vital sign data and medical imaging data. During periods of high risk of regional infectious diseases, it prioritizes collecting environmental factor data and symptom description data to implement flexible data collection strategies for different scenarios.
[0035] The dynamic sampling unit in the data acquisition module is designed to flexibly adjust the frequency of data acquisition and the priority of data types according to the dynamic changes of the patient's condition or the external environment through intelligent and adaptive mechanisms, thereby ensuring that the most valuable information can be collected in different scenarios and providing high-quality support for subsequent analysis and prediction.
[0036] The implementation of the dynamic sampling unit is based on a real-time monitoring and analysis mechanism. During operation, the unit continuously monitors multimodal data from patients and the environment, including but not limited to vital sign data, medical imaging data, symptom description data, and external environmental factors such as air quality index and infectious disease prevalence. When the system detects a significant change in the patient's condition, such as a sharp increase in body temperature, a decrease in blood oxygen saturation, or a new abnormal shadow in the medical image, the dynamic sampling unit will immediately increase the frequency of collecting vital sign data and medical imaging data. This rapid response mechanism enables the system to capture details of acute changes in the condition with a higher temporal resolution by prioritizing the scheduling of sensors, imaging devices, or data interfaces, thereby supporting timely diagnosis and intervention.
[0037] On the other hand, during periods of high risk of regional infectious diseases, the focus of the dynamic sampling unit will be adjusted according to fluctuations in environmental data. For example, when external data sources show that the air quality index in a certain area continues to deteriorate or the risk of infectious disease transmission increases, the system will increase the frequency of collecting environmental factor data while giving priority to collecting symptom description data. The core of this strategy is to focus limited collection capabilities on the most critical variables at the moment by dynamically adjusting resource allocation and collection plans, ensuring prediction accuracy in complex environments.
[0038] The flexibility of the dynamic sampling unit is also reflected in the dynamic priority sorting of data types. For example, when the patient's condition is stable but the epidemic risk in the area increases, the system can reduce the frequency of collecting vital sign data and medical imaging data and transfer resources to the acquisition of environmental factor data. In addition, for environmental factor data with periodic changes (such as daily fluctuations in the air quality index), the dynamic sampling unit can intelligently reduce redundant sampling based on known change patterns, thereby improving the overall efficiency of system operation.
[0039] This adaptive adjustment mechanism is supported by multi-layer decision logic. The dynamic sampling unit calculates the collection priority and resource allocation weight by analyzing the difference between real-time data and historical baseline data. In the implementation process, the dynamic rule set embedded in the unit can not only adapt to conventional collection needs, but also be expanded according to user-defined policies or external input conditions. For example, in a public health emergency, the dynamic sampling unit can receive risk level information released by the government through an external interface and quickly adjust the sampling strategy to adapt to new scenario requirements.
[0040] Through this collection optimization based on the dynamic changes of patients and the environment, the dynamic sampling unit not only ensures the timeliness and pertinence of the data, but also reduces unnecessary collection and storage burdens. This flexible and efficient design greatly improves the system's adaptability in complex scenarios and provides strong technical support for the realization of precision medicine and public health management.
[0041] The feature extraction module 102 is used to extract physiological index feature vectors from the clinical data, extract image feature vectors from the medical image data, and extract environmental impact feature vectors from the environmental factor data.
[0042] The feature extraction module 102 is one of the key technical modules of the prediction system provided in this embodiment. Its function is to extract useful feature vectors from the multimodal data provided by the data acquisition module to provide high-quality input for subsequent data fusion and prediction modeling. This module ensures that the extracted features are highly representative in semantics and achieves the best balance between the expressiveness of the features and the computational efficiency by designing special feature extraction methods for different data types.
[0043] Module 102 uses a hierarchical processing method for feature extraction of clinical data. First, for vital sign data (such as body temperature, pulse, blood pressure, blood oxygen saturation, etc.), the module extracts statistical features through time series analysis, including mean value, fluctuation amplitude, trend change, etc., and uses wavelet transform to extract frequency domain features to capture subtle features of pathological changes. For blood test data, the module automatically parses laboratory indicators, such as white blood cell count, C-reactive protein level, etc., and uses normalization and dimensionality reduction techniques to compress multidimensional indicators into a comprehensive feature vector representing the patient's inflammatory state. Symptom description data is processed by natural language processing (NLP) technology, first converting unstructured text data into structured semantic vectors, and then combining sentiment analysis and context modeling technology (such as Transformer-based language model) to extract key semantic features related to pulmonary infectious diseases, such as cough frequency, chest pain severity, etc.
[0044] For feature extraction of medical imaging data, module 102 uses deep learning technology to perform feature encoding on chest X-rays, CT scan images, and ultrasound images through a pre-trained convolutional neural network (CNN) model. The module supports multi-level feature extraction, extracting basic image features such as edges and textures at a low level, while extracting semantic features that can characterize the lesion area at a high level, such as the morphology, density, and position of abnormal shadows in lung images. To improve the robustness of the extracted features, the module also integrates data enhancement technology, including image rotation, flipping, noise addition, etc., to ensure the generalization ability of the model under limited sample conditions. In addition, the module can dynamically adjust the network structure according to the image type. For example, for CT scan images, a three-dimensional convolution layer is added to extract three-dimensional spatial features; for ultrasound images, a specific filter is introduced to eliminate artifact interference.
[0045] For feature extraction of environmental factor data, module 102 combines statistical analysis and time series modeling technology. After standardized processing, information such as air quality index (AQI), meteorological data (temperature, humidity, wind speed), etc., extracts time series features, such as changing trends, periodic fluctuations, etc. At the same time, the module identifies potential correlation features of disease outbreaks in the region through cluster analysis of regional infectious disease prevalence data, such as the characteristic distribution of high transmission risk areas. By introducing graph network analysis methods, the module can also extract spatial correlation features from regional disease transmission networks to further improve the comprehensiveness of feature expression.
[0046] The feature extraction module 102 not only supports the extraction of the above-mentioned standard features, but also can expand customized feature extraction strategies according to specific application scenarios. For example, the module can preliminarily align features from different modalities in the extraction stage through a joint attention mechanism to reduce the complexity of data fusion in subsequent modules. This method effectively reduces the redundancy between high-dimensional features while enhancing the complementarity of features. In addition, the module supports a feature importance evaluation mechanism that can dynamically adjust the priority of feature extraction according to specific prediction targets, thereby improving the adaptability and intelligence level of the system.
[0047] Through the above method, the feature extraction module 102 can not only efficiently and accurately extract the core features in the multimodal data, but also has a high degree of flexibility and scalability, laying a solid foundation for the multimodal data fusion and disease prediction of the entire system.
[0048] Furthermore, the feature extraction module includes a dynamic parameter processing unit, which is used to normalize and filter noise on physical sign data and blood test data according to the patient's physiological characteristics and medical history; the dynamic parameter processing unit is used to dynamically adjust the feature extraction criteria in combination with the patient's age, gender and known health status to ensure that the extracted physiological indicator feature vector can fully reflect the patient's pathological change trend.
[0049] The dynamic parameterization processing unit in the feature extraction module is an important part of this embodiment. It aims to dynamically adjust the processing standard of feature extraction by combining the patient's individual physiological characteristics and previous medical history, so as to ensure that the normalization processing and noise filtering of physical sign data and blood test data can accurately adapt to the specific health status of each patient. In the feature extraction process, this unit uses a series of intelligent analysis and optimization steps to make the extracted physiological indicator feature vector more accurately reflect the patient's pathological change trend.
[0050] The working mechanism of the dynamic parameterized processing unit is centered on individualization. After receiving the patient's vital signs data (such as body temperature, pulse, blood pressure and blood oxygen saturation) and blood test data (such as white blood cell count, C-reactive protein level, etc.), the system first extracts the patient's key physiological characteristics, including age, gender, and previous health records. This personalized information is used as parameter input to adjust the feature processing standards. For example, for elderly patients, the system may relax the normalized range of blood pressure fluctuations, while for pregnant women, the system may give priority to the dynamic changes in blood oxygen saturation to ensure that data processing can meet the special needs of different physiological states.
[0051] In terms of normalization processing, the dynamic parameterization processing unit combines the above personalized parameters to dynamically adjust the normalization range by calculating the deviation between each patient's physiological baseline value and the current collected data. For example, based on the patient's age and health status, the reference range of normal body temperature is determined, and then the body temperature data is linearly transformed to map it to the standardized feature space. This dynamically adjusted normalization process not only retains the individual differences in vital sign data, but also minimizes information loss caused by a single standardization rule.
[0052] At the same time, the unit also has built-in noise filtering for vital sign data and blood test data. By combining the patient's health records and the change pattern of the current data, the system can identify possible acquisition errors or equipment noise. For example, when a short-term abnormally low value appears in the blood oxygen saturation data, the system will determine whether it is acquisition noise through historical trend analysis; if it is confirmed to be noise, it will be filtered or corrected. In addition, for abnormal values in blood test data, the system can combine the patient's previous test results to determine whether the current data is reasonable, and correct small deviations through interpolation or statistical methods.
[0053] The core innovation of the dynamic parameterized processing unit is the ability to dynamically adjust the feature extraction criteria over time. When the patient's health status changes significantly (for example, from a healthy state to a sub-healthy state), the system recalculates the normalization and noise filtering parameters to adapt to the new health status. For example, when a patient is detected to have a persistent fever or an increased inflammatory response, the system adjusts the feature extraction rules for white blood cell counts, giving priority to the features of abnormal changes, and ensuring that the final generated physiological indicator feature vector can fully reflect the current condition.
[0054] Through this dynamic adjustment based on personalized parameters, the dynamic parameterized processing unit not only improves the accuracy and applicability of feature extraction, but also ensures that the extracted physiological indicator feature vector can accurately capture the subtle trends of the patient's pathological changes, providing high-quality input for subsequent dynamic fusion and predictive modeling. This design enables the system to operate efficiently in complex clinical scenarios while meeting the actual needs of personalized medicine.
[0055] Furthermore, the feature extraction module includes a multi-scale feature extraction unit for extracting features of different resolutions from chest X-rays, CT scan images and ultrasound images; the multi-scale feature extraction unit generates an image feature vector that can distinguish between normal tissue and lesion areas by performing block processing, edge enhancement and texture analysis on the image data, and can also identify potential lung infection areas.
[0056] The multi-scale feature extraction unit in the feature extraction module is an important technical link of the present invention, which is specially designed for efficiently extracting features of different resolutions from medical image data. Its core goal is to generate high-precision image feature vectors through multi-scale processing of chest X-rays, CT scan images and ultrasound images, providing strong support for subsequent prediction and analysis. Through the synergistic effect of block processing, edge enhancement and texture analysis, this unit can not only distinguish normal tissue from lesion areas, but also identify possible lung infection areas.
[0057] The work of the multi-scale feature extraction unit begins with the block processing of the input image data. Chest X-rays, CT scan images, and ultrasound images usually have different resolutions and data structures, so the system first divides the image into multiple image blocks with fixed or adaptive sizes. This block strategy can capture the details of local features in the image while ensuring computational efficiency. For example, for high-resolution CT images, block processing can focus on microscopic areas of small lung nodules or abnormal shadows, while for low-resolution X-rays, a larger range of tissue structures can be covered to ensure the extraction of global features.
[0058] After the block processing, the multi-scale feature extraction unit performs edge enhancement on each image block. This process highlights the edge features of tissue boundaries and lesion areas by increasing the contrast of the edge areas in the image. For example, using an algorithm based on gradient calculation, the system can clearly depict the boundaries of the lung lobes, the morphology of the trachea, and the contours of the lesion area. For CT scan images, edge enhancement can reveal subtle lesion edges, such as the range and morphology of ground-glass shadows; while for ultrasound images, edge enhancement helps to distinguish different tissue interfaces, thereby improving the accuracy of feature extraction.
[0059] After completing edge enhancement, the system further performs texture analysis to capture hidden patterns and detail features in the image data. The focus of texture analysis is to use the grayscale distribution, directionality, and repetitive patterns of the image to extract high-level features that can characterize the differences between normal tissue and lesion areas. For example, by calculating texture indicators such as local entropy and grayscale co-occurrence matrix of the image, the system can identify whether there are typical infectious features in the lung image, such as uneven tissue density or irregular texture patterns. This texture analysis can extract predictive features from the global information of the image and generate a high-dimensional image feature vector.
[0060] Another key innovation of the multi-scale feature extraction unit is its ability to identify potential infection areas. By combining the results of block processing, edge enhancement, and texture analysis, the system is able to mark areas where infection may occur and generate independent feature vectors for these areas. This approach can significantly improve the system's ability to detect infectious lung diseases at an early stage. For example, in CT images, the system can mark areas with a high risk of infection, such as areas with irregular boundaries and high-density textures; in X-rays, it can focus on abnormally bright or shadowed areas; in ultrasound images, it can detect irregular patterns of tissue echoes.
[0061] Through the above steps, the multi-scale feature extraction unit can not only adapt to the characteristics of different types of medical images, but also comprehensively extract key features from a multi-scale perspective. These feature vectors are further processed and fused to provide accurate and rich input data for subsequent predictive modeling.
[0062] Furthermore, the feature extraction module includes a spatiotemporal correlation analysis unit, which is used to extract an environmental impact feature vector reflecting the patient's exposure risk by combining the patient's geographic location, air quality index and regional infectious disease prevalence data; the spatiotemporal correlation analysis unit is used to dynamically construct an infectious disease transmission risk network based on geographic location, and predict the specific impact of environmental factors on the patient's infection risk by identifying the association between high-risk areas and patient locations.
[0063] The spatiotemporal correlation analysis unit in the feature extraction module aims to extract environmental impact feature vectors that reflect the patient's exposure risk by dynamically analyzing the spatial and temporal information of the patient's environment, providing key support for the accurate prediction of pulmonary infectious diseases. The unit combines the patient's geographic location data, air quality index, regional infectious disease prevalence and other environmental variables to comprehensively assess the patient's infection risk under specific spatiotemporal conditions, and dynamically construct an infectious disease transmission risk network to identify potential health threats.
[0064] During the implementation process, the spatiotemporal correlation analysis unit first receives the patient's real-time geographic location information, which is usually obtained through the global positioning system (GPS) or network positioning technology, and the accuracy can meet the needs of the city level or even the community level. Combined with this location information, the system obtains environmental data of the corresponding area from external data sources (such as air quality monitoring platforms or public health databases), including air quality index, meteorological conditions (such as temperature, humidity, wind speed) and the prevalence of infectious diseases in the current area (such as the number of confirmed cases, transmission index, etc.). These data are mapped through unified spatiotemporal coordinates for subsequent correlation analysis.
[0065] The spatiotemporal correlation analysis unit further processes and analyzes the environmental data to extract key spatiotemporal features. For example, for the air quality index, the system evaluates the health risks of patients exposed to low air quality environments for a long time by analyzing its historical change trends and abnormal fluctuations in the current value; for the prevalence of regional infectious diseases, the system combines the infectious disease transmission model to dynamically estimate the potential spread of the virus in the region and identify areas where high-risk people gather. Through the integration of this multi-dimensional data, the system can accurately assess the health threats in the patient's environment.
[0066] The spatiotemporal correlation analysis unit is able to dynamically construct a geographic location-based infectious disease transmission risk network. The network is centered on the patient's location and connected to high-risk nodes in the surrounding areas. Each node represents a specific geographic area, and the edge weights reflect the mobility or transmission intensity between regions. For example, by combining traffic flow data, population density, and case distribution, the system is able to identify potential paths of transmission from high-risk areas to the patient's area. The dynamic construction of the risk network allows the system to update the transmission relationship in real time, especially in the case of rapidly changing infectious disease epidemics, providing an accurate basis for predicting the patient's infection risk.
[0067] Finally, the spatiotemporal correlation analysis unit generates environmental impact feature vectors based on the above analysis results. These feature vectors not only contain the environmental data of the patient's current location, but also comprehensively consider the temporal dynamics and regional transmission risks. For example, for a patient who has been in a highly polluted air environment for a long time, the feature vector will particularly highlight the negative impact of air quality on health; for a patient in an area with a high risk of infectious disease transmission, the feature vector will show the number of cases in the area, the speed of transmission, and the impact of the virus type.
[0068] Through this design, the spatiotemporal correlation analysis unit can extract accurate feature information when environmental data is complex and dynamically changing, providing key support for subsequent multimodal data fusion and disease prediction modeling. This analysis method based on spatiotemporal correlation can effectively solve the problem that traditional systems are difficult to dynamically identify environmental risks, while greatly improving the ability to predict patient infection risks.
[0069] The dynamic fusion module 103 is used to use the attention mechanism to adaptively assign weights to the physiological indicator feature vector, the image feature vector and the environmental impact feature vector, and fuse them into a unified multimodal feature representation.
[0070] The dynamic fusion module 103 is one of the core technical modules of the prediction system provided in this embodiment. Its main function is to fuse the multimodal feature vectors from the feature extraction module, and adaptively assign the importance weight of each feature by using the attention mechanism, so as to generate a unified multimodal feature representation. This module not only realizes the effective integration of different modal data, but also significantly improves the quality of data expression and the prediction ability of the system.
[0071] The design of this module is based on the attention mechanism in deep learning. First, each feature vector is analyzed independently. Specifically, the physiological indicator feature vector, image feature vector, and environmental impact feature vector are processed by attention networks respectively. These networks generate attention weights related to their semantics by contextual modeling of the input features. For example, in certain situations, such as the early stages of a disease, physiological indicator features may dominate, while image features increase in weight when the lesion area is obvious, and environmental impact features have a higher influence when seasonal infectious diseases are prevalent.
[0072] The module is implemented using an architecture based on a multi-head attention mechanism. Each head processes a set of features independently, extracts feature importance information from different angles, and then integrates the output results of multiple heads into a unified fusion feature representation through linear transformation and weighted summation operations. The advantage of this design is that it can capture complex associations between multimodal data, such as the temporal consistency of fluctuations in physiological indicators and lesion diffusion in imaging features, or the potential impact of environmental changes on the aggravation of clinical symptoms.
[0073] During the fusion process, the module further introduces residual connections and layer normalization techniques to avoid the gradient vanishing problem and improve the training stability of the model. In addition, to enhance the interpretability of feature representation, the module supports visualization-based attention distribution display functions, such as using heat maps to display the weight changes of a specific modality in a specific time period. This function facilitates medical experts to verify the fusion results and enhances the acceptability of the system in clinical applications.
[0074] In order to improve the real-time performance and efficiency of dynamic fusion, the module also supports an incremental feature update mechanism. When new data flows in, there is no need to recalculate all fusion features. The module can efficiently update the fusion results based on the existing results. This feature is particularly suitable for scenarios that require continuous monitoring, such as real-time monitoring of changes in a patient's condition.
[0075] The dynamic fusion module 103 is also highly flexible. For example, the module can dynamically adjust the parameters of the attention mechanism to meet the needs of different application scenarios, such as strengthening the weight of environmental factors in high-risk environments and highlighting the impact of long-term physiological indicators in chronic disease management. In addition, the module supports multi-task learning and can simultaneously optimize multiple objective functions during the fusion process, such as taking into account both the prediction of disease occurrence probability and the analysis of disease deterioration trend.
[0076] Through these implementations, the dynamic fusion module 103 achieves efficient integration and dynamic weight adjustment of multimodal data, which not only improves the prediction performance of the system, but also has high scalability and practical application value.
[0077] Furthermore, the dynamic fusion module is specifically used for:
[0078] Obtain physiological index feature vector , image feature vector And the environmental impact feature vector is , these feature vectors are generated from the feature extraction modules respectively.
[0079] According to the following formula 1, the initial modal attention weight corresponding to the physiological indicator feature vector is calculated :
[0080] ;
[0081] in, is the sigmoid activation function; It is a trainable query vector used to select core information related to physiological indicator features; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector;
[0082] According to the following formula 2, the initial modal attention weight corresponding to the image feature vector is calculated :
[0083] ;
[0084] in, is the sigmoid activation function; is a trainable query vector; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector;
[0085] According to the following formula 3, the initial modal attention weight corresponding to the environmental impact feature vector is calculated :
[0086] ;
[0087] in, is the sigmoid activation function; is a trainable query vector; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector;
[0088] According to the following formula 4-6, calculate the modified modal attention weight corresponding to the physiological indicator feature vector , corresponding to the modified modal attention weight of the image feature vector and the modified modal attention weight corresponding to the environmental impact feature vector :
[0089] ;
[0090] ;
[0091] ;
[0092] in, is the correction factor, the recommended value is 2;
[0093] According to the following formula 7, the fused multimodal feature representation is calculated :
[0094] ;
[0095] in, is the weight matrix corresponding to the physiological index feature vector; is the weight matrix corresponding to the image feature vector; is the weight matrix corresponding to the environmental impact eigenvector; is the fused bias vector; As the activation function, the following formula 8 is used:
[0096] ;
[0097] in, is the input variable.
[0098] The prediction modeling module 104 performs a time series analysis on the multimodal feature representation to obtain the time series characteristics of the multimodal feature representation; and predicts the probability value of the patient suffering from a lung infectious disease and the possible disease progression trend based on the time series characteristics.
[0099] The prediction modeling module 104 is the final key module of the prediction system provided in this embodiment, which is responsible for using the unified multimodal feature representation generated from the dynamic fusion module to perform time series analysis to predict the probability value of the patient suffering from a pulmonary infectious disease and the possible progression trend of the disease. The design of the module is based on advanced time series modeling technology to ensure that the dynamic pattern of feature changes over time can be captured efficiently and accurately, providing a scientific basis for disease prediction.
[0100] This module introduces recurrent neural networks (RNNs) and their variants (such as long short-term memory networks (LSTMs) or gated recurrent units (GRUs)) to model time series of multimodal features. These models can preserve data correlations over long time spans and capture subtle patterns of feature changes over time. For example, the LSTM network can effectively handle long-term dependencies by introducing memory units, which is suitable for analyzing the long-term trend of patient conditions, while the GRU structure is relatively lightweight and suitable for scenarios with high real-time requirements.
[0101] The module first segments the multimodal feature representation on the time axis, and each segment of data corresponds to a feature flow within a fixed time window. For example, for patients with rapidly changing conditions, the time window can be set to several hours, while for patients with chronic diseases, the window can be adjusted to several days or even weeks. After normalization, the time series data is input into the modeling network, and the network automatically extracts dynamic features in the time series, including patterns such as trends, fluctuations, and periodicity. By introducing an adaptive learning rate optimization algorithm, such as the Adam optimizer, the module can dynamically adjust the learning rate of the model to improve convergence speed and prediction performance.
[0102] In addition, in order to enhance the ability to predict the trend of disease progression, the module also combines sequence-to-sequence (Seq2Seq) modeling technology. In this framework, the input captures historical time series features, and the output generates disease probability prediction values and trend curves in future time periods. For example, in the early stages of the disease, the system can predict whether the risk of infection will increase sharply in the next few days based on existing data, and the clinical course stage that the patient may enter.
[0103] The module also integrates a multi-task learning mechanism, which enables it to optimize multiple prediction targets at the same time. For example, the module can simultaneously predict the probability of disease occurrence and progression rate, and even further predict the key nodes where the disease may occur (such as severe or recovery stages). This multi-task learning framework improves modeling efficiency by sharing the underlying network structure, while ensuring collaborative optimization between different prediction tasks.
[0104] To further enhance the interpretability of the model, the module also introduces an attention mechanism to identify the time segments and feature dimensions that contribute most to the prediction results. For example, by visualizing the distribution of attention weights, medical professionals can intuitively understand which time periods of data or which features (such as temperature fluctuations in clinical indicators or changes in air quality in environmental factors) play a key role in disease prediction. This design that enhances interpretability has greatly improved the credibility and acceptance of the system in actual medical applications.
[0105] The module also supports a real-time updated prediction mechanism, making it suitable for dynamically changing patient data. For example, when new clinical data or imaging data flows in, the module is able to quickly update the prediction model through an incremental learning algorithm without completely retraining. This feature is particularly suitable for clinical scenarios with high-frequency monitoring, such as patient management in intensive care units.
[0106] The predictive modeling module is able to perform personalized prediction optimization based on scenario requirements. For example, for pediatric patients, the module can adjust model parameters based on age and physiological differences; for regional infectious disease outbreak scenarios, the module can increase the weight of environmental factors to more accurately predict large-scale infection trends. This flexible configuration capability makes the module suitable not only for individual patient predictions, but also for public health management and infectious disease prevention and control.
[0107] Through these implementations, the predictive modeling module 104 can fully exploit the temporal information of multimodal features to achieve high-precision prediction of disease occurrence and development, while having high scalability and adaptability.
[0108] Furthermore, the prediction modeling module performs time series analysis on the multimodal feature representation based on a neural network model, and ultimately predicts the probability value of the patient suffering from a pulmonary infectious disease and the possible disease progression trend, wherein the neural network model includes a time feature extraction part, a dynamic context modeling part, and a result prediction part;
[0109] The input of the temporal feature extraction part is a multimodal feature representation sequence, and the output is a temporal feature representation sequence; the temporal feature extraction part includes a multi-layer hierarchical convolution unit, wherein each layer extracts features of different time granularities through a convolution operation of a fixed time window; the first layer receives the original feature sequence and captures the feature change pattern in a short time, and the subsequent layers recursively aggregate the feature output of the previous layer to capture longer-term dependencies;
[0110] The input of the dynamic context modeling part is the time feature representation sequence provided by the time feature extraction part, and the output is the context feature representation; the dynamic context modeling part is composed of an autoregressive module with a cross-attention mechanism, and the autoregressive module is used to capture the dynamic association between historical information and the current state in the time dimension, wherein the cross-attention mechanism calculates the correlation weights between the current moment feature and all historical features, and strengthens the highly correlated features, so that this part can accurately model the time dependency and mutation events in the development of complex diseases;
[0111] The input of the result prediction part is the context feature representation provided by the dynamic context modeling part, and the output is the probability value of the patient suffering from a pulmonary infectious disease and the disease progression trend; the result prediction part includes a multi-task output unit, which is based on the design of shared parameters and optimizes the following two target tasks at the same time:
[0112] The first task is to predict the probability of disease occurrence, using a fully connected network with an activation function to generate probability values;
[0113] The second task is to predict the progression of the disease and use a recursive network to generate a disease trend curve for several future time steps.
[0114] The design of the prediction modeling module is based on the neural network model. Through phased processing and multi-task optimization, a comprehensive time series analysis of multimodal feature representation is performed, and finally the probability prediction of patients suffering from pulmonary infectious diseases and the analysis of disease progression trends are realized. The module consists of a time feature extraction part, a dynamic context modeling part, and a result prediction part. Each part has unique functions and technical implementations, and they work together through a close logical relationship.
[0115] The temporal feature extraction part is the first stage of the entire model, responsible for extracting feature representations of different time granularities from the multimodal feature representation sequence. This part processes the input feature sequence through multiple layers of hierarchical convolutional units, each of which operates with a fixed time window size to capture the feature change pattern within a specific time span. The first layer directly receives the original multimodal feature sequence and focuses on extracting fast-changing features in a short period of time, such as short-term fluctuations in the patient's body temperature or blood oxygen saturation. Subsequent layers recursively aggregate the output features of the previous layer, allowing this part to gradually capture trend information with longer-term dependencies, such as long-term changes in lesion expansion in lung images or the temporal evolution trend of regional infectious disease epidemics. This hierarchical convolution design ensures that the temporal feature extraction part can cover everything from short-term dynamics to long-term trends.
[0116] The dynamic context modeling part receives the temporal feature representation sequence generated by the temporal feature extraction part, and captures the dynamic association between historical information and the current state in the temporal dimension through an autoregressive module with a cross-attention mechanism. The autoregressive module gradually integrates the features of the current moment with the features of all historical moments through recursive processing to generate contextual feature representations. The cross-attention mechanism is particularly critical in the implementation of this part. It calculates the correlation weights between the features of the current moment and all historical features, and strengthens the historical features that are highly relevant to the current moment as part of the contextual feature representation. For example, when a patient's recent vital sign data and medical images show abnormal changes, the cross-attention mechanism can automatically identify and focus on these highly relevant features, thereby capturing potential key patterns of disease occurrence or progression. In addition, the mechanism is specially designed for the time dependency of complex diseases, and can accurately model the impact of sudden pathological events, such as the sudden outbreak of infectious diseases or acute deterioration of lesions.
[0117] The result prediction part receives the context feature representation provided by the dynamic context modeling part and completes two core tasks simultaneously through a multi-task output unit. The first task is to predict the probability of a patient suffering from a lung infectious disease. This task is implemented through a fully connected network with an activation function, which uses the high-dimensional information represented by the context feature to generate the probability value of the disease occurrence, thereby providing a direct basis for medical decision-making. The second task is to predict the progression trend of the disease. This task generates a trend curve of the disease in several future time steps through a recursive network. For example, the trend curve can describe the further changes in the patient's body temperature, the expansion speed of the lesion area, or the dynamic changes in the risk of regional infectious diseases. The multi-task output unit ensures that the optimization objectives of the two tasks work together in the model through the design of shared parameters, thereby improving the accuracy and robustness of the overall prediction.
[0118] Through the organic combination of time feature extraction, dynamic context modeling and outcome prediction, the predictive modeling module can comprehensively analyze the multimodal data of patients and accurately capture the dynamic characteristics of disease development in the time dimension. This phased, collaboratively optimized design enables the module to adapt to the complex situations of different patients and environments, providing strong technical support for disease prediction and trend analysis, while greatly improving the practicality and accuracy of the model.
[0119] Furthermore, the temporal feature extraction part includes a time window adaptive adjustment unit, which is used to dynamically adjust the time window size of each layer of hierarchical convolutional units according to the changes in the input multimodal feature representation sequence; wherein the time window adaptive adjustment unit selects a shorter time window for rapidly changing features to capture fine-grained dynamic patterns, and selects a longer time window for slowly changing features to enhance the modeling capability of long-term dependencies by analyzing the rate and amplitude of feature changes in the input sequence.
[0120] The time window adaptive adjustment unit of the temporal feature extraction part is an intelligent dynamic optimization component, which is designed to flexibly adjust the time window size of each layer of hierarchical convolutional units according to the actual changes in the input multimodal feature representation sequence. This design enables the temporal feature extraction part to achieve a balance between capturing rapidly changing short-term dynamic patterns and long-term stable trends, thereby significantly improving the accuracy and robustness of predictions.
[0121] The core of the time window adaptive adjustment unit is to analyze the change rate and amplitude of the input feature sequence in real time to determine the optimal time window size. The input sequence may include multimodal data such as the patient's vital signs data, medical imaging features, and environmental factors. These data usually show different change characteristics. For example, the body temperature and pulse in the vital signs data may fluctuate violently in a short period of time, while the medical imaging features reflect the long-term trend of lesion changes; environmental factors such as the air quality index may have medium-frequency changes. This heterogeneity determines that a single fixed time window is difficult to adapt to the dynamic characteristics of multimodal features at the same time.
[0122] The time window adaptive adjustment unit dynamically optimizes the time window of each layer of hierarchical convolutional units by analyzing the change pattern in the feature sequence. When a high rate of change in the feature sequence is detected, the unit assigns a shorter time window to the corresponding convolutional unit. For example, for rapidly changing vital sign data, such as when a patient's body temperature rises sharply, the reduction of the time window can more sensitively capture important details of changes in the short term. In contrast, when a smaller amplitude and slower rate of change in the feature sequence is detected, the unit assigns a longer time window so that the convolutional unit can integrate features over a longer time span, thereby effectively modeling long-term dependencies, such as chronic lesion trends in patient lung images or slow growth patterns of regional infectious disease transmission.
[0123] In order to achieve this dynamic adjustment process, the time window adaptive adjustment unit has a built-in real-time monitoring and decision-making mechanism. This mechanism generates quantitative indicators of feature changes by calculating the local gradient, change amplitude and frequency distribution of the input feature sequence. Based on these indicators, the adjustment unit optimizes the time window. For example, for sequences with large changes, the system automatically narrows the time window to focus on high-frequency details; for sequences with low frequency but continuous changes, the system selects a larger time window to capture global trends. The results of this dynamic adjustment are directly passed to the corresponding hierarchical convolutional unit to ensure that the feature extraction operation of each layer can adapt to the dynamic characteristics of the current input features.
[0124] In addition, the time window adaptive adjustment unit ensures the continuity and consistency of the output of the entire time feature extraction part through multi-level collaborative optimization. Even if there are significant differences between multimodal feature sequences, the adjustment unit can independently adjust the time window according to the characteristics of each modality while maintaining the synergistic relationship between cross-modal features. For example, in the feature extraction process of vital sign data, medical imaging data, and environmental factor data, the system can allocate a shorter time window to vital sign data to capture rapid changes, a medium-length time window to medical imaging data to focus on medium-term changes, and a longer time window to environmental factor data to capture long-term trends.
[0125] Through this design, the time window adaptive adjustment unit realizes efficient processing of multimodal features in the temporal feature extraction part, enabling it to flexibly respond to feature changes of different modalities in scenes with complex and changing temporal dynamics. This not only enhances the model's ability to model rapid changes and long-term dependencies, but also provides higher-quality input features for subsequent dynamic context modeling and result prediction.
[0126] Furthermore, the cross-attention mechanism of the dynamic context modeling part is implemented by the following steps:
[0127] According to the following formulas 11-13, linear projection is performed on the time feature representation sequence to generate the query matrix, key matrix and value matrix, where:
[0128] ;
[0129] ;
[0130] ;
[0131] in, Indicates time The query vector of
[0132] is the projection weight matrix that maps the temporal features to the query space;
[0133] Indicates that the time feature sequence is at the current moment The eigenvector of
[0134] For historical moments The key vector of ;
[0135] is the projection weight matrix that maps the temporal features to the key space;
[0136] Represents the time feature sequence at a historical moment The eigenvector of
[0137] For historical moments A vector of values of ; is the projection weight matrix that maps temporal features to the space of values;
[0138] According to the following formula 14, the current time is calculated by the following formula The query vector and all historical moments Relevance weight of key vector :
[0139] ;
[0140] in, is the key vector Dimensions;
[0141] is a nonlinear compensation term based on feature differences, which is used to enhance historical features with significant feature differences and is calculated according to the following formula 15:
[0142] ;
[0143] in, is the activation function;
[0144] is the weight matrix of the difference features;
[0145] According to the following formula 16, the preliminary context feature representation of the current moment is calculated :
[0146] ;
[0147] in, For historical moments A vector of values of ; is the total duration of the time series;
[0148] According to the following formula 17, the context feature representation is calculated :
[0149] ;
[0150] in, is the weight matrix adjusted for dynamic weights; is the bias term for dynamic weight adjustment; is the sigmoid activation function.
[0151] The following is the reference implementation code of the neural network model:
[0152] import torch
[0153] import torch.nn as nn
[0154] import torch.nn.functional as F
[0155] # Define the time feature extraction part
[0156] class TimeFeatureExtractor(nn.Module):
[0157] def __init__(self, input_dim, num_layers, hidden_dim):
[0158] super(TimeFeatureExtractor, self).__init__()
[0159] self.layers = nn.ModuleList()
[0160] self.hidden_dim = hidden_dim
[0161] # Dynamic time window adjustment parameters
[0162] self.time_windows = nn.ParameterList([nn.Parameter(torch.tensor(1.0)) for _ in range(num_layers)])
[0163] for i in range(num_layers):
[0164] self.layers.append(nn.Conv1d(input_dim if i == 0 elsehidden_dim, hidden_dim, kernel_size=1))
[0165] def forward(self, x):
[0166] """
[0167] Input: x - sequence of multimodal feature representations [batch_size, time_steps, feature_dim]
[0168] Output: Time feature representation sequence
[0169] """
[0170] for i, layer in enumerate(self.layers):
[0171] # Dynamically adjust the time window based on input
[0172] time_window = torch.clamp(self.time_windows[i], min=1, max=x.size(1))
[0173] time_window = int(time_window.item()) # Convert to integer
[0174] x = F.pad(x, (0, 0, time_window - 1, 0)) # Dynamic time window padding
[0175] x = layer(x.transpose(1, 2)).transpose(1, 2) # convolution operation
[0176] x = F.relu(x) # activation function
[0177] return x
[0178] # Define the dynamic context modeling part of the cross-attention mechanism
[0179] class CrossAttention(nn.Module):
[0180] def __init__(self, feature_dim, attention_dim):
[0181] super(CrossAttention, self).__init__()
[0182] self.query_weight = nn.Linear(feature_dim, attention_dim)
[0183] self.key_weight = nn.Linear(feature_dim, attention_dim)
[0184] self.value_weight = nn.Linear(feature_dim, attention_dim)
[0185] self.diff_weight = nn.Linear(feature_dim, attention_dim)
[0186] self.sigmoid = nn.Sigmoid()
[0187] def forward(self, x):
[0188] """
[0189] Input: x - sequence of temporal feature representations [batch_size, time_steps, feature_dim]
[0190] Output: context feature representation
[0191] """
[0192] query = self.query_weight(x) # Generate query vector
[0193] key = self.key_weight(x) # Generate key vector
[0194] value = self.value_weight(x) # Generate value vector
[0195] # Calculate relevance weights
[0196] diff = torch.relu(self.diff_weight(x.unsqueeze(2) -x.unsqueeze(1)))
[0197] attention_scores = torch.matmul(query, key.transpose(-1, -2)) / torch.sqrt(torch.tensor(query.size(-1), dtype=torch.float))
[0198] attention_scores += torch.sum(diff, dim=-1)
[0199] attention_weights = F.softmax(attention_scores, dim=-1)
[0200] # Calculate context features
[0201] context = torch.matmul(attention_weights, value)
[0202] return context
[0203] # Define the result prediction part
[0204] class ResultPredictor(nn.Module):
[0205] def __init__(self, feature_dim, hidden_dim, num_steps):
[0206] super(ResultPredictor, self).__init__()
[0207] self.probability_predictor = nn.Sequential(
[0208] nn.Linear(feature_dim, hidden_dim),
[0209] nn.ReLU(),
[0210] nn.Linear(hidden_dim, 1),
[0211] nn.Sigmoid() # Output the probability of disease occurrence )
[0213] self.trend_predictor = nn.LSTM(feature_dim, hidden_dim, batch_first=True)
[0214] self.trend_output = nn.Linear(hidden_dim, num_steps)
[0215] def forward(self, context):
[0216] """
[0217] Input: context - context feature representation [batch_size, time_steps, feature_dim]
[0218] Output: Disease occurrence probability, disease progression trend
[0219] """
[0220] # Predict the probability of disease occurrence
[0221] prob = self.probability_predictor(context[:, -1, :]) # Use the last time step
[0222] # Predict disease progression trend
[0223] trend, _ = self.trend_predictor(context)
[0224] trend = self.trend_output(trend)
[0225] return prob, trend
[0226] # Define the overall neural network model
[0227] class DiseasePredictionModel(nn.Module):
[0228] def __init__(self, input_dim, time_feature_dim, attention_dim, hidden_dim, num_layers, num_steps):
[0229] super(DiseasePredictionModel, self).__init__()
[0230] self.time_feature_extractor = TimeFeatureExtractor(input_dim,num_layers, time_feature_dim)
[0231] self.cross_attention = CrossAttention(time_feature_dim, attention_dim)
[0232] self.result_predictor = ResultPredictor(attention_dim, hidden_dim, num_steps)
[0233] def forward(self, x):
[0234] """
[0235] Input: x - raw multimodal feature representation sequence [batch_size, time_steps,feature_dim]
[0236] Output: Disease occurrence probability, disease progression trend
[0237] """
[0238] # Time feature extraction
[0239] time_features = self.time_feature_extractor(x)
[0240] Dynamic context modeling
[0241] context = self.cross_attention(time_features)
[0242] # Result prediction
[0243] prob, trend = self.result_predictor(context)
[0244] return prob, trend
[0245] The process of training the above neural network model needs to be based on a standard supervised learning framework, mainly relying on labeled data sets and optimization algorithms to adjust the parameters of the model. First, a high-quality data set containing multimodal data should be prepared. Each record in the data set should include the patient's physiological indicators, medical imaging features, environmental factor features, and corresponding disease probability labels and disease progression trend labels.
[0246] Before training begins, the data needs to be preprocessed to ensure that the data of different modalities have the same time step and proper normalization. This can be achieved by standardizing the numerical data and interpolating or filling in missing values. Subsequently, the multimodal features are combined to form the input data of the model, and the distribution of the training dataset and the validation dataset is ensured to be consistent.
[0247] During the training process, the loss function is first defined. For the prediction of the probability of disease occurrence, the binary cross entropy loss function can be used; for the prediction of the disease progression trend, the mean square error loss function can be used. By combining the loss values of these two target tasks, a total loss function is formed to guide the optimization of the model.
[0248] Next, select a suitable optimization algorithm, such as the Adam optimizer, and set a suitable learning rate. Input the data into the model in batches, calculate the predicted value through forward propagation, and then perform backpropagation based on the result of the loss function to update the model parameters. The entire training process requires multiple rounds of iterations, each of which includes a complete traversal of the training data set, called an epoch. After each round, use the validation data set to evaluate the model performance, and use the loss value or accuracy of the validation set to determine whether the model is overfitting.
[0249] During the training process, the performance of the validation set can be monitored by setting an early stopping mechanism. When the performance of the validation set no longer improves, the training is terminated to prevent overfitting. After the training is completed, the data that was not involved in the training is used for testing to evaluate the comprehensive performance of the model in predicting the probability of disease occurrence and the trend of disease progression.
[0250] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A pulmonary infectious disease prediction system based on multimodal data fusion, characterized in that: include: A data acquisition module, for collecting clinical data, medical imaging data and environmental factor data of patients in real time, wherein the clinical data includes physical sign data, blood test data and symptom description data, the medical imaging data includes chest X-rays, CT scan images and ultrasound images, and the environmental factor data includes air quality index, meteorological data and regional infectious disease prevalence data; A feature extraction module, used to extract physiological index feature vectors from the clinical data, image feature vectors from the medical imaging data, and environmental impact feature vectors from the environmental factor data; A dynamic fusion module, used to use an attention mechanism to adaptively assign weights to the physiological indicator feature vector, the image feature vector, and the environmental impact feature vector, and fuse them into a unified multimodal feature representation; A prediction modeling module performs a time series analysis on the multimodal feature representation to obtain the time series features of the multimodal feature representation; and predicts the probability value of a patient suffering from a pulmonary infectious disease and a possible disease progression trend based on the time series features; The dynamic fusion module is specifically used for: Obtain physiological index feature vector , image feature vector And the environmental impact feature vector is ; According to the following formula 1, the initial modal attention weight corresponding to the physiological indicator feature vector is calculated : ; in, is the sigmoid activation function; is a trainable query vector; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector; According to the following formula 2, the initial modal attention weight corresponding to the image feature vector is calculated : ; in, is the sigmoid activation function; is a trainable query vector; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector; According to the following formula 3, the initial modal attention weight corresponding to the environmental impact feature vector is calculated : ; in, is the sigmoid activation function; is a trainable query vector; is the self-modal weight mapping matrix; and is the cross-modal interaction weight mapping matrix; is the bias vector; According to the following formula 4-6, calculate the modified modal attention weight corresponding to the physiological indicator feature vector , corresponding to the modified modal attention weight of the image feature vector and the modified modal attention weight corresponding to the environmental impact feature vector : ; ; ; in, is the correction factor; According to the following formula 7, the fused multimodal feature representation is calculated : ; in, is the weight matrix corresponding to the physiological index feature vector; is the weight matrix corresponding to the image feature vector; is the weight matrix corresponding to the environmental impact eigenvector; is the fused bias vector; As the activation function, the following formula 8 is used: ; in, is the input variable.
2. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The data acquisition module includes a data verification unit, which is used to perform real-time consistency check and integrity verification on the collected clinical data, medical imaging data and environmental factor data during the data acquisition process, specifically including: By comparing historical data trends, identifying outliers and detecting data loss, the accuracy and reliability of the collected data can be ensured. At the same time, it supports automatic generation of predicted values for missing data to supplement incomplete records, thereby improving the overall quality of the data.
3. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The data acquisition module includes a dynamic sampling unit, which is used to adaptively adjust the frequency of data acquisition and the priority of data types according to the patient's condition changes or the dynamic fluctuations of environmental factors, specifically including: When the patient's condition changes dramatically, the dynamic sampling unit prioritizes increasing the frequency of collecting vital sign data and medical imaging data. During periods of high risk of regional infectious diseases, it prioritizes collecting environmental factor data and symptom description data to implement flexible data collection strategies for different scenarios.
4. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The feature extraction module includes a dynamic parameterization processing unit, which is used to normalize and filter noise on physical sign data and blood test data according to the patient's physiological characteristics and medical history; the dynamic parameterization processing unit is used to dynamically adjust the feature extraction standard in combination with the patient's age, gender and known health status to ensure that the extracted physiological indicator feature vector can fully reflect the patient's pathological change trend.
5. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The feature extraction module includes a multi-scale feature extraction unit for extracting features of different resolutions from chest X-rays, CT scan images and ultrasound images; the multi-scale feature extraction unit generates an image feature vector that can distinguish normal tissue from lesion areas by performing block processing, edge enhancement and texture analysis on the image data, and can also identify potential lung infection areas.
6. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The feature extraction module includes a spatiotemporal correlation analysis unit for extracting an environmental impact feature vector reflecting the patient's exposure risk in combination with the patient's geographic location, air quality index, and regional infectious disease prevalence data; The spatiotemporal correlation analysis unit is used to dynamically construct an infectious disease transmission risk network based on geographic location, and predict the specific impact of environmental factors on the patient's infection risk by identifying the association between high-risk areas and patient locations.
7. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 1, characterized in that: The prediction modeling module performs time series analysis on the multimodal feature representation based on a neural network model, and ultimately predicts the probability value of the patient suffering from a pulmonary infectious disease and the possible disease progression trend, wherein the neural network model includes a time feature extraction part, a dynamic context modeling part, and a result prediction part; The input of the temporal feature extraction part is a multimodal feature representation sequence, and the output is a temporal feature representation sequence; the temporal feature extraction part includes a multi-layer hierarchical convolution unit, wherein each layer extracts features of different time granularities through a convolution operation of a fixed time window; the first layer receives the original feature sequence and captures the feature change pattern in a short time, and the subsequent layers recursively aggregate the feature output of the previous layer to capture longer-term dependencies; The input of the dynamic context modeling part is the time feature representation sequence provided by the time feature extraction part, and the output is the context feature representation; the dynamic context modeling part is composed of an autoregressive module with a cross-attention mechanism, and the autoregressive module is used to capture the dynamic association between historical information and the current state in the time dimension, wherein the cross-attention mechanism calculates the correlation weights between the current moment feature and all historical features, and strengthens the highly correlated features, so that this part can accurately model the time dependency and mutation events in the development of complex diseases; The input of the result prediction part is the context feature representation provided by the dynamic context modeling part, and the output is the probability value of the patient suffering from a pulmonary infectious disease and the disease progression trend; the result prediction part includes a multi-task output unit, which is based on the design of shared parameters and optimizes the following two target tasks at the same time: The first task is to predict the probability of disease occurrence, using a fully connected network with an activation function to generate probability values; The second task is to predict the progression of the disease and use a recursive network to generate a disease trend curve for several future time steps.
8. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 7, characterized in that: The temporal feature extraction part includes a time window adaptive adjustment unit, which is used to dynamically adjust the time window size of each layer of hierarchical convolutional units according to the changes in the input multimodal feature representation sequence; wherein the time window adaptive adjustment unit selects a shorter time window for rapidly changing features to capture fine-grained dynamic patterns, and selects a longer time window for slowly changing features to enhance the modeling capability of long-term dependencies by analyzing the rate and amplitude of feature changes in the input sequence.
9. The pulmonary infectious disease prediction system based on multimodal data fusion according to claim 7, characterized in that: The cross-attention mechanism of the dynamic context modeling part is implemented by the following steps: According to the following formulas 11-13, linear projection is performed on the time feature representation sequence to generate the query matrix, key matrix and value matrix, where: ; ; ; in, Indicates time The query vector of is the projection weight matrix that maps the temporal features to the query space; Indicates that the time feature sequence is at the current moment The eigenvector of is the key vector at the historical moment τ; is the projection weight matrix that maps the temporal features to the key space; Represents the time feature sequence at a historical moment The eigenvector of is the value vector of the historical moment τ; is the projection weight matrix that maps temporal features to the space of values; According to the following formula 14, the current time is calculated by the following formula The query vector and all historical moments Relevance weight of key vector : ; in, is the key vector Dimensions; is a nonlinear compensation term based on feature differences, which is used to enhance historical features with significant feature differences and is calculated according to the following formula 15: ; in, is the activation function; is the weight matrix of the difference features; According to the following formula 16, the preliminary context feature representation of the current moment is calculated : ; in, is the value vector of the historical moment τ; is the total duration of the time series; According to the following formula 17, the context feature representation is calculated : ; in, is the weight matrix adjusted for dynamic weights; is the bias term for dynamic weight adjustment; is the sigmoid activation function.
Citation Information
Patent Citations
Disease prediction method, device based on machine learning
CN113707309A
Aspect-level multi-modal sentiment analysis method based on collaborative attention fusion
CN115293170A
Personalized difference analysis method for retinopathy
CN117877692A
Cited By
Ultra-wide-angle fundus multi-mode medical image intelligent diagnosis and data management system
CN121414713A