Intelligent analysis method for pneumonia based on multi-modal data
By employing intelligent analysis methods for multimodal data, the problems of high misdiagnosis rate and prognostic uncertainty in interstitial pneumonia have been solved. Efficient data annotation and feature fusion have been achieved, improving the accuracy of diagnosis and prognostic analysis and supporting personalized treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUILIN MEDICAL UNIVERSITY
- Filing Date
- 2026-02-03
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies suffer from high misdiagnosis rates, significant uncertainty in prognostic analysis, lack of personalized tools, semantic gaps in multimodal data fusion, and poor generalization ability of AI models in the diagnosis and classification of interstitial pneumonia.
We employ intelligent analysis methods for multimodal data. By connecting the hospital information system and the electronic medical record system, we perform data cleaning and format unification. We combine expert knowledge bases and automated tools for fine-grained annotation, use a dynamic contrastive learning framework to achieve semantic association of multimodal features, combine cross-attention and gating mechanisms for feature fusion, and finally fine-tune the model to improve the accuracy of diagnosis and prognosis analysis.
It significantly improves the accuracy of interstitial pneumonia identification, reduces labor costs, increases annotation efficiency, enhances the effect of multimodal data fusion, provides a high-quality data foundation for subsequent models, and supports the development of personalized treatment plans.
Smart Images

Figure CN122224531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent pneumonia analysis technology, and in particular to an intelligent pneumonia analysis method based on multimodal data. Background Technology
[0002] Interstitial pneumonia (ISP) is a lung disease characterized by interstitial lung inflammation and progressive pulmonary fibrosis. Its etiology is complex and multifaceted, encompassing environmental, drug-related, and autoimmune factors, leading to different subtypes with varying prognoses. Common ISP and nonspecific ISP are the most common subtypes. Currently, the diagnosis of ISP primarily relies on high-resolution CT, but the imaging manifestations of unexplained interstitial pneumonia (UIP) and nonspecific interstitial pneumonia (NSIP) partially overlap, making differential diagnosis extremely difficult. The diagnostic challenge is further amplified when lesions in both lungs are atypical or mixed. Due to the significant differences in pathological manifestations and biological behavior among different subtypes, clinical treatment strategies and prognoses also vary considerably. Blindly treating patients without clearly identifying the ISP subtype can easily lead to unpredictable complications. Similarly, in terms of prognostic analysis, the progression of ISP is highly uncertain, lacking effective prognostic assessment methods and making it difficult to accurately predict disease progression. Traditional prognostic analyses are often based on clinical experience and limited indicators, which cannot fully consider the complexity of the disease, making it difficult to develop personalized treatment plans and affecting patients' treatment outcomes and quality of life.
[0003] In recent years, the rapid development of deep learning technology and artificial intelligence (AI) has brought about new changes and opportunities to the medical field, and has also provided new ideas for the intelligent diagnosis and prognostic analysis of interstitial pneumonia. AI models, with their powerful data analysis capabilities, can not only dynamically process various types of medical data, but also perform exceptionally well in the analysis of high-dimensional and complex data. At the same time, the country attaches great importance to technological innovation and development in the medical and health field. Driven by policy, the integration of AI and medicine has become a development trend. However, there are still technological shortcomings in the field of interstitial pneumonia that urgently need to be overcome.
[0004] Intelligent diagnosis of interstitial pneumonia faces challenges: UIP and NSIP subtypes overlap significantly in imaging, traditional diagnosis relies on physician experience, and is prone to misdiagnosis; prognostic analysis is highly uncertain, and personalized tools are lacking. Existing AI models are mostly based on single-modal data (e.g., HRCT only), with limited feature extraction capabilities, and multimodal fusion suffers from semantic gaps. While international research employs deep learning, the models have poor generalization ability, and domestic systems are functionally limited, failing to form a complete closed-loop process. There is an urgent need for a method that can integrate multimodal data and achieve end-to-end intelligent analysis. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent pneumonia analysis method based on multimodal data, which solves the technical problems of poor generalization ability and low data processing quality in existing pneumonia identification methods.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A pneumonia intelligent analysis method based on multimodal data, the method comprising the following steps:
[0008] Step 1: Connect the hospital information system and electronic medical record system to obtain pneumonia-related data;
[0009] Step 2: Perform quality cleaning and format standardization on pneumonia-related data;
[0010] Step 3: Employ a three-level collaborative mechanism of retrieval, generation, and verification, combined with an expert knowledge base and automated tools, to achieve refined data annotation. By retrieving reference annotations and conducting human-machine collaborative verification, annotation efficiency and consistency are improved.
[0011] Step 4: Based on the dynamic contrastive learning framework, construct a quality-aware alignment algorithm. By evaluating the dimensions of data clarity and resolution, dynamically weight samples of different quality, realize the semantic association of multimodal features, and reduce fusion bias.
[0012] Step 5: Use the data labeled in Step 3 to perform a VIT-based mask self-supervised pre-training model, and then extract features from the data in Step 4;
[0013] Step 6: Achieve fusion of image features and text features based on the cross-attention gating mechanism;
[0014] Step 7: Fine-tune the trained model, and then perform pneumonia identification and analysis on the fine-tuned model to obtain the analysis results.
[0015] Further, the specific process of step 1 is as follows: Set up the corresponding RPA data collection script, and through the RPA automated process, collect patient basic information, medical record data, examination and test results, professional knowledge, and Q&A data periodically or in real time. The relevant data collection includes medical information system data, pulmonary professional knowledge system data, and pulmonary disease consultation Q&A data. Medical information system data includes medical information data, laboratory information system focused test data, and medical image data managed by the medical image archiving and communication system. The medical information data integrates hospital administrative, financial, and medical business data. Laboratory information system focused test data includes test items, results, and reference ranges. Medical image data managed by the medical image archiving and communication system includes X-ray, CT, and MRI images. This provides information for the entire pneumonia diagnosis and treatment process, including medical history, diagnosis, treatment plan, examination and test results, and medical orders, constructing a basic record for pneumonia health care services and providing original data for subsequent consultations.
[0016] The lung professional knowledge system integrates clinical guidelines and expert consensus issued by authoritative organizations such as the WHO and the Chinese Medical Association to ensure that the diagnosis and treatment recommendations are scientific and effective. It also gathers cutting-edge research results from authoritative medical databases such as PubMed and Embase to provide knowledge support for complex diseases. The lung disease consultation Q&A data comes from real hospital consultation records, Q&A on health science popularization platforms and discussion posts on medical forums, reflecting actual needs, helping the public understand knowledge, and providing diverse references for large models.
[0017] Furthermore, the specific process of step 2 is as follows: image artifacts and text redundancy are processed based on heuristic methods to ensure data quality. Considering that the training corpus contains sensitive or personal information that threatens privacy and security, an algorithm based on named entity recognition is used to detect, delete or replace personal information such as names and addresses. Combined with the Transformer model and machine translation method, the privacy and security of pneumonia health data are guaranteed. For voice data in the field of pneumonia health, voice recognition technology is used to convert it into text and image data of examination reports. Text information is extracted through OCR technology, and the data is unified into text form for easy integration and processing in the knowledge management module of the platform.
[0018] Furthermore, the specific process of step 3 is as follows: a brain data annotation and governance framework is set up. The framework realizes intelligent retrieval and reference annotation of medical cases through a three-stage collaborative mechanism of retrieval, generation and verification. Combined with cross-modal association analysis technology, deep semantic analysis is performed on several dimensions of imaging features. A human-computer collaborative verification mechanism is introduced so that clinicians only need to review and fine-tune the standardized annotation results generated by the system to ensure the clinical accuracy and reliability of the annotation data.
[0019] The specific process of annotation is as follows: Enhanced image generation is performed using patient CT images and medical records. Similar cases are identified from the medical records database through retrieval and sorting to generate preliminary annotation results. Medical experts conduct final review and correction of the generated results. Cross-modal association annotation involves linking the preliminary annotations of CT images with corresponding diagnostic reports, automatically supplementing other modal information from dynamic MRI and clinical medical records, and comprehensively verifying lung images and medical record information to ensure consistency across multiple modalities. Diffusion synthesis utilizes a generative model to synthesize a diverse number of lung images. The synthesized results undergo quality review by medical experts to ensure their medical rationality and usability. The quality of the generated or processed images in terms of detail and clarity is evaluated, the logical consistency of the annotation results in terms of anatomical structure and pathological features is checked, and the semantic matching and alignment of different modal data is measured.
[0020] Furthermore, in step 4, a quality perception and contrast alignment several modal learning frameworks are set up. The frameworks construct a quality evaluation system that covers several dimensions of factors such as sharpness, resolution, annotation quality and modality specificity. A dynamic weighted contrast learning algorithm is developed, which adopts an adaptive weight allocation strategy to achieve a combination of reinforcement learning of samples and deep mining of low-quality samples.
[0021] The specific process involves inputting computed tomography (CT) images, dynamic magnetic resonance imaging (MRI), and structured or unstructured text. Sample construction and contrastive learning are then employed. The raw data is constructed into pairs of samples for contrastive learning. First, high-quality sample pairs are used, representing effective pairings where text and non-text features are aligned and closely correlated. Then, low-quality samples are compared. By bringing the features of high-quality sample pairs closer together and pushing the features of low-quality sample pairs further away, the quality of each sample pair is automatically evaluated and distinguished. After contrastive learning evaluation, samples are filtered and directed to their corresponding positions based on quality. Those identified as high-quality sample pairs are assigned weights greater than a predetermined limit. The weighted sample information is then aggregated into the loss function module. The loss value between the model's prediction and the actual situation is calculated to guide model optimization. After quality-aware sample screening and weighting, the calculated loss is more accurate and effective, directly driving the model to learn more robust multimodal feature representations.
[0022] Further, in step 5, texture and morphological features of lung tissue are extracted from CT images, and signal intensity and boundary features of lung lesion areas are extracted from MRI images. For medical record text feature extraction, preprocessing is performed first to desensitize patient information, followed by structured extraction to extract key fields from free text. Then, a BERT-based word segmenter is used to process the clinical medical record text. The BERT word segmenter can decompose text into appropriate sub-word units, which can better preserve the semantic information of the text. The processed sub-word sequence is input into the BERT model. The BERT model, through a bidirectional Transformer architecture, captures the contextual relationships between words in the text and generates text feature vectors with semantic information. Each medical record text, after being processed by the BERT model, will obtain a fixed-dimensional feature vector containing key information about the patient's symptoms, diagnosis results, and treatment history.
[0023] Furthermore, in step 6, the cross-attention mechanism enables features from different modalities to pay attention to each other and obtain the correlation between modalities. The gating mechanism can adaptively control the fusion ratio of features from different modalities. During the fusion process, the features of the three modalities are first aligned by the projection layer. Then, the cross-attention gating mechanism is used to allow each modality to dynamically capture the key information of other modalities, resulting in feature-level fusion. By calculating the attention weights, the model adaptively obtains more important image features and text features, thereby performing targeted fusion. The gating unit is introduced to learn the contribution of image features and text features respectively. The fused multimodal feature vector obtained under the synergistic effect of the gating mechanism and the cross-attention mechanism fully integrates image and text information, providing a more comprehensive and accurate feature representation for subsequent diagnosis and prognostic analysis.
[0024] The weights for the MRI modality are calculated as follows:
[0025]
[0026] Wherein, F_MRI represents the feature vector of MRI images, F_CT represents the feature vector of CT images, F_clinical represents the feature vector of clinical medical record text, W_a is a learnable weight matrix, b_a is a bias term that determines which vectors are concatenated into a longer joint feature vector, and finally, decision-level feature fusion is performed. The fused features and intra-modal features will enter the ViT decoder, and the final recognition and analysis will be performed through a gating mechanism and a fully connected layer.
[0027] Furthermore, in step 7, for the diagnostic task, the labels are divided into two categories: common interstitial pneumonia and nonspecific interstitial pneumonia. The model is fine-tuned using labeled data, with the cross-entropy loss function as the optimization objective. The cross-entropy loss function can effectively measure the difference between the model's prediction results and the true labels, making it very suitable for classification tasks. For the prognostic task, the labels are optimized into a binary classification task with the prediction of survival within 3 years as the boundary. Similarly, the cross-entropy loss function is used as the optimization objective, and the model is fine-tuned using the corresponding labeled data. During the fine-tuning process, the model will learn the correlation between features and patient prognosis. By continuously optimizing the parameters, the model's ability to predict patient prognosis is improved.
[0028] The present invention, by adopting the above-described technical solution, has the following beneficial effects:
[0029] This invention significantly improves the accuracy of analysis and recognition through hierarchical fine-tuning and cross-modal attention. It overcomes the problem of fragmented medical data through a standardized process integrating data cleaning, annotation standardization, and quality assessment, filling the gap in domestic interstitial pneumonia specialty datasets. The intelligent collaborative annotation mechanism improves annotation efficiency by more than 60% and reduces labor costs. The mode contrast learning alignment technology enhances the effect of multimodal data fusion, providing a high-quality data foundation for subsequent model recognition. Attached Figure Description
[0030] Figure 1 This is a framework diagram of the intelligent analysis and prediction model for pneumonia of this invention;
[0031] Figure 2 This is a diagram of the lung unstructured data annotation system of the present invention;
[0032] Figure 3 This is a flowchart of the modal alignment process for dynamic comparative learning in this invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are merely to provide the reader with a thorough understanding of one or more aspects of the present invention, and these aspects of the invention can be implemented even without these specific details.
[0034] like Figure 1-3 As shown, a pneumonia intelligent analysis method based on multimodal data is described, the method comprising the following steps:
[0035] Step 1: Connect the hospital information system and electronic medical record system to obtain pneumonia-related data. Set up corresponding RPA data collection scripts to collect patient basic information, medical record data, examination and test results, professional knowledge, and Q&A data on a regular or real-time basis through the RPA automated process. The relevant data collection includes data from the medical information system, data from the pulmonary professional knowledge system, and Q&A data from pulmonary disease consultations. The medical information system data includes medical information data, laboratory information system focused test data, and medical image data managed by the medical image archive and communication system. The medical information data integrates data from the hospital's administrative, financial, and medical operations. The laboratory information system focused test data includes test items, results, and reference ranges. The medical image archive and communication system managed medical image data includes X-ray, CT, and MRI images. This provides information for the entire process of pneumonia diagnosis and treatment, including medical history, diagnosis, treatment plan, examination and test results, and medical orders, constructing a basic record of pneumonia health care services and providing original data for subsequent consultations.
[0036] The lung professional knowledge system integrates clinical guidelines and expert consensus issued by authoritative organizations such as the WHO and the Chinese Medical Association to ensure that the diagnosis and treatment recommendations are scientific and effective. It also gathers cutting-edge research results from authoritative medical databases such as PubMed and Embase to provide knowledge support for complex diseases. The lung disease consultation Q&A data comes from real hospital consultation records, Q&A on health science popularization platforms and discussion posts on medical forums, reflecting actual needs, helping the public understand knowledge, and providing diverse references for large models.
[0037] Step 2: Quality cleaning and format standardization of pneumonia-related data. Heuristic methods are used to handle image artifacts and text redundancy to ensure data quality. Considering the inclusion of sensitive or personal information in the training corpus, posing a threat to privacy, named entity recognition algorithms are used to detect, delete, or replace personal information such as names and addresses. This is combined with Transformer models and machine translation methods to protect the privacy of pneumonia health data. For voice data in the pneumonia health field, speech recognition technology is used to convert it into text; for image data in examination reports, OCR technology is used to extract text information, unifying the data into text format for easy integration and processing in the platform's knowledge management module.
[0038] Step 3: Employing a three-tiered collaborative mechanism of retrieval, generation, and verification, combined with an expert knowledge base and automated tools, refined data annotation is achieved. Through retrieval of reference annotations and human-machine collaborative verification, annotation efficiency and consistency are improved. A brain data annotation and governance framework is established. This framework, through a three-stage collaborative mechanism of retrieval, generation, and verification, enables intelligent retrieval and reference annotation of medical cases. Combined with cross-modal association analysis technology, deep semantic analysis of multi-dimensional data of imaging features is performed. The introduction of a human-machine collaborative verification mechanism allows clinicians to only need to review and fine-tune the standardized annotation results generated by the system, ensuring the clinical accuracy and reliability of the annotated data.
[0039] The specific process of annotation is as follows: Enhanced image generation is performed using patient CT images and medical records. Similar cases are identified from the medical records database through retrieval and sorting to generate preliminary annotation results. Medical experts conduct final review and correction of the generated results. Cross-modal association annotation involves linking the preliminary annotations of CT images with corresponding diagnostic reports, automatically supplementing other modal information from dynamic MRI and clinical medical records, and comprehensively verifying lung images and medical record information to ensure consistency across multiple modalities. Diffusion synthesis utilizes a generative model to synthesize a diverse number of lung images. The synthesized results undergo quality review by medical experts to ensure their medical rationality and usability. The quality of the generated or processed images in terms of detail and clarity is evaluated, the logical consistency of the annotation results in terms of anatomical structure and pathological features is checked, and the semantic matching and alignment of different modal data is measured.
[0040] Step 4: Based on the dynamic contrastive learning framework, a quality-aware alignment algorithm is constructed. This algorithm dynamically weights samples of different quality levels by evaluating data clarity and resolution, achieving semantic association of multimodal features and reducing fusion bias. Several modal learning frameworks for quality-aware and contrastive alignment are established. These frameworks construct a quality evaluation system covering several dimensions, including clarity, resolution, annotation quality, and modality specificity. A dynamically weighted contrastive learning algorithm is developed, employing an adaptive weight allocation strategy to combine reinforcement learning of samples with deep mining of low-quality samples.
[0041] The specific process involves inputting computed tomography (CT) images, dynamic magnetic resonance imaging (MRI), and structured or unstructured text. Sample construction and contrastive learning are then employed. The raw data is constructed into pairs of samples for contrastive learning. First, high-quality sample pairs are used, representing effective pairings where text and non-text features are aligned and closely correlated. Then, low-quality samples are compared. By bringing the features of high-quality sample pairs closer together and pushing the features of low-quality sample pairs further away, the quality of each sample pair is automatically evaluated and distinguished. After contrastive learning evaluation, samples are filtered and directed to their corresponding positions based on quality. Those identified as high-quality sample pairs are assigned weights greater than a predetermined limit. The weighted sample information is then aggregated into the loss function module. The loss value between the model's prediction and the actual situation is calculated to guide model optimization. After quality-aware sample screening and weighting, the calculated loss is more accurate and effective, directly driving the model to learn more robust multimodal feature representations.
[0042] Step 5: Use the data labeled in Step 3 to perform a VIT-based masked self-supervised pre-training model, and then extract features from the data in Step 4. Texture and morphological features of lung tissue are extracted from CT images, and signal intensity and boundary features of lung lesion areas are extracted from MRI images. For medical record text feature extraction, preprocessing is performed first to desensitize patient information, followed by structured extraction to extract key fields from free text. Then, a BERT-based word segmenter is used to process the clinical medical record text. The BERT word segmenter can decompose text into appropriate sub-word units, which better preserve the semantic information of the text. The processed sub-word sequence is input into the BERT model. The BERT model, through a bidirectional Transformer architecture, captures the contextual relationships between words in the text and generates semantic text feature vectors. Each medical record text, after being processed by the BERT model, will obtain a fixed-dimensional feature vector containing key information about the patient's symptoms, diagnosis, and treatment history.
[0043] Step 6: Achieve fusion of image and text features based on cross-attention gating mechanism. The cross-attention mechanism enables features from different modalities to pay attention to each other and obtain the correlation between modalities. The gating mechanism can adaptively control the fusion ratio of features from different modalities. During the fusion process, the features of the three modalities are first aligned by projection layer mapping. Then, the cross-attention gating mechanism is used to allow each modality to dynamically capture the key information of other modalities, resulting in feature-level fusion results. By calculating attention weights, the model adaptively acquires more important image and text features, thus performing targeted fusion. The introduction of gating units to learn the contribution of image and text features respectively, and the fusion of multimodal feature vectors obtained under the synergistic effect of gating and cross-attention mechanisms, fully integrates image and text information, providing a more comprehensive and accurate feature representation for subsequent diagnosis and prognostic analysis.
[0044] The weights for the MRI modality are calculated as follows:
[0045]
[0046] Wherein, F_MRI represents the feature vector of MRI images, F_CT represents the feature vector of CT images, F_clinical represents the feature vector of clinical medical record text, W_a is a learnable weight matrix, b_a is a bias term that determines which vectors are concatenated into a longer joint feature vector, and finally, decision-level feature fusion is performed. The fused features and intra-modal features will enter the ViT decoder, and the final recognition and analysis will be performed through a gating mechanism and a fully connected layer.
[0047] Step 7: Fine-tune the trained model, and then perform pneumonia identification analysis on the fine-tuned model to obtain the analysis results. For the diagnostic task, the labels are divided into two categories: ordinary interstitial pneumonia and non-specific interstitial pneumonia. The model is fine-tuned using labeled data, with the cross-entropy loss function as the optimization objective. The cross-entropy loss function can effectively measure the difference between the model's prediction results and the true labels, and is very suitable for classification tasks. For the prognostic task, the labels are optimized into a binary classification task with the prediction of survival within 3 years as the boundary. The cross-entropy loss function is also used as the optimization objective. The model is fine-tuned using the corresponding labeled data. During the fine-tuning process, the model will learn the correlation between features and patient prognosis. By continuously optimizing the parameters, the model's ability to predict patient prognosis is improved.
[0048] Application Examples
[0049] The application scenario is a hospital radiology department, so the design aims to cater to the daily needs of radiologists. The user interface will implement visualization of 3D images in common formats (such as DICOM and NIFTI), visualization of regions of interest (ROIs), login interface design and management, and adjustment of window width, window level, and label transparency. In the interaction with the large model, the main functions to be implemented include automatic segmentation, intelligent diagnosis and prognostic analysis, calling the large model, quantitative analysis and display of ROIs, and generation and display of radiology reports.
[0050] The login interface design aims to build a responsive login interface based on Vue.js and the ElementUI component library, employing a form validation mechanism to verify the validity of the user's entered username and password. HTTP communication with the backend Flask server is implemented through the Axios library, using JWT (JSON Web Token) for user authentication. After successful login, the JWT is stored in the browser's LocalStorage and carried in the header of subsequent requests, maintaining user state and controlling access. Furthermore, users can easily modify their account passwords on the personal center page, effectively ensuring account privacy and system security.
[0051] The system visualizes lung images using the VTK.js library, providing 3D visualization capabilities. Following mainstream methods, it reconstructs CT and MRI images into three planes: transverse, sagittal, and coronal. The transverse plane cuts along the horizontal plane of the body, displaying the upper and lower layers of organs; the sagittal plane cuts along the left and right directions, showing the anterior-posterior relationship of organs; and the coronal plane cuts along the posterior-posterior direction, showing the left-right distribution of organs. When the user clicks the "Load Medical Image" button, a local file loading dialog box pops up, supporting the selection of DICOM and NIfTI format medical images. Once selected, the 3D structure of the image will be reconstructed in the interface. Users can interact with the 3D model using the mouse to rotate, scale, and translate, and can also set parameters such as transparency and color mapping, allowing doctors to observe lung structures and lesion areas from different angles.
[0052] ROI visualization, a crucial step in image quantification, transforms abstract image information into quantifiable and actionable structured data. This serves both clinical precision diagnosis and treatment and provides fundamental support for AI model training and medical research. In assisted diagnostic systems, the import and display of ROI labels are indispensable, allowing radiologists to quantify lesions and analyze disease progression. In this project, ROI primarily refers to lung and pulmonary fibrosis areas. The system mainly implements the following three functions: ① ROI Import: Clicking the "Load Labels" button brings up a file path loading dialog box. After selecting the label file, the ROI is displayed on the original image plane in pseudo-color overlay, defaulting to red; ② ROI Transparency Adjustment: The system supports ROI transparency adjustment and includes a slider for easy adjustment by doctors. The default transparency is 50%, which is beneficial for observing the overlap between the label and anatomical structures; ③ ROI Quantitative Analysis: Clicking the "Generate Report" button automatically calculates the length, width, height, and volume of the lesion area and generates a text report. This helps improve diagnostic efficiency and assists doctors in quantitative tumor analysis.
[0053] Matters not covered in this invention are common knowledge.
[0054] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A pneumonia intelligent analysis method based on multimodal data, characterized in that, The method includes the following steps: Step 1: Connect the hospital information system and electronic medical record system to obtain pneumonia-related data; Step 2: Perform quality cleaning and format standardization on pneumonia-related data; Step 3: Employ a three-level collaborative mechanism of retrieval, generation, and verification, combined with an expert knowledge base and automated tools, to achieve refined data annotation. By retrieving reference annotations and conducting human-machine collaborative verification, annotation efficiency and consistency are improved. Step 4: Based on the dynamic contrastive learning framework, construct a quality-aware alignment algorithm. By evaluating the dimensions of data clarity and resolution, dynamically weight samples of different quality, realize the semantic association of multimodal features, and reduce fusion bias. Step 5: Use the data labeled in Step 3 to perform a VIT-based mask self-supervised pre-training model, and then extract features from the data in Step 4; Step 6: Achieve fusion of image features and text features based on the cross-attention gating mechanism; Step 7: Fine-tune the trained model, and then perform pneumonia identification and analysis on the fine-tuned model to obtain the analysis results.
2. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: Step 1 involves setting up a corresponding RPA data collection script. Through the RPA automated process, basic patient information, medical records, examination and test results, professional knowledge, and Q&A data are collected periodically or in real time. The relevant data collection includes medical information system data, pulmonary professional knowledge system data, and pulmonary disease consultation Q&A data. The medical information system data includes medical information data, laboratory information system focused test data, and medical image data managed by the medical image archiving and communication system. The medical information data integrates data from the hospital's administrative, financial, and medical operations. The laboratory information system focused test data includes test items, results, and reference ranges. The medical image archiving and communication system managed medical image data includes X-ray, CT, and MRI images. This provides information for the entire process of pneumonia diagnosis and treatment, including medical history, diagnosis, treatment plan, examination and test results, and medical orders, constructing a basic record of pneumonia health care services and providing original data for subsequent consultations. The lung professional knowledge system integrates clinical guidelines and expert consensus issued by authoritative organizations such as the WHO and the Chinese Medical Association to ensure that the diagnosis and treatment recommendations are scientific and effective. It also gathers cutting-edge research results from authoritative medical databases such as PubMed and Embase to provide knowledge support for complex diseases. The lung disease consultation Q&A data comes from real hospital consultation records, Q&A on health science popularization platforms and discussion posts on medical forums, reflecting actual needs, helping the public understand knowledge, and providing diverse references for large models.
3. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: Step 2 involves processing image artifacts and text redundancy using heuristic methods to ensure data quality. Considering that the training corpus contains sensitive or personal information that threatens privacy and security, an algorithm based on named entity recognition is used to detect, delete, or replace personal information such as names and addresses. Combined with the Transformer model and machine translation methods, the privacy and security of pneumonia health data are protected. For voice data in the field of pneumonia health, speech recognition technology is used to convert it into text and image data of examination reports. Text information is extracted using OCR technology, and the data is unified into text format for easy integration and processing in the platform's knowledge management module.
4. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: Step 3 involves setting up a brain data annotation and governance framework. The framework uses a three-stage collaborative mechanism of retrieval, generation, and verification to achieve intelligent retrieval and reference annotation of medical cases. It also combines cross-modal association analysis technology to perform deep semantic analysis on several dimensions of imaging features and introduces a human-computer collaborative verification mechanism. This allows clinicians to only need to review and fine-tune the standardized annotation results generated by the system to ensure the clinical accuracy and reliability of the annotation data. The specific process of annotation is as follows: Enhanced image generation is performed using patient CT images and medical records. Similar cases are identified from the medical records database through retrieval and sorting to generate preliminary annotation results. Medical experts conduct final review and correction of the generated results. Cross-modal association annotation involves linking the preliminary annotations of CT images with corresponding diagnostic reports, automatically supplementing other modal information from dynamic MRI and clinical medical records, and comprehensively verifying lung images and medical record information to ensure consistency across multiple modalities. Diffusion synthesis utilizes a generative model to synthesize a diverse number of lung images. The synthesized results undergo quality review by medical experts to ensure their medical rationality and usability. The quality of the generated or processed images in terms of detail and clarity is evaluated, the logical consistency of the annotation results in terms of anatomical structure and pathological features is checked, and the semantic matching and alignment of different modal data is measured.
5. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: In step 4, a quality perception and contrast alignment learning framework for several modalities is set up. The framework constructs a quality evaluation system that covers several dimensions of factors such as sharpness, resolution, annotation quality and modality specificity. A dynamic weighted contrast learning algorithm is developed, which adopts an adaptive weight allocation strategy to achieve a combination of reinforcement learning of samples and deep mining of low-quality samples. The specific process involves inputting computed tomography (CT) images, dynamic magnetic resonance imaging (MRI), and structured or unstructured text. Sample construction and contrastive learning are then employed. The raw data is constructed into pairs of samples for contrastive learning. First, high-quality sample pairs are used, representing effective pairings where text and non-text features are aligned and closely correlated. Then, low-quality samples are compared. By bringing the features of high-quality sample pairs closer together and pushing the features of low-quality sample pairs further away, the quality of each sample pair is automatically evaluated and distinguished. After contrastive learning evaluation, samples are filtered and directed to their corresponding positions based on quality. Those identified as high-quality sample pairs are assigned weights greater than a predetermined limit. The weighted sample information is then aggregated into the loss function module. The loss value between the model's prediction and the actual situation is calculated to guide model optimization. After quality-aware sample screening and weighting, the calculated loss is more accurate and effective, directly driving the model to learn more robust multimodal feature representations.
6. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: In step 5, texture and morphological features of lung tissue are extracted from CT images, and signal intensity and boundary features of lung lesion areas are extracted from MRI images. For medical record text feature extraction, preprocessing is performed first to desensitize patient information, followed by structured extraction to extract key fields from free text. Then, a BERT-based word segmenter is used to process the clinical medical record text. The BERT word segmenter can decompose text into appropriate sub-word units, which can better preserve the semantic information of the text. The processed sub-word sequence is input into the BERT model. The BERT model, through a bidirectional Transformer architecture, captures the contextual relationships between words in the text and generates text feature vectors with semantic information. Each medical record text, after being processed by the BERT model, will obtain a fixed-dimensional feature vector containing key information about the patient's symptoms, diagnosis results, and treatment history.
7. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: In step 6, the cross-attention mechanism enables features from different modalities to pay attention to each other and obtain the correlation between modalities. The gating mechanism can adaptively control the fusion ratio of features from different modalities. During the fusion process, the features of the three modalities are first aligned by the projection layer. Then, the cross-attention gating mechanism is used to allow each modality to dynamically capture the key information of other modalities, resulting in feature-level fusion results. By calculating attention weights, the model adaptively obtains more important image features and text features, thereby performing targeted fusion. The gating unit is introduced to learn the contribution of image features and text features respectively. The fused multimodal feature vector obtained under the synergistic effect of the gating mechanism and the cross-attention mechanism fully integrates image and text information, providing a more comprehensive and accurate feature representation for subsequent diagnosis and prognostic analysis. The weights for the MRI modality are calculated as follows: Wherein, F_MRI represents the feature vector of MRI images, F_CT represents the feature vector of CT images, F_clinical represents the feature vector of clinical medical record text, W_a is a learnable weight matrix, b_a is a bias term that determines which vectors are concatenated into a longer joint feature vector, and finally, decision-level feature fusion is performed. The fused features and intra-modal features will enter the ViT decoder, and the final recognition and analysis will be performed through a gating mechanism and a fully connected layer.
8. The intelligent pneumonia analysis method based on multimodal data according to claim 1, characterized in that: In step 7, for the diagnostic task, the labels are divided into two categories: common interstitial pneumonia and nonspecific interstitial pneumonia. The model is fine-tuned using labeled data, with the cross-entropy loss function as the optimization objective. The cross-entropy loss function can effectively measure the difference between the model's prediction results and the true labels, making it very suitable for classification tasks. For the prognostic task, the labels are optimized into a binary classification task with the prediction of survival within 3 years as the boundary. The cross-entropy loss function is also used as the optimization objective. The model is fine-tuned using the corresponding labeled data. During the fine-tuning process, the model learns the correlation between features and patient prognosis. By continuously optimizing the parameters, the model's ability to predict patient prognosis is improved.