Children MPP auxiliary diagnosis system based on multi-modal time series data modeling
Through multimodal timing data modeling and progressive cross-modal semantic alignment technology, combined with the fine-tuning method of GRPO reinforcement learning, the existing AI is solved inadequate multimodal fusion and low clinical adaptability in children's MPP-assisted diagnosis, and a high accuracy and clinically applicable MPP-assisted diagnosis system is achieved.
Patent Information
- Application Number
- CN202510503307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AI has problems such as insufficient multimodal fusion, limited dynamic course modeling and low clinical adaptability in the diagnosis of MPP in children, which is difficult to meet the needs of accuracy, efficiency and clinical application.
Multimodal timing data modeling is adopted, and multimodal data such as images, electronic medical records and inspection indicators are integrated, and timing feature embedding technology is used to capture the evolution trend of the disease course. Using the progressive cross-modal semantic alignment strategy, the deep fusion of data in different modalities is achieved, and combined with the fine-tuning method dominated by GRPO reinforcement learning, the diagnostic accuracy and robustness of the model are optimized.
It improves the accurate capture and diagnosis accuracy of MPP disease course, enhances the reliability of diagnostic decisions and clinical applicability of models, reduces the work burden of doctors, and promotes the implementation of intelligent diagnosis and treatment of children's MPP.
Smart Images

Figure CN120015296A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical artificial intelligence technology, and specifically relates to a children's MPP auxiliary diagnosis system based on multimodal time series data modeling. Background Art
[0002] In the field of modern medicine, respiratory tract infections in children are a major public health problem worldwide. Mycoplasma pneumonia (MP pneumonia, MPP) caused by Mycoplasma pneumoniae (MP) infection has become the main pathogen of community-acquired pneumonia (CAP) in children. Research data show that among CAP patients aged 5 years and above in China, MPP accounts for as high as 39%, while this proportion is also 30%~40% worldwide. In recent years, clinical findings have shown that the population infected with MPP is showing a trend of younger age, and the incidence of MPP in children under 5 years old has increased significantly. At the same time, the MP resistance rate exceeds 70%, which has led to a gradual increase in the incidence of severe and refractory cases, posing a serious threat to children's health. Severe MPP can lead to serious complications inside and outside the lungs, even endangering life, and may lead to long-term sequelae such as obliterative bronchitis. Therefore, improving the accurate diagnosis and early warning capabilities of MPP is an urgent problem to be solved in the medical community.
[0003] The current clinical diagnosis of MPP faces three major contradictions: First, although pathogen culture is considered the "gold standard" for MP infection, it is complex and time-consuming (usually more than 14 days), which cannot meet the needs of rapid clinical diagnosis; second, existing rapid detection methods, such as immunocolloidal gold detection, are easily affected by factors such as sample quality and operating specifications, and can only detect the presence of MP antibodies, but cannot determine whether it is a current infection; finally, the clinical classification of MPP (mild, severe, and critical) lacks a standardized quantitative evaluation system and relies heavily on the personal experience of doctors. Especially in my country's primary medical system, there is a shortage of pediatrician resources, and it is common for non-professional doctors to make the first visit, which leads to frequent misdiagnosis and missed diagnosis of MPP in children, causing the disease to prolong. Therefore, building an intelligent and precise MPP auxiliary diagnosis system, especially a lightweight intelligent diagnosis and treatment plan suitable for primary medical institutions, has become a key direction for promoting the precise prevention and control of children's respiratory diseases.
[0004] The rapid development of artificial intelligence (AI) technology has provided a new paradigm for precision medical diagnosis and treatment. In the field of intelligent auxiliary diagnosis of pneumonia, researchers at home and abroad have developed a variety of deep learning models using multimodal clinical data such as electronic medical records and medical images to achieve multiple applications such as automatic classification diagnosis of pneumonia, discharge time prediction, and mortality prediction. However, the existing AI-assisted diagnosis models still have three core limitations:
[0005] 1. Insufficient modality collaboration: Most methods rely on a single or a small amount of modality data for auxiliary diagnosis, lack deep semantic fusion across modalities, and fail to fully utilize the collaborative value of multi-source heterogeneous medical data.
[0006] 2. Limitations of static assessment: Current AI models are mainly based on static data analysis and fail to fully incorporate the dynamic evolution data of the disease course, making it difficult to accurately simulate the actual clinical diagnosis process.
[0007] 3. Insufficient clinical adaptability: Most existing methods do not consider the problem of missing data in clinical practice, resulting in limited generalization ability in real application scenarios and difficulty in effectively integrating into the doctor's decision-making process.
[0008] For AI-assisted diagnosis of MPP, early methods mainly focused on single-modality clinical test index data or biochemical index modeling, such as building a classification model through blood parameters and C-reactive protein (CRP) levels to distinguish MPP from non-MPP infections. In recent years, with the accumulation of medical imaging data, the application of deep learning technologies such as convolutional neural networks (CNN) in pneumonia image analysis has gradually increased, such as identifying different types of lung infections such as MPP and bacterial pneumonia through chest X-rays or CT images. However, these methods still have problems such as single data modality and difficulty in dynamically tracking the course of the disease.
[0009] In recent years, multimodal data fusion technology has gradually become a research hotspot for intelligent diagnosis of pneumonia. For example, the joint learning method based on X-ray images and clinical text data can improve the accuracy of pneumonia diagnosis. At the same time, there are also studies that develop more clinically adaptable pneumonia AI diagnosis systems by integrating clinical medical records, laboratory test results and imaging features. However, in terms of severity assessment and severe disease prediction of MPP, relevant AI research is still relatively scarce. Some studies have initially explored MPP severe disease prediction models based on radiomics features, but the generalization ability and clinical practicality of the model still need to be improved.
[0010] With the development of large language model (LLM) technology, its application in the field of medical diagnosis and prediction has gradually increased. For example, the intelligent question-answering system based on the medical big model can assist doctors in making diagnostic decisions, automatically generate imaging reports, and improve medical work efficiency. However, the application of existing large language models in the medical field is still mainly single-modal optimization, and the cross-modal deep reasoning ability is limited, which makes it difficult to meet the clinical needs for multimodal data fusion. At the same time, the adaptability of general medical big models to specific diseases (such as MPP) is still insufficient. How to use domain knowledge for efficient fine-tuning to better serve the precise diagnosis and treatment of specific diseases is still the focus of current research.
[0011] In summary, although AI has made some progress in the field of MPP-assisted diagnosis, there are still problems such as insufficient multimodal fusion, limited dynamic course modeling, and low clinical adaptability. In response to these challenges, the present invention proposes to build a large model of pediatric MPP-specific disease with multimodal time series data modeling and cross-modal semantic alignment as the core, aiming to provide an accurate, efficient and clinically adaptable intelligent auxiliary diagnosis system, and promote the in-depth application of AI technology in the diagnosis and treatment of pediatric respiratory diseases. Summary of the invention
[0012] In view of the deficiencies in the prior art, the present invention provides a pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling to improve the auxiliary diagnosis, clinical classification and severe disease prediction capabilities of Mycoplasma pneumoniae pneumonia and reduce the workload of doctors.
[0013] This invention innovatively proposes an MPP dynamic characterization method based on multimodal time series modeling. By integrating multimodal data such as images, electronic medical records, and test indicators, and using time series feature embedding technology, the evolution trend of the disease can be effectively captured. Furthermore, by using a progressive cross-modal semantic alignment strategy, the deep fusion of data of different modalities is achieved, solving the problem of insufficient diagnostic accuracy caused by incompatible data modalities in traditional methods.
[0014] In order to achieve the above object, the present invention provides the following technical solutions:
[0015] A pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling includes a data processing module, a multimodal feature extraction module, a cross-modal fusion module, a decoder module, a diagnostic reasoning module, an interactive display module and a model optimization module.
[0016] The data processing module is used for data preprocessing, format conversion, and screening of input data.
[0017] The multimodal feature extraction module uses the constructed modality-specific encoder to extract features from data of different modalities and embed timestamp embedding vectors.
[0018] The cross-modal fusion module maps the extracted features to a unified space through a progressive cross-modal alignment strategy, and then achieves deep fusion of images, texts, inspection indicators and temporal information through a modal attention gating mechanism.
[0019] The decoder module uses a language model to decode the multimodal features fused by the cross-modal fusion module to achieve collaborative reasoning of diagnostic logic and data evidence.
[0020] The diagnostic reasoning module uses a large disease-specific model composed of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to analyze the input data, and generates auxiliary diagnostic prompt text associated with time series events in natural language, covering the patient's disease progression, suspected lesion areas and key evidence.
[0021] The interactive display module uses a visual method to display text for auxiliary diagnosis prompts, supporting doctors' interactive adjustments and decisions.
[0022] The model optimization module adopts the GRPO (Group Relative Policy Optimization) reinforcement learning-led fine-tuning method to train and optimize the large disease-specific model composed of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to ensure the diagnostic accuracy of the system.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1. The system of the present invention can accurately capture the evolution of MPP disease course and improve the accuracy of diagnosis through multimodal time series data modeling;
[0025] 2. The present invention adopts a progressive cross-modal semantic alignment strategy to enable collaborative optimization of image, text and test index data, thereby improving the reliability of diagnostic decisions;
[0026] 3. The present invention combines the group relative policy optimization (GRPO) algorithm to optimize the reasoning process and improve the robustness of the model;
[0027] 4. The present invention introduces the thinking chain technology to improve the interpretability of the model, making the auxiliary diagnosis system easier for doctors to accept and apply;
[0028] 5. This invention has been clinically verified in multiple centers to ensure its applicability in different medical institutions and promote the implementation of intelligent diagnosis and treatment of MPP in children. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0030] Figure 1 A schematic diagram of an X-ray image encoder provided by an embodiment of the present invention.
[0031] Figure 2 A schematic diagram of a CT image encoder provided in an embodiment of the present invention.
[0032] Figure 3 A schematic diagram of an electronic medical record encoder provided in an embodiment of the present invention.
[0033] Figure 4 A schematic diagram of an indicator data encoder provided by an embodiment of the present invention.
[0034] Figure 5 A schematic diagram of timestamp embedding provided by an embodiment of the present invention.
[0035] Figure 6 A schematic diagram of progressive cross-modal alignment provided by an embodiment of the present invention.
[0036] Figure 7 A schematic diagram of the two-stage disease-specific large model training provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0038] A pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling includes a data processing module, a multimodal feature extraction module, a cross-modal fusion module, a decoder module, a diagnostic reasoning module, an interactive display module and a model optimization module.
[0039] The data processing module is used for data preprocessing, format conversion, and screening of input data.
[0040] The data include images, electronic medical records, and test index data of patients with Mycoplasma pneumonia (MPP) and non-MPP patients. All acquired data have undergone ethical review, and cases that meet the diagnostic criteria for MPP have been screened to ensure data quality. To protect patient privacy, desensitization has been performed during the data collection stage as a prerequisite for data acquisition.
[0041] For imaging data, X-ray and CT imaging data were collected, and metadata was recorded to ensure data consistency. For electronic medical record data (text mode), redundant characters were removed. For test index data (table mode), laboratory test indicators were collected, including blood routine, C-reactive protein (CRP), procalcitonin (PCT), lactate dehydrogenase (LDH) and D-dimer, and the test time, unit and instrument model were recorded to ensure that the data source was traceable. All data were preprocessed and stored in a format suitable for deep learning models (CT was stored as NIfTI, X-ray was stored as PNG, and electronic medical records and test index data were stored as CSV), providing a high-quality data foundation for subsequent analysis.
[0042] The multimodal feature extraction module uses a modality-specific encoder to extract features from data of different modalities, that is, different encoders are used for different modalities, and combined with time series information embedding technology to capture the dynamic changes of the disease course and improve the prediction accuracy of the model.
[0043] The Mamba-Transformer hybrid architecture is used to process X-ray images, and a multi-scale 3D convolutional network is used to extract CT image features. The electronic medical records are parsed based on the hierarchical Transformer architecture, and the attention mechanism is used to extract the pathological association of the test indicator data.
[0044] For image data, the feature extraction of X-ray images (X-ray image encoder) adopts the Mamba-Transformer hybrid architecture, in which the MambaVision Mixer module recognizes subtle lesions through efficient feature capture and uses the Selective Scan State Space Model (SSM) to obtain key area features, while the Self-Attention module uses the self-attention mechanism to analyze the global information of the image to model long-distance dependencies and improve the detection ability of lesion areas. In the feature extraction process, given an input X-ray image , after preprocessing by the Stem module (two 3×3 convolutions), the initial feature map is obtained :
[0045] ;
[0046] Then, the MambaVision Mixer module is used to enhance the ability to capture local features, and the self-attention mechanism is combined to optimize the global information modeling ability. The processing flow is as follows:
[0047] ;
[0048] ;
[0049] ;
[0050] in is the intermediate feature representation, are the features processed by the MambaVision Mixer module. Indicates that the input dimension is , the output dimension is The linear layer, Indicates a selective scan operation, is the activation function, and Represent one-dimensional convolution and concatenation operations respectively. Next, the features processed by the MambaVision Mixer module are input into the Self-Attention module to analyze the global information of the image:
[0051] ;
[0052] ;
[0053] in, is the final extracted X-ray image feature, Through linear transformation respectively, we get:
[0054] ;
[0055] For feature extraction of CT images (CT image encoder), a multi-scale 3D convolutional network is used to enhance the model's ability to recognize deep lesions by fusing features of different scales. For the input CT image, features are extracted through 3D convolution kernels of different scales:
[0056] ;
[0057] in, They are 3D convolution kernels of different scales. Next, the features obtained at three different scales are concatenated and fused to obtain the final multi-scale feature representation. :
[0058] ;
[0059] For electronic medical record data (text modality), the electronic medical record encoder uses the BioMamba model to process unstructured text. , represents the sequence length, is the word embedding dimension, output features through the BioMamba text encoder :
[0060] ;
[0061] Finally, take the [CLS] token corresponding vector As a global representation.
[0062] For the test index data (tabular mode), a test index data encoder based on a multi-head attention mechanism is designed to capture the high-order correlations between different clinical indicators. , transformed through a fully connected layer, where is the weight matrix, is the bias vector. The embedding representation of the feature for:
[0063] ;
[0064] After obtaining the feature embedding representation Finally, the multi-head attention mechanism is used to capture the complex relationship between the test indicator data. Through linear projection, we obtain the query, key, and value vectors respectively, and perform attention interaction:
[0065] ;
[0066] ;
[0067] Next, the output feature vector is obtained through normalization and feed-forward network:
[0068] ;
[0069] ;
[0070] ;
[0071] in, is layer normalization, is a feed-forward network, and is the weight matrix, .and is the bias vector, is the hidden layer dimension, is the final output feature vector, i.e., the tabular modal feature.
[0072] Since medical data usually has time series characteristics, after obtaining the initial feature representation of each modality, timestamp embedding is added according to the original timestamp information to improve the model's perception of disease progression. In order to preserve the periodic characteristics of timestamps, timestamp embedding is mapped to a high-dimensional space through a sine-cosine function to obtain the corresponding embedding of each modality (i.e. The timestamp embedding vectors of the same dimension are concatenated after the corresponding modal embeddings. This process ensures that the timestamp information can work together with the image features in the same space. The principle of timestamp encoding is as follows:
[0073] ;
[0074] ;
[0075] in is a time series index, is the index of the encoding dimension, is the dimension of the embedding vector.
[0076] The cross-modal fusion module adopts a progressive hierarchical alignment strategy to solve the modality inconsistency problem in the multimodal medical data fusion process and improve the generalization ability of the model. The strategy includes three key stages: bimodal feature fine-tuning, trimodal fusion alignment, and quadmodal joint optimization.
[0077] The first stage is the fine tuning of bimodal features. In the initial stage, fine-grained alignment is performed on X-ray images and electronic medical records to ensure that the image data and the corresponding electronic medical records can form an effective semantic association. The specific implementation method is as follows:
[0078] 1) Construct a shared projection layer to map the X-ray image and electronic medical record features extracted using the modality-specific encoder into a unified feature space:
[0079] ;
[0080] in, They are the learnable projection matrices corresponding to the X-ray image and electronic medical record features, respectively.
[0081] 2) A two-way contrastive learning method is used to optimize feature representation by calculating cross-modal cosine similarity. The positive sample pairs are from the same patient. of and , negative sample pairs are different patients and of and , then the cosine similarity for:
[0082] ;
[0083] 3) Design a bimodal contrast loss function to minimize the alignment error between the image and text modalities and achieve efficient matching. It is expressed as follows:
[0084] ;
[0085] ;
[0086] ;
[0087] in is the temperature coefficient.
[0088] The second stage is the three-modal fusion alignment. On the basis of completing the alignment of X-ray images and electronic medical records, the inspection index data is introduced to build a hierarchical attention mechanism to achieve three-modal collaborative expression. The specific steps are as follows:
[0089] 1) Extracting tabular modal features using a modality-specific encoder Finally, the tabular modal features are also mapped to a unified feature space through a shared projection layer:
[0090] ;
[0091] 2) Design a cross-modal multi-head attention mechanism to dynamically calculate the association weights between inspection indicator data and image / text features to enhance the interactive information between modalities and optimize the representation quality of the tri-modal embedding space. Represents the output of cross attention:
[0092] ;
[0093] 3) Through the trimodal contrast loss function, the distance between the same modality pairs is minimized and the distance between the different modality pairs is maximized. , its contrast loss :
[0094] ;
[0095] Then the trimodal contrast loss function is It can be expressed as:
[0096] ;
[0097] The third stage is the four-modality joint optimization. In order to further improve the accuracy of multi-modality alignment, on the basis of three-modality alignment, the CT modality is added to Mapped to the four-modal shared embedding space through a shared projection layer. The optimization objective at this stage is composed of the four-modal contrast loss and the distribution matching loss, including:
[0098] 1) Constructing a four-modal contrast loss function In order to maintain the alignment between the modes, the optimization goal is:
[0099] ;
[0100] in, is the weight of the alignment loss for each modality, =[1-6].
[0101] 2) To further measure the distribution matching degree of different modalities, the consistency loss based on Wasserstein distance is introduced. Based on the complementarity and correlation of modalities, two groups of modalities, X-ray and CT, and electronic medical records and test index data, are selected to calculate the Wasserstein distance consistency loss. :
[0102] ;
[0103] in, Represents the Wasserstein distance.
[0104] The final optimization goal is:
[0105] ;
[0106] The cross-modal fusion module includes a shared projection layer and a Modality Attention Gate (MAG) mechanism. The Modality Attention Gate mechanism is introduced after the shared projection layer. The dynamic importance score of each modality feature is calculated through a learnable weight matrix. The formula is expressed as:
[0107] ;
[0108] in For the The eigenvectors of the modes, and is a trainable parameter. The final fusion feature is the weighted sum of each modal feature, ensuring that the model can adaptively adjust the modal contribution.
[0109] The decoder module uses a language model to decode the multimodal features fused by the cross-modal fusion module, and the fused multimodal features are input to the language model decoder of the language model part through a cross-attention mechanism. During the decoding process, the language model generates auxiliary diagnosis prompt text based on the input instructions and multimodal features through the decoder autoregression.
[0110] The diagnostic reasoning module uses a large disease-specific model composed of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to analyze the input data, and generates auxiliary diagnostic prompt text associated with time series events in natural language, covering the patient's disease progression, suspected lesion areas and key evidence.
[0111] The interactive display module uses a visual method to display text for auxiliary diagnosis prompts, supporting doctors' interactive adjustments and decisions.
[0112] The model training of the model optimization module adopts the fine-tuning method dominated by GRPO (Group Relative Policy Optimization) reinforcement learning. This method is divided into two stages: supervised fine-tuning and reinforced fine-tuning to optimize the adaptability of the large model in auxiliary diagnosis of Mycoplasma pneumonia.
[0113] The first stage is supervised fine-tuning. In this stage, LoRA (Low-Rank Adaptation) technology is used to adjust the weight of the large model in a lightweight way by adding a LoRA adaptation layer to the projection layer of the language model decoder to improve the adaptability of the model to medical field knowledge. The LoRA technical formula is as follows:
[0114] ;
[0115] in, is the original weight of the large model, The weight of the large model after fine-tuning of LoRA, and are trainable weights.
[0116] A multi-task loss function is used to update the weights of the LoRA adaptation layer, and a triple collaborative optimization system including diagnosis classification loss, thought chain reasoning loss and clinical rule loss is constructed. Adopt standard cross entropy function to optimize disease classification accuracy; Mindchain sequence prediction loss Dynamic focus weights are used to enhance model reasoning capabilities; clinical rule loss The medical diagnosis and treatment standards are transformed into differentiable constraints to ensure the compliance of the diagnosis conclusion. The formula is as follows:
[0117] ;
[0118] in is the total number of rules, is the KL divergence, For the Clinical rules for samples The expected distribution of For the model to sample The predicted distribution of categories.
[0119] Finally, we get the joint loss function :
[0120] ;
[0121] The second stage is enhanced fine-tuning, that is, based on supervised fine-tuning, the Chain-of-Thought strategy is used to build a GRPO optimization framework to improve the model's reasoning ability.
[0122] First, multiple reasoning paths are generated through Monte Carlo sampling, and rewards are calculated for each path, including clinical accuracy rewards (calculated by medical experts' annotations) and logical consistency rewards (evaluated using propositional logic verifiers). Then, a relative advantage function update strategy is adopted to ensure the stability of gradient optimization, and the KL divergence constraint strategy is combined to update the amplitude. Finally, reinforcement learning is used to optimize the reasoning path, so that the model prioritizes the orthogonal reasoning path of "image feature analysis → laboratory indicator association → differential diagnosis exclusion", while suppressing irrational reasoning patterns such as "direct conclusion jump".
[0123] The present invention also provides a method for constructing a large model for Mycoplasma pneumonia. This method integrates multimodal data such as images, electronic medical records, and test indicators, and uses time series feature embedding technology to accurately capture the evolution trend of the disease course. At the same time, it uses a progressive cross-modal semantic alignment strategy to achieve deep fusion of data of different modalities, solving the problem of insufficient diagnostic accuracy caused by incompatible data modalities in traditional methods. The specific implementation methods are as follows:
[0124] Step 1: Data collection and organization, building a multimodal time series data set covering image sequences, electronic medical records and test indicators.
[0125] The data sources of this embodiment include multiple medical institutions, including images, electronic medical records, and test index data of Mycoplasma pneumonia (MPP) patients and non-MPP patients. All acquired data have undergone ethical review, and cases that meet the MPP diagnostic criteria are manually screened to ensure data quality. In order to protect patient privacy, desensitization processing has been performed during the data collection stage as a prerequisite for data acquisition.
[0126] For imaging data, X-ray and CT imaging data are collected, and metadata such as image acquisition time and equipment parameters are recorded to ensure data consistency. For electronic medical record data (text mode), text data such as patient medical history, doctor's notes, and examination reports are collected, and redundant characters are removed. For test index data (table mode), laboratory test indicators are collected, including blood routine, C-reactive protein (CRP), procalcitonin (PCT), lactate dehydrogenase (LDH), and D-dimer, and the test time, unit, and instrument model are recorded to ensure that the data source is traceable. All data are preprocessed and stored in a format suitable for deep learning models (CT is stored as NIfTI, X-ray is stored as PNG, and electronic medical records and test index data are stored as CSV), providing a high-quality data foundation for subsequent analysis.
[0127] Step 2: Multimodal feature extraction module with temporal migration.
[0128] The present invention adopts a modality-specific encoder to extract features from data of different modalities, that is, different encoders are used for different modalities, and combined with time series information embedding technology, to capture the dynamic changes of the disease course and improve the prediction accuracy of the model.
[0129] The Mamba-Transformer hybrid architecture is used to process X-ray images, and a multi-scale 3D convolutional network is used to extract CT image features. The hierarchical Transformer architecture is used to parse electronic medical records, and the attention mechanism is used to extract pathological associations of test index data.
[0130] For image data, the feature extraction of X-ray images (X-ray image encoder) adopts the Mamba-Transformer hybrid architecture, in which the MambaVision Mixer module recognizes subtle lesions through efficient feature capture and uses the Selective Scan State Space Model (SSM) to obtain key area features, while the Self-Attention module uses the self-attention mechanism to analyze the global information of the image to model long-distance dependencies and improve the detection ability of lesion areas. In the feature extraction process, given an input X-ray image , after preprocessing by the Stem module (two 3×3 convolutions), the initial feature map is obtained :
[0131] ;
[0132] Then, the MambaVision Mixer module is used to enhance the ability to capture local features, and the self-attention mechanism is combined to optimize the global information modeling ability. The processing flow is as follows:
[0133] ;
[0134] ;
[0135] ;
[0136] in is the intermediate feature representation, are the features processed by the MambaVision Mixer module. Indicates that the input dimension is , the output dimension is The linear layer, Indicates a selective scan operation, is the activation function, and Represent one-dimensional convolution and concatenation operations respectively. Next, the features processed by the MambaVision Mixer module are input into the Self-Attention module to analyze the global information of the image:
[0137] ;
[0138] ;
[0139] in, is the final extracted X-ray image feature, Through linear transformation respectively, we get:
[0140] ;
[0141] The feature extraction of this Mamba-Transformer hybrid architecture can take into account both local and global information, which helps the model improve its fine-grained image parsing capabilities.
[0142] For feature extraction of CT images (CT image encoder), a multi-scale 3D convolutional network is used to enhance the model's ability to recognize deep lesions by fusing features of different scales. For the input CT image, features are extracted through 3D convolution kernels of different scales:
[0143] ;
[0144] in, They are 3D convolution kernels of different scales. Next, the features obtained at three different scales are concatenated and fused to obtain the final multi-scale feature representation. :
[0145] ;
[0146] This multi-scale feature extraction can focus on information at multiple resolutions, which is helpful to improve the model's ability to identify key lesion areas.
[0147] For electronic medical record data (text modality), the electronic medical record encoder uses the BioMamba model to process unstructured text. , represents the sequence length, is the word embedding dimension, output features through the BioMamba text encoder :
[0148] ;
[0149] Finally, take the [CLS] token corresponding vector As a global representation, the model combines SSM and deep semantic encoding technology to model long-range dependencies and extract key medical information in medical record text.
[0150] For the test index data (table modality), this paper designs a test index data encoder based on a multi-head attention mechanism to capture the high-order correlations between different clinical indicators. , transformed through a fully connected layer, where is the weight matrix, is the bias vector. The embedding representation of the feature for:
[0151] ;
[0152] After obtaining the feature embedding representation Finally, the multi-head attention mechanism is used to capture the complex relationship between the test indicator data. Through linear projection, we obtain the query, key, and value vectors respectively, and perform attention interaction:
[0153] ;
[0154] ;
[0155] Next, the output feature vector is obtained through normalization and feed-forward network:
[0156] ;
[0157] ;
[0158] ;
[0159] in, is layer normalization, is a feed-forward network, and is the weight matrix, .and is the bias vector, is the hidden layer dimension, is the final output feature vector, i.e., the tabular modal feature.
[0160] Since medical data usually has time series characteristics, after obtaining the initial feature representation of each modality, timestamp embedding is added according to the original timestamp information to improve the model's perception of disease progression. In order to preserve the periodic characteristics of timestamps, timestamp embedding is mapped to a high-dimensional space through a sine-cosine function to obtain the corresponding embedding of each modality (i.e. ) timestamp embedding vectors of the same dimension are concatenated after the corresponding embedding of each modality. This process ensures that the timestamp information can work together with the image features in the same space. The principle of timestamp encoding is as follows:
[0161] ;
[0162] ;
[0163] in is a time series index, is the index of the encoding dimension, is the dimension of the embedding vector. In this way, the temporal encoding can show periodicity in different dimensions and effectively distinguish elements at different positions.
[0164] Step 3: Progressive cross-modal alignment.
[0165] A progressive hierarchical alignment strategy is adopted to solve the modality inconsistency problem in the multimodal medical data fusion process and improve the generalization ability of the model. The strategy includes three key stages: bimodal feature fine-tuning, trimodal fusion alignment, and quadmodal joint optimization.
[0166] The first stage is the fine tuning of bimodal features. In the initial stage, fine-grained alignment is performed on X-ray images and electronic medical records to ensure that the image data and the corresponding electronic medical records can form an effective semantic association. The specific implementation method is as follows:
[0167] 1) Construct a shared projection layer to map the X-ray image and electronic medical record features extracted using the modality-specific encoder into a unified feature space:
[0168] ;
[0169] in, They are the learnable projection matrices corresponding to the X-ray image and electronic medical record features, respectively.
[0170] 2) A two-way contrastive learning method is used to optimize feature representation by calculating cross-modal cosine similarity. The positive sample pairs are from the same patient. of and , negative sample pairs are different patients and of and , then the cosine similarity for:
[0171] ;
[0172] 3) Design a bimodal contrast loss function to minimize the alignment error between the image and text modalities and achieve efficient matching. It is expressed as follows:
[0173] ;
[0174] ;
[0175] ;
[0176] in is the temperature coefficient.
[0177] The second stage is the three-modal fusion alignment. On the basis of completing the alignment of X-ray images and electronic medical records, the inspection index data is introduced to build a hierarchical attention mechanism to achieve three-modal collaborative expression. The specific steps are as follows:
[0178] 1) Extracting tabular modal features using a modality-specific encoder Finally, the tabular modal features are also mapped to a unified feature space through a shared projection layer:
[0179] ;
[0180] 2) Design a cross-modal multi-head attention mechanism to dynamically calculate the association weights between inspection indicator data and image / text features to enhance the interactive information between modalities and optimize the representation quality of the tri-modal embedding space. Represents the output of cross attention:
[0181] ;
[0182] 3) Through the trimodal contrast loss function, the distance between the same modality pairs is minimized and the distance between the different modality pairs is maximized. , its contrast loss :
[0183] ;
[0184] Then the trimodal contrast loss function is It can be expressed as:
[0185] ;
[0186] The third stage is the four-modality joint optimization. In order to further improve the accuracy of multi-modality alignment, on the basis of three-modality alignment, the CT modality is added to Mapped to the four-modal shared embedding space through a shared projection layer. The optimization objective at this stage is composed of the four-modal contrast loss and the distribution matching loss, including:
[0187] 1) Constructing a four-modal contrast loss function In order to maintain the alignment between the modes, the optimization goal is:
[0188] ;
[0189] in, is the weight of the alignment loss for each modality.
[0190] 2) To further measure the distribution matching degree of different modalities, the consistency loss based on Wasserstein distance is introduced. Based on the complementarity and correlation of modalities, two groups of modalities, X-ray and CT, and electronic medical records and test index data, are selected to calculate the Wasserstein distance consistency loss. :
[0191] ;
[0192] in, Represents the Wasserstein distance.
[0193] The final optimization goal is:
[0194] ;
[0195] Through the above-mentioned progressive cross-modal alignment strategy, the present invention effectively alleviates the modality inconsistency problem in medical multimodal data fusion, improves the collaborative expression ability of images, texts, test indicators and time series information, and ultimately improves the accuracy and robustness of the diagnosis of Mycoplasma pneumonia in children.
[0196] Step 4: Two-stage disease-specific large model training.
[0197] A multimodal large model refers to a large-scale artificial intelligence model that can process multiple types of data (such as text, images, audio, etc.). It integrates information from different modalities and mines cross-modal associations and semantics to achieve more comprehensive understanding and generation capabilities.
[0198] In this step, we construct a large MPP disease-specific model, which includes a modality-specific encoder, a cross-modal fusion module, and a language model. The cross-modal fusion module includes a shared projection layer and a modality attention gate (MAG) mechanism. The modality attention gate mechanism is introduced after the shared projection layer, and the dynamic importance score of each modality feature is calculated through a learnable weight matrix. The formula is expressed as:
[0199] ;
[0200] in For the The eigenvectors of the modes, and is a trainable parameter. The final fusion feature is the weighted sum of each modal feature, ensuring that the model can adaptively adjust the modal contribution. The fused multimodal features are input to the language model decoder of the language model through the cross-attention mechanism. During the decoding process, the language model generates auxiliary diagnosis prompt text based on the input instructions and multimodal features through the decoder autoregression.
[0201] The model training adopts the fine-tuning method dominated by GRPO (Group Relative Policy Optimization) reinforcement learning. This method is divided into two stages: supervised fine-tuning and reinforced fine-tuning, to optimize the adaptability of the large model in auxiliary diagnosis of Mycoplasma pneumonia.
[0202] The first stage is supervised fine-tuning. In this stage, LoRA (Low-Rank Adaptation) technology is used to adjust the weight of the large model in a lightweight way by adding a LoRA adaptation layer to the projection layer of the language model decoder to improve the adaptability of the model to medical field knowledge. The LoRA technical formula is as follows:
[0203] ;
[0204] in, is the original weight of the large model, The weight of the large model after fine-tuning of LoRA, and are trainable weights.
[0205] Traditional large model fine-tuning uses a single cross entropy loss function, which has low data efficiency. In order to efficiently utilize limited data, this paper adopts a multi-task loss function to update the weights of the LoRA adaptation layer and construct a triple collaborative optimization system including diagnostic classification loss, thought chain reasoning loss and clinical rule loss. Adopt standard cross entropy function to optimize disease classification accuracy; Mindchain sequence prediction loss Dynamic focus weights are used to enhance model reasoning capabilities; clinical rule loss The medical diagnosis and treatment standards are transformed into differentiable constraints to ensure the compliance of the diagnosis conclusion. The formula is as follows:
[0206] ;
[0207] in is the total number of rules, is the KL divergence, For the Clinical rules for samples The expected distribution of For the model to sample The predicted distribution of categories.
[0208] Finally, we get the joint loss function :
[0209] ;
[0210] The second stage is enhanced fine-tuning, that is, based on supervised fine-tuning, the Chain-of-Thought strategy is used to build a GRPO optimization framework to improve the model's reasoning ability.
[0211] First, multiple reasoning paths are generated through Monte Carlo sampling, and rewards are calculated for each path, including clinical accuracy rewards (calculated by medical experts' annotations) and logical consistency rewards (evaluated using propositional logic verifiers). Then, a relative advantage function update strategy is adopted to ensure the stability of gradient optimization, and the KL divergence constraint strategy is combined to update the amplitude. Finally, reinforcement learning is used to optimize the reasoning path, so that the model prioritizes the orthogonal reasoning path of "image feature analysis → laboratory indicator association → differential diagnosis exclusion", while suppressing irrational reasoning patterns such as "direct conclusion jump".
[0212] Through the above-mentioned GRPO reinforcement learning-led staged fine-tuning strategy, the present invention achieves efficient adaptation of large models in the medical field and improves the accuracy and clinical applicability of the diagnosis of Mycoplasma pneumonia.
[0213] Step 5: System construction and verification.
[0214] After completing the model training, the present invention further conducts system construction to ensure that the model can be efficiently applied in clinical diagnosis scenarios and conducts comprehensive performance verification.
[0215] The present invention constructs an intelligent diagnosis system of a large model of Mycoplasma pneumonia-specific disease based on a microservice architecture to ensure the scalability and efficiency of the system. The system mainly includes a data processing module, a multimodal feature extraction module, a cross-modal fusion module, a decoder module, a diagnostic reasoning module, an interactive display module and a model optimization module. The data processing module is responsible for data preprocessing, format conversion, and screening of input data. The multimodal feature extraction module uses the constructed modality-specific encoder to extract features from data of different modalities and embed timestamp embedding vectors. The cross-modal fusion module maps the extracted features to a unified space through a progressive cross-modal alignment strategy, and then realizes the deep fusion of images, texts, test indicators and time series information by the modal attention gating mechanism. The decoder module uses a language model to decode the multimodal features fused by the cross-modal fusion module to achieve collaborative reasoning of diagnostic logic and data evidence. The diagnostic reasoning module uses a large disease-specific model consisting of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to analyze the input data, and generates auxiliary diagnostic prompt text associated with time series events in natural language, covering the patient's disease progression, suspected lesion areas, and key evidence for reference by professionals. The interactive display module uses a visual method to display text for auxiliary diagnostic prompts, supporting doctors to make interactive adjustments and decisions. The model optimization module uses a fine-tuning method led by GRPO (Group Relative Policy Optimization) reinforcement learning to optimize the training of a large disease-specific model consisting of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to ensure the accuracy of the system's auxiliary diagnosis.
[0216] In order to evaluate the accuracy, robustness and interpretability of the model of the present invention, a multi-dimensional testing method was adopted, including retrospective validation, prospective validation and clinical application testing. Retrospective validation selected historical MPP patient data from multiple medical institutions to evaluate the classification performance of the model, and used indicators such as accuracy, recall, precision, F1-score and AUC (Area Under Curve) to measure the classification effect. In the prospective validation, the model was predicted for MPP patient data in an actual clinical environment, and compared with the doctor's diagnosis results, and the consistency index (Cohen's Kappa) was statistically analyzed to analyze the consistency between the model prediction results and the expert diagnosis. The clinical application test used real-world data (RWD) for systematic evaluation, and invited pneumonia diagnosis experts to conduct blind tests to record the auxiliary value of the model in clinical decision-making, including diagnosis time, misdiagnosis rate and key lesion identification.
[0217] Through the construction and verification of the above system, the present invention has realized a high-precision, multi-modal fusion Mycoplasma pneumonia disease model. In the future, the system will be further optimized and expanded to other types of pneumonia, such as bacterial pneumonia and viral pneumonia, to improve the versatility of the model; integrate multi-source medical knowledge bases to enhance the interpretability of the model; optimize the user interaction interface to improve the doctor's experience and diagnostic efficiency.
[0218] The above description is only by way of illustration of certain exemplary embodiments of the present invention. It is undoubted that those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling, characterized in that: It includes data processing module, multimodal feature extraction module, cross-modal fusion module, decoder module, diagnostic reasoning module, interactive display module and model optimization module; The data processing module is used for data preprocessing, format conversion, and screening of input data; The multimodal feature extraction module uses the constructed modality-specific encoder to extract features from data of different modalities and embed timestamp embedding vectors; The cross-modal fusion module maps the extracted features to a unified space through a progressive cross-modal alignment strategy, and then achieves deep fusion of images, texts, inspection indicators and time series information through a modal attention gating mechanism; The decoder module uses a language model to decode the multimodal features fused by the cross-modal fusion module to achieve collaborative reasoning of diagnostic logic and data evidence; The diagnostic reasoning module uses a large disease-specific model composed of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to analyze the input data, and generates auxiliary diagnostic prompt text associated with time series events in natural language, covering the patient's disease progression, suspected lesion areas, and key evidence; The interactive display module uses a visual method to display text used for auxiliary diagnosis prompts, supporting doctors' interactive adjustments and decisions; The model optimization module adopts the GRPO reinforcement learning-led fine-tuning method to train and optimize the large disease-specific model composed of a multimodal feature extraction module, a cross-modal fusion module, and a decoder module to ensure the diagnostic accuracy of the system.
2. According to claim 1, a pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling is characterized in that: The data processing module is used for data preprocessing, format conversion, and screening of input data; The data included images, electronic medical records, and test index data of MPP patients and non-MPP patients with Mycoplasma pneumoniae pneumonia. All acquired data were subject to ethical review, and cases that met the MPP diagnostic criteria were screened to ensure data quality. To protect patient privacy, desensitization was performed during the data collection stage as a prerequisite for data acquisition. For imaging data, X-ray and CT imaging data are collected, and metadata is recorded to ensure data consistency; for electronic medical record data, redundant characters are removed; for test index data, laboratory test indicators are collected, including blood routine, C-reactive protein, procalcitonin, lactate dehydrogenase and D-dimer, and the test time, unit and instrument model are recorded to ensure that the data source is traceable; all data are preprocessed and stored in a format suitable for deep learning models.
3. The pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling according to claim 2 is characterized in that: The multimodal feature extraction module uses a modality-specific encoder to extract features from different modal data, that is, different encoders are used for different modalities, and combined with time series information embedding technology to capture the dynamic changes of the disease course and improve the prediction accuracy of the model; The Mamba-Transformer hybrid architecture is used to process X-ray images, and a multi-scale 3D convolutional network is used to extract CT image features. The hierarchical Transformer architecture is used to parse electronic medical records, and the attention mechanism is used to extract pathological associations of test index data. For image data, the feature extraction of X-ray images adopts the Mamba-Transformer hybrid architecture, in which the MambaVision Mixer module recognizes subtle lesions through efficient feature capture and uses the selective scanning state space model to obtain key area features, while the Self-Attention module uses the self-attention mechanism to analyze the global information of the image to model long-distance dependencies and improve the detection ability of lesion areas. In the feature extraction process, given an input X-ray image , after preprocessing by the Stem module, the initial feature map is obtained : ; Then, the MambaVision Mixer module is used to enhance the ability to capture local features, and the self-attention mechanism is combined to optimize the global information modeling ability; The processing flow is as follows: ; ; ; in is the intermediate feature representation, are the features processed by the MambaVision Mixer module. Indicates that the input dimension is , the output dimension is The linear layer, Indicates a selective scan operation, is the activation function, and They represent one-dimensional convolution and concatenation operations respectively. Then, the features processed by the MambaVision Mixer module are input into the Self-Attention module to analyze the global information of the image: ; ; in, is the final extracted X-ray image feature. Through linear transformation respectively, we get: ; For feature extraction of CT images, a multi-scale 3D convolutional network is used to enhance the model's ability to recognize deep lesions by fusing features of different scales. For the input CT image, features are extracted through 3D convolution kernels of different scales: ; in, They are 3D convolution kernels of different scales respectively; next, the features obtained at three different scales are concatenated and fused to obtain the final multi-scale feature representation : ; For electronic medical record data, the BioMamba model is used to process unstructured text; for input text sequences , represents the sequence length, is the word embedding dimension, output features through the BioMamba text encoder : ; Finally, take the [CLS] token corresponding vector As a global representation; For the test index data, a test index data encoder based on a multi-head attention mechanism is designed to capture the high-order correlations between different clinical indicators; for the input tabular modal features , transformed through a fully connected layer, where is the weight matrix, is the bias vector; the embedding representation of the feature for: ; After obtaining the feature embedding representation Finally, the multi-head attention mechanism is used to capture the complex relationship between the test indicator data; the feature embedding representation Through linear projection, we obtain the query, key, and value vectors respectively, and perform attention interaction: ; ; Next, the output feature vector is obtained through normalization and feed-forward network: ; ; ; in, is layer normalization, is a feed-forward network, and is the weight matrix, ;and is the bias vector, is the hidden layer dimension, is the final output feature vector, i.e., the tabular modal feature; Since medical data usually has time series characteristics, after obtaining the initial feature representation of each modality, timestamp embedding is added according to the original timestamp information to improve the model's perception of disease progression; in order to retain the periodic characteristics of timestamps, timestamp embedding is mapped to a high-dimensional space through a sine-cosine function to obtain a timestamp embedding vector with the same dimension as the corresponding modality embedding, which is then concatenated with the corresponding modality embedding; the principle of timestamp encoding is as follows: ; ; in is a time series index, is the index of the encoding dimension, is the dimension of the embedding vector.
4. The pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling according to claim 3 is characterized in that: The cross-modal fusion module adopts a progressive hierarchical alignment strategy to solve the modality inconsistency problem in the multimodal medical data fusion process and improve the generalization ability of the model; the strategy includes three key stages, namely, bimodal feature fine-tuning, trimodal fusion alignment and quadmodal joint optimization; The first stage is the fine tuning of bimodal features; In the initial stage, fine-grained alignment is performed on X-ray images and electronic medical records to ensure that the image data and the corresponding electronic medical records can form an effective semantic association; the specific implementation method is as follows: 1) Construct a shared projection layer to map the X-ray image and electronic medical record features extracted using the modality-specific encoder into a unified feature space: ; in, are the learnable projection matrices corresponding to X-ray image and electronic medical record features, respectively; 2) A two-way contrastive learning method is used to optimize feature representation by calculating cross-modal cosine similarity; the positive sample pairs are from the same patient of and , negative sample pairs are different patients and of and , then the cosine similarity for: ; 3) Design a bimodal contrast loss function to minimize the alignment error between image and text modalities and achieve efficient matching; Bimodal contrast loss function It is expressed as follows: ; ; ; in is the temperature coefficient; The second stage is the three-modal fusion alignment. On the basis of completing the alignment of X-ray images and electronic medical records, the test index data is introduced to build a hierarchical attention mechanism to achieve three-modal collaborative expression. The specific steps are as follows: 1) Extracting tabular modal features using a modality-specific encoder Finally, the tabular modal features are also mapped to a unified feature space through a shared projection layer: ; 2) Design a cross-modal multi-head attention mechanism to dynamically calculate the association weights between the inspection index data and the image / text features to enhance the interactive information between the modalities and optimize the representation quality of the tri-modal embedding space; Represents the output of cross attention: ; 3) Through the trimodal contrast loss function, the distance between the same modality pairs is minimized and the distance between the different modality pairs is maximized; for the modality , its contrast loss : ; Then the trimodal contrast loss function is It can be expressed as: ; The third stage is the four-modality joint optimization. To further improve the accuracy of multi-modality alignment, the CT modality is added on the basis of the three-modality alignment. Mapped to the four-modal shared embedding space through a shared projection layer; the optimization objective at this stage is composed of the four-modal contrast loss and the distribution matching loss, including: 1) Constructing a four-modal contrast loss function In order to maintain the alignment relationship between modes, the optimization goal is: ; in, is the weight of the alignment loss for each modality, =[1-6]; 2) To further measure the distribution matching degree of different modalities, the consistency loss based on Wasserstein distance is introduced; based on the complementarity and correlation of modalities, two groups of modalities, X-ray and CT, and electronic medical records and test index data, are selected to calculate the Wasserstein distance consistency loss : ; in, represents the Wasserstein distance; The final optimization goal is: ; The cross-modal fusion module includes a shared projection layer and a modal attention gating mechanism. The modal attention gating mechanism is introduced after the shared projection layer. The dynamic importance score of each modal feature is calculated through a learnable weight matrix. The formula is expressed as: ; in For the The eigenvectors of the modes, and is a trainable parameter; the final fusion feature is the weighted sum of the features of each modality, ensuring that the model adaptively adjusts the modal contribution.
5. The pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling according to claim 4 is characterized in that: The decoder module uses a language model to decode the multimodal features fused by the cross-modal fusion module, and the fused multimodal features are input into the language model decoder of the language model part through a cross-attention mechanism; during the decoding process, the language model generates auxiliary diagnosis prompt text based on the input instructions and multimodal features through decoder autoregression.
6. The pediatric MPP auxiliary diagnosis system based on multimodal time series data modeling according to claim 5 is characterized in that: The model training of the model optimization module adopts the fine-tuning method dominated by GRPO reinforcement learning; this method is divided into two stages: supervised fine-tuning and reinforcement fine-tuning, so as to optimize the adaptability of the large model in auxiliary diagnosis of Mycoplasma pneumoniae pneumonia; The first stage is supervised fine-tuning; In this stage, LoRA technology is used to add a LoRA adaptation layer to the projection layer of the language model decoder to adjust the weight of the large model in a lightweight way to improve the adaptability of the model to medical field knowledge; the LoRA technology formula is as follows: ; in, is the original weight of the large model, The weight of the large model after fine-tuning of LoRA, and are trainable weights; A multi-task loss function is used to update the weights of the LoRA adaptation layer, and a triple collaborative optimization system including diagnosis classification loss, thinking chain reasoning loss and clinical rule loss is constructed; diagnosis classification loss Using standard cross entropy function; Thinking Chain sequence prediction loss Dynamic focus weighting is adopted; clinical rule loss Convert medical diagnosis and treatment norms into differentiable constraints: ; in is the total number of rules, is the KL divergence, For the Clinical rules for samples The expected distribution of For the model to sample The predicted distribution of categories; Finally, we get the joint loss function : ; The second stage is enhanced fine-tuning, that is, based on supervised fine-tuning, the thinking chain strategy is used to build a GRPO optimization framework to improve the model's reasoning ability; Firstly, multiple reasoning paths were generated through Monte Carlo sampling, and rewards were calculated for each path, including clinical accuracy rewards and logical consistency rewards. Then, a relative advantage function update strategy was adopted to ensure the stability of gradient optimization, and the KL divergence constraint strategy was combined to update the amplitude. Finally, reinforcement learning was used to optimize the reasoning path, so that the model gave priority to the orthogonal reasoning path of "image feature analysis→laboratory indicator association→differential diagnosis exclusion", while suppressing irrational reasoning patterns.
Citation Information
Cited By
Complex equipment fault diagnosis method and system based on multi-modal knowledge graph
CN120217264A
A Complex Equipment Fault Diagnosis Method and System Based on a Multimodal Knowledge Graph
CN120217264B
Power battery disassembly path optimization method based on autoregression generation strategy
CN120218382A
Child multi-modal data output method and system
CN120296704A
Industrial image anomaly detection method based on noise suppression modal fusion alignment
CN120298399A