Method and system for fusing few types of multi-modal medical data, electronic equipment and readable storage medium
Through the unified projection layer and multi-task Prefix Tokens strategy, the modal differences and multi-task collaborative learning problems of multimodal medical data are solved, and efficient multimodal data fusion and robustness enhancement are achieved, which is suitable for stable performance improvement in multi-task scenarios.
Patent Information
- Application Number
- CN202510763761.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies cannot effectively solve the problems of modal differences in multimodal medical data, multi-task collaborative learning, and robustness in high-noise scenarios, resulting in poor fusion effects and unstable model performance.
A unified projection layer, a multi-task Prefix Tokens strategy and a robustness enhancement mechanism are adopted. The unified projection layer is used to map different modal data into a shared latent feature space. The multi-head self-attention mechanism is combined for feature fusion. Prefix Tokens are used to achieve multi-task differentiation and adaptive learning, thereby improving the robustness of the model and its multi-task collaborative learning capabilities.
It achieves efficient feature alignment and fusion of multimodal data, supports multi-task collaborative learning, improves the stability and generalization ability of the model in small sample and high noise scenarios, and reduces system complexity.
Smart Images

Figure CN120636846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, system, electronic device, and readable storage medium for fusing a small number of types, multiple types, and multi-modal medical data, and belongs to the technical field of medical data processing. Background Art
[0002] With the continuous advancement of medical technology and data collection methods, the types and scale of data generated in clinical medicine have gradually expanded, encompassing multimodal data such as text, time series, and images. These data collectively reflect the comprehensive characteristics of patients, from physiological to pathological. Fusion and analysis of these data can provide more comprehensive information for accurate disease diagnosis, treatment optimization, and health status prediction. However, existing technical approaches still cannot fully tap the potential of multimodal data, especially in multimodal alignment and fusion, and multi-task collaborative learning, which still face significant challenges.
[0003] First, clinical multimodal data has significant modal differences. For example, text data (such as medical records) are mainly based on discrete symbols and semantic information, emphasizing the ability to understand natural language; time series data (such as electrocardiograms and blood glucose monitoring curves) reflect dynamic processes and focus on capturing temporal patterns; and medical images (such as CT, MRI, and X-rays) are centered on spatial features, showing anatomical and pathological information. This difference in characteristics between modalities makes semantic alignment and feature fusion of multimodal data extremely difficult. Traditional feature splicing or simple weighting methods often lack the ability to deeply understand modal associations, resulting in poor fusion effects.
[0004] Secondly, the tasks in medical scenarios are diverse and complex, including text generation (such as automatically generating medical records), time series prediction (such as blood pressure trend prediction), classification tasks (such as disease diagnosis), and anomaly detection (such as critical event warning). These tasks share commonalities but also exhibit significant differences. Supporting multi-task collaborative learning within a unified model framework while avoiding conflicts between tasks remains a significant research challenge. Existing methods are typically optimized for a single task and are unable to achieve a balanced performance in multi-task scenarios.
[0005] Furthermore, clinical data is limited by factors such as collection costs and privacy protection, and often suffers from problems such as small sample sizes, high noise, and incomplete data. For example, electrocardiograms can produce abnormal signals due to equipment interference, medical images can be affected by artifacts or lighting changes, and text records can contain inconsistent or ambiguous descriptions. These issues place higher demands on model robustness. However, existing models often struggle to maintain performance in small sample sizes and high noise scenarios, especially when modal data is missing or of poor quality.
[0006] In this context, there is an urgent need for a novel approach that can achieve deep fusion of data from different modalities, support multi-task learning, and demonstrate excellent robustness and generalization capabilities in complex clinical environments. This invention aims to fill this technological gap by introducing a large-scale model for the fusion of multimodal data from a small number of categories, providing an integrated solution for the efficient processing and application of multimodal medical data.
[0007] Existing research on the analysis and fusion of multimodal data focuses on three aspects: modal feature extraction, multimodal alignment, and multi-task learning. However, these methods still have many shortcomings in their performance in clinical data scenarios.
[0008] In terms of modal feature extraction, mainstream technologies are usually optimized for a single modality. Large language models (such as BERT and GPT series) perform well in natural language processing tasks, but they are mainly trained on general-domain texts and fail to optimize for the professional and fine-grained semantic requirements of medical texts. Time series analysis methods (such as LSTM and GRU) are good at processing single-variable or low-dimensional sequence data, but have limited performance in high-dimensional and long-term dependent time series analysis. In addition, convolutional neural networks (CNNs) are often used in medical image analysis. Although they have significant advantages in extracting local features, they lack the ability to model complex anatomical structures and multi-scale features.
[0009] When it comes to multimodal alignment and fusion, traditional methods typically use independent encoders to extract features from data of different modalities, followed by fusion through simple feature concatenation. However, this approach fails to address the heterogeneity between modalities, resulting in a lack of semantic consistency in the fused features. For example, there are significant differences in feature dimensions between time series and text data, while images and text lack a natural correlation. Furthermore, existing alignment methods are highly sensitive to noise and anomalous data, resulting in unstable performance in high-noise scenarios.
[0010] Despite significant progress in multimodal data fusion and multi-task learning technologies in recent years, existing methods still face significant shortcomings in modality alignment, multi-task optimization, and robustness in high-noise scenarios. This paper comprehensively addresses the limitations of existing technologies by introducing a unified projection layer, a multi-task Prefix Tokens strategy, and a robustness enhancement mechanism. Summary of the Invention
[0011] To address the shortcomings of existing methods, the present invention provides a method, system, electronic device, and readable storage medium for fusing a small number of multi-modal medical data. The present invention comprehensively addresses the limitations of existing technologies by introducing a unified projection layer, a multi-task Prefix Tokens strategy, and a robustness enhancement mechanism. The present invention significantly improves the efficiency and quality of modality fusion.
[0012] The technical solution of the present invention is: a method for fusing a small number of types and multiple modalities of medical data, the method comprising:
[0013] Step 1: Extract features from input text data, time series data, and image data;
[0014] Step 2: Align and fuse the extracted single-modal features;
[0015] Step 3: Implement multi-task learning using a task differentiation strategy based on prefix tokens.
[0016] Step 4: Load the corresponding prefix token according to the task requirements of the input sample, and use the decoding method corresponding to the task to implement reasoning.
[0017] Furthermore, the Step 1 includes:
[0018] (1) Feature extraction of text data: For the input text data, directly use the text encoder to generate a high-dimensional representation vector; the encoding process is expressed as:
[0019] h text =LLM Encoder (X text )
[0020] Among them, X text ={x1,x2,...x n} is the text input sequence, h text ∈R n×d is a high-dimensional semantic vector representation, d is the hidden layer dimension;
[0021] (2) Feature extraction of time series data: The TimesFM model is used to generate a feature representation vector for time series data through a decoder; it is expressed as:
[0022] h time =TimesFM(X time )
[0023] Among them, X time ={t1,t2,...t m} is the time series data input, h time ∈R m×d is the generated feature representation vector;
[0024] (3) Feature extraction of image data: ViT is used to encode image data and extract multi-scale feature information; it is expressed as:
[0025] himg =ViT(X img )
[0026] Among them, X img is the image data input, h img ∈R p×d is the multi-scale feature extracted after ViT, where p represents the sequence length after image segmentation.
[0027] Furthermore, the Step 2 includes:
[0028] Step 2.1. Use the projection layer to map the features of the time series and image modalities to the latent feature space, which is convenient for multimodal splicing and alignment with text. Assume that the projection layer is a multi-layer perceptron, and the mapping formula is:
[0029] h img '=MLP(h img ),h time '=MLP(h time )
[0030] Step 2.2: Generate a fusion representation by concatenating the aligned features:
[0031] h input =Concat(h text ,h time ',h img ')
[0032] Among them, h input is the multimodal representation after fusion;
[0033] Step 2.3: The fused multimodal representation is further integrated with the features of each modality using the multi-head self-attention mechanism; it is expressed as:
[0034] h output =LLM(h input ).
[0035] Furthermore, in Step 3, the task differentiation strategy based on prefix tokens is adopted, and each text generation and time series prediction task is represented by a prefix token; the representation of the input data is:
[0036] h task =[p i ;h input ]
[0037] Among them, pi is the task t i Prefix Token;
[0038] The overall loss function is:
[0039]
[0040] Among them, α i is the loss weight of task i, l is the multi-task loss, and the multi-task loss l includes the generation loss l gen , classification loss lcls and time series prediction loss l pred .
[0041] Furthermore, the Step 4 includes:
[0042] (1) Generate text data: Use autoregressive decoding to generate text output;
[0043] (2) Predict and interpolate time series data: Based on the continuity requirements of the sequence characteristics, use sliding windows to predict future trends or fill in missing values;
[0044] (3) Detect anomalies: Based on the similarity of the output features of the contrast loss, a threshold is used to determine whether it is an anomaly point.
[0045] The present invention also provides a system for fusing medical data of a small number of categories, multiple types, and multiple modalities, the system comprising: a module for executing the above-mentioned method for fusing medical data of a small number of categories, multiple types, and multiple modalities.
[0046] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for fusing a small number of types, multiple types, and multimodal medical data described above is implemented.
[0047] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for fusing a small number of types, multiple types, and multi-modal medical data.
[0048] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for fusing a small number of types, multiple types, and multi-modal medical data.
[0049] The beneficial effects of the present invention are:
[0050] Through refined design, this invention achieves the following technical advantages in multimodal data analysis and multi-task learning:
[0051] (1) Efficient multimodal feature alignment and fusion: This paper uses a unified projection layer to map data features from different modalities into a shared latent feature space. Combined with a multi-head self-attention mechanism, this method achieves semantic alignment and deep feature interaction between modalities. This approach significantly improves the efficiency and quality of modal fusion and is suitable for complex analysis tasks involving heterogeneous multimodal data.
[0052] (2) Flexible and efficient multi-task collaborative learning: By introducing a task differentiation strategy based on prefix tokens, this paper achieves collaborative optimization of multiple tasks such as text generation, time series prediction, classification, and anomaly detection in a single model. At the same time, dynamic task sampling and adaptive learning rate adjustment ensure balanced performance across tasks.
[0053] (3) Excellent robustness and generalization: Leveraging pre-trained TimesFM and ViT models, the proposed method demonstrates remarkable robustness when dealing with small sample sizes, high noise, and incomplete data. Furthermore, the model is able to effectively adapt to diverse clinical data scenarios and maintain stable performance even in highly heterogeneous and abnormal data.
[0054] (4) Efficient end-to-end implementation: The present invention adopts an end-to-end design architecture, integrating feature extraction and task reasoning, reducing system complexity while improving the ease of use and deployability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0056] Example 1: Figure 1 As shown, a method for fusing a small number of types and multiple modalities of medical data is provided; the method comprises:
[0057] Step 1: Extract features from input text data, time series data, and image data;
[0058] Step 1 includes:
[0059] (1) Feature extraction of text data: For the input text data, the model does not need additional preprocessing and directly uses the text encoder to generate a high-dimensional representation vector; the text encoder in the large language model (LLM) is used to represent the text data. Let the text input sequence be X text ={x1,x2,...x n}, the encoder converts the input into a high-dimensional semantic vector representation h text ∈R n×d , where d is the hidden dimension; the encoding process is expressed as:
[0060] h text =LLMEncoder (X text )
[0061] (2) Feature extraction of time series data: The TimesFM model is used to generate feature representation vectors for time series data through a decoder. This paper introduces the TimesFM model as a time series feature extractor. TimesFM is a decode-only basic model that has been pre-trained on a large and diverse time series dataset containing 100 billion time points and can extract potential features of time series in multiple fields. The model generates representation vectors for time series data through a decoder to capture dynamic trends and patterns in the time series. Time series data input X time ={t1,t2,...t m}, use the pre-trained TimesFM model to extract its temporal features and generate the feature representation vector h time ∈R m×d ; which is expressed as:
[0062] h time =TimesFM(X time )
[0063] (3) Feature extraction of image data: ViT is used to encode image data and extract multi-scale feature information; it is expressed as:
[0064] h img =ViT(X img )
[0065] Among them, X img is the image data input, h img ∈R p×d is the multi-scale feature extracted by ViT, where p represents the sequence length after image segmentation. ViT demonstrates strong performance in image feature extraction and classification tasks. For clinical image data (CT, X-rays, and facial expressions), ViT can extract multi-scale feature information, providing high-quality representation for subsequent multimodal fusion.
[0066] Step 2: Align and fuse the extracted single-modal features;
[0067] Furthermore, the Step 2 includes:
[0068] Step 2.1. After completing the single-modal feature extraction, the present invention designs a unified alignment method based on the data characteristics of different modalities. The projection layer is used to map the features of the time series and image modalities to the latent feature space, which facilitates multimodal splicing and alignment with the text. The projection layer is assumed to be a multi-layer perceptron (MLP), and the mapping formula is:
[0069] h img '=MLP(h img ),h time '=MLP(h time )
[0070] Step 2.2: Generate a fusion representation by concatenating the aligned features:
[0071] h input =Concat(h text ,h time ',h img ')
[0072] Among them, h input is the multimodal representation after fusion;
[0073] Step 2.3, the fused multimodal representation h input Entering a large language model (such as LLaMA-3 or Qwen-2), the multi-head self-attention mechanism is used to further integrate the features of each modality; it is expressed as:
[0074] h output =LLM(h input ).
[0075] Step 3: Use a task differentiation strategy based on prefix tokens to implement multi-task learning. Each task (such as text generation and time series prediction) is represented by a unique prefix token.
[0076] Furthermore, in Step 3, the task differentiation strategy based on prefix tokens is adopted, and each text generation and time series prediction task is represented by a prefix token; the representation of the input data is:
[0077] h task =[p i ;h input ]
[0078] Among them, pi is the task t i Prefix Token;
[0079] The overall loss function is:
[0080]
[0081] Among them, α i is the loss weight of task i, which is dynamically adjusted according to the importance of the task, l is the multi-task loss, and the multi-task loss l includes the generation loss l gen, classification loss lcls and time series prediction loss l pred .
[0082] During the training process, a dynamic task sampling strategy is adopted. In each iteration, tasks are randomly sampled and input is performed according to the prefix token corresponding to the task. Through adaptive learning rate adjustment and gradient clipping, the balanced performance of the model on each task is ensured.
[0083] Step 4: Load the corresponding Prefix Token according to the task requirements of the input sample, and use the decoding method corresponding to the task to implement inference.
[0084] Furthermore, the Step 4 includes:
[0085] (1) Generate text data: Use autoregressive decoding to generate text output;
[0086] (2) Predict and interpolate time series data: Based on the continuity requirements of the sequence characteristics, use sliding windows to predict future trends or fill in missing values;
[0087] (3) Detect anomalies: Based on the similarity of the output features of the contrast loss, a threshold is used to determine whether it is an anomaly point.
[0088] The present invention also provides a system for fusing a small number of types of multi-modal medical data, the system comprising:
[0089] Feature extraction module, used to extract features from input text data, time series data, and image data;
[0090] Alignment and fusion module, used to align and fuse the extracted single-modal features;
[0091] Multi-task learning module, which is used to implement multi-task learning using a task differentiation strategy based on prefix tokens;
[0092] The inference module is used to load the corresponding prefix token according to the task requirements of the input sample and implement inference using the decoding method corresponding to the task.
[0093] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for fusing a small number of types, multiple types, and multimodal medical data described above is implemented.
[0094] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for fusing a small number of types, multiple types, and multi-modal medical data.
[0095] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for fusing a small number of types, multiple types, and multi-modal medical data.
[0096] The large model of the present invention for fusion of a small number of multi-modal data has wide application value in the following medical scenarios:
[0097] (1) Automatic generation of medical records: By integrating text, time series and medical imaging data, complete medical record documents are generated for patients, significantly reducing the paperwork burden of medical staff.
[0098] (2) Assisting in early disease screening: Utilizing the fused multimodal representation, it provides auxiliary support for early screening of chronic diseases such as cardiovascular disease and diabetes.
[0099] (3) Real-time anomaly detection: Real-time analysis of electrocardiogram and physiological data in intensive care units (ICUs) to assist in the detection of potential critical events and improve patient safety.
[0100] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A method for fusing medical data of a small number of categories, multiple types, and multiple modalities, characterized by: The method comprises: Step 1: Extract features from input text data, time series data, and image data; Step 2: Align and fuse the extracted single-modal features; Step 3: Implement multi-task learning by using a task differentiation strategy based on prefix tokens. Step 4: Load the corresponding prefix token according to the task requirements of the input sample, and use the decoding method corresponding to the task to implement reasoning.
2. The method for fusion of medical data of a small number of categories, multiple types and multiple modalities according to claim 1, characterized in that: Step 1 includes: (1) Feature extraction of text data: For the input text data, directly use the text encoder to generate a high-dimensional representation vector; the encoding process is expressed as: h text =LLM Encoder (X text ) Among them, X text ={x1,x2,...x n } is the text input sequence, h text ∈R n×d is a high-dimensional semantic vector representation, d is the hidden layer dimension; (2) Feature extraction of time series data: The TimesFM model is used to generate a feature representation vector for time series data through a decoder; it is expressed as: h time =TimesFM(X time ) Among them, X time ={t1,t2,...t m } is the time series data input, h time ∈R m×d is the generated feature representation vector; (3) Feature extraction of image data: ViT is used to encode image data and extract multi-scale feature information; it is expressed as: h img =ViT(X img ) Among them, X img is the image data input, h img ∈R p×d is the multi-scale feature extracted after ViT, where p represents the sequence length after image segmentation.
3. The method for fusion of medical data of a small number of categories, multiple types and multiple modalities according to claim 1 is characterized by: Step 2 includes: Step 2.
1. Use the projection layer to map the features of the time series and image modalities to the latent feature space, which is convenient for multimodal splicing and alignment with text. Assume that the projection layer is a multi-layer perceptron, and the mapping formula is: h img '=MLP(h img ),h time '=MLP(h time ) Step 2.2: Generate a fusion representation by concatenating the aligned features: h input =Concat(h text ,h time ',h img ') Among them, h input is the multimodal representation after fusion; Step 2.3: The fused multimodal representation is further integrated with the features of each modality using the multi-head self-attention mechanism; it is expressed as: h output =LLM(h input )。 4. The method for fusion of medical data of a small number of categories, multiple types and multiple modalities according to claim 1, characterized in that: In Step 3, the task differentiation strategy based on prefix token is adopted. Each text generation and time series prediction task is represented by prefix token. The representation of input data is: h task =[p i ;h input ] Among them, pi is the task t i Prefix Token; The overall loss function is: Among them, αi is the loss weight of task i, For multi-task loss, multi-task loss Includes generation loss Classification loss and time series prediction loss 5. The method for fusion of medical data of a small number of categories, multiple types and multiple modalities according to claim 1 is characterized by: Step 4 includes: (1) Generate text data: Use autoregressive decoding to generate text output; (2) Predict and interpolate time series data: Based on the continuity requirements of the sequence characteristics, use sliding windows to predict future trends or fill in missing values; (3) Detect anomalies: Based on the similarity of the output features of the contrast loss, a threshold is used to determine whether it is an anomaly point.
6. A system for fusion of medical data of few categories, multiple types and multiple modalities, characterized by: The system comprises: a module for executing the method according to any one of claims 1-5.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the method for fusing a small number of types, multiple types, and multi-modal medical data as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for fusing a small number of categories, multiple types and multi-modal medical data as described in any one of claims 1 to 5 is implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for fusing a small number of categories, multiple types and multi-modal medical data as described in any one of claims 1 to 5 is implemented.