Raman image fusion multi-modal fusion classification method
By introducing a visual Transformer and a multi-scale residual convolutional network into Raman image analysis, and combining them with a multimodal fusion algorithm, a multimodal fusion classification model was constructed. This model solved the accuracy problems across batches, devices, and environments, and achieved cell classification with high robustness and high recognition rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO SINGLE CELL BIOTECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-21
AI Technical Summary
Existing Raman image multimodal fusion classification methods have limited accuracy across batches, devices, and environments, making it difficult to maintain high robustness.
We employ the visual cue adjustment mechanism from the Visual Transformer to inject a unique cue vector P into each domain. Combined with a multi-scale residual convolutional network and batch normalization, we construct a multi-modal fusion classification model through a multi-modal fusion algorithm to eliminate the influence of different acquisition conditions on visual and spectral features.
It significantly improves the average recognition rate of data under different acquisition conditions, reduces the amount of computation, maintains high stability of the model in modal degradation or equipment failure scenarios, and quickly adapts to unknown data through high-confidence pseudo-label self-training.
Smart Images

Figure CN121904752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Raman image analysis technology, and in particular to a Raman image fusion multimodal fusion classification method. Background Technology
[0002] Raman spectroscopy is a non-destructive, label-free analytical technique that acquires chemical and structural information from biological samples. Learn more about the applications of Raman imaging in cell biology, proteomics and genomics, tissue and biofluid diagnostics, and pathology. With the development of machine vision and artificial intelligence, the analysis of Raman images and Raman spectroscopy is gradually moving towards an intelligent era.
[0003] Chinese patent CN114660040A discloses a method for identifying single-cell microorganisms. The method includes comparing and analyzing collected single-cell Raman spectral data with Raman spectral data in a reference spectral database to screen out Raman spectral data that meets certain criteria; using the screened Raman spectral data as samples, calculating the minimum number of spectral detections for each sample based on specific spectral feature values in the spectral data samples; collecting spectral data corresponding to the number of spectral detections, and standardizing the spectral data using a calibration transfer model; storing the standardized spectral data and real-time collected single-cell image data in an omics database; and performing multimodal feature fusion on the image and spectral feature values based on cell images and spectral data in the single-cell phenomics database to achieve single-cell phenotypic data identification, increasing data completeness and thus improving the accuracy of single-cell species identification.
[0004] The aforementioned multimodal feature fusion method employs a static fusion approach based on fixed weights, the accuracy of which is limited by differences in the batches of Raman images and Raman spectra acquired, the acquisition equipment used, and the acquisition environment. Therefore, there is an urgent need for a Raman image fusion multimodal fusion classification method that can maintain high accuracy and robustness across batches, devices, and environments. Summary of the Invention
[0005] To address the shortcomings of related technologies, this invention provides a Raman image fusion multimodal fusion classification method. Based on the visual cue adjustment mechanism in the Visual Transformer, a unique cue vector P is injected into each domain to eliminate the influence of different acquisition conditions on visual features. Using a multi-scale residual convolutional network and batch normalization processing, the influence of different acquisition conditions on spectral features is eliminated. A multimodal fusion algorithm is used to obtain fused data of visual and spectral features, and a multimodal fusion classification model is constructed accordingly to integrate the advantages of images and spectra and improve the average recognition rate for data under different acquisition conditions.
[0006] This invention provides a Raman image fusion multimodal fusion classification method, which uses a multimodal fusion classification model for cell classification. The method for constructing the multimodal fusion classification model is as follows: S1. Acquire multiple sets of data, each set including cell images and corresponding Raman spectra; S2. Divide all data into multiple domains. Data within the same domain are acquired under the same conditions, while data in different domains are acquired under different conditions. Acquisition conditions include acquisition batch, acquisition device, and acquisition environment. Input cell images from different domains into a visual Transformer to obtain visual features. During the iteration of the visual Transformer, a visual prompting adjustment mechanism is introduced to inject a unique prompt vector P into each domain. Input Raman spectra from different domains into a multi-scale residual convolutional network to obtain spectral features. During each iteration of the multi-scale residual convolutional network, batch normalization is performed on each spectral feature channel. S3. Obtain fused data of visual features and spectral features based on multimodal fusion algorithm; S4. The fused data is calibrated, and a multimodal fusion classification model is constructed based on the neural network, with the fused data as input and cell type as output.
[0007] In some embodiments, S3 obtains visual features based on a cross-attention mechanism by fusing spectral information. and spectral features of fused image information ;for and Each data point is assigned a weight and then summed together to obtain the merged data.
[0008] In some embodiments, S3 and The superposition, through accomplish, Fusion features representing the same cellular characteristics; The confidence prediction model is obtained using a neural network-based approach. and As input, with This is the output.
[0009] In some embodiments, S3 further includes: processing the fusion data of each cell based on a self-attention mechanism so that each fusion feature contains information of all fusion features belonging to the same cell; and S4 calibrating the fusion data processed based on the self-attention mechanism.
[0010] In some embodiments, a multi-head self-attention mechanism is introduced during the visual Transformer iteration process to process all visual feature vectors of the same iteration layer in parallel, so that each visual feature vector is fused with the information of all visual feature vectors of the same iteration layer.
[0011] In some embodiments, each iteration of the visual Transformer undergoes layer normalization.
[0012] In some embodiments, an Adapter module is inserted into the iteration layer of the visual Transformer.
[0013] In some embodiments, the cell image is transformed into a cell image under different acquisition conditions through a mapping function, and the cell image transformed by the mapping function is used as a training sample for the visual Transformer; the mapping function includes at least one random transformation among color transformation, exposure adjustment, Gaussian blur, and grayscale conversion.
[0014] In some embodiments, S4 uses KL divergence, combined with maximum mean difference (MMD), correlation alignment (CORAL), and center loss to constrain and optimize the fused data, reducing the distribution differences of the fused data in different domains; during the iteration of the multimodal fusion classification model, all normalized statistics, including layer normalized statistics and batch normalized statistics, are dynamically updated, and high-confidence pseudo-labels are generated based on the fused data. The high-confidence pseudo-labels guide the optimization of the multimodal fusion classification model in a self-training manner.
[0015] In some embodiments, in S1, images of the same cell at different focal planes are acquired, and all focal plane images of the same cell are fused into a single cell image based on an attention network; in S1, spectra of the same cell at different locations are acquired, and all spectra of the same cell are fused into a single Raman spectrum based on an attention network.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention is based on the visual cue adjustment mechanism in the visual Transformer, injecting a dedicated cue vector P into each domain to eliminate the influence of different acquisition conditions on visual features; based on a multi-scale residual convolutional network, combined with batch normalization processing, it eliminates the influence of different acquisition conditions on spectral features; based on a multimodal fusion algorithm, it obtains fused data of visual features and spectral features, and constructs a multimodal fusion classification model to integrate the advantages of images and spectra and improve the average recognition rate of data under different acquisition conditions.
[0017] 2. The visual Transformer does not require full model adjustment. It only needs to update the cue vector P and Adapter parameters corresponding to the acquisition conditions. The update amount is less than 5% of the total parameters of the visual Transformer, which can significantly improve the adaptability of the visual Transformer to cell images under different acquisition conditions and greatly reduce the amount of computation.
[0018] 3. Through cross-attention mechanism, confidence prediction model, KL divergence, maximum mean difference (MMD), correlation alignment (CORAL) and center loss, the multimodal fusion classification model can maintain high stability in modal degradation or equipment failure scenarios, and achieve complementary enhancement of image and spectral information.
[0019] 4. During the iteration process of the inference phase, the multimodal fusion classification model dynamically updates all normalized statistics, including layer normalized statistics and batch normalized statistics, and generates high-confidence pseudo-labels based on the fused data. The high-confidence pseudo-labels guide the multimodal fusion classification model to optimize in a self-training manner, so as to quickly adapt to unknown data in the absence of human intervention. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating the construction process of the multimodal fusion classification model in a specific embodiment of the present invention; Figure 2 This is the confusion matrix for cell classification prediction based on cell images; Figure 3 This is the confusion matrix for cell classification prediction based on Raman spectroscopy. Figure 4 This is a confusion matrix for cell classification prediction using the multimodal fusion classification model in a specific embodiment of the present invention. Detailed Implementation
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0022] In the description of this invention, it should be understood that the terms "center", "lateral", "longitudinal", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0023] The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more of that feature.
[0024] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0025] like Figure 1-4 As shown in the illustrative embodiment of the Raman image fusion multimodal fusion classification method of the present invention, this Raman image fusion multimodal fusion classification method uses a multimodal fusion classification model for cell classification. The method for constructing the multimodal fusion classification model is as follows: S1. Acquire multiple sets of data, each set including cell images and corresponding Raman spectra; S2. Divide all data into multiple domains. Data collection conditions are the same in the same domain, and data collection conditions are different in different domains. Acquisition conditions include acquisition batch, acquisition equipment, and acquisition environment; cell images from different domains are input into the visual Transformer to obtain visual features. During the iteration process of the visual Transformer, a visual prompting tuning mechanism is introduced to inject a unique prompt vector P into each domain; Visual features The matrix is used to divide the cell image into uniformly sized image patches, and each image patch is converted into a visual feature vector. d represents the number of visual feature vectors, and d represents the dimension of the visual feature vectors; P is a k*e matrix, where k represents the dimension of the visual feature vectors, and e represents the cue length, which is manually set.
[0026] Raman spectra from different domains are input into a multi-scale residual convolutional network to obtain spectral features. During each iteration, the multi-scale residual convolutional network performs batch normalization on each spectral feature channel.
[0027] Spectral characteristics are The Raman spectrum is divided into bands of uniform size, and each band is converted into an optical eigenvector. represents the number of spectral eigenvectors, and f represents the dimension of the eigenvectors.
[0028] S3. Obtain fused data of visual features and spectral features based on multimodal fusion algorithm; S4. The fused data is calibrated, and a multimodal fusion classification model is constructed based on the neural network, with the fused data as input and cell type as output.
[0029] This embodiment uses the visual cue adjustment mechanism in the Visual Transformer to inject a unique cue vector P into each domain, eliminating the influence of different acquisition conditions on visual features. Based on a multi-scale residual convolutional network and batch normalization processing, it eliminates the influence of different acquisition conditions on spectral features. Based on a multimodal fusion algorithm, it obtains fused data of visual and spectral features, and constructs a multimodal fusion classification model to integrate the advantages of images and spectra and improve the average recognition rate for data under different acquisition conditions.
[0030] In some embodiments, S3 obtains visual features based on a cross-attention mechanism by fusing spectral information. and spectral features of fused image information ;for and Each data point is assigned a weight and then summed together to obtain the merged data.
[0031] The cross-attention mechanism enables mutual learning and reinforcement of visual and spectral features to form more comprehensive fused data, which helps to deepen the understanding of cross-modal relationships in multimodal fusion classification models and improve their discrimination ability.
[0032] In some embodiments, S3 and The superposition, through accomplish, Fusion features representing the same cellular characteristics; The confidence prediction model is obtained using a neural network-based approach. and As input, with This is the output.
[0033] Adding a confidence prediction model can dynamically balance the contributions of visual features and spectral features, preventing a single modality from significantly biasing the overall decision.
[0034] In some embodiments, S3 further includes: processing the fusion data of each cell based on a self-attention mechanism so that each fusion feature contains information of all fusion features belonging to the same cell; and S4 calibrating the fusion data processed based on the self-attention mechanism.
[0035] Self-attention mechanisms can dynamically capture the complex dependencies between all fusion features of the same cell, focusing on key information and further improving the recognition ability of multimodal fusion classification models.
[0036] In some embodiments, a multi-head self-attention mechanism is introduced during the iteration of the visual Transformer to process all visual feature vectors of the same iteration layer in parallel, which significantly improves the training and inference efficiency of the visual Transformer. At the same time, it enables each visual feature vector to fuse information from all visual feature vectors of the same iteration layer to comprehensively understand the local and global information in the cell image.
[0037] In some embodiments, each iterative layer of the visual Transformer undergoes layer normalization. By normalizing the input of each sample, the covariate shift within the visual Transformer is reduced, making the gradients more stable and preventing gradient vanishing or exploding. This process is computationally efficient and memory-saving, making it particularly suitable for processing large images or long sequences. Layer normalization enables stable training of the visual Transformer, accelerates convergence, and enhances its image processing and generalization capabilities.
[0038] In some embodiments, an Adapter module is inserted into the iterative layers of the visual Transformer to achieve efficient fine-tuning with extremely low parameter increments while maintaining model performance. This significantly reduces training costs, avoids catastrophic forgetting, and supports continuous learning by freezing the parameters of the pre-trained visual Transformer model and fine-tuning only the Adapter.
[0039] In some embodiments, cell images are transformed into cell images under different acquisition conditions through a mapping function. The cell images transformed by the mapping function are used as training samples for the visual Transformer, enabling the visual Transformer to learn a wider range of image distributions and improving the versatility and anti-interference ability of the visual Transformer. The mapping function includes at least one random transformation among color transformation, exposure adjustment, Gaussian blur, and grayscale conversion.
[0040] Furthermore, this includes all random transformations such as color transformation, exposure adjustment, Gaussian blur, and grayscale conversion.
[0041] In some embodiments, S4 uses KL divergence to guide the learning of low-confidence features from visual and spectral features with high confidence, achieving feature complementarity and knowledge sharing. Based on this, the fused data is constrained and optimized by combining maximum mean difference (MMD), correlation alignment (CORAL), and center loss to reduce the distribution differences of fused data from different domains. During the iteration process of the multimodal fusion classification model, all normalized statistics, including layer normalized statistics and batch normalized statistics, are dynamically updated, and high-confidence pseudo-labels are generated based on the fused data. These high-confidence pseudo-labels guide the optimization of the multimodal fusion classification model through self-training, improving the model's adaptability to unknown data and solving the data distribution drift problem during the model training and testing phases.
[0042] In some embodiments, in S1, images of the same cell at different focal planes are acquired. Based on an attention network, all focal plane images of the same cell are fused into a single cell image. This helps the visual Transformer dynamically focus on key regions of the cell, enriching the details of visual features, improving the robustness of the visual Transformer, reducing the impact of single focal plane blur or noise on subsequent classification results, and enhancing the generalization ability of the visual Transformer. In S1, spectra of the same cell at different locations are acquired. Based on an attention network, all spectra of the same cell are fused into a single Raman spectrum. This helps the multi-scale residual convolutional network dynamically focus on key regions of the cell, enriching the details of spectral features, improving the robustness of the multi-scale residual convolutional network, reducing the impact of single-band noise on subsequent classification results, and enhancing the generalization ability of the multi-scale residual convolutional network.
[0043] Furthermore, before fusing all spectra of the same cell into a single Raman spectrum, the spectra input to the attention network undergo quality control and normalization. Quality control improves spectral quality and typically includes baseline correction, noise reduction, and smoothing. These steps effectively reduce instrument or environmental errors, making the data more reliable. Normalization adjusts the spectra to a uniform scale, facilitating comparisons between different samples. This step eliminates intensity differences caused by concentration or experimental conditions, resulting in more consistent analysis.
[0044] The Raman image fusion multimodal fusion classification method described in the preferred embodiment of the present invention includes all the features in all the above embodiments.
[0045] Figure 2-4 The confusion matrix for different models targeting the same three types of human cells. Figure 2 and Figure 3The confusion matrix of a single-modal classification model. Figure 2 Using an image classification model, Figure 3 A spectral classification model is used. Figure 4 The multimodal fusion classification model described in this embodiment is adopted.
[0046] Figure 2 In the study, the image classification model performed well in classifying human hepatic stellate cells, but its performance was poor in classifying liver cancer cells and immune subtype breast cancer cells.
[0047] Figure 3 In the study, the spectral classification model performed well in classifying liver cancer cells and immune subtype breast cancer cells, but performed poorly in classifying human hepatic stellate cells.
[0048] Figure 4 The multimodal fusion classification model integrates the advantages of image and spectral methods, demonstrating superior classification performance for all three cell types. The average recognition rate of the multimodal fusion classification model described in this embodiment is approximately 10%–15% higher than that of the single-modal classification model.
[0049] Through the description of several embodiments of the Raman image fusion multimodal fusion classification method of the present invention, it can be seen that the embodiments of the Raman image fusion multimodal fusion classification method of the present invention have at least one or more of the following advantages: 1. This invention is based on the visual cue adjustment mechanism in the visual Transformer, injecting a dedicated cue vector P into each domain to eliminate the influence of different acquisition conditions on visual features; based on a multi-scale residual convolutional network, combined with batch normalization processing, it eliminates the influence of different acquisition conditions on spectral features; based on a multimodal fusion algorithm, it obtains fused data of visual features and spectral features, and constructs a multimodal fusion classification model to integrate the advantages of images and spectra and improve the average recognition rate of data under different acquisition conditions.
[0050] 2. The visual Transformer does not require full model adjustment. It only needs to update the cue vector P and Adapter parameters corresponding to the acquisition conditions. The update amount is less than 5% of the total parameters of the visual Transformer, which can significantly improve the adaptability of the visual Transformer to cell images under different acquisition conditions and greatly reduce the amount of computation.
[0051] 3. Through cross-attention mechanism, confidence prediction model, KL divergence, maximum mean difference (MMD), correlation alignment (CORAL) and center loss, the multimodal fusion classification model can maintain high stability in modal degradation or equipment failure scenarios, and achieve complementary enhancement of image and spectral information.
[0052] 4. During the iteration process of the inference phase, the multimodal fusion classification model dynamically updates all normalized statistics, including layer normalized statistics and batch normalized statistics, and generates high-confidence pseudo-labels based on the fused data. The high-confidence pseudo-labels guide the multimodal fusion classification model to optimize in a self-training manner, so as to quickly adapt to unknown data in the absence of human intervention.
[0053] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0054] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A Raman image fusion multimodal fusion classification method, characterized in that, Cell classification is performed using a multimodal fusion classification model. The method for constructing the multimodal fusion classification model is as follows: S1. Acquire multiple sets of data, each set including cell images and corresponding Raman spectra; S2. Divide all data into multiple domains. Data within the same domain are acquired under the same conditions, while data in different domains are acquired under different conditions. Acquisition conditions include acquisition batch, acquisition device, and acquisition environment. Input cell images from different domains into a visual Transformer to obtain visual features. During the iteration of the visual Transformer, a visual prompting adjustment mechanism is introduced to inject a unique prompt vector P into each domain. Input Raman spectra from different domains into a multi-scale residual convolutional network to obtain spectral features. During each iteration of the multi-scale residual convolutional network, batch normalization is performed on each spectral feature channel. S3. Obtain fused data of visual features and spectral features based on multimodal fusion algorithm; S4. The fused data is calibrated, and a multimodal fusion classification model is constructed based on the neural network, with the fused data as input and cell type as output.
2. The Raman image fusion multimodal fusion classification method according to claim 1, characterized in that, S3 uses a cross-attention mechanism to obtain visual features that fuse spectral information. and spectral features of fused image information ;for and Each data point is assigned a weight and then summed together to obtain the merged data.
3. The Raman image fusion multimodal fusion classification method according to claim 2, characterized in that, S3 and The superposition, through accomplish, Fusion features representing the same cellular characteristics; The confidence prediction model is obtained using a neural network-based approach. and As input, with This is the output.
4. The Raman image fusion multimodal fusion classification method according to claim 2, characterized in that, S3 also includes: processing the fusion data of each cell based on the self-attention mechanism so that each fusion feature contains information of all fusion features belonging to the same cell; S4 calibrating the fusion data processed based on the self-attention mechanism.
5. The Raman image fusion multimodal fusion classification method according to any one of claims 1-4, characterized in that, During the iteration of the visual Transformer, a multi-head self-attention mechanism is introduced to process all visual feature vectors of the same iteration layer in parallel, so that each visual feature vector is fused with the information of all visual feature vectors of the same iteration layer.
6. The Raman image fusion multimodal fusion classification method according to claim 5, characterized in that, Each iteration of the visual Transformer undergoes layer normalization.
7. The Raman image fusion multimodal fusion classification method according to claim 6, characterized in that, Insert the Adapter module into the iteration layer of the Visual Transformer.
8. The Raman image fusion multimodal fusion classification method according to any one of claims 1-4, characterized in that, Cell images are transformed into cell images under different acquisition conditions through a mapping function, and the cell images transformed by the mapping function are used as training samples for the visual Transformer; the mapping function includes at least one random transformation among color transformation, exposure adjustment, Gaussian blur, and grayscale conversion.
9. The Raman image fusion multimodal fusion classification method according to any one of claims 1-4, characterized in that, S4 uses KL divergence, combined with maximum mean difference (MMD), correlation alignment (CORAL), and center loss to constrain and optimize the fused data, reducing the distribution differences of the fused data in different domains. During the iteration of the multimodal fusion classification model, all normalized statistics, including layer normalized statistics and batch normalized statistics, are dynamically updated, and high-confidence pseudo-labels are generated based on the fused data. The high-confidence pseudo-labels guide the optimization of the multimodal fusion classification model in a self-training manner.
10. The Raman image fusion multimodal fusion classification method according to any one of claims 1-4, characterized in that, S1 acquires images of the same cell at different focal planes, and uses an attention network to fuse all the focal plane images of the same cell into a single cell image; S1 also acquires spectra of the same cell at different locations, and uses an attention network to fuse all the spectra of the same cell into a single Raman spectrum.
Citation Information
Patent Citations
Microorganism single cell species identification method, device, medium and equipment
CN114660040A