Lightweight and efficient multi-modal data fusion method

Through an improved multimodal data fusion method, combined with specific preprocessing and dynamic weight fusion mechanism, the problems of single detection method and difficulty in model transplantation of traditional detection equipment are solved, and efficient and accurate disease detection is achieved.

CN120656025APending Publication Date: 2025-09-16GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510623177.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional fluorescent quantitative PCR detection equipment has a single detection method, which makes it difficult to accurately capture complete information about pathogens, and has low detection efficiency. In addition, deep learning network models are difficult to transplant on the equipment.

Method used

A lightweight and efficient multimodal data fusion method is adopted. Through the improved MobileNetV3-Small network and ResNet50 teacher model, combined with specific preprocessing steps and dynamic weight fusion mechanism, the microfluorescence and microfluidic fluorescence image features are fused to enhance the target signal and reduce noise interference.

Benefits of technology

It improves the accuracy and efficiency of disease detection, reduces computational complexity and memory usage, and enables efficient feature extraction and information fusion on the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656025A_ABST
    Figure CN120656025A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image fusion, and discloses a lightweight and efficient multi-modal data fusion method. According to the method, an improved MobileNetV3-Small network is adopted as a basic feature extractor, parallel branch design in a training stage is optimized through a structure re-parameterization technology, and the parallel branch design is combined into single convolution in a reasoning stage, so that the calculation cost is remarkably reduced; resNet50 is introduced as a teacher model, a soft label is generated through temperature scaling, a student model is guided in combination with KL divergence loss and cosine feature alignment loss, and high-precision feature extraction of a lightweight network is realized; a dynamic weight fusion mechanism is designed, microscopic fluorescence and microfluidic fluorescence image features are fused, and multi-modal information contribution is balanced; aiming at the characteristics of different modal images, preprocessing methods such as guided filtering, CLAHE contrast enhancement, time sequence difference artifact removal and the like are respectively adopted, and effective signals are enhanced. According to the algorithm, the pathogen detection precision is ensured, the model parameter quantity is reduced, and an efficient solution capable of being deployed in an embedded mode is provided for livestock and poultry epidemic disease screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image fusion technology, and in particular to a lightweight and efficient multimodal data fusion method. Background Art

[0002] In the livestock and poultry industry, rapid disease detection and precise prevention and control are crucial for ensuring the healthy development of animal husbandry. Traditional fluorescence quantitative PCR detection equipment suffers from a single detection method, making it difficult to accurately capture complete pathogen information, and has low detection efficiency. Therefore, it is necessary to design a multimodal fusion algorithm to fuse information from different views of the data to obtain more complete and accurate features.

[0003] Artificial intelligence algorithms can fuse and analyze multimodal data to improve detection accuracy and efficiency. However, traditional feature extraction network models are usually large, and there is a trade-off between size and performance. Considering the size limitations of the device, some deep and high-precision networks cannot be well transplanted to the device. Therefore, a method with acceptable size and performance that can fuse multimodal data is needed. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the present invention provides a lightweight and efficient multimodal data fusion method, which has the advantages of lightweight and efficient multimodal data fusion and solves the above technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a lightweight and efficient multimodal data fusion method, comprising the following steps:

[0006] S1: Collect microscopic fluorescence images and microfluidic fluorescence images, and preprocess them respectively to obtain preprocessed microscopic fluorescence images and microfluidic fluorescence images. The specific steps are as follows:

[0007] S101: Use bilinear interpolation for microfluidic fluorescence images and Lanczos interpolation for microscopic images;

[0008] S102: Guided filtering is used for microscopic fluorescence images, and adaptive Canny edge detection is adopted with dynamic double threshold detection. Non-maximum suppression and morphological closing operation are combined to connect broken edges, and CLAHE contrast enhancement is used.

[0009] S103: Gaussian blur is applied to estimate the background of the microfluidic fluorescence image, and the background estimate is subtracted to retain the target signal. Temporal differencing is used to remove artifacts, and the absolute difference between adjacent frames is calculated. The difference map is thresholded and segmented to extract the dynamic target region. The threshold is automatically determined using the Otsu algorithm, and a binary mask M is generated based on the threshold. That is, for each pixel in the image, pixels above the threshold are set to 1, and pixels below the threshold are set to 0. Mask M is obtained. The target signal is enhanced using binary mask M. Multi-frame averaging is used for enhancement, and a weighted average of five consecutive frames is performed.

[0010] S2: Construct a teacher model and a student model and train them separately. The teacher model is Resnet50, and the student model is the improved MobileNetV3-Small model. Specifically, the bottleneck module in the MobileNetV3-Small network is improved. The depthwise separable convolution with a kernel size of 3 and the SE layer are moved before the convolution layer with a kernel size of 1. A parallel branch of 1x1 depthwise separable convolution and a parallel residual connection are added. During the inference phase, the branches are fused into a 3x3 depthwise separable convolution.

[0011] S3: The processed microscopic fluorescence image and microfluidic fluorescence image output by S1 are input into the teacher model and the student model. The student model performs feature extraction on the input image and outputs the data results to obtain features. These features contain key information related to disease detection in the image.

[0012] S4: The features obtained by the student model are fed into the RPN network to generate candidate regions, obtain the coordinates of the upper left corner and the lower right corner of the candidate regions, and unify the candidate regions through ROIPooling in the RPN network;

[0013] S5: After obtaining the candidate region, the features of the microscopic image and the microfluidic fluorescence image are fused using the maximum and minimum weight fusion mechanism;

[0014] The specific expression of the maximum and minimum weight fusion mechanism is as follows:

[0015] F i =α·max(F1,F2)+(1-α)·min(F1,F2)

[0016] Where F1 and F2 represent the feature matrices of the two images respectively, and Fi represents the fused feature matrix. The fusion weight α is set as a trainable parameter and automatically optimized through backpropagation (initial α = 0.7).

[0017] As a preferred technical solution of the present invention, the expression of the dual threshold is as follows:

[0018] Tlow =0.1·max(I)

[0019] T high =0.3·max(I)

[0020] Where I represents the grayscale value of the image.

[0021] As a preferred technical solution of the present invention, the specific calculation expression of the absolute difference between adjacent frames is as follows:

[0022] D t (x,y)=|I t (x,y)-I t-1 (x,y)|

[0023] Where, It(x,y) represents the pixel value at the coordinate (x,y) position at time t (the t-th frame image), and Dt(x,y) represents the absolute value of the difference in pixel values ​​between two adjacent frames (frame t and frame (t-1)) at the coordinate (x,y) position at time t.

[0024] The specific expression for automatically determining the threshold using the Otsu algorithm is as follows:

[0025]

[0026] Where T represents the threshold; Represents the inter-class variance, which is used to measure the variance between the two parts after the image is divided into foreground and background by the threshold T;

[0027] The specific expression for enhancing the target signal by the binary mask M is as follows

[0028]

[0029] Among them, I enhanced Represents the enhanced image; I clean It represents the image after the previous processing to remove the background fluorescence; M represents the threshold;

[0030] Using multi-frame average enhancement, the specific expression for weighted averaging of 5 consecutive frames of images is as follows:

[0031]

[0032] Where n is the number of consecutive frames, I i Represents the i-th frame image; ω i Represents the weight coefficient corresponding to the i-th frame image;

[0033] Weight coefficient ω i The weight is allocated based on the time decay factor γ = 0.8, and the specific expression is as follows:

[0034] ω i =γ n-i (i=1,2,...,n)

[0035] Where γ represents the time decay factor; i represents a loop index from 1 to n, which is used to refer to each image in turn;

[0036] Compared with the existing technology, the present invention provides a lightweight and efficient multimodal data fusion method with the following beneficial effects:

[0037] The present invention designs special preprocessing processes for different modal image characteristics, such as edge operators and CLAHE contrast enhancement of microscopic images, temporal difference artifact removal and multi-frame weighted fusion of microfluidic images, which effectively suppresses noise interference and strengthens target signals, so that subsequent network feature extraction can focus more on key areas; through the improved MobileNetV3-Small network, combined with structural reparameterization technology, parallel branches are introduced in the training stage to enhance learning ability, and the inference stage is merged into a single convolution, which significantly reduces computational complexity and memory usage. A dynamic weight adjustment mechanism is used to fuse microscopic fluorescence and microfluidic fluorescence image features, replacing the traditional maximum fusion method, and balancing the maximum and minimum feature contributions through learnable weight parameters to avoid single modality dominance or information loss; ResNet50 is used as the teacher model, soft labels are generated through temperature scaling and combined with KL divergence loss and cosine feature alignment loss, so that the student model approaches the characterization ability of the teacher model under a lightweight architecture, significantly improving feature extraction accuracy while reducing the number of model parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a structural diagram of the multimodal data fusion algorithm prediction method of the present invention;

[0039] Figure 2 This is the bottleneck structure diagram of the improved MobileNetV3-Small of the present invention;

[0040] Figure 3 It is the network structure diagram of the multimodal data fusion algorithm. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] See also Figure 1-3, a lightweight and efficient multimodal data fusion method, including the following steps:

[0043] S1: Collect microscopic fluorescence images and microfluidic fluorescence images, and preprocess them respectively to obtain preprocessed microscopic fluorescence images and microfluidic fluorescence images. The specific steps are as follows:

[0044] S101: Use bilinear interpolation for microfluidic fluorescence images and Lanczos interpolation for microscopic images. Adjust the image resolution to 224×224. Use bilinear interpolation for microfluidic images to preserve details to balance computational efficiency and dynamic artifact suppression. Use Lanczos interpolation for microscopic images to preserve high-frequency details and reduce downsampling aliasing effects.

[0045] S102: Guided filtering is used on the microscopic fluorescence image, along with adaptive Canny edge detection and a dynamic double threshold for detection. Non-maximum suppression and morphological closing operations are used to connect broken edges, and CLAHE contrast enhancement is used. Microscopic fluorescence images are characterized by high resolution and complex background fluorescence, requiring precise segmentation of target structures. Guided filtering is used with a kernel size of 3×3 and a standard deviation of σ=1.5. The guide image is the original image to suppress sensor noise. Adaptive Canny edge detection is used, and the threshold is dynamically adjusted: T low =0.1·max(I), T high =0.3·max(I), T low T high Represent two thresholds, I represents the image grayscale value, max represents the maximum function, and min represents the minimum function. Non-maximum suppression (NMS) and morphological closing operation (kernel 5×5) are combined to connect the broken edges. CLAHE contrast enhancement is used, the histogram distribution is cropped, the block size is 8×8, the contrast limit is 2.0, and the fluorescence signal is locally enhanced.

[0046] S103: Microfluidic fluorescence images are characterized by dynamic flow artifacts, low signal-to-noise ratio, and temporal correlation. Gaussian blur is applied to estimate the background and eliminate background fluorescence interference. The background estimate is subtracted from the original image to retain the target signal. Temporal difference is used to remove artifacts and suppress flow smearing effects (such as signal blurring caused by fluid motion). The absolute difference between adjacent frames is calculated, and the difference map is thresholded to extract the dynamic target area. Inter-frame difference calculation: D t (x,y)=|I t (x,y)-I t-1 (x,y)|, using Otsu algorithm to automatically determine the threshold argmax T Represents the maximum value index, generates a binary mask M based on the threshold, and enhances the target signal through the mask: Use multi-frame average enhancement to improve the signal-to-noise ratio of weak signals, and perform weighted averaging on 5 consecutive frames of images (the weight decays with time). Weight distribution: distribute weights according to the time decay factor γ = 0.8: ω i =γ n-i (i=1,2,...,n), n is the number of consecutive frames (the default is 5), weighted fusion formula:

[0047]

[0048] For each original image I i Multiply by its corresponding weight coefficient ω i Then sum it up, that is, weighted sum each image according to its weight, The whole formula is equivalent to dividing the weighted image sum by the weight sum to obtain the fused image, which ensures the normalization of the weights and makes the fusion result more reasonable. t (x,y) represents the t-th frame image, the coordinates are the pixel values ​​at the (x,y) position, I t-1 (x,y) represents the t-1 frame image, the coordinates are the pixel values ​​at the (x,y) position, D t (x,y) represents the absolute value of the difference between the pixel values ​​of the t-frame image and the t-1-frame image at the coordinate position (x,y) at time t, and || represents the absolute value;

[0049] S2: Construct a teacher model and a student model and train them separately. The teacher model is Resnet50, and the student model is the improved MobileNetV3-Small model. Specifically, the bottleneck module in the MobileNetV3-Small network is improved. The depthwise separable convolution with a kernel size of 3 and the SE layer are moved before the convolution layer with a kernel size of 1. A parallel branch of 1x1 depthwise separable convolution and a parallel residual connection are added. During the inference phase, the branches are fused into a 3x3 depthwise separable convolution.

[0050] S201: Use ResNet50 as the teacher model and input microscopy and microfluidics images for fine-tuning. Based on the ImageNet pre-trained weights, fine-tuning is performed using cross-entropy loss and bounding box regression loss. The teacher model guides the student model by extracting feature maps from the ResNet50 backbone network for feature alignment loss and saving the teacher model's soft labels for the training data (for subsequent distillation loss calculation). The soft labels are generated as follows, with an initial temperature parameter of t = 5, and the teacher model output is smoothed:

[0051]

[0052] S202: Use the improved MobileNetV3-Small as the basic feature extraction network, Figure 2 As shown in Figure 1, during the training phase, parallel branches are added to the bottleneck module: 1×1 depthwise separable convolution and residual connection. During the inference phase, it is merged into a single 3×3 depthwise separable convolution to eliminate redundant calculations (structural reparameterization). ReLU activation function is used in shallow layers and h-swish activation function is used in deep layers to balance speed and accuracy. The specific structure is shown in Figure 1. Figure 2 As shown, the improved network is represented as MobileNetV3-Small-1;

[0053] S203. Combining the loss functions includes KL divergence loss (initial temperature coefficient t=5, gradually decreasing to t=1) and intermediate feature alignment loss (cosine similarity loss). The specific expression is:

[0054] L total =L distillation +λL align

[0055] Among them, L total Represents the total loss value of the model training process, which is ultimately used for backpropagation to calculate the gradient to update the comprehensive loss index of the model parameters; L distillation Represents the distillation loss, which is used to measure the difference between the output results (such as predicted probability distribution, etc.) of the student model and the teacher model; L align Represents the intermediate feature alignment loss, which is used to measure the similarity or alignment between the intermediate layer features of the teacher model and the student model.

[0056] S203, training phase: training is performed in combination with the teacher model. Data is input into both the teacher model and the student model. The initial learning rate is set to 1e-4. The iteration is 80 epochs. The input data is pre-processed multimodal images (microscopy + microfluidics). The distillation loss and feature alignment loss are jointly optimized to update the model parameters.

[0057] S3: The processed microscopic fluorescence image and microfluidic fluorescence image output by S1 are input into the teacher model and the student model. The student model performs feature extraction on the input image and outputs the data results to obtain features. These features contain key information related to disease detection in the image.

[0058] The scanned image is input into the teacher model and the soft label is output. In the process of generating the soft label, the output probability of the teacher model is smoothed by adjusting the temperature coefficient. At the same time, the image is input into the student model branch. The output of the student model also needs to generate a soft label through the Softmax function and temperature adjustment, using the Kullback-Leibler (KL) divergence. As a metric, in order to compensate for the influence of the temperature parameter, the KL divergence needs to be multiplied by the square of the temperature parameter (t 2 );

[0059] S4: The features obtained by the student model are fed into the RPN network (Region Proposal Network), which is used to generate candidate boxes. It is proposed by Faster R-CNN to generate candidate regions, obtain the coordinates of the upper left corner and the lower right corner of the candidate region, and unify the candidate regions through ROIPooling in the RPN network;

[0060] S401: RPN generates candidate regions and NMS filters candidate boxes.

[0061] S402, ROIPooling unifies the size of the candidate regions and outputs them.

[0062] S5: After obtaining the candidate region, the features of the microscopic image and the microfluidic fluorescence image are fused using the maximum and minimum weighted fusion mechanism. The fused features are input into the MLP Head for classification and regression. The fused features are input into the MLP Head, which outputs the classification probability (Softmax) and the bounding box regression parameters (Smooth L1 Loss).

[0063] The specific expression of the maximum and minimum weight fusion mechanism is as follows:

[0064] F i =α·max(F1,F2)+(1-α)·min(F1,F2)

[0065] Among them, F1 and F2 represent the feature matrices of the two images respectively, and Fi represents the fused feature matrix;

[0066] Next, we optimized the model for inference deployment: We removed parallel branches from the student model training phase and merged them into a single convolutional layer (structural reparameterization). Microscopic images were processed on a single frame basis, and the microfluidic image buffer was used to cache five consecutive frames for multi-frame fusion. The preprocessing module was deployed to the GPU (CUDA acceleration).

[0067] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A lightweight and efficient multimodal data fusion method, characterized by: The following steps are involved: S1: Collect microscopic fluorescence images and microfluidic fluorescence images, and preprocess them respectively to obtain preprocessed microscopic fluorescence images and microfluidic fluorescence images. The specific steps are as follows: S101: Use bilinear interpolation for microfluidic fluorescence images and Lanczos interpolation for microscopic images; S102: Guided filtering is used for microscopic fluorescence images, and adaptive Canny edge detection is adopted with dynamic double threshold detection. Non-maximum suppression and morphological closing operation are combined to connect broken edges, and CLAHE contrast enhancement is used. S103: Apply Gaussian blur to estimate the background of the microfluidic fluorescence image, subtract the background estimate, retain the target signal, use temporal difference to remove artifacts, calculate the absolute difference between adjacent frames, perform threshold segmentation on the difference map, extract the dynamic target area, and automatically determine the threshold using the Otsu algorithm. Generate a binary mask M based on the threshold. That is, for each pixel in the image, set it to 1 if it is greater than the threshold and set it to 0 if it is less than the threshold. This generates the binary mask M. The target signal is enhanced using the binary mask M. Multi-frame averaging is used for enhancement, and a weighted average is performed on 5 consecutive frames. S2: Construct a teacher model and a student model. The teacher model is Resnet50, and the student model is the improved MobileNetV3-Small model. Specifically, the bottleneck module in the MobileNetV3-Small network is improved. The depthwise separable convolution with a kernel size of 3 and the SE layer are moved before the convolution layer with a kernel size of 1. A parallel branch of 1x1 depthwise separable convolution and a parallel residual connection are added. During the inference phase, the branches are fused into a 3x3 depthwise separable convolution. S3: The processed microscopic fluorescence image and microfluidic fluorescence image output by S1 are input into the teacher model and the student model, and the student model performs feature extraction on the input image and outputs the data result to obtain the feature; S4: The features obtained by the student model are fed into the RPN network to generate candidate regions, obtain the coordinates of the upper left corner and the lower right corner of the candidate regions, and unify the candidate regions through ROIPooling in the RPN network; S5: After obtaining the candidate region, the features of the microscopic image and the microfluidic fluorescence image are fused using the maximum and minimum weight fusion mechanism; The specific expression of the maximum and minimum weight fusion mechanism is as follows: F i =α·max(F1,F2)+(1-α)·min(F1,F2) Among them, F1 and F2 represent the feature matrices of the two images respectively, F i Represents the fused feature matrix. The fusion weight α is set as a trainable parameter and automatically optimized through backpropagation. max represents the maximum value function, and min represents the minimum value function.

2. The lightweight and efficient multimodal data fusion method according to claim 1, characterized in that: The expression of the dual threshold is as follows: T low =0.1·max(I) T high =0.3·max(I) Among them, T low T high Represent two thresholds respectively, I represents the image grayscale value, max represents the maximum value function, and min represents the minimum value function.

3. The lightweight and efficient multimodal data fusion method according to claim 1, characterized in that: The specific calculation expression of the absolute difference between adjacent frames is as follows: D t (x,y)=|I t (x,y)-I t-1 (x,y)| Among them, I t (x,y) represents the t-th frame image, the coordinates are the pixel values ​​at the (x,y) position, I t-1 (x,y) represents the t-1 frame image, the coordinates are the pixel values ​​at the (x,y) position, D t (x,y) represents the absolute value of the difference between the pixel values ​​of the t-frame image and the t-1-frame image at the coordinate position (x,y) at time t, and || represents the absolute value; The specific expression for automatically determining the threshold using the Otsu algorithm is as follows: Where T represents the threshold; Represents the inter-class variance, which is used to measure the variance between the two parts after the image is divided into foreground and background by the threshold T. T Indicates the maximum value index; The specific expression for enhancing the target signal by the binary mask M is as follows Among them, I enhanced Represents the enhanced image, I clean It represents the image after the previous processing to remove the background fluorescence, M represents the threshold, represents element-wise product; Using multi-frame average enhancement, the specific expression for weighted averaging of 5 consecutive frames of images is as follows: Where n is the number of consecutive frames, I i Represents the i-th frame image; ω i Represents the weight coefficient corresponding to the i-th frame image, For each original image I i Multiply by its corresponding weight coefficient ω i Then sum it up, that is, weighted sum each image according to its weight, is the sum of all weight coefficients; Weight coefficient ω i The specific expression is as follows: oh i =c n-i (i=1,2,...,n) Wherein, γ represents the time decay factor; i represents the loop index from 1 to n, which is used to refer to each image in turn.

Citation Information

Cited By

  • A pipe piece liquid surface pouring detection method and device based on multi-modal perception

    CN122365307A