Lightweight data fusion method based on dynamic multi-scale dual-channel attention
Through the lightweight data fusion method of dynamic multi-scale dual-channel attention, the problems of high computing complexity and low efficiency in multi-modal data fusion are solved, and efficient and accurate multi-modal information processing is achieved, suitable for intelligent security, autonomous driving and multimedia content analysis.
Patent Information
- Application Number
- CN202510707045.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The prior art has scenarios in multimodal data fusion with high computational complexity and low efficiency, which cannot meet the requirements of real-time. The traditional methods cannot fully consider the complex correlation between modals, resulting in rough fusion effects and loss of information.
A lightweight data fusion method with dynamic multi-scale dual-channel attention is adopted, and efficient fusion of RGB images and point cloud data is achieved through phased feature extraction, cross-modal similarity calculation and global context cascade structure, combined with attention mechanism.
It realizes efficient and accurate multimodal information processing, improves the expression ability and semantic richness of features, and is suitable for fields such as intelligent security, autonomous driving and multimedia content analysis.
Smart Images

Figure CN120236174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of lightweight data fusion, and particularly to a lightweight data fusion method based on dynamic multi-scale dual-channel attention. Background Art
[0002] In the era of digital information explosion, data is presented in various modal forms, including text, images, audio, and video, etc., each carrying unique information value; images can intuitively display visual features, text can accurately convey semantics, and audio can transmit sound characteristics. However, single-modal information processing can no longer meet the needs of many fields for high-precision and deep-level information mining. Traditional information fusion methods have obvious limitations. The early fusion methods based on simple rules only mechanically stitch or add different modal data with fixed weights, without fully considering the complex correlations between modalities, resulting in rough fusion effects and information loss. The subsequent fusion methods based on shallow models, although attempting to capture modal correlations, are limited by the model representation ability and are difficult to handle large-scale and highly complex data scenarios. Especially when dealing with massive and real-time information, it is difficult to balance efficiency and accuracy.
[0003] With the development of deep learning, convolutional neural networks have performed excellently in single-modal tasks. However, in multi-modal fusion scenarios, standard convolutional neural networks lack a dynamic and adaptive fusion mechanism for different modal features. Although the self-attention mechanism can focus on the key parts of the data, when directly applied to large-scale multi-modal data, the computational complexity is high and the processing efficiency is low, which cannot meet the requirements of high-real-time scenarios. Summary of the Invention
[0004] Based on this, it is necessary to provide a lightweight data fusion method based on dynamic multi-scale dual-channel attention, which includes: S1: Obtain image data and point cloud data, and perform staged feature extraction on the point cloud data and the preprocessed image data respectively to obtain point cloud features in four stages and image features in five stages; S2: Align the dimension of the point cloud features in the first stage with the dimension of the image features; based on the cosine similarity calculation formula, calculate the cross-modal similarities between the aligned point cloud features in the first stage and the image features in the first stage and the second stage respectively, enhance the aligned point cloud features in the first stage based on the two cross-modal similarities respectively, splice the image features in the first stage, the image features in the second stage, and the two enhanced point cloud feature maps, and normalize the splicing result through layer normalization to obtain the first fusion feature map; S3: Correspond the remaining stage image features and the remaining stage point cloud features in sequence one by one, and perform feature fusion on the three groups of corresponding image features and point cloud features respectively through step S2 to obtain the second fusion feature map, the third fusion feature map, and the fourth fusion feature map; S4: Pass the four fused feature maps through the global context cascading structure to obtain the first global context feature, the second global context feature, the third global context feature, and the fourth global context feature respectively; S5: Pass the four global context features through the attention mechanism respectively to obtain the fused and optimized data at four levels.
[0005] Beneficial effects: This method respectively performs phased feature extraction on the point cloud data and the preprocessed image data to obtain the point cloud features in four stages and the image features in five stages; align the dimensions of the point cloud features in the first stage with the dimensions of the image features; calculate the cross-modal similarities between the aligned point cloud features in the first stage and the image features in the first and second stages respectively, and enhance the aligned point cloud features in the first stage based on the two cross-modal similarities, splice the image features in the first stage, the image features in the second stage, and the two enhanced point cloud feature maps, and normalize the splicing result through layer normalization to obtain the first fused feature map; correspond the remaining stage of image features and the remaining stage of point cloud features in sequence one by one, perform feature fusion on the corresponding three groups of image features and point cloud features respectively to obtain the second fused feature map, the third fused feature map, and the fourth fused feature map; pass the four fused feature maps through the global context cascading structure, and pass the four obtained global context features through the attention mechanism respectively to obtain the fused and optimized data at four levels. This method effectively solves the deficiencies of the existing technology, meets the urgent needs of efficient and accurate multi-modal information processing in fields such as intelligent security, autonomous driving, and multimedia content analysis, and provides strong support for the intelligent upgrade of the industry. Description of the Drawings
[0006] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0007] Figure 1 It is a flowchart of the lightweight data fusion method based on dynamic multi-scale dual-channel attention in the embodiments of the present application. Detailed Embodiments
[0008] To make the above objects, features, and advantages of the present application more apparent and understandable, the following describes the specific embodiments of the present application in detail with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0009] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0010] As Figure 1 shown, this embodiment provides a lightweight data fusion method based on dynamic multi-scale dual-channel attention. The method includes: S1: Obtain RGB image data and point cloud data, and perform staged feature extraction on the point cloud data and the preprocessed image data respectively to obtain point cloud features in four stages and image features in five stages.
[0011] In this embodiment, the preprocessing of the image data includes performing adaptive histogram equalization and Gaussian filtering on the obtained image data.
[0012] Furthermore, the staged feature extraction of the preprocessed image data includes: Successively passing the preprocessed image data through the shallow and deep layers of the ShuffleNetV2 network, an EfficientNet-Lite network, and two MobileNetV2 networks to obtain image features at each stage respectively.
[0013] Specifically, input the preprocessed image data into the ShuffleNetV2 network, pass through the first convolutional layer and the max pooling layer in the ShuffleNetV2 network, and output the image features of the first stage; Pass the image features of the first stage through the subsequent convolutional layers and the channel shuffle module in the ShuffleNetV2 network to output the image features of the second stage.
[0014] The ShuffleNetV2 network is a lightweight convolutional neural network designed specifically for mobile and embedded devices. It optimizes computational efficiency through channel shuffle and pointwise grouped convolutions. Its core features include: channel shuffle, which addresses the "channel isolation" problem of grouped convolutions and enhances the efficiency of feature flow through cross-group information interaction; balanced computational complexity, with an equal-width channel design to reduce memory access costs and improve inference speed; and a multi-branch structure that combines pointwise grouped convolutions to maintain feature representation ability at low computational costs.
[0015] Input the image features of the second stage into the EfficientNet-Lite network to output the image features of the third stage.
[0016] The EfficientNet-Lite network is a lightweight version of Google's EfficientNet series. Based on a compound scaling strategy (simultaneously optimizing network depth, width, and resolution), it achieves a balance between accuracy and efficiency and is suitable for edge device deployment. Its core features include: compound scaling, which avoids the performance bottleneck of single-dimensional optimization by uniformly adjusting network dimensions; lightweight design, which simplifies the structure (such as reducing the number of layers) and combines depthwise separable convolutions and squeeze-and-excitation attention mechanisms to enhance feature sensitivity; and multi-scale adaptability, which has good generalization ability for input images of different sizes.
[0017] Input the image features of the third stage into the MobileNetV2 network to output the image features of the fourth stage.
[0018] Input the image features of the fourth stage into the MobileNetV2 network to output the image features of the fifth stage.
[0019] The MobileNetV2 network is a classic lightweight convolutional neural network proposed by Google. Based on inverted residual blocks and linear bottleneck design, it addresses the information loss problem in low-dimensional feature extraction. Its core features include: inverted residual blocks, which first expand the dimension and then reduce it, extracting spatial features through depthwise convolutions to avoid feature sparsity in shallow networks; linear bottleneck, which uses a linear activation at the end of the residual block to retain the linear information of low-dimensional features and improve representation ability; and efficient computation, based on depthwise separable convolutions, which significantly reduces the number of parameters and computational costs.
[0020] Furthermore, the staged feature extraction of point cloud data includes: Based on the features of all voxels divided from the point cloud data, obtain the point cloud features of the first stage; based on the point cloud features of the first stage, successively perform three aggregations of point cloud features with gradually decreasing neighborhoods to obtain the point cloud features of the remaining three stages.
[0021] Specifically, divide the three-dimensional space of the point cloud data into multiple voxels of the same size; for each voxel, calculate the average coordinates of all the points contained therein, and use the average coordinates as the voxel feature of the corresponding voxel. Input all the voxel features into a convolutional neural network to output the point cloud features in the first stage.
[0022] This stage mainly extracts the low-level features of the point cloud data, such as the spatial distribution and local geometric shape of the point cloud. These features provide a basis for subsequent feature extraction and fusion.
[0023] For the point cloud features in the first stage, define a set of kernel points (such as the center points of local regions or key anchor points) in the point cloud data, calculate the Euclidean distances between each kernel point and the remaining points, and calculate the point cloud features in the second stage based on the distances and the point cloud features in the first stage. The calculation formula is as follows: ; Where, represents the point cloud feature in the second stage of the i th kernel point; represents the neighborhood point set of the i th kernel point; represents the Gaussian kernel function, which performs a non-linear mapping on the distance. The smaller the distance, the closer the mapping result is to 1. When the distance exceeds the preset threshold, the mapping result approaches 0, which can implicitly define the neighborhood range; represents the i th kernel point and its j th neighboring point; represents the point cloud feature in the first stage of the j th neighboring point; Traverse all the kernel points in the point cloud data, and replace the corresponding point cloud features in the first stage with the point cloud features in the second stage of each kernel point to obtain the point cloud features in the second stage.
[0024] This stage continues to deepen the feature extraction, enhancing the expression ability and semantic information of the features. The point cloud features in the second stage are further reduced in resolution, but the semantic information is more concentrated and abstract. For example, in a complex scene containing multiple objects, the point cloud features in the second stage can more accurately identify the local features of different objects.
[0025] For the point cloud features in the second stage, randomly sample the points in the point cloud data. For each sampled point, calculate the second distance between each sampled point and the corresponding neighboring point, and calculate the point cloud features in the third stage based on the second distance and the point cloud features in the second stage. The calculation formula is as follows: ; Where, represents the point cloud feature in the third stage of the g th sampled point; Denote the neighborhood point set of the g th sampling point; Denote the Gaussian kernel function; Denote the g th sampling point and the second distance between it and its h th neighboring point; Denote the point cloud feature of the second stage of the h th neighboring point; Traverse all sampling points, and replace the corresponding point cloud features of the second stage with the point cloud features of the third stage of each sampling point to obtain the point cloud features of the third stage.
[0026] The features extracted in this stage reach a higher semantic level and contain richer semantic information in the point cloud data. The resolution of the point cloud features in the third stage is further reduced, but the semantic information contained in each feature point is more abundant. For example, in a complex scene containing multiple objects, the point cloud features in the third stage can more accurately identify the category and shape features of different objects.
[0027] For the point cloud features in the third stage, randomly sample the points in the point cloud data of all sampling points again. For each second sampling point, calculate the third distance between each second sampling point and its corresponding neighboring point, and calculate the point cloud features of the fourth stage based on the third distance and the point cloud features of the third stage. The calculation formula is as follows: ; where, Denote the point cloud feature of the fourth stage of the u th sampling point; Denote the neighborhood point set of the u th sampling point; Denote the Gaussian kernel function; Denote the u th sampling point and the third distance between it and its v th neighboring point; Denote the point cloud feature of the third stage of the v th neighboring point; Traverse all second sampling points, and replace the corresponding point cloud features of the third stage with the point cloud features of the fourth stage of each second sampling point to obtain the point cloud features of the fourth stage.
[0028] The features extracted in this stage reach the highest semantic level and contain the most core semantic information in the point cloud data. The resolution of the point cloud features in the fourth stage is the lowest, but each feature point condenses the most critical semantic information in the point cloud data. For example, in a complex scene containing multiple objects, the point cloud features in the fourth stage can extract the features that best represent the core semantics of the entire scene, providing key support for the final data fusion and decision-making.
[0029] In this embodiment, the point cloud features in the four stages all use the same type of Gaussian kernel function. By only adjusting the neighborhood range or the kernel function parameters, the feature scales of different stages can be adapted. For example, in the first stage, fine geometric features (such as edges and surface normal vectors) are extracted through a small neighborhood range, and in the subsequent stages, semantic-level features (such as object parts and scene structures) are aggregated through a large neighborhood range, realizing a hierarchical expression of point cloud features from local to global.
[0030] S2: Align the dimension of the point cloud features in the first stage with the dimension of the image features; based on the cosine similarity calculation formula, calculate the cross-modal similarities between the aligned point cloud features in the first stage and the image features in the first and second stages respectively. Based on the two cross-modal similarities, enhance the aligned point cloud features in the first stage respectively, splice the image features in the first stage, the image features in the second stage, and the two enhanced point cloud feature maps, and normalize the splicing result through layer normalization to obtain the first fusion feature map.
[0031] Specifically, the process of obtaining the first fusion feature map includes: In the cross-modal attention module (CMAM, Cross - Modal Attention Module), Align the dimension of the point cloud features in the first stage with the dimension of the image features; Convolve the image features in the first and second stages and the aligned point cloud features in the first stage through a convolutional layer containing a 1×1 convolutional kernel; Based on the cosine similarity calculation formula, calculate the cross-modal similarities between the aligned point cloud features in the first stage and the convolved image features in the first and second stages respectively, and obtain the first similarity matrix and the second similarity matrix respectively. The calculation formulas are as follows: ; ; Among them, represents the first similarity matrix; represents the second similarity matrix; represents the first 1×1 convolutional kernel; represents the second 1×1 convolutional kernel; represents the modulus; represents the element in the a th row and b th column of the image features in the first stage; represents the element in the a th row and b th column of the image features in the second stage; represents the element in the k th row and l th column of the aligned point cloud features in the first stage; Based on the first similarity matrix and the point cloud features in the aligned first stage, the first fused point cloud feature map is obtained, and the calculation formula is: ; where, represents the first fused point cloud feature map; represents the softmax activation function; C represents the number of channels of the point cloud features in the first stage; represents the point cloud features in the aligned first stage; Based on the second similarity matrix and the point cloud features in the aligned first stage, the second fused point cloud feature map is obtained, and the calculation formula is: ; where, represents the second fused point cloud feature map; Concatenate the image features in the first stage, the image features in the second stage, the first fused point cloud feature map and the second fused point cloud feature map, and normalize the concatenated feature map through layer normalization to obtain the first fused feature map.
[0032] CMAM is used to fuse the low-level and mid-level features of the RGB image with the low-level features of the point cloud, laying a foundation for subsequent feature extraction and fusion, ensuring the consistency and integrity of basic details, and at the same time introducing some mid-level features to enhance the expression ability.
[0033] S3: One-to-one correspond the image features in the remaining stages with the point cloud features in the remaining stages in order, and respectively perform feature fusion on the three groups of corresponding image features and point cloud features through step S2 to obtain the second fused feature map, the third fused feature map, and the fourth fused feature map.
[0034] Specifically, repeat step S2 based on the image features in the remaining stages and the point cloud features in the remaining stages; Align the dimension of the point cloud features in the second stage with the dimension of the image features; calculate the cross-modal similarity between the image features in the third stage and the point cloud features in the aligned second stage; based on the cross-modal similarity, fuse the point cloud features in the second stage with the image features in the third stage to obtain the second fused feature map.
[0035] CMAM is used to fuse the high-level semantic features of the RGB image with the concentrated abstract features of the point cloud, realizing the complementarity and enhancement of semantic information and improving the semantic richness of the features.
[0036] Align the dimension of the point cloud features in the third stage with the dimension of the image features; calculate the cross-modal similarity between the image features in the fourth stage and the point cloud features in the aligned third stage; based on the cross-modal similarity, fuse the point cloud features in the third stage with the image features in the fourth stage to obtain the third fused feature map.
[0037] Fuse the semantic centralized abstract features of the RGB image and the higher semantic features of the point cloud through CMAM to further enhance the semantic expression ability and spatial consistency of the features.
[0038] Align the dimension of the point cloud features in the fourth stage with the dimension of the image features; calculate the cross-modal similarity between the image features in the fifth stage and the aligned point cloud features in the fourth stage; fuse the point cloud features in the fourth stage and the image features in the fifth stage based on the cross-modal similarity to obtain the fourth fusion feature map.
[0039] Fuse the highest semantic features of the RGB image and the point cloud through CMAM to achieve the complementarity and enhancement of the core semantic information, and provide the most powerful feature support for the final multi-modal application tasks.
[0040] Input the stage features extracted from the RGB image and the point cloud data respectively into CMAM in the specified cascade order, and use the cross-modal similarity modeling, dynamic feature weighting and multi-granularity alignment mechanisms of CMAM to achieve the cross-modal fusion of each level of features, give full play to the advantages of the RGB image and the point cloud data at different levels, enhance the expression ability and semantic richness of the fused features, and lay a foundation for subsequent tasks.
[0041] By inputting the stage features of each level of the RGB image and the point cloud data into CMAM for fusion in the specified cascade order, the advantageous features of both at different levels can be fully utilized. The low-level fusion ensures the consistency of the basic details, the middle-level fusion enhances the expression ability of the features, and the high-level fusion realizes the complementarity and enhancement of the semantic information. This step-by-step fusion method enables the final fused features to contain both rich detail information and powerful semantic expression ability, and can better support subsequent multi-modal application tasks such as object detection, classification and segmentation.
[0042] S4: Pass the four fusion feature maps through the global context cascade structure to obtain the first global context feature, the second global context feature, the third global context feature, and the fourth global context feature respectively.
[0043] Specifically, this step includes: The global context cascade structure is a multi-level feature fusion mechanism designed to integrate multi-scale information through feature interactions at different levels, enhancing the model's understanding and perception of images. This structure consists of multiple GCM modules, each responsible for processing features at a specific level and passing high-level global information to low levels through upsampling and feature fusion operations to enrich the semantic representation of low-level features. Specifically, the global context cascade structure usually includes the fifth global context module, the fourth global context module, the third global context module, and the second global context module, which are connected in sequence from high level to low level to form a top-down feature enhancement path.
[0044] In the fifth global context module, the fourth fused feature map is successively passed through dilated convolution and global average pooling to extract global context information; the global context information is adjusted in terms of channel number and dimension through the second convolutional layer to obtain the first global context feature.
[0045] In the fourth global context module, the first global context feature is upsampled (such as by bilinear interpolation or transposed convolution) to align its resolution with the third fused feature map; the upsampled first global context feature and the third fused feature map are fused through element-wise addition or concatenation operations, and the first fusion result is successively refined through the third convolutional layer and activation function to obtain the second global context feature.
[0046] In the third global context module, the second global context feature is upsampled, and the first global context feature is upsampled again to align the resolutions of the first global context feature and the second global context feature with the second fused feature map; the second fused feature map is fused with the upsampled second global context feature and the again-upsampled first global context feature through concatenation operations, and the second fusion result is successively enhanced through the fourth convolutional layer and activation function to obtain the third global context feature.
[0047] In the second global context module, the third global context feature is upsampled, the second global context feature is upsampled again, and the first global context feature is upsampled for the third time to align the resolutions of the first global context feature, the second global context feature, and the third global context feature with the first fused feature map; the first fused feature map is fused with the upsampled third global context feature, the again-upsampled second global context feature, and the third-upsampled first global context feature through concatenation operations, and the third fusion result is successively refined and enhanced through the fifth convolutional layer and activation function to obtain the fourth global context feature.
[0048] In this embodiment, the above activation functions are all ReLU activation functions.
[0049] The global context cascading structure transfers high-level global context information step by step to low levels through a top-down feature fusion method, and integrates it with low-level detailed features. This interaction and fusion mechanism of multi-scale features has the following functions and advantages: 1. Combination of global and local information: High-level features provide the overall semantics and context information of the image, while low-level features contain rich details and pixel-level information. Through the cascading structure, the model can utilize both types of information simultaneously, improving the understanding and perception ability of complex scenes.
[0050] 2. Feature enhancement and refinement: Each GCM module enhances and refines the features to a certain extent, removes irrelevant or noisy information, and highlights key features. This step-by-step feature optimization process helps improve the robustness and accuracy of the model.
[0051] 3. Multi-scale feature representation: The cascading structure integrates features at different levels, forming a feature representation containing multi-scale information. This is crucial for capturing objects and structures of different sizes and shapes in the image, and can improve the generalization ability and adaptability of the model.
[0052] 4. Efficient information transfer: Through upsampling and feature fusion operations, high-level global information can be effectively transferred to low levels, avoiding information loss during downsampling, while reducing computational complexity and the number of parameters, and improving the efficiency of the model.
[0053] In summary, the global context cascading structure is an effective multi-scale feature fusion mechanism. Through orderly module connection and feature interaction, it provides a more comprehensive and accurate feature representation for the model, thereby enhancing its performance in various computer vision tasks.
[0054] S5: Pass the four global context features through the attention mechanism respectively to obtain the fused and optimized data at four levels.
[0055] Specifically, this step includes: In the dynamic multi-scale dual-channel attention module, For any global context feature, Pass it through global average pooling, adaptive average pooling, and local maximum pooling respectively, and fuse the three pooling results through a concatenation operation to obtain the fused pooling result; Pass the fused pooling result through the first multi-layer perceptron for dimensionality reduction to obtain the dimensionality reduction result; Pass the dimensionality-reduced feature through the ReLU activation function to obtain the first activation result; Pass the first activation result through the second multi-layer perceptron for feature extraction to obtain the first extracted feature; The first extracted feature is passed through the ReLU activation function to obtain the second activation result; The number of channels of the second activation result is restored through the third-layer multi-layer perceptron to obtain the restoration result; The restoration result is passed through the sigmoid activation function to output the channel attention weight; The channel attention weight is multiplied by the first adjustment coefficient to obtain the updated channel attention weight; Based on the updated channel attention weight, channel enhancement is performed on the selected global context features to obtain channel-enhanced features; Deformable convolution is performed on the channel-enhanced features to obtain the convolved features; The convolved features are passed through max pooling and average pooling, and the edge features are detected through the Canny operator. The features after max pooling and average pooling are fused with the detected edge features through a splicing operation to obtain the fused spatial features; The fused spatial features are sequentially passed through the sixth convolutional layer and the activation function for spatial enhancement to obtain the spatially enhanced features; The spatially enhanced features are passed through the self-attention mechanism to obtain the self-attention features; The spatially enhanced features and the self-attention features are fused through a splicing operation to obtain the features combined with the self-attention mechanism; The features combined with the self-attention mechanism are passed through a convolutional operation to output the convolutional features; The convolutional features are upsampled by the bilinear interpolation upsampling method to obtain the fused and optimized data corresponding to the selected global context features; All global context features are traversed, and all global context features are respectively passed through the attention mechanism to obtain the fused and optimized data corresponding to each global context feature.
[0056] The dynamic multi-scale dual-channel attention module further includes: 1. Adjust the attention calculation order; according to the characteristics of the input multi-modal data and the specific task requirements, in some cases, spatial attention calculation is performed first, and then channel attention calculation is performed. This can enable the model to first focus on the key spatial positions in the data, and then enhance the features of these positions in the channel dimension, so as to more accurately capture complex multi-modal information and provide more valuable feature representations for subsequent improvement of data generation.
[0057] 2. Multi-scale attention mechanism; construct a multi-scale feature pyramid and apply a dynamic multi-scale two-channel attention module to all global context features respectively. The dynamic multi-scale two-channel attention module on low-resolution features focuses on global semantic information, and the dynamic multi-scale two-channel attention module on high-resolution features focuses on local detail features; then, the features combined with the self-attention mechanism obtained by passing each global context feature through the dynamic multi-scale two-channel attention module are fused through a splicing operation, enabling the model to comprehensively utilize multi-scale information, better adapt to multi-modal data with different characteristics, and generate more comprehensive fusion-optimized data based on the fused features.
[0058] 3. Reduce the amount of computation and the number of parameters; Low-rank approximation optimization: For the fully connected layer in the channel attention, assume the original weight matrix of the fully connected layer is , is the dimension of the weight of the fully connected layer; through the low-rank approximation method, it is decomposed into the product of two low-dimensional matrices and , where . In this way, the original matrix multiplication operation amount is reduced from to , significantly reducing the number of parameters and the amount of computation with almost no loss of performance.
[0059] Pruning operation: Evaluate the importance of the connection weight by calculating its gradient . Generally, a method based on the absolute value of the gradient can be used, that is, calculate the absolute value of the gradient of each connection weight . Set a threshold , and perform pruning on the connections where . During the pruning process, ensure that the performance loss of the model is within an acceptable range to ensure that the pruned model can still maintain good performance.
[0060] Through the above methods, the dynamic multi-scale two-channel attention module can reduce the consumption of computing resources while ensuring the feature optimization effect, improving the efficiency and scalability of the model.
[0061] The dynamic multi-scale two-channel attention module realizes the fine-grained optimization of features through strategies such as adjusting the attention calculation order, optimizing the channel attention calculation, enhancing the spatial attention calculation, combining the self-attention mechanism, introducing the multi-scale attention mechanism, dynamically adjusting the attention weights, and reducing the amount of computation and the number of parameters.
[0062] The advantages of the fused features are reflected here as: The fused features generated by the cross-modal attention module and the dynamic multi-scale dual-channel attention module significantly enhance the complementarity and discriminability of the features. Specifically: Semantic and geometric alignment: The multi-granularity alignment mechanism (such as high-level semantic alignment, middle-level component alignment, and low-level pixel-point alignment) enables the fused features to retain both the texture details of the RGB image and the three-dimensional spatial structure of the point cloud data, reducing feature ambiguity, thus reducing confusion during model training and accelerating convergence.
[0063] Multi-scale feature enhancement: The dynamic multi-scale dual-channel attention module generates features that can capture key information at different resolution levels through a multi-scale pyramid structure and dynamic weight adjustment, improving the model's adaptability to complex scenes and further reducing the MSE and SSIM loss values.
[0064] Attention-guided optimization: The attention weights (such as channel attention and spatial attention) in the fused features directly guide the loss function to focus on key regions, such as object edges or sparse point cloud regions, thereby optimizing the model's ability to model local details.
[0065] In this embodiment, the first adjustment coefficient is dynamically learned by an auxiliary network according to the input data; the input of the auxiliary network is any global context feature, and the output is the first adjustment coefficient.
[0066] In this embodiment, it further includes: calculating a loss function based on the fused optimized data and its corresponding original expected data, performing backpropagation based on the loss function, and adjusting the network parameters until the loss function converges; the network parameters include the weights and biases of each convolutional layer.
[0067] Furthermore, the process of constructing the loss function includes: Calculating the mean squared error loss based on the fused optimized data and its corresponding original expected data through the mean squared error loss function calculation formula; Calculating the structural similarity loss based on the fused optimized data and its corresponding original expected data through the structural similarity loss function calculation formula; Constructing a loss function based on the mean squared error loss and the structural similarity loss, and the expression of the loss function is: ; where represents the loss function; represents the weight coefficient; represents the mean squared error loss; represents the structural similarity loss.
[0068] This embodiment also includes: establishing a database; and properly saving the fusion optimization data finally generated. Select a suitable storage method according to the data type. For structured data, it can be stored in a relational database, such as MySQL, and a reasonable table structure can be designed to store the data, and necessary indexes can be established to speed up the query speed; for unstructured or semi-structured data, a document database such as MongoDB can be used. At the same time, the data is backed up regularly to ensure the security and integrity of the data, forming a complete database to facilitate future searches and comparisons.
[0069] The lightweight data fusion method based on dynamic multi-scale dual-channel attention provided in this embodiment has the following beneficial effects: In the multimodal feature extraction stage, the RGB image and point cloud data are subjected to phased feature extraction to mine features at different levels; secondly, in terms of cross-modal feature fusion, the cross-modal attention module is innovatively introduced to achieve complementary advantages between RGB images and point cloud data through cross-modal similarity modeling, dynamic feature weighting and multi-granularity alignment mechanism; thirdly, in terms of data fusion optimization, a dynamic multi-scale dual-channel attention module is proposed to achieve refined optimization of features through strategies such as adjusting the order of attention calculation, optimizing channel attention calculation, enhancing spatial attention calculation, combining self-attention mechanism, introducing multi-scale attention mechanism, dynamically adjusting attention weights, and reducing the amount of calculation and number of parameters; finally, in terms of cascade structure design, the cascade structure of the global context module is adopted to transmit high-level global information from top to bottom, enrich the semantic representation of low-level features, and enhance the model's understanding and perception of images. These innovations work together to effectively solve the shortcomings of existing technologies, meet the urgent needs of intelligent security, autonomous driving, multimedia content analysis and other fields for efficient and accurate multimodal information processing, and provide strong support for the intelligent upgrading of the industry.
[0070] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0071] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.
Claims
1. A lightweight data fusion method based on dynamic multi-scale dual-channel attention, characterized in that Including: S1: Obtain image data and point cloud data, and perform phased feature extraction on the point cloud data and the preprocessed image data respectively to obtain point cloud features in four phases and image features in five phases; Based on the features of all voxels divided from the point cloud data, obtain the point cloud features in the first phase; on the basis of the point cloud features in the first phase, successively perform three times of point cloud feature aggregation with gradually decreasing neighborhoods to obtain the point cloud features in the remaining three phases; Successively pass the preprocessed image data through the shallow and deep layers of the ShuffleNetV2 network, an EfficientNet-Lite network, and two MobileNetV2 networks to obtain the image features in each phase respectively; S2: Align the dimensions of the point cloud features in the first phase with the dimensions of the image features; based on the cosine similarity calculation formula, calculate the cross-modal similarities between the aligned point cloud features in the first phase and the image features in the first and second phases respectively, enhance the aligned point cloud features in the first phase based on the two cross-modal similarities, splice the image features in the first phase, the image features in the second phase, and the two enhanced point cloud feature maps, and pass the splicing result through layer normalization to obtain the first fusion feature map; S3: Correspond the image features in the remaining phases and the point cloud features in the remaining phases in sequence one by one, and respectively pass the corresponding three groups of image features and point cloud features through step S2 for feature fusion to obtain the second fusion feature map, the third fusion feature map, and the fourth fusion feature map; S4: Pass the four fusion feature maps through the global context cascade structure to obtain the first global context feature, the second global context feature, the third global context feature, and the fourth global context feature respectively; S5: Pass the four global context features through the attention mechanism respectively to obtain four levels of fusion and optimization data.
2. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, wherein The preprocessing of the image data includes performing adaptive histogram equalization and Gaussian filtering on the acquired image data.
3. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, wherein The phased feature extraction of the preprocessed image data includes: Input the preprocessed image data into the ShuffleNetV2 network, pass through the first convolutional layer and the max pooling layer in the ShuffleNetV2 network, and output the image features in the first phase; Pass the image features in the first phase through the subsequent convolutional layers and the channel shuffle module in the ShuffleNetV2 network, and output the image features in the second phase; Input the image features in the second phase into the EfficientNet-Lite network, and output the image features in the third phase; Input the image features in the third phase into the MobileNetV2 network, and output the image features in the fourth phase; Input the image features in the fourth phase into the MobileNetV2 network, and output the image features in the fifth phase.
4. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, wherein The phased feature extraction of the point cloud data includes: Divide the three-dimensional space of the point cloud data into multiple voxels of the same size; for each voxel, calculate the average coordinates of all points contained therein, and use the average coordinates as the voxel features of the corresponding voxel; Input all voxel features into a convolutional neural network to output the point cloud features in the first stage; For the point cloud features in the first stage, define a set of kernel points in the point cloud data, calculate the distances between each kernel point and the remaining points, and calculate the point cloud features in the second stage based on the distances and the point cloud features in the first stage. The calculation formula is as follows: ; Among them, represents the point cloud feature of the second stage of the i th nuclear point; represents the neighborhood point set of the i th nuclear point; represents the Gaussian kernel function; represents the i th nuclear point and the j th neighbor point thereof; represents the point cloud feature of the first stage of the j th neighbor point; traverse all nuclear points in the point cloud data, and replace the corresponding point cloud feature of the first stage with the point cloud feature of the second stage of each nuclear point to obtain the point cloud feature of the second stage; For the point cloud features in the second stage, randomly sample the points in the point cloud data. For each sampled point, calculate the second distances between each sampled point and its corresponding neighboring points, and calculate the point cloud features in the third stage based on the second distances and the point cloud features in the second stage. The calculation formula is as follows: ; Among them, represents the point cloud feature of the third stage of the g th sampling point; represents the neighborhood point set of the g th sampling point; represents the second distance between the g th sampling point and its h th adjacent point; represents the point cloud feature of the second stage of the h th adjacent point; Traverse all sampling points, and replace the corresponding point cloud features of the second stage with the point cloud features of the third stage of each sampling point to obtain the point cloud features of the third stage; For the point cloud features in the third stage, randomly sample the points in the point cloud data of all sampled points again. For each second sampled point, calculate the third distances between each second sampled point and its corresponding neighboring points, and calculate the point cloud features in the fourth stage based on the third distances and the point cloud features in the third stage. The calculation formula is as follows: ; Among them, represents the point cloud feature of the fourth stage at the u th sampling point; represents the neighborhood point set of the u th sampling point; represents the third distance between the u th sampling point and its v th neighboring point; represents the point cloud feature of the third stage of the v th neighboring point; Traverse all the second sampling points, and replace the corresponding point cloud features of the third stage with the point cloud features of the fourth stage of each second sampling point to obtain the point cloud features of the fourth stage.
5. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, characterized in that, The process of obtaining the first fusion feature map includes: Align the dimensions of the point cloud features in the first stage with the dimensions of the image features; Convolve the image features in the first and second stages and the aligned point cloud features in the first stage through a convolutional layer containing a 1×1 convolutional kernel; Based on the cosine similarity calculation formula, calculate the cross-modal similarities between the aligned point cloud features in the first stage and the convolved image features in the first and second stages respectively, and obtain the first similarity matrix and the second similarity matrix respectively. The calculation formulas are as follows: ; ; Among them, represents the first similarity matrix; represents the second similarity matrix; represents the first 1×1 convolution kernel; represents the second 1×1 convolution kernel; represents the modulus; represents the a -th row and b -th column element in the image features of the first stage; represents the a -th row and b -th column element in the image features of the second stage; represents the k -th row and l -th column element in the point cloud features of the first stage after alignment; Based on the first similarity matrix and the aligned point cloud features in the first stage, obtain the first fused point cloud feature map. The calculation formula is: ; Among them, represents the first fused point cloud feature map; represents the softmax activation function; C represents the number of channels of the point cloud features in the first stage; represents the aligned point cloud features in the first stage; Based on the second similarity matrix and the aligned point cloud features in the first stage, obtain the second fused point cloud feature map. The calculation formula is: ; Among them, represents the second fused point cloud feature map; Concatenate the image features in the first stage, the image features in the second stage, the first fused point cloud feature map, and the second fused point cloud feature map, and normalize the concatenated feature map through layer normalization to obtain the first fusion feature map.
6. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, wherein S4 It includes: The global context cascade structure includes the fifth global context module, the fourth global context module, the third global context module, and the second global context module; In the fifth global context module, pass the fourth fusion feature map through dilated convolution and global average pooling in sequence to extract the global context information; Pass the global context information through the second convolutional layer to adjust its number of channels and dimensions to obtain the first global context feature; In the fourth global context module, upsample the first global context feature to align its resolution with the third fusion feature map; fuse the upsampled first global context feature and the third fusion feature map through element-wise addition or concatenation operations, and refine the first fusion result through the third convolutional layer and the activation function in sequence to obtain the second global context feature; In the third global context module, the second global context feature is upsampled, and the first global context feature is upsampled again to align the resolutions of the first global context feature and the second global context feature with the second fusion feature map; the second fusion feature map is fused with the upsampled second global context feature and the upsampled-again first global context feature through a concatenation operation, and the second fusion result is enhanced successively through a fourth convolutional layer and an activation function to obtain a third global context feature; In the second global context module, the third global context feature is upsampled, the second global context feature is upsampled again, and the first global context feature is upsampled for the third time to align the resolutions of the first global context feature, the second global context feature, and the third global context feature with the first fusion feature map; the first fusion feature map is fused with the upsampled third global context feature, the upsampled-again second global context feature, and the third-times-upsampled first global context feature through a concatenation operation, and the third fusion result is refined and enhanced successively through a fifth convolutional layer and an activation function to obtain a fourth global context feature.
7. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 1, wherein S5 Including: For any global context feature, it is respectively subjected to global average pooling, adaptive average pooling, and local maximum pooling, and the three pooling results are fused through a concatenation operation to obtain a fused pooling result; the fused pooling result is dimension-reduced through a first-layer multi-layer perceptron to obtain a dimension-reduced result; the dimension-reduced feature is passed through a ReLU activation function to obtain a first activation result; the first activation result is subjected to feature extraction through a second-layer multi-layer perceptron to obtain a first extracted feature; the first extracted feature is passed through a ReLU activation function to obtain a second activation result; the second activation result is passed through a third-layer multi-layer perceptron to restore the number of channels to obtain a restored result; the restored result is passed through a sigmoid activation function to output a channel attention weight; the channel attention weight is multiplied by a first adjustment coefficient to obtain an updated channel attention weight; Based on the updated channel attention weight, channel enhancement is performed on the selected global context feature to obtain a channel-enhanced feature; Deformable convolution is performed on the channel-enhanced feature to obtain a convolved feature; The convolved feature is subjected to max pooling and average pooling, and edge features are detected through a Canny operator. The features after max pooling and average pooling are fused with the detected edge features through a concatenation operation to obtain a fused spatial feature; The fused spatial feature is spatially enhanced successively through a sixth convolutional layer and an activation function to obtain a spatially enhanced feature; The spatially enhanced feature is passed through a self-attention mechanism to obtain a self-attention feature; The spatially enhanced feature is fused with the self-attention feature through a concatenation operation to obtain a feature combined with the self-attention mechanism; The feature combined with the self-attention mechanism is subjected to a convolution operation to output a convolutional feature; The convolutional feature is upsampled through a bilinear interpolation upsampling method to obtain a fused and optimized data corresponding to the selected global context feature; Traverse all global context features, pass all global context features through the attention mechanism respectively, and obtain the fusion and optimization data corresponding to each global context feature respectively.
8. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 7, characterized in that, The first adjustment coefficient is dynamically learned by the auxiliary network according to the input data; the input of the auxiliary network is any one of the global context features, and the output is the first adjustment coefficient.
9. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 7, characterized in that Calculate the loss function based on the fusion and optimization data and its corresponding original expected data, perform backpropagation based on the loss function, and adjust the network parameters until the loss function converges; the network parameters include the weights and biases of each convolutional layer.
10. The lightweight data fusion method based on dynamic multi-scale dual-channel attention according to claim 9, wherein, The process of constructing the loss function includes: Based on the fusion and optimization data and its corresponding original expected data, calculate the mean square error loss through the calculation formula of the mean square error loss function; Based on the fusion and optimization data and its corresponding original expected data, calculate the structural similarity loss through the calculation formula of the structural similarity loss function; Construct a loss function based on the mean square error loss and the structural similarity loss, and the expression of the loss function is: ; Among them, represents the loss function; represents the weight coefficient; represents the mean squared error loss; represents the structural similarity loss.
Citation Information
Patent Citations
Three-dimensional target detection method based on multi-modal fusion and deformable attention
CN117975436A
Line channel inflammable tree species identification method based on laser point cloud and image fusion
CN118072177A
Road disease detection method and system based on deep learning and ground-air cooperation
CN119723290A
ECLA-HSFPN-fused lightweight fatigue driving detection method
CN119785327A
Multi-modal fusion three-dimensional target detection method based on bidirectional cross attention
CN119964143A
Cited By
Laser radar and vision fusion target detection method based on space-time registration
CN120491093A
Multi-modal dynamic fusion method based on graph attention network low-rank decomposition
CN121527584A
A multi-modal dynamic fusion method based on graph attention network low-rank decomposition
CN121527584B