An adaptive feature fusion method and apparatus based on optimal transmission

By using an adaptive feature fusion method based on optimal transmission theory, the problem of inaccurate image feature map fusion in existing technologies is solved, and intelligent fusion of local and global features is achieved, thereby improving the accuracy of image feature maps and the performance of complex visual tasks.

CN120510143BActive Publication Date: 2025-10-31厦门工学院
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510992942.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

In existing technologies, image feature map fusion methods cannot adaptively weigh the importance of element features, resulting in low accuracy of the fused image feature map, which in turn affects the performance of complex vision tasks.

Method used

An adaptive feature fusion method based on optimal transmission theory is adopted. By constructing a transmission cost matrix, iteratively optimizing to obtain the optimal transmission scheme matrix, calculating the Wasserstein distance, and mapping it to adaptive fusion gating weights, intelligent fusion of local and global features is achieved.

Benefits of technology

It improves the accuracy of image feature maps and the performance of complex visual tasks, especially significantly improving the performance of object detection under high precision requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510143B_ABST
    Figure CN120510143B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of deep learning and computer vision, and discloses an adaptive feature fusion method and apparatus based on optimal transmission. The method includes: acquiring a first image feature map and a second image feature map to be fused; flattening the two image feature maps in spatial dimensions to obtain a first feature vector set and a second feature vector set; constructing a transmission cost matrix based on the pairing distance between feature vectors in the two feature vector sets; iteratively optimizing the transmission cost matrix until an optimal transmission scheme matrix is ​​obtained; obtaining the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and mapping the Wasserstein distance to adaptive fusion gating weights; and fusing the image feature maps using the adaptive fusion gating weights. This invention, based on optimal transmission theory, provides an adaptive feature fusion method to achieve a principled, distribution-aware, and more intelligent feature fusion mechanism, improving the accuracy of the fused image feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning and computer vision technology, and specifically relates to an adaptive feature fusion method and apparatus based on optimal transmission. Background Technology

[0002] In modern deep neural network architectures, fusing image feature maps from different layers or branches is a common and effective way to improve performance, especially for complex vision tasks. For example, in image object detection tasks, it is often necessary to fuse shallow local feature maps containing fine geometric details and deep global feature maps containing rich semantic information.

[0003] In existing technologies, image feature map fusion methods typically involve element-wise operations on image feature maps. Specifically, different image feature maps can be concatenated or added element-wise to obtain a fused image feature map. While these image fusion methods are simple, they cannot adaptively balance the importance of element features, resulting in low accuracy of the fused image feature map. Consequently, the performance of complex visual tasks performed based on the fused image feature map is poor; for example, it may be unable to accurately detect targets in an image. Summary of the Invention

[0004] The purpose of this invention is to address the technical problem of low accuracy of fused image feature maps in existing technologies. It proposes a new adaptive feature fusion method based on optimal transmission theory to achieve a principled, distribution-aware, and more intelligent feature fusion mechanism, thereby greatly improving the accuracy of the fused image feature maps.

[0005] In a first aspect, embodiments of the present invention provide an adaptive feature fusion method based on optimal transmission, the method comprising:

[0006] Obtain the first image feature map and the second image feature map to be fused;

[0007] Flatten the first image feature map and the second image feature map in the spatial dimension to obtain the first feature vector set and the second feature vector set;

[0008] Based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set, a transmission cost matrix is ​​constructed. The elements in the transmission cost matrix are used to characterize the transmission cost required to move a first feature vector to a second feature vector.

[0009] The transmission cost matrix is ​​iteratively optimized using an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thus obtaining the optimal transmission scheme matrix. The marginal distribution constraint is used to constrain the uniform distribution of the row sum and column sum of the transmission cost matrix.

[0010] The Wasserstein distance is obtained through the transmission cost matrix and the optimal transmission scheme matrix, and then mapped to an adaptive fusion gating weight; the Wasserstein distance is used to characterize the degree of distribution difference between the first image feature map and the second image feature map;

[0011] The first image feature map and the second image feature map are fused using the adaptive fusion gating weights to obtain the adaptively fused image feature map.

[0012] Optionally, constructing the transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set includes:

[0013] For each first target feature vector in the first feature vector set, calculate the Euclidean distance between the first target feature vector and any second target feature vector in the second feature vector set;

[0014] The square of the Euclidean distance between the first target feature vector and the second target feature vector is determined as the transmission cost required to move the first target feature vector to the second target feature vector.

[0015] All calculated transmission costs are identified as elements of the transmission cost matrix.

[0016] Optionally, the step of iteratively optimizing the transmission cost matrix using an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and minimizes the total transmission cost, thereby obtaining the optimal transmission scheme matrix, includes:

[0017] The optimal transmission scheme matrix is ​​iteratively approximated by repeatedly performing row and column normalization operations on the transmission cost matrix using the Sinkhorn-Knopp iterative optimization algorithm to solve the optimal transmission problem.

[0018] Optionally, obtaining the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and mapping the Wasserstein distance to adaptive fusion gating weights, includes:

[0019] Multiply the transmission cost matrix and the optimal transmission scheme matrix to obtain the Wasserstein distance;

[0020] The Wasserstein distance is mapped to the interval (0, 1) by a learnable mapping network and activation function to obtain the final adaptive fusion weights.

[0021] Optionally, the first image feature map is a local feature map generated by a module capable of capturing the geometric details of the image, and the second image feature map is a global feature map generated by a module capable of capturing the long-range dependencies of the image. The data structures of the local feature map and the global feature map are both four-dimensional tensors.

[0022] Optionally, fusing the first image feature map and the second image feature map using the adaptive fusion gating weights to obtain the adaptively fused image feature map includes:

[0023] The adaptive fusion gate weight is used as the first weight corresponding to the local feature map, and the weight value obtained by subtracting the adaptive fusion gate weight is used as the second weight corresponding to the global feature map.

[0024] The local feature map and the global feature map are weighted and summed using the first weight and the second weight to obtain the adaptively fused image feature map.

[0025] Secondly, embodiments of the present invention provide an adaptive feature fusion apparatus based on optimal transmission, the apparatus comprising:

[0026] The image feature map acquisition module is used to acquire the first image feature map and the second image feature map to be fused.

[0027] The image feature map flattening module is used to flatten the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set.

[0028] The transmission cost matrix construction module is used to construct a transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set. The elements in the transmission cost matrix are used to characterize the transmission cost required to move a first feature vector to a second feature vector.

[0029] The optimal transmission scheme matrix determination module is used to iteratively optimize the transmission cost matrix through an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thereby obtaining the optimal transmission scheme matrix. The marginal distribution constraint is used to constrain the uniform distribution of the row sum and column sum of the transmission cost matrix.

[0030] The adaptive fusion gate weight calculation module obtains the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and maps the Wasserstein distance to adaptive fusion gate weights; the Wasserstein distance is used to characterize the degree of distribution difference between the first image feature map and the second image feature map;

[0031] The image feature map fusion module is used to fuse the first image feature map and the second image feature map through the adaptive fusion gating weights to obtain the adaptively fused image feature map.

[0032] Thirdly, embodiments of the present invention provide an electronic device, including:

[0033] At least one processor;

[0034] Memory for storing the at least one processor-executable instruction;

[0035] The at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0036] Fourthly, embodiments of the present invention provide a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method described in the first aspect.

[0037] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0038] The technical solution provided by this invention involves obtaining a first image feature map and a second image feature map to be fused, and flattening the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set; constructing a transmission cost matrix based on the pairing distance between the first feature vector and the second feature vector in the first feature vector set and the second feature vector set; and iteratively optimizing the transmission cost matrix through an iterative optimization algorithm to obtain an optimal transmission scheme matrix that satisfies the marginal distribution constraints while minimizing the total transmission cost.

[0039] Next, the Wasserstein distance is calculated based on the transmission cost matrix and the optimal transmission scheme matrix. Since the Wasserstein distance is a rigorous mathematical measure of the difference between probability distributions, it accurately captures the geometric structure of the entire distribution of image feature maps, contributing to more accurate feature fusion. Finally, the Wasserstein distance is mapped to adaptive fusion gating weights. Because the Wasserstein distance can characterize the real-time distribution differences of the input image feature maps, the feature fusion mechanism of this embodiment has high content adaptability. When multiple image feature maps are highly different and complementary, the fusion strategy is automatically adjusted; conversely, it is adjusted when they are less complementary. This allows the neural network to intelligently determine the weights of multiple image feature maps during feature fusion based on different scenes and objects, thereby improving the accuracy of feature fusion. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the overall technical solution of the present invention;

[0041] Figure 2 A flowchart illustrating an adaptive feature fusion method based on optimal transmission, provided in an embodiment of the present invention;

[0042] Figure 3(a) is a heatmap of the "source-target distribution" of the Sinkhorn-Knopp iterative algorithm after 0 iterations and 1 iteration provided in the embodiment of the present invention; Figure 3(b) is a heatmap of the "source-target distribution" of the Sinkhorn-Knopp iterative algorithm after 2 iterations and 3 iterations provided in the embodiment of the present invention; Figure 3(c) is a histogram of row sums and column sums.

[0043] Figure 4 A schematic diagram of the structure of an adaptive feature fusion device based on optimal transmission provided in an embodiment of the present invention;

[0044] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0045] The present invention will be described in detail below through embodiments.

[0046] In modern deep neural network architectures, fusing image feature maps from different layers or branches is a common and effective way to improve performance, especially for complex vision tasks. For example, in image object detection tasks, it is often necessary to fuse shallow local feature maps containing fine geometric details and deep global feature maps containing rich semantic information.

[0047] In existing technologies, image feature map fusion methods typically involve element-wise operations on image feature maps. Specifically, different image feature maps can be concatenated or added element-wise to obtain a fused image feature map. While these image fusion methods are simple, they cannot adaptively balance the importance of element features, resulting in low accuracy of the fused image feature map. Consequently, the performance of complex visual tasks performed based on the fused image feature map is poor; for example, it may be unable to accurately detect targets in an image.

[0048] Therefore, how to design a feature fusion method that can measure feature differences from a more fundamental level (such as probability distribution) and use this as a basis for adaptive and principled feature fusion is a technical problem that urgently needs to be solved in this field.

[0049] This invention relates to the fields of deep learning and computer vision technology. Specifically, it relates to a technique for fusing multi-source, multi-scale feature information in neural networks, which is particularly suitable for complex visual tasks that require the coordinated processing of local details and global semantics, such as object detection and instance segmentation.

[0050] This invention aims to address the technical problem that existing feature fusion methods cannot measure feature differences at a high-dimensional distribution level, leading to blind fusion strategies and inaccurate feature fusion. By providing a novel adaptive feature fusion method and apparatus based on optimal transport theory, it aims to achieve a principled, distribution-aware, and more intelligent feature fusion mechanism.

[0051] To ensure clarity, the overall technical solution of this invention will be described in detail below. Figure 1 As shown, it may include the following steps:

[0052] Step 1: Receive multi-source image feature maps. Specifically, receive image feature map data from at least two different sources that need to be fused, for example, a local feature map. L and a global feature map G . L and G The data structures are all four-dimensional tensors. A four-dimensional tensor can include: B, C, H and W , B For batch processing, this represents the number of image samples processed at one time. C Indicates the channel dimension; H Indicates the height of the feature map. W Indicates the width of the feature map.

[0053] Step 2, Feature Vectorization. Flatten the local and global feature maps from Step 1 in spatial dimensions to obtain two sets of feature vectors. l and gThis step transforms the two-dimensional image feature data structure into a one-dimensional point set data structure required for subsequent probability distribution measurements. The following example will illustrate this. l and g Let's illustrate with examples.

[0054] Step 3: Calculate the cost matrix. Specifically, this step processes the two sets of eigenvectors from Step 2. l and g .Will l and g Each feature vector in the dataset is considered as a sample point in a high-dimensional space. This is achieved by calculating the values ​​of the two sets. l and g A cost matrix C is constructed using the pairwise distances (e.g., Euclidean distances) between sample points. Each element C of this cost matrix C... ij This represents the feature vector l i Move to the feature vector g j The transmission cost. The cost matrix C, which quantifies the microscopic differences between the two feature map distributions, can be stored in memory.

[0055] Step 4: Solve for the optimal transmission scheme. Specifically, this step deals with the cost matrix C. An iterative optimization algorithm, such as the Sinkhorn-Knopp algorithm, is used to approximate the entropy-regularized optimal transmission problem. This algorithm converges to an optimal transmission scheme matrix P by repeatedly performing row and column normalization operations on the cost matrix C.

[0056] Step 5: Calculate the Wasserstein distance. Specifically, this step deals with the cost matrix C and the optimal transmission scheme matrix P. First, we can calculate the distance using the formula w = ∑P. ij C ij We calculate a scalar value w, which is an approximate Wasserstein distance.

[0057] Step 6: Calculate the adaptive fusion gate weights. Input the Wasserstein distance w into a default learnable mapping function (such as a small MLP network) to map it into a scalar gate signal, which is an adaptive fusion gate weight in the interval (0,1). α .

[0058] Step 7, Feature Weighted Fusion. Specifically, this step processes the adaptive fusion weights. α and the local feature map in step 1 L and global feature map G A weighted summation operation using the following formula yields an output feature map that has undergone adaptive fusion.

[0059]

[0060] Where Fused_feature is the output feature map of adaptive fusion.

[0061] As described above, this invention, by leveraging optimal transport theory, accurately aligns local and global features, overcoming the coarseness of traditional fusion methods, achieving refined fusion of differentiated features, enhancing the model's ability to model complex feature relationships, and providing an efficient solution for tasks requiring multi-scale feature collaboration (such as image understanding and pattern recognition).

[0062] The technical solution proposed in this invention innovatively introduces optimal transmission theory into image feature fusion in neural networks. By calculating the Wasserstein distance to drive an adaptive gating signal, it solves the problem of low accuracy of the fused image feature map in existing technologies. The technical solution provided by this invention can bring at least the following significant beneficial effects:

[0063] 1. Principled, distribution-aware feature fusion is achieved. The core technical feature of this invention lies in using Wasserstein distance as the basis for measuring the similarity between two features. The fundamental reason is that Wasserstein distance is a rigorous mathematical measure of the difference between probability distributions. It can capture the geometric structure of the entire distribution, which is far more profound and accurate than traditional attention mechanisms based on simple statistics (such as the mean) or dot product similarity calculations. This makes the fusion decision more principled, and the fused image feature map is more accurate.

[0064] 2. Enhanced adaptability and intelligence of the fusion process. Since the adaptive fusion gate weight α is a direct function of the real-time distribution differences of the input features, the fusion mechanism of this invention possesses high content adaptability. When local and global features differ significantly and are highly complementary, the fusion strategy automatically adjusts; conversely, it does not. This enables the neural network to intelligently determine whether to prioritize local details or global semantics based on different scenes and objects.

[0065] 3. Enhanced performance of the final visual task, especially under high-precision requirements. A superior fusion mechanism inevitably leads to stronger feature representation, thereby improving the performance of the final task. The principle is that distributed-aware fusion can retain the effective components of multi-source feature maps to the maximum extent, reducing feature loss during the fusion process. Verification shows that simply introducing this invention can improve the mAP50-95 metric of the baseline model by 8.9 percentage points, demonstrating a significant enhancement in its accurate localization capability under high IoU requirements. Here, mAP50 is a core metric for measuring model performance in object detection, representing the average accuracy when the IoU of the predicted bounding box is ≥0.5; mAP95 represents the average accuracy when the IoU of the predicted bounding box is ≥0.95.

[0066] After a detailed description of the overall technical solution of the present invention, the following will describe in detail an adaptive feature fusion method based on optimal transmission provided by an embodiment of the present invention. For example... Figure 2 As shown in the figure, an adaptive feature fusion method based on optimal transmission provided by an embodiment of the present invention may include the following steps:

[0067] S210, Obtain the first image feature map and the second image feature map to be fused.

[0068] Specifically, this step receives multi-source feature maps. In practical applications, at least two types of input feature maps can be received from different parts of the neural network, namely a first image feature map and a second image feature map.

[0069] As one implementation of the present invention, the first image feature map can be a local feature map generated by a module capable of capturing the geometric details of the image, and the second image feature map can be a global feature map generated by a module capable of capturing the long-range dependence of the image. The data structures of both the local feature map and the global feature map are four-dimensional tensors.

[0070] For example, a local feature map can be represented as The global feature map can be represented as .in, B For batch processing, this represents the number of image samples processed at one time. C Indicates the channel dimension; H Indicates the height of the feature map. W Indicates the width of the feature map.

[0071] S220, flatten the first image feature map and the second image feature map in the spatial dimension to obtain the first feature vector set and the second feature vector set.

[0072] Specifically, after obtaining the first image feature map and the second image feature map, the first image feature map is flattened in the spatial dimension to obtain the first feature vector set, and the second image feature map is flattened in the spatial dimension to obtain the second feature vector set.

[0073] Using the first image feature map as the local feature map L The second image feature map is a global feature map. G For example, the first set of feature vectors is a local feature vector set, which can be represented as: The second set of feature vectors is the global set of feature vectors, which can be represented as: ,in, N = H × W.

[0074] S230, construct the transmission cost matrix based on the pairing distance between each first eigenvector in the first eigenvector set and each second eigenvector in the second eigenvector set.

[0075] The elements in the transmission cost matrix represent the transmission cost required to move the first eigenvector to the second eigenvector.

[0076] Specifically, after obtaining the first set of eigenvectors and the second set of eigenvectors, the pairing distance between any first eigenvector in the first set of eigenvectors and any second eigenvector in the second set of eigenvectors is calculated, and all the calculated pairing distances are used to form the transmission cost matrix.

[0077] As one implementation of this invention, S230, constructing a transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set may include the following steps, namely steps a1 to a3:

[0078] Step a1: For each first target feature vector in the first feature vector set, calculate the Euclidean distance between the first target feature vector and any second target feature vector in the second feature vector set.

[0079] Step a2: The square of the Euclidean distance between the first target feature vector and the second target feature vector is determined as the transmission cost required to move the first target feature vector to the second target feature vector.

[0080] Step a3: Determine all the calculated transmission costs as elements of the transmission cost matrix.

[0081] In this embodiment, taking the local feature vector set and the global feature vector set in step S220 as an example, the two vector sets are... N The C-dimensional eigenvectors are considered as probability distributions. N There are 10 sample points. The transmission cost between these sample points is calculated. The square of the Euclidean distance is used as the transmission cost, resulting in a cost matrix C with dimensions 1. B × N × N Specifically, the elements in cost matrix C are: .

[0082]

[0083] in, Indicates the first in the batch b The first sample in the set of local feature vectors i1 eigenvector; Indicates the first in the batch b The first sample in the global feature vector set j 1 eigenvector.

[0084] S240 uses an iterative optimization algorithm to iteratively optimize the transmission cost matrix until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thus obtaining the optimal transmission scheme matrix.

[0085] Among them, the marginal distribution constraint is used to constrain the uniform distribution of the row sum and column sum of the transmission cost matrix.

[0086] In practical applications, directly solving the optimal transmission problem is computationally difficult. This invention employs an iterative optimization algorithm to solve the optimal transmission problem, iteratively optimizing the transmission cost matrix to achieve the optimal transmission scheme matrix that satisfies marginal distribution constraints while minimizing the total transmission cost. Marginal distribution constraints, in probability theory and statistics, refer to the constraints imposed on the marginal distributions (i.e., the probability distributions of individual variables) of a random variable. These constraints are derived from joint distributions and reflect the independent distribution characteristics of the variable when it is unaffected by other variables.

[0087] As one implementation of this invention, the transmission cost matrix is ​​iteratively optimized using an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thus obtaining the optimal transmission scheme matrix. This can include the following steps:

[0088] The Sinkhorn-Knopp iterative optimization algorithm, which solves the optimal transmission problem, iteratively approximates the optimal transmission scheme matrix by repeatedly performing row and column normalization operations on the transmission cost matrix.

[0089] Specifically, an efficient Sinkhorn-Knopp iterative algorithm is employed to approximate the entropy-regularized optimal transport problem. This algorithm aims to find a transport scheme matrix P with dimensions of... B × N × N This allows P to minimize the total transmission cost while satisfying the edge distribution constraint (i.e., both row sums and column sums are uniformly distributed).

[0090] Figures 3(a), 3(b), and 3(c) show the visualization results of the iterative process of the Sinkhorn-Knopp iterative algorithm, presenting the dynamic matching process from initialization to convergence to verify the effectiveness of the scheme in the distribution alignment task. Figures 3(a) and 3(b) are heatmaps of the "source-target distribution" from iteration 0 to iteration 3, using a 5×6 grid to show the matching probability of the source distribution (5 classes, vertical axis) and the target distribution (6 classes, horizontal axis) under different iteration steps. The color bars update with iteration (0.01-0.05 for iteration 0, 0.05-0.4 for iteration 1, and 0.05-0.35 for iterations 2-3), intuitively showing the process of distribution matching from discrete to convergent in the optimization of the Sinkhorn-Knopp iterative algorithm; Figure 3(c) is a histogram of row sums and column sums, statistically analyzing the total row sum of the source distribution (5 classes) and the total column sum of the target distribution (6 classes), verifying that the algorithm satisfies the marginal distribution constraints and ensures the consistency of the source-target distribution. Multiple iterations and marginal statistics clearly demonstrate the convergence and constraint satisfaction of the algorithm, providing key experimental evidence for the optimal transmission feature fusion module in this invention. This highlights the scientific nature and robustness of achieving accurate distribution alignment in feature matching and fusion tasks, demonstrating that the solution of this invention has strong advantages in complex feature processing scenarios.

[0091] S250 obtains the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and maps the Wasserstein distance to adaptive fusion gating weights.

[0092] The Wasserstein distance is used to characterize the degree of difference in the distribution of the first image feature map and the second image feature map.

[0093] As one implementation of this invention, S250, obtaining the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and mapping the Wasserstein distance to adaptive fusion gating weights, may include the following steps, namely steps b1 and b2:

[0094] Step b1: Multiply the transmission cost matrix and the optimal transmission scheme matrix to obtain the Wasserstein distance.

[0095] Step b2: The Wasserstein distance is mapped to the interval (0, 1) using a learnable mapping network and activation function to obtain the final adaptive fusion weights.

[0096] Specifically, to accurately determine the degree of difference in the distribution of the first image feature map and the second image feature map, after obtaining the transmission cost matrix and the optimal transmission scheme matrix, the transmission cost matrix and the optimal transmission scheme matrix can be multiplied to obtain the Wasserstein distance, which can be expressed as: w bAssume the first image feature map is a local image feature map, and the second image feature map is a global image feature map. w b It can be done To represent, where, These are elements in the transmission scheme matrix P. It is an element in cost matrix C; w b It can accurately measure the degree of distribution difference between local image feature maps and global image feature maps.

[0097] Then, the Wasserstein distance is input into a small learnable mapping network f(·) (e.g., a multilayer perceptron) and mapped to the (0,1) interval through a Sigmoid activation function to obtain the final adaptive fusion gating weights. α b It can be expressed by the following formula:

[0098] α b =sigmoid( f ( w b ))

[0099] S260, the first image feature map and the second image feature map are fused by adaptive fusion gating weights to obtain the adaptively fused image feature map.

[0100] Specifically, after obtaining the adaptive fusion gating weights in step S250, the first image feature map and the second image feature map can be weighted and summed using these adaptive fusion gating weights to achieve accurate adaptive fusion of the first image feature map and the second image feature map.

[0101] As one implementation of this invention, S260, fusing the first image feature map and the second image feature map through adaptive fusion gating weights to obtain the adaptively fused image feature map, may include the following two steps, namely step c1 and step c2:

[0102] Step c1: Use the adaptive fusion gate weight as the first weight corresponding to the local feature map, and use the weight value obtained by subtracting the adaptive fusion gate weight as the second weight corresponding to the global feature map.

[0103] Step c2: The local feature map and the global feature map are weighted and summed using the first weight and the second weight to obtain the adaptively fused image feature map.

[0104] Specifically, the adaptive fusion gating weights generated in S250 are used. αb The original local feature map L and global feature map G are weighted and summed to obtain the final output feature map as follows:

[0105] Fused_feature= α b * L +(1- α b )* G

[0106] The output feature map takes into account both local details and global semantics, and can be used by the neural network for subsequent processing.

[0107] To verify the beneficial effects of this invention, we conducted ablation experiments on the YOLOv8n baseline model, using the DOTA v1.0 dataset as the testing platform. The experimental results are as follows:

[0108] 1. Baseline model (not using the technical solution of this invention): mAP50 score is 27.9%, and mAP50-95 score is 15.8%. Here, mAP50 is a core metric for evaluating model performance in object detection, representing the average accuracy when the intersection-union ratio (IU) between the predicted and ground truth boxes is ≥0.5. Similarly, mAP95 represents the average accuracy when the IU between the predicted and ground truth boxes is ≥0.95.

[0109] 2. Baseline model + technical solution of this invention: mAP50 index is improved to 40.6% (an absolute improvement of 12.7 percentage points), and mAP50-95 index is improved to 24.7% (an absolute improvement of 8.9 percentage points).

[0110] The experimental data above demonstrate that the fusion method based on optimal transmission proposed in this invention can achieve more intelligent and efficient information integration at the distribution level, thereby significantly improving the final performance of complex visual tasks.

[0111] The technical solution provided by this invention involves obtaining a first image feature map and a second image feature map to be fused, and flattening the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set; constructing a transmission cost matrix based on the pairing distance between the first feature vector and the second feature vector in the first feature vector set and the second feature vector set; and iteratively optimizing the transmission cost matrix through an iterative optimization algorithm to obtain an optimal transmission scheme matrix that satisfies the marginal distribution constraints while minimizing the total transmission cost.

[0112] Next, the Wasserstein distance is calculated based on the transmission cost matrix and the optimal transmission scheme matrix. Since the Wasserstein distance is a rigorous mathematical measure of the difference between probability distributions, it accurately captures the geometric structure of the entire distribution of image feature maps, contributing to more accurate feature fusion. Finally, the Wasserstein distance is mapped to adaptive fusion gating weights. Because the Wasserstein distance can characterize the real-time distribution differences of the input image feature maps, the feature fusion mechanism of this embodiment has high content adaptability. When multiple image feature maps are highly different and complementary, the fusion strategy is automatically adjusted; conversely, it is adjusted when they are less complementary. This allows the neural network to intelligently determine the weights of multiple image feature maps during feature fusion based on different scenes and objects, thereby improving the accuracy of feature fusion.

[0113] This invention provides an adaptive feature fusion device 40 based on optimal transmission, such as... Figure 4 As shown, the device includes:

[0114] The image feature map acquisition module 410 is used to acquire the first image feature map and the second image feature map to be fused.

[0115] The image feature map flattening module 420 is used to flatten the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set.

[0116] The transmission cost matrix construction module 430 is used to construct a transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set. The elements in the transmission cost matrix are used to characterize the transmission cost required to move the first feature vector to the second feature vector.

[0117] The optimal transmission scheme matrix determination module 440 is used to iteratively optimize the transmission cost matrix through an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thereby obtaining the optimal transmission scheme matrix. The marginal distribution constraint is used to constrain the row sum and column sum of the transmission cost matrix to be evenly distributed.

[0118] The adaptive fusion gate weight calculation module 450 obtains the Wasserstein distance through the transmission cost matrix and the optimal transmission scheme matrix, and maps the Wasserstein distance to adaptive fusion gate weights; the Wasserstein distance is used to characterize the degree of distribution difference between the first image feature map and the second image feature map;

[0119] The image feature map fusion module 460 is used to fuse the first image feature map and the second image feature map through the adaptive fusion gating weight to obtain the adaptively fused image feature map.

[0120] Thirdly, embodiments of the present invention provide an electronic device 500, such as... Figure 5 As shown, it includes:

[0121] At least one processor 501;

[0122] Memory 502 for storing the at least one processor-executable instruction;

[0123] The at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0124] Fourthly, embodiments of the present invention provide a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method described in the first aspect.

[0125] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0126] The technical solution provided by this invention involves obtaining a first image feature map and a second image feature map to be fused, and flattening the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set; constructing a transmission cost matrix based on the pairing distance between the first feature vector and the second feature vector in the first feature vector set and the second feature vector set; and iteratively optimizing the transmission cost matrix through an iterative optimization algorithm to obtain an optimal transmission scheme matrix that satisfies the marginal distribution constraints while minimizing the total transmission cost.

[0127] Next, the Wasserstein distance is calculated based on the transmission cost matrix and the optimal transmission scheme matrix. Since the Wasserstein distance is a rigorous mathematical measure of the difference between probability distributions, it accurately captures the geometric structure of the entire distribution of image feature maps, contributing to more accurate feature fusion. Finally, the Wasserstein distance is mapped to adaptive fusion gating weights. Because the Wasserstein distance can characterize the real-time distribution differences of the input image feature maps, the feature fusion mechanism of this embodiment has high content adaptability. When multiple image feature maps are highly different and complementary, the fusion strategy is automatically adjusted; conversely, it is adjusted when they are less complementary. This allows the neural network to intelligently determine the weights of multiple image feature maps during feature fusion based on different scenes and objects, thereby improving the accuracy of feature fusion.

[0128] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.

Claims

1. An adaptive feature fusion method based on optimal transmission, characterized in that, The method includes: Obtain a first image feature map and a second image feature map to be fused; the first image feature map is a local feature map generated by a module capable of capturing image geometric details, and the second image feature map is a global feature map generated by a module capable of capturing long-range image dependencies; the data structures of the local feature map and the global feature map are both four-dimensional tensors. Flatten the first image feature map and the second image feature map in the spatial dimension to obtain the first feature vector set and the second feature vector set; Based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set, a transmission cost matrix is ​​constructed. The elements in the transmission cost matrix are used to characterize the transmission cost required to move a first feature vector to a second feature vector. The transmission cost matrix is ​​iteratively optimized using an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thus obtaining the optimal transmission scheme matrix. The marginal distribution constraint is used to constrain the uniform distribution of the row sum and column sum of the transmission cost matrix. The Wasserstein distance is obtained by multiplying the transmission cost matrix and the optimal transmission scheme matrix; the Wasserstein distance is then mapped to the interval (0, 1) using a learnable mapping network and activation function to obtain the final adaptive fusion weights; the Wasserstein distance is used to characterize the degree of distribution difference between the first image feature map and the second image feature map; The first image feature map and the second image feature map are fused using the adaptive fusion gating weights to obtain the adaptively fused image feature map.

2. The method according to claim 1, characterized in that, The step of constructing a transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set includes: For each first target feature vector in the first feature vector set, calculate the Euclidean distance between the first target feature vector and any second target feature vector in the second feature vector set; The square of the Euclidean distance between the first target feature vector and the second target feature vector is determined as the transmission cost required to move the first target feature vector to the second target feature vector. All calculated transmission costs are identified as elements of the transmission cost matrix.

3. The method according to claim 1, characterized in that, The step of iteratively optimizing the transmission cost matrix using an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and minimizes the total transmission cost, thereby obtaining the optimal transmission scheme matrix, includes: The optimal transmission scheme matrix is ​​iteratively approximated by repeatedly performing row and column normalization operations on the transmission cost matrix using the Sinkhorn-Knopp iterative optimization algorithm to solve the optimal transmission problem.

4. The method according to claim 1, characterized in that, The step of fusing the first image feature map and the second image feature map using the adaptive fusion gating weights to obtain the adaptively fused image feature map includes: The adaptive fusion gate weight is used as the first weight corresponding to the local feature map, and the weight value obtained by subtracting the adaptive fusion gate weight is used as the second weight corresponding to the global feature map. The local feature map and the global feature map are weighted and summed using the first weight and the second weight to obtain the adaptively fused image feature map.

5. An adaptive feature fusion device based on optimal transmission, characterized in that, The device includes: The image feature map acquisition module is used to acquire a first image feature map and a second image feature map to be fused; the first image feature map is a local feature map generated by a module capable of capturing image geometric details, and the second image feature map is a global feature map generated by a module capable of capturing long-range dependencies of the image; the data structure of the local feature map and the global feature map is a four-dimensional tensor. The image feature map flattening module is used to flatten the first image feature map and the second image feature map in the spatial dimension to obtain a first feature vector set and a second feature vector set. The transmission cost matrix construction module is used to construct a transmission cost matrix based on the pairing distance between each first feature vector in the first feature vector set and each second feature vector in the second feature vector set. The elements in the transmission cost matrix are used to characterize the transmission cost required to move a first feature vector to a second feature vector. The optimal transmission scheme matrix determination module is used to iteratively optimize the transmission cost matrix through an iterative optimization algorithm until the transmission cost matrix satisfies the marginal distribution constraint and the total transmission cost is minimized, thereby obtaining the optimal transmission scheme matrix. The marginal distribution constraint is used to constrain the uniform distribution of the row sum and column sum of the transmission cost matrix. An adaptive fusion gated weight calculation module is used to multiply the transmission cost matrix and the optimal transmission scheme matrix to obtain the Wasserstein distance; the Wasserstein distance is mapped to the interval (0, 1) through a learnable mapping network and activation function to obtain the final adaptive fusion weight; the Wasserstein distance is used to characterize the degree of distribution difference between the first image feature map and the second image feature map; The image feature map fusion module is used to fuse the first image feature map and the second image feature map through the adaptive fusion gating weights to obtain the adaptively fused image feature map.

6. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method as described in any one of claims 1-4.

8. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Feature fusion method based on gating mechanism and target detection network

    CN113537396A

  • Similar node dynamic perception traffic flow prediction method with space-time heterogeneous attention

    CN118430229A

  • High-compression-ratio image compression method based on optimal transmission mapping

    CN119110084A

  • Global structure sensing vector quantization method based on optimal transmission

    CN120123628A