Heterogeneous backbone network adaptive fusion method
Through the adaptive fusion method of heterogeneous backbone network, the feature maps of Transformer and CNN networks are weighted and summed, which solves the problem of insufficient feature extraction capabilities of the excavator target detection model, realizes dynamic fusion of local and global features, and improves the detection effect.
Patent Information
- Application Number
- CN202510382161.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the backbone network feature extraction capability of the excavator target detection model is affected, resulting in poor feature extraction capability and poor effect.
The heterogeneous backbone network adaptive fusion method is adopted, and the feature map adaptive fusion of the Transformer network structure (Swin-T backbone network) and the CNN network structure (ConvNeXt-T backbone network) is adaptively integrated, and the fusion module is used to perform weighted summing of the feature maps to achieve dynamic fusion of local and global features.
It enhances the feature expression ability of the backbone network, provides a good foundation for subsequent object detection, and improves detection efficiency and accuracy.
Smart Images

Figure CN120339765A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of excavators, and specifically relates to a method for adaptive fusion of heterogeneous backbone networks. Background Art
[0002] When an excavator undergoes factory quality inspection, it is necessary to detect whether there are any missing or misinstalled components in the vehicle configuration items, such as whether the labels on the appearance are missing, whether the rearview mirrors, windshield wipers, working lights, etc. are missing or misinstalled, as Figure 1 shown. To improve the detection efficiency and reduce false or missed detections, in an intelligent quality inspection system for excavators based on computer vision technology, industrial cameras or dome cameras are used to take pictures of the excavator, and then a target detection model is used for automatic detection and recognition of configuration items. Since there are significant differences in the size, color, texture, shape, and other characteristics of various configuration items of the excavator, this poses high requirements for the feature extraction ability of the backbone network of the target detection model.
[0003] In the prior art, due to the relatively wide range of application scenarios of the excavator, the feature extraction ability of the backbone network of the target detection model will be affected, resulting in poor feature extraction ability and unsatisfactory results. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for adaptive fusion of heterogeneous backbone networks to solve the technical problem that the feature extraction ability of the backbone network of the target detection model is affected, resulting in poor feature extraction ability and unsatisfactory results, and to achieve the purpose of improving the local and global feature extraction capabilities of the backbone network through the adaptive fusion of Transformer and CNN feature maps.
[0005] To solve the above technical problems, the present invention provides a method for adaptive fusion of heterogeneous backbone networks, including: A Transformer network structure and a CNN network structure; The Transformer network structure adopts a Swin-T backbone network, and the CNN network structure adopts a ConvNeXt-T backbone network; Among them, the Transformer network structure includes T-stem, Stage1, Stage2, Stage3, and Stage4; The CNN network structure includes C-stem, res2, res3, res4, and res5; The output images of the Transformer network structure and the CNN network structure have the same size; It further includes a fusion module for fusing the Transformer network structure and the CNN network structure.
[0006] Further, the steps are as follows: Step 1: Input the feature maps of Stage1 and res2 into the fusion module, where represents the feature map of Swin-T, that is, the Stage1 feature map, represents the feature map of ConvNeXt-T, that is, the res2 feature map; Step 2: and are subjected to feature mapping and enhancement through a convolutional layer with a size of 3*3, a stride of 1, and a padding of 1; Step 3: Perform feature dimensionality reduction on and through the global pooling layer, and output the feature vector; Step 4: Then, through two fully connected layers (FC), batch normalization (BN), and ReLU activation function processing, two feature vectors a and b are obtained respectively. The number of nodes in the first fully connected layer (FC) is , and the number of nodes in the second fully connected layer (FC) is , is the scaling factor, and its value is 8; Step 5: The number of elements in both feature vectors a and b is , which respectively correspond to the mapping results of each channel of the input feature maps of Swin-T and ConvNeXt-T. The relative magnitudes of their values reflect the relative importance of the features. Therefore, perform softmax on the two feature vectors a and b bit by bit and update their values, that is , , ; In Step 5, the updated a and b represent the relative weights of each channel of the Swin-T and ConvNeXt-T feature maps. Multiply a and element by element, that is, multiply and the th channel feature map of each element to obtain . Similarly, multiply b and element by element to obtain , and then add the two feature maps to obtain .
[0007] Further, in Step 1, the sizes of the feature maps of Stage1 and res2 are both , respectively representing the height, width, and number of channels of the feature map.
[0008] Further, the T-stem and C-stem are composed of convolutional layers; The convolution kernel size of the convolutional layer is 4*4, and the stride of the convolutional layer is 4.
[0009] Further, the size of the output images of the Transformer network structure and the CNN network structure is H*W*3, where the height is H, the width is W, and the number of channels is 3.
[0010] Further, the feature map sizes of Stage1 and res2 are H / 4*W / 4*C, the feature map sizes of Stage2 and res3 are H / 8*W / 8*C, the feature map sizes of Stage3 and res4 are H / 16*W / 16*C, and the feature map sizes of Stage4 and res5 are H / 32*W / 32*C.
[0011] The beneficial effects of the present invention are: 1. It can adaptively fuse the feature maps of two heterogeneous networks, namely Swin-T and ConvNeXt-T, to achieve dynamic fusion of local features and global features, enhance the feature expression ability of the backbone network, and provide a good foundation for subsequent object detection.
[0012] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. Description of the Drawings
[0013] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0014] Figure 1 It is a schematic flowchart of the method for adaptive fusion of the heterogeneous backbone network of the present invention; Figure 2 It is a schematic structural diagram of an excavator of the method for adaptive fusion of the heterogeneous backbone network of the present invention. Specific Embodiments
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0016] Embodiment: As Figures 1 to 2 shown, a heterogeneous backbone network adaptive fusion method includes: a Transformer network structure and a CNN network structure; wherein, the Transformer network structure adopts a Swin-T backbone network, and the CNN network structure adopts a ConvNeXt-T backbone network.
[0017] Among them, the Transformer network structure includes T-stem, Stage1, Stage2, Stage3, and Stage4, and the CNN network structure includes C-stem, res2, res3, res4, and res5; in this embodiment, T-stem and C-stem are composed of convolutional layers; the convolutional kernel size of the convolutional layer is 4*4, and the stride of the convolutional layer is 4.
[0018] The output images of the Transformer network structure and the CNN network structure have the same size; in this embodiment, the size of the output images of the Transformer network structure and the CNN network structure is H*W*3, where the height is H, the width is W, and the number of channels is 3. Among them, the feature map sizes of Stage1 and res2 are H / 4*W / 4*C, the feature map sizes of Stage2 and res3 are H / 8*W / 8*C, the feature map sizes of Stage3 and res4 are H / 16*W / 16*C, and the feature map sizes of Stage4 and res5 are H / 32*W / 32*C.
[0019] As Figure 1 shown, it further includes a fusion module, and the fusion module fuses the Transformer network structure and the CNN network structure.
[0020] The steps of the adaptive fusion method of this application are as follows: Step 1: Input the feature maps of Stage1 and res2 into the fusion module, where represents the feature map of Swin-T, that is, the Stage1 feature map, represents the feature map of ConvNeXt-T, that is, the res2 feature map; In Step 1, the feature map sizes of Stage1 and res2 are both , representing the height, width, and number of channels of the feature map respectively.
[0021] Step 2: and go through a convolutional layer with a size of 3*3, a stride of 1, and a padding of 1 for feature mapping and enhancement; Step 3: For and Feature dimensionality reduction is performed through the Global Pooling layer, and the output feature vector; Step 4: After passing through two fully connected layers (FC), batch normalization (BN), and ReLU activation function processing, two feature vectors a and b are obtained respectively. The number of nodes in the first fully connected layer (FC) is , and the number of nodes in the second fully connected layer (FC) is , is the scaling factor, and its value is 8; Step 5: The number of elements in both feature vectors a and b is , which respectively correspond to the mapping results of each channel of the input feature maps of Swin-T and ConvNeXt-T. The relative magnitudes of their values reflect the relative importance of the features. Therefore, softmax is performed on the two feature vectors a and b bit by bit, and their values are updated, that is , , ; In Step 5, the updated a and b represent the relative weights of each channel of the Swin-T and ConvNeXt-T feature maps. Multiply a and element by element, that is, multiply and each element of the th channel feature map to obtain . Similarly, multiply b and element by element to obtain , and then add the two feature maps to obtain .
[0022] The feature fusion module essentially uses the attention mechanism to calculate the relative importance of the corresponding channels of the two input feature maps, and uses this as the weight to perform weighted summation on the two input feature maps to obtain the fused feature map. The feature maps of the four intermediate layers of Swin-T and ConvNeXt-T all perform the above fusion operation to obtain the fused feature map, and then use this feature map for subsequent object detection. For example, in Faster-RCNN, a feature pyramid can be constructed based on the Feature Pyramid Network (FPN) using the fused feature map, and then sent to the detection head to output the detection results of each object.
[0023] In summary: It can adaptively fuse the feature maps of two heterogeneous networks, Swin-T and ConvNeXt-T, realize the dynamic fusion of local features and global features, enhance the feature expression ability of the backbone network, and provide a good foundation for subsequent object detection.
[0024] All the components selected in this application are common standard components or components known to those skilled in the art, and their structures and principles can be learned from technical manuals or obtained through conventional experimental methods by those skilled in the art.
[0025] In the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0026] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0027] Based on the above inspiration from the ideal embodiments of the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. An adaptive fusion method for heterogeneous backbone networks, characterized in that, Including: Transformer network structure and CNN network structure; The Transformer network structure adopts a Swin-T backbone network, and the CNN network structure adopts a ConvNeXt-T backbone network; Among them, the Transformer network structure includes T-stem, Stage1, Stage2, Stage3, and Stage4; The CNN network structure includes C-stem, res2, res3, res4, and res5; The output images of the Transformer network structure and the CNN network structure have the same size; It further includes a fusion module, and the fusion module is used to fuse the Transformer network structure and the CNN network structure.
2. The heterogeneous backbone network adaptive fusion method according to claim 1, characterized in that The steps are as follows: Step 1: Input the feature maps of Stage1 and res2 into the fusion module, where represents the feature map of Swin-T, that is, the Stage1 feature map, represents the feature map of ConvNeXt-T, that is, the res2 feature map; Step 2: and perform feature mapping and enhancement through a convolutional layer with a size of 3*3, a stride of 1, and a padding of 1; Step 3: Perform feature dimensionality reduction on and through the global pooling layer, and output the feature vector; Step 4: After processing through two fully connected layers (FC), batch normalization (BN), and ReLU activation function, two feature vectors a and b are obtained respectively. The number of nodes in the first fully connected layer (FC) is , and the number of nodes in the second fully connected layer (FC) is , is the scaling factor, and its value is 8; Step Five: The number of elements in both eigenvectors a and b is , which respectively correspond to the mapping results of each channel of the input feature maps of Swin-T and ConvNeXt-T. The relative magnitudes of their values reflect the relative importance of the features. Therefore, perform softmax on the two eigenvectors a and b bit by bit and update their values, that is , , ; In step five, the updated a and b represent the relative weights of each channel of the Swin-T and ConvNeXt-T feature maps. Multiply a and element-wise, that is, multiply and each element of the feature map of the th channel to obtain . Similarly, multiply b and element-wise to obtain , and then add the two feature maps to obtain .
3. The heterogeneous backbone network adaptive fusion method according to claim 2, characterized in that In Step 1, the feature map sizes of Stage1 and res2 are both , representing the height, width, and number of channels of the feature map, respectively.
4. The heterogeneous backbone network adaptive fusion method according to claim 3, characterized in that The T-stem and C-stem are composed of convolutional layers; The convolutional kernel size of the convolutional layer is 4*4, and the stride of the convolutional layer is 4.
5. The heterogeneous backbone network adaptive fusion method according to claim 4, characterized in that The size of the output images of the Transformer network structure and the CNN network structure is H*W*3, where the height is H, the width is W, and the number of channels is 3.
6. The heterogeneous backbone network adaptive fusion method according to claim 5, characterized in that Among them, The feature map sizes of Stage1 and res2 are H / 4*W / 4*C, the feature map sizes of Stage2 and res3 are H / 8*W / 8*C, the feature map sizes of Stage3 and res4 are H / 16*W / 16*C, and the feature map sizes of Stage4 and res5 are H / 32*W / 32*C.