An object detection method and system based on angle adaptive fusion

By introducing an angle adaptive fusion pyramid network into the object detection model, extracting and fusing multi-angle feature information, the problem of insufficient robustness in the existing model when dealing with multi-angle distribution of objects is solved, and a more efficient object detection effect is achieved.

CN114743015BActive Publication Date: 2025-05-27SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210343569.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-02
Publication Date
2025-05-27
Estimated Expiration
2042-04-02

AI Technical Summary

Technical Problem

The existing object detection model lacks robustness when processing multi-angle distribution of objects, especially the limited ability of convolutional neural networks to handle rotation invariance, resulting in a decrease in detection effect.

Method used

A target detection method based on angle adaptive fusion is proposed. Through the combination of the base network, feature pyramid network and angle adaptive fusion pyramid network, multi-angle feature information is extracted and fused to generate a feature map that is robust to the multi-angle changes of the target.

Benefits of technology

The robustness of the object detection model to multi-angle distribution of objects is improved, the detection effect is enhanced, and the calculation efficiency is high, the application range is wide, and it is not limited by data set expansion and feature map size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114743015B_ABST
    Figure CN114743015B_ABST
Patent Text Reader

Abstract

The present invention discloses a target detection method and system based on angle adaptive fusion. The method includes: extracting features from an input image through a backbone network, and outputting corresponding first feature maps from the output results of convolutional layers at different stages; performing lateral connection dimensionality reduction and top-down fusion processing on the original feature maps through a feature pyramid network to obtain second feature maps; performing target multi-angle feature extraction and fusion processing on the second feature maps through an angle adaptive fusion pyramid network to obtain third feature maps; classifying and regressing the third feature maps through a detector head network to output target detection results. The present invention has a wide range of applications and high computational efficiency, and can be widely applied to the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular to an object detection method and system based on angle adaptive fusion. Background Art

[0002] With the continuous development of society, object detection has many application scenarios in our daily life, improving our daily production and lifestyle. As one of the basic problems in the field of computer vision, the problem to be solved by object detection is to judge whether there are instances in the predefined categories in the input image. If so, the category and spatial position information of the instance need to be returned. That is to say, object detection includes two subtasks: classification and regression. With the substantial improvement of computer computing power, object detection has also changed from the traditional image object detection stage to the deep learning-based object detection stage, and the performance of object detection is constantly being refreshed and improved.

[0003] Although the performance of object detectors is constantly improving, there are still some problems to be solved, such as the problem of diverse object directions. In daily life, objects do not always maintain a single pose and orientation, resulting in the objects to be detected being distributed at arbitrary angles in the image. We know that convolutional neural networks play a very good role in image feature extraction, and because convolutional layers have the characteristic of parameter sharing, convolutional neural networks have good translational invariance to the input. However, the usual convolution kernels and convolution operations are not designed for the angle characteristics of objects, so they do not have natural rotational invariance. Although subsequent pooling operations can make convolutional neural networks have a certain rotational invariance to small-angle changes of objects, overall, the processing ability of convolutional neural networks for important local and global image rotations is still limited. In order to improve the robustness of object detection algorithms to object rotation, existing work has made attempts in this regard, and the main improvement directions are three: one is to randomly rotate the input image to expand the dataset, which is one of the commonly used data augmentation methods; one is to rotate the feature map or convolution kernel at multiple angles, extract the feature information of the object in multiple directions and fuse it to obtain a model that is robust to the object direction, such as the RFN network and the ORN network; the other is to rotate the regression anchor box to obtain a rotated object representation, such as SCRDet and OBB. Exploring a better way to represent objects in multiple directions is also one of the ways to improve the object detection effect.

[0004] For the problem of diverse object directions in object detection, the main idea of the solution is how to make the model successfully extract features that are less sensitive to object directions. Although some related solutions have achieved certain effects at present, they also have corresponding defects.

[0005] 1. The method of directly performing random rotation on the input image dataset directly expands the dataset, doubling the computational amount of the model. If the training is insufficient, it may also lead to underfitting of the model.

[0006] 2. The method of rotating the feature map or convolution kernel usually has restrictions on the size and rotation angle of the feature map. Sometimes, interpolation operations need to be introduced to assist in rotation, which also increases the computational amount.

[0007] 3. The method of rotating the regression anchor box is only applicable to the anchor box-based object detection model and cannot be applied to all types of object detection models. Summary of the Invention

[0008] In view of this, embodiments of the present invention provide an object detection method and system based on angle adaptive fusion, which have a wide application range and high computational efficiency.

[0009] One aspect of the present invention provides an object detection method based on angle adaptive fusion, including:

[0010] Performing feature extraction on the input image through a backbone network, and outputting corresponding first feature maps from the output results of convolutional layers at different stages;

[0011] Performing horizontal connection dimensionality reduction and top-down fusion processing on the original feature map through a feature pyramid network to obtain a second feature map;

[0012] Performing target multi-angle feature extraction and fusion processing on the second feature map through an angle adaptive fusion pyramid network to obtain a third feature map;

[0013] Performing classification and regression on the third feature map through a detector head network, and outputting an object detection result.

[0014] Optionally, the performing target multi-angle feature extraction and fusion processing on the second feature map through an angle adaptive fusion pyramid network to obtain a third feature map includes:

[0015] Processing the input second feature map through a multi-angle feature extraction module to obtain multiple fourth feature maps with different angle direction information and retaining shallow details;

[0016] Performing weighted fusion on the fourth feature maps through a multi-feature fusion module to obtain a final third feature map with angle adaptive fusion.

[0017] Optionally, the processing the input second feature map through a multi-angle feature extraction module to obtain multiple fourth feature maps with different angle direction information and retaining shallow details includes:

[0018] Perform rotational convolution processing on the second feature map using a 3×3 convolutional kernel, rotate the positions corresponding to the weight parameters counterclockwise by the rotation angle, and output the fifth feature map;

[0019] Perform bottom-up connection fusion on the fifth feature map to obtain the fourth feature map.

[0020] Optionally, the step of performing rotational convolution processing on the second feature map using a 3×3 convolutional kernel, rotating the positions corresponding to the weight parameters counterclockwise by the rotation angle, and outputting the fifth feature map is specifically:

[0021] Configure four offset convolutional branches with angle offsets of 0°, 90°, 180°, and 270°, use the four offset convolutional branches to extract feature information in different directions of the input feature map respectively, and output the fifth feature map.

[0022] Optionally, the step of configuring four offset convolutional branches with angle offsets of 0°, 90°, 180°, and 270°, using the four offset convolutional branches to extract feature information in different directions of the input feature map respectively, and outputting the fifth feature map includes:

[0023] Take a convolutional branch with a kernel size of 3×3 as the reference branch, and use the reference branch as the offset convolutional branch with an angle offset of 0°;

[0024] Rotate the convolutional kernels of the reference branch by 90°, 180°, and 270° respectively to obtain three offset convolutional branches with specific angle offsets. Among them, the weights of the four offset convolutional branches are shared with each other;

[0025] Output the fifth feature map through each convolutional branch.

[0026] Optionally, the step of performing bottom-up connection fusion on the fifth feature map to obtain the fourth feature map includes:

[0027] Construct a bottom-up connection fusion path between the results of the same angle offset convolution of adjacent feature maps;

[0028] For the feature map of the lowest layer, directly output it as the fourth feature map;

[0029] For non-lowest layer feature maps, first downsample the input low-layer feature map to the same size as the input high-layer feature map, splice the downsampled feature map and the input high-layer feature map along the channel dimension to obtain the spliced feature map, and use a 1×1 convolutional layer to reduce the channels of the spliced feature map, so that the fused feature map is output as the fourth feature map;

[0030] Among them, the downsampling is a 3×3 convolution operation.

[0031] Optionally, the weighted fusion of the fourth feature map by the multi-feature fusion module to obtain the third feature map with final angle adaptive fusion includes:

[0032] Obtain three feature map sets from the fourth feature map, where each feature map set contains four feature maps of the same size but with different target angle direction information; and perform the following operations on each feature map set:

[0033] Add the different feature maps pixel by pixel to obtain the fusion result of the feature maps;

[0034] Perform a global average pooling operation on the fused feature map to generate a first feature vector of size 1×1×C; where each first feature vector represents the global information on each channel;

[0035] Generate a second feature vector by passing the first feature vector through a fully connected layer;

[0036] Calculate the channel weighting values of different input feature maps according to the second feature vector;

[0037] Perform adaptive weighted fusion on the feature maps in the feature map set according to the channel weighting values, and output the third feature map.

[0038] Another aspect of the embodiments of the present invention further provides an object detection system based on angle adaptive fusion, including:

[0039] A first module for extracting features from an input image through a backbone network and outputting corresponding first feature maps from the output results of convolutional layers at different stages;

[0040] A second module for performing lateral connection dimensionality reduction and top-down fusion processing on the original feature map through a feature pyramid network to obtain a second feature map;

[0041] A third module for performing target multi-angle feature extraction and fusion processing on the second feature map through an angle adaptive fusion pyramid network to obtain a third feature map;

[0042] A fourth module for classifying and regressing the third feature map through a detector head network and outputting the object detection result.

[0043] Another aspect of the embodiments of the present invention further provides an electronic device, including a processor and a memory;

[0044] The memory is used to store a program;

[0045] The processor executes the program to implement the method as described above.

[0046] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing a program, which when executed by a processor implements the method described above.

[0047] An embodiment of the present invention also discloses a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the computer device executes the method described above.

[0048] In an embodiment of the present invention, feature extraction is performed on an input image based on a backbone network, and a corresponding first feature map is output from the output results of convolutional layers at different stages; horizontal connection dimensionality reduction and top-down fusion processing are performed on the original feature map through a feature pyramid network to obtain a second feature map; through an angle-adaptive fusion pyramid network, multi-angle feature extraction and fusion processing of the target are performed on the second feature map to obtain a third feature map; classification and regression are performed on the third feature map through a detector head network to output a target detection result. The present invention has a wide application range and high computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0050] Figure 1 It is a schematic structural diagram of a target detection network based on angle-adaptive fusion provided by an embodiment of the present invention;

[0051] Figure 2 It is a schematic diagram of a 3×3 convolution kernel rotated by 8 angles provided by an embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of a multi-angle feature extraction module provided by an embodiment of the present invention;

[0053] Figure 4 It is a schematic diagram of a multi-feature adaptive fusion module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] In view of the problems existing in the prior art, an embodiment of the present invention provides an object detection method based on angle adaptive fusion, including:

[0056] Performing feature extraction on the input image through a backbone network, and outputting corresponding first feature maps from the output results of convolutional layers at different stages;

[0057] Performing horizontal connection dimensionality reduction and top-down fusion processing on the original feature maps through a feature pyramid network to obtain second feature maps;

[0058] Performing object multi-angle feature extraction and fusion processing on the second feature maps through an angle adaptive fusion pyramid network to obtain third feature maps;

[0059] Performing classification and regression on the third feature maps through a detector head network, and outputting object detection results.

[0060] Optionally, the performing object multi-angle feature extraction and fusion processing on the second feature maps through an angle adaptive fusion pyramid network to obtain third feature maps includes:

[0061] Processing the input second feature maps through a multi-angle feature extraction module to obtain multiple fourth feature maps with different angle direction information and retaining shallow details;

[0062] Performing weighted fusion on the fourth feature maps through a multi-feature fusion module to obtain the finally angle-adaptively fused third feature maps.

[0063] Optionally, the processing the input second feature maps through a multi-angle feature extraction module to obtain multiple fourth feature maps with different angle direction information and retaining shallow details includes:

[0064] Performing rotational convolution processing on the second feature maps using a 3×3 convolutional kernel, rotating the positions corresponding to the weight parameters counterclockwise by the rotation angle, and outputting fifth feature maps;

[0065] Performing bottom-up connection fusion on the fifth feature maps to obtain fourth feature maps.

[0066] Optionally, the performing rotational convolution processing on the second feature maps using a 3×3 convolutional kernel, rotating the positions corresponding to the weight parameters counterclockwise by the rotation angle, and outputting fifth feature maps specifically is:

[0067] Configuring four offset convolutional branches with angle offsets of 0°, 90°, 180°, and 270°, and respectively extracting feature information in different directions of the input feature maps using the four offset convolutional branches, and outputting fifth feature maps.

[0068] Optionally, the configuration includes four offset convolution branches with angle offsets of 0°, 90°, 180°, and 270°. The four offset convolution branches are used to extract feature information in different directions of the input feature map and output a fifth feature map, including:

[0069] Taking a convolution branch with a convolution kernel size of 3×3 as the reference branch, and using the reference branch as the offset convolution branch with an angle offset of 0°;

[0070] Rotating the convolution kernel of the reference branch by 90°, 180°, and 270° respectively to obtain three offset convolution branches with specific angle offsets. Among them, the weights of the four offset convolution branches are shared with each other;

[0071] Outputting the fifth feature map through each convolution branch.

[0072] Optionally, the bottom-up connection fusion of the fifth feature map to obtain a fourth feature map includes:

[0073] Constructing a bottom-up connection fusion path between the results of the same-angle offset convolution of adjacent feature maps;

[0074] For the feature map of the lowest layer, directly output it as the fourth feature map;

[0075] For non-lowest layer feature maps, first downsample the input low-layer feature map to the same size as the input high-layer feature map, splice the downsampled feature map and the input high-layer feature map along the channel dimension to obtain a spliced feature map, and use a 1×1 convolution layer to reduce the channels of the spliced feature map so that the fused feature map is output as the fourth feature map;

[0076] Among them, the downsampling is a 3×3 convolution operation.

[0077] Optionally, the weighted fusion of the fourth feature map through a multi-feature fusion module to obtain a third feature map with final angle adaptive fusion includes:

[0078] Obtaining three feature map sets from the fourth feature map. Among them, each feature map set contains four feature maps of the same size but with different target angle direction information; and performing the following operations on each feature map set:

[0079] Adding different feature maps pixel by pixel to obtain a fusion result of the feature maps;

[0080] Performing global average pooling operation on the fused feature map to generate a first feature vector of size 1×1×C; where each first feature vector represents the global information on each channel;

[0081] Generate a second feature vector by passing the first feature vector through a fully connected layer;

[0082] Calculate the channel weighting values of different input feature maps according to the second feature vector;

[0083] Perform adaptive weighted fusion on the feature maps in the feature map set according to the channel weighting values, and output a third feature map.

[0084] Another aspect of the embodiments of the present invention further provides an object detection system based on angle adaptive fusion, including:

[0085] A first module for extracting features from an input image through a backbone network and outputting corresponding first feature maps from the output results of convolutional layers at different stages;

[0086] A second module for performing lateral connection dimensionality reduction and top-down fusion processing on the original feature maps through a feature pyramid network to obtain second feature maps;

[0087] A third module for performing target multi-angle feature extraction and fusion processing on the second feature maps through an angle adaptive fusion pyramid network to obtain third feature maps;

[0088] A fourth module for classifying and regressing the third feature map through a detector head network and outputting object detection results.

[0089] Another aspect of the embodiments of the present invention further provides an electronic device, including a processor and a memory;

[0090] The memory is used to store programs;

[0091] The processor executes the program to implement the method as described above.

[0092] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method as described above.

[0093] The embodiments of the present invention also disclose a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method as described above.

[0094] The following combines the description of the accompanying drawings of the specification to elaborate in detail the specific implementation principle of the present invention:

[0095] Starting from the problem of diverse target angles in object detection, based on existing solutions, this invention proposes an object detection algorithm based on angle - adaptive fusion. By setting up multi - angle convolutional branches and an adaptive multi - branch fusion module, it extracts feature representations with low direction sensitivity from the input feature map. Compared with existing work, the method proposed in this chapter improves the model's ability to represent multi - angle targets without expanding the dataset and without restricting the size of the input feature map, and is applicable to various object detection models.

[0096] The object detection network based on angle - adaptive fusion proposed in this invention is constructed based on the existing baseline network and mainly consists of four parts: the backbone network, the feature pyramid construction module, the angle - adaptive fusion pyramid network proposed in this embodiment, and the head network of the detector. The overall structure is as Figure 1 shown. First, the backbone network extracts features from the input image and selects appropriate feature maps as outputs from the output results of convolutional layers at different stages. Then, the feature pyramid network performs lateral connection dimensionality reduction by 1×1 convolution and top - down fusion on the feature maps {C3, C4, C5} from the backbone network to obtain feature maps {F3, F4, F5}. Then, these feature maps pass through the angle - adaptive fusion pyramid network of this embodiment to obtain feature maps {R3, R4, R5, R6, R7} that are more robust to target angle transformation. Finally, the detector head network performs classification and regression on the feature maps after angle - adaptive fusion.

[0097] The most core module of this invention is the angle - adaptive fusion pyramid network, which contains two modules: the multi - angle feature extraction module and the multi - feature adaptive fusion module. In this embodiment, each input feature map is first sent into the multi - angle feature extraction module, which sets multiple offset convolutional branches with different angle offsets and a bottom - up connection and fusion path, and can obtain multiple feature maps with different angle direction information and retaining shallow details. For example, for the input F3, a set of feature maps extracted after the convolutional kernels are offset by 0°, 90°, 180°, and 270° is obtained Then, these multi - angle feature map sets are sent into the multi - feature fusion module, and weighted fusion is performed using channel attention to obtain the final angle - adaptive fusion features {R3, R4, R5, R6, R7}. Note that R6 and R7 are obtained by upsampling R5 and R6 respectively, and the upsampling method is convolution. The specific structures of the multi - angle feature extraction module and the multi - feature fusion module will be introduced in detail later.

[0098] 1. Multi - angle feature extraction module

[0099] The main function of this module is to enhance the ability of the feature map to represent the target from multiple angles. To achieve this goal, in this embodiment, the convolution kernel is rotated at different angles to form multiple offset convolution branches, so as to extract feature information from different angles of the input. Relevant experiments have proved that rotating the convolution kernel can achieve the same multi-angle representation effect as rotating the feature map, but the rotating convolution kernel has more advantages. First, the operation of the rotating convolution kernel does not limit the size of the input feature map, and the convolution process is similar to ordinary convolution. Second, the operation of the rotating convolution kernel is more flexible. Without introducing interpolation operations, it can be rotated at multiple different angles, which reduces the computational complexity of the model. For example, a 3×3 convolution kernel can have 8 different rotation angles: 0°, 45°, 90°, 135°, 180°, 225°, 270°, 335°. Third, the rotating convolution kernel can reduce the number of parameters of the model because the multiple offset convolution branches share weights and do not introduce new parameters, which helps to avoid the problem of overfitting, especially for small-scale training sets.

[0100] For the feature maps from the feature pyramid, the low-level feature maps have obtained certain semantic information through top-down fusion. If a large convolution kernel is used for rotating convolution operations, it will not only destroy the local detail information of the low-level feature maps but also increase the number of parameters of the model. Therefore, a 3×3 convolution kernel is used here for rotating convolution. The operation of rotating the convolution kernel is as Figure 2 shown. It can be seen that the operation of the rotating convolution kernel is actually the counterclockwise inward rotation of the positions corresponding to the weight parameters according to the rotation angle.

[0101] The structure of the multi-angle feature extraction module is as Figure 3 shown. It is mainly divided into two parts: multi-angle offset convolution feature extraction and bottom-up connection. In the multi-angle offset convolution feature extraction part, this embodiment sets four offset convolution branches with angle offsets of 0°, 90°, 180°, and 270° to extract feature information of different directions of the input feature map. Specifically, taking a convolution branch with a convolution kernel size of 3×3 as a reference, regarding it as an offset convolution branch with an angle offset of 0°, and then rotating the convolution kernel of the reference branch by 90°, 180°, and 270° respectively according to the rotation method introduced above to obtain three offset convolution branches with specific angle offsets. The weights of these four convolution branches are shared. Express this process with the formula as follows:

[0102]

[0103] In the formula, n = {3, 4, 5} represents the layer number of the input feature map, θ = {0, 90, 180, 270} represents the rotation angle of the convolution kernel, represents the nth layer feature map of the input, The convolution kernel for performing an offset convolution operation with a rotation angle of θ on the feature map of the nth layer, b n The bias value when performing a convolution operation on the feature map of the nth layer, Denotes the output feature map after performing an offset convolution operation with a rotation angle of θ on the input feature map of the nth layer. Here, all convolution branches are set to convolutions with a kernel size of 3×3, a stride of 1, and a padding of 0, and the size of the output feature map is the same as that of the input feature map.

[0104] Then, for the output feature maps from the multi-angle offset convolution branches, in this embodiment, a bottom-up connection fusion path is constructed between the results of the same-angle offset convolution of adjacent feature maps, which is to shorten the path from the low layer to the high layer and enrich the detailed information of the high-layer feature maps. The specific connection is as Figure 3 shown, and this process can be expressed by the following formula:

[0105]

[0106] For the lowest-layer feature map This embodiment directly outputs it to obtain For the bottom-up fusion process of other feature maps, in this embodiment, the input low-layer feature map is first downsampled to the same size as the input high-layer feature map , that is, from w n-1 ×h n-1 ×c to w n ×h n ×c. The sampling method selects a 3×3 convolution operation. Then, the downsampled feature map and the input high-layer feature map are concatenated along the channel dimension to obtain a concatenated feature map with a size of w n ×h n ×2c. Finally, a 1×1 convolutional layer is used to reduce the channels of the concatenated feature map, so that the size of the fused feature map is changed to w n ×h n ×c.

[0107] 2. Multi-feature Adaptive Fusion Module Based on Channel Attention

[0108] After the feature maps {F3, F4, F5} from the feature pyramid network pass through the multi-angle feature extraction module, three sets of feature maps can be obtained, namely Each feature map set contains four feature maps of the same size but with different angular direction information of the target. Then, in this embodiment, feature fusion is performed on each feature map set. This embodiment designs a multi-feature adaptive fusion module with reference to the structure of SKNet. This module can adaptively adjust the direction representation of the target according to the target direction representation information of each feature map in the feature map set to obtain the final angle-adaptive fusion feature map. The specific structure of the multi-feature adaptive fusion module is as shown in Figure 4 shown.

[0109] The process of multi-feature adaptive fusion is mainly divided into two steps: feature aggregation and feature selection. In the process of feature aggregation, this embodiment first fuses the feature maps of different branches, that is, adds them pixel by pixel:

[0110]

[0111] Then this embodiment performs global average pooling operation on the fused feature map to generate a feature vector s of size 1×1×C, which represents the global information on each channel:

[0112]

[0113] In Equation (2-2), W, H, and C are the width, height, and number of channels of the feature map respectively. Then s is passed through a fully connected layer to generate a d×1 vector z, as shown in Equation (2-3), which can more accurately and adaptively guide feature selection.

[0114] z = F fc (s) = σ(BN(Ws)) (2-3)

[0115] In the formula, δ represents the ReLU function, BN represents the batch normalization function, and W ∈ R d×C is the weight of the fully connected layer. To learn the influence of the parameter d on the effect of the model, this embodiment introduces a decay ratio r to control its value:

[0116] d = max(C / r, L) (2-4)

[0117] In the above formula, C represents the number of channels, and L refers to the minimum value that d can take. In the experiments of this embodiment, L = 32.

[0118] After feature aggregation, this embodiment uses channel attention to select features. Specifically, this embodiment passes the global feature description z through a c , b c , d c , f cA function is used to calculate the channel weighting values of different input feature maps, and the function expression is shown in Equation (2-5):

[0119]

[0120] where A, B, D, F ∈ R C×d , and a, b, d, f are respectively channel soft attention vectors of c ∈ R 1×d is the c-th row of A, a c is the c-th element of a, and the definitions of other matrix rows and vector elements are similar. To implement weight setting for multi-branch feature maps, this embodiment uses the softmax function to limit a c + b c + d c + f c = 1. Finally, the generated function value is multiplied by the original and summed, as shown in Equation (2-6), to obtain the finally output angle adaptive fusion feature map R n .

[0121]

[0122] In summary, the angle adaptive fusion pyramid network proposed by the present invention can improve the model's ability to represent targets at multiple angles when added to a common object detection framework. The angle adaptive fusion pyramid network includes two modules: one is a multi-angle feature extraction module, which sets multiple offset convolution branches with different angle offsets and a bottom-up connection path, can extract multi-directional representations of the input target, and shortens the path to transfer shallow detail features to the high layer; the other is a multi-feature adaptive fusion module, which uses channel attention to adaptively weight and fuse multiple feature maps containing information about different directions of the target, and can obtain a feature map with stronger robustness to target direction changes. Under the same backbone network and object detector settings, this embodiment conducts a comparative experiment on the angle adaptive fusion pyramid network and the feature pyramid network on two publicly available object detection datasets (PASCAL VOC, COCO). The experimental results show that the angle adaptive fusion pyramid network has better detection effects and obtains higher detection accuracy.

[0123] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are envisioned in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.

[0124] In addition, while the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art will be able to implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0125] If the functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0126] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence of executable instructions for implementing a logical function, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.

[0127] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, deciphering, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0128] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0129] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0130] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0131] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A target detection method based on angle - adaptive fusion, characterized in that, it includes: Performing feature extraction on the input image through a backbone network, and outputting corresponding first feature maps from the output results of convolutional layers at different stages; Performing horizontal connection dimensionality reduction and top - down fusion processing on the original feature maps through a feature pyramid network to obtain second feature maps; Performing target multi - angle feature extraction and fusion processing on the second feature maps through an angle - adaptive fusion pyramid network to obtain third feature maps; Performing classification and regression on the third feature maps through a detector head network to output target detection results; The step of performing target multi - angle feature extraction and fusion processing on the second feature maps through the angle - adaptive fusion pyramid network to obtain third feature maps includes: Processing the input second feature maps through a multi - angle feature extraction module to obtain multiple fourth feature maps with different angle - direction information and retaining shallow - layer details; Performing weighted fusion on the fourth feature maps through a multi - feature fusion module to obtain the final third feature maps with angle - adaptive fusion; The step of processing the input second feature maps through the multi - angle feature extraction module to obtain multiple fourth feature maps with different angle - direction information and retaining shallow - layer details includes: Performing rotational convolution processing on the second feature maps using a 3×3 convolutional kernel, and rotating the positions corresponding to the weight parameters counter - clockwise by the rotation angle to output fifth feature maps; Performing bottom - up connection fusion on the fifth feature maps to obtain fourth feature maps; The step of performing rotational convolution processing on the second feature maps using a 3×3 convolutional kernel, and rotating the positions corresponding to the weight parameters counter - clockwise by the rotation angle to output fifth feature maps is specifically: Configuring four offset convolutional branches with angle offsets of 0°, 90°, 180°, and 270°, and using the four offset convolutional branches to extract feature information in different directions of the input feature maps respectively to output fifth feature maps; The step of configuring four offset convolutional branches with angle offsets of 0°, 90°, 180°, and 270°, and using the four offset convolutional branches to extract feature information in different directions of the input feature maps respectively to output fifth feature maps includes: Taking a convolutional branch with a convolutional kernel size of 3×3 as a reference branch, and using the reference branch as the offset convolutional branch with an angle offset of 0°; Rotating the convolutional kernels of the reference branch by 90°, 180°, and 270° respectively to obtain three offset convolutional branches with specific angle offsets, where the weights of the four offset convolutional branches are shared with each other; Outputting fifth feature maps through each convolutional branch; The step of performing bottom - up connection fusion on the fifth feature maps to obtain fourth feature maps includes: Constructing a bottom - up connection fusion path between the results of the same - angle - offset convolution of adjacent feature maps; For the feature maps of the lowest layer, directly outputting them as fourth feature maps; For the feature maps that are not in the lowest layer, first downsample the input low-level feature maps to the same size as the input high-level feature maps. Concatenate the downsampled feature maps and the input high-level feature maps along the channel dimension to obtain the concatenated feature maps. Use a 1×1 convolutional layer to reduce the channels of the concatenated feature maps, so that the fused feature maps are output as the fourth feature maps; wherein, the downsampling is a 3×3 convolutional operation; The weighted fusion of the fourth feature maps by the multi-feature fusion module to obtain the third feature maps with final angle adaptive fusion includes: Obtain three sets of feature maps from the fourth feature maps, where each set of feature maps contains four feature maps of the same size but with different target angle direction information; and perform the following operations on each set of feature maps: Add the different feature maps pixel by pixel to obtain the fusion result of the feature maps; Perform global average pooling operation on the fused feature maps to generate the first feature vectors of size 1×1×C; where each first feature vector represents the global information on each channel; Pass the first feature vectors through a fully connected layer to generate the second feature vectors; Calculate the channel weighting values of different input feature maps according to the second feature vectors; Perform adaptive weighted fusion on the feature maps in the set of feature maps according to the channel weighting values, and output the third feature maps.

2. A system for implementing the object detection method based on angle adaptive fusion as described in claim 1, characterized in that, it includes: The first module is used to extract features from the input image through the backbone network and output the corresponding first feature maps from the output results of convolutional layers at different stages; The second module is used to perform lateral connection dimensionality reduction and top-down fusion processing on the original feature maps through the feature pyramid network to obtain the second feature maps; The third module is used to perform target multi-angle feature extraction and fusion processing on the second feature maps through the angle adaptive fusion pyramid network to obtain the third feature maps; The fourth module is used to classify and regress the third feature maps through the detector head network and output the object detection results.

3. An electronic device, characterized in that, it includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method as described in claim 1.

4. A computer-readable storage medium, characterized in that, the storage medium stores a program, and the program is executed by the processor to implement the method as described in claim 1.