A target detection method, device, equipment and storage medium
By normalizing and fusion of feature sizes of the target detection network, the problem of insufficient feature fusion in the prior art is solved and the detection accuracy is improved.
Patent Information
- Application Number
- CN202211460829.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-11-17
AI Technical Summary
During the feature fusion process, the existing target detection network only upsamples and simple fusion of the lowest features, resulting in limited improvement in high-resolution feature learning and affecting detection accuracy.
By normalizing the feature size of the target detection network, combining channel fusion and spatial fusion characteristics, the final features of the output layer are determined, and prediction network is used to improve detection accuracy.
By introducing channel fusion and spatial fusion features, the feature expression ability is enhanced and the detection accuracy of target detection is improved.
Smart Images

Figure CN115713683B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to an object detection method, device, equipment and storage medium. Background Art
[0002] Object detection is used to locate and identify objects of interest in images. With the improvement of hardware computing power, the development of deep learning, and the disclosure of high-quality data sets in recent years, object detection has made great progress in recent years. Object detection functions can be roughly divided into two categories: one-stage object detection and two-stage object detection. One-stage object detection is represented by retinanet and the YOLO (You Only Look Once) series, and two-stage object detection is represented by Faster R-CNN (Faster Region-CNN, Fast Region Convolutional Neural Network) and Cascade R-CNN (Cascade Region-CNN, Cascade Region Convolutional Neural Network).
[0003] Most object detection networks can be divided into several main parts: a backbone network, a bottleneck layer network, and a prediction head. Among them, the more important part is the bottleneck layer network. After feature extraction by the backbone network, features of different resolutions and sizes at different layers are generally obtained. The expression capabilities of features at different levels of the feature map are different. Shallow features mainly reflect details such as brightness and edges, while deep features reflect a more comprehensive overall structure. Most detections in the bottleneck layer network part use a feature pyramid network to fuse multi-level features. However, during the fusion process, only the features of the bottom layer are upsampled and single-convolved, and simply fused with the bottom layer features to obtain high-resolution features, resulting in limited improvement in learning for high-resolution features, thereby affecting the detection accuracy of object detection. Therefore, improvement is urgently needed. Summary of the Invention
[0004] The present invention provides an object detection method, device, equipment and storage medium to improve the detection accuracy of object detection.
[0005] According to one aspect of the present invention, there is provided an object detection method, including:
[0006] Using the feature extraction network of the object detection network, performing feature extraction on the image to be detected to obtain initial features output by at least one output layer;
[0007] According to the initial features output by the output layer, respectively performing feature size normalization on the initial features output by at least one output layer to obtain at least one normalized feature corresponding to the output layer;
[0008] Determine the channel fusion features corresponding to the output layer and the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer;
[0009] Determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features;
[0010] Use the prediction network of the object detection network to predict the final features corresponding to at least one output layer, and obtain the prediction result of the image to be detected.
[0011] According to another aspect of the present invention, there is provided an object detection device, including:
[0012] An initial feature determination module, configured to use the feature extraction network of the object detection network to extract features from the image to be detected, and obtain the initial features output by at least one output layer;
[0013] A normalization feature determination module, configured to perform feature size normalization on the initial features output by at least one output layer respectively according to the initial features output by the output layer, and obtain the normalized features of at least one output layer corresponding to the output layer;
[0014] A fusion feature determination module, configured to determine the channel fusion features corresponding to the output layer and the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer;
[0015] A final feature determination module, configured to determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features;
[0016] A prediction result determination module, configured to use the prediction network of the object detection network to predict the final features corresponding to at least one output layer, and obtain the prediction result of the image to be detected.
[0017] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0018] At least one processor; and
[0019] A memory communicatively connected to at least one processor; wherein,
[0020] The memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor so that at least one processor can execute the object detection method of any embodiment of the present invention.
[0021] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the object detection method according to any embodiment of the present invention when executed.
[0022] In the technical solution of the embodiment of the present invention, by using the feature extraction network of the object detection network, feature extraction is performed on the image to be detected to obtain initial features output by at least one output layer; according to the initial features output by the output layer, feature size normalization is respectively performed on the initial features output by at least one output layer to obtain at least one normalized feature corresponding to the output layer; according to the initial features of the output layer and at least one normalized feature corresponding to the output layer, the channel fusion feature corresponding to the output layer is determined, and the spatial fusion feature corresponding to the output layer is determined; according to the channel fusion feature and the spatial fusion feature, the final feature corresponding to the output layer is determined; the prediction network of the object detection network is used to predict the final features corresponding to at least one output layer to obtain the prediction result of the image to be detected. In the above technical solution, the introduction of the channel fusion feature can better characterize the characteristics between different channels, and the introduction of the spatial fusion feature can better characterize the correlation between different output levels, so that the final feature obtained based on the channel fusion feature and the spatial fusion feature has better expression ability, thereby making the prediction result obtained according to the final feature more accurate and improving the detection accuracy of object detection.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0025] Figure 1A is a flowchart of an object detection method according to Embodiment 1 of the present invention;
[0026] Figure 1B is a schematic diagram of an object detection process according to Embodiment 1 of the present invention;
[0027] Figure 2A is a flowchart of an object detection method according to Embodiment 2 of the present invention;
[0028] Figure 2BIt is a schematic diagram of the determination process of channel fusion features provided in Embodiment 2 of the present invention;
[0029] Figure 3A It is a flowchart of a target detection method provided in Embodiment 3 of the present invention;
[0030] Figure 3B It is a schematic diagram of the determination process of spatial fusion features provided in Embodiment 3 of the present invention;
[0031] Figure 4 It is a schematic structural diagram of a target detection device provided in Embodiment 4 of the present invention;
[0032] Figure 5 It is a schematic structural diagram of an electronic device for implementing the target detection method of the embodiments of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "target", "initial" and "final" in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] In addition, it should be noted that in the technical solutions of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of the images to be detected and the like comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0036] Embodiment 1
[0037] Figure 1A It is a flowchart of a target detection method provided in Embodiment 1 of the present invention, Figure 1BA schematic diagram of a target detection process provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of performing target detection on an image. This method can be executed by a target detection device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. The electronic device can be an embedded device, such as various online open-source dataset platforms.
[0038] As Figure 1A and Figure 1B shown, the method includes:
[0039] S101. Use the feature extraction network of the target detection network to extract features from the image to be detected, and obtain initial features output by at least one output layer.
[0040] Among them, the target detection network refers to a pre-trained neural network for performing target detection; the target detection network includes a feature extraction network and a prediction network; among them, the feature extraction network is used to extract features from the image to be detected. Optionally, the feature extraction network includes at least one network layer; for example, it can be a backbone network and / or a bottleneck layer network. The prediction network is used to perform predictions on the image to be detected.
[0041] Among them, the image to be detected can refer to an image that needs to perform target detection. The initial feature can refer to the feature output by each output layer after the feature extraction network extracts features from the image to be detected, and can be represented in the form of a vector or a matrix. It should be noted that the size of the initial features output by different output layers may be the same or different.
[0042] Specifically, the image to be detected can be input into the feature extraction network of the target detection network, and through network learning, initial features output by at least one output layer can be obtained. A specific example is as Figure 1B shown. The bottleneck layer in the feature extraction network outputs 4 output layers. After the image to be detected is input into the feature extraction network, initial features output by 4 output layers can be obtained respectively.
[0043] S102. According to the initial features output by the output layer, perform feature size normalization on the initial features output by at least one output layer respectively, and obtain the normalized features of at least one output layer corresponding to the output layer.
[0044] Among them, the normalized feature can refer to the feature obtained after performing feature scale normalization processing on the initial feature, and can be represented in the form of a vector or a matrix; optionally, the size of the normalized features of at least one output layer corresponding to each output layer is the same as the size of the initial features output by this output layer.
[0045] A specific example is as Figure 1BAs shown, the output layer has 4 layers. Taking output layer 0 as an example, assume that the feature size of the initial features output by output layer 0 is (H, W), where H represents the height of the feature and W represents the width of the feature. Feature size normalization processing is respectively performed on the initial features of output layer 1, output layer 2, and output layer 3 associated with output layer 0, that is, the feature sizes of output layer 1, output layer 2, and output layer 3 are respectively adjusted to (H, W). Specifically, for the output layers in output layer 1, output layer 2, and output layer 3 whose feature sizes are greater than (H, W), an adaptive pooling operation is used to reduce the feature sizes of these output layers to (H, W); for the output layers in output layer 1, output layer 2, and output layer 3 whose feature sizes are less than or equal to (H, W), an upsampling operation is used to expand the feature sizes of these output layers to (H, W), thereby obtaining 4 normalized features corresponding to output layer 0, denoted as R0. Similarly, 4 normalized features corresponding to output layer 1 can be obtained, denoted as R1; 4 normalized features corresponding to output layer 2 can be obtained, denoted as R2; 4 normalized features corresponding to output layer 3 can be obtained, denoted as R3. Among them, the adaptive pooling operation outputs the input feature size according to the specified output feature size. The upsampling operation is used to convert the input small feature size into a feature size that meets the conditions.
[0046] It can be understood that according to the initial features output by the output layer, feature size normalization is respectively performed on the initial features output by at least one output layer, thereby obtaining the normalized features of at least one output layer corresponding to the output layer, increasing the richness of semantics between different output layers.
[0047] S103. Determine the channel fusion feature corresponding to the output layer and determine the spatial fusion feature corresponding to the output layer according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer.
[0048] Among them, the channel fusion feature can refer to the feature obtained by fusing different channel features and can be represented in the form of a vector or a matrix. The spatial fusion feature can refer to the feature obtained by fusing the features of different output layers in the same space or different spaces and can be represented in the form of a vector or a matrix. The channel fusion feature can be obtained through a second-order feature attention module; the spatial fusion feature can be obtained through a multi-size spatial fusion module.
[0049] Specifically, according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer, based on a preset calculation rule, the channel fusion feature and the spatial fusion feature corresponding to the output layer are determined.
[0050] A specific example is as Figure 1BAs shown, taking the output layer 0 as an example, based on the initial features of the output layer 0 and the four normalized features corresponding to the output layer 0 (denoted as R0), the channel fusion features corresponding to the output layer 0 are determined through SOFA (second order feature attention module), and the spatial fusion features corresponding to the output layer 0 are determined through MSSF (Multi-scale spatial fusion). Similarly, the channel fusion features and spatial fusion features corresponding to the output layer 1, output layer 2, and output layer 3 in Figure 1B can be obtained. Figure 1B The channel fusion features and spatial fusion features corresponding to the output layer 1, output layer 2, and output layer 3.
[0051] S104. Determine the final features corresponding to the output layer according to the channel fusion features and spatial fusion features.
[0052] Among them, the final features can refer to the features obtained after specific processing of the channel fusion features and spatial fusion features corresponding to the output layer, and can be represented in the form of vectors or matrices.
[0053] A specific example is as Figure 1B shown. The channel fusion features and spatial fusion features corresponding to each output layer can be added to obtain the final features corresponding to each output layer.
[0054] S105. Use the prediction network of the object detection network to predict the final features corresponding to at least one output layer to obtain the prediction result of the image to be detected.
[0055] Among them, the prediction network is used to predict the image to be detected according to the final features corresponding to each input output layer. The prediction network can be a one-stage prediction network or a two-stage prediction network.
[0056] Specifically, input the final features corresponding to each output layer into the prediction network of the object detection network, and through network learning, obtain the prediction result of the image to be detected. If it is a one-stage prediction network, directly input the final features corresponding to the output layer into the prediction head to obtain the prediction result of the image to be detected; if it is a two-stage prediction network, it is necessary to input the final features corresponding to the output layer into the target candidate boxes extracted by the convolutional neural network through the region candidate network; then input the final features in the target candidate boxes into the prediction head to obtain the prediction result of the image to be detected.
[0057] A specific example is as Figure 1B shown. Input the final features corresponding to the output layer 0, output layer 1, output layer 2, and output layer 3 into the prediction network of the object detection network, and through network learning, obtain the prediction result of the image to be detected.
[0058] In the technical solution of the embodiment of the present invention, by using the feature extraction network of the target detection network, feature extraction is performed on the image to be detected, and initial features output by at least one output layer are obtained; according to the initial features output by the output layer, feature size normalization is respectively performed on the initial features output by at least one output layer, and at least one normalized feature corresponding to the output layer is obtained; according to the initial features of the output layer and at least one normalized feature corresponding to the output layer, the channel fusion feature corresponding to the output layer is determined, and the spatial fusion feature corresponding to the output layer is determined; according to the channel fusion feature and the spatial fusion feature, the final feature corresponding to the output layer is determined; the prediction network of the target detection network is used to predict the final features corresponding to at least one output layer, and the prediction result of the image to be detected is obtained. In the above technical solution, by introducing the channel fusion feature, the characteristics between different channels can be better represented, and by introducing the spatial fusion feature, the correlation between different output levels can be better represented, so that the final feature obtained based on the channel fusion feature and the spatial fusion feature has better expression ability, thereby making the prediction result obtained according to the final feature more accurate and improving the detection accuracy of target detection.
[0059] Embodiment 2
[0060] Figure 2A The flowchart of a target detection method provided by Embodiment 2 of the present invention. On the basis of the above embodiment, this embodiment further optimizes "determining the channel fusion feature corresponding to the output layer according to the initial feature of the output layer and at least one normalized feature corresponding to the output layer", and provides an optional implementation solution. It should be noted that for the parts not detailed in the embodiments of the present invention, reference may be made to the relevant descriptions of the foregoing embodiments. As Figure 2A and 2B shown, the method includes:
[0061] S201. Use the feature extraction network of the target detection network to perform feature extraction on the image to be detected, and obtain the initial features output by at least one output layer.
[0062] S202. According to the initial features output by the output layer, perform feature size normalization on the initial features output by at least one output layer respectively, and obtain at least one normalized feature corresponding to the output layer.
[0063] S203. Determine the channel score weights of at least one channel of the output layer according to the normalized features of at least one output layer.
[0064] Among them, the channel score weight is used to represent the importance degree of each channel feature of the output layer to the channel fusion feature of the output layer.
[0065] Specifically, based on a preset rule, at least one channel score weight of the output layer is determined according to the normalized features of at least one output layer.
[0066] Exemplarily, the normalized features of at least one output layer can be added to obtain a sum feature corresponding to the output layer; the covariance of the sum feature is calculated to obtain a hierarchical correlation feature corresponding to the output layer; the hierarchical correlation feature is orthogonally decomposed to obtain at least one channel feature corresponding to the hierarchical correlation feature; at least two convolution operations are performed on the channel feature to obtain the channel score weight corresponding to the channel.
[0067] Among them, the sum feature can refer to the feature obtained by adding the normalized features of each output layer, and can be represented in the form of a vector or a matrix. The covariance is used to measure the overall error of the normalized features of multiple output layers to obtain the correlation between the features of each output layer. The hierarchical correlation feature is used to represent the features related between different hierarchical networks, and can be obtained by calculating the covariance of the sum feature corresponding to the output layer, and can be represented in the form of a vector or a matrix.
[0068] Specifically, the normalized features of at least one output layer can be added to obtain a sum feature corresponding to the output layer, and then the covariance of the sum feature is calculated to obtain a hierarchical correlation feature corresponding to the output layer. Exemplarily, the sum feature is a feature of H×W×C, where C represents the number of channels, H represents the height of the feature, and W represents the width of the feature; the sum feature is re-denoted as C×s, s = H×W, and the hierarchical correlation feature can be determined by the following formula:
[0069]
[0070] Among them, Σ represents the hierarchical correlation feature, I represents the s×s identity matrix, 1 represents the all-ones matrix, X T represents the transpose matrix of X.
[0071] It should be noted that using the standard covariance of the feature space to characterize the correlation between different channels can better characterize the characteristics of different channels and the relationship between different channels than the existing operations such as global average pool, and can better promote the fusion at the channel level.
[0072] After obtaining the hierarchical correlation feature, the hierarchical correlation feature is orthogonally decomposed to obtain at least one channel feature corresponding to the hierarchical correlation feature. Exemplarily, the channel feature can be determined by the following formula:
[0073] Y = ∑ a = U∧ a U T ,
[0074] Among them, Y = [y1, …, y C represents the channel features corresponding to each output layer, Σ represents the hierarchical correlation features, a represents the positive real power, and here a is selected U represents an orthogonal matrix, U T represents the transpose matrix of U, ∧ represents the diagonal matrix of eigenvalues in non-increasing order (λ1, …, λ C ), and C represents the number of channels.
[0075] After that, average pooling is performed on the channel features of each output layer to obtain the average channel features of each output layer. Exemplarily, the average channel features of each output layer can be determined by the following formula:
[0076]
[0077] Among them, Z C represents the average channel features of each output layer, C represents the number of channels, and y i represents the features of the i-th (i = 0, 1, …, C - 1) channel.
[0078] After obtaining Z C , at least two convolution operations are performed on Z C to obtain the channel score weights corresponding to the channels. Exemplarily, the channel score weights corresponding to the channels can be determined by the following formula:
[0079] w = f(W a δ(W b Z C ))
[0080] Among them, w represents the channel score weight, f() is used to calculate the channel score weights of each channel, W a represents the convolution weight of convolution operation a (denoted as conv_a), W b represents the convolution weight of convolution operation b (denoted as conv_b), δ() represents the relu and normalise functions, and relu and normalise are activation functions.
[0081] A specific example is as Figure 2B shown. Taking the output layer 0 as an example, the normalised features of the 4 output layers corresponding to the output layer 0 are added to obtain the sum feature corresponding to the output layer 0; the covariance of the sum feature corresponding to the output layer 0 is calculated to obtain the hierarchical correlation feature corresponding to the output layer 0; the hierarchical correlation feature is orthogonally decomposed to obtain the channel features of at least one channel corresponding to the hierarchical correlation feature; the convolution weight for the channel features is W aconv_a to extract the important parameters of the channel features; then, after relu and normalise operations, the channel features are converted from linear to non-linear to maintain the stability of model training; then, convolution with weight W b is performed on conv_b to obtain the channel score weights corresponding to each channel of output layer 0.
[0082] The above example provides a method for calculating the channel score weights, obtaining the correlation between the features of each output layer by calculating the covariance of the summation features; at the same time, by performing orthogonal decomposition on the hierarchy-related features, the important features of each channel in each output layer are extracted, better representing the characteristics of different channels and the relationships between different channels, and better promoting the fusion of the channel hierarchy.
[0083] S204. Determine the channel fusion features corresponding to the output layer according to the initial features of the output layer and the channel score weights of at least one channel of the output layer.
[0084] Exemplarily, the initial features of the output layer can be multiplied by the corresponding channel score weights of at least one channel respectively to obtain the channel fusion features corresponding to the output layer.
[0085] Specifically, multiply the initial features of each output layer by the corresponding channel score weights of at least one channel to obtain the channel fusion features corresponding to each output layer.
[0086] A specific example is as Figure 2B shown. Taking output layer 0 as an example, multiply the initial features of output layer 0 by the corresponding channel score weights of at least one channel. This operation is denoted as x to obtain the channel fusion features corresponding to output layer 0.
[0087] The above example provides a method for calculating the channel fusion features, which can accurately obtain the importance degree of each channel, extract the important features of each channel, so that more abundant channel features can be learned, laying a foundation for subsequent object detection.
[0088] S205. Determine the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the normalised features of at least one output layer corresponding to the output layer.
[0089] Specifically, according to the initial features of the output layer and the normalised features of at least one output layer corresponding to the output layer, based on a preset calculation rule, generate the spatial fusion features corresponding to the output layer.
[0090] S206. Determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features.
[0091] S207. Use the prediction network of the object detection network to predict the final features corresponding to at least one output layer, and obtain the prediction result of the image to be detected.
[0092] In the technical solution of the embodiment of the present invention, according to the normalized features of at least one output layer, determine the channel score weights of at least one channel of the output layer; according to the initial features of the output layer and the channel score weights of at least one channel of the output layer, determine the channel fusion features corresponding to the output layer. The above technical solution provides a method for determining channel fusion features, adding all the output layer features to obtain a feature map with different receptive fields. This feature map has more semantic information and representation information than the existing single-layer feature map. By calculating the channel score weights to clarify the importance of each channel feature, and according to the initial features of the output layer and the channel score weights, determine the channel fusion features corresponding to the output layer, so that the obtained channel fusion features are more accurate, and further make the final features corresponding to the output layer determined according to the channel fusion features and the spatial fusion features more accurate, thereby improving the detection accuracy of object detection; this method has second-order differentiability and can reduce the phenomenon of loss explosion.
[0093] Embodiment III
[0094] Figure 3A It is a flowchart of an object detection method provided by Embodiment III of the present invention. On the basis of the above embodiment, this embodiment further optimizes "determine the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer", and provides an optional implementation solution. It should be noted that for the parts not detailed in the embodiments of the present invention, reference may be made to the relevant descriptions of the foregoing embodiments. Such as Figure 3A and Figure 3B shown, the method includes:
[0095] S301. Use the feature extraction network of the object detection network to extract features from the image to be detected, and obtain the initial features output by at least one output layer.
[0096] S302. According to the initial features output by the output layer, perform feature size normalization on the initial features output by at least one output layer respectively, and obtain the normalized features of at least one output layer corresponding to the output layer.
[0097] S303. According to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer, determine the channel fusion features corresponding to the output layer.
[0098] S304. According to the normalized features of at least one output layer, determine the spatial hierarchical features corresponding to the output layer.
[0099] Among them, the spatial hierarchical features are used to represent the features related to different spaces in the same output layer, which can be obtained by using the normalized features of each output layer and can be represented in the form of vectors or matrices.
[0100] Specifically, based on the three-dimensional convolution operation, the spatial hierarchical features corresponding to the output layer can be determined according to the normalized features of at least one output layer.
[0101] Optionally, the normalized features of at least one output layer are respectively dimension-expanded to obtain the dimension-expanded normalized features of at least one output layer; the dimension-expanded normalized features of at least one output layer are merged to obtain the merged features corresponding to the output layer; and three-dimensional convolution is performed on the merged features to obtain the spatial hierarchical features corresponding to the output layer.
[0102] Specifically, for each output layer, the normalized features of each output layer corresponding to the output layer are respectively dimension-expanded. Among them, the normalized features of each output layer can be respectively features of H×W×C. Specifically, a dimension (denoted as the level axis) is added to the normalized features of the output layer, so as to obtain the dimension-expanded normalized features of each output layer, which can be denoted as features of 1×H×W×C; the dimension-expanded normalized features of each output layer are merged along the expanded dimension (denoted as the level axis) to obtain the merged features corresponding to the output layer, which can be denoted as features of L×H×W×C, where L represents the dimension of the merged features. Exemplarily, if the number of output layers is 4, then L is 4; furthermore, at least two three-dimensional convolution operations are performed on the merged features corresponding to the output layer to obtain the processed merged features, and then the reshape function is used to perform dimension reduction processing on the processed merged features, so as to obtain the spatial hierarchical features corresponding to the output layer, which can be denoted as features of H×W×C.
[0103] A specific example is as Figure 3B shown. Taking output layer 0 as an example, the normalized features of 4 output layers corresponding to output layer 0 are dimension-expanded (denoted as unsqueeze) to obtain the dimension-expanded normalized features of 4 output layers corresponding to output layer 0, which are respectively denoted as uR0, uR1, uR2, and uR3; the dimension-expanded normalized features of 4 output layers corresponding to output layer 0 are merged to obtain the merged features corresponding to output layer 0; the merged features corresponding to output layer 0 are subjected to three-dimensional convolution operation (denoted as 3Dconv block) along the level axis, and then a three-dimensional normalized convolution operation is performed. The AvgPool3d() function is used to perform global pooling on the merged features along the level axis to obtain the processed merged features, and then the reshape function is used to perform dimension reduction processing on the processed merged features, so as to obtain the spatial hierarchical features corresponding to output layer 0.
[0104] The above example provides a method for obtaining the spatial hierarchical features corresponding to each output layer. By means of dimension expansion and three-dimensional convolution operations, the correlation of features between different output layers is enhanced, and the fusion along the dimension change of convolution is taken into account.
[0105] S305. Determine the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the spatial hierarchical features corresponding to the output layer.
[0106] Specifically, specific operations can be performed on the initial features of the output layer and the spatial hierarchical features corresponding to the output layer, so as to obtain the spatial fusion features corresponding to the output layer.
[0107] Optionally, merge the initial features of the output layer and the spatial hierarchical features corresponding to the output layer; perform a convolution operation on the merged spatial hierarchical features to obtain the spatial fusion features.
[0108] Specifically, merge the initial features of the output layer and the spatial hierarchical features corresponding to the output layer to obtain features of H×W×2C, that is, the merged spatial hierarchical features; perform a convolution operation with a convolution kernel of 1×1 on the merged spatial hierarchical features to obtain the spatial fusion features corresponding to the output layer.
[0109] A specific example is as Figure 3B shown. Taking output layer 0 as an example, merge the initial features of output layer 0 and the spatial hierarchical features corresponding to output layer 0 to obtain the merged spatial hierarchical features; perform a convolution operation with a convolution kernel of 1×1 on the merged spatial hierarchical features to obtain the spatial fusion features corresponding to output layer 0.
[0110] The above example provides a method for calculating spatial fusion features. By using three-dimensional convolution operations, stereo convolution operations are performed on the features of different-sized output layers and different spaces, increasing the fusion of the interaction and relationship between different spaces of the same output layer and the same space of different output layers, thus laying a foundation for subsequent object detection.
[0111] S306. Determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features.
[0112] S307. Use the prediction network of the object detection network to predict the final features corresponding to at least one output layer to obtain the prediction result of the image to be detected.
[0113] In the technical solution of the embodiment of the present invention, according to the normalized features of at least one output layer, the spatial hierarchical features corresponding to the output layer are determined; according to the initial features of the output layer and the spatial hierarchical features corresponding to the output layer, the spatial fusion features corresponding to the output layer are determined. The above technical solution provides a method for obtaining spatial fusion features. By expanding dimensions and three-dimensional convolution operations, spatial fusion features are obtained, making the spatial fusion features better able to represent the correlation of features between different output layers, making the final features corresponding to the output layer determined according to the channel fusion features and spatial fusion features more accurate, and thus improving the detection accuracy of object detection; introducing three-dimensional convolution into two-dimensional image object detection, replacing the four-layer output layer with the time sequence of the original three-dimensional convolution, so that the spatial feature layer obtained by three-dimensional convolution not only has spatial features, but also has semantic information between different output layers.
[0114] Embodiment 4
[0115] Figure 4 FIG. is a schematic structural diagram of an object detection device provided in Embodiment 4 of the present invention. This embodiment is applicable to the situation of object detection for images. The device can be implemented in the form of hardware and / or software, and can be configured in an electronic device, which can be an embedded device, such as various online open-source dataset platforms. As Figure 4 shown, the device includes:
[0116] An initial feature determination module 401, configured to use a feature extraction network of an object detection network to perform feature extraction on an image to be detected, and obtain initial features output by at least one output layer;
[0117] A normalization feature determination module 402, configured to perform feature size normalization on the initial features output by at least one output layer respectively according to the initial features output by the output layer, and obtain the normalized features of at least one output layer corresponding to the output layer;
[0118] A fusion feature determination module 403, configured to determine the channel fusion features corresponding to the output layer and determine the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the normalized features of at least one output layer corresponding to the output layer;
[0119] A final feature determination module 404, configured to determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features;
[0120] A prediction result determination module 405, configured to use a prediction network of an object detection network to perform prediction on the final features corresponding to at least one output layer, and obtain a prediction result of the image to be detected.
[0121] In the technical solution of the embodiment of the present invention, an initial feature acquisition module is used to acquire initial features output by at least one output layer; a normalized feature acquisition module is used to acquire normalized features of at least one output layer corresponding to the output layer; a fusion feature determination module is used to acquire channel fusion features and spatial fusion features corresponding to the output layer; a final feature determination module is used to acquire final features corresponding to the output layer; a prediction result acquisition module is used to acquire a prediction result of an image to be detected. In the above technical solution, by introducing channel fusion features, the characteristics between different channels can be better represented, and by introducing spatial fusion features, the correlation between different output levels can be better represented, so that the final features obtained based on the channel fusion features and spatial fusion features have better expression ability, thereby making the prediction result obtained based on the final features more accurate and improving the detection accuracy of object detection.
[0122] Optionally, the fusion feature determination module 403 includes:
[0123] A channel score weight determination unit for determining channel score weights of at least one channel of the output layer according to the normalized features of at least one output layer;
[0124] A channel fusion feature determination unit for determining channel fusion features corresponding to the output layer according to the initial features of the output layer and the channel score weights of at least one channel of the output layer.
[0125] Optionally, the channel score weight determination unit is specifically used for:
[0126] Adding the normalized features of at least one output layer to obtain a sum feature corresponding to the output layer; calculating the covariance of the sum feature to obtain a hierarchical correlation feature corresponding to the output layer; performing orthogonal decomposition on the hierarchical correlation feature to obtain channel features of at least one channel corresponding to the hierarchical correlation feature; performing at least two convolution operations on the channel features to obtain channel score weights corresponding to the channels.
[0127] Optionally, the channel fusion feature determination unit is specifically used for:
[0128] Multiplying the initial features of the output layer by the channel score weights of the corresponding at least one channel respectively to obtain channel fusion features corresponding to the output layer.
[0129] Optionally, the fusion feature determination module 403 includes:
[0130] A spatial hierarchical feature determination unit for determining spatial hierarchical features corresponding to the output layer according to the normalized features of at least one output layer;
[0131] A spatial fusion feature determination unit is configured to determine the spatial fusion feature corresponding to the output layer according to the initial feature of the output layer and the spatial hierarchical feature corresponding to the output layer.
[0132] Optionally, the spatial hierarchical feature determination unit is specifically configured to:
[0133] Respectively expand the normalized features of at least one output layer to obtain the expanded normalized features of at least one output layer; merge the expanded normalized features of at least one output layer to obtain the merged feature corresponding to the output layer; perform three-dimensional convolution on the merged feature to obtain the spatial hierarchical feature corresponding to the output layer.
[0134] Optionally, the spatial fusion feature determination unit is specifically configured to:
[0135] Merge the initial feature of the output layer and the spatial hierarchical feature corresponding to the output layer; perform a convolution operation on the merged spatial hierarchical feature to obtain the spatial fusion feature.
[0136] The target detection device provided by the embodiments of the present invention can execute the target detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing each target detection method.
[0137] Embodiment 5
[0138] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0139] As Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0140] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0141] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the object detection method.
[0142] In some embodiments, the object detection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the object detection method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the object detection method by any other appropriate means (e.g., by means of firmware).
[0143] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0144] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0145] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0147] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0148] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0150] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A target detection method, characterized in that, Including: Using the feature extraction network of the object detection network to extract features from the image to be detected, and obtaining the initial features output by at least one output layer; According to the initial features output by the output layer, respectively performing feature size normalization on the initial features output by at least one output layer to obtain the normalized features of at least one output layer corresponding to the output layer; Adding the normalized features of at least one output layer to obtain the summation feature corresponding to the output layer; Calculating the covariance of the summation feature to obtain the hierarchical correlation feature corresponding to the output layer; Performing orthogonal decomposition on the hierarchical correlation feature to obtain the channel features of at least one channel corresponding to the hierarchical correlation feature; Performing at least two convolution operations on the channel features to obtain the channel score weights corresponding to the channel; Determining the channel fusion feature corresponding to the output layer according to the initial feature of the output layer and the channel score weights of at least one channel of the output layer; Respectively expanding the dimensions of the normalized features of at least one output layer to obtain the expanded normalized features of at least one output layer; Merging the expanded normalized features of at least one output layer to obtain the merged feature corresponding to the output layer; Performing three-dimensional convolution on the merged feature to obtain the spatial hierarchical feature corresponding to the output layer; Determining the spatial fusion feature corresponding to the output layer according to the initial feature of the output layer and the spatial hierarchical feature corresponding to the output layer; Determining the final feature corresponding to the output layer according to the channel fusion feature and the spatial fusion feature; Using the prediction network of the object detection network to predict the final features corresponding to at least one output layer to obtain the prediction result of the image to be detected.
2. The method according to claim 1, characterized in that, Determining the channel fusion feature corresponding to the output layer according to the initial feature of the output layer and the channel score weights of at least one channel of the output layer, including: Respectively multiplying the initial feature of the output layer by the channel score weights of the corresponding at least one channel to obtain the channel fusion feature corresponding to the output layer.
3. The method according to claim 1, wherein The determining the spatial fusion feature corresponding to the output layer according to the initial feature of the output layer and the spatial hierarchical feature corresponding to the output layer, including: Merging the initial feature of the output layer and the spatial hierarchical feature corresponding to the output layer; Performing a convolution operation on the merged spatial hierarchical feature to obtain the spatial fusion feature.
4. A target detection device, characterized in that, Including: An initial feature determination module, configured to use the feature extraction network of the object detection network to extract features from the image to be detected, and obtain the initial features output by at least one output layer; A normalized feature determination module, configured to respectively perform feature size normalization on the initial features output by at least one output layer according to the initial features output by the output layer to obtain the normalized features of at least one output layer corresponding to the output layer; A fusion feature determination module, configured to add the normalized features of at least one output layer to obtain the summation feature corresponding to the output layer; calculating the covariance of the summation feature to obtain the hierarchical correlation feature corresponding to the output layer; Perform orthogonal decomposition on the hierarchical correlation features to obtain the channel features of at least one channel corresponding to the hierarchical correlation features; Perform at least two convolutional operations on the channel features to obtain the channel score weights corresponding to the channels; determine the channel fusion features corresponding to the output layer according to the initial features of the output layer and the channel score weights of at least one channel of the output layer; Respectively expand the normalized features of at least one output layer to obtain the expanded normalized features of at least one output layer; Merge the expanded normalized features of at least one output layer to obtain the merged features corresponding to the output layer; Perform three-dimensional convolution on the merged features to obtain the spatial hierarchical features corresponding to the output layer; determine the spatial fusion features corresponding to the output layer according to the initial features of the output layer and the spatial hierarchical features corresponding to the output layer; A final feature determination module, configured to determine the final features corresponding to the output layer according to the channel fusion features and the spatial fusion features; A prediction result determination module, configured to use the prediction network of the object detection network to predict the final features corresponding to at least one output layer to obtain the prediction result of the image to be detected; 5. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the object detection method according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the processor to implement the object detection method according to any one of claims 1-3 when executed.
Citation Information
Patent Citations
Target detection method and device, storage medium and electronic equipment
CN114399696A
Target detection method and device, equipment and storage medium
CN114419410A