Target detection method, device, computer equipment and storage medium

By performing convolution and scale transformation on multi-level feature maps, the false alarm and inaccurate warning problems of traditional blind spot obstacle detection technology in complex environments are solved, and high-accuracy target detection is achieved.

CN115147784BActive Publication Date: 2025-10-03CHANGSHA INTELLIGENT DRIVING INST CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110332658.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-29
Publication Date
2025-10-03
Estimated Expiration
2041-03-29

AI Technical Summary

Technical Problem

Traditional blind spot obstacle detection technology is easily interfered with in complex environments, resulting in false alarms and inaccurate warnings, and poor detection results.

Method used

By obtaining multi-level feature maps of the image to be detected, performing convolution and scale transformation fusion, combining the information of feature maps at different levels, performing target frame detection, and obtaining the target detection result.

Benefits of technology

The accuracy of target detection is improved, effective detection of obstacles in blind spots is achieved, and good detection results are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147784B_ABST
    Figure CN115147784B_ABST
Patent Text Reader

Abstract

The present application relates to a target detection method, device, computer equipment and storage medium. The method includes: obtaining a multi-level feature map of an image to be detected, determining a current-level feature map and a next-level feature map based on the multi-level feature map, convolving the current-level feature map to obtain a convolved feature map, fusing the convolved feature map and the next-level feature map through scale transformation, convolving the fused feature map to obtain a fused feature map, using the fused feature map as a new convolved feature map, determining a new next-level feature map based on the multi-level feature map, returning to the step of fusing the convolved feature map and the next-level feature map through scale transformation, realizing the fusion of adjacent-level feature maps in the multi-level feature map to obtain a fused multi-level feature map, performing target frame detection on each layer of fused feature map in the fused multi-level feature map, and obtaining a target detection result. This method can achieve good detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a target detection method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of computer technology, blind spot obstacle detection technology has emerged. Blind spot obstacle detection technology is mainly used to detect the inherent blind spots of vehicles such as muck trucks and sanitation trucks during transportation, in order to reduce the occurrence of vehicle accidents.

[0003] In traditional technologies, blind spot obstacle detection technology mainly uses detection methods such as detecting the vehicle's blind spots based on gyroscopes and ultrasonic modules, and warning of vehicle speeding by monitoring the throttle control switch.

[0004] However, traditional technologies that use gyroscopes and ultrasonic modules to detect vehicle blind spots are easily susceptible to interference and produce false alarms in complex environments. Monitoring the throttle control switch to warn of vehicle speeding behavior can also result in inaccurate warnings, and both have the problem of poor detection results. Summary of the Invention

[0005] Based on this, it is necessary to provide a target detection method, device, computer equipment and storage medium that can achieve good detection effects in response to the above technical problems.

[0006] A target detection method, comprising:

[0007] Obtain a multi-level feature map of the image to be detected;

[0008] According to the multi-level feature map, determine the current level feature map and the next level feature map, the level of the next level feature map is lower than the current level feature map;

[0009] Convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map;

[0010] The fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map. The level of the new next-level feature map is lower than that of the current next-level feature map.

[0011] Return to the step of fusing the convolved feature map and the next-level feature map by scaling until there is no new next-level feature map in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0012] The target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0013] A target detection device, comprising:

[0014] An acquisition module is used to obtain a multi-level feature map of the image to be detected;

[0015] A first processing module is configured to determine a feature map of a current level and a feature map of a next level according to the multi-level feature map, wherein the level of the feature map of the next level is lower than that of the feature map of the current level;

[0016] The fusion module is used to convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map;

[0017] An update module is used to use the fused feature map as a new convolved feature map and determine a new next-level feature map based on the multi-level feature map, where the level of the new next-level feature map is lower than that of the current next-level feature map;

[0018] The second processing module is used to return to the step of fusing the convolved feature map and the next-level feature map through scale transformation until there is no new next-level feature map in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0019] The detection module is used to perform target frame detection on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0020] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0021] Obtain a multi-level feature map of the image to be detected;

[0022] According to the multi-level feature map, determine the current level feature map and the next level feature map, the level of the next level feature map is lower than the current level feature map;

[0023] Convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map;

[0024] The fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map. The level of the new next-level feature map is lower than that of the current next-level feature map.

[0025] Return to the step of fusing the convolved feature map and the next-level feature map by scaling until there is no new next-level feature map in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0026] The target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0027] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0028] Obtain a multi-level feature map of the image to be detected;

[0029] According to the multi-level feature map, determine the current level feature map and the next level feature map, the level of the next level feature map is lower than the current level feature map;

[0030] Convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map;

[0031] The fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map. The level of the new next-level feature map is lower than that of the current next-level feature map.

[0032] Return to the step of fusing the convolved feature map and the next-level feature map by scaling until there is no new next-level feature map in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0033] The target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0034] The above-mentioned target detection method, device, computer equipment and storage medium obtain a multi-level feature map of the image to be detected, determine the current level feature map and the next level feature map according to the multi-level feature map, convolve the current level feature map to obtain a convolved feature map, fuse the convolved feature map and the next level feature map through scale transformation, convolve the fused feature map to obtain a fused feature map, and use the fused feature map as a new convolved feature map, determine a new next level feature map according to the multi-level feature map, and return the convolved feature map obtained by scale transformation. The step of fusing the feature map with the feature map of the next level can realize the fusion of the feature maps of adjacent levels in the multi-level feature map to obtain a fused multi-level feature map, so that the target detection result can be obtained by performing target frame detection on each layer of the fused feature map in the fused multi-level feature map. In the whole process, by fusing the features of the multi-level feature map and combining the image features of different levels, the feature information of the fused multi-level feature map used for target detection is enriched, which can improve the accuracy of the target detection result, realize the detection of blind spot obstacles, and achieve good detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 1 is a flow chart of a target detection method according to an embodiment;

[0036] Figure 2 is a schematic diagram of a target detection method in one embodiment;

[0037] Figure 3 is a schematic diagram of a target detection method in another embodiment;

[0038] Figure 4 is a schematic diagram of a target detection method in yet another embodiment;

[0039] Figure 5 Schematic diagram of a target detection method in yet another embodiment;

[0040] Figure 6 is a flow chart of a target detection method in another embodiment;

[0041] Figure 7 1 is a flow chart of a target detection method in another embodiment;

[0042] Figure 8 is a structural block diagram of a target detection device in one embodiment;

[0043] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0045] In one embodiment, Figure 1 As shown, a target detection method is provided. This embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0046] Step 102: Obtain a multi-level feature map of the image to be detected.

[0047] The image to be detected refers to a pre-collected image that requires target detection. For example, when detecting obstacles in a blind spot while a vehicle is driving, the image to be detected can specifically refer to a road image captured by the vehicle. A multi-level feature map refers to feature maps at multiple levels of different scales, i.e., each level of the feature map has a different scale.

[0048] Specifically, when it is necessary to detect obstacles in the blind spot, the server will obtain the image to be detected, perform feature extraction on the image to be detected, and use convolution layers of different sizes to obtain feature maps of different scales. Furthermore, the server can extract features from the image to be detected using an image feature extraction network including convolution layers of different sizes. For example, the image feature extraction network can be DarkNet53, whose network structure is as follows: Figure 2 As shown in the figure, the image to be detected is input and passes through CBL, Res1, Res2, Res8, Res8, and Res4 to obtain three feature maps: level 1, level 2, and level 3. The CBL module consists of a convolutional layer, a batch normalization layer, and an activation layer. The Res* module (including Res1, Res2, Res8, Res8, and Res4) consists of a CBL module and N Res units. The input of the Res unit module is calculated by two CBL modules and then added to the original input.

[0049] Step 104: Determine a feature map of a current level and a feature map of a next level based on the multi-level feature maps, where the level of the feature map of the next level is lower than that of the feature map of the current level.

[0050] Among them, the current level feature map refers to the feature map with the highest level among the multi-level feature maps, and the next level feature map refers to the feature map after the current level feature map and with a level lower than the current level feature map.

[0051] Specifically, after obtaining the multi-level feature map, the server will sort the multi-level feature map according to the number of channels of each layer of the feature map in the multi-level feature map to obtain the hierarchical order of the feature map. According to the hierarchical order of the feature map, the current level feature map and the next level feature map can be determined. It should be noted here that when sorting the multi-level feature map according to the number of channels of each layer of feature map, it can be sorted in ascending order or in descending order. For example, when sorting the multi-level feature map in ascending order according to the number of channels, the current level feature map refers to the feature map with the least number of channels. For another example, when sorting the multi-level feature map in descending order according to the number of channels, the current level feature map refers to the feature map with the largest number of channels.

[0052] In step 106, the feature map of the current level is convolved to obtain a convolved feature map, the convolved feature map is fused with the feature map of the next level through scale transformation, and the fused feature map is convolved to obtain a fused feature map.

[0053] Among them, scale transformation refers to transforming the scale of the convolved feature map to be the same as the feature map of the next level. The scale transformation methods include sampling and dimension transformation.

[0054] Specifically, because the scale of each feature map in the multi-level feature map is different, when performing feature fusion, it is necessary to scale the feature map by sampling and dimensionality transformation to achieve the fusion of feature maps of the same scale. After obtaining the convolved feature map, the server will scale the convolved feature map according to the size parameters of the feature map of the next level, and scale the convolved feature map to a feature map of the same size as the feature map of the next level, so that it can be fused by direct addition, and the fused feature map is convolved to obtain a fused feature map. For example, when the size of the convolved feature map is 32*32*512, and the size of the feature map of the next level is also 32*32*512, the fusion is to directly add the corresponding values, and the final result is a fused feature map of size 32*32*512.

[0055] Specifically, when the convolved feature map is scaled according to the size parameters of the next-level feature map, if the number of channels of the next-level feature map is smaller than the convolved feature map, the sampling method adopted is upsampling, and the dimensionality transformation adopted is dimensionality reduction. If the number of channels of the next-level feature map is greater than the convolved feature map, the sampling method adopted is downsampling, and the dimensionality transformation adopted is dimensionality increase. When upsampling, bilinear interpolation can be used for sampling, or other methods can be used. This embodiment does not make specific limitations here.

[0056] In this embodiment, before performing the scale transformation, the current level feature map needs to be convolved to obtain richer level semantic information. Furthermore, the current level feature map can be convolved through a preset preprocessing convolutional network to obtain a convolved feature map. The preprocessing convolutional network can be specifically composed of convolutional networks of different sizes, including at least two convolution kernels of different sizes. For example, the convolution kernel sizes in the preprocessing convolutional network are 1*1 and 3*3, which not only ensures the stability of the network parameters, but also uses large convolution kernels to increase the receptive field area so that the network can learn better feature information.

[0057] In step 108 , the fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map. The level of the new next-level feature map is lower than that of the current next-level feature map.

[0058] Specifically, after obtaining the fused feature map, the server will use the fused feature map as a new convolved feature map, and determine a new next-level feature map with a level lower than the current next-level feature map based on the multi-level feature map to continue feature fusion.

[0059] Step 110, returning to the step of fusing the convolved feature map and the next-level feature map through scale transformation until there is no new next-level feature map in the multi-level feature map, and obtaining a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time.

[0060] Specifically, after updating the convolved feature map and the next-level feature map, the server will return to the step of fusing the convolved feature map and the next-level feature map through scale transformation to obtain a new fused feature map until there is no new next-level feature map in the multi-level feature map, indicating that each layer of feature map in the multi-level feature map has participated in the fusion, and the fused multi-level feature map is obtained based on the convolved feature map obtained for the first time and the fused feature map obtained each time.

[0061] For example, the embodiment is described by taking the multi-level feature map including the first scale feature map, the second scale feature map and the third scale feature map, and the number of channels of the second scale feature map is smaller than that of the first scale feature map, and the number of channels of the third scale feature map is smaller than that of the second scale feature map. Figure 3As shown in the figure, the level 3 feature map is the first-scale feature map, the level 2 feature map is the second-scale feature map, and the level 1 feature map is the third-scale feature map. The level 3 feature map is the current-level feature map, the level 2 feature map is the next-level feature map, and the level 1 feature map is the new next-level feature map. The server calculates the level 3 feature map through the preprocessing convolutional network to obtain the F3 feature map (i.e., the convolved feature map). The size of the F3 feature map calculated by the preprocessing convolutional network is consistent with the level 3 feature map. The F3 feature map is upsampled using bilinear interpolation to obtain a feature map F3_Up that is consistent with the width and height of the level 2 feature map. The feature map F3_Up is dimensionality-reduced to make it consistent with the channel dimension of the level 2 feature map. It is then directly added to the level 2 feature map to obtain the feature map temp2. ​​This greatly reduces the number of channels in the feature map after the fusion of features at different levels, making the computational complexity of the entire feature map fusion process very small. The feature map temp2 is calculated through the preprocessing convolutional network to obtain the F2 feature map (i.e., the fused feature map, the new convolved feature map). The F2 feature map is upsampled by bilinear interpolation to obtain the feature map F2_Up with the same width and height as the level1 feature map. The feature map F2_Up is dimensionally reduced to make it consistent with the channel dimension of the level1 feature map. The addition method is used to achieve fusion with the level1 feature map to obtain the feature map temp1. Similarly, the feature map temp1 is calculated through the preprocessing convolutional network to obtain the F1 feature map. According to the initially obtained convolved feature map (i.e., the F3 feature map) and the fused feature maps obtained each time (i.e., the F2 feature map and the F1 feature map), the fused multi-level feature map is obtained.

[0062] In step 112, target frame detection is performed on each fused feature map in the multi-level fused feature maps to obtain a target detection result.

[0063] Specifically, the server can obtain the corresponding target detection result by performing target frame detection on each layer of the fused feature map in the fused multi-level feature map. For example, the server can obtain the target detection network through training, and obtain the corresponding target detection result by using the target detection network to perform target frame detection on each layer of the fused feature map in the fused multi-level feature map. Furthermore, the server can obtain the target detection network by training by obtaining training samples and the initial target prediction network, and optimizing the initial target prediction network according to the training samples. Before using the target detection network to perform target frame detection on each layer of the fused feature map in the fused multi-level feature map, the server can also perform sparse training and pruning on the target detection network to further optimize the target detection network and improve the network processing speed. It should be noted that when pruning, in order to further improve the processing speed of target detection, in addition to pruning the target detection network, the server can also simultaneously prune the above-mentioned image feature extraction network, preprocessing convolution network, etc.

[0064] The above-mentioned target detection method obtains a multi-level feature map of the image to be detected, determines the current level feature map and the next level feature map according to the multi-level feature map, convolves the current level feature map to obtain a convolved feature map, fuses the convolved feature map and the next level feature map through scale transformation, convolves the fused feature map to obtain a fused feature map, and uses the fused feature map as a new convolved feature map, determines a new next level feature map according to the multi-level feature map, and returns to the step of fusing the convolved feature map and the next level feature map through scale transformation. It can realize the fusion of adjacent level feature maps in the multi-level feature map to obtain a fused multi-level feature map, so that the target detection result can be obtained by performing target box detection on each layer of the fused feature map in the fused multi-level feature map. In the whole process, by fusing the features of the multi-level feature map and combining the image features of different levels, the feature information of the fused multi-level feature map used for target detection is enriched, which can improve the accuracy of the target detection result, realize the detection of blind area obstacles, and achieve good detection effect.

[0065] In one embodiment, the convolved feature map and the next-level feature map are fused by scaling, and the fused feature map is convolved to obtain a fused feature map including:

[0066] The convolved feature map is sampled and dimensionally transformed in sequence to obtain a transformed feature map;

[0067] The transformed feature map and the next-level feature map are fused by direct addition, and the fused feature map is convolved to obtain the fused feature map.

[0068] Specifically, the server will first sample the convolved feature map to obtain a sampled feature map with the same width and height as the feature map of the next level, then perform dimension transformation on the sampled feature map to obtain a transformed feature map with the same number of channels as the feature map of the next level, and finally fuse the transformed feature map and the feature map of the next level by direct addition, and convolve the fused feature map to obtain a fused feature map.

[0069] In this embodiment, by sampling and dimensionality transforming the convolved feature map in sequence, scale transformation can be achieved to obtain a transformed feature map, and then by directly adding the transformed feature map and the feature map of the next level for fusion, and convolving the fused feature map, a fused feature map can be obtained.

[0070] In one embodiment, convolving the feature map of the current level to obtain the convolved feature map includes:

[0071] The current level feature map is convolved through a preset preprocessing convolutional network to obtain a convolved feature map. The preprocessing convolutional network includes at least two convolution kernels of different sizes.

[0072] Among them, the preprocessing convolutional network is used to perform partial convolution operations on the feature map to deepen the depth of the network and obtain richer hierarchical semantic features. For example, the preprocessing convolutional network can be specifically composed of convolutional networks of different sizes. Furthermore, the preprocessing convolutional network can be specifically composed of convolutional networks of different sizes appearing alternately. For example, the network structure of the preprocessing convolutional network can be as follows Figure 4 As shown in the figure, it consists of CBL1 to CBL5, and the sizes of the convolution kernels of CBL1 to CBL5 are 1*1 and 3*3, respectively, alternating with a step size of 1. The alternating sizes of the convolution kernels 1*1 and 3*3 ensure the stability of the network parameters and also use large convolution kernels to increase the receptive field area, allowing the network to learn better feature information. Each CBL is composed of the conv+bn+relu mode.

[0073] Specifically, the server convolves the current level feature map through a preprocessing convolutional network including at least two convolution kernels of different sizes to obtain a convolved feature map.

[0074] In this embodiment, the current level feature map is convolved through a preset preprocessing convolutional network to obtain a convolved feature map with rich semantic information.

[0075] In one embodiment, target frame detection is performed on each fused feature map in the multi-level fused feature map, and the target detection result obtained includes:

[0076] Obtain training samples and an initial target prediction network, optimize the initial target prediction network based on the training samples, and obtain a target detection network;

[0077] According to the target detection network, target frame detection is performed on each layer of fused feature maps in the multi-level fused feature maps to obtain the target detection result.

[0078] The training samples refer to images collected in advance and used as a training set for constructing a target detection model. For example, the training samples may specifically refer to images collected in advance, including blind spot obstacles, captured by a camera installed on the vehicle while the vehicle is driving. The initial target prediction network refers to a target detection network that has not yet undergone parameter adjustment. The target detection network can be obtained by adjusting the parameters of the initial target prediction network. The initial target detection network is used to predict the target frame of each layer of the fused feature map in the multi-level fused feature map, so that the model parameters in the initial target detection network are adjusted according to the predicted target frame and the real target frame carried by the training samples to obtain the target detection network.

[0079] Specifically, when obtaining the target detection result, the server will first obtain the training sample and the initial target prediction network, optimize the initial target prediction network according to the training sample to obtain the target detection network, and then use the target detection network to perform target frame detection on each layer of the fused feature map in the fused multi-level feature map to obtain the target detection result. Among them, when optimizing the initial target prediction network according to the training sample to obtain the target detection network, the server needs to first extract features from the training sample to obtain a multi-level feature map corresponding to the training sample, and then fuse each layer of feature maps in the multi-level feature map to obtain a fused multi-level feature map. The fused multi-level feature map is input into the initial target prediction network so that the initial target prediction network outputs a predicted target frame corresponding to the fused multi-level feature map. Then, based on the predicted target frame and the true target frame carried by the training sample, the model parameters in the initial target detection network are tuned so that the predicted target frame is closer to the true target frame, thereby obtaining the target detection network. It should be noted that the number of predicted target boxes is not limited in this embodiment, but preferably, the number of predicted target boxes corresponds to the size of the feature map to be predicted, which is an integer multiple N of the sum of the pixels of the feature map to be predicted, that is, N boxes are predicted for each pixel on the feature map. More preferably, the value of N can be 3.

[0080] In this embodiment, by obtaining training samples and an initial target prediction network, the initial target prediction network is optimized according to the training samples to obtain a target detection network. The target detection network can be used to perform target frame detection on each layer of the fused multi-level feature map to obtain a target detection result.

[0081] In one embodiment, obtaining a training sample and an initial target prediction network, optimizing the initial target prediction network based on the training sample, and obtaining a target detection network includes:

[0082] Obtain training samples with true target frames and an initial target prediction network;

[0083] The target box of the training sample is predicted through the initial target prediction network to obtain the predicted target box;

[0084] Determine the IOU value between the predicted target box and the true target box;

[0085] According to the IOU value and the preset positive sample threshold and negative sample threshold, the positive and negative sample ratio corresponding to the initial target prediction network is determined;

[0086] According to the positive and negative sample ratio, the initial target prediction network is adjusted to obtain the target detection network.

[0087] The real target frame refers to the rectangular frame that marks the target detection object in the training sample. For example, the real target frame can specifically refer to the rectangular frame that marks the blind spot obstacle in the training sample image. The real target frame carried by the training sample refers to the rectangular frame that has been marked in the training sample and includes the target detection object. The IOU value refers to the ratio of the intersection and union between the predicted target frame and the real target frame, such as Figure 5 As shown, Where A and B are the areas of the predicted target box and the true target box, respectively. The positive-negative sample ratio refers to the ratio of the number of positive samples to the number of negative samples at the prediction box level. A positive sample at the prediction box level refers to a predicted target box that overlaps with the true target box and the degree of overlap is greater than the preset positive sample threshold. A negative sample refers to a predicted target box that overlaps with the true target box and the degree of overlap is less than the preset negative sample threshold.

[0088] Specifically, the server will obtain training samples carrying real target frames and an initial target prediction network, perform feature extraction on the training samples, obtain a multi-level feature map corresponding to the training samples, and then fuse each layer of feature maps in the multi-level feature map to obtain a fused multi-level feature map. The fused multi-level feature map is input into the initial target prediction network so that the initial target prediction network outputs a predicted target frame corresponding to the fused multi-level feature map, and then calculates the IOU value between the predicted target frame and the real target frame carried by the training sample. According to the IOU value and the preset positive sample threshold, the number of positive samples corresponding to the initial target prediction network is determined. According to the IOU value and the preset negative sample threshold, the number of negative samples corresponding to the initial target prediction network is determined. According to the number of positive samples and the number of negative samples, the positive-negative sample ratio corresponding to the initial target prediction network is determined. The positive-negative sample ratio is used to adjust the model parameters of the initial target prediction network so that the positive-negative sample ratio meets the preset positive-negative sample ratio requirement, and the target detection network is obtained.

[0089] Among them, the preset positive-negative sample ratio requirement can specifically be one of the following conditions: the positive-negative sample ratio is greater than the preset positive-negative sample ratio threshold, the number of positive samples is greater than the preset positive sample number threshold, and the number of negative samples is less than the preset negative sample number threshold, or a combination of several conditions. This embodiment does not specifically limit the positive-negative sample ratio requirement. Preferably, the positive sample threshold can be 0.6, and the negative sample threshold can be 0.3. In this embodiment, different thresholds are used to distinguish between positive and negative samples, avoiding the direct use of a single threshold to distinguish between positive and negative samples, so that positive samples may be classified as negative samples, affecting the distribution ratio of positive and negative samples, as well as the optimization effect and convergence speed of the network.

[0090] In this embodiment, by first obtaining a training sample carrying a true target frame and an initial target prediction network, the initial target prediction network performs target frame prediction on the training sample to obtain a predicted target frame, and then determines the IOU value between the predicted target frame and the true target frame. The positive and negative sample ratio is calculated using the IOU value, and the positive and negative sample ratio can be used to adjust the initial target prediction network to obtain a target detection network.

[0091] In one embodiment, according to the target detection network, target frame detection is performed on each fused feature map in the multi-level fused feature map, and the target detection results obtained include:

[0092] Performing sparse training on the target detection network to obtain a sparse target detection network;

[0093] Adaptively prune the sparse object detection network according to preset channel pruning thresholds and convolutional layer pruning thresholds to obtain a pruned object detection network, and retrain the pruned object detection network;

[0094] According to the retrained pruned object detection network, a trained object detection network is obtained;

[0095] According to the trained target detection network, target box detection is performed on each layer of fused feature maps in the multi-level fused feature maps to obtain the target detection result.

[0096] The preset channel pruning threshold refers to a preset threshold for channel pruning, specifically the threshold for the gamma value of the normalization layer in the convolutional layer. The preset convolutional layer pruning threshold refers to a preset threshold for convolutional layer pruning, specifically the threshold for the number of channels in the convolutional layer.

[0097] Specifically, the server will perform sparsification training on the target detection network to obtain a sparse target detection network, perform channel-adaptive pruning on the sparse target detection network according to a preset channel pruning threshold, and then adaptively prune the channel-pruned model according to the convolutional layer pruning threshold on the basis of channel pruning to obtain a pruned target detection network, and retrain the pruned target detection network using the training sample images. Based on the retrained pruned target detection network, a trained target detection network is obtained. Based on the trained target detection network, target box detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0098] In this embodiment, the method of sparse training is not specifically limited. As long as sparse training can be achieved so that the weights corresponding to unimportant neurons are zero or close to zero, the method of retraining the pruned object detection model is the same as the method of training the object detection network. Furthermore, obtaining a trained target detection network based on the retrained pruned target detection network means that after obtaining the retrained pruned target detection network, a model evaluation is first performed on the retrained pruned target detection network. When the model evaluation passes, the retrained pruned target detection network is directly used as the trained target detection network. When the model evaluation fails, the channel pruning threshold and the convolutional layer pruning threshold are adjusted. According to the adjusted channel pruning threshold and the convolutional layer pruning threshold, the sparse target detection network is pruned again to obtain a new retrained pruned target detection network. The new retrained pruned target detection network is evaluated again. If the model evaluation passes, the trained target detection network is obtained. If the model evaluation fails, the channel pruning threshold and the convolutional layer pruning threshold are adjusted again and pruned again until the model evaluation passes and the trained target detection network is obtained.

[0099] In this embodiment, by performing sparse training, channel pruning, and convolutional layer pruning on the target detection network, the model can be optimized to obtain a lightweight trained target detection network. The trained target detection network can then be used to perform target box detection on each layer of the fused feature map in the multi-level fused feature map to quickly obtain the target detection result.

[0100] In one embodiment, the sparse object detection network is adaptively pruned according to a preset channel pruning threshold and a convolutional layer pruning threshold to obtain a pruned object detection network including:

[0101] Get the gamma value of the normalization layer in all convolutional layers in the sparse object detection network;

[0102] Perform channel pruning on all convolutional layers in the sparse object detection network according to the channel pruning threshold and gamma value to obtain a channel-pruned network.

[0103] Get the current number of channels of all convolutional layers in the channel-pruned network;

[0104] All convolutional layers in the channel-pruned network are pruned according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned object detection network.

[0105] Convolutional layers all include a normalization layer. This ensures that the eigenvalues ​​in the convolution process conform to a normal distribution, facilitating training. The formula used by the normalization layer is yi = γxi + β, where γ is the gamma value. Each channel in the feature map has a corresponding gamma value, which is consistent with the number of output channels of the convolution kernel. Using the gamma value to implement channel pruning involves pruning channels using a channel pruning threshold and the corresponding gamma value.

[0106] Specifically, the server obtains the gamma values ​​of the normalization layers in all convolutional layers in the sparsified object detection network. It then performs channel pruning by comparing the channel pruning threshold with the gamma value, filtering out channels with gamma values ​​less than the channel pruning threshold. These channels are then pruned, resulting in a channel-pruned network. After channel pruning, some convolutional layers may have only a small number of channels. Therefore, further convolutional layer pruning is performed using the convolutional layer pruning threshold to balance speed and accuracy. Specifically, the server obtains the current number of channels in all convolutional layers in the channel-pruned network. It then compares the convolutional layer pruning threshold with the current number of channels, filtering out convolutional layers with channel numbers less than the convolutional layer pruning threshold. These convolutional layers are then pruned, resulting in a pruned object detection network. Furthermore, since the dimensionality of the gamma value is consistent with the number of output channels of the convolutional kernel in a convolutional layer, during channel pruning, the server records the index number corresponding to the gamma value of the deleted convolutional layer and simultaneously deletes the channel corresponding to the index number on the output channel of the convolutional kernel in that convolutional layer.

[0107] In this embodiment, the sparse object detection network is adaptively pruned according to the preset channel pruning threshold and convolutional layer pruning threshold, so as to obtain a pruned object detection network.

[0108] In one embodiment, obtaining a trained object detection network based on the retrained pruned object detection network includes:

[0109] Obtain a test dataset and perform model evaluation on the retrained pruned object detection network based on the test dataset;

[0110] When the model evaluation passes, the trained object detection network is obtained;

[0111] When the model evaluation fails, the channel pruning threshold and the convolutional layer pruning threshold are adjusted, and the process of adaptively pruning the sparse object detection network to obtain a pruned object detection network is returned. The pruned object detection network is then retrained until the model evaluation passes and a trained object detection network is obtained.

[0112] Among them, the test data set refers to the data collected in advance for model evaluation of the pruned object detection network. The test data in the test data set all carry detection box labels.

[0113] Specifically, the server will obtain a test data set, and perform a model evaluation on the retrained pruned target detection network based on the test data set. When the model evaluation passes, it indicates that the model has been optimized and a trained target detection network is obtained. When the model evaluation fails, the channel pruning threshold and the convolutional layer pruning threshold are adjusted, and the step of adaptively pruning the sparse target detection network is returned to obtain a pruned target detection network, and the pruned target detection network is retrained, and the optimization is performed again until the model evaluation passes, and a trained target detection model is obtained. In particular, when using the test data set to perform a model evaluation on the pruned target detection network, the evaluation criteria used are pre-set as needed. For example, the evaluation criteria can specifically be that the target box prediction accuracy of the pruned target detection network on the test data set reaches a preset accuracy threshold, etc. This embodiment does not make specific limitations here.

[0114] In this embodiment, by performing model evaluation on the retrained pruned object detection network, model optimization can be achieved to obtain a trained object detection network.

[0115] In one embodiment, the target detection method of the present application can be implemented by pre-building a target detection model, using the target detection model to process the image to be detected, and obtaining the target detection result. Figure 6 As shown, a flowchart is used to illustrate a method for building a target detection model. The method for building a target detection model specifically includes the following steps:

[0116] At the beginning of training, the server obtains a training sample image and an initial target detection model. The training sample image carries the real target box information. Through the initial target detection model, the multi-level feature map of the training sample image is obtained, and the multi-level feature map is feature fused to obtain a fused multi-level feature map. The target box is predicted for each layer of the feature map in the fused multi-level feature map to obtain the predicted target box. According to the predicted target box and the real target box information, the target detection model to be optimized is obtained (i.e., model evaluation, selecting a model with better effect). The target detection model to be optimized is sparsely trained and adaptively pruned to obtain a pruned model. The pruned model is retrained and the retrained pruned model is evaluated. When the retrained pruned model meets the requirements, the training is terminated and the trained target detection model is obtained. When the retrained pruned model does not meet the requirements, the preset pruning threshold is adjusted, the sparsely trained model is pruned again, and retrained until the retrained pruned model meets the requirements, thereby obtaining a trained target detection model. Requirements can be set as needed, primarily to evaluate the effectiveness of the pruned model, specifically the preset accuracy. For example, a test dataset can be obtained and used to evaluate the pruned model to determine whether it meets the requirements. Pruning thresholds include channel and convolutional layer thresholds, and the thresholds can be set as needed.

[0117] It should be noted that after obtaining the trained target detection model, by inputting the image to be detected into the trained target detection model, the trained target detection model can be used to extract features of the image to be detected to obtain a multi-level feature map, and feature fusion is performed on the multi-level feature map to obtain a fused multi-level feature map. Target box detection is performed on each layer of feature map in the fused multi-level feature map to obtain a target detection result corresponding to the image to be detected.

[0118] In one embodiment, Figure 7 As shown, a flow chart is used to illustrate the target detection method of the present application, which specifically includes the following steps:

[0119] Step 702: Obtain a multi-level feature map of the image to be detected;

[0120] Step 704: determining a feature map of a current level and a feature map of a next level based on the multi-level feature map, wherein the level of the feature map of the next level is lower than that of the feature map of the current level;

[0121] Step 706: Convolve the feature map of the current level through a preset preprocessing convolutional network to obtain a convolved feature map, where the preprocessing convolutional network includes at least two convolution kernels of different sizes.

[0122] Step 708: Sampling and dimension transformation are performed on the convolved feature map in sequence to obtain a transformed feature map;

[0123] Step 710, fusing the transformed feature map and the next-level feature map by direct addition, and performing convolution on the fused feature map to obtain a fused feature map;

[0124] Step 712: Use the fused feature map as a new convolved feature map, and determine a new next-level feature map based on the multi-level feature map, where the level of the new next-level feature map is lower than the current next-level feature map.

[0125] Step 714, returning to step 708, until there is no new next-level feature map in the multi-level feature map, obtaining a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0126] Step 716: Obtain training samples with true target frames and an initial target prediction network;

[0127] Step 718: Perform target box prediction on the training sample through the initial target prediction network to obtain a predicted target box;

[0128] Step 720: determine the IOU value between the predicted target box and the true target box;

[0129] Step 722: Determine the positive-negative sample ratio corresponding to the initial target prediction network based on the IOU value and the preset positive sample threshold and negative sample threshold;

[0130] Step 724: Adjust the initial target prediction network based on the positive-negative sample ratio to obtain a target detection network.

[0131] Step 726, performing sparse training on the target detection network to obtain a sparse target detection network;

[0132] Step 728, obtaining the gamma value of the normalization layer in all convolutional layers in the sparse object detection network;

[0133] Step 730 , performing channel pruning on all convolutional layers in the sparse object detection network according to the channel pruning threshold and gamma value to obtain a channel-pruned network;

[0134] Step 732, obtaining the current number of channels of all convolutional layers in the channel-pruned network;

[0135] Step 734: prune all convolutional layers in the channel-pruned network according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned object detection network, and retrain the pruned object detection network.

[0136] Step 736: Obtain a test dataset and perform model evaluation on the retrained pruned object detection network based on the test dataset. If the model evaluation passes, jump to step 738; if the model evaluation fails, jump to step 740.

[0137] Step 738: Get the trained object detection network and jump to step 742;

[0138] Step 740: Adjust the channel pruning threshold and the convolutional layer pruning threshold, and return to step 728 until the model evaluation passes, thereby obtaining a trained object detection network.

[0139] In step 742, according to the trained target detection network, target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain a target detection result.

[0140] It should be understood that, although the various steps in the various flow charts that the above-described embodiments relate to are shown in sequence according to the indications of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the various flow charts that the above-described embodiments relate to can include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.

[0141] In one embodiment, Figure 8 As shown, a target detection device is provided, including: an acquisition module 802, a first processing module 804, a fusion module 806, an update module 808, a second processing module 810 and a detection module 812, wherein:

[0142] An acquisition module 802 is used to acquire a multi-level feature map of the image to be detected;

[0143] A first processing module 804 is configured to determine a feature map of a current level and a feature map of a next level according to the multi-level feature map, wherein the level of the feature map of the next level is lower than that of the feature map of the current level;

[0144] A fusion module 806 is configured to convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through a scale transformation, and convolve the fused feature map to obtain a fused feature map.

[0145] An updating module 808 is configured to use the fused feature map as a new convolved feature map, and determine a new next-level feature map based on the multi-level feature map, where the level of the new next-level feature map is lower than that of the current next-level feature map;

[0146] The second processing module 810 is configured to return to the step of fusing the convolved feature map and the next-level feature map by scaling until no new next-level feature map exists in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature maps obtained each time.

[0147] The detection module 812 is used to perform target frame detection on each layer of the fused feature map in the multi-level fused feature map to obtain a target detection result.

[0148] The above-mentioned target detection device obtains a multi-level feature map of the image to be detected, determines the current level feature map and the next level feature map according to the multi-level feature map, convolves the current level feature map to obtain a convolved feature map, fuses the convolved feature map and the next level feature map through scale transformation, convolves the fused feature map to obtain a fused feature map, and uses the fused feature map as a new convolved feature map, determines a new next level feature map according to the multi-level feature map, and returns to the step of fusing the convolved feature map and the next level feature map through scale transformation. It can realize the fusion of adjacent level feature maps in the multi-level feature map to obtain a fused multi-level feature map, so that the target detection result can be obtained by performing target frame detection on each layer of the fused feature map in the fused multi-level feature map. In the whole process, by fusing the features of the multi-level feature map and combining the image features of different levels, the feature information of the fused multi-level feature map used for target detection is enriched, the accuracy of the target detection result can be improved, the detection of blind area obstacles can be realized, and good detection effect can be achieved.

[0149] In one embodiment, the fusion module is also used to sample and dimensionally transform the convolved feature map in sequence to obtain a transformed feature map, fuse the transformed feature map and the next-level feature map by direct addition, and convolve the fused feature map to obtain a fused feature map.

[0150] In one embodiment, the fusion module is further used to convolve the current level feature map through a preset preprocessing convolutional network to obtain a convolved feature map, and the preprocessing convolutional network includes at least two convolution kernels of different sizes.

[0151] In one embodiment, the detection module is also used to obtain training samples and an initial target prediction network, optimize the initial target prediction network based on the training samples to obtain a target detection network, and perform target box detection on each layer of fused feature maps in the multi-level fused feature maps based on the target detection network to obtain a target detection result.

[0152] In one embodiment, the detection module is also used to obtain training samples carrying real target frames and an initial target prediction network, perform target frame prediction on the training samples through the initial target prediction network to obtain a predicted target frame, determine the IOU value between the predicted target frame and the real target frame, and determine the positive and negative sample ratio corresponding to the initial target prediction network based on the IOU value and the preset positive sample threshold and negative sample threshold. According to the positive and negative sample ratio, the initial target prediction network is adjusted to obtain a target detection network.

[0153] In one embodiment, the detection module is also used to perform sparsification training on the target detection network to obtain a sparse target detection network, adaptively prune the sparse target detection network according to a preset channel pruning threshold and a convolutional layer pruning threshold to obtain a pruned target detection network, and retrain the pruned target detection network. Based on the retrained pruned target detection network, a trained target detection network is obtained. Based on the trained target detection network, target box detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain a target detection result.

[0154] In one embodiment, the detection module is further used to obtain the gamma value of the normalization layer in all convolutional layers in the sparse target detection network, perform channel pruning on all convolutional layers in the sparse target detection network according to the channel pruning threshold and the gamma value to obtain a channel-pruned network, obtain the current number of channels of all convolutional layers in the channel-pruned network, perform convolutional layer pruning on all convolutional layers in the channel-pruned network according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned target detection network.

[0155] In one embodiment, the detection module is further used to obtain a test data set, perform model evaluation on the retrained pruned target detection network based on the test data set, and when the model evaluation passes, obtain the trained target detection network; when the model evaluation fails, adjust the channel pruning threshold and the convolutional layer pruning threshold, return to the step of adaptively pruning the sparse target detection network to obtain the pruned target detection network, and retrain the pruned target detection network until the model evaluation passes and the trained target detection network is obtained.

[0156] For the specific definition of the target detection device, please refer to the definition of the target detection method above, which will not be repeated here. Each module in the above-mentioned target detection device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0157] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as training sample images and test data sets. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a target detection method is implemented.

[0158] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0159] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0160] Obtain a multi-level feature map of the image to be detected;

[0161] According to the multi-level feature map, determine the current level feature map and the next level feature map, the level of the next level feature map is lower than the current level feature map;

[0162] Convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map;

[0163] The fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map. The level of the new next-level feature map is lower than that of the current next-level feature map.

[0164] Return to the step of fusing the convolved feature map and the next-level feature map by scaling until there is no new next-level feature map in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time;

[0165] The target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain the target detection result.

[0166] In one embodiment, when the processor executes the computer program, it also implements the following steps: sampling and dimensionally transforming the convolved feature map in sequence to obtain a transformed feature map, fusing the transformed feature map and the next-level feature map by direct addition, and convolving the fused feature map to obtain a fused feature map.

[0167] In one embodiment, when the processor executes the computer program, it further implements the following steps: convolving the feature map of the current level through a preset preprocessing convolutional network to obtain a convolved feature map, and the preprocessing convolutional network includes at least two convolution kernels of different sizes.

[0168] In one embodiment, when the processor executes the computer program, it also implements the following steps: obtaining training samples and an initial target prediction network, optimizing the initial target prediction network based on the training samples to obtain a target detection network, and performing target box detection on each layer of the fused feature map in the multi-level fused feature map based on the target detection network to obtain a target detection result.

[0169] In one embodiment, when the processor executes the computer program, it further implements the following steps: obtaining training samples carrying real target frames and an initial target prediction network, performing target frame prediction on the training samples through the initial target prediction network to obtain a predicted target frame, determining an IOU value between the predicted target frame and the real target frame, determining a positive-negative sample ratio corresponding to the initial target prediction network based on the IOU value and a preset positive sample threshold and a negative sample threshold, and adjusting the initial target prediction network based on the positive-negative sample ratio to obtain a target detection network.

[0170] In one embodiment, when the processor executes the computer program, it further implements the following steps: performing sparsification training on the target detection network to obtain a sparse target detection network, adaptively pruning the sparse target detection network according to a preset channel pruning threshold and a convolutional layer pruning threshold to obtain a pruned target detection network, and retraining the pruned target detection network, obtaining a trained target detection network based on the retrained pruned target detection network, and performing target box detection on each layer of the fused feature map in the multi-level fused feature map based on the trained target detection network to obtain a target detection result.

[0171] In one embodiment, when the processor executes the computer program, it further implements the following steps: obtaining the gamma value of the normalization layer in all convolutional layers in the sparse object detection network, performing channel pruning on all convolutional layers in the sparse object detection network according to the channel pruning threshold and the gamma value to obtain a channel-pruned network, obtaining the current number of channels of all convolutional layers in the channel-pruned network, performing convolutional layer pruning on all convolutional layers in the channel-pruned network according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned object detection network.

[0172] In one embodiment, when the processor executes the computer program, it further implements the following steps: obtaining a test data set, performing a model evaluation on the retrained pruned object detection network based on the test data set, obtaining a trained object detection network when the model evaluation passes, and adjusting the channel pruning threshold and the convolutional layer pruning threshold when the model evaluation fails, returning to the steps of adaptively pruning the sparsed object detection network to obtain a pruned object detection network, and retraining the pruned object detection network until the model evaluation passes to obtain a trained object detection network.

[0173] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0174] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0175] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0176] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A target detection method, characterized in that: The method comprises: Obtain a multi-level feature map of the image to be detected; Determining a current-level feature map and a next-level feature map according to the multi-level feature map, where the level of the next-level feature map is lower than that of the current-level feature map; Convolving the feature map of the current level to obtain a convolved feature map, fusing the convolved feature map and the feature map of the next level through scale transformation, and convolving the fused feature map to obtain a fused feature map; The fused feature map is used as a new convolved feature map, and a new next-level feature map is determined based on the multi-level feature map, where the level of the new next-level feature map is lower than the current next-level feature map; Returning to the step of fusing the convolved feature map and the next-level feature map by scaling until no new next-level feature map exists in the multi-level feature map, obtaining a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time; According to the trained target detection network, target frame detection is performed on each layer of the fused feature map in the multi-level fused feature map to obtain a target detection result; The trained target detection network is trained in the following way: Obtain training samples and an initial target prediction network, optimize the initial target prediction network according to the training samples to obtain a target detection network, perform sparsification training on the target detection network to obtain a sparse target detection network, obtain the gamma value of the normalization layer in all convolutional layers in the sparse target detection network, perform channel pruning on all convolutional layers in the sparse target detection network according to a channel pruning threshold and the gamma value to obtain a channel-pruned network, obtain the current number of channels of all convolutional layers in the channel-pruned network, perform convolutional layer pruning on all convolutional layers in the channel-pruned network according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned target detection network, retrain the pruned target detection network, and obtain a trained target detection network based on the retrained pruned target detection network.

2. The method according to claim 1, characterized in that The step of fusing the convolved feature map and the next-level feature map through scale transformation and convolving the fused feature map to obtain the fused feature map includes: The convolved feature map is sampled and dimensionally transformed in sequence to obtain a transformed feature map; The transformed feature map and the next-level feature map are fused by direct addition, and the fused feature map is convolved to obtain a fused feature map.

3. The method according to claim 1, characterized in that Convolving the current level feature map to obtain a convolved feature map includes: The current-level feature map is convolved through a preset preprocessing convolutional network to obtain a convolved feature map, wherein the preprocessing convolutional network includes at least two convolution kernels of different sizes.

4. The method according to claim 1, wherein The acquiring of training samples and an initial target prediction network, and optimizing the initial target prediction network according to the training samples to obtain a target detection network comprises: Obtain training samples with true target frames and an initial target prediction network; Performing target box prediction on the training sample through the initial target prediction network to obtain a predicted target box; Determine the IOU value between the predicted target box and the true target box; Determine the positive-negative sample ratio corresponding to the initial target prediction network according to the IOU value and the preset positive sample threshold and negative sample threshold; The initial target prediction network is adjusted according to the positive-negative sample ratio to obtain a target detection network.

5. The method according to claim 1, wherein The step of obtaining a trained object detection network based on the retrained pruned object detection network includes: Obtain a test dataset, and perform model evaluation on the retrained pruned object detection network based on the test dataset; When the model evaluation passes, the trained object detection network is obtained; When the model evaluation fails, the channel pruning threshold and the convolutional layer pruning threshold are adjusted, and the step of adaptively pruning the sparse object detection network to obtain a pruned object detection network is returned to, and the pruned object detection network is retrained until the model evaluation passes and a trained object detection network is obtained.

6. A target detection device, characterized in that: The device comprises: An acquisition module is used to obtain a multi-level feature map of the image to be detected; A first processing module is configured to determine a current-level feature map and a next-level feature map based on the multi-level feature map, where the level of the next-level feature map is lower than that of the current-level feature map; A fusion module is used to convolve the feature map of the current level to obtain a convolved feature map, fuse the convolved feature map with the feature map of the next level through scale transformation, and convolve the fused feature map to obtain a fused feature map; An updating module, configured to use the fused feature map as a new convolved feature map, and determine a new next-level feature map based on the multi-level feature map, where the level of the new next-level feature map is lower than the current next-level feature map; A second processing module is configured to return to the step of fusing the convolved feature map and the next-level feature map by scaling until no new next-level feature map exists in the multi-level feature map, and obtain a fused multi-level feature map based on the initially obtained convolved feature map and the fused feature map obtained each time; A detection module is used to perform target frame detection on each layer of the fused multi-level feature map according to the trained target detection network to obtain a target detection result; The detection module is further used to obtain training samples and an initial target prediction network, optimize the initial target prediction network according to the training samples to obtain a target detection network, perform sparsification training on the target detection network to obtain a sparse target detection network, obtain the gamma value of the normalization layer in all convolutional layers in the sparse target detection network, perform channel pruning on all convolutional layers in the sparse target detection network according to the channel pruning threshold and the gamma value to obtain a channel-pruned network, obtain the current number of channels of all convolutional layers in the channel-pruned network, perform convolutional layer pruning on all convolutional layers in the channel-pruned network according to the current number of channels and the convolutional layer pruning threshold to obtain a pruned target detection network, and retrain the pruned target detection network to obtain a trained target detection network based on the retrained pruned target detection network.

7. The device according to claim 6, characterized in that The fusion module is also used to sample and dimensionally transform the convolved feature map in sequence to obtain a transformed feature map, fuse the transformed feature map and the next-level feature map by direct addition, and convolve the fused feature map to obtain a fused feature map.

8. The device according to claim 6, characterized in that The fusion module is further used to convolve the current level feature map through a preset preprocessing convolutional network to obtain a convolved feature map, wherein the preprocessing convolutional network includes at least two convolution kernels of different sizes.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Target detection method and device and storage medium

    CN111079623A