Target detection method, device and electronic equipment
Through the combination of target scenario classifier and corresponding target detectors, the problem that traditional methods are difficult to deal with multiple traffic scenarios is solved, and efficient target detection in various traffic scenarios is achieved.
Patent Information
- Application Number
- CN202411267056.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-10
AI Technical Summary
Traditional target detection methods are difficult to deal with different traffic scenarios at the same time, resulting in poor detection results in some traffic scenarios.
The target scene classifier determines the scene category to which the image to be detected belongs, and selects the corresponding target detector for object detection according to the scene category.
The detection effect in multiple traffic scenarios is improved, ensuring that excellent detection effect can be maintained in various traffic scenarios.
Smart Images

Figure CN118781475B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a target detection method, device and electronic equipment. Background Art
[0002] The main task of target detection is to determine the type and location of each target in the image to be detected, which is widely used in traffic scenes.
[0003] Traditional target detection methods usually optimize a single target detection model, and then use the optimized target detection model to perform target detection to enhance the target detection model's target detection capability in a complex traffic scenario. For example, by increasing the network depth and width in the target detection model, fusing various network layers, optimizing the loss function, etc., the target detection model's ability to detect targets in a complex traffic scenario can be enhanced.
[0004] However, traffic scenes are numerous and complex, such as urban roads, bridges, highways, elevated roads, tunnels and other different traffic scenes; and the targets involved in traffic are also diverse, such as people, motor vehicles, and non-motor vehicles; at the same time, the height and angle of the camera taking the image will also cause the size of the target to be detected in the image to be greatly different; in addition, the lighting changes in different traffic scenes vary greatly, such as the night lighting conditions in urban traffic scenes are good, but the night lighting conditions in high-speed traffic scenes are poor. Therefore, traffic scenes are complex and diverse, making it difficult to effectively apply the optimization of a single target detection model to target detection in various traffic scenes, making it difficult for the target detection model to cope with various traffic scenes at the same time, resulting in poor detection results in some traffic scenes. Summary of the invention
[0005] This application provides a target detection method, device and electronic device to solve the problem that traditional target detection methods are difficult to deal with different traffic scenes at the same time, resulting in poor detection effect in certain traffic scenes. The specific implementation scheme is as follows:
[0006] In a first aspect, the present application provides a target detection method, the method comprising:
[0007] Based on the target scene classifier, determine the scene category to which the image to be detected belongs;
[0008] Determining an object detector according to the scene category;
[0009] The target detector processes the image to be detected to determine information about the target in the image to be detected.
[0010] Through the above application embodiment, the scene category to which the image to be detected belongs is determined based on the target scene classifier, and then after the scene category is determined, the target detector of the scene category is selected to complete the target detection of the image to be detected, so that in the scene category, an excellent detection effect is achieved. At the same time, the target detector selected for target detection according to the scene category can cope with various traffic scenes at the same time, so as to improve the detection effect in multiple traffic scenes.
[0011] In a possible implementation, determining the scene category to which the image to be detected belongs based on the target scene classifier includes:
[0012] Acquire the image to be detected;
[0013] Extracting a first image feature of the image to be detected through a backbone network of a main detector;
[0014] Inputting the extracted first image features into the target scene classifier to obtain a scene classification result;
[0015] The scene category to which the image to be detected belongs is determined according to the scene classification result.
[0016] Through the above application embodiments, the scene category to which the image to be detected belongs is determined efficiently and accurately, providing a selection basis for the subsequent selection of the target detector, so as to increase the detection effect of the image to be detected.
[0017] In a possible implementation, the first image feature includes first sub-image features of different scales output by multiple network layers, and the step of inputting the extracted first image feature into the target scene classifier to obtain a scene classification result includes:
[0018] Unifying the scales of the first sub-image features into the same scale, to obtain a plurality of second sub-image features with the same scale;
[0019] fusing a plurality of the second sub-image features to obtain a fused feature;
[0020] The fusion features are processed based on the global average pooling layer GAP and the activation function to obtain the scene classification result.
[0021] Through the above-mentioned application embodiment, the scales of each first sub-image feature in the first image feature are unified, so that the scales of multiple sub-image features in the first image feature are the same, so that the multiple sub-image features in the first image feature can be better fused. Then, through the fusion of each sub-image feature (i.e., the second sub-image feature after the scale is unified), it is helpful to improve the target detection model's understanding of the scene, making the scene classification result more accurate. Then, based on the processing of the fused fusion features after GAP and the activation function, the classification features of the scene are extracted, so that the classification features contain more spatial information and semantic information, which is helpful to improve the classification effect of the scene, thereby further making the obtained scene classification results more accurate.
[0022] In a possible implementation, determining the target detector according to the scene category includes:
[0023] If it is determined that the scene category is a simple scene, the first detector used for performing target detection in the simple scene is determined to be the target detector; wherein the simple scene is a scene without small targets, the number of targets is less than a target number threshold, and the occlusion degree of the targets is less than a first occlusion threshold; the small target is a target whose size accounts for a proportion of the image to be detected that is less than a preset proportion;
[0024] If it is determined that the scene category is a small target scene, determining that a second detector for performing target detection in the small target scene is the target detector; wherein the small target scene is a scene in which the small target exists;
[0025] If it is determined that the scene category is a severe occlusion scene, then the third detector used for target detection in the severe occlusion scene is determined to be the target detector; wherein the severe occlusion scene is a scene without the small target and the occlusion degree of the target is greater than the second occlusion threshold.
[0026] Through the above-mentioned application embodiments, when the scene category is a simple scene, the first detector used for target detection in a simple scene is used as the target detector; when the scene category is a small target scene, the second detector used for target detection in a small target scene is used as the target detector; when the scene category is a severe occlusion scene, the third detector used for detection in a severe occlusion scene is used as the target detector, thereby selecting the corresponding target detector according to the corresponding scene category, so that the target detector has the best detection effect on the image to be detected.
[0027] In a possible implementation, the processing the image to be detected by the target detector to determine the information of the target in the image to be detected includes:
[0028] When the target detector is a main detector, the first image feature is processed by the detection head of the main detector to obtain information of the target in the image to be detected; wherein the main detector is the first detector;
[0029] When the target detector is the first sub-detector or the second sub-detector, the second image features of the image to be detected are extracted through the backbone network of the target detector, and the second image features are detected by the detection head of the target detector to obtain information about the target in the image to be detected; wherein the first sub-detector is the second detector, and the second sub-detector is the third detector.
[0030] Through the above-mentioned application embodiment, when the target detector is the main detector (i.e., the first detector), the first image feature that has been extracted by the main detector is directly used, and the first image feature is processed by the detection head of the main detector to obtain the information of the target in the image to be detected, thereby avoiding repeated extraction of image features, thereby improving the efficiency of target detection and realizing the detection of targets in simple scenes.
[0031] At the same time, when the main detector is the first detector used for target detection in simple scenes, compared with the main detector being the second detector and the third detector used for target detection in complex scenes (i.e., small target scenes and severely occluded scenes), the extraction rate of the first image feature in the image to be detected is the highest, thereby making the scene classification result determination speed the fastest, and thus making the scene classification efficiency the highest.
[0032] In addition, when the target detector is the first sub-detector (i.e., the second detector) or the second sub-detector (i.e., the third detector), the second image features of the image to be detected are first extracted by the target detector, and then the second image features are detected by the detection head of the target detector, thereby accurately and efficiently realizing target detection in small target scenes and severely occluded scenes.
[0033] In a possible implementation manner, before determining the scene category to which the image to be detected belongs based on the target scene classifier, the method further includes:
[0034] Freezing the weights of the first detector, the second detector, and the third detector, and training the weights of the scene classifier to obtain the target scene classifier; and
[0035] If it is determined that the scene annotation category of the training image in the training data set is the simple scene, iteratively training the first detector, the second detector, and the third detector simultaneously;
[0036] If it is determined that the scene annotation category is the small object scene, iteratively training the second detector;
[0037] If it is determined that the scene annotation category is the severely occluded scene, the third detector is iteratively trained.
[0038] Through the above application embodiments, various training strategies of scene classifiers and various detectors are flexibly used, thereby maximally ensuring the classification effect of the scene classifier and the detection effect of various detectors.
[0039] In a second aspect, the present application further provides a target detection device, the device comprising:
[0040] A scene classification module is used to determine the scene category to which the image to be detected belongs based on the target scene classifier;
[0041] A selection module, used for determining a target detector according to the scene category;
[0042] The processing module is used to process the image to be detected by the target detector to determine the information of the target in the image to be detected.
[0043] In one possible implementation, the scene classification module is specifically used to obtain the image to be detected; extract the first image feature of the image to be detected through the backbone network of the main detector; input the extracted first image feature into the target scene classifier to obtain a scene classification result; and determine the scene category to which the image to be detected belongs based on the scene classification result.
[0044] In a possible implementation, the scene classification module is specifically used to unify the scales of each of the first sub-image features into the same scale to obtain multiple second sub-image features of the same scale; fuse multiple second sub-image features to obtain a fused feature; and process the fused feature based on a global average pooling layer GAP and an activation function to obtain the scene classification result.
[0045] In a possible implementation, the selection module is specifically used to determine that the first detector used for target detection in the simple scene is the target detector if it is determined that the scene category is a simple scene; wherein the simple scene is a scene without small targets, the number of targets is less than a target number threshold, and the occlusion degree of the targets is less than a first occlusion threshold; the small target is a target whose size accounts for a proportion of the image to be detected that is less than a preset proportion; if the scene category is determined to be a small target scene, then the second detector used for target detection in the small target scene is determined to be the target detector; wherein the small target scene is a scene in which the small target exists; if the scene category is determined to be a severe occlusion scene, then the third detector used for target detection in the severe occlusion scene is determined to be the target detector; wherein the severe occlusion scene is a scene without the small target and the occlusion degree of the target is greater than the second occlusion threshold.
[0046] In one possible implementation, the processing module is specifically used to, when the target detector is a main detector, process the first image feature through the detection head of the main detector to obtain information about the target in the image to be detected; wherein the main detector is the first detector; when the target detector is a first sub-detector or a second sub-detector, extract the second image feature of the image to be detected through the backbone network of the target detector, and detect the second image feature through the detection head of the target detector to obtain information about the target in the image to be detected; wherein the first sub-detector is the second detector, and the second sub-detector is the third detector.
[0047] In a possible implementation, the device further includes a training module, and before the first image feature of the image to be detected is extracted through the main detector, the training module is specifically used to freeze the weights of the first detector, the second detector and the third detector, and train the weights of the scene classifier to obtain the target scene classifier; and if it is determined that the scene annotation category of the training image in the training data set is the simple scene, the first detector, the second detector and the third detector are iteratively trained at the same time; if it is determined that the scene annotation category is the small target scene, the second detector is iteratively trained; if it is determined that the scene annotation category is the severely occluded scene, the third detector is iteratively trained.
[0048] In a third aspect, the present application provides an electronic device, including:
[0049] Memory, used to store computer programs;
[0050] The processor is used to implement the above-mentioned target detection method steps when executing the computer program stored in the memory.
[0051] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned target detection method steps are implemented.
[0052] For each aspect from the second to the fourth aspect and the technical effects that may be achieved by each aspect, please refer to the above description of the technical effects that can be achieved by the first aspect or various possible schemes in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1a A schematic diagram of a small target scene provided in an embodiment of the present application;
[0054] Figure 1b A schematic diagram of a severe occlusion scenario provided in an embodiment of the present application;
[0055] Figure 2 A schematic diagram of a target detection method provided in an embodiment of the present application;
[0056] Figure 3a Schematic diagram 1 of a scaling method provided in an embodiment of the present application;
[0057] Figure 3b Schematic diagram of the scale scaling method provided in the embodiment of the present application Figure 2 ;
[0058] Figure 4 A schematic diagram showing a comparison of the scales before and after the GAP processing provided in the embodiment of the present application;
[0059] Figure 5 A schematic diagram of GAP and FC provided in an embodiment of the present application;
[0060] Figure 6 A schematic diagram of the processing process of the target detection method provided in the embodiment of the present application;
[0061] Figure 7 A schematic diagram of a target detection device provided in an embodiment of the present application;
[0062] Figure 8 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0063] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The specific operating methods in the method embodiments can also be applied to device embodiments or system embodiments. It should be noted that in the description of the present application, "multiple" is understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A is connected to B, which can represent: A is directly connected to B and A is connected to B through C. In addition, in the description of the present application, words such as "first" and "second" are only used to distinguish the purpose of description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0064] The embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0065] Traditional target detection methods usually optimize a single target detection model, and then perform target detection based on the optimized target detection model to enhance the detection effect of the target detection model in a certain traffic scene. However, traffic scenes are various, complex and diverse, which makes it difficult for a single target detection model to be optimized to cope with various traffic scenes at the same time, resulting in poor detection results in some traffic scenes.
[0066] Therefore, in order to cope with various traffic scenarios at the same time, the present application proposes a target detection method, which uses a target scene classifier to determine the scene category to which the image to be detected belongs, and then selects a target detector corresponding to the scene category to determine the information of the target in the image to be detected, thereby completing the target detection, which is conducive to maintaining excellent detection effects in multiple traffic scenarios.
[0067] In the embodiment of the present application, the scene categories may include three categories, namely simple scenes, small target scenes and severe occlusion scenes, but are not limited thereto and may be adjusted according to specific application scenarios.
[0068] Among them, a simple scene is a scene where there are no small targets, the number of targets is less than the target number threshold, and the occlusion degree of the targets is less than the first occlusion threshold; a small target is a target whose size accounts for a proportion of the image to be detected (i.e., the entire image) that is less than a preset proportion (the preset proportion may be 20%). In other words, a simple scene is a scene where there are no small targets, the number of targets is small, and there is no occlusion or slight occlusion between them.
[0069] The size of the target may be the size of a target frame of the target in the image to be detected.
[0070] The occlusion degree of the target may be the occlusion degree of the target between target frames in the image to be detected.
[0071] A small target scene is a scene with a small target. Figure 1a shown.
[0072] A severely occluded scene is a scene where there is no small target and the occlusion degree of the target is greater than the second occlusion threshold. In other words, a severely occluded scene is a scene where there is no small target and the occlusion is severe. The severely occluded scene may also be a scene with many and complex targets. The aforementioned second occlusion threshold may be the same as or different from the first occlusion threshold. For example, a severely occluded scene such as Figure 1b shown.
[0073] The above-mentioned small target scenes and severe occlusion scenes are complex scenes.
[0074] The detector corresponding to the simple scene is the first detector, that is, the first detector is used to detect objects in simple scenes. The first detector can use the detection network with the best detection effect in simple scenes as a detection tool, such as a high-performance detection network as a detection tool, such as the YOLO (You Only Look Once) series of small models, but is not limited to this.
[0075] The detector corresponding to the above-mentioned small target scene is the second detector, that is, the second detector is used for target detection in the small target scene. The second detector can be a detector whose network input size and output size are consistent, or whose output size is slightly smaller than the input size, so as to reduce the feature loss during small target detection and ensure the high resolution of the features, so as to obtain the best detection effect in the small target scene. For example, the second detector can be a target detector based on a U-net neural network (U-Net), or it can be an end-to-end target detector (English: Detection Transformer, abbreviated as DERT) based on a transformer model (Transformer), but it is not limited to this.
[0076] The detector corresponding to the above-mentioned severe occlusion scene is the third detector, that is, the third detector is used for target detection in severe occlusion scenes. The third detector can be selected as the detector with the best detection effect in severe occlusion scenes, such as a detector with occlusion perception ability, such as adding an occlusion perception module to the traditional YOLO series model or the Transformer-based DETR detector, but is not limited to this. The occlusion perception module can be an attention module, such as a self-supervised equivariant attention mechanism (SEAM).
[0077] The first detector, the second detector and the third detector all include at least their corresponding backbone network and detection head. The backbone network can extract the image features of the image to be detected, and the detection head processes the extracted image features to detect the information of the target in the image to be detected, such as the target coordinate frame (i.e., target frame), type, confidence and other information. The target's position can be determined by the target's coordinate frame.
[0078] Therefore, in an embodiment of the present application, the target detection model includes at least a target scene classifier, a first detector, a second detector, and a third detector.
[0079] Reference Figure 2 As shown, a target detection method provided in an embodiment of the present application includes:
[0080] S201, determining the scene category to which the image to be detected belongs based on a target scene classifier.
[0081] In order to cope with various traffic scenarios, the present application first uses a scene classifier to determine the scene category to which the image to be detected belongs.
[0082] Before determining the scene category to which the image to be detected belongs through the scene classifier, the scene classifier needs to be trained first, so that the scene classifier that determines the scene category to which the image to be detected belongs is the target scene classifier with the best classification effect. The first detector, the second detector, and the third detector can also be trained to obtain the trained first detector, the second detector, and the third detector for subsequent use.
[0083] Before training, it is also necessary to determine the training data set. Specifically, first obtain the training image set after image preprocessing. The image preprocessing includes conventional image preprocessing such as image enhancement and noise removal.
[0084] Then, in each training image in the training image set, the target frame and category of the target are marked; then according to the size of the marked target frame and the degree of occlusion between the target frames, the scene category corresponding to the image is determined, and the scene category corresponding to the image is marked, thereby generating a scene classification data set, that is, a training data set.
[0085] After the training data set is determined, the scene classifier and the multiple detectors are trained based on the training data set to obtain an optimal target scene classifier and an optimal first detector, a second detector, and a third detector.
[0086] In the embodiment of the present application, the training of the above-mentioned scene classifier and detector can be trained using the training strategies of minimization training and full network architecture training, but is not limited thereto.
[0087] The above-mentioned minimization training freezes the weights of the first detector, the second detector, and the third detector, and only trains the weight of the scene classifier. The freezing of the weights of the first detector, the second detector, and the third detector is to set the weights of the first detector, the second detector, and the third detector as non-trainable during the training process, that is, these weights remain unchanged during the training process, so as to save computing resources.
[0088] The above-mentioned full network architecture training is to select different detectors for training according to the scene annotation categories of the training images in the training data set. The scene annotation category is the scene category annotated in the training image.
[0089] If it is determined that the scene annotation category of the training image is a simple scene, the first detector, the second detector and the third detector are iteratively trained simultaneously.
[0090] If it is determined that the scene annotation category of the training image is a small object scene, the second detector is iteratively trained.
[0091] If it is determined that the scene annotation category of the training image is a severely occluded scene, the third detector is iteratively trained.
[0092] Through the above method, various training strategies of scene classifiers and various detectors are flexibly used, thereby maximally ensuring the classification effect of the scene classifier and the detection effect of various detectors.
[0093] After training the scene classifier and multiple detectors, the scene category to which the image to be detected belongs is determined by the target scene classifier obtained after training.
[0094] Specifically, first, an image to be detected is acquired; then, the first image feature of the image to be detected is extracted through the backbone network of the main detector; then, the extracted first image feature is input into the target scene classifier, and then the scene category to which the image to be detected belongs is determined according to the scene classification result output by the target scene classifier, thereby efficiently and accurately determining the scene category to which the image to be detected belongs (such as a simple scene, or a small target scene, or a severely occluded scene), providing a selection basis for the subsequent selection of the target detector, so as to increase the detection effect of the image to be detected.
[0095] The above-mentioned image to be detected is an image obtained after the collected original image is preprocessed.
[0096] The first image features include first sub-image features of different scales output by multiple network layers. The multiple network layers are network layers in the backbone network of the main detector. Through the backbone network, the image to be detected can be converted into a more representative feature (i.e., the first image feature) for representation, which is conducive to better scene classification through the feature (i.e., the first image feature).
[0097] The main detector can be any one of the first detector, the second detector, and the third detector. However, optimally, the main detector is the first detector. Compared with the second detector and the third detector corresponding to the complex scene, the first detector corresponding to the simple scene is used as the main detector, and the extraction rate of the first image feature in the image to be detected is the highest, so that the scene classification result is determined the fastest, and the efficiency of scene classification is the highest.
[0098] The above scene classification result can be any one of a first value (such as 0), a second value (such as 1), and a third value (such as 2). The first value indicates that the scene category is a simple scene, the second value indicates that the scene category is a small target scene, and the third value indicates that the scene category is a severe occlusion scene, so that the scene category to which the image to be detected belongs can be determined according to the scene classification result.
[0099] In the embodiment of the present application, in the above-mentioned inputting the extracted first image feature into the target scene classifier to obtain the scene classification result, the specific processing process of the target scene classifier on the first image feature can be as follows:
[0100] First, the scales of the first sub-image features in the first image feature are unified to the same scale to obtain multiple second sub-image features with the same scale. Then, multiple second sub-image features are fused to obtain fused features. The fused features are then processed based on the global average pooling layer (English: Global Average Pooling, abbreviated as GAP) and the activation function to obtain the scene classification result.
[0101] In the embodiment of the present application, the scales of the first sub-image features in the first image feature are unified into the same scale to obtain a plurality of second sub-image features with the same scale. The scale unification can be performed in the following manner:
[0102] Specifically, a target scale is first determined. The target scale is the same scale to which the scales of the first sub-image features are unified. Then, according to the difference between the own scale of each first sub-image feature and the target scale, a corresponding scale scaling method is selected to accurately convert the scale of the corresponding first sub-image feature into the target scale, so that the scales of the first sub-image features are the same, and then a plurality of second sub-image features with the same scales are obtained.
[0103] In the embodiment of the present application, the scale scaling method can be determined according to the difference between the specific scale of the first sub-image feature and the target scale, so as to flexibly adjust the scale scaling method to avoid the loss of important features after the scale of each first sub-image feature is converted to the target scale.
[0104] Exemplarily, the first sub-image features included in the first image feature may be f1 and f2, and the target scale is [H, W, C].
[0105] like Figure 3a As shown, if the size of the first sub-image feature f1 is [h1, w1, c1], and h1<H, w1<W, the first sub-image feature f1 is sent to a 1*1 convolution layer, and then the size of the first sub-image feature f1 is expanded by upsampling (Upsample), so that the height h1 is H, the width w1 is W, and then sent to a 1*1 convolution layer, so that the number of channels c1 of the first sub-image feature f1 is C, and finally the scale of the first sub-image feature f1 is [H, W, C]. The first sub-image feature f1 with a scale of [H, W, C] is the second sub-image feature f'1.
[0106] like Figure 3b As shown, if the size of the first sub-image feature f2 is [h2, w2, c2], and h2>H, w2>W, the first sub-image feature f2 is sent to a 1*1 convolution layer, and then the size of the first sub-image feature f2 is reduced by a 3*3 convolution layer and a maximum pooling layer in sequence, so that the height h2 of the first sub-image feature f2 is H, and the width is W, and then sent to a 1*1 convolution layer, so that the number of channels c2 of the first sub-image feature f2 is C, and finally the scale of the first sub-image feature f2 is [H, W, C]. The first sub-image feature f2 with a scale of [H, W, C] is the second sub-image feature f'2.
[0107] The above-mentioned fusion step of fusing multiple second sub-image features may be to add the feature values at the same position, so that the obtained fusion feature contains the information of each second sub-image feature.
[0108] Exemplarily, the features of the second sub-images are f'1, f'2, f'3, and f'4, and the scales are all [H, W, C]. Then, the feature values of f'1, f'2, f'3, and f'4 at each length h (h is a positive integer less than or equal to H) and each width w (w is a positive integer less than or equal to W) in each channel number c (c is a positive integer less than or equal to C) are added.
[0109] The above-mentioned processing of the fusion features based on GAP and the activation function to obtain the scene classification result can be that the fusion features are first processed by GAP, and then the activation function is used to generate the scene classification result, thereby obtaining the scene classification result. The activation function can be a normalized exponential function (softmax).
[0110] In addition, the above GAP can transform the scale [H, W, C] of the fusion feature into a scale with fixed values for both length and width, which can be 1, that is, [1, 1, C]. Figure 4 As shown in FIG. 1 , assuming that the scale of the fused feature is [6, 6, 3], the scale can be converted to [1, 1, 3] by processing the fused feature through GAP.
[0111] The above GAP also reduces the dimension by replacing the fully connected layer (English: Fully Connected, abbreviated as FC), that is, using the pooling layer to retain the spatial information and / or semantic information extracted by each convolutional layer and pooling layer before GAP, so that the effect is more obvious than that of the fully connected layer in practical applications. Figure 5 shown.
[0112] In addition, GAP removes the restriction on input size and has important applications in convolution visualization, such as Gradient Weighted Class Activation Mapping (Grad-CAM).
[0113] S202, determining a target detector according to the scene category.
[0114] After the scene category to which the image to be detected belongs is determined in step S201, an object detector is determined according to the scene category.
[0115] Specifically, if the scene category is determined to be a simple scene, the first detector used for target detection in the simple scene is determined as the target detector, which is beneficial to improving the detection effect and efficiency of target detection in the simple scene, and avoiding the waste of resources when using complex models (such as the second detector and the third detector) for target detection in the simple scene.
[0116] If the scene category is determined to be a small target scene, then the second detector used for performing target detection in the small target scene is determined to be the target detector, which is beneficial to improving the detection effect and efficiency of target detection in the small target scene.
[0117] If the scene category is determined to be a severe occlusion scene, then the third detector used for performing target detection in the severe occlusion scene is determined to be the target detector, which is beneficial to improving the detection effect and efficiency of performing target detection in the severe occlusion scene.
[0118] Through the above method, the corresponding detector is selected as the target detector according to different scene categories, so that the target detector has the best detection effect on the image to be detected, so as to obtain excellent detection results in various traffic scenes, and at the same time balance the model performance and accuracy.
[0119] In addition, it should be noted that since the output of the target scene classifier is unique, the scene category corresponding to an image to be detected is unique, so an image to be detected will only be matched with one target detector, such as the first detector, or the second detector, or the third detector.
[0120] S203, processing the image to be detected by a target detector to determine information of the target in the image to be detected.
[0121] After determining the target detector according to the scene category in step S202, the target detector processes the image to be detected to determine the target information in the output image. The target information can be the information of all targets in the image to be detected, such as target frame, type, confidence, etc. The position of the target can be determined through the target coordinate frame.
[0122] Specifically, when the target detector is the main detector, the first image feature is directly processed by the detection head of the main detector to obtain information about the target in the image to be detected, thereby avoiding repeated extraction of image features and improving the efficiency of target detection.
[0123] When the target detector is the first sub-detector, that is, the target detector is not the main detector, the third image feature in the image to be detected is first extracted through the backbone network of the first sub-detector, and then the third image feature is processed through the detection head of the first sub-detector to obtain the information of the target in the image to be detected.
[0124] When the target detector is the second sub-detector, that is, the target detector is not the main detector, the fourth image feature in the image to be detected is first extracted through the backbone network of the second sub-detector, and then the fourth image feature is processed through the detection head of the second sub-detector to obtain the information of the target in the image to be detected.
[0125] That is to say, when the target detector is the first sub-detector or the second sub-detector (that is, when the target detector is not the main detector), the second image feature in the image to be detected is first extracted through the backbone network of the target detector (the second image feature is the third image feature or the fourth image feature), and then the second image feature is processed by the detection head of the target detector to obtain the information of the target in the image to be detected, so that the extraction of image features is more accurate, thereby further improving the detection effect of target detection in the corresponding scene.
[0126] It should be noted that when the main detector is the first detector, the first sub-detector may be the second detector, and the second sub-detector may be the third detector. Thus, the above target detector can achieve accurate target detection in simple scenes, small target scenes, and severely occluded scenes.
[0127] Therefore, even if traffic scenes are diverse and complex, different traffic scenes can be classified according to the size of the target, the degree of occlusion of the target frame, etc., and then the above steps S201-S203 can be used for target detection, thereby realizing target detection in various traffic scenes and having excellent detection effects in various traffic scenes. Furthermore, on the premise of improving detection effects and balancing performance, the target detection model's ability to cope with various traffic scene detection tasks can be improved.
[0128] For example, (1) there are many and complex traffic scenes, such as urban roads, bridges, highways, elevated roads, tunnels and other different traffic scenes; (2) the targets involved in traffic are also different, such as people, motor vehicles, and non-motor vehicles; (3) the lighting changes in different traffic scenes vary greatly, such as the night lighting conditions in urban traffic scenes are good, but the night lighting conditions in high-speed traffic scenes are poor; (4) the camera angle and height cause the size of the target to vary greatly, resulting in the diversity and complexity of traffic scenes. In the embodiment of the present application, the traffic scenes in the above (1)-(4) situations can be classified according to the size of the target, the degree of occlusion of the target frame, etc., and then the target detection is performed through the above steps S201-S203, thereby realizing target detection in various traffic scenes and maintaining excellent detection effects in various traffic scenes.
[0129] In summary, the target detection method proposed in this application, in order to cope with various traffic scenarios at the same time, adaptively selects a suitable target detector through scene classification to complete target detection (that is, determines the scene category to which the image to be detected belongs based on the target scene classifier, and then selects the target detector according to the scene category to complete the target detection), which improves the detection effect of the target detection model, achieves a balance between the performance and accuracy of the target detection method, and improves the efficiency of application implementation.
[0130] In addition, scene classifiers (such as target scene classifiers) can share the first image features extracted by the backbone network of the main detector. By fusing the multi-scale first sub-image features in the first image features, the target detection model's understanding of the scene is improved. Then, GAP is used to replace the traditional linear layer to extract classification features, so that the classification features contain more spatial and semantic information, thereby improving the scene classification effect.
[0131] In addition, the target detection method proposed in the embodiment of the present application can flexibly freely combine detectors such as convolution-based detectors and Transformer-based detectors to maximize the target detection capability of the framework, while flexibly using various training strategies for multiple detectors and scene classifiers to maximize the detection effect of the detector and the classification effect of the scene classifier.
[0132] Compared with the traditional single target detection model, the target detection method proposed in the embodiment of the present application has more customized solutions for traffic scenarios, thereby maintaining excellent detection effects in various traffic scenarios.
[0133] The technical solution of this application is further explained below in conjunction with a specific application process.
[0134] like Figure 6 The figure shows a processing diagram of the target detection method. First, in the training module, a training data set transmitted from the data preprocessing module is received. The training images in the training data set are annotated with the scene annotation categories corresponding to the training images. Then, according to the training strategies of minimization training and full network architecture training, the scene classifier and multiple detectors are trained to obtain the target scene classifier with the best scene classification effect, as well as the first detector, the second detector and the third detector with the best detection effect in the corresponding scene category. The first detector and the target scene classifier are then transmitted to the scene classification module so that the scene classification module determines the scene category to which the image to be detected belongs.
[0135] Then, in the scene classification module, firstly, the image to be detected transmitted by the data preprocessing module, as well as the first detector and the target scene classifier transmitted by the training module are received. Secondly, the first image feature of the image to be detected is extracted through the backbone network of the first detector. Then, the first image feature is input into the target scene classifier to obtain the scene classification result. Then, according to the scene classification result, the scene category to which the image to be detected belongs is determined, that is, the scene classification result is converted into the scene category to which the image to be detected belongs. And the scene category is transmitted to the target detector selection module, so that the target detector selection module determines the target detector according to the scene category.
[0136] In the object detector selection module, the scene category of the class transmitted by the scene classification module is first received.
[0137] If the scene category is determined to be a simple scene, the first detector used for performing object detection in the simple scene is determined to be the object detector.
[0138] If the scene category is determined to be a small object scene, the second detector used for performing object detection in the small object scene is determined to be the object detector.
[0139] If the scene category is determined to be a severe occlusion scene, the third detector used for performing object detection in the severe occlusion scene is determined to be the object detector.
[0140] The determined target detector is transmitted to the detection module so that the detection module performs target detection.
[0141] In the detection module, the target detector transmitted by the target detector selection module is first received.
[0142] When the target detector is the first detector, the trained first detector is obtained from the training module, and the first image features extracted by the backbone network of the first detector are obtained from the scene classification module. Then, the first image features are processed by the detection head of the first detector to obtain information about the target in the image to be detected (such as the target frame, target category and confidence information).
[0143] When the target detector is the second detector, the trained second detector is obtained from the training module, and the image to be detected is obtained from the data preprocessing module. Then, the third image feature of the image to be detected is extracted through the backbone network of the second detector, and the third image feature is processed by the detection head of the second detector to obtain the information of the target in the image to be detected.
[0144] When the target detector is the third detector, the trained third detector is obtained from the training module, and the image to be detected is obtained from the data preprocessing module. Then, the fourth image features of the image to be detected are extracted through the backbone network of the third detector, and then the fourth image features are processed by the detection head of the fourth detector to obtain the information of the target in the image to be detected.
[0145] Based on the same inventive concept, an object detection device is also provided in the embodiment of the present application, such as Figure 7 The figure is a schematic diagram of the structure of a target detection device provided by the present application, the device comprising:
[0146] A scene classification module 701 is used to determine the scene category to which the image to be detected belongs based on a target scene classifier;
[0147] A selection module 702 is used to determine a target detector according to the scene category;
[0148] The processing module 703 is used to process the image to be detected by using the target detector to determine the information of the target in the image to be detected.
[0149] In one possible implementation, the scene classification module 701 is specifically used to obtain the image to be detected; extract the first image feature of the image to be detected through the backbone network of the main detector; input the extracted first image feature into the target scene classifier to obtain a scene classification result; and determine the scene category to which the image to be detected belongs based on the scene classification result.
[0150] In a possible implementation, the scene classification module 701 is specifically used to unify the scales of each of the first sub-image features into the same scale to obtain multiple second sub-image features with the same scale; fuse multiple of the second sub-image features to obtain a fused feature; and process the fused feature based on the global average pooling layer GAP and the activation function to obtain the scene classification result.
[0151] In one possible implementation, the selection module 702 is specifically used to, if it is determined that the scene category is a simple scene, determine that the first detector used for target detection in the simple scene is the target detector; wherein the simple scene is a scene without small targets, the number of targets is less than the target number threshold, and the occlusion degree of the target is less than the first occlusion threshold; the small target is a target whose size accounts for a proportion of the image to be detected that is less than a preset proportion; if it is determined that the scene category is a small target scene, determine that the second detector used for target detection in the small target scene is the target detector; wherein the small target scene is a scene in which the small target exists; if it is determined that the scene category is a severe occlusion scene, determine that the third detector used for target detection in the severe occlusion scene is the target detector; wherein the severe occlusion scene is a scene without the small target and the occlusion degree of the target is greater than the second occlusion threshold.
[0152] In one possible implementation, the processing module 703 is specifically used to, when the target detector is a main detector, process the first image feature through the detection head of the main detector to obtain information about the target in the image to be detected; wherein the main detector is the first detector; when the target detector is a first sub-detector or a second sub-detector, extract the second image feature of the image to be detected through the backbone network of the target detector, and detect the second image feature through the detection head of the target detector to obtain information about the target in the image to be detected; wherein the first sub-detector is the second detector, and the second sub-detector is the third detector.
[0153] In a possible implementation, the device further includes a training module, and before the first image feature of the image to be detected is extracted through the main detector, the training module is specifically used to freeze the weights of the first detector, the second detector and the third detector, and train the weights of the scene classifier to obtain the target scene classifier; and if it is determined that the scene annotation category of the training image in the training data set is the simple scene, the first detector, the second detector and the third detector are iteratively trained at the same time; if it is determined that the scene annotation category is the small target scene, the second detector is iteratively trained; if it is determined that the scene annotation category is the severely occluded scene, the third detector is iteratively trained.
[0154] Based on the same inventive concept, an electronic device is also provided in the embodiment of the present application. The electronic device can realize the functions of the aforementioned target detection device. Figure 8 , the electronic equipment includes:
[0155] At least one processor 801, and a memory 802 connected to the at least one processor 801. The specific connection medium between the processor 801 and the memory 802 is not limited in the embodiment of the present application. Figure 8 In the example, the processor 801 and the memory 802 are connected via a bus 800. Figure 8 The connections between other components are shown in bold lines, and are not intended to be limiting. The bus 800 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 8 Only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 801 can also be called a controller, and there is no limitation on the name.
[0156] In the embodiment of the present application, the memory 802 stores instructions that can be executed by at least one processor 801. The at least one processor 801 can execute the target detection method discussed above by executing the instructions stored in the memory 802. The processor 801 can implement Figure 7 The functions of each module in the device shown.
[0157] Among them, the processor 801 is the control center of the device, and can use various interfaces and lines to connect the various parts of the entire control device. By running or executing instructions stored in the memory 802 and calling the data stored in the memory 802, the various functions of the device and process data, the device can be monitored as a whole.
[0158] In one possible design, the processor 801 may include one or more processing units, and the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 801. In some embodiments, the processor 801 and the memory 802 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on separate chips.
[0159] Processor 801 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the target detection method disclosed in the embodiments of the present application can be directly embodied as a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0160] The memory 802 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 802 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (English: Random Access Memory, abbreviated as RAM), static random access memory (English: Static Random Access Memory, abbreviated as SRAM), programmable read-only memory (English: Programmable Read Only Memory, abbreviated as PROM), read-only memory (English: Read Only Memory, abbreviated as ROM), electrically erasable programmable read-only memory (English: Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), magnetic memory, disk, optical disk, etc. The memory 802 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 802 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, used to store program instructions and / or data.
[0161] By designing and programming the processor 801, the code corresponding to the target detection method described in the above embodiment can be fixed into the chip, so that the chip can execute the target detection method when running. Figure 2 The steps of the target detection method of the embodiment shown are as follows: How to design and program the processor 801 is a technique known to those skilled in the art and will not be described in detail here.
[0162] Based on the same inventive concept, an embodiment of the present application further provides a storage medium, which stores computer instructions. When the computer instructions are executed on a computer, the computer executes the target detection method discussed above.
[0163] In some possible implementations, various aspects of the target detection method provided by the present application may also be implemented in the form of a program product, which includes a program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of the target detection method according to various exemplary implementations of the present application described above in this specification.
[0164] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0165] The present application is described with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0166] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0168] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A target detection method, characterized in that: include: Freeze the weights of the first detector, the second detector, and the third detector, and train the weights of the scene classifier to obtain a target scene classifier; as well as If it is determined that the scene annotation category of the training image in the training data set is a simple scene, iteratively training the first detector, the second detector, and the third detector simultaneously; If it is determined that the scene annotation category is a small object scene, iteratively training the second detector; If it is determined that the scene annotation category is a severely occluded scene, iteratively training the third detector; Based on the target scene classifier, determining the scene category to which the image to be detected belongs; Determine a target detector according to the scene category; wherein the target detector is one of the first detector, the second detector, and the third detector; The target detector processes the image to be detected to determine information about the target in the image to be detected.
2. The method according to claim 1, characterized in that The step of determining the scene category to which the image to be detected belongs based on the target scene classifier includes: Acquire the image to be detected; Extracting a first image feature of the image to be detected through a backbone network of a main detector; Inputting the extracted first image features into the target scene classifier to obtain a scene classification result; The scene category to which the image to be detected belongs is determined according to the scene classification result.
3. The method according to claim 2, characterized in that The first image features include first sub-image features of different scales output by multiple network layers, and the extracted first image features are input into the target scene classifier to obtain a scene classification result, including: Unifying the scales of the first sub-image features into the same scale, to obtain a plurality of second sub-image features with the same scale; fusing a plurality of the second sub-image features to obtain a fused feature; The fusion features are processed based on the global average pooling layer GAP and the activation function to obtain the scene classification result.
4. The method according to claim 1, characterized in that The step of determining a target detector according to the scene category includes: If it is determined that the scene category is a simple scene, the first detector used for performing target detection in the simple scene is determined to be the target detector; wherein the simple scene is a scene without small targets, the number of targets is less than a target number threshold, and the occlusion degree of the targets is less than a first occlusion threshold; the small target is a target whose size accounts for a proportion of the image to be detected that is less than a preset proportion; If it is determined that the scene category is a small target scene, determining that a second detector for performing target detection in the small target scene is the target detector; wherein the small target scene is a scene in which the small target exists; If it is determined that the scene category is a severe occlusion scene, then the third detector used for target detection in the severe occlusion scene is determined to be the target detector; wherein the severe occlusion scene is a scene without the small target and the occlusion degree of the target is greater than the second occlusion threshold.
5. The method according to claim 4, characterized in that The step of processing the image to be detected by the target detector to determine information of the target in the image to be detected includes: When the target detector is a main detector, the first image feature is processed by the detection head of the main detector to obtain information of the target in the image to be detected; wherein the main detector is the first detector; When the target detector is the first sub-detector or the second sub-detector, the second image features of the image to be detected are extracted through the backbone network of the target detector, and the second image features are detected by the detection head of the target detector to obtain information about the target in the image to be detected; wherein the first sub-detector is the second detector, and the second sub-detector is the third detector.
6. A target detection device, characterized in that: include: A scene classification module, used to freeze the weights of the first detector, the second detector and the third detector, and train the weights of the scene classifier to obtain a target scene classifier; And if it is determined that the scene annotation category of the training image in the training data set is a simple scene, the first detector, the second detector and the third detector are iteratively trained at the same time; if it is determined that the scene annotation category is a small target scene, the second detector is iteratively trained; if it is determined that the scene annotation category is a severe occlusion scene, the third detector is iteratively trained; based on the target scene classifier, the scene category to which the image to be detected belongs is determined; A selection module, configured to determine a target detector according to the scene category; wherein the target detector is one of the first detector, the second detector, and the third detector; The processing module is used to process the image to be detected by the target detector to determine the information of the target in the image to be detected.
7. The device according to claim 6, characterized in that The selection module is further used to determine that the first detector used for target detection in the simple scene is the target detector if it is determined that the scene category is a simple scene; wherein the simple scene is a scene without small targets, the number of targets is less than a target number threshold, and the occlusion degree of the target is less than a first occlusion threshold; the small target is a target whose size accounts for a proportion of the image to be detected that is less than a preset proportion; If it is determined that the scene category is a small target scene, determining that a second detector for performing target detection in the small target scene is the target detector; wherein the small target scene is a scene in which the small target exists; If it is determined that the scene category is a severe occlusion scene, then the third detector used for target detection in the severe occlusion scene is determined to be the target detector; wherein the severe occlusion scene is a scene without the small target and the occlusion degree of the target is greater than the second occlusion threshold.
8. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to implement the method steps of any one of claims 1 to 5 when executing the computer program stored in the memory.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Road target detection method and device, equipment and vehicle
CN111985378A
Adaptive target detection method based on scene complexity pre-classification
CN114022705A