A target detection model training method, a target detection method, a device and equipment

CN122597906APending Publication Date: 2026-08-18CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510141415.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本申请提供一种目标检测模型训练方法、目标检测方法、装置及设备,可以解决现有技术中存在的DINO算法的查询效果不佳的技术问题

Benefits of technology

[0044] The process involves obtaining multi-scale feature maps of training images based on the target detection model to be trained, encoding these multi-scale feature maps to obtain encoded features, generating noise samples, decoding the noise samples and encoded features based on the query vector, outputting multiple candidate prediction boxes based on the decoded features, matching all candidate prediction boxes with the ground truth of the training images, calculating the loss, and then training the target detection model based on the loss to obtain an optimized target detection model for target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597906A_ABST
    Figure CN122597906A_ABST
Patent Text Reader

Abstract

A target detection model training method, a target detection method, a device and equipment, relate to the technical field of computer vision algorithms. The above method comprises: obtaining a multi-scale feature map of a training image according to a target detection model to be trained, and encoding the multi-scale feature map to obtain encoded features; generating noise samples, decoding the noise samples and the encoded features based on a query vector, and outputting a plurality of candidate prediction boxes according to the obtained decoded features; the query vector is obtained according to anchor box parameters and multi-scale feature vectors of the multi-scale feature map, and the anchor box parameters include at least one learnable point coordinate generated based on the multi-scale feature vectors; all candidate prediction boxes are matched with the true value of the training image, and the loss is calculated, and then the target detection model is trained based on the loss. The present application realizes the combination of features and the positions reflecting the features, and increases key point sampling, thereby realizing better query effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision algorithm technology, specifically to a target detection model training method, target detection method, apparatus and equipment. Background Technology

[0002] Currently, object detection is a fundamental task in computer vision. DETR (Detection Transformer) is a novel object detection algorithm based on Transformer, and the DETR series of algorithms has outperformed classic 2D detectors (such as Faster-RCNN). Unlike previous 2D detectors, DETR models object detection as an ensemble prediction task and assigns labels through bipartite graph matching. Therefore, DETR simplifies the design process of object detection algorithms and achieves end-to-end detection.

[0003] Among related technologies, the DINO (DETR with Improved DeNoising Anchor Boxes) algorithm, while retaining the basic architecture of DETR, adds several optimized structures and is one of the best 2D detection algorithms currently available. The query mechanism of the DINO algorithm generally initializes a query, which is a high-dimensional feature obtained by processing randomly initialized anchor boxes through an MLP (Multilayer Perceptron). The randomly initialized anchor boxes are represented by their width and height (l, w) and center point coordinates (x, y).

[0004] However, the DINO algorithm's query mechanism only includes anchor boxes, resulting in low query accuracy. Summary of the Invention

[0005] This application provides a method for training an object detection model, an object detection method, an apparatus, and a device, which can solve the technical problem of poor query performance of the DINO algorithm in the prior art.

[0006] Firstly, this application provides a method for training an object detection model, the method comprising:

[0007] The multi-scale feature maps of the training images are obtained based on the target detection model to be trained, and the multi-scale feature maps are encoded to obtain the encoded features.

[0008] Noise samples are generated, and based on the query vector, the noise samples and the encoded features are decoded, and multiple candidate prediction boxes are output according to the decoded features. The query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map. The anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector.

[0009] All candidate predicted boxes are matched with the ground truth values ​​of the training images, and the loss is calculated. The object detection model is then trained based on the loss.

[0010] In conjunction with the first aspect, in one embodiment, before decoding the aforementioned noise samples and the aforementioned encoded features, the method further includes:

[0011] The anchor frame position vector is obtained based on the above anchor frame parameters, and a high-dimensional position query vector is generated based on the above position vector.

[0012] The above multi-scale feature vectors are interacted with the above position vectors to obtain instance feature vectors;

[0013] Based on the example feature vector and the high-dimensional location query vector, the query vector is obtained.

[0014] In conjunction with the first aspect, in one implementation, the number of learnable points is determined based on the number of key points of the target;

[0015] The above anchor frame parameters also include the coordinates of at least one fixed point, as well as the length and width of the anchor frame.

[0016] In conjunction with the first aspect, in one implementation, matching all candidate prediction boxes with the ground truth values ​​of the aforementioned training images specifically includes:

[0017] The above truth values ​​are divided into truth boxes and ignore regions;

[0018] Obtain the intersection-union ratio (IUCN) of all candidate predicted bounding boxes and ignored regions;

[0019] Delete candidate predicted boxes with an intersection-union ratio greater than a threshold, and match the remaining candidate predicted boxes with multiple ground truth boxes.

[0020] In conjunction with the first aspect, in one implementation, encoding the aforementioned multi-scale feature map specifically includes:

[0021] A multi-layered cascaded encoder is constructed to encode the above multi-scale feature maps;

[0022] For the first M training cycles, the output features of each layer are concatenated to obtain the encoded features of each training cycle.

[0023] For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle.

[0024] The above M is a preset percentage of the preset number of training cycles.

[0025] In conjunction with the first aspect, in one implementation, encoding the aforementioned multi-scale feature map specifically includes:

[0026] A multi-layered cascaded encoder is constructed to encode the above multi-scale feature maps;

[0027] For the first M training cycles, the output features of each layer are concatenated, mapped through a linear layer, and then standardized to obtain the encoded features of each training cycle.

[0028] For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle.

[0029] The above M is a preset percentage of the preset number of training cycles.

[0030] Secondly, this application provides an object detection model training device, which includes:

[0031] The encoding module is used to obtain multi-scale feature maps of training images based on the target detection model to be trained, and to encode the multi-scale feature maps to obtain encoded features.

[0032] The output module is used to generate noise samples, decode the noise samples and the encoded features based on the query vector, and output multiple candidate prediction boxes according to the obtained decoded features; the query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map, and the anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector.

[0033] The training module is used to match all candidate prediction boxes with the ground truth values ​​of the training images and calculate the loss, and then train the object detection model based on the loss.

[0034] Thirdly, this application provides a target detection method, which includes:

[0035] Acquire the image to be detected;

[0036] The target detection model trained using the above-mentioned target detection model training method is used to detect the above-mentioned image to obtain the detected target.

[0037] Fourthly, this application provides a target detection device, which includes:

[0038] The acquisition module is used to acquire the image to be detected;

[0039] The detection module is used to detect the target in the image to be detected by using the target detection model trained by the above-mentioned target detection model training method.

[0040] Fifthly, this application provides an electronic device, which includes a processor, a memory, and a target detection model training program or a target detection program stored in the memory and executable by the processor, wherein...

[0041] When the above-mentioned object detection model training program is executed by the above-mentioned processor, the steps of the above-mentioned object detection model training method are implemented.

[0042] When the aforementioned target detection program is executed by the aforementioned processor, it implements the steps of the aforementioned target detection method.

[0043] The beneficial effects of the technical solution provided in this application include:

[0044] The process involves obtaining multi-scale feature maps of training images based on the target detection model to be trained, encoding these multi-scale feature maps to obtain encoded features, generating noise samples, decoding the noise samples and encoded features based on the query vector, outputting multiple candidate prediction boxes based on the decoded features, matching all candidate prediction boxes with the ground truth of the training images, calculating the loss, and then training the target detection model based on the loss to obtain an optimized target detection model for target detection.

[0045] Since the query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map, and the anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector, the query combines the features and the positions reflecting the features, and adds key point sampling to achieve better query results, thus solving the technical problem of poor query results in related technologies. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating an embodiment of the object detection model training method of this application;

[0047] Figure 2 This is a flowchart illustrating another embodiment of the object detection model training method of this application;

[0048] Figure 3 This is the encoder structure according to an embodiment of this application;

[0049] Figure 4This is a flowchart illustrating another encoding method according to an embodiment of this application;

[0050] Figure 5 This is a matching diagram of an embodiment of this application.

[0051] Figure 6 This is a schematic diagram of the functional modules of an embodiment of the target detection model training device of this application. Detailed Implementation

[0052] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0053] Firstly, embodiments of this application provide a method for training an object detection model.

[0054] In one embodiment, reference is made to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the object detection model training method of this application. The above-mentioned object detection model training method includes:

[0055] S1. Obtain multi-scale feature maps of training images based on the target detection model to be trained, and encode the multi-scale feature maps to obtain encoded features;

[0056] S2. Generate noise samples, decode the noise samples and the encoded features based on the query vector, and output multiple candidate prediction boxes based on the obtained decoded features; wherein, the query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map, and the anchor box parameters include at least one learnable point coordinate generated based on the multi-scale feature vector.

[0057] S3. Match all candidate prediction boxes with the ground truth values ​​of the training images above, calculate the loss, and then train the object detection model based on the loss.

[0058] In this embodiment, multi-scale feature maps of training images are obtained based on the target detection model to be trained, and the multi-scale feature maps are encoded to obtain encoded features. Then, noise samples are generated, and the noise samples and encoded features are decoded based on the query vector. Multiple candidate prediction boxes are output based on the decoded features. Finally, all candidate prediction boxes are matched with the ground truth of the training images, and the loss is calculated. The target detection model is then trained based on the loss to obtain an optimized target detection model for target detection.

[0059] Since the query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map, and the anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector, the query combines the features and the positions reflecting the features, and adds key point sampling to achieve better query results, thus solving the technical problem of poor query results in related technologies.

[0060] Based on the above embodiments, in this embodiment, before decoding the noise samples and the encoded features in step S2, the following further step is included:

[0061] First, obtain the anchor frame position vector based on the above anchor frame parameters, and then generate a high-dimensional position query vector based on the above position vector;

[0062] Then, the multi-scale feature vectors and the position vectors are interacted to obtain the instance feature vectors;

[0063] Finally, based on the above instance feature vector and high-dimensional location query vector, the above query vector is obtained.

[0064] Furthermore, in one embodiment, the number of learnable points is determined based on the number of key points of the target.

[0065] Preferably, the number of learnable points is 1-4, and the same as the number of key points of the target.

[0066] In this embodiment, the aforementioned learnable points are generated by passing multi-scale feature vectors through a learnable MLP network.

[0067] Furthermore, in this embodiment, the anchor frame parameters also include at least one fixed point coordinate, as well as the length and width of the anchor frame.

[0068] In this embodiment, by modifying the anchor frame parameter structure, sampling of key areas can be increased.

[0069] Optionally, the number of fixed points mentioned above can be 1-5. When there is only one fixed point, it is the center point of the anchor frame. When there are more than one fixed point, the fixed points include the center point of the anchor frame and the midpoints of some edges.

[0070] In this embodiment, there are 5 fixed points, namely the center point of the anchor frame and the midpoints of the four sides of the anchor frame.

[0071] In this embodiment, through the design of anchor+featurer, the anchor box is still a rectangle, but it does not adopt the fixed x, y, w, l form. Instead, it is composed of 5 fixed points and 4 learnable points, as well as w and l. The instance feature vector is obtained through the initial anchor and multi-scale features and is continuously updated to realize implicit and explicit queries, thereby achieving better query results.

[0072] Furthermore, in one embodiment, step S3 above, matching all candidate prediction boxes with the ground truth values ​​of the training images, specifically includes:

[0073] First, the above truth values ​​are divided into truth boxes and ignore regions.

[0074] Then, obtain the intersection-union ratio of all candidate predicted boxes and the ignored regions.

[0075] Finally, candidate predicted boxes with an intersection-union ratio greater than the threshold are deleted, and the remaining candidate predicted boxes are matched with multiple ground truth boxes.

[0076] In this embodiment, by dividing the ground truth into ground truth boxes and ignored regions, and redesigning the matching process between candidate prediction boxes and ground truth, it is possible to ignore target boxes in the ground truth that should be ignored, so that they do not affect the gradient, and optimize the problem of unstable bipartite graph matching, thereby improving the detection effect of distant targets and severely occluded targets.

[0077] Furthermore, in one embodiment, the encoding of the multi-scale feature map in step S1 specifically includes:

[0078] First, a multi-layered cascaded encoder is constructed to encode the aforementioned multi-scale feature maps;

[0079] Then, for the first M training cycles, the output features of each layer are concatenated in each training cycle to obtain the encoded features of the current training cycle; for the remaining training cycles, the output features of the last level are used as the encoded features of the current training cycle.

[0080] Where M is a preset percentage of the preset number of training cycles.

[0081] Optionally, M is 20%-25% of the preset number of training cycles.

[0082] Preferably, a six-layer cascaded encoder is constructed to encode multi-scale feature maps.

[0083] In this embodiment, by concatenating the output features of each layer of the six-layer encoder together, the coded features are input into the decoder, making full use of the output features of each layer of the encoder. This is equivalent to increasing the number of supervisions for each encoder, which is conducive to faster convergence.

[0084] Furthermore, in one embodiment, the encoding of the multi-scale feature map in step S1 specifically includes:

[0085] First, a multi-layered cascaded encoder is constructed to encode the aforementioned multi-scale feature maps;

[0086] Then, for the first M training cycles, in each training cycle, the output features of each layer are concatenated, mapped through a linear layer, and standardized to obtain the encoded features of the current training cycle; for the remaining training cycles, the output features of the last level are used as the encoded features of the current training cycle.

[0087] Where M is a preset percentage of the preset number of training cycles.

[0088] Optionally, M is 20%-25% of the preset number of training cycles.

[0089] Preferably, a six-layer cascaded encoder is constructed to encode multi-scale feature maps.

[0090] In this embodiment, by designing a fusion module in the encoder stage, the features output by each layer of the encoder are fused in a dense connection manner, thereby making fuller use of the pre-trained backbone network to increase supervision of each layer in the decoder stage.

[0091] like Figure 2 As shown, specifically, the target detection model training method of this embodiment uses the basic DINO network structure and improves it. The method includes the following steps:

[0092] A1. Feature extraction is performed through the backbone network and the neck network (Backbone+FPN) to extract multi-scale feature maps of the image;

[0093] In this embodiment, the backbone network is ResNet-50, which is pre-trained on the ImageNet dataset; the neck network is FPN (Feature Pyramid Network). The image size input to the backbone network is c*w*h, and the size parameters are represented by letters. After feature extraction by ResNet-50, the features of the three layers S3, S4, and S5 are extracted and input into FPN for multi-scale feature fusion.

[0094] The dimensions of S3, S4, and S5 are [512, 135, 240], [1024, 68, 120], and [2048, 34, 60], respectively. After passing through FPN, the output features maps of four scales are [256, 135, 240], [256, 68, 120], [256, 34, 60], and [256, 17, 30], respectively.

[0095] A2. Perform patching and positional encoding on the multi-scale feature map to obtain encoded features;

[0096] The backbone and neck network use model weights pre-trained on large-scale datasets such as ImageNet, which have good feature extraction capabilities. The encoder part has randomly initialized weights. Therefore, in the early training stage, the main parts affecting the model's feature extraction are the encoder and decoder. In order to accelerate the model's early convergence speed, mainly the convergence speed of the encoder, it is necessary to make full use of the pre-trained backbone and neck network.

[0097] like Figure 3 As shown, in this embodiment, a six-layer cascaded encoder is designed, including encoder 1, encoder 2, encoder 3, encoder 4, encoder 5, and encoder 6. In the early training phase, i.e., the first M training cycles, the six-layer cascaded encoder is used, directly concatenating the output features of each of the six encoder layers and inputting them into the decoder. This utilizes the output features of each layer, effectively increasing the number of supervision cycles for each encoder, which is beneficial for faster convergence.

[0098] like Figure 4 As shown, in this embodiment, another method for inputting the six-layer features into the decoder is also designed, that is, the six-layer output features, namely features 1 to features 6, are concatenated, mapped and standardized through a linear layer, and then input into the decoder.

[0099] A3. Generate noise samples, decode the encoded features, and output candidate prediction boxes based on the decoded features;

[0100] The decoder takes as input the encoded features generated in A2 and the pseudo-samples (noise samples) generated by the denoising module, and outputs the top K predicted candidate boxes, which include position and category.

[0101] Optionally, the denoising module follows the original DINO design, generating two positive and negative sample denoising groups based on the ground truth. Each positive and negative sample denoising group generates 2N positive and negative sample noises based on N ground truths. For positive samples, a smaller noise is typically used to make it closer to the ground truth. Conversely, for negative samples, a larger noise is used, but it will not exceed a certain range to avoid the contrastive denoising learning being affected by distant detection boxes.

[0102] In this embodiment, the query mechanism is redesigned, and the query and positive / negative sample noise are cross-attentioned together in the decoder. Finally, during loss calculation, the denoising group performs bipartite graph matching and loss calculation separately.

[0103] In some embodiments, the query mechanism simply initializes a query, which is a high-dimensional feature obtained by passing a randomly initialized anchor box through an MLP. The randomly initialized anchor box is represented by its length, width, and center point coordinates (i.e., l, w, x, y).

[0104] The query mechanism in this embodiment redesigns the anchor box, which is represented by five fixed points, four learnable points, and a length and width. The five fixed points are the center point and the center points of the four sides of the anchor box. The four learnable points are generated by the multi-scale feature vector (i.e., the feature output by backbone + FPN) and the learnable MLP network. Each point is represented by two coordinates, with the length and width each represented by a number. Therefore, the new anchor box has a total of (5+4)*2+1+1=20 positions. q anchor boxes are initialized according to the clustering algorithm, which is a q*20 vector, i.e., a position vector. Then, the position vector is processed by an MLP to obtain a high-dimensional q*512 vector. This q*512 vector serves as the implicit query part of the query, i.e., the high-dimensional position query vector. In addition, the above position vector interacts with the feature output by backbone + FPN to obtain the corresponding instance feature vector. This instance feature vector and the q*512 vector together constitute the query. The above query method includes features, the location reflecting the features, and key point sampling. In addition, the instance feature vector can also serve as the basis for time series development, and the query can be input into the next time step for reuse.

[0105] A4. Match the candidate predicted boxes with the ground truth and calculate the loss;

[0106] Specifically, for the top K predicted candidate boxes output in step A3, the top q candidate predicted boxes can be selected and matched with the ground truth. Optionally, K is greater than q.

[0107] In some embodiments, if a predicted bounding box matches a ground truth bounding box, it is considered a correct detection box (TP); if a predicted bounding box does not match a ground truth bounding box, it is considered a false detection box (FP). After matching, positive sample loss is calculated for detection boxes that match ground truth bounding boxes, and negative sample loss is calculated for those that do not. The positive and negative sample losses are weighted and used as part of the final loss. However, this approach has a drawback: the predicted bounding box must be either background or target data.

[0108] Furthermore, during the matching process, the Intersection over Union (IoU) is calculated between each ground truth bounding box and each predicted bounding box. Then, all IoU scores are sorted, and each ground truth bounding box sequentially selects a predicted bounding box with the highest IoU score. If that predicted bounding box has already been matched by another ground truth bounding box, the predicted bounding box with the second highest score is selected, and so on, ensuring that each ground truth bounding box matches at least one predicted bounding box (usually the number of ground truth bounding boxes is less than the number of predicted bounding boxes). The remaining predicted bounding boxes are then considered background boxes. This one-to-one matching method, where a predicted bounding box matches at most one ground truth bounding box and a ground truth bounding box is matched at most one predicted bounding box, does not consider an IoU threshold but rather matches based on score sorting.

[0109] like Figure 5 As shown, this embodiment designs a new matching method, which divides the ground truth into non-negligible ground truth boxes and negligible ground truth boxes, with the negligible ground truth boxes being the ignored regions. If a predicted box matches an ignored region, the predicted box is not subject to loss calculation. This allows the model to learn a portion of the region without treating it as either background or target, thus avoiding the calculation of loss for that portion. In other words, the model is not allowed to learn this part of the content, but only to learn more obvious features. This adds the possibility of "negative-zero and non-1" to the original "either 0 or 1" matching method, making the model's learning more flexible.

[0110] Furthermore, before performing the above matching, the Intersection over Union (IoU) between all candidate predicted boxes and the ignored regions can be calculated. If the IoU is greater than the threshold s, the candidate predicted box is matched with the ignored region, resulting in multiple candidate predicted boxes being matched with the ignored regions. Then, candidate predicted boxes with an IoU greater than the threshold are deleted, and the remaining candidate predicted boxes are matched with the ground truth boxes according to the sorting method. Finally, a one-to-one match between the ground truth boxes and the candidate predicted boxes is obtained, while the match between the predicted boxes and the ignored regions is many-to-one.

[0111] In this embodiment, the above matching method enables the function of not learning the ignored regions.

[0112] A5. Train and update the object detection model using the training loss until the final object detection model is obtained.

[0113] Specifically, when calculating the loss, the loss for positive and negative samples and the loss for the pseudo-true value group in the denoising training part are calculated separately. The loss calculation function for the pseudo-true value group loss is as follows:

[0114]

[0115] Where y is the set of pseudo-truth values. Let c be the set of predicted values. b represents the category of false true values ​​and the category of predicted values, respectively. These represent the false truth box and the predicted box, respectively; BCE is the cross-entropy loss, L... box This is the L1 loss.

[0116] The loss function for positive and negative sample loss is:

[0117]

[0118] The classification loss is given by formula 1 above, L. cls For class loss, L box For L1 loss, L giou For giou loss.

[0119] The target detection model is supervised by calculating the above loss function, and the final target detection model is obtained after convergence.

[0120] The training method in this embodiment obtains high-quality instance features by sampling multiple fixed points of the anchor box and multiple learnable points, and optimizes the process of the Hungarian matching algorithm, thereby ignoring the target boxes that should be ignored in the ground truth, further improving the convergence speed of the DINO algorithm and significantly improving its detection accuracy, especially for the detection of small targets and severely occluded targets, where its indicators show a significant increase.

[0121] Secondly, embodiments of this application also provide a target detection model training device.

[0122] In one embodiment, reference is made to Figure 6 , Figure 6 This is a functional block diagram of an embodiment of the object detection model training device of this application. The object detection model training device includes an encoding module, an output module, and a training module.

[0123] The above-mentioned encoding module is used to obtain multi-scale feature maps of training images based on the target detection model to be trained, and to encode the multi-scale feature maps to obtain encoded features.

[0124] The output module generates noise samples, decodes the noise samples and encoded features based on the query vector, and outputs multiple candidate prediction boxes based on the decoded features. The query vector is obtained based on anchor box parameters and multi-scale feature vectors of the multi-scale feature map. The anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vectors.

[0125] The training module described above is used to match all candidate prediction boxes with the ground truth values ​​of the training images and calculate the loss, and then train the object detection model based on the loss.

[0126] Furthermore, in one embodiment, the above-mentioned target detection model training device further includes a query vector module, which is used for:

[0127] The anchor frame position vector is obtained based on the above anchor frame parameters, and a high-dimensional position query vector is generated based on the above position vector.

[0128] The above multi-scale feature vectors are interacted with the above position vectors to obtain instance feature vectors;

[0129] Based on the example feature vector and the high-dimensional location query vector, the query vector is obtained.

[0130] Furthermore, in one embodiment, the number of learnable points is determined based on the number of key points of the target;

[0131] The above anchor frame parameters also include the coordinates of at least one fixed point, as well as the length and width of the anchor frame.

[0132] Furthermore, in one embodiment, the training module is also used for:

[0133] The above truth values ​​are divided into truth boxes and ignore regions;

[0134] Obtain the intersection-union ratio (IUCN) of all candidate predicted bounding boxes and ignored regions;

[0135] Delete predicted boxes with an intersection-union ratio greater than a threshold, and match the remaining predicted boxes with multiple ground truth boxes.

[0136] Furthermore, in one embodiment, the encoding module is also used for:

[0137] A multi-layered cascaded encoder is constructed to encode the above multi-scale feature maps;

[0138] For the first M training cycles, the output features of each layer are concatenated to obtain the encoded features of each training cycle.

[0139] For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle.

[0140] The above M is a preset percentage of the preset number of training cycles.

[0141] Furthermore, in one embodiment, the encoding module is also used for:

[0142] Encoding the above multi-scale feature maps specifically includes:

[0143] A multi-layer cascaded encoder is constructed to encode the multi-scale feature map;

[0144] For the first M training cycles, the output features of each layer are concatenated, mapped through a linear layer, and then standardized to obtain the encoded features of each training cycle.

[0145] For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle.

[0146] The above M is a preset percentage of the preset number of training cycles.

[0147] The functions of each module in the above-mentioned target detection model training device correspond to the steps in the above-mentioned target detection model training method embodiment, and their functions and implementation processes will not be described in detail here.

[0148] Thirdly, embodiments of this application provide a target detection method, the method comprising:

[0149] First, acquire the image to be detected;

[0150] Then, the target detection model trained using the above target detection model training method is used to detect the above image to obtain the detected target.

[0151] In this embodiment, the target detection model trained using the above-mentioned target detection model training method is used to detect the above-mentioned image to be detected, and the detection accuracy is significantly improved, especially for the detection of small targets and severely occluded targets, the detection index is significantly improved.

[0152] Fourthly, embodiments of this application provide a target detection device, which includes an acquisition module and a detection module.

[0153] The acquisition module described above is used to acquire the image to be detected.

[0154] The aforementioned detection module is used to detect the target in the image to be detected by using the target detection model trained by the aforementioned target detection model training method.

[0155] The functions of each module in the target detection device correspond to the steps in the target detection method embodiment, and their functions and implementation processes will not be described in detail here.

[0156] Fifthly, embodiments of this application provide an electronic device, which may be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.

[0157] In this embodiment of the application, the electronic device may include a processor, a memory, a communication interface, and a communication bus.

[0158] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0159] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces used to interconnect devices within an electronic device, as well as interfaces used to interconnect the electronic device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.

[0160] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0161] The processor can be a general-purpose processor, which can call the object detection model training program or object detection program stored in the memory and execute the object detection model training method or object detection method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the object detection model training program is called can refer to the various embodiments of the object detection model training method of this application, and the method executed when the object detection program is called can refer to the various embodiments of the object detection method of this application, which will not be repeated here.

[0162] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0163] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0164] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0165] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0166] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0167] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0168] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for training an object detection model, characterized in that, The method includes: The multi-scale feature map of the training image is obtained based on the target detection model to be trained, and the multi-scale feature map is encoded to obtain the encoded features; Noise samples are generated, and the noise samples and the encoded features are decoded based on the query vector. Multiple candidate prediction boxes are output based on the obtained decoded features. The query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map. The anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector. All candidate predicted boxes are matched with the ground truth values ​​of the training images, and the loss is calculated. The target detection model is then trained based on the loss.

2. The target detection model training method as described in claim 1, characterized in that, Before decoding the noise samples and the encoded features, the method further includes: The anchor frame position vector is obtained based on the anchor frame parameters, and a high-dimensional position query vector is generated based on the position vector. The multi-scale feature vector is interacted with the position vector to obtain the instance feature vector; The query vector is obtained based on the instance feature vector and the high-dimensional location query vector.

3. The target detection model training method as described in claim 1, characterized in that: The number of learnable points is determined based on the number of key points of the target; The anchor frame parameters also include the coordinates of at least one fixed point, as well as the length and width of the anchor frame.

4. The target detection model training method as described in claim 1, characterized in that, Matching all candidate predicted bounding boxes with the ground truth values ​​of the training images, specifically including: The truth value is divided into a truth box and an ignore region; Obtain the intersection-union ratio (IUCN) of all candidate predicted bounding boxes and ignored regions; Delete candidate predicted boxes with an intersection-union ratio greater than a threshold, and match the remaining candidate predicted boxes with multiple ground truth boxes.

5. The target detection model training method as described in claim 1, characterized in that, Encoding the multi-scale feature map specifically includes: A multi-layer cascaded encoder is constructed to encode the multi-scale feature map; For the first M training cycles, the output features of each layer are concatenated to obtain the encoded features of each training cycle. For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle. M is a preset percentage of the preset number of training cycles.

6. The target detection model training method as described in claim 1, characterized in that, Encoding the multi-scale feature map specifically includes: A multi-layer cascaded encoder is constructed to encode the multi-scale feature map; For the first M training cycles, the output features of each layer are concatenated, mapped through a linear layer, and then standardized to obtain the encoded features of each training cycle. For the remaining training cycles, the output features of the last level are used as the encoding features for each training cycle. M is a preset percentage of the preset number of training cycles.

7. A target detection model training device, characterized in that, The device includes: The encoding module is used to obtain multi-scale feature maps of training images based on the target detection model to be trained, and to encode the multi-scale feature maps to obtain encoded features. The output module is used to generate noise samples, decode the noise samples and the encoded features based on the query vector, and output multiple candidate prediction boxes according to the obtained decoded features; the query vector is obtained based on the anchor box parameters and the multi-scale feature vector of the multi-scale feature map, and the anchor box parameters include the coordinates of at least one learnable point generated based on the multi-scale feature vector; The training module is used to match all candidate prediction boxes with the ground truth values ​​of the training images, calculate the loss, and then train the object detection model based on the loss.

8. A target detection method, characterized in that, It includes: Acquire the image to be detected; The target detection model trained using the target detection model training method according to any one of claims 1-6 is used to detect the image to be detected, thereby obtaining the detected target.

9. A target detection device, characterized in that, The device includes: The acquisition module is used to acquire the image to be detected; The detection module is used to detect the image to be detected using the target detection model trained by the target detection model training method according to any one of claims 1-6, and obtain the detected target.

10. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a target detection model training program or a target detection program stored in the memory and executable by the processor, wherein... When the target detection model training program is executed by the processor, it implements the steps of the target detection model training method as described in any one of claims 1 to 6; When the target detection program is executed by the processor, it implements the steps of the target detection method as described in claim 8.