Target detection method and device, computer equipment and vehicle

By introducing uncertainty quantification method in Transformer object detection network DION, the uncertainty graph based on variance is extracted and model training is assisted, the problem of inaccurate Transformer object detection in complex scenarios is solved, and the accuracy and reliability of detection are improved.

CN120014232APending Publication Date: 2025-05-16CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510087141.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The Transformer-based object detection method is inaccurate in complex, unknown, unfamiliar or out-of-domain scenarios, especially in extreme car accident scenarios, with few sample training, resulting in the perception algorithm outputting unreliable perceptual results in the autonomous driving edge scenario.

Method used

An uncertainty quantization method is introduced into the Transformer object detection network DION, which assists in model training to improve detection accuracy by extracting variance-based uncertainty maps and inputting them into the encoding layer together with feature maps and position encodings.

Benefits of technology

By evaluating the uncertainty of the feature map, more prior information is provided, the direction of model parameter adjustment is improved, and the accuracy and reliability of the overall network detection target are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014232A_ABST
    Figure CN120014232A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle target detection, and discloses a target detection method and device, computer equipment and a vehicle, and the method comprises the steps: obtaining a to-be-detected image, and extracting a feature map of the to-be-detected image; extracting an uncertainty map based on variance from the feature map; converting the position code of the feature map, the uncertainty map and the feature map into a first feature sequence, and inputting the first feature sequence into a coding layer; the middle classification prediction layer processes the second feature sequence output by the coding layer to obtain a plurality of prediction query vectors; processing the second feature sequence based on a decoding layer fusing the prediction query vector and the decoding layer query vector to obtain a third feature sequence; inputting the third feature sequence into a Hungary matching layer, and matching to obtain a target feature sequence; and processing the target feature sequence through the classification detection layer and the regression detection layer to obtain a target position and a target category in the to-be-detected image. According to the target detection method based on the Transform, the accuracy of the target detection method based on the Transform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle target detection, and in particular to a target detection method, device, computer equipment and vehicle. Background Art

[0002] At present, target detection based on images taken by vehicles based on deep learning models is a research hotspot. Performing tasks such as driver reminders and automatic obstacle avoidance based on the detected targets can further improve driving and automation and safety. However, as a black box model, the deep learning model has relatively weak interpretability, and its internal operating principle also has certain uncertainties. With the trend of advanced artificial intelligence algorithms being widely used in target detection, the lightweight and interpretable design of the model and the safety assessment are important research directions for the actual implementation and application of artificial intelligence algorithms. In response to these requirements, related technologies have proposed to apply uncertainty quantification methods to the YOLO series detectors to improve the accuracy of the YOLO series detectors. Uncertainty quantification refers to quantifying the confidence of deep learning model detection through specific indicators to characterize the quality of deep learning model target detection. At the same time, uncertainty quantification indicators can further assist the training of deep learning models to improve the accuracy of deep learning models.

[0003] At present, the target detection algorithm based on the Transformer model has also achieved relatively good results on many large public data sets. At present, the YOLO series target detectors are mainly improved based on the uncertainty quantification method. There is a lack of uncertainty-related research that can be applied to the Transformer target detection algorithm. The existing Transformer target detection method cannot obtain reliable information in some complex, unknown, unfamiliar or out-of-domain scenes. For example, there are fewer image samples in some extreme car accident scenes, and there are fewer target detection samples for such scenes. As a result, the perception algorithm often cannot output reliable perception results in such autonomous driving edge scenes. It can be seen that how to quantify uncertainty for the Transformer target detection method and improve the accuracy of target detection in images is a problem worth studying. Summary of the invention

[0004] In view of this, the present invention provides a target detection method, apparatus, computer equipment and vehicle to solve the problem of inaccurate target detection based on the Transformer target detection method.

[0005] In a first aspect, the present invention provides a target detection method, which includes: obtaining an image to be detected; extracting a feature map of the image to be detected based on a backbone network; extracting a variance-based uncertainty map from the feature map; converting the position code of the feature map, the uncertainty map and the feature map into a first feature sequence and inputting the first feature sequence into a coding layer; obtaining a second feature sequence output by the coding layer, and processing the second feature sequence through an intermediate classification prediction layer to obtain multiple prediction query vectors; initializing a decoding layer query vector; processing the second feature sequence based on a decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence; inputting the third feature sequence into a Hungarian matching layer to obtain a target feature sequence by matching; processing the target feature sequence through a classification detection layer and a regression detection layer respectively to obtain a target position and a target category in the image to be detected.

[0006] According to the above technical means, the present invention is improved on the basis of the latest Transformer target detection network DION (DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, improving the DETR of the denoising anchor box), introducing the uncertainty quantification method into the DION network, and assisting the training of the DION network. The first improvement point of the present invention is carried out on the backbone network in the DION structure. For the feature map slices extracted by the backbone network, the uncertainty map with the same number of feature maps is further extracted from the feature map slices based on the variance sampling method, and then the position encoding of the feature map, the uncertainty map and the feature map are converted into the first feature sequence as the test sample and input into the coding layer for subsequent training and detection steps, and the ability of the discrete degree of the model is evaluated based on the variance, so that the uncertainty of each feature map can be evaluated, and the information is used as the prior information of the model training and detection together with the feature map input model, so that the model knows more information, improves the direction of guiding the adjustment of model parameters, and improves the accuracy of the model, so that the accuracy of the overall network detection target can be accurately improved.

[0007] In some optional embodiments, a variance-based uncertainty map is extracted from a feature map, including: extracting an RGB three-channel image for each feature map; calculating the mean of the RGB three-channel image of each feature map to obtain a mean map corresponding to each feature map; calculating the variance of the RGB three-channel image of each feature map to obtain a variance map corresponding to each feature map; initializing a standard Gaussian distribution matrix; using the standard Gaussian distribution matrix to randomly sample from the mean map and variance map corresponding to each feature map to obtain the same number of random sampling maps as the feature maps; and extracting a number of uncertainty maps from the random sampling maps.

[0008] According to the above technical means, the present invention first extracts the RGB three-channel image from each feature map to perform mean and variance operations to obtain the mean map and variance map of each feature map. After that, the current feature map is randomly sampled for its mean map and variance map using a standard Gaussian distribution matrix to obtain a matrix with variance properties. Each feature map corresponds to a matrix, which can be used as an uncertainty map, thereby realizing the uncertainty measurement of the output features of the backbone network.

[0009] In some optional embodiments, the backbone network is also trained by a first loss function, and the loss value of the first loss function is obtained by calculating the sum of a binary cross entropy loss and a KL divergence. The binary cross entropy loss is used to calculate the loss between the uncertainty map and the corresponding true target image, and the KL divergence is used to calculate the distribution difference value between the distribution of the uncertainty map and the distribution of the corresponding true target image.

[0010] According to the above technical means, the present invention can also use the uncertainty map to optimize the parameters of the backbone network, and specifically create a first loss function. The first loss function is divided into two parts. The first part calculates and obtains the uncertainty map and the real target image corresponding to the uncertainty map. The real target image includes the label coordinates of the real target, and the accuracy of the output features of the backbone network is measured by the loss between the two; the second part uses KL divergence to measure the difference between the probability distribution of the uncertainty map and the probability distribution of the corresponding real target image. The accuracy of the output features of the backbone network is evaluated by the loss function fused by the above two parts, which can guide the parameter adjustment direction of the backbone network.

[0011] In some optional embodiments, the second feature sequence is processed based on a decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence, including: calculating a first uncertainty quantification index based on the distance between each prediction query vector and the target category included in the image to be detected; screening the prediction query vector by the first uncertainty quantification index to obtain an optimized query vector; calculating optimized position information corresponding to the optimized query vector; embedding the optimized position information, the decoding layer query vector, the position encoding of the decoding layer query vector and the uncertainty quantification index into the decoding layer; processing the second feature sequence through the decoding layer to obtain a third feature sequence.

[0012] According to the above technical means, the present invention proposes a second improvement point, which adds a step of screening operation on the prediction query vector between the encoder and the decoder. Specifically, the uncertainty of each prediction query vector is measured by calculating the distance between each prediction query vector and each target category in the image. When the distance between the prediction query vector and each target category in the image is large, it means that the prediction query vector is not close to any target category and its uncertainty is large, thereby realizing the quantification of the uncertainty of the prediction query vector. Based on the quantification method, an optimized query vector with small uncertainty is screened out from the prediction query vector, and the decoding layer query vector initialized with the position information corresponding to the optimized query vector as prior information is input into the decoding layer together, and then the second feature sequence is processed to make the query direction of the decoding layer clearer, and further improve the accuracy of the third feature sequence output by the prediction query vector.

[0013] In some optional embodiments, an uncertainty quantification index is calculated based on the distance between each prediction query vector and the target category included in the image to be detected, including: projecting each prediction query vector to the feature space through a preset mapping matrix; determining the centroid vector of each target category in the feature space; calculating the exponential distance between the projection of each prediction query vector and each centroid vector; and extracting the maximum exponential distance corresponding to each prediction query vector from all the calculated exponential distances as the first uncertainty quantification index corresponding to each prediction query vector.

[0014] In some optional implementations, the prediction query vector is screened by using a first uncertainty quantification index to obtain an optimized query vector, including: sorting the maximum exponential distances corresponding to each prediction query vector from large to small; selecting a preset number of target maximum exponential distances from one side of the maximum value in the sorting result; and determining the prediction query vector corresponding to the target maximum exponential distance as the optimized query vector.

[0015] According to the above technical means, the present invention proposes an uncertainty quantification method for calculating the exponential distance between the prediction query vector and the centroid vector of each target category in the feature space through the radial basis function, wherein the exponential distance is opposite to some commonly used distances, and the larger the exponential distance is, the greater the correlation between the prediction query vector and the target category is, and the higher the credibility is, so that the maximum exponential distance corresponding to each prediction query vector can represent the credibility of each prediction query vector, and the screening of the prediction query vectors with higher credibility of the preset number is realized according to the sorting of credibility, and the accuracy of the prior information of the query vector of the decoding layer is further improved. At the same time, the centroid vector in the feature space has the ability to be adjusted, which is not only convenient for screening the prediction query vector according to the first uncertainty quantification index, but also can continuously adjust the centroid during the model training process to improve the accuracy of the first uncertainty quantification index.

[0016] In some optional embodiments, the third feature sequence is input into the Hungarian matching layer, and matching to obtain the target feature sequence includes: inputting the third feature sequence into the multilayer perceptron, and projecting the third feature sequence into the feature space; determining the centroid vector of each target category in the feature space; calculating the exponential distance between the projection of each predicted feature in the third feature sequence and each centroid vector; extracting the maximum exponential distance corresponding to each predicted feature from all the calculated exponential distances as the second uncertainty quantification index corresponding to each predicted feature; classifying the third feature sequence based on the Softmax classification module to obtain the classification confidence corresponding to each predicted feature; fusing the classification confidence and the second uncertainty quantification index to obtain the third uncertainty quantification index corresponding to each predicted feature in the third feature sequence; when processing the third feature sequence through the Hungarian matching layer, calculating the third uncertainty quantification index in the matching cost to obtain the target feature sequence.

[0017] According to the above technical means, the present invention proposes a third improvement point, which realizes uncertainty quantification for the detection head of the DINO network, and optimizes the matching process of the Hungarian matching layer based on the uncertainty quantification result. Among them, the uncertainty index of each third feature sequence is calculated by radial basis function, and is integrated with the classification confidence output by the traditional classification layer to obtain the credibility of each predicted feature in the third feature sequence, and the credibility of each predicted feature is calculated in the matching cost of the Hungarian matching algorithm. The higher the credibility, the smaller the corresponding matching cost, thereby improving the matching accuracy of the Hungarian matching algorithm, and then improving the accuracy of the final target detection frame position and classification.

[0018] In some optional embodiments, the centroid vector of each target category is determined in the feature space, including: randomly initializing the centroid vector during the first round of training; inputting a third feature sequence into a multilayer perceptron during the training process after the second round, and projecting the third feature sequence into the feature space; calculating the exponential distance between the projection of each predicted feature in the third feature sequence and each centroid vector; calculating the binary cross entropy between each exponential distance and the one-hot encoding of the true category label; and updating the centroid vector with the binary cross entropy minimized.

[0019] According to the above technical means, the present invention also creates a second loss function to optimize the centroid vector in the aforementioned feature space, projects each predicted feature in the third feature sequence output by the decoding layer into the feature space, and calculates the exponential distance between the projection and each centroid vector. The second loss function calculates the binary cross entropy between each exponential distance and the unique hot encoding of the true category label; with the minimum binary cross entropy as the condition, back propagation is used to update the centroid vector, thereby further improving the accuracy of the centroid vector.

[0020] In a second aspect, the present invention provides a target detection device, which includes: an image acquisition module for acquiring an image to be detected; a feature extraction module for extracting a feature map of the image to be detected based on a backbone network; an uncertainty feature extraction module for extracting a variance-based uncertainty map from the feature map; an encoding module for converting the position encoding of the feature map, the uncertainty map and the feature map into a first feature sequence and inputting the sequence into the encoding layer; a prediction query vector generation module for acquiring a second feature sequence output by the encoding layer, and processing the second feature sequence through an intermediate classification prediction layer to obtain multiple prediction query vectors; a random initialization query vector module for initializing a decoding layer query vector; a hybrid query module for processing the second feature sequence based on a decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence; a Hungarian matching module for inputting the third feature sequence into the Hungarian matching layer to match and obtain a target feature sequence; and a target result output module for processing the target feature sequence through a classification detection layer and a regression detection layer, respectively, to obtain a target position and a target category in the image to be detected.

[0021] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing any one of the methods in the first aspect by executing the computer instructions.

[0022] In a fourth aspect, the present invention provides a vehicle in which the computer device provided in the third aspect is installed.

[0023] The technical solution provided by the present invention has the following advantages:

[0024] (1) According to the above technical means, the present invention improves on the latest Transformer target detection network DION (DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, DETR with improved denoising anchor boxes), introduces uncertainty quantification method into DION network, and assists in the training of DION network. The first improvement of the present invention is carried out on the backbone network in the DION structure. For the feature map slices extracted by the backbone network, uncertainty maps with the same number of feature maps are further extracted from the feature map slices based on variance sampling, and then the position encoding of the feature map, the uncertainty map and the feature map are converted into the first feature sequence as test samples and input into the coding layer for subsequent training and detection steps. The ability of the discrete degree of the model is evaluated based on the variance, so that the uncertainty of each feature map can be evaluated. The information is used as the prior information for model training and detection and input into the model together with the feature map, so that the model knows more information, improves the direction of guiding the adjustment of model parameters, and improves the accuracy of the model, so as to accurately improve the accuracy of the overall network detection target.

[0025] (2) According to the above technical means, the present invention first extracts the RGB three-channel image from each feature map and performs mean and variance operations to obtain the mean map and variance map of each feature map. After that, the current feature map is randomly sampled with respect to its mean map and variance map using a standard Gaussian distribution matrix, thereby obtaining a matrix with variance properties. Each feature map corresponds to a matrix, which can be used as an uncertainty map, thereby realizing the measurement of the uncertainty of the output features of the backbone network.

[0026] (3) According to the above technical means, the present invention can also use the uncertainty map to optimize the parameters of the backbone network. Specifically, a first loss function is created. The first loss function is divided into two parts. The first part calculates and obtains the uncertainty map and the real target image corresponding to the uncertainty map. The real target image includes the label coordinates of the real target, and the accuracy of the output features of the backbone network is measured by the loss between the two. The second part uses KL divergence to measure the difference between the probability distribution of the uncertainty map and the probability distribution of the corresponding real target image. The accuracy of the output features of the backbone network is evaluated by the loss function fused by the above two parts, and then the parameter adjustment direction of the backbone network can be guided.

[0027] (4) According to the above technical means, the present invention proposes a second improvement point, which adds a step of screening operation on the prediction query vector between the encoder and the decoder. Specifically, the uncertainty of each prediction query vector is measured by calculating the distance between each prediction query vector and each target category in the image. When the distance between the prediction query vector and each target category in the image is large, it means that the prediction query vector is not close to any target category and its uncertainty is large, thereby realizing the uncertainty quantification of the prediction query vector. Based on the quantization method, an optimized query vector with small uncertainty is screened out from the prediction query vector, and the position information corresponding to the optimized query vector is input into the decoding layer together with the decoding layer query vector initialized by the decoding layer as the prior information, and then the second feature sequence is processed to make the query direction of the decoding layer clearer, thereby further improving the accuracy of the third feature sequence output by the prediction query vector.

[0028] (5) According to the above technical means, the present invention proposes an uncertainty quantification method for calculating the exponential distance between the prediction query vector and the centroid vector of each target category in the feature space through the radial basis function, wherein the exponential distance is opposite to some commonly used distances, and the larger the exponential distance is, the greater the correlation between the prediction query vector and the target category is, and the higher the credibility is, so that the maximum exponential distance corresponding to each prediction query vector can represent the credibility of each prediction query vector, and the prediction query vector with a higher credibility is selected according to the sorting of credibility, further improving the accuracy of the prior information of the query vector in the decoding layer. At the same time, the centroid vector in the feature space has the ability to be adjusted, which is not only convenient for screening the prediction query vector according to the first uncertainty quantification index, but also can continuously adjust the centroid during the model training process to improve the accuracy of the first uncertainty quantification index.

[0029] (6) Based on the above technical means, the present invention proposes a third improvement point, which realizes uncertainty quantification for the detection head of the DINO network, and optimizes the matching process of the Hungarian matching layer based on the uncertainty quantification result. Among them, the uncertainty index of each third feature sequence is calculated by the radial basis function, and is integrated with the classification confidence output by the traditional classification layer to obtain the credibility of each predicted feature in the third feature sequence. The credibility of each predicted feature is calculated in the matching cost of the Hungarian matching algorithm. The higher the credibility, the smaller the corresponding matching cost, thereby improving the matching accuracy of the Hungarian matching algorithm, and then improving the accuracy of the final target detection frame position and classification.

[0030] (7) According to the above technical means, the present invention also creates a second loss function for optimizing the centroid vector in the aforementioned feature space, projects each predicted feature in the third feature sequence output by the decoding layer into the feature space, and calculates the exponential distance between the projection and each centroid vector. The second loss function calculates the binary cross entropy between each exponential distance and the unique-hot encoding of the true category label; with the minimum binary cross entropy as the condition, back propagation is used to update the centroid vector, thereby further improving the accuracy of the centroid vector. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 It is a structural diagram of the DINO network according to the related technology;

[0033] Figure 2 is a schematic flow chart of a target detection method according to an embodiment of the present invention;

[0034] Figure 3 is a schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0035] Figure 4 is another schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0036] Figure 5 is another schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0037] Figure 6 is another schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0038] Figure 7 is another schematic diagram of a network structure of a target detection method according to an embodiment of the present invention;

[0039] Figure 8 is a schematic diagram of detection effect of a target detection method according to an embodiment of the present invention;

[0040] Fig. 9 is a schematic structural diagram of a target detection device according to an embodiment of the present invention;

[0041] Fig.10 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0043] At present, the DINO (DETR with Improved DeNoising Anchor Boxes for End-to-EndObject Detection) network is a relatively new Transformer target detection network, which has a good performance in the field of vehicle target detection. However, reliable information is often not obtained in some complex, unknown, unfamiliar or out-of-domain scenes, such as some extreme scenes of car accidents. There are fewer training samples for such scenes, so the uncertainty of the detected target results is relatively large, and the detection error rate may be higher, which is not conducive to the safety of autonomous driving.

[0044] Based on this, if uncertainty quantification can be introduced into the DINO network, it will not only be able to judge the quality of the DINO network, but also increase the prior information in the DINO network training and actual prediction process, thereby making the prediction results more accurate.

[0045] Before introducing the technical solution provided by the present invention, in order to enable readers to more fully understand the solution, the structure and detection principle of the DINO network are briefly described below.

[0046] like Figure 1 As shown in Figure 1, it is the network structure of the DINO network, which includes a backbone network, an encoding layer, an intermediate classification prediction layer, a decoding layer, a Hungarian matching layer ( Figure 1 in the matching), classification detection layer and regression detection layer (the classification detection layer and regression detection layer are in Figure 1 After the match in Figure 1Not shown in the figure). For an image to be detected, the complete image must first be divided into several small images, for example, a picture is divided into 81 small images of 9*9 in total, and each small image corresponds to a position code, which is used to indicate the position of the small image in the complete image. After that, the features of each small image are extracted through the backbone network to obtain multiple feature maps, and then each feature map is serialized to obtain a first feature sequence, and the first feature sequence is input into the encoding layer together with the position code corresponding to each feature map. The encoding layer is processed by the self-attention mechanism to output the encoded feature sequence to obtain the second feature sequence. Each vector in the second feature sequence is used to indicate which feature maps may contain targets and need special attention, and also indicates the position where the target is most likely to appear in the feature map. Afterwards, a certain number of query vectors (query) will be randomly initialized at the entrance of the decoding layer. The function of the query vector is to query each vector in the encoded feature sequence. Each query vector is accompanied by a position code to limit the image area that the query vector pays more attention to. The query vector, the position code of the query vector and the second feature sequence are input into the decoding layer together, so that the decoding layer can find the feature vector that is most likely to include the target according to the prompt of the encoding layer after certain processing, and output it in the form of a sequence to obtain the third feature sequence. Among them, the position code of the filtered query vector is also input into the decoding layer together with the initialized query vector. The position code of the filtered query vector is obtained by making a rough prediction of the second feature sequence through the intermediate classification prediction layer, and roughly predicting which vectors in the second feature sequence include the target, thereby obtaining a number of predicted query vectors (predicted query). The position codes corresponding to these predicted query vectors can tell the decoding layer which query vectors in the initialized query vectors are better, thereby increasing the prior information of the decoding layer, making the decoding layer analysis more sufficient, and the output third feature sequence more accurate. The third feature sequence is usually used to represent hundreds of target detection frames, and it is necessary to find the frame that is closest to the real target in the picture, so as to use the Hungarian matching algorithm to process and match the target feature sequence. The purpose is to match the frame that is closest to the real target in the picture from the third feature sequence. At this time, the target feature sequence is not in the form of a frame, and needs to go through the final classification detection layer and regression detection layer to determine the coordinates, size and category of the frame.

[0047] The above process describes the processing logic of the DINO network when performing target detection. The calculation process, number of convolutions, and network connection method within the backbone network, encoding layer, intermediate classification prediction layer, decoding layer, Hungarian matching layer, classification detection layer, and regression detection layer are all prior arts and will not be described in detail in the embodiments of the present invention. The implementation principle can refer to the basic knowledge materials of the DINO network and the DETR network.

[0048] The technical solution provided by the present invention has made three improvements on the basis of the above-mentioned DINO network, and has performed uncertainty quantification in the backbone network part, between the encoder and the decoder, and between the decoding layer and the Hungarian matching layer, and has input the results of the uncertainty quantification into the network structure to increase the prior information of the model, thereby improving the quality of model training and improving the accuracy of the model detection target. The specific improvement scheme is as follows.

[0049] According to an embodiment of the present invention, an embodiment of a target detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0050] In this embodiment, a target detection method is provided. Figure 2 : is a flow chart of a target detection method according to an embodiment of the present invention, the flow chart includes the following steps:

[0051] Step S201, obtaining an image to be detected;

[0052] Step S202, extracting a feature map of the image to be detected based on the backbone network;

[0053] Step S203, extracting a variance-based uncertainty map from the feature map;

[0054] Step S204, converting the position code of the feature map, the uncertainty map and the feature map into a first feature sequence and inputting the first feature sequence into a coding layer;

[0055] Step S205, obtaining a second feature sequence output by the encoding layer, and processing the second feature sequence through the intermediate classification prediction layer to obtain multiple prediction query vectors;

[0056] Step S206, initializing a decoding layer query vector;

[0057] Step S207, processing the second feature sequence based on the decoding layer that combines the prediction query vector and the decoding layer query vector to obtain a third feature sequence;

[0058] Step S208, inputting the third feature sequence into the Hungarian matching layer to obtain a target feature sequence through matching;

[0059] Step S209, processing the target feature sequence through the classification detection layer and the regression detection layer respectively to obtain the target position and target category in the image to be detected.

[0060] Specifically, the embodiment of the present invention is first described with respect to the first improvement point, such as Figure 3As shown, the first improvement point is located in the backbone network of the DINO network. The above steps S201 to S209 are basically the same as the overall processing flow of the DINO network. The same parts will not be explained again. The relevant principles can refer to the content of the aforementioned DINO network. The embodiment of the present invention is mainly for further processing of the feature map extracted by the backbone network, and extracting a variance-based uncertainty map from the feature map. For example, one way is that each feature map includes three RGB color channels, which includes three feature sub-maps. According to the variance of the elements at corresponding positions in the three feature sub-maps, a processed feature map can be obtained, and the feature map can be used as a variance-based uncertainty map.

[0061] Through the embodiments of the present invention, for the feature map slices extracted by the backbone network, uncertainty maps with the same number as the feature maps are further extracted from the feature map slices based on the variance sampling method, and then the position coding of the feature map, the uncertainty map and the feature map are converted into a first feature sequence as test samples and input into the coding layer for subsequent training and detection steps. The ability of the discrete degree of the model is evaluated based on the variance, so that the uncertainty of each feature map can be evaluated. This information is used as prior information for model training and detection and is input into the model together with the feature map, so that the model knows more information, improves the direction of guiding the adjustment of model parameters, improves the model accuracy, and thus can accurately improve the accuracy of the overall network detection target.

[0062] In some optional implementation manners, the above step S203 specifically includes:

[0063] Step a1, extracting RGB three-channel images for each feature map;

[0064] Step a2, calculating the mean of the RGB three-channel image of each feature map to obtain the mean map corresponding to each feature map;

[0065] Step a3, calculating the variance of the RGB three-channel image of each feature map to obtain a variance map corresponding to each feature map;

[0066] Step a4, initializing the standard Gaussian distribution matrix;

[0067] Step a5, using a standard Gaussian distribution matrix to randomly sample the mean map and variance map corresponding to each feature map, to obtain the same number of random sampling maps as the feature maps;

[0068] Step a6, extracting several uncertainty graphs from the random sampling graph.

[0069] Specifically, Figure 4As shown, it is a schematic diagram of the specific process of calculating the uncertainty map of an embodiment of the present invention. The uncertainty quantification part of the feature map is constructed as a probability representation model. At this stage, the network performs pixel-level mean and variance calculations for the feature map extracted by each backbone network, and then uses the standard Gaussian distribution N(0,1) to randomly sample the above mean and variance. The two are fused according to the sampling result, and the obtained uncertainty map is used to represent the uncertainty score of the corresponding pixel. The uncertainty score is extracted from the learned distribution. This method decomposes the direct sampling operation into a trainable part and a random sampling part.

[0070] Specifically, for each feature map extracted by the backbone network F1, the corresponding RGB three-channel image is first extracted, and then divided into two paths and input into the probability module F2. Each feature map has three feature sub-maps corresponding to the three color channels of R, G, and B. Taking the current feature map as an example, the first path of the probability module F2 calculates the average values ​​of the pixels of the three feature sub-maps according to the corresponding positions of the pixels of the three color channels, and the three channels are fused into one channel to obtain the mean map corresponding to the current feature map. Similarly, taking the current feature map as an example, the second path of the probability module F2 calculates the variances of the pixels of the three feature sub-maps according to the corresponding positions of the pixels of the three color channels, and the three channels are fused into one channel to obtain the variance map corresponding to the current feature map, for example Figure 4 The dimension of the image to be detected is H×W×3. After slicing, the dimension of each feature map is h×w×3. After calculating the mean map and variance map, the RGB three channels are combined into one, and the dimension becomes h×w×1. The formula is expressed as follows:

[0071] Given an input image P∈R H×W×3 , the backbone network F1 extracts the c-dimensional feature embedding F = F1(P)∈R h×w×3 , and then the probability module F2 further transforms the features into mean μ and variance σ 2 :

[0072] μ=F 2μ (F),σ=F 2σ (F)

[0073] where μ∈R h×w×1 and σ∈R h×w×1 Represent the mean and variance plots respectively.

[0074] Then initialize the matrix of the labeled Gaussian distribution, and randomly sample the mean map and variance map by the labeled Gaussian distribution matrix. The specific operation method is to multiply the labeled Gaussian distribution matrix and the variance map and then add it to the mean map. Each feature map is processed as above to obtain c random sampling maps. Finally, K extractions are performed in the c random sampling maps to obtain K uncertainty maps. The obtained uncertainty map can be used as the prior information of the encoder and each feature map. Figure 1 The vector is converted into a sequence and input into the encoding layer to participate in the subsequent model training and prediction.

[0075] Through the technical solution provided by the embodiment of the present invention, uncertainty Figure 1 The input coding layer allows the neural network to include more parameters and adds more prior information for subsequent training and prediction. The uncertainty graph can evaluate the discreteness of the model based on the variance, thereby evaluating the uncertainty of the features and significantly improving the reliability of the model.

[0076] In addition, in an optional implementation, the uncertainty map is essentially a matrix. The embodiment of the present invention can also simply regard the extracted K uncertainty maps as empirical samples of the approximate prediction distribution, and measure the confidence of the model in its prediction by calculating the variance:

[0077] U=Norm(Var(C init ))

[0078] Where U represents the quantitative uncertainty indicator, Norm() represents the maximum mean normalization operation, and Var() represents the operation of calculating the variance.

[0079] This quantitative indicator can be used to objectively evaluate the credibility of the DINO network at the abstract level and reflect it in the number of iterations of model training.

[0080] In some optional embodiments, the present invention further creates a first loss function that can be used to train the backbone network. The loss value of the first loss function is obtained by calculating the sum of the binary cross entropy loss and the KL divergence. The binary cross entropy loss is used to calculate the loss between the uncertainty map and the corresponding real target image. The KL divergence is used to calculate the distribution difference value between the distribution of the uncertainty map and the distribution of the corresponding real target image. The specific formula is as follows:

[0081] L UD =L BCE (c k ,c gt )+αD KL (N(μ,σ)||N(0,1))

[0082] Where, L UD represents the loss value of the first loss function; α represents the combination weight, which is usually set to 0.1 to emphasize the prediction of the model, but can also be adjusted according to user needs; c k represents an uncertainty map randomly drawn from the learning distribution; c gt represents the true label, c gt It is a matrix containing the coordinate parameters of the real target (such as people, animals, etc.), with the same dimension as c kSame; L BCE represents the binary cross entropy loss, the formula of the binary cross entropy loss is the prior art, and will not be described in detail in this embodiment; D KL Represents the KL divergence, which is used to measure the distribution difference between the distribution N(μ,σ) of the uncertainty map and the distribution N(0,1) of the corresponding true target image.

[0083] When training the network, the influence of the uncertainty graph is taken into account, and L UD The parameters of the backbone network are adjusted to minimize the optimization target and improve the accuracy of model training.

[0084] In some optional implementations, the above step S207 includes:

[0085] Step b1, calculating a first uncertainty quantification index based on the distance between each predicted query vector and the target category included in the image to be detected;

[0086] Step b2, screening the prediction query vector by using the first uncertainty quantification index to obtain an optimized query vector;

[0087] Step b3, calculating the optimized position information corresponding to the optimized query vector;

[0088] Step b4, embedding the optimization position information, the decoding layer query vector, the position encoding of the decoding layer query vector and the uncertainty quantification index into the decoding layer;

[0089] Step b5, processing the second feature sequence through a decoding layer to obtain a third feature sequence.

[0090] Specifically, the DINO network processes the second feature sequence output by the encoding layer using the initialized decoding layer query vector and the position information of the prediction query vector at the decoding layer. This step is called a hybrid query in the DINO network. On this basis, the embodiment of the present invention provides a second improvement point, which is to perform an additional step of screening on each prediction query vector by calculating the uncertainty quantification index of each prediction query vector, select a better optimized query vector, and then perform a hybrid query using the optimized position information corresponding to the optimized query vector, the decoding layer query vector, the position encoding of the decoding layer query vector, and the uncertainty quantification index, so as to further improve the accuracy of the decoding layer output.

[0091] First, the corresponding centroid vector is defined for the target category (such as human, dog, car, etc.) included in the image to be detected, and then the distance between the centroid vector of the target category and the predicted query vector is calculated, such as Euclidean distance, Chebyshev distance, etc. The calculated distance is used to measure the uncertainty of each predicted query vector. When the distance between the predicted query vector and each target category in the image is large, it means that the predicted query vector is not close to any target category and its uncertainty is large. When the distance between the predicted query vector and some target categories is small, it means that the current predicted query vector is more sensitive to the target category close to the query and its uncertainty is small. In this way, the uncertainty of the predicted query vector is quantified, and the optimized query vector with smaller uncertainty is selected from the predicted query vector based on the quantification method. Then, based on the initialization decoding layer query vector and the decoding layer query vector position encoding of the original input of DINO, the position information and uncertainty quantification index corresponding to the optimized query vector are used as prior information, and the initialization decoding layer query vector and the decoding layer query vector position encoding are input into the decoding layer together, and then the decoding layer is used to process the second feature sequence.

[0092] Through this screening mechanism, the position information corresponding to the optimized query vector is calculated, thereby telling the decoding layer which positions in the initialized query vector are more likely to contain targets and which positions are more likely to be backgrounds, making the query direction of the decoding layer clearer and further improving the accuracy of the predicted query vector output third feature sequence.

[0093] In some optional implementations, step b1 includes:

[0094] Step c1, projecting each prediction query vector into the feature space through a preset mapping matrix;

[0095] Step c2, determining the centroid vector of each target category in the feature space;

[0096] Step c3, calculating the exponential distance between the projection of each predicted query vector and each centroid vector;

[0097] Step c4: extracting the maximum exponential distance corresponding to each prediction query vector from all the calculated exponential distances as the first uncertainty quantification indicator corresponding to each prediction query vector.

[0098] Specifically, Figure 5 As shown, the present invention adopts the uncertainty quantification method of DUQ (Deep Uncertainty Quantification), which is completed by calculating the kernel function (distance function) of each predicted query vector and the centroid vector of each target category, and the obtained quantification result is used as one of the main indicators for screening the predicted query vector.

[0099] First, through a mapping matrix W c Project the prediction query vector to the feature dimension R required by the feature space hw×c , define the learnable category centroid e corresponding to each target category in the feature space c , the category centroid e of each category c It exists in the form of a weight matrix.

[0100] The embodiment of the present invention calculates each prediction query vector and the category centroid e based on the RBF radial basis function c The exponential distance between is as follows:

[0101]

[0102] Where W c :R d →R c , R d is the space of input dimensions, R c is the space of output dimensions, e c is the centroid of different categories c, and σ is a hyperparameter. This function uses the category-related weight matrix calculation to allow insensitivity to features, thereby minimizing the possibility of feature collapse. The numerator calculates the distance from each predicted query vector data f to each centroid, and then undergoes an exponential process. The meaning is that the smaller the distance, the smaller the K c (f,e c ) is larger.

[0103] Because the distance between each prediction query vector and each centroid is calculated, a prediction query vector currently contains multiple exponential distances. For a prediction query vector, you only need to take the closest category, that is, the largest K c (f,e c )(maximum exponential distance) to represent the uncertainty of a prediction query vector, thereby obtaining the first uncertainty quantification index y of each prediction query vector f , as shown below.

[0104] u f = argmaxK c (f,e c )

[0105] In some optional implementations, step b2 includes:

[0106] Step c5, sorting the maximum exponential distances corresponding to the prediction query vectors from large to small;

[0107] Step c6, selecting a preset number of target maximum index distances from one side of the maximum value in the sorting results;

[0108] Step c7: determine the predicted query vector corresponding to the target maximum exponential distance as the optimized query vector.

[0109] Specifically, Figure 6 As shown, the uncertainty quantification index u obtained based on the above quantification in the embodiment of the present invention is f In the second step, the prediction query vector (prediction query) after the encoder is screened. Each prediction query vector is sorted from large to small according to their respective maximum exponential distances, and then a preset number of prediction query vectors are selected. Usually, the preset number is set to 900. In other words, 900 high-quality optimized query vectors are screened from f prediction query vectors. The corresponding optimized query vectors are passed through the bounding box detection head to obtain the optimized position information, which is passed to the next decoding layer to enhance the prior information of the model. Finally, the optimized position information, the initialized decoding layer query vector, the position encoding of the initialized decoding layer query vector, and the first uncertainty quantification index are input into the decoding layer together.

[0110] The embodiment of the present invention proposes an uncertainty quantification method for calculating the exponential distance between the prediction query vector and the centroid vector of each target category in the feature space through the radial basis function, wherein the exponential distance is opposite to some commonly used distances, and the larger the exponential distance is, the greater the correlation between the prediction query vector and the target category is, and the higher the credibility is, so that the maximum exponential distance corresponding to each prediction query vector can represent the credibility of each prediction query vector, and the prediction query vector with a higher credibility of a preset number is screened according to the sorting of credibility, further improving the accuracy of the prior information of the query vector of the decoding layer. At the same time, the centroid vector in the feature space has the ability to be adjusted, which is not only convenient for screening the prediction query vector according to the first uncertainty quantification index, but also can continuously adjust the centroid during the model training process to improve the accuracy of the first uncertainty quantification index.

[0111] In some optional implementations, the above step S208 includes:

[0112] Step d1, inputting the third feature sequence into a multilayer perceptron, and projecting the third feature sequence into a feature space;

[0113] Step d2, determining the centroid vector of each target category in the feature space;

[0114] Step d3, calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence;

[0115] Step d4, extracting the maximum exponential distance corresponding to each prediction feature from all the calculated exponential distances as a second uncertainty quantification indicator corresponding to each prediction feature;

[0116] Step d5, classifying the third feature sequence based on the Softmax classification module to obtain the classification confidence corresponding to each prediction feature;

[0117] Step d6, fusing the classification confidence and the second uncertainty quantification index to obtain a third uncertainty quantification index corresponding to each prediction feature in the third feature sequence;

[0118] Step d7, when processing the third feature sequence through the Hungarian matching layer, the third uncertainty quantification index is calculated in the matching cost to obtain the target feature sequence.

[0119] According to the above technical means, the present invention proposes a third improvement point, which realizes uncertainty quantification for the detection head of the DINO network, and optimizes the matching process of the Hungarian matching layer based on the uncertainty quantification result. Among them, the uncertainty index of each third feature sequence is calculated by radial basis function, and is integrated with the classification confidence output by the traditional classification layer to obtain the credibility of each predicted feature in the third feature sequence, and the credibility of each predicted feature is calculated in the matching cost of the Hungarian matching algorithm. The higher the credibility, the smaller the corresponding matching cost, thereby improving the matching accuracy of the Hungarian matching algorithm, and then improving the accuracy of the final target detection frame position and classification.

[0120] Among them, Figure 7 As shown, the multilayer perceptron is used to project the third feature sequence into the feature space, and its role is the same as that of the mapping matrix in the aforementioned step. The principles of steps d1 to d4 can refer to the relevant descriptions of steps c1 to c4 above. The principles of the two are the same and will not be repeated here.

[0121] After the uncertainty of the third feature sequence is quantified, the Softmax classification module (i.e., the original classification detection head of the DINO network) is used to classify the third feature sequence, thereby obtaining the classification confidence corresponding to different prediction features in the third feature sequence. Finally, for each prediction feature in the third feature sequence, the corresponding classification confidence is multiplied by the second uncertainty quantification index to obtain the third uncertainty quantification index, which is used to quantify the credibility of the third feature sequence output by the decoding layer.

[0122] In some optional implementations, the above steps c2 and d2 include:

[0123] Step e1, randomly initialize the centroid vector during the first round of training;

[0124] Step e2, inputting the third feature sequence into the multilayer perceptron during the training process after the second round, and projecting the third feature sequence into the feature space;

[0125] Step e3, calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence;

[0126] Step e4, calculate the binary cross entropy between each exponential distance and the one-hot encoding of the true category label;

[0127] Step e5, update the centroid vector based on the condition of minimizing the binary cross entropy.

[0128] Specifically, the embodiment of the present invention also provides a loss function for optimizing the centroid vector of the target category in the graph, thereby improving the correlation between the centroid vector and the true target category to improve the accuracy of the centroid vector. This is specifically achieved by the binary cross entropy between the exponential distance calculated at the output position of the decoding layer and the one-hot encoding of the true category label.

[0129] Except for the first round of training process which requires random initialization of the centroid vector, after the second round of training, the decoder can output the third feature sequence to predict the target in the image. The target predicted by the third feature sequence is closer to the actual situation, thus using the third feature sequence to predict the centroid vector has higher reliability.

[0130] After calculating the exponential distance K between the projection of each predicted feature and each centroid vector in the third feature sequence c (f,e c ) and then K c (f,e c ) into the following formula.

[0131]

[0132] Where L(x,y) represents the binary cross entropy, (x,y) is the data point in the third feature sequence f, and y c represents the unique hot encoding of the true category label corresponding to category c. During training, the losses of small batches of data points are averaged and K c (f,e c ) is learned by stochastic gradient descent, and the category centroid e c Update using the exponential moving average of the feature vectors of the data points belonging to this class. If the model parameters and the multilayer perceptron parameters remain unchanged, this update rule will form the centroid e c The closed-form solution of the centroid e can be completed by minimizing the above loss through this closed-form solution. c In the subsequent training process, the updated centroid e c Back to the above steps c2 and d2, which can significantly improve the centroid e c improve the accuracy of the model.

[0133] Through the technical solution provided by the embodiment of the present invention, three innovations for the DINO network are proposed. Based on the above three innovations, the front, middle and back parts of the DINO process of the Transformer target detection algorithm are improved respectively, thereby improving the accuracy of target detection.

[0134] The feature map uncertainty quantification network considers the backbone network to further analyze the results of image feature extraction, quantifies the data uncertainty in the feature map, and embeds the uncertainty results into the Transformer encoder output sequence, so that the encoder can have a deeper processing of the feature data representing different position areas of the image; the RBF network-based uncertainty screening module uses the DUQ method to innovate the query screening module between the encoder and decoder of the Transformer, quantifies the uncertainty of all queries processed by the encoder, and based on the results, selects the position information corresponding to high-quality queries as the position feature of the decoder query. At the same time, the corresponding uncertainty results are also embedded in the input query to enhance the prior information of this part of the input; the uncertainty-based detection head network adds an uncertainty output module to the classification and regression modules of the original model, and also sets the uncertainty loss function based on the DUQ algorithm, so that the results after model training are more reliable.

[0135] like Figure 8As shown in FIG. 1 , it is a schematic diagram of the detection effect. The first example selects a dense target scene with obvious foreground and background. For the very obvious foreground target in the picture, such as the horse in the picture, the present invention and other methods can obtain accurate detection frames, with confidence levels of 83.7% and 86.1% respectively, and the present method gives a higher confidence level; and for controversial small targets at the back of the picture, such as an umbrella with only half of the edge of the picture exposed, other methods and the present method also obtain confidence levels of 52.2% and 47.0% respectively, indicating that the present method will give higher confidence levels to targets that are obviously more certain, and will output lower confidence levels for targets that are uncertain and prone to errors, indicating the reliability of the present method; the second and third examples select more complex traffic scenes, and it can be seen that compared with other methods, the present method will output higher confidence levels for some obvious large targets. High confidence, and can still maintain the correct output frame results for blurred small targets, and there will be more rational confidence judgments, and output more conservative results, which verifies that this method has better output reliability in complex traffic scenes; the fourth example selects a close-up scene of an athlete after the background is blurred. Other methods and this method can detect the athlete itself, with confidence levels of 88.5% and 92.3% respectively, and this method is better; however, other methods and this method have output detection results for the audience after the background is blurred, which is a target that is not marked in the true value image, but this method outputs fewer frames and has a lower confidence level. Although humans can understand the audience behind the blurred background through knowledge, the model should have more cautious and reliable judgments for such blurred targets, so that users of artificial intelligence algorithms can be more assured of the output results.

[0136] According to the visual analysis and comparison of the above four figures, it can be seen that this method can accurately detect the target, and compared with the original method, it has a higher confidence level for obvious targets and a cautious judgment for blurred targets. At the same time, the experimental results prove that this method has higher detection accuracy and will not have the problem of missed detection, which illustrates the performance superiority and reliability of this method.

[0137] In this embodiment, a target detection device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0138] This embodiment provides a target detection device, such as Fig. 9 As shown, including:

[0139] The image acquisition module 801 is used to acquire the image to be detected;

[0140] A feature extraction module 802 is used to extract a feature map of the image to be detected based on the backbone network;

[0141] The uncertainty feature extraction module 803 is used to extract the variance-based uncertainty map from the feature map;

[0142] An encoding module 804 is used to convert the position code of the feature map, the uncertainty map and the feature map into a first feature sequence and input it into the encoding layer;

[0143] A prediction query vector generation module 805 is used to obtain a second feature sequence output by the encoding layer, and process the second feature sequence through an intermediate classification prediction layer to obtain multiple prediction query vectors;

[0144] A random initialization query vector module 806 is used to initialize a decoding layer query vector;

[0145] A hybrid query module 807, configured to process the second feature sequence based on a decoding layer that combines the prediction query vector and the decoding layer query vector to obtain a third feature sequence;

[0146] A Hungarian matching module 808 is used to input the third feature sequence into the Hungarian matching layer to obtain a target feature sequence through matching;

[0147] The target result output module 809 is used to process the target feature sequence through the classification detection layer and the regression detection layer respectively to obtain the target position and target category in the image to be detected.

[0148] In some optional implementations, the uncertainty feature extraction module 803 includes:

[0149] A three-channel image extraction unit, used for extracting an RGB three-channel image for each feature map;

[0150] A mean calculation unit is used to calculate the mean of the RGB three-channel image of each feature map to obtain a mean map corresponding to each feature map;

[0151] A variance calculation unit is used to calculate the variance of the RGB three-channel image of each feature map to obtain a variance map corresponding to each feature map;

[0152] Standard distribution initialization unit, used to initialize the standard Gaussian distribution matrix;

[0153] A random sampling unit, used to randomly sample from the mean map and variance map corresponding to each feature map using a standard Gaussian distribution matrix to obtain random sampling maps with the same number as the feature maps;

[0154] The uncertainty graph generation unit is used to extract a number of uncertainty graphs from the random sampling graph.

[0155] In some optional implementations, the hybrid query module 807 includes:

[0156] A first uncertainty quantification unit, configured to calculate a first uncertainty quantification index based on the distance between each predicted query vector and the target category included in the image to be detected;

[0157] A screening unit, configured to screen the prediction query vector by using a first uncertainty quantification indicator to obtain an optimized query vector;

[0158] An optimization position unit, used to calculate optimization position information corresponding to the optimization query vector;

[0159] A decoding layer input unit, used for embedding the optimization position information, the decoding layer query vector, the position encoding of the decoding layer query vector and the uncertainty quantification indicator into the decoding layer;

[0160] The decoding layer output unit is used to process the second feature sequence through the decoding layer to obtain a third feature sequence.

[0161] In some optional implementations, the first uncertainty quantification unit includes:

[0162] A projection unit, used for projecting each prediction query vector into a feature space through a preset mapping matrix;

[0163] A centroid determination unit, used to determine the centroid vector of each target category in the feature space;

[0164] An exponential distance unit, used to calculate the exponential distance between the projection of each predicted query vector and each centroid vector;

[0165] The first indicator generating unit is used to extract the maximum exponential distance corresponding to each prediction query vector from all the calculated exponential distances as the first uncertainty quantification indicator corresponding to each prediction query vector.

[0166] In some optional embodiments, the screening unit comprises:

[0167] A sorting unit, used to sort the maximum index distances corresponding to each prediction query vector from large to small;

[0168] A maximum distance determination unit, used for selecting a preset number of target maximum index distances from one side of the maximum value in the sorting result;

[0169] The screening subunit is used to determine the predicted query vector corresponding to the target maximum exponential distance as the optimized query vector.

[0170] In some optional implementations, the Hungarian matching module 808 includes:

[0171] A projection unit, used for inputting the third feature sequence into the multi-layer perceptron, and projecting the third feature sequence into a feature space;

[0172] A centroid determination unit, used to determine the centroid vector of each target category in the feature space;

[0173] A distance calculation unit, used for calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence;

[0174] A second index generating unit is used to extract the maximum index distance corresponding to each prediction feature from all the calculated index distances as a second uncertainty quantification index corresponding to each prediction feature;

[0175] A classification detection unit, used for classifying the third feature sequence based on a Softmax classification module to obtain a classification confidence corresponding to each prediction feature;

[0176] A fusion unit, used to fuse the classification confidence and the second uncertainty quantification index to obtain a third uncertainty quantification index corresponding to each prediction feature in the third feature sequence;

[0177] The matching unit is used to calculate the third uncertainty quantification index into the matching cost when processing the third feature sequence through the Hungarian matching layer to obtain the target feature sequence.

[0178] In some optional implementations, the centroid determination unit comprises:

[0179] The random unit is used to randomly initialize the centroid vector during the first round of training;

[0180] A training projection unit, used for inputting the third feature sequence into the multilayer perceptron in the training process after the second round, and projecting the third feature sequence into the feature space;

[0181] A training distance calculation unit, used for calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence;

[0182] The loss calculation unit is used to calculate the binary cross entropy between each exponential distance and the one-hot encoding of the true class label;

[0183] The centroid update unit is used to update the centroid vector under the condition of minimizing the binary cross entropy.

[0184] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0185] The target detection device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0186] The embodiment of the present invention also provides a computer device having the above Fig. 9 The target detection device shown.

[0187] See also Fig.10 , Fig.10 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig.10 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.10 A processor 10 is taken as an example.

[0188] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0189] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0190] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0191] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0192] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0193] The present invention also provides a vehicle, in which the computer device provided by the above embodiment is installed.

[0194] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0195] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0196] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A target detection method, characterized in that: The method comprises: Acquire the image to be detected; Extracting a feature map of the image to be detected based on a backbone network; extracting a variance-based uncertainty map from the feature map; Convert the position code of the feature map, the uncertainty map and the feature map into a first feature sequence and input the first feature sequence into a coding layer; Acquire a second feature sequence output by the encoding layer, and process the second feature sequence through an intermediate classification prediction layer to obtain multiple prediction query vectors; Initialize the decoding layer query vector; Processing the second feature sequence based on a decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence; Input the third feature sequence into the Hungarian matching layer to obtain a target feature sequence through matching; The target feature sequence is processed respectively by the classification detection layer and the regression detection layer to obtain the target position and target category in the image to be detected.

2. The method according to claim 1, characterized in that: The extracting a variance-based uncertainty map from the feature map comprises: Extract RGB three-channel images for each feature map; Calculate the mean of the RGB three-channel image of each feature map to obtain the mean map corresponding to each feature map; Calculate the variance of the RGB three-channel image of each feature map to obtain the variance map corresponding to each feature map; Initialize the standard Gaussian distribution matrix; Using the standard Gaussian distribution matrix, randomly sampling from the mean map and the variance map corresponding to each feature map to obtain the same number of random sampling maps as the feature maps; A number of the uncertainty graphs are extracted from the random sampling graph.

3. The method according to claim 2, characterized in that The backbone network is also trained by a first loss function, the loss value of which is obtained by calculating the sum of a binary cross entropy loss and a KL divergence, the binary cross entropy loss is used to calculate the loss between the uncertainty map and the corresponding true target image, and the KL divergence is used to calculate the distribution difference value between the distribution of the uncertainty map and the distribution of the corresponding true target image.

4. The method according to claim 1, characterized in that The step of processing the second feature sequence based on the decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence includes: Calculating a first uncertainty quantification index based on the distance between each predicted query vector and the target category included in the image to be detected; Filtering the prediction query vector by using the first uncertainty quantification index to obtain an optimized query vector; Calculating optimized position information corresponding to the optimized query vector; embedding the optimization position information, the decoding layer query vector, the position encoding of the decoding layer query vector, and the uncertainty quantification indicator into a decoding layer; The second feature sequence is processed by the decoding layer to obtain a third feature sequence.

5. The method according to claim 4, characterized in that The calculating a first uncertainty quantification index based on the distance between each predicted query vector and the target category included in the image to be detected includes: Project each prediction query vector into the feature space through a preset mapping matrix; Determining a centroid vector for each target category in the feature space; Calculate the exponential distance between the projection of each predicted query vector and each centroid vector; The maximum exponential distance corresponding to each prediction query vector is extracted from all calculated exponential distances as a first uncertainty quantification indicator corresponding to each prediction query vector.

6. The method according to claim 5, characterized in that The step of screening the prediction query vector by using the first uncertainty quantification indicator to obtain an optimized query vector includes: Sort the maximum index distances corresponding to each prediction query vector from large to small; Select a preset number of target maximum index distances from one side of the maximum value in the sorting results; The predicted query vector corresponding to the target maximum exponential distance is determined as the optimized query vector.

7. The method according to claim 1, characterized in that Inputting the third feature sequence into the Hungarian matching layer to obtain a target feature sequence by matching includes: Inputting the third feature sequence into a multilayer perceptron, and projecting the third feature sequence into a feature space; Determining a centroid vector for each target category in the feature space; Calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence; Extracting the maximum exponential distance corresponding to each prediction feature from all the calculated exponential distances as the second uncertainty quantification index corresponding to each prediction feature; Classify the third feature sequence based on the Softmax classification module to obtain the classification confidence corresponding to each prediction feature; The classification confidence and the second uncertainty quantification index are combined to obtain a third uncertainty quantification index corresponding to each prediction feature in the third feature sequence; When the third feature sequence is processed by the Hungarian matching layer, the third uncertainty quantification index is calculated in the matching cost to obtain the target feature sequence.

8. The method according to claim 5 or 7, characterized in that: Determining the centroid vector of each target category in the feature space includes: Randomly initialize the centroid vector during the first round of training; In the training process after the second round, the third feature sequence is input into a multilayer perceptron, and the third feature sequence is projected into a feature space; Calculating the exponential distance between the projection of each predicted feature and each centroid vector in the third feature sequence; Calculate the binary cross entropy between each exponential distance and the one-hot encoding of the true class label; The centroid vector is updated on the condition that the binary cross entropy is minimized.

9. A target detection device, characterized in that: The device comprises: An image acquisition module, used for acquiring an image to be detected; A feature extraction module, used to extract a feature map of the image to be detected based on a backbone network; An uncertainty feature extraction module, configured to extract a variance-based uncertainty map from the feature map; An encoding module, used for converting the position encoding of the feature map, the uncertainty map and the feature map into a first feature sequence and inputting the first feature sequence into an encoding layer; A prediction query vector generation module, used for obtaining a second feature sequence output by the encoding layer, and processing the second feature sequence through an intermediate classification prediction layer to obtain a plurality of prediction query vectors; Randomly initialize query vector module, used to initialize the query vector of the decoding layer; a hybrid query module, configured to process the second feature sequence based on a decoding layer that fuses the prediction query vector and the decoding layer query vector to obtain a third feature sequence; A Hungarian matching module, used for inputting the third feature sequence into a Hungarian matching layer to obtain a target feature sequence through matching; The target result output module is used to process the target feature sequence through the classification detection layer and the regression detection layer respectively to obtain the target position and target category in the image to be detected.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 8 by executing the computer instructions.

11. A vehicle, characterized in that: The computer device provided in claim 10 is installed in the vehicle.

Citation Information

Cited By

  • Scene perception method, training method, program product, medium, and electronic device

    CN120877253A

  • Behavior recognition and model training method and device, equipment, medium and product

    CN121053702A