Target detection method and device, computer equipment and storage medium
By jointly training the visual self-attention model, and introducing a random mask into the attention mechanism, the problem of limited accuracy of the existing object detection algorithm in small targets and dense scenes is solved, and the robustness and noise resistance of the model are improved.
Patent Information
- Application Number
- CN202510288421.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-24
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-16
AI Technical Summary
The existing object detection algorithms are limited in small-objective and dense scenarios, and the Transformer-based model is prone to overfitting, damage to attention weight distribution and loss of semantic information.
By jointly training the preset visual self-attention model, encoder, decoder and one-to-one set matching model, a random mask with the same shape as the K matrix in the visual self-attention model is generated, and it is added element by element to the K matrix to obtain the processed K matrix to improve the robustness of the attention mechanism.
It improves the model's performance in diversified feature learning and noise resistance, reduces the dependence on redundant features, and enhances the accuracy of object detection.
Smart Images

Figure CN120014385A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of model detection technology, and in particular, to a target detection method, apparatus, computer equipment and storage medium. Background Art
[0002] In the field of computer vision, traditional object detection algorithms such as YOLO (an efficient object detection algorithm) achieve real-time detection through a single-stage design. It divides the image into grids and directly regresses the bounding box and category probability. It has the characteristics of fast speed, but the accuracy is limited in small objects and dense scenes. Transformer-based models such as ViT (Vision Transformer, visual self-attention model) model the global context through the self-attention mechanism, breaking through the local receptive field limitation of CNN (traditional convolutional neural network), but it is easy to overfit on small data sets, and existing regularization methods have problems such as attention weight distribution destruction and semantic information loss. ViT-FRCNN, which combines ViT with Faster R-CNN (Fast Regional Convolutional Neural Network), uses Transformer to enhance the feature extraction capability of the backbone network and improves the accuracy of semantic understanding in complex scenes, but it is still limited by the inherent defects of Faster R-CNN's two-stage process, such as the quality dependence of candidate boxes and the problem of redundant box mistaken deletion caused by NMS (non-maximum suppression). DETR (a target detection algorithm based on Transformer architecture) is an end-to-end Transformer detector that directly outputs detection results through an encoder-decoder architecture, avoiding the RPN (region proposal network) and NMS steps. It performs well in multi-target scenarios, but its global attention mechanism has difficulty distinguishing subtle differences when there is occlusion or target overlap, resulting in missed detections or false detections. Summary of the invention
[0003] The embodiments of the present application provide a target detection method, apparatus, computer device and storage medium.
[0004] A first aspect of an embodiment of the present application provides a target detection method, comprising:
[0005] Jointly training a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder, and a one-to-one set matching model;
[0006] The image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder are respectively input into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model to obtain the target object in the image to be detected and the position information of the target object.
[0007] Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes:
[0008] Generate a random mask with the same shape as the K matrix in the visual self-attention model;
[0009] Add the random mask to the K matrix element by element to obtain the processed K matrix;
[0010] The first image feature output from the visual self-attention model is determined according to the Q matrix, the V matrix and the processed K matrix in the visual self-attention model. In an optional embodiment of the present application, the visual self-attention model, the encoder, the decoder and the one-to-one set matching model are jointly trained to obtain the trained visual self-attention model, the encoder, the decoder and the one-to-one set matching model, including:
[0011] Inputting a preset image into a preset visual self-attention model to obtain a fourth image feature outputted from the preset visual self-attention model;
[0012] Inputting the fourth image feature into a preset encoder to obtain a fifth image feature output by the encoder;
[0013] Inputting the fifth image feature into a preset multi-scale feature extraction model to obtain an output sixth image feature;
[0014] Inputting the sixth image feature into each preset auxiliary head model to obtain a seventh image feature output by each preset auxiliary head model;
[0015] Inputting the position code output by the preset multi-scale feature extraction model and the seventh image feature into a preset decoder to obtain an eighth image feature output by the decoder;
[0016] Inputting the eighth image feature into a preset one-to-one set matching model to obtain a first prediction set of the preset image for the target object;
[0017] Taking the seventh image features respectively output by each preset auxiliary head model as the second prediction set, and constructing a loss function according to the first prediction set, the second prediction set, the first preset actual set, and the second preset actual set;
[0018] According to the loss function value, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model.
[0019] In an optional embodiment of the present application, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted according to the loss function value to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model, including:
[0020] Adjust the parameters of each auxiliary head model respectively according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set;
[0021] Based on the adjusted parameters of each auxiliary head model, adjusting parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, and the one-to-one set matching model according to the loss function value constructed between the first prediction set and the first preset actual set;
[0022] Determine a new first prediction set based on the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and the adjusted parameters of each auxiliary head model, and determine a new loss function value corresponding to the new first prediction set and the first preset actual set;
[0023] When the difference between the loss function values corresponding to the current training round and the previous training round is less than a preset error, the adjusted parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model are used as the parameters of the trained visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model;
[0024] When the difference between the loss function values corresponding to the current training round and the previous training round is greater than or equal to the preset error, the step of adjusting the parameters of each auxiliary head model according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set is re-executed.
[0025] In an optional embodiment of the present application, the fifth image feature is input into a preset multi-scale feature extraction model to obtain an output sixth image feature, including:
[0026] The fifth image feature is sampled using convolution kernels of different scales of a preset multi-scale feature extraction model, where the size of the convolution kernel is H×W, there are C input channels, the convolution kernel is K, and the size of the convolution kernel is k. h ×k w , C′ output channels, convolution step size is s, padding size is p, the sixth image feature output from the preset multi-scale feature extraction model is O, size is H′×W′, where
[0027]
[0028] Among them, O c′ (i, j) is the value of the C′th channel of the sixth image feature at position (i, j), K c′,c (m, n) represents the (m, n) element of the C′th output channel and the Cth input channel of the convolution kernel, I c (i·s+mp, j·s+np) represents the pixel value of the Cth channel of the input feature map in the area covered by the convolution operation, b c′ is the bias term of the C′th output channel, and the sixth image feature is O={O1,O2,...,O n}.
[0029] In an optional embodiment of the present application, the sixth image feature is input into each preset auxiliary head model through the following expression to obtain the seventh image feature output by each preset auxiliary head model:
[0030] Q aux =W q (σ(FCN(Flatten(O)))+α·P)+W a A
[0031] Among them, Q aux is a positive query, O is the sixth image feature, σ is the activation function, Flatten is the flattening function, FCN is the fully connected linear function, α is the learnable scalar weight, P is the position code, A is the auxiliary information vector, and W q and W a are all learnable vector matrices.
[0032] In an optional embodiment of the present application, the position code output by the preset multi-scale feature extraction model and the seventh image feature are input into a preset decoder to obtain an eighth image feature output by the decoder through the following expression:
[0033]
[0034] Final=LayerNorm(FCN(Attention(p q )+Q aux ))
[0035] Where M is the number of sampling points of deformable self-attention, ΔP is the offset, and p q To query the location, P sample = {P m |m=1,2,...,M} is the set of sampling points, From the feature map The features of the corresponding positions are obtained. From the feature map The features of the corresponding position obtained, w ij is the weight, Attention(p q ) is the query position p q attention, To query the position p q is the positive query, LayerNorm is the normalization function, and Final is the eighth image feature output from the decoder.
[0036] A second aspect of an embodiment of the present application provides a target detection device, including:
[0037] A joint training module, used for jointly training a preset visual self-attention model, a preset encoder, a preset decoder and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder and a one-to-one set matching model;
[0038] An input module is used to input the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model, respectively, to obtain the target object in the image to be detected and the position information of the target object.
[0039] Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes:
[0040] Generate a random mask with the same shape as the K matrix in the visual self-attention model;
[0041] Add the random mask to the K matrix element by element to obtain the processed K matrix;
[0042] The first image feature output from the visual self-attention model is determined according to the Q matrix, V matrix and processed K matrix in the visual self-attention model.
[0043] According to a third aspect of an embodiment of the present application, a computer device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above target detection methods are implemented.
[0044] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above target detection methods are implemented.
[0045] The above technical solution provided by the embodiment of the present application has at least some or all of the following advantages compared with the prior art:
[0046] The target detection method described in the embodiment of the present application jointly trains a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder, and a one-to-one set matching model; the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder are respectively input into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model to obtain the target object in the image to be detected and the position information of the target object, wherein, when the image to be detected is input into the trained In the visual self-attention model, obtaining the first image feature output by the visual self-attention model includes: generating a random mask with the same shape as the K matrix in the visual self-attention model; adding the random mask to the K matrix element by element to obtain the processed K matrix; determining the first image feature output from the visual self-attention model according to the Q matrix, V matrix and the processed K matrix in the visual self-attention model, adding the random mask to the K matrix element by element in the visual self-attention model, fusing the random mask with the attention calculation, improving the robustness of the attention mechanism, reducing the dependence on redundant features, thereby promoting diversified feature learning and improving the model's anti-noise ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0048] Figure 1 A flowchart of a target detection method provided by one embodiment of the present application;
[0049] Figure 2 A schematic diagram of the structure of a target detection device provided in one embodiment of the present application;
[0050] Figure 3 A schematic diagram of the computer device structure provided for one embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the technical solutions and advantages in the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0052] See also Figure 1The target detection method provided in the embodiment of the present application includes the following steps 100 to 200:
[0053] Step 100: jointly train a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder, and a one-to-one set matching model;
[0054] Step 200: input the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model, respectively, to obtain the target object in the image to be detected and the position information of the target object.
[0055] Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes:
[0056] Generate a random mask with the same shape as the K matrix in the visual self-attention model;
[0057] Add the random mask to the K matrix element by element to obtain the processed K matrix;
[0058] The first image feature output from the visual self-attention model is determined according to the Q matrix, V matrix and processed K matrix in the visual self-attention model.
[0059] In an optional embodiment of the present application, when the second image feature output by the encoder is input into a trained decoder, the output of the decoder during the training process is also input into the decoder during the detection process as prior knowledge, wherein the output of the decoder during the training process is the output corresponding to the training round of the trained decoder.
[0060] In an optional embodiment of the present application, generating a random mask having the same shape as the K matrix in the visual self-attention model includes:
[0061] Performing instance segmentation on the image to be detected to obtain a target mask matrix;
[0062] Segmenting the target mask matrix based on a preset side length to obtain a plurality of segmented regions;
[0063] Counting the pixel quantities of the plurality of segmented areas to obtain a block mask corresponding to each of the segmented areas;
[0064] A mask matrix of the adapted region self-attention mechanism is generated based on the block masks corresponding to each of the segmented regions, wherein the mask matrix has the same shape as the K matrix in the visual self-attention model.
[0065] In an optional embodiment of the present application, generating a random mask having the same shape as the K matrix in the visual self-attention model further includes:
[0066] Set the target mask matrix to a two-dimensional array A, set the preset side length to size, and set the threshold to threshold.
[0067] Create two a*b all-zero matrices, namely Co matrix as shown in formula (1) and Cr matrix as shown in formula (2), where the value of a is A.shape(0) / size and the value of b is A.shape(1) / size.
[0068]
[0069] The Co matrix is used to record the number of pixels in the plurality of segmented regions, where co ij is the value corresponding to the matrix element, co ij The calculation formula is (3):
[0070]
[0071] The Cr matrix is used to mask irrelevant areas, where cr ij is the value corresponding to the matrix element, cr ij The calculation formula is (4):
[0072]
[0073] Among them, the Cr matrix is the mask matrix of the adaptation region self-attention mechanism, wherein the mask matrix has the same shape as the K matrix in the visual self-attention model.
[0074] In an optional embodiment of the present application, the Q matrix, V matrix and processed K matrix in the visual self-attention model are obtained by the following steps:
[0075] Performing image embedding processing on the image to be detected to obtain an embedding vector X;
[0076] The embedding vector X and the trained weight matrix W q , W k , W v Perform matrix multiplication to obtain the query matrix Q, key matrix K and value matrix V. The formula is:
[0077] Q=X*W q ,
[0078] K=X*W k ,
[0079] V=X*Wv .
[0080] In an optional embodiment of the present application, determining a first image feature output from the visual self-attention model according to a Q matrix, a V matrix, and a processed K matrix in the visual self-attention model includes:
[0081] The query matrix Q is multiplied by the transpose of the processed K matrix K' to obtain the score matrix QK T , the formula is Q*K' T =QK T ;
[0082] The score matrix QK is adjusted by the attention adjustment mechanism T and the mask matrix to generate the attention adjustment result;
[0083] The value matrix V is matrix-multiplied by the attention adjustment result to generate the attention feature vector Z' as the first image feature.
[0084] In an optional embodiment of the present application, the mask matrix is set to Mask, and the score matrix QK is adjusted by the attention adjustment mechanism. T and the mask matrix to generate the attention adjustment result, including:
[0085] The mask matrix Mask is converted into one dimension to obtain a one-dimensional mask matrix Mask_1D;
[0086] Transpose the one-dimensional mask matrix Mask_1D to obtain the transposed mask matrix (Mask_1D) T ;
[0087] The one-dimensional mask matrix Mask_1D and the transposed mask matrix (Mask_1D) are combined. T Perform matrix multiplication to obtain the mask autocorrelation matrix M, the formula is: Mask_1D*(Mask_1D) T =M;
[0088] The score matrix QK T Subtract the product of the mask autocorrelation matrix M and the preset parameter k to obtain the new QK T , the formula is: New QK T =QK T -k*M;
[0089] According to the new QK T Determine attentional modulation outcomes.
[0090] Among them, according to the new QK TIdentify attention regulation outcomes, including:
[0091] The new QK T As QK i T , the attention adjustment result is determined by the following expression:
[0092]
[0093] Among them, DA i is the attention adjustment result of the i-th row, QK i T The result of multiplying the i-th row of the K matrix by the Q matrix. The result of multiplying all rows of the K matrix by the Q matrix is the new QK T , is the scaling factor, k is the preset parameter, M is the mask autocorrelation matrix, and N is the number of rows of the K matrix. When the background attention needs to be weakened, the value interval of the preset parameter k is (-∞, 0), and when the foreground attention needs to be enhanced, the value interval of the preset parameter k is (0, +∞).
[0094] In an optional embodiment of the present application, before determining the first image feature output from the visual self-attention model according to the Q matrix, the V matrix and the processed K matrix in the visual self-attention model, the method further includes:
[0095] Get the training dataset;
[0096] Based on the training data set, the pre-selected visual large model framework is trained to obtain the weight matrix W q , W k , W v .
[0097] In an optional embodiment of the present application, the first image feature output from the visual self-attention model is determined according to the Q matrix, the V matrix and the processed K matrix in the visual self-attention model by the following expression:
[0098]
[0099] FCN(DA i )=max(0,DA i ·W1+b1)W2+b2
[0100] Output i =LayerNorm(FCN(DA i )+DA i )
[0101] Output = f(Output l-1 W1+b1)W2+b2
[0102] Among them, DA i The output of self-attention after random masking is introduced for the i-th layer of the visual self-attention model. is the scaling factor, K i , Q i and V i is the input of the visual self-attention model, W1, b1, W1, b2 are all learnable parameters, f(·) is the activation function, Output i Output is the residual connection and normalized output. l-1 is the output of the previous layer of the last layer of the visual self-attention model, and Output is the first image feature output from the visual self-attention model.
[0103] In an optional embodiment of the present application, the output of the last layer of the visual self-attention model is input into the multi-layer perceptron module of the visual self-attention model for final feature extraction and classification. The dimension of the obtained vector is related to the number of classification categories. Residual connections and layer normalization are added after each self-attention and feedforward neural network layer to alleviate the gradient vanishing problem of the deep network. After the fully connected linear layer (FCN), the DA i The captured features are mapped to a higher dimension. After extracting the high-dimensional features, they need to be mapped to the original dimension. The RELU activation function is used in the fully connected linear layer to activate the obtained features. The output vector is used as the output of the visual self-attention model. Its dimension is the same as the number of classification categories. Then, a layer of encoder is used to repeat the above operations between these categories to extract features, and finally it is used as the input of the decoder and the preset multi-scale feature extraction model. When it is used as the input of the decoder, it is used as the input of the cross attention, and as the input of the preset multi-scale feature extraction model, it is used to extract features from different scales and input into the auxiliary head.
[0104] In an optional embodiment of the present application, when calculating the self-attention weights, a random mask is applied to the K matrix, setting some of the elements to zero. This mask is randomly generated at each forward propagation to ensure that the model does not overly rely on a specific key during training. At each forward propagation, a random mask with the same shape as the K matrix is generated. The elements in this mask are set to infinity with a certain probability, and the remaining elements remain at 0. The probability follows the Bernoulli distribution.
[0105]
[0106] Where n is the element in Mask, and Mask is a mask matrix with the same shape as the K matrix.
[0107] In an optional embodiment of the present application, in step 100, the visual self-attention model, the encoder, the decoder and the one-to-one set matching model are jointly trained to obtain the trained visual self-attention model, the encoder, the decoder and the one-to-one set matching model, including:
[0108] Inputting a preset image into a preset visual self-attention model to obtain a fourth image feature outputted from the preset visual self-attention model;
[0109] Inputting the fourth image feature into a preset encoder to obtain a fifth image feature output by the encoder;
[0110] Inputting the fifth image feature into a preset multi-scale feature extraction model to obtain an output sixth image feature;
[0111] Inputting the sixth image feature into each preset auxiliary head model to obtain a seventh image feature output by each preset auxiliary head model;
[0112] Inputting the position code output by the preset multi-scale feature extraction model and the seventh image feature into a preset decoder to obtain an eighth image feature output by the decoder;
[0113] Inputting the eighth image feature into a preset one-to-one set matching model to obtain a first prediction set of the preset image for the target object;
[0114] Taking the seventh image features respectively output by each preset auxiliary head model as the second prediction set, and constructing a loss function according to the first prediction set, the second prediction set, the first preset actual set, and the second preset actual set;
[0115] According to the loss function value, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model.
[0116] In an optional embodiment of the present application, the sixth image feature is input into each preset auxiliary head model to obtain the seventh image feature output by each preset auxiliary head model, including:
[0117] The sixth image feature is input into the preset multi-scale feature extraction model to obtain the output feature of the preset multi-scale feature extraction model, and the output feature of the preset multi-scale feature extraction model is input into each preset auxiliary head model to obtain the seventh image feature output by each preset auxiliary head model.
[0118] In an optional embodiment of the present application, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted according to the loss function value to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model, including:
[0119] Adjust the parameters of each auxiliary head model respectively according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set;
[0120] Based on the adjusted parameters of each auxiliary head model, adjusting parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, and the one-to-one set matching model according to the loss function value constructed between the first prediction set and the first preset actual set;
[0121] Determine a new first prediction set based on the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and the adjusted parameters of each auxiliary head model, and determine a new loss function value corresponding to the new first prediction set and the first preset actual set;
[0122] When the difference between the loss function values corresponding to the current training round and the previous training round is less than a preset error, the adjusted parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model are used as the parameters of the trained visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model;
[0123] When the difference between the loss function values corresponding to the current training round and the previous training round is greater than or equal to the preset error, the step of adjusting the parameters of each auxiliary head model according to the loss function values constructed according to the second prediction set corresponding to each auxiliary head model and the second preset actual set is re-executed.
[0124] In an optional embodiment of the present application, the fifth image feature is input into a preset multi-scale feature extraction model to obtain an output sixth image feature, including:
[0125] The fifth image feature is sampled using convolution kernels of different scales of a preset multi-scale feature extraction model, where the size of the convolution kernel is H×W, there are C input channels, the convolution kernel is K, and the size of the convolution kernel is k. h ×k w , C′ output channels, convolution step size is s, padding size is p, the sixth image feature output from the preset multi-scale feature extraction model is O, size is H′×W′, where
[0126]
[0127] Among them, Q c′ (i, j) is the value of the C′th channel of the sixth image feature at position (i, j), K c′,c (m, n) represents the (m, n) element of the C′th output channel and the Cth input channel of the convolution kernel, I c (i·s+mp, j·s+np) represents the pixel value of the Cth channel of the input feature map in the area covered by the convolution operation, b c′ is the bias term of the C′th output channel, and the sixth image feature is O = {O1, O2, ..., O n}.
[0128] In an optional embodiment of the present application, the feature map obtained by convolution kernels of different scales is O={O1, O2, ..., O n}, these multi-scale features are input into the auxiliary head i Feature extraction, prediction and customization of positive query are performed in, and i is the number of auxiliary heads.
[0129] In an optional embodiment of the present application, the sixth image feature is input into each preset auxiliary head model through the following expression to obtain the seventh image feature output by each preset auxiliary head model:
[0130] Q aux =W q (σ(FCN(Flatten(O)))+α·P)+W a A
[0131] Among them, Q aux is a positive query, O is the sixth image feature, σ is the activation function, Flatten is the flattening function, FCN is the fully connected linear function, α is the learnable scalar weight, P is the position code, A is the auxiliary information vector, and W q and W a are all learnable vector matrices.
[0132] In an optional embodiment of the present application, the position encoding P and the auxiliary information vector A are usually derived from prior information (such as category label embedding or task-related parameters). Finally, the positive query vector Q is generated aux , a learnable scalar weight α is used to control the contribution of the position encoding, which converts the prior positive query Q aux The output of the encoder is used as the input of the decoder, where Q aux As the Q of the decoder, the output of the encoder is used as the K and V of the encoder, and feature fusion is performed through deformable cross self-attention in the decoder.
[0133] In an optional embodiment of the present application, the position encoding of the preset multi-scale feature extraction model and the seventh image feature are input into a preset decoder through the following expression to obtain an eighth image feature output from the decoder:
[0134]
[0135] Final=LayerNσrm(FCN(Attention(p q )+Q aux ))
[0136] Where M is the number of sampling points of deformable self-attention, ΔP is the offset, and p q To query the location, P sample = {P m |m=1,2,...,M} is the set of sampling points, From the feature map The features of the corresponding positions are obtained. From the feature map The features of the corresponding position obtained, w ij is the weight, Attention(p q ) is the query position p q attention, To query the position p q is the positive query, LayerNorm is the normalization function, and Final is the eighth image feature output from the decoder.
[0137] In an optional embodiment of the present application, the seventh image feature output by each preset auxiliary head model is used as the second prediction set, and a loss function is constructed according to the first prediction set, the second prediction set, the first preset actual set, and the second preset actual set through the following expression:
[0138]
[0139] in, For the first or second prediction set The i-th element in y j The first or second preset actual set Y = {y1, y2, ..., y n}, is the loss function value, λ1 and λ2 are preset values, L class is the classification loss function, L box is the regression loss function.
[0140] The target detection method of this application encourages the model to learn more robust feature representations and prevents overfitting by randomly suppressing some key information. This mechanism is simple and effective to implement and can improve the performance of the model without increasing the complexity of the model. In practical applications, hyperparameters can be adjusted according to the characteristics of specific tasks and data sets to obtain the best performance improvement effect.
[0141] It should be understood that, although the various steps in the flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0142] See also Figure 2 , an embodiment of the present application provides a target detection device 200, including:
[0143] A joint training module 210 is used to jointly train a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder, and a one-to-one set matching model;
[0144] The input module 220 is used to input the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model, respectively, to obtain the target object in the image to be detected and the position information of the target object.
[0145] Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes:
[0146] Generate a random mask with the same shape as the K matrix in the visual self-attention model;
[0147] Add the random mask to the K matrix element by element to obtain the processed K matrix;
[0148] The first image feature output from the visual self-attention model is determined according to the Q matrix, V matrix and processed K matrix in the visual self-attention model.
[0149] For the specific definition of the above-mentioned device 200, please refer to the definition of the target detection method above, which will not be repeated here. Each module in the above-mentioned device 200 can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0150] In one embodiment, a computer device is provided, the internal structure diagram of which can be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a target detection method as described above is implemented. It includes: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, any step in the target detection method as described above is implemented.
[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any step in the above target detection method can be implemented.
[0152] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0153] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0154] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0156] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0157] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A target detection method, characterized in that: include: Jointly training a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder, and a one-to-one set matching model; The image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder are respectively input into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model to obtain the target object in the image to be detected and the position information of the target object. Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes: Generate a random mask with the same shape as the K matrix in the visual self-attention model; Add the random mask to the K matrix element by element to obtain the processed K matrix; The first image feature output from the visual self-attention model is determined according to the Q matrix, V matrix and processed K matrix in the visual self-attention model.
2. The method according to claim 1, characterized in that: The visual self-attention model, encoder, decoder and one-to-one set matching model are jointly trained to obtain the trained visual self-attention model, encoder, decoder and one-to-one set matching model, including: Inputting a preset image into a preset visual self-attention model to obtain a fourth image feature outputted from the preset visual self-attention model; Inputting the fourth image feature into a preset encoder to obtain a fifth image feature output by the encoder; Inputting the fifth image feature into a preset multi-scale feature extraction model to obtain an output sixth image feature; Inputting the sixth image feature into each preset auxiliary head model to obtain a seventh image feature output by each preset auxiliary head model; Inputting the position code output by the preset multi-scale feature extraction model and the seventh image feature into a preset decoder to obtain an eighth image feature output by the decoder; Inputting the eighth image feature into a preset one-to-one set matching model to obtain a first prediction set of the preset image for the target object; Taking the seventh image features respectively output by each preset auxiliary head model as the second prediction set, and constructing a loss function according to the first prediction set, the second prediction set, the first preset actual set, and the second preset actual set; According to the loss function value, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model.
3. The method according to claim 2, characterized in that According to the loss function value, the parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model are adjusted to obtain the trained visual self-attention model, the encoder, the multi-scale feature extraction model, each auxiliary head model, the decoder, and the one-to-one set matching model, including: Adjust the parameters of each auxiliary head model respectively according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set; Based on the adjusted parameters of each auxiliary head model, adjusting parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, and the one-to-one set matching model according to the loss function value constructed between the first prediction set and the first preset actual set; Determine a new first prediction set based on the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and the adjusted parameters of each auxiliary head model, and determine a new loss function value corresponding to the new first prediction set and the first preset actual set; When the difference between the loss function values corresponding to the current training round and the previous training round is less than a preset error, the adjusted parameters of the visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model are used as the parameters of the trained visual self-attention model, the encoder, the multi-scale feature extraction model, the decoder, the one-to-one set matching model, and each auxiliary head model; When the difference between the loss function values corresponding to the current training round and the previous training round is greater than or equal to the preset error, the step of adjusting the parameters of each auxiliary head model according to the loss function values constructed according to the second prediction set corresponding to each auxiliary head model and the second preset actual set is re-executed.
4. The method according to claim 2, characterized in that: Inputting the fifth image feature into a preset multi-scale feature extraction model to obtain an output sixth image feature includes: The fifth image feature is sampled using convolution kernels of different scales of a preset multi-scale feature extraction model, where the size of the convolution kernel is H×W, there are C input channels, the convolution kernel is K, and the size of the convolution kernel is k. h ×k w , C′ output channels, convolution step size is s, padding size is p, the sixth image feature output from the preset multi-scale feature extraction model is O, size is H′×W′, where Among them, O c′ (i, j) is the value of the C′th channel of the sixth image feature at position (i, j), K c′,c (m, n) represents the (m, n) element of the C′th output channel and the Cth input channel of the convolution kernel, I c (i·s+mp, j·s+np) represents the pixel value of the Cth channel of the input feature map in the area covered by the convolution operation, b c′ is the bias term of the C′th output channel, and the sixth image feature is O = {O1, O2, ..., O n }.
5. The method according to claim 2, characterized in that: The sixth image feature is input into each preset auxiliary head model through the following expression to obtain the seventh image feature output by each preset auxiliary head model: Q aux =W q (σ(FCN(Flatten(O)))+α·P)+W a A Among them, Q aux is a positive query, O is the sixth image feature, σ is the activation function, Flatten is the flattening function, FCN is the fully connected linear function, α is the learnable scalar weight, P is the position code, A is the auxiliary information vector, and W q and W a are all learnable vector matrices.
6. The method according to claim 2, characterized in that By using the following expression, the position code and the seventh image feature output by the preset multi-scale feature extraction model are input into the preset decoder to obtain the eighth image feature output by the decoder: Final=LayerNorm(FCN(Attention(p q )+Q aux )) Where M is the number of sampling points of deformable self-attention, ΔP is the offset, and p q To query the location, P sample = {P m |m=1,2,...,M} is the set of sampling points, From the feature map The features of the corresponding positions are obtained. From the feature map The features of the corresponding position obtained, w ij is the weight, Attention(p q ) is the query position p q attention, To query the position p q is the positive query, LayerNorm is the normalization function, and Final is the eighth image feature output from the decoder.
7. A target detection device, characterized in that: include: A joint training module, used for jointly training a preset visual self-attention model, a preset encoder, a preset decoder and a preset one-to-one set matching model to obtain a trained visual self-attention model, an encoder, a decoder and a one-to-one set matching model; An input module is used to input the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder into the trained visual self-attention model, the encoder, the decoder, and the one-to-one set matching model, respectively, to obtain the target object in the image to be detected and the position information of the target object. Among them, when the image to be detected is input into the trained visual self-attention model, the first image feature output by the visual self-attention model includes: Generate a random mask with the same shape as the K matrix in the visual self-attention model; Add the random mask to the K matrix element by element to obtain the processed K matrix; The first image feature output from the visual self-attention model is determined according to the Q matrix, V matrix and processed K matrix in the visual self-attention model.
8. A computer device comprising: The method comprises a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the target detection method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the target detection method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Ground penetrating radar B-scan image denoising method
CN112819732A
Extraction model training method, image processing method and related device
CN118071993A
System and method for video instance segmentation via multi-scale spatio-temporal split attention transformer
US20240161334A1