Decoder training methods, object detection methods, devices, and storage media
By introducing relational attention and cross-attention modules into the DETR model, a salient query feature set is generated and a loss function is constructed, which solves the problem of lack of semantic relationships between query features and improves the accuracy and precision of video action detection.
Patent Information
- Application Number
- CN202210788886.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-07-06
Smart Images

Figure CN115063666B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method for training a decoder, a method for detecting objects, an apparatus, and a storage medium. Background Technology
[0002] With the increasing volume of video data, the demand for video data analysis and processing is growing. For example, in scenarios such as live streaming content security detection and short video dangerous action detection, video action detection methods are needed to identify risky actions in video data. Currently, DETR (Bidirectional Encoder Representations from Transformer) models are commonly used for target detection. The DETR model utilizes the Transformer structure to achieve query-based two-dimensional image target detection. The Transformer structure is a network structure based on an attention mechanism; building a model using Transformers can effectively improve the performance of video action detection methods. In the process of realizing this invention, the inventors discovered that the DETR model predicts a fixed number of detection targets through an encoder-decoder approach. In the decoder, a dense self-attention mechanism is typically used to determine the correlation between query features. Since the semantic relationship between the video segments corresponding to each query feature is not considered, invalid query features can interfere with the prediction results, and the prediction results for query features may be inaccurate. Summary of the Invention
[0003] In view of this, one technical problem to be solved by the present invention is to provide a decoder training method, a target detection method, an apparatus, and a storage medium.
[0004] According to a first aspect of this disclosure, a method for training a decoder is provided, wherein the decoder includes: a relational attention module and a cross-attention module; the training method includes: using the relational attention module and based on query features, generating a set of salient query features corresponding to the query features, for updating the query features using the relational attention module and based on the set of salient query features; using the cross-attention module and based on the updated query features, obtaining predicted segment quality information corresponding to the updated query features, and constructing a segment quality loss function based on the predicted segment quality information; obtaining segment relationship features between predicted video segments corresponding to the query features, and constructing a segment relationship loss function; and adjusting the relational attention module and the cross-attention module according to the segment quality loss function and the segment relationship loss function.
[0005] Optionally, generating a set of salient query features corresponding to the query features includes: using the relational attention module and based on the query features, obtaining similarity information between each query feature and segment relationship feature information between video segments corresponding to each query feature; generating a set of similar features corresponding to the query features based on the similarity information; generating a set of relational features corresponding to the query features based on the segment relationship feature information; and generating the set of salient query features based on the set of similar features, the set of relational features, and the query features themselves.
[0006] Optionally, generating a set of similar features corresponding to the query feature based on the similarity information includes: obtaining similar query features of the query feature based on the similarity information; wherein the similarity between the query feature and the similar query features is greater than a preset similarity threshold; and generating the set of similar features based on the similar query features.
[0007] Optionally, the fragment relationship feature information includes: fragment intersection-union ratio; generating a set of relationship features corresponding to the query feature based on the fragment relationship feature information includes: obtaining the relationship query feature of the query feature based on the fragment intersection-union ratio; wherein, the fragment intersection-union ratio between the query feature and the relationship query feature is greater than a preset intersection-union ratio threshold; generating the set of relationship features based on the relationship query feature.
[0008] Optionally, generating the significant query feature set based on the similarity feature set, the relation feature set, and the query feature itself includes: obtaining the relative complement of the similarity feature set with respect to the relation feature set; and taking the union of the relative complement and the query feature itself as the significant query feature set.
[0009] Optionally, the predicted segment quality information includes: a predicted segment quality score; the step of using the cross-attention module and based on the updated query features to obtain the predicted segment quality information corresponding to the updated query features includes: determining the predicted segment corresponding to the updated query features, and obtaining the video segment corresponding to the predicted segment; determining the prediction distance between the midpoint of the predicted segment and the midpoint of the video segment, and the prediction intersection-union ratio between the predicted segment and the video segment; and generating the predicted segment quality score based on the prediction distance and the prediction intersection-union ratio.
[0010] Optionally, constructing the segment quality loss function based on the predicted segment quality information includes: determining the segment distance between the midpoint of the predicted segment and the midpoint of the video segment, and the segment intersection-over-union ratio (CIU) between the predicted segment and the video segment; and constructing the segment quality loss function based on the deviation information between the predicted distance, the predicted CIU, and the corresponding segment distance and CIU.
[0011] Optionally, the segment relationship feature includes: predicted segment intersection-union ratio; obtaining the segment relationship feature between predicted video segments corresponding to the query feature and constructing the segment relationship loss function includes: determining the predicted segment intersection-union ratio between predicted segments corresponding to the updated query feature; and constructing the segment relationship loss function based on the cumulative information of the predicted segment intersection-union ratio.
[0012] Optionally, the step of using the relational attention module and updating the query features based on the salient query feature set includes: using the relational attention module to perform self-attention calculation on the features within the salient query feature set to update the query features.
[0013] Optionally, the decoder module includes a decoder based on the Transformer architecture.
[0014] According to a second aspect of this disclosure, a target detection method is provided, comprising: acquiring a trained decoder; wherein the decoder is trained by the training method described above; using the decoder and based on query features, generating a classification confidence score, regression information for characterizing the target location, and a predicted segment quality score; and determining a prediction score based on the classification confidence score and the predicted segment quality score.
[0015] According to a third aspect of this disclosure, a training apparatus for a decoder is provided, wherein the decoder includes: a relational attention module and a cross-attention module; the training apparatus includes: a query set acquisition module, configured to use the relational attention module and based on query features to generate a set of salient query features corresponding to the query features; a query feature update module, configured to use the relational attention module and based on the set of salient query features to update the query features; a segment quality determination module, configured to use the cross-attention module and based on the updated query features to acquire predicted segment quality information corresponding to the updated query features, and construct a segment quality loss function based on the predicted segment quality information; a prediction loss determination module, configured to determine segment relationship features between predicted video segments corresponding to the query features, and construct a segment relationship loss function; and a module adjustment module, configured to adjust the relational attention module and the cross-attention module according to the segment quality loss function and the segment relationship loss function.
[0016] Optionally, the query set acquisition module includes: a feature information acquisition unit, used to acquire similarity information between various query features and segment relationship feature information between video segments corresponding to each query feature using the relationship attention module and based on query features; a similarity set acquisition unit, used to generate a similar feature set corresponding to the query feature based on the similarity information; a relationship set acquisition unit, used to generate a relationship feature set corresponding to the query feature based on the segment relationship feature information; and a salient set acquisition unit, used to generate the salient query feature set based on the similar feature set, the relationship feature set, and the query feature itself.
[0017] Optionally, the similarity set acquisition unit is specifically used to acquire similar query features of the query feature based on the similarity information; wherein the similarity between the query feature and the similar query features is greater than a preset similarity threshold; and to generate the similar feature set based on the similar query features.
[0018] Optionally, the fragment relationship feature information includes: fragment intersection-union ratio; the relationship set acquisition unit is specifically used to acquire the relationship query feature of the query feature based on the fragment intersection-union ratio; wherein, the fragment intersection-union ratio between the query feature and the relationship query feature is greater than a preset intersection-union ratio threshold; and the relationship feature set is generated based on the relationship query feature.
[0019] Optionally, the salient set acquisition unit is specifically used to acquire the relative complement of the similar feature set with respect to the relation feature set; and to take the union of the relative complement and the query feature itself as the salient query feature set.
[0020] Optionally, the predicted segment quality information includes: a predicted segment quality score; the segment quality determination module includes: a segment quality determination unit, configured to determine the predicted segment corresponding to the updated query features, and obtain the video segment corresponding to the predicted segment; determine the predicted distance between the midpoint of the predicted segment and the midpoint of the video segment, and the predicted intersection-union ratio between the predicted segment and the video segment; and generate the predicted segment quality score based on the predicted distance and the predicted intersection-union ratio.
[0021] Optionally, the segment quality determination module includes: a quality loss determination unit, used to determine the segment distance between the midpoint of the predicted segment and the midpoint of the video segment, and the segment intersection-union ratio between the predicted segment and the video segment; and to construct the segment quality loss function based on the deviation information between the predicted distance, the predicted intersection-union ratio and the corresponding segment distance and segment intersection-union ratio.
[0022] Optionally, the fragment relationship features include: predicted fragment intersection-union ratio; the prediction loss determination module is specifically used to determine the predicted fragment intersection-union ratio between predicted fragments corresponding to the updated query features; and to construct the fragment relationship loss function based on the cumulative information of the predicted fragment intersection-union ratio.
[0023] Optionally, the query feature update module is specifically used to perform self-attention calculation on the features within the salient query feature set using the relational attention module, in order to update the query features.
[0024] Optionally, the decoder module includes a decoder based on the Transformer architecture.
[0025] According to a fourth aspect of this disclosure, a training apparatus for a decoder is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method described above based on instructions stored in the memory.
[0026] According to a fifth aspect of this disclosure, a target detection apparatus is provided, comprising: a model acquisition module for acquiring a trained decoder; wherein the decoder is trained using the training method described above; a detection processing module for using the decoder and based on query features to generate a classification confidence score, regression information characterizing the target location, and a predicted segment quality score; and a prediction score module for determining a prediction score based on the classification confidence score and the predicted segment quality score.
[0027] According to a sixth aspect of this disclosure, a target detection apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method described above based on instructions stored in the memory.
[0028] According to a seventh aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions which are executed by a processor as described above.
[0029] The decoder training method, object detection method, device, and storage medium disclosed herein construct a salient query feature set based on the relationships between query features, and perform self-attention processing on the query features within the salient query feature set to reduce the interference of invalid query features on prediction; by acquiring the quality information of newly added prediction segments and constructing a segment quality loss function, redundant prediction results can be suppressed, and the accuracy of detection results can be improved; by constructing a segment relationship loss function, redundant predictions can be suppressed, making the prediction results more accurate. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a schematic flowchart of an embodiment of the decoder training method according to the present disclosure;
[0032] Figure 2 This is a schematic diagram of the network framework structure of one embodiment of the decoder disclosed herein;
[0033] Figure 3 This is a schematic diagram of the process for generating a set of significant query features in one embodiment of the decoder training method according to the present disclosure;
[0034] Figure 4 This is a diagram illustrating the relationships between query features;
[0035] Figure 5 This is a schematic diagram of the process for generating a predicted segment quality score in one embodiment of the decoder training method according to the present disclosure;
[0036] Figure 6 This is a schematic diagram illustrating the processing of query features in one embodiment of the decoder training method according to the present disclosure;
[0037] Figure 7This is a schematic diagram of the process of constructing a segment quality loss function in one embodiment of the decoder training method according to the present disclosure;
[0038] Figure 8 This is a schematic diagram of the process of constructing a segment relation loss function in one embodiment of the decoder training method according to the present disclosure;
[0039] Figure 9 This is a flowchart illustrating an embodiment of the target detection method according to the present disclosure;
[0040] Figure 10 A schematic diagram of a module of a training apparatus for a decoder according to this disclosure;
[0041] Figure 11 This is a schematic diagram of a query set acquisition module in one embodiment of a training apparatus for a decoder according to the present disclosure;
[0042] Figure 12 A schematic diagram of a segment quality determination module in one embodiment of a decoder training apparatus according to the present disclosure;
[0043] Figure 13 A schematic diagram of a module for another embodiment of a training apparatus for a decoder according to the present disclosure;
[0044] Figure 14 This is a schematic diagram of a module of an embodiment of the target detection device according to the present disclosure;
[0045] Figure 15 This is a schematic diagram of a module of another embodiment of the target detection apparatus according to the present disclosure. Detailed Implementation
[0046] The present disclosure will now be described more fully with reference to the accompanying drawings, which illustrate exemplary embodiments of the present disclosure. The technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present disclosure. The technical solutions of the present disclosure will be described in various aspects below with reference to the various figures and embodiments.
[0047] In the related technologies known to the inventors, the DETR model includes an encoder and decoder based on a Transformer structure, namely a Transformer encoder and a Transformer decoder. The original video sequence is processed by a backbone network (e.g., a convolutional neural network) to extract temporal and spatial feature maps and add location encoding information to synthesize an embedding vector, which is then input into the Transformer encoder. The Transformer encoder extracts image-encoded features through a self-attention mechanism and inputs these image-encoded features along with query features into the Transformer decoder. The Transformer decoder outputs a target query vector, which is then processed by a classification head and a regression head constructed from fully connected layers and multi-layer perceptron layers to output the location and category of the detected target. The detected target can be actions such as walking or running.
[0048] The Transformer architecture exhibits superior performance in feature representation, and building models using Transformers can effectively improve the performance of video action detection methods. A Transformer encoder comprises multiple encoder layers, typically consisting of a multi-head self-attention layer, two normalization layers, and a feedforward neural network layer. Similarly, a Transformer decoder comprises multiple decoder layers, typically consisting of two multi-head self-attention layers, three normalization layers, and a feedforward neural network layer.
[0049] The DETR method takes a fixed number of N learnable query features as input. Each query feature is adaptively sampled from pixels in a 2D image by the network, and information interaction between query features occurs through self-attention. Ultimately, each query feature is used to predict the location and category of a detection box. In the field of temporal action detection, a fixed number of targets are predicted using an encoder-decoder approach. When detecting targets, a Transformer structure based on sparse sampling is used to extract temporal segment features.
[0050] For the decoder, K trainable query features are used as input. Each query feature is a learnable vector that can extract temporal features from a specific moment based on learned statistical information. Self-attention is used to facilitate information exchange between all query features. Each query feature can be passed through a fully connected layer to predict the normalized coordinates of k sampled points across N time dimensions, and features are extracted from video features based on these sampled points to update the query features. For example, another fully connected layer is used to predict k weights from the input query features, and the sampled k features are summed using weighted averages. The updated query features then predict the location and type of the action using a regression head and a classification head, respectively. The regression head consists of a three-layer fully connected layer, and the classification head consists of a one-layer fully connected layer. The regression head predicts the normalized coordinates of the start and end of the action, while the classification head predicts the action's classification and confidence score.
[0051] The decoder in the existing DETR model usually uses a dense self-attention mechanism to obtain the correlation between query features, without considering the semantic relationship between the video segments corresponding to each query feature. Therefore, invalid query segments will interfere with the prediction results of each query feature. Furthermore, due to the lack of constraints between query features, redundant prediction results are easily caused, resulting in inaccurate prediction scores.
[0052] Figure 1 This is a flowchart illustrating an embodiment of a training method for a dialogue generation model according to the present disclosure. The decoder includes a relational attention module and a cross-attention module, such as... Figure 1 As shown:
[0053] Step 101: Using the relational attention module and based on the query features, generate a set of salient query features corresponding to the query features, so as to use the relational attention module and based on the set of salient query features to update the query features.
[0054] In one embodiment, the query feature can be a query vector generated by an existing Transformer encoder, etc. The decoder module includes a decoder based on the Transformer architecture, i.e., a Transformer decoder. For example... Figure 2 As shown, the Transformer decoder includes a relational attention module, a cross-attention module, two normalization layers, and a feedforward network. The normalization layers and feedforward network can use various existing implementations. The input to the Transformer decoder is a fixed number of trainable query features. The relational attention module is an optimized version of the self-attention module in the existing Transformer decoder, used for non-dense attention processing of the query features.
[0055] Step 102: Using the cross-attention module and based on the updated query features, obtain the predicted fragment quality information corresponding to the updated query features, and construct a fragment quality loss function based on the predicted fragment quality information.
[0056] In one embodiment, a cross-attention module is used and, based on updated query features, a feedforward network, along with a classification head, a regression head, and a segment quality head, is used to generate classification confidence, regression information characterizing the target location, and a predicted segment quality score. Here, the target is an action in the video, the classification confidence can be a classification confidence score, and the regression information can be the start and end information of the action.
[0057] The cross-attention module is an optimized version of the self-attention module in the existing Transformer decoder. It adds a segment quality head to obtain a predicted segment quality score. During prediction, this predicted segment quality score is multiplied by the classification confidence score to obtain the final predicted score for the query feature.
[0058] Step 103: Obtain the segment relationship features between the predicted video segments corresponding to the query features, and construct the segment relationship loss function.
[0059] Step 104: Adjust the relation attention module and cross attention module according to the fragment quality loss function and fragment relation loss function.
[0060] In one embodiment, various existing model adjustment methods can be used to adjust the parameters of modules such as the relational attention module and the cross-attention module based on the fragment quality loss function and the fragment relation loss function, so that the function values of the fragment quality loss function and the fragment relation loss function are within the allowed range.
[0061] In one embodiment, a variety of methods can be used to generate a set of salient query features corresponding to query features. Figure 3 This is a schematic diagram illustrating the process of generating a set of salient query features in one embodiment of the decoder training method according to the present disclosure, as follows: Figure 3 As shown:
[0062] Step 301: Using the relational attention module and based on the query features, obtain the similarity information between each query feature and the segment relationship feature information between the video segments corresponding to each query feature.
[0063] In one embodiment, various existing methods can be used to calculate the similarity information between each query feature, such as cosine similarity. Various existing methods can also be used to calculate the segment relationship feature information between video segments corresponding to each query feature, including segment intersection-union ratio (IUU).
[0064] Step 302: Generate a set of similar features corresponding to the query features based on the similarity information.
[0065] In one embodiment, similar query features are obtained based on similarity information. The similarity between the query feature and the similar query features is greater than a preset similarity threshold. The similarity can be cosine similarity, etc. A set of similar features is generated based on the similar query features.
[0066] For example, a relational attention module can be used to model the relationships between query features. Figure 4 In the query feature set, the query features include the true label 311, the reference query fragment 321, significantly similar fragments 331, 332, and 333, significantly dissimilar fragments 341 and 342, and redundant fragments 351. After entering the relational attention module, each query feature predicts its corresponding time segment through a fully connected layer. For the reference query fragment 321, the corresponding significant query feature set includes significantly similar fragments 331, 332, and 333, etc. The query features in the similar feature set have characteristics such as semantic similarity and non-redundancy in the time dimension.
[0067] Based on the similarity information between each query feature, a similarity matrix A∈ Where Lq is a fixed number of query features, and A is a similarity matrix representing the pairwise similarity between the Lq query features. Each element in the similarity matrix A is the cosine similarity between two query features. A similar feature set WEI is constructed based on the similarity threshold γ∈[-1,1].
[0068] E sim ={(i,j)|A[i,j]-γ>0} (1-1);
[0069] Where A[i,j] is the similarity between the i-th query feature and the j-th query feature, and γ is a predefined similarity threshold before training; E sim E is a set of similar features constructed based on the similarity between features, which can correspond to the query features. sim There are multiple [items / items].
[0070] Step 303: Generate a set of relational features corresponding to the query features based on the fragment relational feature information.
[0071] In one embodiment, the fragment relationship feature information includes the fragment intersection-over-union (IoU) ratio. Relationship query features are obtained based on the IoU ratio, provided that the IoU ratio between the query features and the relationship query features is greater than a preset IoU threshold. A set of relationship features is then generated based on the relationship query features.
[0072] For example, the Intersection over Union (IoU) ratio is used to represent the length of the intersection of two segments divided by the length of their union. An IoU matrix is constructed based on the IoU ratio. Each element in matrix B is the IoU value between the video segments (which can be reference feature segments) corresponding to the two query features. Based on the intersection-union ratio threshold v∈[0,1], a set of relational features is constructed:
[0073] E IoU ={(i,j)|B[i,j]-τ>0} (1-2);
[0074] Among them, E IoU τ is the set of relational features constructed based on IoU relationships; B[i,j] is the IoU relationship between the i-th query feature and the j-th query feature, that is, B[i,j] is the IoU value between the video segments corresponding to the i-th query feature and the j-th query feature; τ is the predefined crossover ratio threshold before training.
[0075] Step 304: Generate a significant query feature set based on the similarity feature set, the relation feature set, and the query feature itself.
[0076] In one embodiment, the relative complement of the similar feature set with respect to the relation feature set is obtained, and the union of the relative complement with the query feature itself is taken as the salient query feature set.
[0077] For example, constructing a set of significant query features:
[0078] E = (E IoU \E sim )∪E self (1-3);
[0079] Where E is the set of significant query features, E self It is a self-join set, representing the connection between the i-th query feature and itself.
[0080] In one embodiment, a relational attention module is used to perform self-attention computation on features within the salient query feature set to update the query features. Existing self-attention computation methods can be used to perform self-attention computation on features within the salient query feature set. Through self-attention computation, more expressive features can be obtained based on existing query features.
[0081] For example, the attention weights are calculated for query features within the salient query feature set, using the following method:
[0082] q′ i =a i V i T (1-4);
[0083]
[0084] Where Q, K, and V are the Query, Key, and Value features of each query feature, respectively, and K... i and V i It is the set of keys and values within the set of significant query features corresponding to the i-th query feature, q′ i It is the updated query feature of the i-th query feature, a i It is the attention weight of the elements in the salient query feature set, which is a row-normalized matrix that is the weighted sum of each feature in the Value set.
[0085] To eliminate the interference of invalid query feature fragments on prediction, the decoder training method disclosed herein is based on two metrics: feature similarity and IoU. It dynamically constructs a salient query feature set for each query feature, replacing the dense attention operation of self-attention. The query feature only calculates attention with other query features within the salient query feature set.
[0086] In one embodiment, various methods can be used to obtain the quality information of the predicted fragments corresponding to the updated query features. Figure 5 This is a schematic diagram of the process for generating predicted segment quality scores in one embodiment of the decoder training method according to the present disclosure. The predicted segment quality information includes predicted segment quality scores, such as... Figure 5 As shown:
[0087] Step 501: Determine the predicted segment corresponding to the updated query features, and obtain the video segment corresponding to the predicted segment.
[0088] Step 502: Determine the prediction distance between the midpoint of the predicted segment and the midpoint of the video segment, and the prediction crossover ratio between the predicted segment and the video segment.
[0089] Step 503: Generate a predicted segment quality score based on the predicted distance and predicted intersection-union ratio.
[0090] There are several methods for constructing a segment quality loss function based on predicted segment quality information. Figure 7This is a schematic diagram illustrating the process of constructing a segment quality loss function in one embodiment of the decoder training method according to the present disclosure, as shown below. Figure 7 As shown:
[0091] Step 701: Determine the segment distance between the midpoint of the predicted segment and the midpoint of the video segment, and the segment intersection-union ratio between the predicted segment and the video segment.
[0092] Step 702: Construct a segment quality loss function based on the deviation information between the predicted distance, the predicted crossover ratio and the corresponding segment distance and segment crossover ratio.
[0093] For example, such as Figure 6 As shown, the query features updated by the relational attention module are input into the transcendental attention module. The transcendental attention module predicts the sampling points in the time dimension and obtains the features of the video segment by weighted summation of the sampled features. The features of the video segment are then fed into each detection head through a feedforward network. In addition to the existing regression head and classification head, a segment quality head is added to estimate the quality of the segments.
[0094] Determine the predicted fragment s corresponding to the updated query features. q s q The corresponding updated query feature f q Define (ζ1,ζ2)=φ(f) q The quality score of a predicted segment is represented by the prediction of two values, ζ1 and ζ2, through a fully connected layer. Here, φ() is a function of a single fully connected layer (φ() can be any function), ζ1 is the predicted distance between the midpoint of the predicted segment and the midpoint of the video segment (action segment), and ζ2 is the predicted intersection-union ratio (IoU) between the predicted segment and the video segment (action segment). The quality score of the predicted segment is defined as ζ = ζ1·ζ2. During training, the segment quality loss function is constructed using the offsets between the midpoint of the predicted segment and its corresponding midpoint of the action segment, as well as their IoU values.
[0095]
[0096] in, The distance between the predicted segment and the nearest ground truth point is the midpoint m of the predicted segment. q The midpoint m of the corresponding video segment (the segment closest to the predicted segment) gt The actual segment distance between them; IoU(s) q ,s gt ) represents the IoU between the predicted fragment and the nearest ground truth, i.e., the predicted fragment s q With the corresponding video clip s gtThe actual segment intersection-union ratio between them.
[0097] During prediction, the classification performance score output by the classification head is multiplied by ζ to obtain the final score of the predicted segment for each query feature. By adding a segment quality head, the product of the deviation and overlap between the predicted segment and the real action is used as the quality score, which is used to jointly determine the predicted segment score during prediction, thereby improving the accuracy of the detection results.
[0098] There are several methods for constructing fragment relationship loss functions. Figure 8 This is a schematic diagram illustrating the process of constructing a segment relation loss function in one embodiment of the decoder training method according to the present disclosure. The segment relation features include the predicted segment intersection-union ratio, such as... Figure 8 As shown:
[0099] Step 801: Determine the crossover ratio (CRO) of the predicted segments corresponding to the updated query features.
[0100] Step 802: Construct a segment relationship loss function based on the cumulative information of the predicted segment intersection-union ratio.
[0101] In one embodiment, during the training phase, an IoU constraint term is introduced to construct the fragment relationship loss function:
[0102]
[0103] Where Lq is the number of query features, s i s j These are the predicted fragments corresponding to the i-th and j-th query features, derived from the output of the regression head predictions from the previous layer; IoU is s i s j The IoU (Intersection over Union) relationship between these two segments is calculated as follows:
[0104]
[0105] By constructing a fragment relationship loss function, redundant query predictions can be suppressed, thereby increasing the probability of obtaining more accurate prediction results.
[0106] Figure 9 This is a flowchart illustrating an embodiment of the target detection method according to the present disclosure, as follows: Figure 9 As shown:
[0107] Step 901: Obtain the trained decoder; wherein the decoder is trained using the training method described above.
[0108] Step 902: Using the decoder and based on the query features, generate classification confidence, regression information to characterize the target location, and predicted fragment quality score.
[0109] In one embodiment, the decoder module includes a Transformer decoder, which comprises a relational attention module, a cross-attention module, two normalization layers, and a feedforward network. The input to the Transformer decoder is a fixed number of trainable query features. The relational attention module performs non-dense attention processing on the query features, and using the cross-attention module and based on the updated query features, generates classification confidence, regression information characterizing the target location, and predicted fragment quality scores via the feedforward network, along with a classification head, a regression head, and a fragment quality head.
[0110] Step 903: Determine the prediction score based on the classification confidence and the predicted segment quality score.
[0111] In one embodiment, the predicted fragment quality score and the classification confidence score are multiplied to determine the final predicted score for each query feature.
[0112] In one embodiment, such as Figure 10 As shown, this disclosure provides a training device 110 for a decoder. The decoder includes a relational attention module and a cross-attention module, etc. The training device 110 for the decoder includes a query set acquisition module 111, a query feature update module 112, a fragment quality determination module 113, a prediction loss determination module 114, and a module adjustment module 115.
[0113] The query set acquisition module 111 uses the relational attention module and, based on the query features, generates a set of salient query features corresponding to the query features. The query feature update module 112 uses the relational attention module and, based on the set of salient query features, updates the query features. For example, the query feature update module 112 uses the relational attention module to perform self-attention calculation on the features within the set of salient query features to update the query features.
[0114] The segment quality determination module 113 uses the cross-attention module and, based on the updated query features, obtains predicted segment quality information corresponding to the updated query features, and constructs a segment quality loss function based on the predicted segment quality information. The prediction loss determination module 114 determines the segment relationship features between the predicted video segments corresponding to the query features and constructs a segment relationship loss function. The module adjustment module 115 adjusts the relationship attention module and the cross-attention module based on the segment quality loss function and the segment relationship loss function.
[0115] In one embodiment, such as Figure 11As shown, the query set acquisition module 111 includes a feature information acquisition unit 1111, a similarity set acquisition unit 1112, a relationship set acquisition unit 1113, and a salient set acquisition unit 1114. The feature information acquisition unit 1111 uses a relationship attention module and, based on the query features, acquires similarity information between various query features and segment relationship feature information between video segments corresponding to each query feature.
[0116] The similarity set acquisition unit 1112 generates a set of similar features corresponding to the query feature based on similarity information. The relation set acquisition unit 1113 generates a set of relation features corresponding to the query feature based on fragment relation feature information. The salient set acquisition unit 1114 generates a set of salient query features based on the similarity feature set, the relation feature set, and the query feature itself.
[0117] In one embodiment, the similarity set acquisition unit 1112 acquires similar query features of the query feature based on similarity information; wherein the similarity between the query feature and the similar query features is greater than a preset similarity threshold. The similarity set acquisition unit 1112 generates a similar feature set based on the similar query features.
[0118] The fragment relationship feature information includes fragment intersection-union ratio (IURR), etc. The relationship set acquisition unit 1113 acquires the relationship query features of the query features based on the fragment IURR; wherein, the fragment IURR between the query features and the relationship query features is greater than a preset IURR threshold. The relationship set acquisition unit 1113 generates a relationship feature set based on the relationship query features.
[0119] The salient set acquisition unit 1114 acquires the relative complement of the similar feature set with respect to the relation feature set. The salient set acquisition unit 1114 takes the union of the relative complement and the query feature itself as the salient query feature set.
[0120] In one embodiment, the predicted fragment quality information includes a predicted fragment quality score; such as Figure 12 As shown, the segment quality determination module 113 includes a segment quality determination unit 1131 and a quality loss determination unit 1132. The segment quality determination unit 1131 determines the predicted segment corresponding to the updated query features and obtains the video segment corresponding to the predicted segment; the segment quality determination unit 1131 determines the predicted distance between the midpoint of the predicted segment and the midpoint of the video segment, and the predicted intersection-union ratio between the predicted segment and the video segment; the segment quality determination unit 1131 generates a predicted segment quality score based on the predicted distance and the predicted intersection-union ratio.
[0121] The quality loss determination unit 1132 determines the segment distance between the midpoint of the predicted segment and the midpoint of the video segment, and the segment intersection-over-union ratio (CIU) between the predicted segment and the video segment. Based on the deviation information between the predicted distance, CIU, and the corresponding segment distance and CIU, the quality loss determination unit 1132 constructs a segment quality loss function.
[0122] In one embodiment, the fragment relationship features include predicted fragment intersection-union ratios (IU / U), and the prediction loss determination module 114 is used to determine the predicted fragment IU / U between predicted fragments corresponding to the updated query features. The prediction loss determination module 114 constructs a fragment relationship loss function based on the cumulative information of the predicted fragment IU / U.
[0123] In one embodiment, such as Figure 13 As shown, this disclosure provides a decoder training apparatus that may include a memory 131, a processor 132, a communication interface 133, and a bus 134. The memory 131 is used to store instructions, and the processor 132 is coupled to the memory 131. The processor 132 is configured to execute the decoder training method described above based on the instructions stored in the memory 131.
[0124] The memory 131 can be a high-speed RAM, non-volatile memory, or a memory array. The memory 131 may also be divided into blocks, and these blocks can be combined into virtual volumes according to certain rules. The processor 132 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the training method of the decoder disclosed herein.
[0125] In one embodiment, this disclosure provides a target detection device 140, including a model acquisition module 141, a detection processing module 142, and a prediction scoring module 143. The model acquisition module 141 acquires a trained decoder; wherein the decoder is trained using the training method described above.
[0126] The detection processing module 142 uses a decoder and, based on query features, generates a classification confidence score, regression information characterizing the target location, and a predicted segment quality score. The prediction score module 143 determines the predicted score based on the classification confidence score and the predicted segment quality score.
[0127] In one embodiment, such as Figure 15As shown, this disclosure provides a target detection device that may include a memory 151, a processor 152, a communication interface 153, and a bus 154. The memory 151 is used to store instructions, and the processor 152 is coupled to the memory 151. The processor 152 is configured to execute the target detection method described above based on the instructions stored in the memory 151.
[0128] The memory 151 can be a high-speed RAM, non-volatile memory, or a memory array. The memory 151 may also be divided into blocks, and these blocks can be combined into virtual volumes according to certain rules. The processor 152 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the target detection method of this disclosure.
[0129] In one embodiment, this disclosure provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method as described in any of the above embodiments.
[0130] The encoder training method, object detection method, device, and storage medium in the above embodiments construct a salient query feature set based on the relationship between query features, and perform self-attention processing on the query features within the salient query feature set to reduce the interference of invalid query features on prediction; by acquiring the quality information of newly added prediction segments and constructing a segment quality loss function, redundant prediction results can be suppressed, and the accuracy of detection results can be improved; by constructing a segment relationship loss function, redundant prediction can be suppressed, making the prediction results more accurate; and the user experience is improved.
[0131] The methods and systems of this disclosure can be implemented in many ways. For example, they can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0132] The description in this disclosure is provided for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit the disclosure to its forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of this disclosure and to enable those skilled in the art to understand this disclosure and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A method for training a decoder, wherein, The decoder includes: Relational attention module and cross-attention module; The training method includes: Using the relational attention module and based on the query features, a set of significant query features corresponding to the query features is generated, which is then used to update the query features using the relational attention module and based on the set of significant query features; Generating a set of significant query features corresponding to the query features includes: Using the relational attention module and based on query features, similarity information between various query features and segment relationship feature information between video segments corresponding to each query feature are obtained; based on the similarity information, a set of similar features corresponding to the query feature is generated; based on the segment relationship feature information, a set of relational features corresponding to the query feature is generated; based on the set of similar features, the set of relational features, and the query feature itself, a set of salient query features is generated; the segment relationship feature information includes: segment intersection-union ratio; Using the cross-attention module and based on the updated query features, predictive fragment quality information corresponding to the updated query features is obtained, and a fragment quality loss function is constructed based on the predictive fragment quality information; wherein, the predictive fragment quality information includes: predictive fragment quality score; Obtain the segment relationship features between predicted video segments corresponding to the query features, and construct a segment relationship loss function; The relationship attention module and the cross-attention module are adjusted based on the fragment quality loss function and the fragment relationship loss function.
2. The method as described in claim 1, wherein generating a set of similar features corresponding to the query feature based on the similarity information comprises: Based on the similarity information, similar query features of the query feature are obtained; wherein the similarity between the query feature and the similar query features is greater than a preset similarity threshold; The similarity feature set is generated based on the similarity query features.
3. The method as described in claim 1, wherein generating a set of relational features corresponding to the query features based on the fragment relational feature information comprises: The relational query feature of the query feature is obtained based on the fragment intersection-union ratio; wherein the fragment intersection-union ratio between the query feature and the relational query feature is greater than a preset intersection-union ratio threshold; The relationship feature set is generated based on the relationship query features.
4. The method as described in claim 1, wherein generating the significant query feature set based on the similarity feature set, the relationship feature set, and the query feature itself comprises: Obtain the relative complement of the similarity feature set with respect to the relation feature set; The union of the relative complement and the query feature itself is taken as the salient query feature set.
5. The method of claim 1, wherein obtaining the predicted fragment quality information corresponding to the updated query features using the cross-attention module and based on the updated query features includes: Determine the predicted segment corresponding to the updated query features, and obtain the video segment corresponding to the predicted segment; Determine the predicted distance between the midpoint of the predicted segment and the midpoint of the video segment, and the predicted intersection-over-union ratio between the predicted segment and the video segment; Based on the predicted distance and the predicted intersection-union ratio, the quality score of the predicted segment is generated.
6. The method of claim 5, wherein constructing the fragment quality loss function based on the predicted fragment quality information comprises: Determine the segment distance between the midpoint of the predicted segment and the midpoint of the video segment, and the segment intersection-union ratio between the predicted segment and the video segment; Based on the deviation information between the predicted distance, the predicted crossover ratio and the corresponding segment distance and segment crossover ratio, the segment quality loss function is constructed.
7. The method as described in claim 1, wherein the fragment relationship features include: Predict the intersection-union ratio of segments; The step of obtaining segment relationship features between predicted video segments corresponding to the query features and constructing a segment relationship loss function includes: Determine the intersection-union ratio of the predicted segments among the predicted segments corresponding to the updated query features; Based on the cumulative information of the predicted segment intersection-union ratio, the segment relationship loss function is constructed.
8. The method of claim 1, wherein updating the query features using the relational attention module and based on the salient query feature set comprises: The relational attention module is used to perform self-attention calculation on the features within the salient query feature set, in order to update the query features.
9. The method according to any one of claims 1 to 8, wherein, The decoder module includes a decoder based on the Transformer architecture.
10. A target detection method, comprising: Obtain a trained decoder; wherein the decoder is trained using the training method described in any one of claims 1 to 9; Using the decoder and based on query features, classification confidence, regression information characterizing the target location, and predicted fragment quality scores are generated. The prediction score is determined based on the classification confidence level and the predicted segment quality score.
11. A training apparatus for a decoder, wherein, The decoder includes: Relational attention module and cross-attention module; The training device includes: The query set acquisition module is used to generate a set of significant query features corresponding to the query features based on the query features and using the relation attention module. The query feature update module is used to update the query features using the relation attention module and based on the salient query feature set; The fragment quality determination module is used to obtain predicted fragment quality information corresponding to the updated query features using the cross-attention module and based on the updated query features, and to construct a fragment quality loss function based on the predicted fragment quality information; The prediction loss determination module is used to determine the segment relationship features between the predicted video segments corresponding to the query features and to construct the segment relationship loss function. The module adjustment module is used to adjust the relationship attention module and the cross attention module according to the segment quality loss function and the segment relationship loss function.
12. A training apparatus for a decoder, comprising: Memory; And a processor coupled to the memory, the processor being configured to perform the method as described in any one of claims 1 to 9 based on instructions stored in the memory.
13. A target detection device, comprising: A model acquisition module is used to acquire a trained decoder; wherein the decoder is trained using the training method described in any one of claims 1 to 9; The detection processing module is used to generate classification confidence, regression information representing the target location, and predicted fragment quality score using the decoder and based on query features; The prediction score module is used to determine the prediction score based on the classification confidence and the predicted segment quality score.
14. A target detection device, comprising: Memory; and a processor coupled to the memory, the processor being configured to perform the method of claim 10 based on instructions stored in the memory.
15. A computer-readable storage medium that non-transitoryly stores computer instructions, which are executed by a processor according to any one of claims 1 to 10.
Citation Information
Patent Citations
Model training method and device
CN110188360A
Target detection method and device based on adaptive decoder
CN114612716A