Video target detection method and system based on alternate decoupling
By using an alternating decoupling method in video object detection, the spatial-time and time-space characteristics of video frames are extracted and coupled, and the problems of limited detection performance and feature collapse in the prior art are solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510101702.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The potential of existing video object detection methods in extracting spatial and temporal information has not been fully tapped, resulting in limited detection performance and may cause feature collapse problems, which in turn leads to false detection or missed detection.
Using the video object detection method based on alternating decoupling, the space-time feature and time-space feature are extracted from the encoded features through the space-time decoupling Transformer decoder and the time-space decoupling Transformer decoder, and coupled through the multi-view structured feature coupling module to obtain the coupled features. At the same time, during the training stage of the object detection model, a text-driven feature imitation learning module is added to alleviate the feature collapse problem.
Through alternating decoupling, the spatial and temporal information of the frame image is more fully mined, which improves the accuracy and robustness of the object detection, alleviates the feature collapse problem, and enhances the training accuracy of the model.
Smart Images

Figure CN120014236A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a video target detection method and system based on alternating decoupling. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In recent years, video object detection has attracted extensive attention in the field of computer vision. Currently, most of the existing video object detection methods either use learnable modules to aggregate spatial and temporal features or follow a continuous space-time paradigm to aggregate features. However, the potential of these methods in extracting spatial and temporal information has not been fully explored, resulting in limited detection performance. More importantly, these feature aggregation methods may lead to feature collapse problems, which in turn cause false detections or missed detections. Summary of the invention
[0004] In order to solve the above problems, the present invention proposes a video target detection method and system based on alternating decoupling, which improves the accuracy of target detection.
[0005] To achieve the above object, the present invention adopts the following technical solution:
[0006] First, a video object detection method based on alternating decoupling is proposed, including:
[0007] Get the video to be detected;
[0008] Divide the video to be detected into frames to obtain multiple frame images;
[0009] The trained target detection model is used to perform target detection on each frame of the image to obtain the target detection result of each frame of the image; wherein, the process of performing target detection on each frame of the image by the trained target detection model is as follows: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
[0010] Furthermore, the target detection model extracts space-time features from the encoded features through a space-time decoupled Transformer decoder; extracts time-space features from the encoded features through a time-space decoupled Transformer decoder; the process of extracting space-time features by the space-time decoupled Transformer decoder includes: extracting spatial features from the encoded features through the space-decoupled Transformer decoder, and extracting time features from the spatial features through the time-decoupled Transformer decoder as space-time features; the process of extracting space-time features from the encoded features by the time-space decoupled Transformer decoder includes: extracting time features from the encoded features through the time-decoupled Transformer decoder, and extracting spatial features from the time features through the space-decoupled Transformer decoder as time-space features.
[0011] Furthermore, the spatially decoupled Transformer decoder includes a spatial mask self-attention mechanism, a spatial deformable self-attention mechanism and a feedforward network module connected in sequence; the features of the input spatially decoupled Transformer decoder are processed through the spatial mask self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the spatial mask self-attention mechanism are interactively fused through the spatial deformable self-attention mechanism to obtain interactively fused features; the interactively fused features obtained by the spatial deformable self-attention mechanism are processed through the feedforward network module, and finally the spatial features are output.
[0012] Furthermore, the time-decoupled Transformer decoder includes a time-masked self-attention mechanism, a time-deformable attention mechanism and a feedforward network module connected in sequence; the features of the input time-decoupled Transformer decoder are processed through the time-masked self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the time-masked self-attention mechanism are interactively fused through the time-deformable attention mechanism to obtain interactively fused features; the interactively fused features obtained by the time-deformable attention mechanism are processed through the feedforward network module, and finally the time features are output.
[0013] Furthermore, the space-time features and the time-space features are coupled through a multi-view structured feature coupling module to obtain coupled features. The process includes:
[0014] Construct multi-view undirected graphs with different viewpoints based on space-time features and time-space features;
[0015] All undirected graphs are merged through the edge-aware multi-graph fuser to obtain the fused graph;
[0016] The vertex features of the fused graph are extracted through a hierarchical adaptive graph convolutional network, and the vertex features are used to update the fused graph to obtain an updated graph;
[0017] Based on the updated graph, the coupling feature is obtained.
[0018] Furthermore, in the target detection model training stage, the category labels pre-annotated in the training data are encoded to obtain text features; a text-driven feature imitation learning module is added to the target detection model, and the input of the text-driven feature imitation learning module is text features and coupling features; the quality score is calculated based on the text features and coupling features; based on the quality score, the coupling features are divided into sample features and non-sample features; based on the loss between non-sample features and sample features belonging to the same category through feature imitation learning, the parameters of the target detection model are updated and trained.
[0019] Secondly, a video object detection system based on alternating decoupling is proposed, including:
[0020] A video acquisition unit, used to acquire the video to be detected;
[0021] A frame division unit, used for dividing the video to be detected into frames to obtain multiple frame images;
[0022] The target detection unit is used to perform target detection on each frame of the image through a trained target detection model to obtain the target detection result of each frame of the image; wherein the process of performing target detection on each frame of the image by the trained target detection model is: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
[0023] In a third aspect, a computer device is provided, the device comprising:
[0024] a processor adapted to execute a computer program;
[0025] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the video target detection method based on alternating decoupling proposed in the first aspect is implemented.
[0026] In a fourth aspect, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the video target detection method based on alternating decoupling proposed in the first aspect.
[0027] In a fifth aspect, a computer program product is proposed, which includes a computer program. When the computer program is executed by a processor, the video object detection method based on alternating decoupling proposed in the first aspect is implemented.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention proposes a video target detection method and system based on alternating decoupling. After acquiring a video to be detected, the method first divides the video to be detected into frames, and then performs target detection on each divided frame image. In the process of performing target detection on each frame image, the space-time features and time-space features of the frame image are respectively extracted, and then the space-time features and the time-space features are coupled to obtain the coupled features of the frame image, and the coupled features are identified and located to obtain the target detection result of the frame image. When obtaining the coupled features, the present invention alternately explores the space and time information in a divide-and-conquer manner, thereby fully mining the space and time information of the frame image, thereby promoting more effective feature aggregation, more fully and comprehensively aggregating the spatiotemporal information, and improving the accuracy of target detection.
[0030] During the target detection model training stage, the method described in the present invention further provides a text-driven feature imitation learning module to generate more discriminative features through the supervision of high-quality features, thereby enhancing the performance of low-quality features, alleviating the feature collapse problem, improving the training accuracy of the target detection model, and further ensuring the accuracy of target detection.
[0031] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0033] Figure 1 A flow chart of a video object detection method based on alternating decoupling disclosed in an embodiment;
[0034] Figure 2 A schematic diagram of the internal structure of the Transformer encoder disclosed in the embodiment;
[0035] Figure 3 A schematic diagram of the internal structure of the alternating decoupled Transformer decoder module disclosed in the embodiment;
[0036] Figure 4A schematic diagram of the internal structure of a spatial-temporal decoupled Transformer decoder disclosed in an embodiment;
[0037] Figure 5 A flow chart of the time-space decoupling Transformer decoder method disclosed in the embodiment;
[0038] Figure 6 A schematic diagram of the internal structure of a multi-view structured feature coupling module disclosed in an embodiment;
[0039] Figure 7 Schematic diagram of the internal structure of the text-driven feature imitation learning module disclosed in the embodiment. DETAILED DESCRIPTION
[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0043] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0044] Example 1
[0045] Since its introduction, the Transformer architecture has become an important breakthrough in the field of deep learning. It can capture global dependencies in sequence data through the self-attention mechanism, and has shown significant advantages over traditional recurrent neural networks and convolutional neural networks when processing long sequence data. However, the training and tuning of the Transformer model usually requires a large amount of data and computing resources, and has high requirements for hardware, which limits its application in scenarios with limited resources.
[0046] Decoupling is a strategy widely used in the fields of computer vision and deep learning. Its core is to split complex problems into multiple independent sub-problems to reduce the complexity of the problem while improving the performance and interpretability of the model.
[0047] In order to take advantage of the Transformer, fully extract temporal and spatial information, and effectively overcome the limitations of the existing technology, the accuracy and robustness of video object detection are improved. This paper proposes a video object detection method based on alternating decoupling, and constructs a new alternating decoupling method from the perspective of feature aggregation, which is used to more fully and comprehensively aggregate spatiotemporal information. In addition, a feature imitation learning module is designed to alleviate the problem of feature collapse, which has not been deeply studied in the field of video object detection.
[0048] The video object detection method based on alternating decoupling disclosed in this embodiment is described in detail. The video object detection method based on alternating decoupling disclosed in this embodiment, first, proposes an alternating decoupling Transformer decoder module, which alternately explores spatial and temporal information in a divide-and-conquer manner, thereby promoting more effective feature aggregation. Secondly, a multi-view structured feature coupling module is designed to adaptively couple spatiotemporal features by taking advantage of graph learning. Thirdly, a text-driven feature imitation learning module is designed to generate more discriminative features through the supervision of high-quality features, thereby enhancing the expressiveness of low-quality features. The feature imitation learning module is only installed in the training phase and will not slow down the testing phase.
[0049] Specifically, the video object detection method based on alternating decoupling disclosed in this embodiment is as follows: Figure 1-Figure 7 As shown, including:
[0050] Get the video to be detected;
[0051] Divide the video to be detected into frames to obtain multiple frame images;
[0052] The trained target detection model is used to perform target detection on each frame of the image to obtain the target detection result of each frame of the image; wherein, the process of performing target detection on each frame of the image by the trained target detection model is as follows: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
[0053] After acquiring the video to be detected, this embodiment inputs all frame images in the video to be detected into the trained target detection model in batches to obtain target detection results for all frame images.
[0054] Among them, the target detection model extracts space-time features from the encoded features through the space-time decoupling Transformer decoder; extracts time-space features from the encoded features through the time-space decoupling Transformer decoder; the process of extracting space-time features by the space-time decoupling Transformer decoder includes: extracting spatial features from the encoded features through the space-decoupling Transformer decoder, and extracting time features from the spatial features through the time-decoupling Transformer decoder as space-time features; the process of extracting space-time features from the encoded features by the time-space decoupling Transformer decoder includes: extracting time features from the encoded features through the time-decoupling Transformer decoder, and extracting spatial features from the time features through the space-decoupling Transformer decoder as time-space features.
[0055] The spatially decoupled Transformer decoder includes a spatial mask self-attention mechanism, a spatial deformable self-attention mechanism and a feedforward network module connected in sequence; the features of the input spatially decoupled Transformer decoder are processed through the spatial mask self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the spatial mask self-attention mechanism are interactively fused through the spatial deformable self-attention mechanism to obtain interactively fused features; the interactively fused features obtained by the spatial deformable self-attention mechanism are processed through the feedforward network module, and finally the spatial features are output.
[0056] The temporal decoupled Transformer decoder includes a temporal masked self-attention mechanism, a temporal deformable attention mechanism and a feedforward network module connected in sequence; the features of the input temporal decoupled Transformer decoder are processed by the temporal masked self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the temporal masked self-attention mechanism are interactively fused by the temporal deformable attention mechanism to obtain interactively fused features; the interactively fused features obtained by the temporal deformable attention mechanism are processed by the feedforward network module, and finally the temporal features are output.
[0057] The space-time features and time-space features are coupled through the multi-view structured feature coupling module to obtain the coupled features. The process includes:
[0058] Construct multi-view undirected graphs with different viewpoints based on space-time features and time-space features;
[0059] All undirected graphs are merged through the edge-aware multi-graph fuser to obtain the fused graph;
[0060] The vertex features of the fused graph are extracted through a hierarchical adaptive graph convolutional network, and the vertex features are used to update the fused graph to obtain an updated graph;
[0061] Based on the updated graph, the coupling feature is obtained.
[0062] The target detection model disclosed in this embodiment is described in detail. The target detection model disclosed in this embodiment includes multiple parallel backbone networks; the input end of each backbone network inputs the frame image to be detected, and the output end of each backbone network outputs the frame features of the corresponding frame image. After the frame features and corresponding position features of each frame image are added, they are input into the Transformer encoder.
[0063] The Transformer encoder outputs the encoded features of each frame feature; the encoded features are input into the alternating decoupled Transformer decoder module.
[0064] The space-time decoupled Transformer decoder inside the alternating decoupled Transformer decoder module outputs space-time features, and the time-space decoupled Transformer decoder outputs time-space features; the space-time features and time-space features are input into the multi-view structured feature coupling module.
[0065] The multi-view structured feature coupling module outputs the coupling features of each frame of the image; the coupling features of each frame of the image are used as the input of the feedforward network, and the output end of the feedforward network outputs the target detection result of each frame of the image.
[0066] This embodiment uses several backbone networks to extract features from several frames of images. One backbone network corresponds to one frame of image, and finally the frame features of all frames are obtained. This embodiment uses ResNet-101 as the backbone network, DETR as the baseline network, and CLIP as the text encoder.
[0067] The frame features are encoded through the Transformer encoder. The process of obtaining the encoded features includes:
[0068] For a given backbone network, the frame features of the nth frame are extracted Among them, c, h, and w represent the dimension, height, and width of the frame feature, respectively;
[0069] Add the frame features and corresponding position features of each frame together to get the added features And the added features Input into the Transformer encoder. The position feature PE is calculated using sine and cosine encoding, and its formula is:
[0070]
[0071] Among them, pso represents the position, i represents the index of the feature dimension, and d model Refers to the dimension of the model, sine encoding is used for even channels and cosine encoding is used for odd channels.
[0072] The Transformer encoder consists of a stack of K = 6 identical layers, such as Figure 2 As shown in the figure. Each layer has two sublayers. The first sublayer is a multi-head self-attention mechanism, and the second sublayer is a feedforward network. Residual connections are used around each of the two sublayers, followed by layer normalization. The output of each sublayer is LayerNorm(x+Sublayer(x)). Among them, Sublayer(x) is the function implemented by the sublayer itself. x represents the vector (or tensor) input to the current sublayer, and LayerNorm(·) represents layer normalization, which is used to normalize the features of each sample.
[0073] like Figure 3 As shown, the alternating decoupled Transformer decoder module includes a space-time decoupled Transformer decoder and a time-space decoupled Transformer decoder; wherein, the space-time decoupled Transformer decoder includes a space-decoupled Transformer decoder and a time-decoupled Transformer decoder connected in sequence; the time-space decoupled Transformer decoder includes a time-decoupled Transformer decoder and a space-decoupled Transformer decoder connected in sequence; N frames of target queries are respectively input into the input end of the space-time decoupled Transformer decoder in the space-time decoupled Transformer decoder and the input end of the time-space decoupled Transformer decoder in the time-space decoupled Transformer decoder; the time-decoupled Transformer decoder in the space-time decoupled Transformer decoder outputs space-time features; the space-decoupled Transformer decoder in the time-space decoupled Transformer decoder outputs time-space features.
[0074] like Figure 4As shown, the spatial-temporal decoupled Transformer decoder includes a spatial decoupled Transformer decoder and a temporal decoupled Transformer decoder. The spatial decoupled Transformer decoder includes N parallel branches, each of which includes a spatial mask self-attention mechanism, a spatial deformable self-attention mechanism, and a feedforward network module connected in sequence; the input end of the spatial mask self-attention mechanism of the nth branch is used to input the target query of the nth frame; the output end of the spatial mask self-attention mechanism of the nth branch is connected to the input end of the spatial deformable self-attention mechanism. The input end of the spatial deformable self-attention mechanism also inputs the encoding features of the nth frame; the output end of the spatial deformable self-attention mechanism is connected to the input end of the feedforward network, and the output end of the feedforward network outputs the spatial features of the nth frame; the spatial features of N frames are input into the time-decoupled Transformer decoder, and the time-decoupled Transformer decoder includes a time-masked self-attention mechanism, a time-deformable attention mechanism and a feedforward network module connected in sequence; the input end of the time-masked self-attention mechanism is used to input the spatial features of N frames; the output end of the time-masked self-attention mechanism is connected to the input end of the time-deformable attention mechanism, and the input end of the time-deformable attention mechanism also inputs the encoding features of N frames; the output end of the time-deformable attention mechanism is connected to the input end of the feedforward network; the output end of the feedforward network outputs the spatial-temporal features of N frames.
[0075] The encoded features are input into the space-time decoupled Transformer decoder, and the process of outputting space-time features includes:
[0076] First, each frame’s target query is input into the spatial mask self-attention layer to explore the spatial dependencies within the frame. The target query is a set of learnable vectors used to generate the final detection results. These query vectors do not depend on the specific content in the image, but interact with the image features through the Transformer structure to obtain the detection information of each target.
[0077] Most adjacent frames contain similar appearance information. Directly using the native self-attention mechanism not only introduces redundant features and reduces model performance, but also brings additional costs in terms of computation and memory. To this end, this embodiment designs a spatial mask self-attention mechanism, which processes the features of the input spatial decoupled Transformer decoder through the spatial mask self-attention mechanism to obtain the processed features. The spatial mask self-attention mechanism SpatMSelfA has the following formula:
[0078]
[0079] Where q∈Ω q represents the query element, k∈Ωk represents the key element, Ω q and Ω k Represents a collection of query elements and key elements respectively. and They represent the features of query element q and key element k respectively, and D is the feature dimension. s represents the index of the attention head, S represents the total number of attention heads, and K represents the total number of key elements. and represents the learnable mapping weight, D v =D / S. The symbol “·” indicates a scalar multiplication operation. sqk ∈{0, 1} represents the mask weight of the kth sampled key element in the sth attention head. When the key element is within the defined local window, its value is 1; otherwise, its value is 0. sqk ∈[0, 1] represents the attention weight of the kth sampled key element in the sth attention head and is normalized over all key elements. sqk The calculation formula is:
[0080]
[0081] Among them, ∝ means proportional to the operation, Represents a transpose operation. and represents the learnable mapping weights.
[0082] Next, the processed features output by the spatial mask self-attention layer are input into the spatial deformable attention layer together with the encoded features, with the goal of fusing the encoded features with the target query by exploring the spatial dependencies between them. In particular, the spatial deformable attention layer interactively fuses the processed features obtained by the spatial mask self-attention mechanism with the encoded features through the spatial deformable self-attention mechanism to obtain the interactively fused features. The spatial deformable self-attention mechanism SpatDeformA can be expressed as:
[0083]
[0084] Among them, p q represents the two-dimensional reference point of the query element q, Δp sqk is the sampling offset of the kth sampling key element in the sth attention head. Since p q +Δp sqk is in fractional form, so bilinear interpolation is used to calculate x(p q +Δp sqk ). Finally, the output of the spatially deformable attention layer is fed into a feed-forward network to generate spatial features.
[0085] Since detection using only spatial information is difficult to cope with challenges such as blur or interference from similar objects, it is necessary to introduce temporal information to enrich the representation of spatial features. To this end, the spatial features of multiple frames are input into the temporal mask self-attention layer to explore the temporal dependencies across frames. Specifically, the temporal mask self-attention layer processes the features of the input temporally decoupled Transformer decoder through the temporal mask self-attention mechanism to obtain the processed features. The temporal mask self-attention mechanism TempMselfA can be defined as:
[0086]
[0087] Where n is the index of the input frame and N is the total number of input frames. snqk ∈{0, 1) represents the mask weight of the kth sampled key element in the sth attention head of the nth frame. When the key element is within the defined local window, the value is 1; otherwise, the value is 0. snqk ∈[0, 1] represents the attention weight of the kth sample key element in the sth attention head of the nth frame. snqk The calculation formula is:
[0088]
[0089] in, Represents the intermediate feature, n represents the index of the input frame, and the interpretation of other symbols can refer to formula (4). NK elements are sampled from the feature map of N frames in a normalized manner, so that multi-frame dependencies can be modeled.
[0090] Subsequently, the output features of the temporal mask self-attention layer are input into the temporal deformable attention layer, and the encoded features are also input into the temporal deformable attention layer, with the goal of fusing the encoded features with the target query by exploring the temporal dependency between them to obtain interactively fused features. In particular, the temporal deformable attention layer interactively fuses the processed features obtained by the temporal mask self-attention mechanism with the encoded features through the temporal deformable attention mechanism to obtain interactively fused features. The temporal deformable attention mechanism TempDeformA can be expressed as:
[0091]
[0092] The meanings of the symbols in formula (8) can refer to those in formulas (6) and (7). Finally, the output of the temporal deformable attention layer is input into the feedforward network to generate spatial-temporal features.
[0093] like Figure 5 As shown, the time-space decoupled Transformer decoder module includes:
[0094] The internal structure of the temporal decoupled Transformer decoder and the spatial decoupled Transformer decoder is the same as above, such as Figure 4 As shown in the figure; the temporal decoupled Transformer decoder and the spatial decoupled Transformer decoder are connected in sequence; the target query of N frames is input into the input end of the temporal decoupled Transformer decoder; the output end of the temporal decoupled Transformer decoder is connected to the input end of the spatial decoupled Transformer decoder; the spatial decoupled Transformer decoder outputs the temporal-spatial features of N frames.
[0095] The encoded features are input into the time-space decoupled Transformer decoder. The process of obtaining time-space features includes:
[0096] The object query for each frame is fed into the temporal decoupled transformer decoder, and then the output of the temporal decoupled transformer decoder is fed into the spatial decoupled transformer decoder. In terms of mining spatial and temporal information, the opposite workflow of the spatial-temporal decoupled transformer decoder is implemented. The goal of the temporal-spatial decoupled transformer decoder is to generate new feature representations that are complementary to the output of the spatial-temporal decoupled transformer decoder.
[0097] The multi-view structured feature coupling module includes a multi-view undirected graph of T different views, an edge-aware multi-graph fuser, and a hierarchical adaptive graph convolutional network connected in sequence; the input of the first multi-view undirected graph is the N-frame spatial-temporal features and temporal-spatial features output by the alternating decoupled Transformer decoder module; the output of the T-th multi-view undirected graph is connected to the input of the edge-aware multi-graph fuser. The first multi-view undirected graph and the T-th multi-view undirected graph share the same set of nodes, but provide diversified relationship representations through different edge sets and adjacency matrices; the edge-aware multi-graph fuser outputs the fused graph and serves as the input of the hierarchical adaptive graph convolutional module, which outputs the updated graph, which contains the information of the enhanced spatial-temporal and temporal-spatial features in different dimensions; finally, the enhanced spatial-temporal features and temporal-spatial features are added element by element to generate coupling features.
[0098] In this embodiment, the time-space feature and the space-time feature are input into the multi-view structured feature coupling module. The specific process of obtaining the coupling feature includes:
[0099] Construct T undirected multi-views with different perspectives based on time-space features and space-time features in represents the vertex set, ε t and A t They represent the edge set and the adjacency matrix in the t-th view, respectively. When constructing multiple views, each input feature is regarded as a vertex, and the edge is determined according to the association between vertices.
[0100] Different from the method of directly building on multiple views and using graph learning operations to integrate features, this embodiment proposes a simple and efficient edge-aware multi-image fuser that merges multiple views into a single image, so that the graph learning operation only needs to be performed once. The calculation of the edge-aware multi-image fuser can be expressed as:
[0101]
[0102] in, Represents the vertex v in the fused graph i and v j The edge between. is a multilayer perceptron, and Concat(·,·,·) represents the concatenation operation. Represents the vertex v in the information view, semantic view and distance view respectively i and v j According to formula (9), using to link the vertices to obtain the final fused graph, whose adjacency matrix can be calculated by using the Softmax function on all edges.
[0103] Based on the fused graph, this embodiment improves the native graph convolution network and designs a hierarchical adaptive graph convolution network to recalculate the adjacency matrix based on the newly learned vertex features of each graph convolution layer. The hierarchical adaptive graph convolution network designed in this patent adopts a two-layer adaptive graph convolution network with a ReLU activation function. The enhanced vertex feature matrix G can be calculated as:
[0104]
[0105] Among them, X represents the input feature matrix, that is, the spatial-temporal features and temporal-spatial features of N frames. and Denote the normalized graph Laplace regularization matrices of the first and second layers, respectively. Θ (1) and θ (2) Denote the weight matrices of the first and second layers respectively. The hierarchical adaptive graph convolutional network outputs the enhanced feature matrix G, which contains the information of the enhanced space-time and time-space features in different dimensions. Finally, the enhanced space-time features and time-space features are added element by element to generate the coupled features.
[0106] After the target detection model is constructed in this embodiment, the constructed target detection model is trained using a training set, wherein the images in the training set are video frame images that have been labeled with target detection labels; the training set is input into the target detection model, and the model is trained. When the loss function value of the model no longer decreases, or the number of iterations reaches a set number, the training is stopped to obtain a trained target detection model.
[0107] The training set refers to: given a video dataset, the dataset contains several videos, each video contains several video frames The video object detection method of this embodiment takes N consecutive frames (the default setting is 30 frames) as input and outputs the detection results of all input frames at one time.
[0108] In the target detection model training stage, this embodiment encodes the category labels pre-annotated in the training data to obtain text features; adds a text-driven feature imitation learning module to the target detection model, and the input of the text-driven feature imitation learning module is text features and coupling features; calculates the quality score according to the text features and the coupling features; divides the coupling features into sample features and non-sample features according to the quality score; and updates and trains the parameters of the target detection model based on the loss between non-sample features and sample features belonging to the same category through feature imitation learning to improve the performance of the model and the discrimination ability of features.
[0109] The text-driven feature imitation learning module includes a feature quality indicator and an embedding space. The input end of the feature quality indicator inputs coupling features and text features. The text features are formed through a text encoder based on the category label data of the video frame. The output end of the feature quality indicator outputs sample features and non-sample features, maps the sample features and non-sample features to the embedding space, and updates the coupling features through back propagation.
[0110] The data processing process of the text-driven feature imitation learning module disclosed in this embodiment includes:
[0111] First, the text features and coupling features are input into the feature quality indicator to generate sample features and non-sample features. In order to better construct sample features, the feature quality indicator adopts a simple and effective metric learning to evaluate the quality score of the features, characterizes the quality of the coupling features by the quality score, and selects the coupling features with a quality score higher than the threshold as sample features to form the sample feature set Ω e , which can be expressed as:
[0112]
[0113]
[0114] Among them, ci represents the i-th coupled feature output by the multi-view structured feature coupling module, y j represents the jth text feature generated by the text encoder processing the category label of the video frame. FQ(·,·) represents the feature quality score, which is an indicator that enables the capture of high-quality coupled features with rich information. θ represents a scalar threshold. γ and ∏ represent the learning parameters of the two linear mapping layers. represents the transposition operation, ρ(·) represents the nonlinear activation function, and Concat(·,·) represents the concatenation operation. According to formula (12), the coupled features with quality scores lower than the threshold θ are regarded as non-sample features, forming a non-sample feature set
[0115] After obtaining the sample features and non-sample features, the text-driven feature imitation learning module uses a single-layer perceptron with negligible cost to map them into the embedding space, and the feature dimension is reduced to 128, which reduces the memory burden. Subsequently, this embodiment designs a feature imitation learning loss function to support non-sample features to be mapped near sample features related to them, while being separated from sample features of other classes, which helps to generate multiple discriminative feature representations, where sample features related to non-sample features refer to sample features that belong to the same category as non-sample features. The feature imitation learning loss can be formulated as:
[0116]
[0117] in, is the loss between the i-th non-sample feature and the j-th sample feature, Ω e and represent sample feature and non-sample feature sets respectively. represents the cosine similarity function, o i and r j They represent the i-th non-sample feature and the j-th sample feature belonging to the same category respectively. τ represents the temperature parameter. Through formula (13), the alternating decoupled Transformer imitation network proposed in the present invention can improve the discriminability of features by using high-quality features to guide low-quality features to learn. The text-driven feature imitation learning module is only installed in the training phase and will not reduce the speed of testing.
[0118] The video object detection method based on alternating decoupling disclosed in this embodiment proposes a new alternating decoupling method from the perspective of feature aggregation through three key technical improvements, which more fully and comprehensively aggregates spatiotemporal information, thereby improving the performance of video object detection. Its overall architecture is shown in the figure below: Figure 1As shown. Given an input video sequence, each video frame is first feature extracted by a weight-sharing backbone network to obtain the corresponding frame features, and then these frame features and the corresponding position features are added and input into the Transformer encoder, which outputs the coded features of each frame. Next, the coded features are input into the alternating decoupling Transformer decoder module, which outputs the space-time features and time-space features. Subsequently, the space-time features and time-space features are input into the multi-view structured feature coupling module to obtain the coupled features of each frame. Finally, the coupled features of each frame are input into the feedforward network for identification and positioning to obtain the target detection results of each frame. In view of the module advantages designed in this embodiment, a parallel detection method is adopted in the framework, which can simultaneously detect targets on all input frames, so that the video target detection method has the ability of real-time reasoning.
[0119] The video target detection method based on alternating decoupling disclosed in this embodiment also designs a feature imitation learning paradigm in the target detection model training stage to alleviate the feature collapse problem, which has not been deeply studied in the field of video target detection.
[0120] The video target detection method based on alternating decoupling disclosed in this embodiment has broad application prospects in multiple fields, especially in scenarios such as intelligent monitoring and virtual reality that have high requirements for high precision and real-time processing. For example, in intelligent monitoring, it can improve the accuracy of target recognition and improve public safety; in virtual reality, it can improve the accuracy of scene recognition and target tracking and enhance user experience. In addition, this technology can also be applied to fields such as robots and drones to improve their autonomy and execution efficiency. Through the application of this technology, these industries will be significantly improved in accuracy, efficiency and safety.
[0121] It should be noted that all data is obtained in compliance with laws and regulations and user consent, and the data is used legally.
[0122] Example 2
[0123] In this embodiment, a video object detection system based on alternating decoupling is disclosed, comprising:
[0124] A video acquisition unit, used to acquire the video to be detected;
[0125] A frame division unit, used for dividing the video to be detected into frames to obtain multiple frame images;
[0126] The target detection unit is used to perform target detection on each frame of the image through a trained target detection model to obtain the target detection result of each frame of the image; wherein the process of performing target detection on each frame of the image by the trained target detection model is: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
[0127] The present invention also discloses a computer device, which includes:
[0128] a processor adapted to execute a computer program;
[0129] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the video object detection method based on alternating decoupling disclosed in Example 1 is implemented.
[0130] The present invention also discloses a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded by a processor and executing the video target detection method based on alternating decoupling disclosed in Example 1.
[0131] The present invention also discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, the video target detection method based on alternating decoupling disclosed in Example 1 is implemented.
[0132] The method disclosed in Example 1 can be directly embodied as a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0133] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0134] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A video object detection method based on alternating decoupling, characterized in that: include: Get the video to be detected; Divide the video to be detected into frames to obtain multiple frame images; The trained target detection model is used to perform target detection on each frame of the image to obtain the target detection result of each frame of the image; wherein, the process of performing target detection on each frame of the image by the trained target detection model is as follows: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
2. The video object detection method based on alternating decoupling as claimed in claim 1, characterized in that: The object detection model extracts space-time features from frame features through a space-time decoupled Transformer decoder; and extracts time-space features from encoded features through a time-space decoupled Transformer decoder; The process of extracting space-time features by the space-time decoupled Transformer decoder includes: extracting spatial features from the encoding features through the space-time decoupled Transformer decoder, and extracting time features from the spatial features through the time-time decoupled Transformer decoder as space-time features; the process of extracting space-time features from frame features by the time-space decoupled Transformer decoder includes: extracting time features from frame features through the time-time decoupled Transformer decoder, and extracting spatial features from the time features through the space-time decoupled Transformer decoder as time-space features.
3. The video object detection method based on alternating decoupling as claimed in claim 2, characterized in that: The spatially decoupled Transformer decoder includes a spatial mask self-attention mechanism, a spatial deformable self-attention mechanism and a feedforward network module connected in sequence; the features of the input spatially decoupled Transformer decoder are processed through the spatial mask self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the spatial mask self-attention mechanism are interactively fused through the spatial deformable self-attention mechanism to obtain interactively fused features; the interactively fused features obtained by the spatial deformable self-attention mechanism are processed through the feedforward network module, and finally the spatial features are output.
4. The video object detection method based on alternating decoupling as claimed in claim 2, characterized in that: The temporal decoupled Transformer decoder includes a temporal masked self-attention mechanism, a temporal deformable attention mechanism and a feedforward network module connected in sequence; the features of the input temporal decoupled Transformer decoder are processed by the temporal masked self-attention mechanism to obtain processed features; the processed features and encoded features obtained by the temporal masked self-attention mechanism are interactively fused by the temporal deformable attention mechanism to obtain interactively fused features; the interactively fused features obtained by the temporal deformable attention mechanism are processed by the feedforward network module, and finally the temporal features are output.
5. The video object detection method based on alternating decoupling as claimed in claim 1, characterized in that: The space-time features and time-space features are coupled through the multi-view structured feature coupling module to obtain the coupled features. The process includes: Construct multi-view undirected graphs with different viewpoints based on space-time features and time-space features; All undirected graphs are merged through the edge-aware multi-graph fuser to obtain the fused graph; The vertex features of the fused graph are extracted through a hierarchical adaptive graph convolutional network, and the vertex features are used to update the fused graph to obtain an updated graph; Based on the updated graph, the coupling feature is obtained.
6. The video object detection method based on alternating decoupling as claimed in claim 1, characterized in that: During the object detection model training phase, the pre-annotated category labels in the training data are encoded to obtain text features; A text-driven feature imitation learning module is added to the target detection model. The input of the text-driven feature imitation learning module is text features and coupling features. The quality score is calculated based on the text features and coupling features. According to the quality score, the coupling features are divided into sample features and non-sample features. The parameters of the target detection model are updated and trained based on the loss between non-sample features and sample features belonging to the same category through feature imitation learning.
7. A video object detection system based on alternating decoupling, characterized in that: include: A video acquisition unit, used to acquire the video to be detected; A frame division unit, used for dividing the video to be detected into frames to obtain multiple frame images; The target detection unit is used to perform target detection on each frame of the image through a trained target detection model to obtain the target detection result of each frame of the image; wherein the process of performing target detection on each frame of the image by the trained target detection model is: extracting frame features of the frame image; encoding the frame features to obtain encoding features; extracting space-time features and time-space features from the encoding features; coupling the space-time features and the time-space features to obtain the coupling features of the frame image; identifying the coupling features to obtain the target detection result of the frame image.
8. An electronic device, characterized in that: The device comprises: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the video target detection method based on alternating decoupling according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the video object detection method based on alternating decoupling according to any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the video object detection method based on alternating decoupling according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Video target segmentation method based on space-time decoupling attention mechanism
CN116416553A
Space-time fusion multi-target tracking method, device, equipment and medium
CN117314965A
Small target detection and trajectory prediction method for substation patrol scene
CN118608119A
Multi-resolution transformer for video quality assessment
WO2023182987A1
Video inpainting method, related apparatus, device and storage medium
WO2023193521A1
Cited By
Video target detection method and system based on multi-scale perception diffusion
CN121190746A