An end-to-end video spatiotemporal visual localization system based on visual language Transformer
Through an end-to-end system based on the visual language Transformer, the detection frame and time information are generated directly from the video frames and text descriptions, solving the limitations and two-stage separation problems of pre-trained detectors in the prior art, and achieving efficient video spatio-temporal positioning effect.
Patent Information
- Application Number
- CN202111100948.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2041-09-18
AI Technical Summary
The prior art requires pre-training of object detectors in video space-time visual positioning tasks, resulting in limited detection capabilities and high training costs. The two-stage method separates the framework into independent timing and spatial positioning tasks, making it difficult to efficiently complete joint positioning in time and space.
The end-to-end system based on the visual language Transformer is adopted, including visual information encoding, text embedding, spatiotemporal visual positioning and spatiotemporal trajectory generation modules. Through cross-modal feature learning and spatiotemporal analysis, detection frames and temporal information are generated directly from video frames and text descriptions to build an end-to-end unified framework.
It realizes efficient video space-time positioning without relying on pre-trained object detectors, improves detection accuracy, avoids high-cost pre-training processes, and completes visual positioning in time and space, which is better than the traditional two-stage method.
Smart Images

Figure CN113849668B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia technology, and in particular to an end-to-end video spatiotemporal visual positioning system based on a visual language Transformer. Background Art
[0002] Video spatiotemporal visual localization is a new and extremely challenging visual language task. Given an unedited video, this task generates a spatiotemporal trajectory block (a series of visual localization boxes) to locate the target in the video based on the required language description of the target. Unlike existing image visual localization tasks, spatiotemporal visual localization requires the detection target to be located in both time and space. In addition, how to efficiently utilize visual and language information to complete cross-modal learning is the key to accurately locating the detection target. Among them, different people performing similar actions in the same scene is a very challenging task scenario.
[0003] Spatial localization in images and videos is a closely related visual localization task. Most existing work first uses pre-trained object detectors to generate a few possible object detection candidates. These methods have certain limitations: 1) spatial detection capabilities are limited by the quality of the object detection candidates; 2) it is difficult to use pre-trained object detectors to generate detection boxes for new classes; and 3) pre-training is expensive. Currently, no work on video localization has attempted to remove the pre-trained detector.
[0004] Furthermore, given that the task of spatiotemporal visual localization in videos requires localization of the detected target in both time and space, existing methods are all two-stage approaches. These first complete temporal visual localization, identifying the start and end times of the detected target. Based on this, spatial visual localization is then performed. However, this two-stage approach results in the overall framework being closer to two independent networks completing their respective subtasks.
[0005] Furthermore, given that the task of spatiotemporal visual localization in videos requires localization of the detected target in both time and space, existing methods are all two-stage approaches. These first complete temporal visual localization, identifying the start and end times of the detected target. Based on this, spatial visual localization is then performed. However, this two-stage approach results in the overall framework being closer to two independent networks completing their respective subtasks.
[0006] Therefore, how to provide a system based on end-to-end visual language translation to complete the video visual positioning task without pre-training target detectors is an urgent problem to be solved in this field. Summary of the Invention
[0007] In view of this, the present invention provides an end-to-end video spatiotemporal visual positioning system based on the visual language Transformer, which can simultaneously complete visual positioning in time and space and learn better feature representation to achieve better positioning effect.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A visual language Transformer-based end-to-end video spatiotemporal visual positioning system comprises a visual information encoding module, a text embedding module, a spatiotemporal visual positioning module and a spatiotemporal trajectory generation module; the visual information encoding module and the text embedding module are connected to the spatiotemporal visual positioning module; the spatiotemporal visual positioning module is connected to the spatiotemporal trajectory generation module; the visual information encoding module is used to obtain visual features of a detection target from a video frame; the text embedding module is used to extract text encoding of the detection target from a query text; the spatiotemporal visual positioning module is used to learn the interaction features between the visual features and the text encoding, and to perform spatial and temporal positioning of the detection target to obtain detection box information and time start and end information; the spatiotemporal trajectory generation module is used to combine the generated detection box information in the time domain and the space domain to obtain a spatiotemporal trajectory block containing the detection target.
[0010] Furthermore, the spatiotemporal visual positioning module includes a cross-modal feature learning module and a spatiotemporal analysis core module; the cross-modal feature learning module acquires text encoding and visual features, generates text-guided visual features and visual-guided text features; the spatiotemporal analysis core module locates the generated text-guided visual features in time and space.
[0011] Furthermore, special text tags ['GLS'] and ['SEP'] are added to the query text and placed at the beginning and end of the query text respectively. The text embedding module obtains the text encoding with special tags based on the query text with special tags; the cross-modal learning module obtains the text encoding with special text tags to obtain global text features.
[0012] Furthermore, the cross-modal feature learning module includes a visual branch module and a text branch module, and the visual branch module and the text branch module interactively learn to obtain text-guided visual features and visually guided text information.
[0013] Furthermore, a spatiotemporal combination decomposition module is constructed in the visual branch module to retain spatial information.
[0014] Furthermore, the spatiotemporal combination decomposition module includes a temporal pooling module, a spatial pooling module, a combination module, a multi-head attention module, a decomposition module, a replication module and a normalization module; the temporal pooling module is used to collect visual features to generate T×C preliminary temporal features, where T represents the number of video frames, C represents the number of feature map channels, H represents height, and W represents width; the spatial pooling module is used to collect visual features to generate preliminary spatial features with a shape of HW×C, and the combination module is used to connect the preliminary temporal features and preliminary spatial features in the feature dimension to form a combined visual feature with a size of (T+HW)×C; the multi-head attention module is used to collect visual features based on the temporal pooling module. Attention operation is performed on the combined visual features and text features to generate preliminary text-guided visual features; the decomposition module is used to generate text-guided temporal features and text-guided spatial features based on the preliminary text-guided visual features; the copy module is used to copy the text-guided temporal features HW times and the text-guided spatial features T times to obtain copied temporal features and copied spatial features with a size of T×HW×C; the normalization module is used to normalize the result of adding the copied temporal features, copied spatial features and visual input features to generate intermediate visual features; the output of the last layer is the text-guided visual features.
[0015] Furthermore, the spatiotemporal analysis core module includes a spatial visual positioning branch module and a temporal visual positioning branch module; the spatial visual positioning module generates the center and size of the detection box based on the text-guided visual features; the temporal visual positioning branch module generates the start and end prediction scores based on the text-guided visual features.
[0016] Furthermore, the spatial visual positioning branch includes a deconvolution layer, a first detection network head and a second detection network head; the deconvolution layer is three-layer, deconvolution is performed on the text-guided visual features, and spatial upsampling is performed to obtain upsampled features; the first detection network head and the second detection network head are both composed of a 3×3 convolution base layer and a 1×1 convolution base layer; the first network head selects to obtain the upsampled features to generate a heat map, and the position with the highest probability value in the heat map is the center of the detection box, and the second network head obtains the size of the detection box by regressing the upsampled features.
[0017] Furthermore, the temporal visual localization branch module includes a spatial pooling layer, a temporal convolution layer, a multi-layer perceptron layer, and a computational activation layer; the spatial pooling layer uses spatial average pooling on the text-guided visual features to obtain text-guided global visual features; the text-guided global visual features are convolved through two parallel temporal convolution blocks to obtain a starting visual feature and an ending visual feature, respectively;
[0018] The multi-layer perception acquires global text features to obtain text features in a C-dimensional common feature space; the calculation activation layer calculates the correlation between the starting visual features and the ending visual features and the text features in the common feature space and activates them using an activation function to obtain the starting and ending prediction scores.
[0019] Furthermore, the spatiotemporal visual localization module combines three focal losses and one L1 loss as loss functions for training.
[0020] The beneficial effects achieved by the present invention are:
[0021] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses an end-to-end video spatiotemporal visual positioning system based on the visual language Transformer, so that the present invention proposes a unified framework for end-to-end video spatiotemporal visual positioning that does not use predictive training target detectors, while realizing visual positioning in time and space to achieve better positioning effects, and also solves the problem that the existing spatial detection capability is limited by the quality of the target detection candidate box, and avoids the high-cost pre-training process of the detector; constructs a spatiotemporal visual positioning module, and retains spatial information in the cross-modal feature learning module. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0023] Figure 1 The accompanying figure is a structural diagram of the end-to-end video spatiotemporal visual positioning system based on the visual language Transformer of the present invention;
[0024] Figure 2 The attached figure is a structural diagram of the cross-modal learning module;
[0025] Figure 3 The accompanying figure is a structural diagram of the spatiotemporal combination decomposition module;
[0026] Figure 4 The attached figure is a structural diagram of the spatiotemporal visual positioning module. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] like Figure 1 The embodiment of the present invention discloses an end-to-end video spatiotemporal visual positioning system (STVGBert) based on a visual language Transformer, including a visual information encoding module (Image Encoder), a word embedding module (WordEmbedding), a spatiotemporal visual positioning module (STVGBert-core) and a spatiotemporal trajectory generation module (TubeGeneration); the visual information encoding module and the word embedding module are connected to the spatiotemporal visual positioning module; the spatiotemporal visual positioning module is connected to the spatiotemporal trajectory generation module; the visual information encoding module obtains visual features from video frames; the word embedding module extracts text codes from query texts; the spatiotemporal visual positioning module interactively learns visual language features of visual features and text codes and generates detection box information and start and end time information; the spatiotemporal trajectory generation module obtains a spatiotemporal trajectory block (Object Tube B) containing a detection target according to the generated detection box information and start and end time information.
[0029] In another embodiment, a video containing K*T frames is divided into K non-overlapping video slices, each video slice contains T frames of video, and the video is defined as in Represents the kth video slice. The visual information encoding module extracts visual features from the video frame. The text embedding module extracts text encoding from the query text. The visual information encoding module uses ResNet-101 as the visual information encoder to extract visual features. For each video frame, a 4-layer residual block is used to convert its shape into HW×C, where H, W, and C represent the height, width, and number of feature map channels, respectively. The visual information encoder stacks the features extracted from each frame to form a video slice feature, which is recorded as The slice features are input as visual features into the spatiotemporal visual localization module.
[0030] In another embodiment, special text tags ['GLS'] and ['SEP'] are added to the query text, placed at the beginning and end of the query text, respectively. The word embedding module maps each word in the query text to a word vector. Each word vector is considered an input textual token. The word vector with the special textual token is the global input textual token.
[0031] like Figure 2 Figure 3 In another embodiment, the spatiotemporal visual positioning module includes a cross-modal learning module (ST-ViLBERT) for interactively learning visual language features for combination. The cross-modal learning module consists of a visual branch module (Visual Branch) and a textual branch module (Textual Branch). Each branch adopts a multi-layer transformer encoding layer structure and learns interactive features through a multi-head attention layer (Multi-headAttention) in the transformer encoding layer. The learning scheme is: the text feature input obtained in the previous layer is used as the key and value in the visual branch module to participate in the multi-head attention calculation, and the visual feature input obtained in the previous layer is used as the key (key) and value (value) in the text branch module to participate in the multi-head attention calculation, and finally generates a text-guided visual feature (Text-guidedvisual feature).
[0032] In another embodiment, a spatio-temporal combination and decomposition module (STCD) is constructed to replace the multi-head attention layer and the add and normalization layer (Add&Norm) in the visual branch module, so that the spatial information is retained during the learning process of the cross-modal feature module. In the STCD module, the visual features are first subjected to spatial average pooling (Spatial Pooling) and temporal average pooling (Temporal Pooling) to produce initial temporal features (initial temporal feature) with a shape of T×C and initial spatial features (initial spatial feature) with a shape of HW×C. The combination module (Combination) connects these two features in the feature dimension to form a combined visual feature (Combined Bisual Feature), the size of which is (T+HW)×C. The combined visual features are then passed to the text branch as vectors, serving as keys and values in the multi-head attention block in the text branch. The combined visual features are also passed to the multi-head attention block in the vision branch, where they participate in an attention operation with the text input (Textual Pooling), generating preliminary text-guided visual features of size (T+HW)×C. The preliminary text-guided visual features are decomposed by the decomposition module (Decomposition) into text-guided temporal features (of size T×C) and text-guided spatial features (of size HW×C). These two features are replicated HW and T times, respectively, by the replication module (Replication) to match the visual feature size, forming replicated temporal features and replicated spatial features. Finally, the replicated temporal features, replicated spatial features, and visual features are summed and normalized (Norm) to generate intermediate visual features of size T×HW×C.
[0033] Among them, the transformer encoding layer is composed of a multi-head attention layer, an addition and normalization layer (Add&Norm), a feedforward neural network layer (Feed Forward), and a second addition and normalization layer connected in sequence.
[0034] In addition, since the cross-modal learning feature module is composed of a multi-layer transformer structure, each layer takes visual features as input and outputs new visual features. The visual features generated in the middle layer are the intermediate visual features, and the outputs of the last layer of the visual branch and the text branch are respectively used as the text-guided visual features F and F. tv ∈R T×HW×C and visually guided text features.
[0035] like Figure 4In another embodiment, the spatiotemporal visual positioning module further includes a spatiotemporal analysis core module; the spatiotemporal analysis core module includes a spatial visual positioning branch and a temporal visual positioning branch.
[0036] In the spatial visual localization branch, the text-guided visual features are first passed through a three-layer deconvolution layer for spatial upsampling by a factor of 8 to generate upsampled features. Two parallel detection network heads obtain the sampled features to generate bounding box centers (BB Centrers) and bounding box sizes (BBSizes) respectively.
[0037] Among them, the two parallel detection network heads are composed of a 3×3 convolutional layer for feature extraction and a 1×1 convolutional layer for dimensionality reduction. Their inputs are all upsampled features. The first detection network head outputs a heat map A∈R 8H ×8W The value in the heat map A represents the probability value of the detection box center at the corresponding position, and the corresponding spatial position with the highest probability value is selected as the center of the detection box. The second detection network head is used to regress the size of the detection box, and the upper left and lower right script values of the detection box are calculated based on the size.
[0038] In the temporal visual localization branch, the start and end times of the target's appearance in the video are detected based on the text-guided visual features and the visual-guided text features. The text-guided visual features are average-pooled to obtain the text-guided global visual features. The text-guided global visual features are then convolved using two parallel temporal convolutional blocks to obtain the start and end visual features, respectively.
[0039] Each temporal convolution block consists of three 1-dimensional convolutional layers with a kernel size of 3. The sizes of the starting and ending visual features are both T×C.
[0040] In another embodiment, the global text input tag is a visually guided text feature generated by a cross-modal feature learning module with a special text tag, namely a global text feature. The global text feature is obtained by a multi-layer perceptron (MLP) to obtain an intermediate text feature in a C-dimensional common feature space. The activation module (Correlation & Sigmoid) calculates the correlation between each starting visual feature or ending visual feature and the intermediate text feature in the common feature space and uses the Sigmoid activation function to activate and obtain a starting prediction score and an ending prediction score. The prediction score represents the probability that each frame in a video slice is a starting frame or an ending frame.
[0041] In another embodiment, the spatiotemporal trajectory generation module combines the detection boxes and the start and end prediction scores in the time domain to construct an initial target spatiotemporal trajectory module. The start and end times of the spatiotemporal trajectory module are selected based on the maximum start and end prediction scores. All target detection boxes before the start time and after the end time are removed. The resulting time boundaries and target detection boxes together constitute the spatiotemporal trajectory prediction result, where the time boundaries are the start and end times.
[0042] In another embodiment, the spatiotemporal visual localization module is trained using three focal losses and one L1 loss as loss functions. The specific method is as follows:
[0043] Randomly sample a series of video slices, each of which contains T consecutive video frames; select video slices containing at least one frame of video with a corresponding target detection frame as training samples, and annotate the detection frame. For each video frame i in each video slice, define the center, width and length of the annotated target detection frame as Similarly, the start and end times of the annotation are marked as Based on the labeled target detection box, using Gaussian kernel A central heat map is generated in, Indicates The value of the space coordinate (x,y) in the middle, Represents the bandwidth parameter, which is determined by the size of the detection target. Similarly, the present invention generates two 1-dimensional time series heat maps for the start and end time of the detection target respectively.
[0044] During the training process, for each video slice, the loss function is defined as the formula:
[0045] Among them, A i ∈R 8H×8W is the predicted heat map, p s ,p e ∈R T It is the prediction score of the start and end time of target appearance. In the i-th frame video, the position The width and length of the predicted target detection box at the center. L c 、L s and L e They are respectively the focal loss for predicting the center of the target detection box, the start and end time of the temporal boundary. sizeIt is the L1 loss used to calculate the regression size of the target detection box. The loss weights are: λ1 = 0.1, λ2 = λ3 = λ4 = 1.
[0046] The technical effects achieved by the present invention are described in detail below:
[0047]
[0048] Table 1 Performance results of different methods on the VidSTG dataset
[0049] method m_vIoU(%) vIoU@0.3(%) vIoU@0.5(%) STVGT 18.15 26.81 9.48 STVGBert 20.42 29.37 11.31
[0050] Table 2 Performance results of different methods on the HC-STVG dataset
[0051] The results for these two datasets are shown in Tables 1 and 2. From these results, we make the following observations. 1) On both datasets, our proposed method significantly outperforms the state-of-the-art methods in all evaluation metrics. 2) On the VidSTG dataset, previous works such as GrounddeR+{·}, STPR+{·}, and WSSTG+{·} first use TALL or L-Net for temporal visual localization and then perform spatial visual localization to obtain the final result. In contrast, our model can simultaneously generate object detection boxes and temporal boundaries to form spatiotemporal tracklets, and its performance significantly outperforms two-stage localization methods, demonstrating the effectiveness of our proposed end-to-end method, STVGBert. 3) In Table 2, both STGVT and our STVGBert employ visual language transformers for cross-modal feature learning, but our proposed video spatiotemporal localization method has significant advantages over STGVT. Furthermore, STGVT requires a pre-trained object detector to generate a set of object detection boxes, while our proposed system, STVGBert, can directly process the input video.
[0052] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0053] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An end-to-end video spatiotemporal visual positioning system based on visual language Transformer, characterized by: The system comprises a visual information encoding module, a text embedding module, a spatiotemporal visual positioning module and a spatiotemporal trajectory generation module; the visual information encoding module and the text embedding module are connected to the spatiotemporal visual positioning module; the spatiotemporal visual positioning module is connected to the spatiotemporal trajectory generation module; the visual information encoding module is used to obtain the visual features of the detection target from the video frame; the text embedding module is used to extract the text encoding of the detection target from the query text; the spatiotemporal visual positioning module is used to learn the interaction features between the visual features and the text encoding, and perform spatial and temporal positioning of the detection target to obtain detection box information and time start and end information; the spatiotemporal trajectory generation module is used to combine the generated detection box information in the time domain and the space domain to obtain a spatiotemporal trajectory block containing the detection target; The spatiotemporal visual localization module includes a cross-modal feature learning module; the cross-modal feature learning module acquires text encoding and visual features to generate text-guided visual features and visual-guided text features; the cross-modal feature learning module includes a visual branch module; a spatiotemporal combination decomposition module is constructed in the visual branch module to retain spatial information; The spatiotemporal combination decomposition module includes a temporal pooling module, a spatial pooling module, a combination module, a multi-head attention module, a decomposition module, a replication module and a normalization module; The temporal pooling module is used to collect visual features to generate T×C preliminary temporal features, where T represents the number of video frames, C represents the number of feature map channels, H represents the height, and W represents the width; The spatial pooling module is used to collect visual features to generate preliminary spatial features with a shape of HW×C. The combination module is used to connect the preliminary temporal features and the preliminary spatial features in the feature dimension to form a combined visual feature with a size of (T+HW)×C; The multi-head attention module is used to perform attention operations based on the combined visual features and text features to generate preliminary text-guided visual features; The decomposition module is used to generate text-guided temporal features and text-guided spatial features based on preliminary text-guided visual features; The copy module is used to copy the text-guided temporal features HW times and the text-guided spatial features T times to obtain copied temporal features and copied spatial features of size T×HW×C; The normalization module is used to normalize the result of adding the copy temporal features, copy spatial features and visual input features to generate intermediate visual features; the output of the last layer is the text-guided visual features.
2. According to the end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 1, it is characterized in that The spatiotemporal visual positioning module also includes a spatiotemporal analysis core module; the spatiotemporal analysis core module generates text-guided visual features for temporal and spatial positioning.
3. According to the end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 1, it is characterized in that Special text tags ['GLS'] and ['SEP'] are added to the query text and placed at the beginning and end of the query text respectively. The text embedding module obtains the text encoding with special tags based on the query text with special tags; the cross-modal learning module obtains the text encoding with special text tags to obtain global text features.
4. The end-to-end video spatiotemporal positioning system based on the visual language Transformer according to claim 1, characterized in that: The cross-modal feature learning module also includes a text branch module, and the visual branch module interacts with the text branch module to learn and obtain text-guided visual features and visually guided text information.
5. The end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 2, characterized in that: The spatiotemporal analysis core module includes a spatial visual positioning branch module and a temporal visual positioning branch module; the spatial visual positioning module generates the center and size of the detection box based on the text-guided visual features; the temporal visual positioning branch module generates the start and end prediction scores based on the text-guided visual features.
6. The end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 5, characterized in that: The spatial visual positioning branch includes a deconvolution layer, a first detection network head, and a second detection network head; the deconvolution layer consists of three layers, which deconvolute the text-guided visual features and perform spatial upsampling to obtain upsampled features; the first detection network head and the second detection network head are both composed of a 3×3 convolutional layer and a 1×1 convolutional layer; the first detection network head selects to obtain the upsampled features to generate a heat map, and the position with the highest probability value in the heat map is the center of the detection box. The second detection network head obtains the size of the detection box by regressing the upsampled features.
7. The end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 6, characterized in that: The temporal visual localization branch module includes a spatial pooling layer, a temporal convolution layer, a multi-layer perceptron layer, and a computational activation layer; the spatial pooling layer uses spatial average pooling on the text-guided visual features to obtain the text-guided global visual features; The text-guided global visual features are convolved through two parallel temporal convolution blocks to obtain starting visual features and ending visual features respectively; the multi-layer perceptron layer obtains the global text features to obtain text features in a C-dimensional common feature space; the computation activation layer calculates the correlation between the starting visual features and the ending visual features and the text features in the common feature space and activates them using an activation function to obtain starting and ending prediction scores.
8. The end-to-end video spatiotemporal visual positioning system based on the visual language Transformer according to claim 1, characterized in that: The spatiotemporal visual localization module combines three focal losses and one L1 loss as the loss function for training.
Citation Information
Patent Citations
Video object positioning method based on weak supervised learning and video spatial and temporal characteristics
CN110765921A