Video event retrieval method based on space-time interleaving
By constructing a spatiotemporal interleaving model of visual and positional features, the temporal and spatial relationships of objects in video frames are extracted, solving the problem of low accuracy in video event retrieval in existing technologies, and realizing efficient and compact hash encoding generation and video event retrieval.
Patent Information
- Application Number
- CN202510877181.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-11
AI Technical Summary
Existing video retrieval methods fail to effectively capture the spatiotemporal relationships of objects in videos, resulting in low event retrieval accuracy and difficulty in distinguishing the differences between different activities.
By constructing a spatiotemporal interleaving model of visual and positional features, the temporal and spatial relationship graphs of objects in video frames are extracted, and video event retrieval is performed by combining hash encoding. The training and retrieval are carried out using a multi-layer spatiotemporal interleaving module and a hash retrieval model.
It improves the accuracy and efficiency of video event retrieval, and generates efficient and compact hash codes that can effectively distinguish different video events.
Smart Images

Figure CN120929638A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a video retrieval method, and more particularly to a video event retrieval method based on spatiotemporal interleaving. Background Technology
[0002] With the explosive growth of video data in various complex scenarios, how to quickly retrieve events has become an urgent problem to solve. However, there is currently a lack of research on event retrieval in videos. Some hash learning methods for videos only encode the video from a holistic perspective, without encoding the activities themselves. Since activities focus on the interactions between multiple objects in a video, a better approach is to classify different activities by capturing the spatiotemporal correlation between location and visual features. Typically, combining location and visual information can approximate an activity. Taking running and jogging as examples, from a visual feature perspective, the objects in these two activities are similar; however, from the perspective of changes in object position, jogging is slower than running. In this case, it is necessary to obtain semantic information from both visual and location features and fuse them to complete the modeling of the activity. However, most current methods directly add location and visual information of different objects, which may lead to information loss and reduce the distinguishability between the two activities. Furthermore, due to the complexity of interactions between multiple objects in an activity, different activities exhibit varying degrees of similarity, such as waiting and talking, or talking and running; however, classification information alone cannot represent the differences between different categories. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a spatiotemporally interleaved video event retrieval method with efficient and compact hash coding and high accuracy in video event retrieval.
[0004] The technical solution adopted by this invention to solve the above-mentioned technical problem is: a video event retrieval method based on spatiotemporal interleaving, comprising the following steps:
[0005] Step 1): Divide the video dataset into multiple video segments, and then select N video segments from all video segments as training samples to form a training set;
[0006] Step 2): Construct a feature extraction model to be trained, including a visual feature extraction module and a position feature extraction module. Randomly shuffle N training samples and input them into the feature extraction model to be trained. Extract the visual features of a single object in the input training samples through the visual feature extraction module. Then extract the position features and temporal information of different objects in each video frame through the position feature extraction module and construct the temporal relationship graph of a single object and the spatial relationship graph between objects in the video frames respectively.
[0007] Step 3): Construct the spatiotemporal interleaving model to be trained, which includes two spatiotemporal interleaving modules. The first spatiotemporal interleaving module interleaves and fuses the temporal relationship graph of a single object with the visual features of a single object to obtain the interleaved fused temporal features of a single object. The second spatiotemporal interleaving module interleaves and fuses the spatial relationship graph between objects in each video frame with the visual features of all single objects in that video frame to obtain the interleaved fused spatial features of that video frame.
[0008] Step 4): Construct the hash retrieval model to be trained, including two pooling layers, two classification layers and one hash layer. Input the interleaved fusion spatial features of all video frames in the input training sample into the first pooling layer and the first classification layer to obtain the action classification prediction of a single object. Input the interleaved fusion spatial features of the video frames into the second pooling layer and the hash layer for dimensionality reduction and binarization to obtain the hash code of the training sample. Finally, input the hash code of the training sample into the second classification layer to obtain the video event classification prediction result of the training sample.
[0009] Step 5): Define the common total loss function for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained according to the similarity preservation principle. Obtain the value of the total loss function by using the spatial relationship graph between objects in each video frame of the training samples, the hash encoding of the training samples, the video event classification prediction results of the training samples, and the action classification prediction of a single object. Update the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained respectively through the backpropagation algorithm. After training, the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model are obtained.
[0010] Step 6): Use all samples in the video dataset except for the training samples as query samples to form a query set, and use the training samples as retrieval samples to form a retrieval set. Use the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model to obtain the hash codes of the query samples in the query set. Use the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model to obtain the hash codes of the retrieval samples in the retrieval set hash code retrieval library composed of the hash codes of all retrieval samples to find the hash code with the Hamming distance closest to the hash code of the query sample, and display the retrieval sample corresponding to the hash code as the retrieval result, thus completing the retrieval process of the query samples.
[0011] Compared with existing technologies, the advantages of this invention are that it first models and vectorizes visual features to obtain the visual features of a single object, and then constructs a temporal relationship graph of a single object based on the positional bounding box information and temporal information in different video frames of a single training sample, and constructs a spatial relationship graph of objects in video frames based on the positional relationships between multiple objects in each video frame. This approach can automatically focus on more prominent individuals in group events under the constraint of the loss function. Subsequently, the visual features, temporal relationship graph, and spatial relationship graph of a single object are input into a spatiotemporal interleaving model to obtain the interleaved fused spatial features of video frames. The spatiotemporal interleaving model emphasizes two levels of spatiotemporal interleaving. The first-level spatiotemporal interleaving module mainly focuses on the visual features of objects. The first layer focuses on the changes in features and their movement in the video; the second layer, the spatiotemporal interleaving module, mainly focuses on the changes in the relative positions of multiple objects and the changes in group visual features; then, a pooling layer, a hash layer, and a classification layer are added after the spatiotemporal interleaving model. The total loss function is obtained from the spatial relationship graph between objects in each video frame of the training samples, the hash encoding of the training samples, the video event classification prediction results of the training samples, and the action classification prediction of a single object. The network is trained under the constraint of the total loss function. It can learn the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model from the input training samples at the same time. This makes the hash encoding output by the model efficient and compact, thus effectively performing video event retrieval.
[0012] Furthermore, the specific process of step 2) is as follows:
[0013] Step 2-1: Construct a visual feature extraction module, including a VGG-16 backbone network, an ROIAlign layer, and a 3D deformable convolutional layer. The VGG-16 backbone network performs multi-scale feature extraction on each video frame of the input training sample to obtain features at multiple different scales. Then, bilinear interpolation is used to fuse the features at multiple different scales to obtain the fused features of the video frames in the training sample. The ROI Align layer extracts visual features from the fused features of the video frames based on the bounding box information of the objects in the fused features of the video frames. The extracted visual features are then temporally modeled and vectorized by the 3D deformable convolutional layer to obtain the visual features of a single object.
[0014] Step 2-2: Construct a location feature extraction module. The location feature extraction module obtains the bounding box information of each object in the input training sample in different video frames. Based on the bounding box information of the same object in the input training sample in different video frames and the temporal information, construct a temporal relationship graph of a single object. Obtain the intersection-union ratio (IUR) of the bounding box information of the same object in the input training sample in different video frames and use the IUR as the corresponding quantization representation in the temporal relationship graph of a single object.
[0015] Steps 2-3: Construct a spatial relationship graph between objects in each video frame of the input training sample, based on the positional relationships between multiple objects in each video frame. Obtain the absolute Euclidean distance between multiple objects in each video frame of the input training sample as the quantized representation of the spatial relationship graph between objects in that video frame.
[0016] Visual features of individual objects are extracted by using a backbone network, ROI Align layers, and 3D deformation convolutional layers. Simultaneously, a temporal relationship graph of individual objects is constructed based on the bounding box information and temporal information of the input training samples in the video frame. A spatial relationship graph between objects in the video frame is constructed based on the positional relationships between multiple objects in each video frame of the input training samples. These methods can effectively extract the visual features of objects in the video, the temporal relationships of individual objects, and the spatial relationships between objects in the video frame.
[0017] Furthermore, the specific process of step 3) is as follows:
[0018] Step 3-1: Construct the first-layer spatiotemporal interleaving module, which consists of three identical first multi-step fusion modules. Each first multi-step fusion module comprises a first normalization layer, a first sparse graph association multi-head attention layer, and a first feedforward neural network. First, the visual features of a single object are normalized by the first normalization layer in the first-layer first multi-step fusion module to obtain the normalized visual features of the single object. Then, the normalized visual features of the single object and the temporal relationship graph of the single object are input into the first sparse graph association multi-head attention layer. The first sparse graph association multi-head attention layer consists of two multi-head attention layers, where the first multi-head attention layer is for a single object. The temporal relationship graph of the objects is used to calculate multi-head graph attention. The second multi-head attention layer calculates multi-head attention for the normalized visual features of a single object. Finally, the multi-head graph attention and the multi-head attention are multiplied by a dot to obtain a fusion attention matrix. Then, the fusion attention matrix is multiplied by the normalized visual features of a single object to obtain the first fusion feature. Finally, the first fusion feature is input into the first feedforward neural network for mapping to obtain the first mapping feature. The first mapping feature is processed by the second layer first multi-step fusion module to obtain the second mapping feature. The second mapping feature is processed by the third layer first multi-step fusion module to output the interleaved fusion temporal features of a single object.
[0019] Step 3-2: Construct the second-layer spatiotemporal interleaving module, which consists of three identical second-multi-step fusion modules. Each second-multi-step fusion module consists of a second normalization layer, a second sparse graph association multi-head attention layer, and a second feedforward neural network. First, the interleaving fusion temporal features of a single object are normalized by the second normalization layer in the first-layer second-multi-step fusion module to obtain normalized interleaving fusion temporal features of a single object. Then, the normalized interleaving fusion temporal features of a single object and the spatial relationship graph between objects in the video frame are input into the second sparse graph association multi-head attention layer for fusion to obtain the second fusion feature. Finally, the second fusion feature is input into the second feedforward neural network for mapping to obtain the third mapping feature. The third mapping feature is processed by the second-layer second-multi-step fusion module to obtain the fourth mapping feature. The fourth mapping feature is processed by the third-layer second-multi-step fusion module to output the interleaving fusion spatial features of the video frame.
[0020] The first-layer spatiotemporal interleaving module extracts the interleaved fusion temporal features of individual objects, which are mainly obtained by fusing the visual features of individual objects with their temporal relationship graphs. This not only reflects the visual features of individual objects but also the position of the same object in different video frames of the input training sample. The second-layer spatiotemporal interleaving module extracts the interleaved fusion spatial features of video frames, which are mainly obtained by fusing the interleaved fusion temporal features of individual objects with their spatial relationship graphs between objects in the video frames. This effectively adds the spatial relationships between objects in the video frames to the interleaved fusion temporal features of individual objects to obtain the final interleaved fusion spatial features of the video frames.
[0021] Furthermore, the specific process of step 5) is as follows:
[0022] Step 5-1: Define the common total loss function L for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained. total L total =L cls +λ1L q +λ2L h Where λ1 and λ2 are hyperparameters, with λ1 ranging from 0.001 to 0.01 and λ2 ranging from 0.1 to 0.5. cls For classification loss, L cls =L acty +0.5L action L acty For the classification loss of video events, For the i-th training sample x i event tags, For x i Event classification prediction results in Laction For the action classification loss of a single object, M represents the total number of objects in a single video frame. For x i The action category label of the j-th object in the dataset. For x i Predict the action classification of the j-th object;
[0023] L q To quantify the loss, Among them, h v For the v-th training sample x v The interleaved fusion spatial features of all video frames, For the u-th training sample x u The transpose of the interleaved fusion spatial features of all video frames, b v For x v hash encoding, For x u The transpose of the hash code;
[0024] L h For hash loss, Where a u To obtain the output of the action classification prediction for a single object and the spatial relationship graph between objects in each video frame, the input is a graph convolution module. u ,b v ) represents a u and b v The degree of similarity,
[0025] Step 5-2: Set the maximum number of iterations. Based on the total loss function, use the Adam optimization algorithm to iteratively optimize the feature extraction model and the hash retrieval model to be trained until the set maximum number of iterations is reached. Then stop the iteration process to obtain the trained feature extraction model and the trained hash retrieval model.
[0026] By constraining the classification loss, feature extraction models, spatiotemporal interleaving models, and hash retrieval models can generate more discriminative features and hash codes for samples of different categories; by constraining the quantization loss, information loss can be reduced when mapping features to the binary encoding space; and by constraining the hash loss, hash codes generated from similar sample pairs can be similar, while hash codes generated from dissimilar sample pairs can be far apart. Attached Figure Description
[0027] Figure 1 This invention relates to the fusion attention matrix of the first sparse graph associated multi-head attention layer in the first multi-step fusion module of the first layer within the VD dataset.
[0028] Figure 2 This invention relates to the fusion attention matrix of the first sparse graph association multi-head attention layer in the second layer first multi-step fusion module within the VD dataset.
[0029] Figure 3 This invention relates to the fusion attention matrix of the first sparse graph association multi-head attention layer in the third layer first multi-step fusion module within the VD dataset.
[0030] Figure 4 This invention relates to the fusion attention matrix of the second sparse graph associated multi-head attention layer in the first layer second multi-step fusion module within the VD dataset.
[0031] Figure 5 This invention relates to the fusion attention matrix of the second sparse graph association multi-head attention layer in the second multi-step fusion module of the second layer within the VD dataset.
[0032] Figure 6 This invention relates to the fusion attention matrix of the second sparse graph association multi-head attention layer in the third layer second multi-step fusion module within the VD dataset. Detailed Implementation
[0033] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0034] A video event retrieval method based on spatiotemporal interleaving includes the following steps:
[0035] Step 1): Divide the video dataset into multiple video segments, each with 10 frames. Then select N video segments from all video segments as training samples to form the training set.
[0036] Step 2): Construct the feature extraction model to be trained, including a visual feature extraction module and a positional feature extraction module. Randomly shuffle N training samples and input them into the feature extraction model. Use the visual feature extraction module to extract the visual features of individual objects in the input training samples. Then use the positional feature extraction module to extract the positional features and temporal information of different objects in each video frame from the input training samples, and construct the temporal relationship graph of individual objects and the spatial relationship graph between objects in the video frames respectively. The specific process is as follows:
[0037] Step 2-1: Construct a visual feature extraction module, including a VGG-16 backbone network, a ROIAlign layer, and a 3D deformable convolutional layer. The VGG-16 backbone network performs multi-scale feature extraction on each video frame of the input training samples to obtain features at multiple different scales. Then, bilinear interpolation is used to fuse the features at multiple different scales to obtain the fused features of the video frames in the training samples. The ROI Align layer extracts visual features from the fused features of the video frames based on the bounding box information of the objects in the fused features of the video frames. The extracted visual features are then temporally modeled and vectorized by the 3D deformable convolutional layer to obtain the visual features of a single object.
[0038] Step 2-2: Construct a location feature extraction module. This module obtains the bounding box information of each object in the input training sample across different video frames. Based on the bounding box information of the same object in the input training sample across different video frames and the temporal information, a temporal relationship graph of a single object is constructed. The intersection-over-union ratio (IoU) of the bounding box information of the same object in each input training sample across different video frames is obtained and used as the corresponding quantization representation in the temporal relationship graph of a single object. The topological structure of the directed graph constructed in this way can well represent the actions and behaviors of an object in the video.
[0039] Steps 2-3: Construct a spatial relationship graph between objects in each video frame of the input training sample, based on the positional relationships between multiple objects. Obtain the absolute Euclidean distance between multiple objects in each video frame of the input training sample as the quantized representation of the spatial relationship graph for that video frame. Different absolute Euclidean distances between multiple objects represent different semantics, and under the same conditions, objects with closer absolute Euclidean distances are more closely connected.
[0040] Step 3): Construct the spatiotemporal interleaving model to be trained, including two layers of spatiotemporal interleaving modules. The first layer of spatiotemporal interleaving modules interleaves the temporal relationship graph of a single object with the visual features of a single object to obtain the interleaved fused temporal features of a single object. The second layer of spatiotemporal interleaving modules interleaves the spatial relationship graph between objects in each video frame with the visual features of all single objects in that video frame to obtain the interleaved fused spatial features of that video frame. The specific process is as follows:
[0041] Step 3-1: Construct the first-layer spatiotemporal interleaving module, which consists of three identical first multi-step fusion modules. Each first multi-step fusion module comprises a first normalization layer, a first sparse graph association multi-head attention layer, and a first feedforward neural network. First, the visual features of a single object are normalized by the first normalization layer in the first-layer first multi-step fusion module to obtain the normalized visual features of the single object. Then, the normalized visual features of the single object and the temporal relationship graph of the single object are input into the first sparse graph association multi-head attention layer. The first sparse graph association multi-head attention layer consists of two multi-head attention layers, where the first multi-head attention layer focuses on the temporal relationship of a single object. The relationship graph calculates multi-head graph attention, and the second multi-head attention layer calculates multi-head attention for the normalized visual features of a single object. These are all existing implementation techniques. Finally, the multi-head graph attention and the multi-head attention are multiplied by a dot to obtain a fused attention matrix. Then, the fused attention matrix is multiplied by the normalized visual features of a single object to obtain the first fused feature. Finally, the first fused feature is input into the first feedforward neural network for mapping to obtain the first mapped feature. The first mapped feature is processed by the second layer first multi-step fusion module to obtain the second mapped feature. The second mapped feature is processed by the third layer first multi-step fusion module to output the interleaved fused temporal features of a single object.
[0042] Step 3-2: Construct the second-layer spatiotemporal interleaving module, which consists of three identical second-multi-step fusion modules. Each second-multi-step fusion module consists of a second normalization layer, a second sparse graph association multi-head attention layer, and a second feedforward neural network. First, the interleaving fusion temporal features of a single object are normalized by the second normalization layer in the first-layer second-multi-step fusion module to obtain normalized interleaving fusion temporal features of a single object. Then, the normalized interleaving fusion temporal features of a single object and the spatial relationship graph between objects in the video frame are input into the second sparse graph association multi-head attention layer for fusion to obtain the second fusion feature. Finally, the second fusion feature is input into the second feedforward neural network for mapping to obtain the third mapping feature. The third mapping feature is processed by the second-layer second-multi-step fusion module to obtain the fourth mapping feature. The fourth mapping feature is processed by the third-layer second-multi-step fusion module to output the interleaving fusion spatial features of the video frame.
[0043] Step 4): Construct the hash retrieval model to be trained, including two pooling layers, two classification layers and one hash layer. Input the interleaved fusion spatial features of all video frames in the input training sample into the first pooling layer and the first classification layer to obtain the action classification prediction of a single object. Input the interleaved fusion spatial features of the video frames into the second pooling layer and the hash layer for dimensionality reduction and binarization to obtain the hash code of the training sample. Finally, input the hash code of the training sample into the second classification layer to obtain the video event classification prediction result of the training sample.
[0044] Step 5): Based on the principle of preserving similarity, define a common total loss function for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained. Utilize the spatial relationship graph between objects in each video frame of the training samples, the hash encoding of the training samples, the video event classification prediction results of the training samples, and the action classification prediction of a single object to obtain the value of the total loss function. Update the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained respectively using the backpropagation algorithm. After training, obtain the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model. The specific process is as follows:
[0045] Step 5-1: Define the common total loss function L for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained. total L total =L cls +λ1L q +λ2L h Where λ1 and λ2 are hyperparameters, with λ1 ranging from 0.001 to 0.01 and λ2 ranging from 0.1 to 0.5. cls For classification loss, L cls =L acty +0.5L action L acty For the classification loss of video events, For the i-th training sample x i event tags, For x i Event classification prediction results in L action For the action classification loss of a single object, M represents the total number of objects in a single video frame. For x i The action category label of the j-th object in the dataset. For x i Predict the action classification of the j-th object;
[0046] L q To quantify the loss, Among them, h v For the v-th training sample x v The interleaved fusion spatial features of all video frames, For the u-th training sample x u The transpose of the interleaved fusion spatial features of all video frames, b v For x v hash encoding, For x uThe transpose of the hash code;
[0047] L h For hash loss, Where a u It is obtained by inputting the action classification prediction of a single object and the spatial relationship graph between objects in each video frame into a graph convolution module and then outputting sim(a) u ,b v ) represents a u and b v The degree of similarity,
[0048] Step 5-2: Set the maximum number of iterations. Based on the total loss function, use the Adam optimization algorithm to iteratively optimize the feature extraction model and the hash retrieval model to be trained until the set maximum number of iterations is reached. Then stop the iteration process to obtain the trained feature extraction model and the trained hash retrieval model.
[0049] Step 6): Use all samples in the video dataset except for the training samples as query samples to form a query set, and use the training samples as retrieval samples to form a retrieval set. Use the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model to obtain the hash codes of the query samples in the query set. Use the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model to obtain the hash codes of the retrieval samples in the retrieval set hash code retrieval library composed of the hash codes of all retrieval samples to find the hash code with the Hamming distance closest to the hash code of the query sample, and display the retrieval sample corresponding to the hash code as the retrieval result, thus completing the retrieval process of the query samples.
[0050] The video event retrieval method based on spatiotemporal interleaving in this embodiment is referred to as "this method". The classification and retrieval performance of the traditional method and this method are compared below through specific comparison examples in Table 1.
[0051] Table 1:
[0052]
[0053]
[0054] Table 1 shows a comparison of the results of video retrieval using this method and traditional methods on three datasets: Volleyballdataset (VD), Collective Activity Dataset (CAD), and Collective Activity Extended Dataset (CAED). As can be seen from Table 1, the classification effect of this method based on the hash encoding of generated video events is still better than the best experimental scheme of traditional methods, and it effectively increases the efficiency and accuracy of video segment retrieval.
[0055] Table 2 below shows the retrieval and video group event classification accuracy of this method on the VD dataset under different hash lengths.
[0056] Table 2:
[0057] mAP@k 16-bit hash length 32-bit hash length 64-bit hash length 128-bit hash length mAP retrieved when K=1 94.69 94.91 94.91 95.99 mAP retrieved when K=5 94.23 95.11 95.07 96.02 mAP retrieved when K=10 93.66 95.10 95.07 96.06 mAP retrieved when K=20 93.82 95.14 95.04 95.67 mAP retrieved when K=50 94.20 95.09 95.03 95.68 Video group event classification accuracy 94.84 94.99 95.21 95.59
[0058] Table 3 below shows the retrieval and video group event classification accuracy of this method on the CAD dataset under different hash lengths.
[0059] Table 3:
[0060] mAP@k 16-bit hash length 32-bit hash length 64-bit hash length 128-bit hash length mAP retrieved when K=1 96.89 96.97 96.97 98.16 mAP retrieved when K=5 97.17 97.07 97.26 98.24 mAP retrieved when K=10 97.73 97.04 97.28 98.10 mAP retrieved when K=20 97.69 97.22 96.79 98.10 mAP retrieved when K=50 96.80 97.21 96.33 97.91 Video group event classification accuracy 97.76 97.37 97.89 98.29
[0061] As can be seen from Tables 2 and 3, this method can achieve good experimental results for different hash lengths.
[0062] Figure 1 , Figure 2 and Figure 3 The inference process of this method on the VD dataset is shown. The fusion attention matrix of the first sparse graph association multi-head attention layer in the first multi-step fusion module of each layer is shown. The Time axis represents the temporal sequence of video segments. In the first spatiotemporal interleaving module, the shallow attention is more focused on visual changes, resulting in stronger attention to image frames. As the first multi-step fusion module increases, the positional information is gradually fused with the visual features, effectively establishing the spatiotemporal association of objects.
[0063] Figure 4 , Figure 5 and Figure 6 The diagram illustrates the fusion attention matrix of the second sparse association multi-head attention layer in each layer of the second multi-step fusion module during the inference process of our method on the VD dataset, where the Person ID axis represents the object number in the video clip. In the second spatiotemporal interleaving module, the second multi-step fusion module mainly considers the correlation between several similar objects, assigning relatively small attention weights to other objects; as the number of second multi-step fusion modules increases, the changes in object position are gradually combined with visual features.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A video event retrieval method based on spatiotemporal interleaving, characterized in that... Includes the following steps: Step 1): Divide the video dataset into multiple video segments, and then select N video segments from all video segments as training samples to form a training set; Step 2): Construct a feature extraction model to be trained, including a visual feature extraction module and a position feature extraction module. Randomly shuffle N training samples and input them into the feature extraction model to be trained. Extract the visual features of a single object in the input training samples through the visual feature extraction module. Then extract the position features and temporal information of different objects in each video frame through the position feature extraction module and construct the temporal relationship graph of a single object and the spatial relationship graph between objects in the video frames respectively. Step 3): Construct a spatiotemporal interleaving model to be trained, including two spatiotemporal interleaving modules. The first spatiotemporal interleaving module interleaves and fuses the temporal relationship graph of a single object with the visual features of a single object to obtain the interleaved fused temporal features of a single object. The second-layer spatiotemporal interleaving module interleaves and fuses the spatial relationship diagram between objects in each video frame with the visual features of all individual objects in that video frame to obtain the interleaved and fused spatial features of that video frame. Step 4): Construct the hash retrieval model to be trained, including two pooling layers, two classification layers and one hash layer. Input the interleaved fusion spatial features of all video frames in the input training sample into the first pooling layer and the first classification layer to obtain the action classification prediction of a single object. Input the interleaved fusion spatial features of the video frames into the second pooling layer and the hash layer for dimensionality reduction and binarization to obtain the hash code of the training sample. Finally, input the hash code of the training sample into the second classification layer to obtain the video event classification prediction result of the training sample. Step 5): Define the common total loss function for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained according to the similarity preservation principle. Obtain the value of the total loss function by using the spatial relationship graph between objects in each video frame of the training samples, the hash encoding of the training samples, the video event classification prediction results of the training samples, and the action classification prediction of a single object. Update the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained respectively through the backpropagation algorithm. After training, the trained feature extraction model, the trained spatiotemporal interleaving model, and the trained hash retrieval model are obtained. Step 6): Use all samples in the video dataset except for the training samples as query samples to form a query set, use the training samples as retrieval samples to form a retrieval set, and use the query samples in the query set to obtain the hash code of the query samples through the trained feature extraction model, the trained spatiotemporal interleaving model and the trained hash retrieval model. The search samples in the search set are processed by a trained feature extraction model, a trained spatiotemporal interleaving model, and a trained hash retrieval model to obtain hash codes for the search samples. The hash code that has the closest Hamming distance to the hash code of the query sample is found in the hash code retrieval library composed of the hash codes of all search samples. The search sample corresponding to the hash code is then displayed as the search result, thus completing the search process for the query sample.
2. The video event retrieval method based on spatiotemporal interleaving according to claim 1, characterized in that... The specific process of step 2) is as follows: Step 2-1: Construct a visual feature extraction module, including a VGG-16 backbone network, an ROIAlign layer, and a 3D deformable convolutional layer. The VGG-16 backbone network performs multi-scale feature extraction on each video frame of the input training sample to obtain features at multiple different scales. Then, bilinear interpolation is used to fuse the features at multiple different scales to obtain the fused features of the video frames in the training sample. The ROI Align layer extracts visual features from the fused features of the video frames based on the bounding box information of the objects in the fused features of the video frames. The extracted visual features are then temporally modeled and vectorized by the 3D deformable convolutional layer to obtain the visual features of a single object. Step 2-2: Construct a location feature extraction module. The location feature extraction module obtains the bounding box information of each object in the input training sample in different video frames. Based on the bounding box information of the same object in the input training sample in different video frames and the temporal information, construct a temporal relationship graph of a single object. Obtain the intersection-union ratio (IUR) of the bounding box information of the same object in the input training sample in different video frames and use the IUR as the corresponding quantization representation in the temporal relationship graph of a single object. Steps 2-3: Construct a spatial relationship graph between objects in each video frame of the input training sample, based on the positional relationships between multiple objects in each video frame. Obtain the absolute Euclidean distance between multiple objects in each video frame of the input training sample as the quantized representation of the spatial relationship graph between objects in that video frame.
3. The video event retrieval method based on spatiotemporal interleaving according to claim 2, characterized in that... The specific process of step 3) is as follows: Step 3-1: Construct the first-layer spatiotemporal interleaving module, which consists of three identical first multi-step fusion modules. Each first multi-step fusion module comprises a first normalization layer, a first sparse graph association multi-head attention layer, and a first feedforward neural network. First, the visual features of a single object are normalized by the first normalization layer in the first-layer first multi-step fusion module to obtain the normalized visual features of the single object. Then, the normalized visual features of the single object and the temporal relationship graph of the single object are input into the first sparse graph association multi-head attention layer. The first sparse graph association multi-head attention layer consists of two multi-head attention layers, where the first multi-head attention layer is for a single object. The temporal relationship graph of the objects is used to calculate multi-head graph attention. The second multi-head attention layer calculates multi-head attention for the normalized visual features of a single object. Finally, the multi-head graph attention and the multi-head attention are multiplied by a dot to obtain a fusion attention matrix. Then, the fusion attention matrix is multiplied by the normalized visual features of a single object to obtain the first fusion feature. Finally, the first fusion feature is input into the first feedforward neural network for mapping to obtain the first mapping feature. The first mapping feature is processed by the second layer first multi-step fusion module to obtain the second mapping feature. The second mapping feature is processed by the third layer first multi-step fusion module to output the interleaved fusion temporal features of a single object. Step 3-2: Construct the second-layer spatiotemporal interleaving module, which consists of three identical second-multi-step fusion modules. Each second-multi-step fusion module consists of a second normalization layer, a second sparse graph association multi-head attention layer, and a second feedforward neural network. First, the interleaving fusion temporal features of a single object are normalized by the second normalization layer in the first-layer second-multi-step fusion module to obtain normalized interleaving fusion temporal features of a single object. Then, the normalized interleaving fusion temporal features of a single object and the spatial relationship graph between objects in the video frame are input into the second sparse graph association multi-head attention layer for fusion to obtain the second fusion feature. Finally, the second fusion feature is input into the second feedforward neural network for mapping to obtain the third mapping feature. The third mapping feature is processed by the second-layer second-multi-step fusion module to obtain the fourth mapping feature. The fourth mapping feature is processed by the third-layer second-multi-step fusion module to output the interleaving fusion spatial features of the video frame.
4. The video event retrieval method based on spatiotemporal interleaving according to claim 3, characterized in that... The specific process of step 5) is as follows: Step 5-1: Define the common total loss function L for the feature extraction model, the spatiotemporal interleaving model, and the hash retrieval model to be trained. total L total =L cls +λ1L q +λ2L h Where λ1 and λ2 are hyperparameters, with λ1 ranging from 0.001 to 0.01 and λ2 ranging from 0.1 to 0.
5. cls For classification loss, L cls =L acty +0.5L action L acty For the classification loss of video events, For the i-th training sample x i event tags, For x i Event classification prediction results in L action For the action classification loss of a single object, M represents the total number of objects in a single video frame. For x i The action category label of the j-th object in the dataset. For x i Predict the action classification of the j-th object; L q To quantify the loss, Among them, h v For the v-th training sample x v The interleaved fusion spatial features of all video frames, For the u-th training sample x u The transpose of the interleaved fusion spatial features of all video frames, b v For x v hash encoding, For x u The transpose of the hash code; L h For hash loss, Where a u It is obtained by inputting the action classification prediction of a single object and the spatial relationship graph between objects in each video frame into a graph convolution module and then outputting sim(a) u ,b v ) represents a u and b v The degree of similarity, Step 5-2: Set the maximum number of iterations. Based on the total loss function, use the Adam optimization algorithm to iteratively optimize the feature extraction model and the hash retrieval model to be trained until the set maximum number of iterations is reached. Then stop the iteration process to obtain the trained feature extraction model and the trained hash retrieval model.