A method for generating a video scene graph based on space-time features
Patent Information
- Application Number
- CN202410796896.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-19
AI Technical Summary
[0052](1)、基于局部可视矩阵和传统注意力机制提出了局部注意力机制,实现了局部空间信息的提取,突破了传统局部注意力机制无法很好地提取局部信息的缺点。
Smart Images

Figure CN118799779B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision semantic technology, and more specifically, relates to a method for generating video scene graphs based on spatiotemporal features. Background Technology
[0002] In video semantic understanding, traditional methods primarily rely on perception. Video classification is achieved through simple object detection and action recognition, often neglecting the rich background knowledge of videos in real-world scenarios. In reality, popular online short videos often face challenges such as changing scenes, numerous objects, and complex actions, making it difficult for traditional video understanding methods to accurately interpret video semantics. Scene graphs possess powerful semantic representation capabilities and play a crucial role in scene understanding. Scene Graph Generation (SGG) refers to automatically mapping images into semantically structured scene graphs, requiring accurate labeling of detected objects and their relationships. Introducing scene graphs allows video semantic understanding to better utilize scene knowledge, thereby increasing its accuracy.
[0003] Scene graph generation was first proposed by Johnson et al. A scene graph is defined as a graph structure representation of the relationships between object instances in a specific scene. Scene graph tasks were initially applied to image processing tasks. The output of an image scene graph is a relation triple, which researchers can use to generate the corresponding scene graph. Typically, a graph structure is used to represent scene graphs, where graph nodes represent object instances, and edge relationships represent the relationships between object instances.
[0004] In the field of video scene graph generation, there are two main methods: frame-based video scene graph generation and track-based video scene graph generation.
[0005] Frame-based video scene graph generation methods do not directly consider the relationships between frames during entity detection; instead, they directly perform object detection on a single frame. This approach often uses Faster R-CNN (Faster Region-based Convolutional Neural Network) for entity detection and then predicts relationships based on the detection results to generate the scene graph. One Transformer-based scene graph generation method first uses Faster R-CNN for object detection, then employs a spatial encoder to acquire spatial information, and simultaneously uses embedding to obtain semantic information. After acquiring spatial features, a temporal decoder is used to obtain temporal relationships for dynamic relationship prediction. The advantage of this method is that it utilizes a simple Transformer structure for video scene graph generation and achieves good results. Many subsequent works have improved upon the STRITRAN model structure to enhance the accuracy of video scene graph generation.
[0006] Trajectory-based video scene graph methods primarily focus on multi-object target detection in videos. Unlike frame-based video scene graph generation methods, trajectory-based methods place greater emphasis on multi-object detection. In this approach, the model incorporates temporal information during object detection, enabling better tracking and prediction of relationships between objects in the video. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a video scene graph generation method based on spatiotemporal features, which extracts structured information from video content in order to better understand and analyze the video content.
[0008] To achieve the above-mentioned objectives, the present invention provides a video scene graph generation method based on spatiotemporal features, characterized by comprising the following steps:
[0009] (1) Data acquisition and preprocessing;
[0010] (1.1) Download the video dataset. The video frames contain multiple targets. Annotate the ground truth bounding box of each target and the relationships between the targets.
[0011] (1.2) Remove video frames that do not contain the specified relationship, then scale the remaining video frames to a uniform size, and finally combine the processed video frames into a video dataset.
[0012] (2) Extract the predicted bounding box and confidence score;
[0013] The Faster R-CNN object detector detects the predicted bounding boxes of each object in each frame of the video, and assigns the RoI feature, confidence score, and object category to each predicted bounding box. Let P be the predicted bounding box of the j-th object in the i-th frame of the video. ij The corresponding RoI features are denoted as F. ij The confidence score is denoted as S. ij The target category is denoted as C. ij ;
[0014] (3) Filter the predicted bounding boxes;
[0015] (3.1) The NMS algorithm filters targets based on their confidence scores in each frame of video.
[0016] Filter out the predicted bounding boxes with confidence scores less than the threshold, and then sort the remaining predicted bounding boxes in descending order of confidence scores.
[0017] (3.2) Filter the sorted predicted bounding boxes using the intersection-union ratio;
[0018] Calculate the intersection-union ratio (IUU) between the sorted predicted bounding boxes and their corresponding ground truth bounding boxes:
[0019]
[0020] Among them, IOU ij P represents the predicted bounding box P of the j-th target in the i-th frame of the video. ij With the corresponding true bounding box The intersection and union ratio;
[0021] Filter out predicted bounding boxes whose intersection-union ratio with the ground truth bounding boxes is less than a threshold, and then store them in the predicted bounding box set. Let ρ be the number of predicted bounding boxes stored in the predicted bounding box set.
[0022] (4) Feature extraction;
[0023] (4.1) Extraction of pose key points and pose features;
[0024] The lightweight OpenPose model is used to extract the pose keypoints and pose features of the targets corresponding to each predicted bounding box in the predicted bounding box set. Here, the pose keypoint set of the j-th target is denoted as K. j The corresponding pose feature is denoted as λ. j j = 1, 2, ..., ρ;
[0025] (4.2) Joint visual feature extraction;
[0026] In the set of predicted bounding boxes, the RoI features corresponding to any two predicted bounding boxes are added together to obtain joint RoI features. Then, the joint visual relationship between the two predicted bounding boxes is extracted by the joint feature extraction module. Finally, the joint RoI features and the joint visual relationship are added together to obtain joint visual features.
[0027] (4.3) Semantic feature extraction;
[0028] By using a global vector model to embed words into the target prediction categories corresponding to each predicted bounding box in the predicted bounding box set, the semantic features of each target are obtained.
[0029] (4.4) Spatial feature extraction;
[0030] (4.4.1) Extraction of global spatial information;
[0031] In the set of predicted bounding boxes, the targets corresponding to any two predicted bounding boxes are denoted as the subject and the object. Then, the RoI features, joint visual features and semantic features corresponding to the subject and the object are input into the global spatial encoding module, and global spatial information is extracted through the global spatial encoding module.
[0032] (4.4.2) Extraction of local spatial information;
[0033] Construct the object category feature matrix M;
[0034] M = [m1 m2…m j …m ρ ]
[0035] Where, m j This represents the RoI feature corresponding to the j-th predicted bounding box;
[0036] Construct the visual matrix V;
[0037]
[0038] Among them, v ij =1 or v ij =0, when v ij =1 indicates that the i-th target is visible relative to the j-th target; otherwise, it is not visible.
[0039] The basis for visibility judgment is: taking the i-th target in the video frame as the reference, the distance between the i-th target and all other targets is calculated by using the pose key points of the target or the coordinates of the center point of the predicted bounding box, and the target with the smallest distance is selected as the visible object;
[0040] (4.4.3) Add the global spatial information and the local spatial information to obtain the spatial features corresponding to a single subject and object. Then repeat steps (4.4.1) to (4.4.2) to extract the spatial features corresponding to the other subjects and objects in this video frame.
[0041] (4.4.4) Repeat steps (3) to (4.4.3) to extract the spatial features corresponding to the subject and object of all video frames in the video dataset;
[0042] (4.5) Spatiotemporal feature extraction;
[0043] The local temporal features corresponding to the subject and object are obtained by averaging the spatial features corresponding to the same subject and object in the current video frame and the corresponding spatial features in the previous video frame.
[0044] Randomly select k video frames from the video dataset, and then take the average of the elements in the k video frames that have the same subject and object to obtain the global temporal features corresponding to the subject and object.
[0045] Adding the local time features corresponding to the subject and object to the global time features yields the spatiotemporal features corresponding to the subject and object.
[0046] (5) Relationship prediction
[0047] The spatiotemporal features corresponding to the subject and object are fused with the pose features of the corresponding subject to obtain the fused features corresponding to the subject and object. The fused features are then input into three relation classifiers FFN to obtain the prediction scores of spatial relations, exchange relations and attention relations. The highest prediction score is then taken as the relationship between the subject and object.
[0048] Finally, all subjects and objects generate relational triples based on the relationships between them. The triple format is <subject-relation-object>, thus obtaining the video scene diagram.
[0049] The objective of this invention is achieved as follows:
[0050] This invention discloses a video scene graph generation method based on spatiotemporal features. First, the video stream is downloaded and each frame is preprocessed to obtain a video dataset. Then, the predicted bounding boxes and confidence scores of each target in each frame are extracted, and the predicted bounding boxes are filtered according to their confidence scores and intersection-union ratio (IU). For the selected and retained predicted bounding boxes, feature extraction is performed sequentially to obtain the spatiotemporal features corresponding to each subject and object in each frame. Finally, based on the extracted features, a relation classifier is used to classify the relationships between the subjects and objects, and a scene graph in triplet <subject-relationship-object> format is generated, thus realizing a structured representation of the video content.
[0051] Meanwhile, the video scene graph generation method based on spatiotemporal features of the present invention also has the following beneficial effects:
[0052] (1) A local attention mechanism was proposed based on the local visibility matrix and the traditional attention mechanism, which realized the extraction of local spatial information and overcame the shortcomings of the traditional local attention mechanism in that it could not extract local information well.
[0053] (2) Feature fusion introduces pose features, which enriches the spatial information and semantics of features compared with the existing method STTran model, and improves the accuracy of model relationship prediction. Attached Figure Description
[0054] Figure 1 This is a flowchart of the video scene graph generation method based on spatiotemporal features of the present invention;
[0055] Figure 2 This is a schematic diagram of the Faster R-CNN architecture;
[0056] Figure 3 This is a schematic diagram of the lightweight OpenPose model;
[0057] Figure 4 This is a schematic diagram of the spatiotemporal feature and attitude feature fusion module;
[0058] Figure 5 It is a scene image from a video; Detailed Implementation
[0059] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.
[0060] Example
[0061] Figure 1 This is a flowchart of the video scene graph generation method based on spatiotemporal features of the present invention.
[0062] In this embodiment, as Figure 1 As shown, the present invention provides a video scene graph generation method based on spatiotemporal features, comprising the following steps:
[0063] S1. Data acquisition and preprocessing;
[0064] S1.1 Download the video dataset. The video frames contain multiple targets. Annotate the ground truth bounding boxes of each target and the relationships between the targets.
[0065] In this embodiment, the standard size of the video frame is 1067*600, the number of targets and target boxes is 100, and the relationships between targets include: person-object relationship, window-table relationship, etc.
[0066] S1.2 Remove video frames that do not contain the specified relationship (such as the person-object relationship), then scale the remaining video frames to a uniform size, and finally combine the processed video frames into a video dataset.
[0067] S2. Extract the predicted bounding box and confidence score;
[0068] like Figure 2 As shown, in the Faster R-CNN (Region-based Convolutional Neural Network) object detector, the predicted bounding boxes of each object in each frame of the video are detected, and the RoI (Region of Interests) features, confidence scores, and object categories are assigned to each predicted bounding box. Here, the predicted bounding box of the j-th object in the i-th frame of the video is denoted as P. ij The corresponding RoI features are denoted as F. ij The confidence score is denoted as S. ij The target category is denoted as C. ij ;
[0069] S3. Filter the predicted bounding boxes;
[0070] S3.1 The NMS (Non-Maximum Suppression) algorithm filters targets based on their confidence scores in each frame of video.
[0071] Filter out the predicted bounding boxes with a confidence score less than 0.3, and then sort the remaining predicted bounding boxes in descending order of their confidence scores.
[0072] S3.2 Filter the sorted predicted bounding boxes using the intersection-union ratio (IUU);
[0073] Calculate the intersection-union ratio (IUU) between the sorted predicted bounding boxes and their corresponding ground truth bounding boxes:
[0074]
[0075] Among them, IOU ij P represents the predicted bounding box P of the j-th target in the i-th frame of the video. ij With the corresponding true bounding box The intersection and union ratio;
[0076] Filter out predicted bounding boxes whose intersection-union ratio with the ground truth bounding boxes is less than 0.5, and then store them in the predicted bounding box set. Let ρ be the number of predicted bounding boxes stored in the predicted bounding box set.
[0077] In this embodiment, it is assumed that ρ = 3 predicted bounding boxes are stored in the i-th frame of the video;
[0078] S4. Feature extraction;
[0079] S4.1, Extraction of pose key points and pose features;
[0080] like Figure 3 As shown, in the lightweight OpenPose model, the pose keypoints and pose features of the target corresponding to each predicted bounding box in the predicted bounding box set are extracted, where the pose keypoint set of the j-th target is denoted as K. j The corresponding pose feature is denoted as λ. j j = 1, 2, ..., ρ;
[0081] S4.2 Joint visual feature extraction;
[0082] In the set of predicted bounding boxes, the RoI features corresponding to any two predicted bounding boxes are added together to obtain joint RoI features. Then, the joint visual relationship between the two predicted bounding boxes is extracted by the joint feature extraction module. Finally, the joint RoI features and the joint visual relationship are added together to obtain joint visual features.
[0083] In this embodiment, let P be the joint visual features extracted by combining the predicted bounding boxes stored in the i-th frame of the video in pairs. 12 P 13 P 23
[0084] S4.3 Semantic Feature Extraction;
[0085] By using a global vector model to embed words into the target prediction categories corresponding to each predicted bounding box in the predicted bounding box set, the semantic features of each target are obtained.
[0086] S4.4), Spatial Feature Extraction;
[0087] S4.4.1, Global Spatial Information Extraction;
[0088] In the set of predicted bounding boxes, the targets corresponding to any two predicted bounding boxes are denoted as the subject and the object. In this embodiment, taking the first and second predicted bounding boxes as examples, the target corresponding to the first predicted bounding box is denoted as the subject, and the target corresponding to the second predicted bounding box is denoted as the object.
[0089] Then, the RoI features, joint visual features, and semantic features corresponding to the subject and object are input into the global spatial encoding module, and global spatial information is extracted through the global spatial encoding module;
[0090] S4.4.2, Local Spatial Information Extraction;
[0091] Construct the object category feature matrix M;
[0092] M = [m1 m2…m j …m ρ ]
[0093] Where, m j This represents the RoI feature corresponding to the j-th predicted bounding box;
[0094] Construct the visual matrix V;
[0095]
[0096] Among them, v ij =1 or v ij =0, when v ij =1 indicates that the i-th target is visible relative to the j-th target; otherwise, it is not visible.
[0097] The basis for visibility judgment is: taking the i-th target in the video frame as the reference, the distance between the i-th target and all other targets is calculated by using the pose key points of the target or the coordinates of the center point of the predicted bounding box, and the target with the smallest distance is selected as the visible object;
[0098] S4.4.3. Add the global spatial information and the local spatial information to obtain the spatial features corresponding to a single subject and object. Then repeat steps S4.4.1 to S4.4.2 to extract the spatial features corresponding to the other subjects and objects in this video frame.
[0099] In this embodiment, three sets of global spatial information can be extracted from the i-th frame of video. These are then added to the local spatial information to obtain three sets of spatial features, denoted as […].
[0100] S4.4.4) Repeat steps S3 to S4.4.3 to extract the spatial features corresponding to the subject and object of all video frames in the video dataset;
[0101] S4.5 Spatiotemporal Feature Extraction;
[0102] The local temporal features corresponding to the subject and object are obtained by averaging the spatial features corresponding to the same subject and object in the current video frame and the corresponding spatial features in the previous video frame.
[0103] Randomly select k frames from the video dataset, then average the elements in the k frames that correspond to the same subject and object to obtain the global temporal features corresponding to the subject and object. For example, if 10 frames are retained and 4 frames are randomly selected, the extracted spatial features for the 7th frame are... Spatial information from video frames 1, 2, 4, and 6 was randomly extracted, and then information containing... Relationships with the same subject-object relationship, assuming the first frame of the video contains The second frame of the video contains The 4th frame of the video contains The 6th frame of the video contains The global temporal feature of the 7th frame of the video is then represented by the average of the corresponding spatial features. That is, the value of the 7th frame... Features and the first frame Averaged, 7th frame With frames 2 and 6 Take the average, the 7th frame With frames 2 and 4 Take the average.
[0104] Adding the local time features corresponding to the subject and object to the global time features yields the spatiotemporal features corresponding to the subject and object.
[0105] S5, Relationship Prediction;
[0106] like Figure 4 As shown, in the spatiotemporal feature and pose feature fusion module, the spatiotemporal features corresponding to the subject and object are fused with the pose features of the corresponding subject. SA represents self-attention mechanism and CA represents cross-attention mechanism. The fused features corresponding to the subject and object are obtained by feature fusion through self-attention and cross-attention. The fused features are then input into three relation classifiers FFN to obtain the prediction scores of spatial relation, exchange relation and attention relation. The highest prediction score is then taken as the relation between the subject and object.
[0107] Finally, all subjects and objects are used to generate relational triples based on the relationships between them. The triple format is <subject-relation-object>, thus obtaining the video scene graph, such as... Figure 5 As shown.
[0108] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.
Claims
1.A method for generating a video scene graph based on spatio-temporal features, characterized in that, The method comprises the following steps: (1) data acquisition and preprocessing; (1.1) download the video dataset, the video frame contains multiple targets, label the real boundary box of each target and the relationship between the targets; (1.2) remove the video frames that do not contain the specified relationship, then scale the remaining video frames to a uniform size, and finally form a video dataset by combining the processed video frames; (2) extract the predicted boundary box and the confidence score; The predicted bounding box of each target in each frame of video is detected by a Faster R-CNN target detector, and RoI features, confidence scores and target categories corresponding to each predicted bounding box are given, wherein the predicted bounding box of the i-th target in the j-th frame of video is denoted as , the corresponding RoI features are denoted as , the confidence score is denoted as , and the target category is denoted as . (3) screen the predicted boundary box; (3.1) the NMS algorithm screens according to the confidence score of each target in each video frame; filter out the predicted boundary box with a confidence score less than the threshold value, and then sort the remaining predicted boundary boxes in descending order according to the confidence score; (3.2) filter the sorted predicted boundary box using the intersection-over-union ratio; calculate the intersection-over-union ratio of the predicted boundary box and the corresponding real boundary box in the sorted predicted boundary box; ; wherein, represents the predicted bounding box of the target in the frame video intersection over union with the corresponding ground truth bounding box The predicted bounding box that has an intersection-over-union with the real bounding box less than a threshold is filtered and stored in a predicted bounding box set, and the number of predicted bounding boxes stored in the predicted bounding box set is denoted as ; (4) feature extraction; (4.1) pose key point and pose feature extraction; The pose key points and the pose features of the target corresponding to each of the prediction bounding boxes in the prediction bounding box set are extracted by using a lightweight OpenPose model, wherein a pose key point set of a first target is denoted as , a pose feature corresponding to the first target is denoted as , ; (4.2) joint visual feature extraction; add the RoI features corresponding to any two predicted boundary boxes in the predicted boundary box set to obtain joint RoI features, extract the joint visual relationship of the two predicted boundary boxes through the joint feature extraction module, and finally add the joint RoI features and the joint visual relationship to obtain the joint visual features; (4.3) semantic feature extraction; use the global vector model to perform word embedding on the target predicted class corresponding to each predicted boundary box in the predicted boundary box set to obtain the semantic features of each target; (4.4) spatial feature extraction; (4.4.1) global spatial information extraction; in the predicted boundary box set, the targets corresponding to any two predicted boundary boxes are denoted as subject and object, then the RoI features, joint visual features and semantic features corresponding to the subject and object are input into the global spatial coding module, and the global spatial information is extracted through the global spatial coding module; (4.4.2) local spatial information extraction; Constructing an object class feature matrix ; ; wherein, represents the RoI feature corresponding to the first prediction bounding box pair. Constructing a visual matrix ; ; wherein = 1 or = 0, when = 1 means that the j-th target has visibility with respect to the i-th target, otherwise means that it has no visibility; = 1 or = 0, when = 1 means that the j-th target has visibility with respect to the i-th target, otherwise means that it has no visibility; = The basis of the visibility judgment is that the distance between the first target and all the other targets is calculated based on the posture key point or the predicted bounding box center point coordinate of the target in the video frame, and the target with the smallest distance is selected as the visible object. The basis of the visibility judgment is that the distance between the first target and all the other targets is calculated based on the posture key point or the predicted bounding box center point coordinate of the target in the video frame, and the target with the smallest distance is selected as the visible object. (4.4.3) add the global spatial information and the local spatial information to obtain the spatial features corresponding to a single subject and object, then repeat steps (4.4.1)-(4.4.2) to extract the spatial features corresponding to the remaining subjects and objects in the current video frame; (4.4.4) repeat steps (3)-(4.4.3) to extract the spatial features corresponding to the subjects and objects in all video frames in the video dataset; (4.5) spatiotemporal feature extraction; take the average of the corresponding elements of the spatial features of the subject and object in the current video frame and the spatial features of the same subject and object in the previous video frame to obtain the local temporal features of the subject and object; Randomly extract in the video dataset Frame video, then take The elements with the same principal and guest object correspondence in the frame video are averaged to obtain the global time characteristics of the principal and guest object correspondence. add the local temporal features of the subject and object to the global temporal features to obtain the spatiotemporal features of the subject and object; (5) relationship prediction; fuse the spatiotemporal features of the subject and object with the pose features of the corresponding subject to obtain the fusion features of the subject and object, then input the fusion features into three relationship classifiers FFN to obtain the prediction scores of the spatial relationship, the exchange relationship and the attention relationship, and then take the highest prediction score as the relationship between the subject and the object. Finally, all the subjects and objects generate relationship triples according to the relationship between the subjects and objects, the triple format is <subject-relationship-object>, and thus the video scene graph is obtained.
Citation Information
Patent Citations
Method and device for detecting human-object interaction relationship in video
CN112464875A
Human body space-time motion detection method, system and equipment based on deep learning
CN116385926A