Dynamic scene graph generation method based on vision transformer spatio-temporal relationship modeling

CN122821058APending Publication Date: 2026-09-25AEROSPACE INFORMATION TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004498.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

动作基因组(Action Genome,AG)中的谓词类分布,会导致模型在训练过程中偏向于常见类别,从而在处理稀有类别时表现不佳

Benefits of technology

本发明公开了一种基于Vision Transformer时空关系建模的动态场景图生成方法,采用双分支并行处理架构,同步实现实体关系精准分类与时序动态关系建模,可以避免单分支结构中分类任务与时序建模任务的特征冲突,兼顾静态帧内语义与动态帧间变化的完整表达;本发明中VOCM模块融合空间几何位置编码与时序正余弦编码,通过堆叠Transformer编码器捕捉帧内与帧间的复杂关联,结合时序 LSTM 强化长时序依赖建模,有效提升实体与关系的分类精度;本发明中TRM模块通过空间与时序双Transformer编码器,分别建模帧内空间拓扑关系与跨帧时序依赖,搭配时序差异聚合器构建差分特征,通过差分注意力强化相邻帧间的关系变化捕捉,解决了现有技术对动态场景中关系时序演变建模薄弱的痛点;还采用可学习门控融合机制自适应融合双分支特征,保证了时序序列的一致性,进而显著提升动态场景图生成的精度与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821058A_ABST
    Figure CN122821058A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of dynamic scene graph generation, and particularly relates to a dynamic scene graph generation method based on Vision Transformer space-time relationship modeling, which specifically comprises the following steps: first, inputting a video frame sequence into a target detector, extracting visual features, normalized bounding boxes and category distributions of entities, constructing subject-object triples and cropping a region of interest; then, through double-branch parallel processing, an VOCM module fuses a Transformer and a time sequence LSTM to complete entity and relationship classification, a TRM module captures space-time features through a time sequence Vision Transformer encoder and a time sequence difference aggregator TDA, generates dynamic time sequence features through the time sequence difference aggregator, and finally fuses double-branch outputs to generate a dynamic scene graph sequence. The application can solve the problems of insufficient capture of time sequence dynamic relationships and insufficient modeling of interframe space-time dependencies in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic scene graph generation technology, and in particular to a dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling. Background Technology

[0002] In understanding complex scenes, effectively modeling entities and their relationships has always been a core research challenge. Scene graphs, a structured representation method, model entities in a scene as nodes and their relationships as edges, providing an intuitive and efficient approach to analyzing semantics and interactions within a scene. This representation method has been widely applied in image retrieval, visual question answering, and autonomous driving, enhancing model interpretability and significantly improving task performance. Dynamic scene graphs extend the concept of static scene graphs by incorporating temporal information, enabling the modeling of scenes that evolve over time. This capability is crucial for applications that require understanding scene changes, such as video analytics, action recognition, and surveillance. By capturing the temporal dynamics of objects and their relationships, dynamic scene graphs provide a richer and more comprehensive understanding of the visual environment.

[0003] In recent years, significant progress has been made in the field of Dynamic Scene Graph Generation (DSGG), primarily due to the superior sequence processing capabilities of transformers, which can effectively aggregate spatial and temporal contextual information from videos. However, existing DSGG methods still face numerous challenges. First, many videos contain substantial contextual noise, and objects in certain frames may be difficult to identify accurately due to occlusion or blurring, leading to inaccurate object recognition. Furthermore, in some blurred video frames, the same relation may correspond to multiple possible object pairs, further increasing the complexity and uncertainty of relation recognition. Second, class imbalance is a particularly prominent issue in datasets, with some objects and relations occurring far more frequently than others. The predicate class distribution in the Action Genome (AG) can cause models to favor common classes during training, resulting in poor performance when dealing with rare classes. Finally, existing DSGG methods remain insufficient in capturing the temporal dependencies of visual relationships between entities, especially exhibiting weaknesses in recognizing fine-grained changes in relationships over time. For example, these methods often lack accurate modeling of changes in relationships between consecutive frames, making it difficult to effectively represent and capture the subtle yet crucial temporal variations in dynamic scenes. Therefore, effectively handling contextual noise, mitigating class imbalance, and enhancing the modeling of temporal dependencies remain important problems that urgently need to be addressed in the DSGG field.

[0004] In summary, this invention proposes a dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling to solve the above problems. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a dynamic scene graph generation method based on spatiotemporal relationship modeling using Vision Transformer. This invention leverages the ability of Vision Transformer to process continuous data and capture long-distance dependencies to understand dynamic scenes and thus generate dynamic scene graphs.

[0006] The technical solution of this invention to solve the technical problem is a dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling, comprising the following steps: S1. Input the video frame sequence into the target detector to detect entities in the video frames, and extract the visual features of the entities, normalized bounding boxes and object category distribution to form triples. Crop the input video frames, retain the regions related to the triples, and obtain the cropped region features and feature maps. S2. The cropped region features and feature maps are processed and then input into the Video Object Relationship Classification (VOCM) module. The VOCM module's processing includes position encoding, Transformer encoder processing, and temporal LSTM processing. The VOCM module classifies entities and relationships in the image to obtain the classified entity and relationship features. Meanwhile, the cropped region features and feature maps are processed and input into the temporal relationship perception module (TRM). The TRM module includes a spatial Vision Transformer encoder, a temporal Vision Transformer encoder, and a temporal difference aggregator (TDA). First, the two encoders capture the relationship features of the spatial and temporal modalities respectively. Then, the TDA constructs the difference features to capture the changes of the relationship features between time steps and generate dynamic temporal relationship features. S3. The outputs of the VOCM module and the TRM module are combined to generate the final dynamic scene graph sequence.

[0007] S1 is as follows: video frame sequence The input is fed into a Faster R-CNN object detector pre-trained on the Action Genome benchmark dataset to detect each video frame. The entity in Indicates the total number of video frames. express index, Indicates the first Each time step of the video frame; video frames The corresponding entity set is represented as , Indicates the first The total number of entities in a video frame at each time step, where each entity is an object. express index, Represents video frames The Middle Each entity is extracted. Visual features, normalized bounding boxes, and object category distribution; For entity collections Exhaustive subject-object pairing is performed on the entities to generate video frames. The set of all subject-object entity pairs Simultaneously, based on the visual features and normalized bounding boxes of each entity, each subject-object entity pair is represented. Derivation of subjective and objective visual features , Normalized bounding boxes for subject and object, same , express The index; for each pair of subject-object entities. Assign a unique relation proposal index , Indicates the first Subject-object entity pair, , This represents the total number of subject-object entity pairs, thus forming a triple. ; triplet The normalized bounding boxes of two entities are combined to form a joint bounding box. ,based on For video frames Perform RoI cropping to obtain the region of interest feature map. and corresponding union box visual features ; In this context, RoI cropping represents the extraction of a specific region of interest from an image or video; the object category distribution of an entity represents the probability of the category to which the entity belongs, which is used to convert it into a semantic embedding vector for each entity, serving as an initial constraint for entity classification.

[0008] The specific operation process of position encoding is as follows: The feature maps of each region of interest are converted into serialized feature tensors and used as input to the VOCM module. After input, spatial and temporal position encoding is performed. Spatial position encoding adopts a geometric encoding strategy based on object detection bounding boxes, while temporal position encoding adopts a fixed encoding form using sine and cosine functions. The generated spatial position encoding and temporal position encoding are concatenated to obtain a position encoding vector, which is then added to the input serialized feature tensor to obtain the enhanced feature tensor.

[0009] The ransformer encoder process is as follows: The enhanced feature tensor is input into the Transformer encoder, which captures the complex relationships between objects within and between frames through a multi-head self-attention mechanism. Each head converts the input enhanced feature tensor into a query, key, and value, and calculates the weight matrix of each head through the self-attention mechanism. The weight matrices of each head are concatenated and combined with a linear transformation matrix to generate a fused weight matrix. Then, the fused weight matrix is ​​passed through residual connections and a feedforward neural network to retain the original input information and perform nonlinear enhancement, resulting in the features output by the Transformer encoder. The VOCM module adopts a multi-layer Transformer encoder stack structure. The output of the previous Transformer encoder is used as the input of the next Transformer encoder, and the output of the last Transformer encoder is the final output feature of the Transformer encoder.

[0010] The specific process of temporal LSTM processing is as follows: The features output by the Transformer encoder are processed by temporal LSTM. The input features correspond to the video frames input by S1. The video frames are temporal sequence data, and the corresponding features are also temporal sequence data. Each feature processed by temporal LSTM includes two internal state vectors, namely the hidden state and the cell state. When the feature of the current time step is processed by temporal LSTM, the hidden state and cell state of the feature of the previous time step are combined to update the feature of the current time step, and the updated feature of the current time step is obtained. This process is repeated multiple times until the updated feature of the last time step is output. The hidden state of the last time step is selected as the final temporal feature vector of each object. The temporal feature vectors output by each object are then passed through a two-layer feedforward neural network to output the classification prediction results of each entity. At the same time, based on the temporal feature vectors corresponding to the subject-object entity pairs, the corresponding relationship classification features and classification prediction results are output through a two-layer feedforward neural network. The VOCM module finally outputs entity classification features and relationship classification features containing intra-frame and inter-frame spatiotemporal dependencies and temporal dynamic information, as well as the corresponding entity and relationship classification prediction results. Calculate the entity cross-entropy loss based on the entity classification prediction results, calculate the relationship cross-entropy loss based on the relationship classification prediction results, and train and optimize VOCM.

[0011] The region features and feature maps after being cropped and input into the TRM module are processed as follows: Based on the visual feature vector of each subject-object entity pair , Visual features of the joint frame Coordinates of the normalized bounding box of the subject and object , and semantic embedding vectors , Perform feature concatenation and dimensional transformation to generate relational features. This leads to the set of relational features. The semantic embedding vector is derived from the semantic embedding vector of each entity. The calculation formula is as follows: , in, , , represent the learnable linear matrices for dimensionality compression; Indicates the flattening operation; This represents an element-wise addition operation; Represents the bounding box feature uniform function; Indicates feature concatenation operation; Indicates the first In the first time step A relational feature vector, .

[0012] The operation process of the spatial Vision Transformer encoder is as follows: Each Set of relational features at each time step The input is fed into the TRM module, where it is processed by the spatial Vision Transformer encoder. First, it is reconstructed using a frame partitioning formula, discretizing the continuous video stream into spatiotemporally consistent segment representations, thus obtaining the relation index set for each frame. ; Then, zero-padded the relation feature matrix of each frame. Expanded to a maximum length per frame Generate batch tensors , This indicates a fill operation. Represents batch processing tensors Dimensions Represents the set of relation feature matrices; simultaneously generates a binary mask matrix. Used to distinguish between valid and invalid positions. , If the first The first video frame of the time step If each location contains a true visual relationship, then Otherwise, it is 1; Then process the tensors in the batch Normalization is performed to obtain the normalized features. Next, the spatial self-attention module captures the spatial topological relationships between objects through dynamic weight allocation, and calculates them through linear projection. The query, key, and value are combined with a binary mask matrix to calculate the output of the self-attention module. This output is then passed through a residual connection and a feedforward network to output intra-frame spatial relationship features.

[0013] The operation process of the timing Vision Transformer encoder is as follows: Based on the object semantic category embedding vector obtained from the object category distribution of entities, intra-frame spatial relationship features are grouped and reorganized, and features of the same category are aggregated along the temporal dimension to form... Each object represents a set of independent temporal feature sequences, and the relation index set for each object is represented as follows: , ; For each temporal feature sequence, a dual position information injection of absolute position index and relative position embedding is performed to obtain a temporal feature sequence after injection of position information. A unique temporal position identifier is generated for each time step through absolute position index, and the temporal distance between adjacent frame keys is mapped into a continuous vector representation through relative position embedding. Then, a dynamic filling strategy is used to dimensionally align the temporal feature sequences after the injection of location information, expanding them to the maximum sequence length allowed for each object category, thus obtaining the temporal feature matrix; Finally, the temporal feature matrix is ​​input into the Transformer unit with the same structure as the spatial Vision Transformer encoder, and the features are processed along the temporal dimension to capture the temporal dependencies across frames, thereby obtaining intra-frame temporal relationship features. These features are then fused with intra-frame spatial relationship features to obtain spatiotemporal fusion relationship features.

[0014] The operation process of the Time Difference Aggregator (TDA) is as follows: The spatiotemporal fusion relationship features are subjected to dimensionality compression and time step pairing to obtain cross-time step feature tensors. A shift sequence mechanism is then used to construct the first... The shift position characteristics of each time step force attention to the local temporal changes of adjacent time steps; The shift position features are concatenated with the cross-time step feature tensor of the current time step along the feature dimension, and then the concatenated part is transformed by a linear transformation. 3D feature projection back The dimension is used to obtain the differential features, which are then used to generate differential attention weights through two layers of nonlinear mapping; The cross-time step feature tensor is enhanced by differential attention weights to obtain the enhanced shift position features. The shift position features are then multiplied element-wise with the differential attention weights generated by the second-layer nonlinear mapping to obtain the enhanced temporal features. Finally, the enhanced temporal features are concatenated with the cross-time step feature tensor to obtain the fused feature vector. Based on the relation proposal index of the subject-object entity pair, the fused feature vector is injected into the feature position of the corresponding relation through the scatter projection operation. The scatter operation injects the fused feature vector into the corresponding position according to the relation proposal index. The TRM module finally generates dynamic temporal relation features enhanced by differential attention.

[0015] S3 is as follows: The entity and relation classification features output by the VOCM module and the dynamic temporal relation features output by the TRM module are dimensionally aligned. Then, a learnable gating fusion mechanism is used to adaptively fuse the two types of features after dimension alignment to obtain a fusion gating vector. Finally, the fusion feature is calculated based on the fusion gating vector. The calculation formula is as follows: , in, Indicates the final fusion characteristics, Represents the fusion gate vector. This indicates the output of the VOCM module. This indicates the output of the TRM module, and ⊙ represents element-wise multiplication. The final fused features are then input into the entity classification head and the relationship classification head, respectively. After passing through the feedforward neural network and the Softmax activation function, the final category prediction results of entities in each video frame and the final category prediction results of the relationship between the subject and object entities are output. Using entities in each video frame as nodes and the predicted relationships between entities as edges, a single-frame scene graph is constructed for each time step of the video frame. All single-frame scene graphs are then concatenated in temporal order according to the video frames to generate the final dynamic scene graph sequence.

[0016] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: This invention discloses a dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling. It employs a dual-branch parallel processing architecture to simultaneously achieve accurate entity relationship classification and temporal dynamic relationship modeling, avoiding feature conflicts between classification and temporal modeling tasks in single-branch structures, and ensuring a complete expression of both static intra-frame semantics and dynamic inter-frame changes. The VOCM module integrates spatial geometric position encoding and temporal sine and cosine encoding, capturing complex intra-frame and inter-frame relationships through stacked Transformer encoders, and combining temporal LSTM to enhance long-term temporal dependency modeling, effectively improving the classification accuracy of entities and relationships. The TRM module uses spatial and temporal dual Transformer encoders to model intra-frame spatial topological relationships and cross-frame temporal dependencies respectively, constructing differential features with a temporal difference aggregator, and strengthening the capture of relationship changes between adjacent frames through differential attention, addressing the weakness of existing technologies in modeling the temporal evolution of relationships in dynamic scenes. Furthermore, a learnable gating fusion mechanism is used to adaptively fuse dual-branch features, ensuring the consistency of the temporal sequence, thereby significantly improving the accuracy and robustness of dynamic scene graph generation.

[0017] In summary, this invention, relying on the strong correlation modeling capabilities of Vision Transformer, fully covers the entire process of dynamic scene graphs from entity detection and relation classification to temporal dynamic modeling. It can be widely applied to scenarios such as video understanding, action recognition, and intelligent monitoring, providing an efficient and reliable technical solution for high-level semantic parsing of dynamic scenes. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0019] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0020] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific implementation methods and in conjunction with the accompanying drawings.

[0021] Example 1 like Figure 1 As shown, a method for generating dynamic scene graphs based on spatiotemporal relationship modeling using Vision Transformer includes the following steps: S1. Input the video frame sequence into the target detector to detect entities in the video frames, and extract the visual features of the entities, normalized bounding boxes and object category distribution to form triples. Crop the input video frames, retain the regions related to the triples, and obtain the cropped region features and feature maps. S2. The cropped region features and feature maps are processed and then input into the Video Object Relationship Classification (VOCM) module. The VOCM module's processing includes position encoding, Transformer encoder processing, and temporal LSTM processing. The VOCM module classifies entities and relationships in the image to obtain the classified entity and relationship features. Meanwhile, the cropped region features and feature maps are processed and input into the temporal relationship perception module (TRM). The TRM module includes a spatial Vision Transformer encoder, a temporal Vision Transformer encoder, and a temporal difference aggregator (TDA). First, the two encoders capture the relationship features of the spatial and temporal modalities respectively. Then, the TDA constructs the difference features to capture the changes of the relationship features between time steps and generate dynamic temporal relationship features. S3. The outputs of the VOCM module and the TRM module are combined to generate the final dynamic scene graph sequence.

[0022] In a specific implementation, S1 is as follows: video frame sequence The input is fed into a Faster R-CNN object detector pre-trained on the Action Genome benchmark dataset to detect each video frame. The entity in Indicates the total number of video frames. express index, Indicates the first Each time step of the video frame; video frames The corresponding entity set is represented as , Indicates the first The total number of entities in a video frame at each time step, where each entity is an object. express index, Represents video frames The Middle Each entity is extracted. Visual features, normalized bounding boxes, and object category distribution; For entity collections Exhaustive subject-object pairing is performed on the entities to generate video frames. The set of all subject-object entity pairs Simultaneously, based on the visual features and normalized bounding boxes of each entity, each subject-object entity pair is represented. Derivation of subjective and objective visual features , Normalized bounding boxes for subject and object, same , express The index; for each pair of subject-object entities. Assign a unique relation proposal index , Indicates the first Subject-object entity pair, , This represents the total number of subject-object entity pairs, thus forming a triple. ; triplet The normalized bounding boxes of two entities are combined to form a joint bounding box. ,based on For video frames Perform RoI cropping to obtain the region of interest feature map. and corresponding union box visual features ; In this context, RoI cropping represents the extraction of a specific region of interest from an image or video; the object category distribution of an entity represents the probability of the category to which the entity belongs, which is used to convert it into a semantic embedding vector for each entity, serving as an initial constraint for entity classification.

[0023] In a specific implementation, the location encoding process is as follows: The feature maps of each region of interest are converted into serialized feature tensors and used as input to the VOCM module. After input, spatial and temporal position encoding is performed. Spatial position encoding adopts a geometric encoding strategy based on object detection bounding boxes, while temporal position encoding adopts a fixed encoding form using sine and cosine functions. The generated spatial position encoding and temporal position encoding are concatenated to obtain a position encoding vector, which is then added to the input serialized feature tensor to obtain the enhanced feature tensor.

[0024] In a specific implementation, the Transformer encoder process is as follows: The enhanced feature tensor is input into the Transformer encoder, which captures the complex relationships between objects within and between frames through a multi-head self-attention mechanism. Each head converts the input enhanced feature tensor into a query, key, and value, and calculates the weight matrix of each head through the self-attention mechanism. The weight matrices of each head are concatenated and combined with a linear transformation matrix to generate a fused weight matrix. Then, the fused weight matrix is ​​passed through residual connections and a feedforward neural network to retain the original input information and perform nonlinear enhancement, resulting in the features output by the Transformer encoder. The VOCM module adopts a multi-layer Transformer encoder stack structure. The output of the previous Transformer encoder is used as the input of the next Transformer encoder, and the output of the last Transformer encoder is the final output feature of the Transformer encoder.

[0025] In a specific implementation, the temporal LSTM processing procedure is as follows: The features output by the Transformer encoder are processed by temporal LSTM. The input features correspond to the video frames input by S1. The video frames are temporal sequence data, and the corresponding features are also temporal sequence data. Each feature processed by temporal LSTM includes two internal state vectors, namely the hidden state and the cell state. When the feature of the current time step is processed by temporal LSTM, the hidden state and cell state of the feature of the previous time step are combined to update the feature of the current time step, and the updated feature of the current time step is obtained. This process is repeated multiple times until the updated feature of the last time step is output. The hidden state of the last time step is selected as the final temporal feature vector of each object. The temporal feature vectors output by each object are then passed through a two-layer feedforward neural network to output the classification prediction results of each entity. At the same time, based on the temporal feature vectors corresponding to the subject-object entity pairs, the corresponding relationship classification features and classification prediction results are output through a two-layer feedforward neural network. The VOCM module finally outputs entity classification features and relationship classification features containing intra-frame and inter-frame spatiotemporal dependencies and temporal dynamic information, as well as the corresponding entity and relationship classification prediction results. Calculate the entity cross-entropy loss based on the entity classification prediction results, calculate the relationship cross-entropy loss based on the relationship classification prediction results, and train and optimize VOCM.

[0026] In a specific implementation, the region features and feature maps input to the TRM module after cropping are processed, and the specific operations are as follows: Based on the visual feature vector of each subject-object entity pair , Visual features of the joint frame Coordinates of the normalized bounding box of the subject and object , and semantic embedding vectors , Perform feature concatenation and dimensional transformation to generate relational features. This leads to the set of relational features. The semantic embedding vector is derived from the semantic embedding vector of each entity. The calculation formula is as follows: , in, , , represent the learnable linear matrices for dimensionality compression; Indicates the flattening operation; This represents an element-wise addition operation; Represents the bounding box feature uniform function; Indicates feature concatenation operation; Indicates the first In the first time step A relational feature vector, .

[0027] In a specific implementation, the operation process of the spatial Vision Transformer encoder is as follows: Each Set of relational features at each time step The input is fed into the TRM module, where it is processed by the spatial Vision Transformer encoder. First, it is reconstructed using a frame partitioning formula, discretizing the continuous video stream into spatiotemporally consistent segment representations, thus obtaining the relation index set for each frame. ; Then, zero-padded the relation feature matrix of each frame. Expanded to a maximum length per frame Generate batch tensors , This indicates a fill operation. Represents batch processing tensors Dimensions Represents the set of relation feature matrices; simultaneously generates a binary mask matrix. Used to distinguish between valid and invalid positions. , If the first The first video frame of the time step If each location contains a true visual relationship, then Otherwise, it is 1; Then process the tensors in the batch Normalization is performed to obtain the normalized features. Next, the spatial self-attention module captures the spatial topological relationships between objects through dynamic weight allocation, and calculates them through linear projection. The query, key, and value are combined with a binary mask matrix to calculate the output of the self-attention module. This output is then passed through a residual connection and a feedforward network to output intra-frame spatial relationship features.

[0028] In a specific implementation, the operation process of the timing Vision Transformer encoder is as follows: Based on the object semantic category embedding vector obtained from the object category distribution of entities, intra-frame spatial relationship features are grouped and reorganized, and features of the same category are aggregated along the temporal dimension to form... Each object represents a set of independent temporal feature sequences, and the relation index set for each object is represented as follows: , ; For each temporal feature sequence, a dual position information injection of absolute position index and relative position embedding is performed to obtain a temporal feature sequence after injection of position information. A unique temporal position identifier is generated for each time step through absolute position index, and the temporal distance between adjacent frame keys is mapped into a continuous vector representation through relative position embedding. Then, a dynamic filling strategy is used to dimensionally align the temporal feature sequences after the injection of location information, expanding them to the maximum sequence length allowed for each object category, thus obtaining the temporal feature matrix; Finally, the temporal feature matrix is ​​input into the Transformer unit with the same structure as the spatial Vision Transformer encoder, and the features are processed along the temporal dimension to capture the temporal dependencies across frames, thereby obtaining intra-frame temporal relationship features. These features are then fused with intra-frame spatial relationship features to obtain spatiotemporal fusion relationship features.

[0029] In a specific implementation, the operation process of the time-difference aggregator (TDA) is as follows: The spatiotemporal fusion relationship features are subjected to dimensionality compression and time step pairing to obtain cross-time step feature tensors. A shift sequence mechanism is then used to construct the first... The shift position characteristics of each time step force attention to the local temporal changes of adjacent time steps; The shift position features are concatenated with the cross-time step feature tensor of the current time step along the feature dimension, and then the concatenated part is transformed by a linear transformation. 3D feature projection back The dimension is used to obtain the differential features, which are then used to generate differential attention weights through two layers of nonlinear mapping; The cross-time step feature tensor is enhanced by differential attention weights to obtain the enhanced shift position features. The shift position features are then multiplied element-wise with the differential attention weights generated by the second-layer nonlinear mapping to obtain the enhanced temporal features. Finally, the enhanced temporal features are concatenated with the cross-time step feature tensor to obtain the fused feature vector. Based on the relation proposal index of the subject-object entity pair, the fused feature vector is injected into the feature position of the corresponding relation through the scatter projection operation. The scatter operation injects the fused feature vector into the corresponding position according to the relation proposal index. The TRM module finally generates dynamic temporal relation features enhanced by differential attention.

[0030] In a specific implementation, S3 is as follows: The entity and relation classification features output by the VOCM module and the dynamic temporal relation features output by the TRM module are dimensionally aligned. Then, a learnable gating fusion mechanism is used to adaptively fuse the two types of features after dimension alignment to obtain a fusion gating vector. Finally, the fusion feature is calculated based on the fusion gating vector. The calculation formula is as follows: , in, Indicates the final fusion characteristics, Represents the fusion gate vector. This indicates the output of the VOCM module. This indicates the output of the TRM module, and ⊙ represents element-wise multiplication. The final fused features are then input into the entity classification head and the relationship classification head, respectively. After passing through the feedforward neural network and the Softmax activation function, the final category prediction results of entities in each video frame and the final category prediction results of the relationship between the subject and object entities are output. Using entities in each video frame as nodes and the predicted relationships between entities as edges, a single-frame scene graph is constructed for each time step of the video frame. All single-frame scene graphs are then concatenated in temporal order according to the video frames to generate the final dynamic scene graph sequence.

[0031] Example 2 To demonstrate the technical effectiveness of this invention, the method was validated on the AG dataset, which is based on the Charades dataset and provides frame-level scene annotations. Specifically, the AG dataset annotates 234,253 images, covering 35 object categories (excluding humans), containing 476,229 bounding boxes and 1,715,568 instances, involving 25 relationship categories. These 25 relationships can be further divided into three types: attention relationships (representing whether a person is looking at an object), spatial relationships (representing the interaction method with an object), and contact relationships (representing the physical contact method with an object). All experiments used the same data partitioning scheme as previous studies. In the experimental setup, the initial training set contained 7,464 videos and 167,068 images. After filtering, 175 videos without human subjects, 87 single-frame videos, and 21,167 images without human annotations were removed, ultimately retaining 7,202 valid videos and 145,901 valid images. The original test set consisted of 1,737 videos with a total of 54,429 frames. After a similar filtering process, the final test set contained 1,669 videos and 45,957 labeled images for evaluation. To generate scene graphs with inferred relation distributions, two strategies were employed: (1) constrained mode: only one predicate is allowed to be predicted for each subject-object pair; (2) unconstrained mode: there is no limit to the number of relation predictions for subject-object pairs. These strategies aim to better understand and predict dynamic scenes in videos by generating scene graphs with inferred relation distributions.

[0032] In predicate classification tasks, since only the predicate categories between entities need to be predicted (essentially a classification problem), most methods can achieve a Recall@50 of over 99% under unconstrained settings, indicating that relation prediction has achieved significant success. However, scene graph classification and scene graph detection tasks require the simultaneous prediction of object categories and predicate categories, which significantly increases the complexity and challenge of the tasks. To explore the difficulties of these tasks in depth, the following two tasks were carried out on the AG dataset: (1) Scene graph classification task: only entity bounding boxes are provided, and the model needs to simultaneously predict the object category and predicate category corresponding to each bounding box; (2) Scene graph detection task: the model needs to detect entities and their bounding boxes from the image, and at the same time identify the entity category and predict the relevant predicate category.

[0033] To maintain consistency with the baseline model, Recall@K (R@K) and average Recall@K (mR@K) were used as evaluation metrics for the STVRM model, with K taking values ​​of [10, 20, 50]. R@K measures the proportion of correct relations among the top K predictions, primarily focusing on the evaluation of high-frequency predicate categories; while mR@K, as a more balanced metric, alleviates the class imbalance problem by calculating the average recall for each category, thus comprehensively evaluating the model's scene graph generation performance across all relation categories. For predicate selection strategies, both "constrained" and "unconstrained" approaches were adopted to ensure comparability with the baseline model. The "constrained" strategy is the most stringent, allowing only one predicate per entity pair, while the "unconstrained" strategy allows multiple relations to be predicted for each entity pair. These settings provide a comprehensive perspective for evaluating the model's performance on the AG dataset, while ensuring fully comparable analysis with existing baseline models.

[0034] The Faster R-CNN model with ResNet-101 as the backbone was chosen as the object detector to extract detection results from the video. First, the detector was trained on the AG training set, with an IoU of 0.5 during training and 0.4 during testing. The non-maximum suppression threshold for people was set to 0.4 to ensure the stability of person detection in a single frame. To ensure fair comparison, this detector was applied to all baseline methods. The proposed method is implemented in PyTorch and adopts a spatiotemporal ViT architecture for scene graph classification and detection tasks. Its encoder-decoder layer count and attention head count are consistent with STTran: the spatial encoder contains one level, and the temporal decoder contains three iterative levels. This architecture can effectively capture the spatiotemporal features of the video and handle complex spatiotemporal relationships.

[0035] The model was trained using the AdamW optimizer with a learning rate of 1e-5, a batch size of 1, and a maximum training epoch of 10. The input data had a dimension of 1936, and the dropout rate was set to 0.1 to prevent overfitting. Internally, the 1936-dimensional input features were first projected to 2048 dimensions through a feedforward network, then non-linearly transformed using the GELU activation function to restore them to 1936 dimensions. Subsequently, the features were expanded to 3872 dimensions through a TDA layer, activated by Tanh, and then projected back to 1936 dimensions for subsequent processing. This process effectively helps the model better capture the relationships between objects in the spatiotemporal dimensions.

[0036] Table 1 compares the Recall@K (%) of the proposed STVRM model with state-of-the-art scene graph generation methods on the Action Genome dataset for SGCLS scene graph classification and SGDET scene graph detection tasks, covering both constrained and unconstrained conditions. Best results are highlighted in bold, with "-" indicating missing data. To ensure a fair comparison, all baseline methods used the same object detector. The most advanced scene graph generation methods include VRD visual relationship detection, M-FREQ multi-level frequency-guided relationship reasoning, MSDN multi-scale dynamic network, VCTREE visual context tree, ReIDN relationship-aware identity network, GPS-Net graph-based location-sensitive network, TRACE Transformer-based relationship attention and context embedding, STTram spatiotemporal Transformer relationship modeling, TPI triple progressive reasoning, TR2 Transformer-based relationship reasoning and optimization, GTR graph Transformer reasoning, UTM Transformer-based unified multi-task network, TD2-Net two-stage dynamic detection and reasoning network, and STABILE spatiotemporal attention-based incremental learning of scene graphs.

[0037] Table 1. Performance Comparison of the Invention Method and the Cutting-Edge Scene Graph Generation Method Table 2 shows the mean-Recall@K (%) performance comparison results of the proposed STVRM method and STTran method on the SGCLS and SGDET tasks on the AG dataset, including both constrained and unconstrained conditions. The best results are marked in bold, and "-" indicates that the data is not yet available.

[0038] Table 2 Performance Comparison of the Invention Method and the STTran Method The proposed STVRM method is comprehensively compared with existing dynamic and static scene graph generation methods. Table 1 shows the R@K performance of the STVRM method in scene graph classification and detection tasks, while Table 2 presents its mR@K metric. To ensure fairness, all methods use the same object detector. The first six static scene graph generation methods generally perform worse than this method because they ignore the temporal continuity between video frames, highlighting the importance of temporal continuity in dynamic scene graph generation tasks.

[0039] Under the constraints, this model improves upon the baseline by 4.1% / 4.2% on the R@10 / 20 metrics for scene graph classification and by 4.7% / 0.3% on the R@10 / 20 metrics for scene graph detection; it also improves by 2% / 2.1% on the mR@10 / 20 metrics for scene graph classification and by 0.3% on the mR@10 metrics for scene graph detection. Notably, when the K value is large in the scene graph detection task, the improvement in R@K and mR@K metrics by STVRM weakens, reflecting the balance between fine-grained prediction and multiple possibility considerations. Under the unconstrained settings, except for R@50 and mR@50 in the scene graph detection task, this model outperforms other methods in all settings. This is because the unconstrained strategy allows for multiple relation predictions for each entity pair, and the R@50 metric provides more room for guessing, leading to reduced stability of the results. Nevertheless, this model still significantly outperforms other methods in terms of R@10 / 20 and mR@10 / 20 metrics, indicating that the model of this invention has higher reliability in scenarios where the guessing space is limited and accurate prediction is required.

Claims

1. A method for generating dynamic scene graphs based on spatiotemporal relationship modeling using Vision Transformer, characterized in that, Includes the following steps: S1. Input the video frame sequence into the target detector to detect entities in the video frames, and extract the visual features of the entities, normalized bounding boxes and object category distribution to form triples. Crop the input video frames, retain the regions related to the triples, and obtain the cropped region features and feature maps. S2. The cropped region features and feature maps are processed and then input into the Video Object Relationship Classification (VOCM) module. The VOCM module's processing includes position encoding, Transformer encoder processing, and temporal LSTM processing. The VOCM module classifies entities and relationships in the image to obtain the classified entity and relationship features. Meanwhile, the cropped region features and feature maps are processed and input into the temporal relationship perception module (TRM). The TRM module includes a spatial Vision Transformer encoder, a temporal Vision Transformer encoder, and a temporal difference aggregator (TDA). First, the two encoders capture the relationship features of the spatial and temporal modalities respectively. Then, the TDA constructs the difference features to capture the changes of the relationship features between time steps and generate dynamic temporal relationship features. S3. The outputs of the VOCM module and the TRM module are combined to generate the final dynamic scene graph sequence.

2. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 1, characterized in that, S1 is as follows: video frame sequence The input is fed into a Faster R-CNN object detector pre-trained on the Action Genome benchmark dataset to detect each video frame. The entity in Indicates the total number of video frames. express index, Indicates the first Each time step of the video frame; video frames The corresponding entity set is represented as , Indicates the first The total number of entities in a video frame at each time step, where each entity is an object. express index, Represents video frames The Middle Each entity is extracted. Visual features, normalized bounding boxes, and object category distribution; For entity collections Exhaustive subject-object pairing is performed on the entities to generate video frames. The set of all subject-object entity pairs Simultaneously, based on the visual features and normalized bounding boxes of each entity, each subject-object entity pair is represented. Derivation of subjective and objective visual features , Normalized bounding boxes for subject and object, same , express The index; for each pair of subject-object entities. Assign a unique relation proposal index , Indicates the first Subject-object entity pair, , This represents the total number of subject-object entity pairs, thus forming a triple. ; triplet The normalized bounding boxes of two entities are combined to form a joint bounding box. ,based on For video frames Perform RoI cropping to obtain the region of interest feature map. and corresponding union box visual features ; In this context, RoI cropping represents the extraction of a specific region of interest from an image or video; the object category distribution of an entity represents the probability of the category to which the entity belongs, which is used to convert it into a semantic embedding vector for each entity, serving as an initial constraint for entity classification.

3. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 2, characterized in that, The specific operation process of position encoding is as follows: The feature maps of each region of interest are converted into serialized feature tensors and used as input to the VOCM module. After input, spatial and temporal position encoding is performed. Spatial position encoding adopts a geometric encoding strategy based on object detection bounding boxes, while temporal position encoding adopts a fixed encoding form using sine and cosine functions. The generated spatial position encoding and temporal position encoding are concatenated to obtain a position encoding vector, which is then added to the input serialized feature tensor to obtain the enhanced feature tensor.

4. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 3, characterized in that, The specific process of the Transformer encoder is as follows: The enhanced feature tensor is input into the Transformer encoder, which captures the complex relationships between objects within and between frames through a multi-head self-attention mechanism. Each head converts the input enhanced feature tensor into a query, key, and value, and calculates the weight matrix of each head through the self-attention mechanism. The weight matrices of each head are concatenated and combined with a linear transformation matrix to generate a fused weight matrix. Then, the fused weight matrix is ​​passed through residual connections and a feedforward neural network to retain the original input information and perform nonlinear enhancement, resulting in the features output by the Transformer encoder. The VOCM module adopts a multi-layer Transformer encoder stack structure. The output of the previous Transformer encoder is used as the input of the next Transformer encoder, and the output of the last Transformer encoder is the final output feature of the Transformer encoder.

5. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 4, characterized in that, The specific process of temporal LSTM processing is as follows: The features output by the Transformer encoder are processed by temporal LSTM. The input features correspond to the video frames input by S1. The video frames are temporal sequence data, and the corresponding features are also temporal sequence data. Each feature processed by temporal LSTM includes two internal state vectors, namely the hidden state and the cell state. When the feature of the current time step is processed by temporal LSTM, the hidden state and cell state of the feature of the previous time step are combined to update the feature of the current time step, and the updated feature of the current time step is obtained. This process is repeated multiple times until the updated feature of the last time step is output. The hidden state of the last time step is selected as the final temporal feature vector of each object. The temporal feature vectors output by each object are then passed through a two-layer feedforward neural network to output the classification prediction results of each entity. At the same time, based on the temporal feature vectors corresponding to the subject-object entity pairs, the corresponding relationship classification features and classification prediction results are output through a two-layer feedforward neural network. The VOCM module finally outputs entity classification features and relationship classification features containing intra-frame and inter-frame spatiotemporal dependencies and temporal dynamic information, as well as the corresponding entity and relationship classification prediction results. Calculate the entity cross-entropy loss based on the entity classification prediction results, calculate the relationship cross-entropy loss based on the relationship classification prediction results, and train and optimize VOCM.

6. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 5, characterized in that, The region features and feature maps after being cropped and input into the TRM module are processed as follows: Based on the visual feature vector of each subject-object entity pair , Visual features of the joint frame Coordinates of the normalized bounding box of the subject and object , and semantic embedding vectors , Perform feature concatenation and dimensionality transformation to generate relational features. This leads to the set of relational features. The semantic embedding vector is derived from the semantic embedding vector of each entity. The calculation formula is as follows: , in, , , represent the learnable linear matrices for dimensionality compression; Indicates the flattening operation; This represents an element-wise addition operation; Represents the bounding box feature uniform function; Indicates feature concatenation operation; Indicates the first In the first time step A relational feature vector, .

7. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 6, characterized in that, spatial... The operation process of the Vision Transformer encoder is as follows: Each Set of relational features at each time step The input is fed into the TRM module, where it is processed by the spatial Vision Transformer encoder. First, it is reconstructed using a frame partitioning formula, discretizing the continuous video stream into spatiotemporally consistent segment representations, thus obtaining the relation index set for each frame. ; Then, zero-padded the relation feature matrix of each frame. Expanded to a maximum length per frame Generate batch tensors , This indicates a fill operation. Represents batch processing tensors Dimensions Represents the set of relation feature matrices; simultaneously generates a binary mask matrix. Used to distinguish between valid and invalid positions. , If the first The first video frame of the time step If each location contains a true visual relationship, then Otherwise, it is 1; Then process the tensors in the batch Normalization is performed to obtain the normalized features. Next, the spatial self-attention module captures the spatial topological relationships between objects through dynamic weight allocation, and calculates them through linear projection. The query, key, and value are combined with a binary mask matrix to calculate the output of the self-attention module. This output is then passed through a residual connection and a feedforward network to output intra-frame spatial relationship features.

8. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 7, characterized in that, The operation process of the timing Vision Transformer encoder is as follows: Based on the object semantic category embedding vector obtained from the object category distribution of entities, intra-frame spatial relationship features are grouped and reorganized, and features of the same category are aggregated along the temporal dimension to form... Each object represents a set of independent temporal feature sequences, and the relation index set for each object is represented as follows: , ; For each temporal feature sequence, a dual position information injection of absolute position index and relative position embedding is performed to obtain a temporal feature sequence after injection of position information. A unique temporal position identifier is generated for each time step through absolute position index, and the temporal distance between adjacent frame keys is mapped into a continuous vector representation through relative position embedding. Then, a dynamic filling strategy is used to dimensionally align the temporal feature sequences after the injection of location information, expanding them to the maximum sequence length allowed for each object category, thus obtaining the temporal feature matrix; Finally, the temporal feature matrix is ​​input into the Transformer unit with the same structure as the spatial Vision Transformer encoder, and the features are processed along the temporal dimension to capture the temporal dependencies across frames, thereby obtaining intra-frame temporal relationship features. These features are then fused with intra-frame spatial relationship features to obtain spatiotemporal fusion relationship features.

9. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 8, characterized in that, temporal sequence... The specific operation process of the Differential Aggregator (TDA) is as follows: The spatiotemporal fusion relationship features are subjected to dimensionality compression and time step pairing to obtain cross-time step feature tensors. A shift sequence mechanism is then used to construct the first... The shift position characteristics of each time step force attention to the local temporal changes of adjacent time steps; The shift position features are concatenated with the cross-time step feature tensor of the current time step along the feature dimension, and then the concatenated part is transformed by a linear transformation. 3D feature projection back The dimension is used to obtain the differential features, which are then used to generate differential attention weights through two layers of nonlinear mapping; The cross-time step feature tensor is enhanced by differential attention weights to obtain the enhanced shift position features. The shift position features are then multiplied element-wise with the differential attention weights generated by the second-layer nonlinear mapping to obtain the enhanced temporal features. Finally, the enhanced temporal features are concatenated with the cross-time step feature tensor to obtain the fused feature vector. Based on the relation proposal index of the subject-object entity pair, the fused feature vector is injected into the feature position of the corresponding relation through the scatter projection operation. The scatter operation injects the fused feature vector into the corresponding position according to the relation proposal index. The TRM module finally generates dynamic temporal relation features enhanced by differential attention.

10. The dynamic scene graph generation method based on Vision Transformer spatiotemporal relationship modeling according to claim 9, characterized in that, S3 is as follows: The entity and relation classification features output by the VOCM module and the dynamic temporal relation features output by the TRM module are dimensionally aligned. Then, a learnable gating fusion mechanism is used to adaptively fuse the two types of features after dimension alignment to obtain a fusion gating vector. Finally, the fusion feature is calculated based on the fusion gating vector. The calculation formula is as follows: , in, Indicates the final fusion characteristics, Represents the fusion gate vector. This indicates the output of the VOCM module. This indicates the output of the TRM module, and ⊙ represents element-wise multiplication. The final fused features are then input into the entity classification head and the relationship classification head, respectively. After passing through the feedforward neural network and the Softmax activation function, the final category prediction results of entities in each video frame and the final category prediction results of the relationship between the subject and object entities are output. Using entities in each video frame as nodes and the predicted relationships between entities as edges, a single-frame scene graph is constructed for each time step of the video frame. All single-frame scene graphs are then concatenated in temporal order according to the video frames to generate the final dynamic scene graph sequence.