Behavior recognition method based on scene semantic understanding

By generating dynamic scene graph dictionary and cross-modal features fusion, the existing methods are solved in complex scenarios, and the efficiency and accuracy of behavior recognition in rail transit stations are achieved, and the real-time performance of safety monitoring is improved.

CN120299084APending Publication Date: 2025-07-11NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373090.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing behavior recognition method based on Transformer architecture has problems such as low dynamics, high latency, and high information redundancy in complex scenarios, and it is difficult to fully consider the scene constraint characteristics of abnormal behavior, resulting in insufficient real-time and accuracy in rail transit passenger safety monitoring.

Method used

A behavior recognition method based on scene semantic understanding is adopted. By obtaining the entity features and relationship features of video frames, a dynamic scene graph dictionary is generated, combining cross-modal feature fusion and spatiotemporal relationship modeling, and using dynamic scene graph dictionary for behavior recognition, reducing redundant information processing, and improving the real-time performance of the model.

Benefits of technology

It significantly improves the accuracy and stability of behavior recognition, can respond quickly in complex scenarios, reduces redundant information interference, and enhances the safety monitoring capabilities in complex environments such as rail transit stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299084A_ABST
    Figure CN120299084A_ABST
Patent Text Reader

Abstract

The invention provides a behavior recognition method based on scene semantic understanding, relates to the technical field of behavior recognition, and aims to enhance logic constraint of a dynamic scene graph in relation change by utilizing fusion of cross-modal features, guiding generation of the dynamic scene graph and utilizing time sequence information to a greater extent, so that the behavior recognition efficiency is improved. Through a dynamic scene graph dictionary, scene semantics are abstracted, key information is intelligently screened, redundant data processing amount is greatly reduced, model operation is accelerated, real-time performance is enhanced, and rapid response in a complex scene is ensured. In the behavior identification, the video frame features are spatially weighted by using field semantics, so that the interference of redundant information on the behavior identification is reduced, and the speed and accuracy of the behavior identification are improved. Meanwhile, according to the method, dynamic scene graph analysis and an MLP-mix layer technology are adopted, multi-angle and refined behavior recognition is achieved, and the accuracy and stability of behavior recognition are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior recognition, and more particularly to a behavior recognition method based on scene semantic understanding. Background Art

[0002] In a highly crowded place such as a rail transit station, the probability of accidents is relatively high, which undoubtedly poses a huge operation management challenge to station operation management personnel. Against this background, how to make full use of existing technologies and resources to effectively ensure the travel safety of passengers has become a key problem to be solved urgently.

[0003] In daily life, existing behavior recognition technologies, such as methods like the two-stream network [1], can perform simple video behavior recognition, i.e., action classification, which only needs to accurately classify a given video into several known action categories to provide general behavior recognition results. However, abnormal behaviors often have scene limitations. For example, a striding action is a normal sports activity on a sports field but is an abnormal behavior above the subway turnstile. In addition, compared with general behaviors in life scenarios, the recognition of abnormal behaviors of rail transit passengers faces greater challenges. The video content and background are more complex and changeable. Different behaviors may be similar, and the same behavior may vary in different environments. Occlusion factors and a large number of people flow result in a large amount of redundant information in a single frame, which puts higher requirements on the real-time performance and accuracy of abnormal behavior recognition.

[0004] In recent years, models based on Transformer have performed well in behavior recognition. Their advantage lies in being able to effectively capture spatio-temporal dependence relationships [2], improving the performance of the model in complex scenarios. Compared with traditional convolutional neural networks (CNNs), multi-modal fusion improves the robustness of behavior recognition by fusing multiple modal information such as RGB, depth, and infrared. Lightweight model design has proposed various lightweight models, such as VideoLightFormer [3] and TinyVIRAT [4], for mobile devices and low-resolution videos. Unsupervised learning learns behavior features through unsupervised methods, reducing the dependence on labeled data. Although Transformer performs well in capturing global dependence relationships, in some tasks, its ability to capture local details may be inferior to that of CNN.

[0005] Although the Transformer architecture has shown unique advantages in the field of behavior recognition, there are still some obvious defects in the current behavior recognition methods based on this architecture. Traditional behavior recognition methods based on the Transformer architecture usually process each pixel in the entire video frame, and this processing method exposes a series of problems when facing complex scenarios. On the one hand, there are problems such as low dynamics, high latency, and high information redundancy. On the other hand, it is difficult for this method to fully consider the scene constraint characteristics of abnormal behaviors, resulting in insufficient adaptability in complex scenarios of real life. In the work of rail transit passenger safety monitoring, complex scenarios such as occlusion and fast movement are common challenges. In these scenarios, traditional behavior recognition methods based on the Transformer architecture are prone to problems such as motion blur and poor real-time performance, which will seriously weaken the performance of the system and affect the accuracy and timeliness of abnormal behavior recognition. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the purpose of the present invention is to propose a behavior recognition method based on scene semantic understanding, including:

[0007] Step 1: Obtain the target video in the target scene, and then obtain the feature map F of each frame of image in the target video t , and the entity features of each entity in each frame of image where t represents the video frame number, t ∈ [1, T], T represents the total number of frames of the target video, i represents the entity index, i ∈ [1, N], N represents the total number of entity indexes, is the entity feature of the i-th entity in the t-th frame of image, is the bounding box coordinate of the i-th entity in the t-th frame of image, is the entity category of the i-th entity in the t-th frame of image, is the visual feature of the i-th entity in the t-th frame of image, and all entity features form an entity feature set

[0008] Step 2: Process all entity features of each frame of image to obtain the relationship feature vector of entities i and j with existing relationships in the image and the enhanced image features;

[0009] Step 3: For each frame of image, process the relationship feature vector and the enhanced image features to obtain the dynamic scene graph dictionary G corresponding to each frame of image t , the dynamic scene graph dictionary includes the relationships between entities in the image, the relationship strengths of the relationships between entities, the bounding box coordinates of the entities constituting the relationships, and the entity categories of the entities constituting the relationships;

[0010] Step 4: According to the dynamic scene graph dictionary G t and the feature map F t , determine the action recognition result of the target video, where the action recognition result indicates whether there is an abnormality in the action in the target video.

[0011] Optionally, step 1 specifically includes:

[0012] Step 1.1: Obtain the target video in the target scene, where the target video includes T frame images, and the T frame images form a frame sequence {Fr t}, where Fr t represents the t-th frame image;

[0013] Step 1.2: Input the frame sequence {Fr t} into the convolutional neural network of the pre-trained Faster R-CNN model to obtain the entity features of each entity in each frame image. All the entity features in the frame sequence {Fr t} form an entity feature set

[0014] Optionally, step 1.2 specifically includes:

[0015] Step 1.2.1: Input each frame image in the frame sequence {Fr t} into the convolutional neural network CNN to obtain the feature map F t of each frame image, which is specifically implemented by the following formula:

[0016] F t = CNN(Fr t );

[0017] where F t represents the feature map of the t-th frame image;

[0018] Step 1.2.2: For the feature map F t of each frame image, process the feature map through the Region Proposal Network RPN to obtain multiple candidate regions;

[0019] Specifically, process the feature map through RPN to obtain the feature mapping of the feature map, which is specifically implemented by the following formula:

[0020] U t = RPN(F t );

[0021] where U t represents the feature mapping of F t ;

[0022] Slide through the sliding window of RNP on the feature mapping of the feature map to obtain multiple candidate regions;

[0023] Step 1.2.3: For each candidate region, project the candidate region onto the feature map corresponding to the candidate region to obtain the feature matrix of the region of interest, and further obtain the feature matrices of multiple regions of interest;

[0024] Step 1.2.4: For the feature matrix of each region of interest, process the feature matrix of the region of interest through the ROIPooling layer to obtain the flattened feature vector, and further obtain multiple flattened feature vectors;

[0025] Step 1.2.5: For each flattened feature vector, input the flattened feature vector into a series of fully connected layers to obtain the target vector corresponding to each candidate region. According to the target vector, determine the entity features of the entities in the candidate region, obtain the entity features of all entities in a frame of image, and further obtain the entity features of each frame of image in the frame sequence. All entity features form an entity feature set

[0026] Optionally, step 2 specifically includes:

[0027] Step 2.1: For all entity features of each frame of image, encode the entity features through a pre-trained image encoder to obtain entity i and entity j with existing relationships in the image, and the joint image features

[0028] Step 2.2: Fuse the joint image features with the entity features of entity i and the entity features of entity j to obtain the relationship feature vector Specifically, it is implemented through the following formula:

[0029]

[0030] Step 2.3: Input the entity category of entity i as and that of entity j into the pre-trained text encoder to extract the joint text features

[0031] Step 2.4: Use the extracted joint text features to perform relationship-aware prompting on the joint image features to obtain the optimized image features Specifically, it is implemented through the following formula:

[0032]

[0033] where α represents the prompting intensity parameter;

[0034] Step 2.5: Based on the knowledge distillation loss function, by comparing the joint image features output by the image encoder and the joint text features output by the text encoder perform knowledge distillation to obtain the knowledge distillation loss The knowledge distillation loss function is represented by the following formula:

[0035]

[0036] where, ‖‖2 represents the Euclidean norm;

[0037] Combine the knowledge distillation loss with the optimized image features to obtain the enhanced image features.

[0038] Optionally, step 3 specifically includes:

[0039] Step 3.1: For each frame of image, input the relationship feature vector and the enhanced image features into the spatial encoder to obtain the spatial relationship features;

[0040] Step 3.2: Input the relationship feature vector and the spatial relationship features into the temporal decoder to obtain the spatio-temporal relationship features

[0041] Step 3.3: Introduce the inter-frame message token Through the feed-forward network, fuse the introduced inter-frame message token and the spatio-temporal relationship features to generate the spatio-temporal relationship features that fuse short-term dynamic information Specifically, it is implemented through the following formula:

[0042]

[0043] where, represents the dynamic weight parameter;

[0044] Step 3.4: Use the relationship classifier to perform relationship classification on the spatio-temporal relationship features that fuse short-term dynamic information to generate the dynamic scene graph dictionary G corresponding to each frame of image t .

[0045] Optionally, step 3.3 further includes:

[0046] Update the inter-frame message token Specifically, it is implemented through the following formula:

[0047]

[0048] Among them, represents the inter-frame message token after update, represents the inter-frame message token of the (t - 1)-th frame, and β is the message update intensity parameter, is the spatio-temporal relationship feature of the t-th frame, is the spatio-temporal relationship feature of the (t - 1)-th frame.

[0049] Optionally, step 4 specifically includes:

[0050] Step 4.1: For the dynamic scene graph dictionary G corresponding to each frame of image t , determine the entities m and n with relationships in the dynamic scene graph dictionary. In the dynamic scene graph dictionary, obtain the entity category of entity m and the entity category of entity n, as well as the bounding box coordinates of entity m and the bounding box coordinates of entity n, and calculate the relationship weight Specifically, it is implemented through the following formula:

[0051]

[0052] Among them, represents the relative position of entity m and entity n;

[0053] According to the relationship weight and the dynamic scene graph dictionary G t , calculate the weighted feature map Specifically, it is implemented through the following formula:

[0054]

[0055] Step 4.2: According to the weighted feature map and the feature map F t , calculate the temporally enhanced frame feature F time ;

[0056] Step 4.3: Input the temporally enhanced frame feature F time into the query classifier. Specifically, use the global average pooling layer to compress the temporally enhanced frame feature F time into a vector to obtain the feature representation of the target video. Through the classifier Classifier and the softmax function, classify the feature representation of the target video to obtain the probability of each category. Specifically, it is implemented through the following formula:

[0057] y = Softmax(Classifier(GlobalPool(F time )));

[0058] Among all categories of probabilities, select the category with the highest probability value as the final behavior recognition result, that is, the behavior recognition result of the target video, and the behavior recognition result characterizes whether there is an abnormality in the behavior in the target video.

[0059] Optionally, step 4.2 specifically includes:

[0060] Step 4.2.1: For each frame of image, fuse the weighted feature map and the feature map F t to obtain the fused feature of each frame of image. Project the fused feature into a query vector Q, a key vector K, and a value vector V. Calculate the dot product of the query vector Q and the key vector K, normalize the dot product of the query vector Q and the key vector K to obtain a normalized vector, and perform a dot product of the normalized vector and the value vector to obtain an enhanced feature Furthermore, obtain the enhanced feature corresponding to each frame of image in the target video All enhanced features constitute a feature sequence I t ;

[0061] Step 4.2.2: Input the feature sequence I t into the MLP-mixer layer for mixing to obtain the mixed frame feature;

[0062] Step 4.2.3: Refine the mixed frame feature through the MLP-mixer layer to obtain the temporally enhanced frame feature

[0063] The beneficial effects produced by adopting the above technical solutions are as follows:

[0064] By utilizing the fusion of cross-modal features, the present invention guides the generation of dynamic scene graphs, makes greater use of temporal information, strengthens the logical constraint of dynamic scene graphs in terms of relationship changes, abstracts scene semantics through a dynamic scene graph dictionary, intelligently screens key information, greatly reduces the amount of redundant data processing, accelerates model operation, enhances real-time performance, and ensures rapid response in complex scenarios. In behavior recognition, field semantics is used to spatially weight the video frame features, thereby reducing the interference of redundant information on behavior recognition and improving the speed and accuracy of behavior recognition. At the same time, the present invention adopts dynamic scene graph parsing and MLP-mixer layer technologies to achieve multi-angle and refined recognition of behaviors, further improving the accuracy and stability of behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a schematic flowchart of a behavior recognition method based on scene semantic understanding in an embodiment of the present invention;

[0066] Figure 2 This is the implementation effect diagram of a behavior recognition method based on scene semantic understanding in the embodiments of the present invention. Among them, (a1) is the first frame of image, and (a2) is the dynamic scene graph dictionary G of (a1). t , (b1) is the second frame of image, and (b2) is the dynamic scene graph dictionary G of (b1). t . Detailed implementation manners

[0067] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0068] Aiming at the problems existing in the prior art, the present invention aims to propose a behavior recognition method based on scene semantic understanding to overcome the defects in the prior art. The inspiration for this method comes from the scene graph generation technology, which is a technology that can abstract visual information into semantic representations and has received extensive attention in the field of computer vision in recent years.

[0069] The definition of scene graph generation is: automatically detecting the objects and their relationships in the input image information and generating a graph structure (i.e., scene graph) composed of a series of <subject-relationship-object> triples. In the scene graph, the objects in the image correspond to the nodes in the graph structure, and the relationships between the objects correspond to the edges in the graph structure.

[0070] Compared with the traditional abnormal behavior recognition method, the present invention has significant advantages. On the one hand, it can help the computer deeply understand the scene semantics. In the process of determining whether the recognized behavior is abnormal, by detecting the human-object interaction and considering the scene constraints, the accuracy of abnormal behavior recognition can be significantly improved. On the other hand, using scene semantics can effectively reduce the processing of redundant information in complex scenes, thereby improving the real-time performance and working efficiency of the model and making it more suitable for complex environments in practical applications.

[0071] Specifically, the present invention can utilize scene semantics when facing a complex subway station scene, reduce the processing of redundant information, and improve the determination accuracy of abnormal behaviors, mainly including scene graph generation and behavior recognition. Scene graph generation can be further divided into: object detection and feature extraction, cross-modal feature guidance and fusion, and spatio-temporal relationship modeling.

[0072] Before implementing the solution of the present invention, certain preparatory work is required. First, collect video data of abnormal behaviors occurring in the subway station, then label the relevant data to make a personalized dataset, and train the parameters in steps 1 to 4 based on the personalized dataset. After the training is completed, in the actual use stage, in combination with Figure 1 , it may include the following steps:

[0073] Step 1: Obtain the target video in the target scenario, and then obtain the feature map F of each frame image in the target video t , as well as the entity features of each entity in each frame image Among them, t represents the video frame number, t ∈ [1, T], T represents the total number of frames of the target video, i represents the entity index, i ∈ [1, N], and N represents the total number of entity indexes is the entity feature of the i-th entity in the t-th frame image is the bounding box coordinates of the i-th entity in the t-th frame image is the entity category of the i-th entity in the t-th frame image is the visual feature of the i-th entity in the t-th frame image, and all entity features form the entity feature set

[0074] Step 1.1: Obtain the target video in the target scenario. The target video contains T frame images, and the T frame images contain the visual information that needs to be detected and feature-extracted. The T frame images form a frame sequence {Fr t}, where Fr t represents the t-th frame image

[0075] Step 1.2: Input the frame sequence {Fr t} into the convolutional neural network of the pre-trained Faster R-CNN model to obtain the entity features of each entity in each frame image. All entity features in the frame sequence {Fr t} form the entity feature set

[0076] Step 1.2.1: Input each frame image in the frame sequence {Fr t} into the convolutional neural network CNN to obtain the feature map F of each frame image t , which is specifically implemented through the following formula

[0077] F t = CNN(Fr t )

[0078] Among them, F t represents the feature map of the t-th frame image

[0079] Step 1.2.2: For the feature map F of each frame image t , process the feature map through the Region Proposal Network (RPN) to obtain multiple candidate regions

[0080] Among them, RPN can predict regions in the image that may contain objects, and assign foreground / background probability scores to each candidate region to help screen out the regions most likely to contain objects.

[0081] Specifically, by processing the feature map through RPN, a feature mapping of the feature map is obtained, which is specifically implemented through the following formula:

[0082] U t =RPN(F t );

[0083] Among them, U t represents the feature mapping of F t ;

[0084] Through the sliding window of RNP, slide on the feature mapping of the feature map to obtain multiple candidate regions;

[0085] Step 1.2.3: For each candidate region, project the candidate region onto the feature map corresponding to the candidate region to obtain the feature matrix of the region of interest, and then obtain multiple feature matrices of the regions of interest;

[0086] Among them, the feature matrix of the region of interest contains the feature information of the object in the candidate region, preparing for subsequent classification and regression tasks.

[0087] Step 1.2.4: For the feature matrix of each region of interest, through the ROIPooling layer, process the feature matrix of the region of interest, scale it to a fixed size, and obtain the flattened feature vector, and then obtain multiple flattened feature vectors;

[0088] Among them, the ROIPooling layer can adapt to regions of interest of different sizes, unify them into feature vectors of a fixed length, and facilitate subsequent processing by the fully connected layer.

[0089] Step 1.2.5: For each flattened feature vector, input the flattened feature vector into a series of fully connected layers to obtain the target vector corresponding to each candidate region. The fully connected layer can learn the complex patterns and relationships in the feature vector, improving the model's expression ability and classification accuracy for target features.

[0090] According to the target vector, determine the entity features of the entities in the candidate region, obtain the entity features of all entities in a frame of image, and then obtain the entity features of each frame of image in the frame sequence. All entity features form an entity feature set

[0091] Step 2: Process all entity features of each frame of image to obtain the relationship feature vector of entities i and j with existing relationships in the image and the enhanced image features;

[0092] Step 2.1: For all entity features of each frame of image, encode the entity features through a pre-trained image encoder to obtain entities i and j with relationships in the image, and joint image features

[0093] Step 2.2: The joint image features are fused with the entity features of entity i and the entity features of entity j to obtain a relationship feature vector Specifically, it is achieved through the following formula:

[0094]

[0095] Step 2.3: The entity category of entity i is and that of entity j are input into a pre-trained text encoder, and using the text prompt library learned during training, the joint text features

[0096] are extracted. Among them, the text encoder can understand the semantic information in the text and convert it into a vector representation compatible with the image features for cross-modal interaction and fusion.

[0097] Step 2.4: Through the extracted joint text features perform relationship-aware prompting on the joint image features to obtain optimized image features to enhance the expression ability of the image features and make them pay more attention to the features related to relationship understanding. Specifically, it is achieved through the following formula:

[0098]

[0099] Among them, α represents the prompting intensity parameter, which is used to control the influence degree of the text features on the image features;

[0100] Step 2.5: Based on the knowledge distillation loss function, perform knowledge distillation by comparing the joint image features output by the image encoder and the output joint text features of the text encoder to obtain the knowledge distillation loss The knowledge distillation loss function is represented by the following formula:

[0101]

[0102] Among them, ‖‖2 represents the Euclidean norm;

[0103] The knowledge distillation loss Combine with the optimized image features to obtain enhanced image features.

[0104] Step 3: For each frame of image, process the relational feature vector and the enhanced image features to obtain the dynamic scene graph dictionary G t corresponding to each frame of image. The dynamic scene graph dictionary contains the relationships between entities in the image, the relationship strengths between the relationships of entities, the bounding box coordinates of the entities constituting the relationships, and the entity categories of the entities constituting the relationships;

[0105] Step 3.1: For each frame of image, input the relational feature vector and the enhanced image features into the spatial encoder to perform information fusion in the spatial dimension for the relational representation, capture the spatial relationships and interactions between the targets, and obtain spatial relational features;

[0106] Among them, the spatial encoder is based on the Transformer architecture. Through multiple layers of processing of the spatial encoder, a relational representation integrating spatial relational features is obtained, enhancing the model's ability to understand the spatial relationships between targets.

[0107] Step 3.2: Input the relational feature vector and the spatial relational features into the temporal decoder to perform information fusion in the spatio-temporal dimension for the relational representation, capture the relationships and dynamic interactions between the relationships changing with time and space, and obtain spatio-temporal relational features

[0108] Among them, the temporal decoder is also based on the Transformer architecture. Through multiple layers of processing of the temporal decoder, a relational feature representation integrating spatio-temporal information is obtained, improving the model's ability to understand and predict the dynamic relationships between targets.

[0109] Step 3.3: Introduce inter-frame message tokens Through the feed-forward network, fuse the introduced inter-frame message tokens and the spatio-temporal relational features to generate spatio-temporal relational features integrating short-term dynamic information Specifically, it is implemented through the following formula:

[0110]

[0111] Among them, represents the dynamic weight parameter, which is used to control the influence degree of the inter-frame message tokens on the spatio-temporal relational features ;

[0112] Update the inter-frame message tokens Specifically, it is implemented through the following formula:

[0113]

[0114] Among them, represents the updated inter-frame message token, represents the inter-frame message token of the (t - 1)-th frame, and β is the message update intensity parameter. is the spatio-temporal relationship feature of the t-th frame, is the spatio-temporal relationship feature of the (t - 1)-th frame.

[0115] Step 3.4: Use the relationship classifier to classify the spatio-temporal relationship features that fuse short-term dynamic information, predict the relationships between entities in the video, and generate the dynamic scene graph dictionary G corresponding to each frame of the image. t Combined with Figure 2 for example, (a1) is the first frame of the image, and (a2) is the dynamic scene graph dictionary G of (a1). t including person - looking at - food - in front of - person, person - lifting - sandwich, person - looking at - cup, (b1) is the second frame of the image, and (b2) is the dynamic scene graph dictionary G of (b1). t including person - looking at - food - in front of - person, person - lifting - sandwich, person - drinking water - cup.

[0116] Step 4: Determine the action recognition result of the target video according to the dynamic scene graph dictionary G t and the feature map F t , where the action recognition result characterizes whether there is an abnormality in the actions in the target video.

[0117] Step 4.1: For the dynamic scene graph dictionary G corresponding to each frame of the image t , determine the entities m and n with relationships in the dynamic scene graph dictionary. In the dynamic scene graph dictionary, obtain the entity category of entity m and the entity category of entity n as well as the bounding box coordinates of entity m

[0118]

[0119] Among them, represents the relative position of entities m and n;

[0120] According to the relationship weight and the dynamic scene graph dictionary G t , calculate the weighted feature map Specifically, it is implemented through the following formula:

[0121]

[0122] Step 4.2: According to the weighted feature map and the feature map F t , calculate the frame feature F after temporal enhancement time ;

[0123] Step 4.2.1: For each frame of the image, fuse the weighted feature map and the feature map F t to obtain the fused feature of each frame of the image. Project the fused feature into a query vector Q, a key vector K, and a value vector V. Calculate the dot product of the query vector Q and the key vector K, normalize the dot product of the query vector Q and the key vector K to obtain a normalized vector, and perform a dot product of the normalized vector and the value vector to obtain an enhanced feature Furthermore, obtain the enhanced feature corresponding to each frame of the image in the target video All the enhanced features constitute the feature sequence I t , that is Figure 1 the spatial weighting in

[0124] Step 4.2.2: Input the feature sequence I t into the MLP-mixer layer for mixing, so that the frame feature can pay attention to the features of all frames in the video, and obtain the mixed frame feature;

[0125] Among them, the MLP-mixer layer contains two MLP sub-layers, namely the Token mixing sub-layer for mixing the frame features to enable interaction between the frame features, and the Channel mixing sub-layer for refining the mixed frame features to enhance the expression ability of the frame features.

[0126] Step 4.2.3: Refine the mixed frame feature through the MLP-mixer layer to obtain the frame feature after temporal enhancement

[0127] Step 4.3: Input the frame feature F after temporal enhancement time into the query classifier. Specifically, use the global average pooling layer to compress the frame feature F after temporal enhancement time into a vector to obtain the feature representation of the target video. Through the classifier Classifier and the softmax function, classify the feature representation of the target video to obtain the probability of each category, which is specifically implemented by the following formula:

[0128] y = Softmax(Classifier(GlobalPool(F time )));

[0129] Among all categories of probabilities, the category with the highest probability value is selected as the final behavior recognition result, that is, the behavior recognition result of the target video. The behavior recognition result indicates whether there is an abnormality in the behavior in the target video. For example, if the target video shows a person crossing the subway turnstile in a subway station, the behavior recognition result is that there is an abnormal behavior.

[0130] By using the fusion of cross-modal features, the present invention guides the generation of dynamic scene graphs, makes greater use of temporal information, strengthens the logical constraint of dynamic scene graphs in terms of relationship changes, abstracts scene semantics through dynamic scene graphs, and spatially weights video frame features using field semantics in behavior recognition, thereby reducing the interference of redundant information on behavior recognition and improving the speed and accuracy of behavior recognition.

[0131] Compared with the prior art, the present invention significantly enhances the accurate recognition of abnormal behaviors, effectively reduces the false positive and false negative rates, and improves the recognition reliability through deep scene semantic parsing and feature optimization techniques. By using the scene semantic filtering mechanism, it intelligently screens key information, greatly reduces the amount of redundant data processing, accelerates model operation, strengthens real-time performance, and ensures rapid response in complex scenarios. For complex environments such as rail transit stations, the present invention incorporates scene constraint characteristics, optimizes the model architecture, enhances its adaptability and stability, and enables it to maintain high accuracy and high efficiency in different scenarios. By upgrading the object detection algorithm and feature extraction technology, and combining the cross-modal feature fusion method, the present invention can deeply mine the key features in video data, realize multi-dimensional feature integration, and significantly improve the model's understanding depth of objects and their mutual relationships. By introducing innovative spatio-temporal relationship modeling techniques, such as the inter-frame message token mechanism, it accurately captures the spatio-temporal associations and dynamic changes between objects, and enhances the model's ability to analyze complex behavior patterns. At the same time, the present invention adopts dynamic scene graph parsing and MLP-mixer layer technology to achieve multi-angle and refined recognition of behaviors, further improving the accuracy and stability of behavior recognition. In practical applications, the present invention can effectively enhance the safety monitoring ability of crowded places such as rail transit stations, timely detect and warn of abnormal behaviors, create a safe travel environment for passengers, and has significant social benefits and practical value.

[0132] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but also covers other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A behavior recognition method based on scenario semantic understanding, characterized in that Including: Step 1: Obtain the target video in the target scenario, and then obtain the feature map F of each frame image in the target video t , as well as the entity features of each entity in each frame image where t represents the video frame number, t ∈ [1, T], T represents the total number of frames of the target video, i represents the entity index, i ∈ [1, N], and N represents the total number of entity indexes is the entity feature of the i-th entity in the t-th frame image is the bounding box coordinate of the i-th entity in the t-th frame image is the entity category of the i-th entity in the t-th frame image is the visual feature of the i-th entity in the t-th frame image, and all entity features form the entity feature set Step 2: Process all entity features of each frame image to obtain the relationship feature vector of entity i and entity j with existing relationships in the image and the enhanced image features; Step 3: For each frame of image, process the relational feature vector and the enhanced image feature to obtain a dynamic scene graph dictionary G corresponding to each frame of image t , where the dynamic scene graph dictionary includes the relationships between entities in the image, the relationship strengths between the relationships of entities, the bounding box coordinates of the entities constituting the relationships, and the entity categories of the entities constituting the relationships; Step 4: Determine the action recognition result of the target video according to the dynamic scene graph dictionary G t and the feature map F t , and determine the action recognition result of the target video, where the action recognition result represents whether there is an abnormality in the action in the target video.

2. The behavior recognition method based on scenario semantic understanding according to claim 1, characterized in that, Step 1 specifically includes: Step 1.1: Obtain a target video in a target scenario, where the target video includes T frame images, and the T frame images form a frame sequence {Fr t}, where Fr t represents the t-th frame image; Step 1.2: Input the frame sequence {Fr t} into the convolutional neural network of the pre-trained Faster R-CNN model to obtain the entity features of each entity in each frame image. All the entity features in the frame sequence {Fr t} form the entity feature set 3. The behavior recognition method based on scenario semantic understanding according to claim 2, characterized in that Step 1.2 specifically includes: Step 1.2.1: Input each frame image in the frame sequence {Fr t} into the convolutional neural network CNN to obtain the feature map F t of each frame image, which is specifically implemented through the following formula: F t = CNN(Fr t ); Among them, F t represents the feature map of the t-th frame image; Step 1.2.2: For the feature map F of each frame of image t , process the feature map through the Region Proposal Network (RPN) to obtain multiple candidate regions; Specifically, the feature map is processed by RPN to obtain the feature mapping of the feature map, which is specifically implemented by the following formula: U t = RPN(F t ); Among them, U t represents the feature mapping of F t ; Through the sliding window of RNP, slide on the feature mapping of the feature map to obtain multiple candidate regions; Step 1.2.3: For each candidate region, project the candidate region onto the feature map corresponding to the candidate region to obtain the feature matrix of the region of interest, and further obtain multiple feature matrices of the regions of interest; Step 1.2.4: For the feature matrix of each region of interest, through the ROIPooling layer, process the feature matrix of the region of interest to obtain the flattened feature vector, and further obtain multiple flattened feature vectors; Step 1.2.5: For each flattened feature vector, input the flattened feature vector into a series of fully connected layers to obtain a target vector corresponding to each candidate region. According to the target vector, determine the entity features of the entities in the candidate region, obtain the entity features of all entities in a frame of image, and further obtain the entity features of each frame of image in the frame sequence. All entity features form an entity feature set 4. The behavior recognition method based on scenario semantic understanding according to claim 1, wherein Step 2 specifically includes: Step 2.1: For all entity features of each frame of image, encode the entity features through a pre-trained image encoder to obtain entity i and entity j with existing relationships in the image, as well as joint image features Step 2.2: Combine the joint image features with the entity features of entity i and the entity features of entity j to obtain a relational feature vector Specifically, it is achieved through the following formula: Step 2.3: The entity category of entity i is and that of entity j are input into the pre-trained text encoder to extract the joint text feature Step 2.4: Through the jointly extracted text features Perform relationship-aware prompting on the jointly extracted image features to obtain optimized image features Specifically, it is implemented through the following formula: Among them, α represents the hint intensity parameter; Step 2.5: Based on the knowledge distillation loss function, by comparing the joint image features output by the image encoder and the joint text features output by the text encoder perform knowledge distillation to obtain the knowledge distillation loss The knowledge distillation loss function is represented by the following formula: Among them, ‖‖2 represents the Euclidean norm; Combine the knowledge distillation loss with the optimized image features to obtain the enhanced image features.

5. A behavior recognition method based on scenario semantic understanding according to claim 1, characterized in that Step 3 specifically includes: Step 3.1: For each frame of image, input the relational feature vector and the enhanced image features into the spatial encoder to obtain spatial relational features; Step 3.2: Input the relational feature vector and the spatial relationship feature into the temporal decoder to obtain the spatio-temporal relationship feature Step 3.3: Introduce the inter-frame message token Through the feed-forward network, the introduced inter-frame message token and the spatio-temporal relationship feature are fused to generate a spatio-temporal relationship feature that integrates short-term dynamic information Specifically, it is implemented through the following formula: Among them, represents the dynamic weight parameter; Step 3.4: Use the relationship classifier to classify the spatio-temporal relationship features that fuse short-term dynamic information to generate a dynamic scene graph dictionary G corresponding to each frame of the image t .

6. The method for behavior recognition based on scenario semantic understanding according to claim 5, characterized in that, Step 3.3 also includes: Update the inter-frame message token through the following formula: Among them, represents the updated inter-frame message token, represents the inter-frame message token of the (t - 1)-th frame, and β is the message update intensity parameter, is the spatio-temporal relationship feature of the t-th frame, is the spatio-temporal relationship feature of the (t - 1)-th frame.

7. A behavior recognition method based on scene semantic understanding according to claim 1, characterized in that, Step 4 specifically includes: Step 4.1: For the dynamic scene graph dictionary G corresponding to each frame of image t , determine the entities m and n with existing relationships in the dynamic scene graph dictionary. In the dynamic scene graph dictionary, obtain the entity category of entity m and the entity category of entity n as well as the bounding box coordinates of entity m and the bounding box coordinates of entity n Calculate the relationship weight Specifically, it is implemented through the following formula: Among them, represents the relative position between entity m and entity n; According to the relationship weight and the dynamic scene graph dictionary G t calculate the weighted feature map Specifically, it is implemented through the following formula: Step 4.2: According to the weighted feature map and the feature map F t , calculate the frame feature F time after temporal enhancement; Step 4.3: Input the time-enhanced frame feature F time into the query classifier. Specifically, use the global average pooling layer to compress the time-enhanced frame feature F time into a vector to obtain the feature representation of the target video. Through the classifier Classifier and the softmax function, classify the feature representation of the target video to obtain the probability of each category, which is specifically implemented through the following formula: y = Softmax(Classifier(GlobalPool(F time ))); Among the probabilities of all categories, select the category with the highest probability value as the final behavior recognition result, that is, the behavior recognition result of the target video, and the behavior recognition result indicates whether there is an abnormality in the behavior in the target video.

8. A behavior recognition method based on scene semantic understanding according to claim 7, characterized in that, Step 4.2 specifically includes: Step 4.2.1: For each frame of image, the weighted feature map and the feature map F t are fused to obtain the fused feature of each frame of image. The fused feature is projected into a query vector Q, a key vector K, and a value vector V. Calculate the dot product of the query vector Q and the key vector K, normalize the dot product of the query vector Q and the key vector K to obtain a normalized vector, and perform a dot product of the normalized vector and the value vector to obtain an enhanced feature Furthermore, the enhanced features corresponding to each frame of image in the target video are obtained All the enhanced features constitute a feature sequence I t ; Step 4.2.2: Input the feature sequence I t into the MLP-mixer layer for mixing to obtain the mixed frame features; Step 4.2.3: Refine the mixed frame features through the MLP-mixer layer to obtain the temporally enhanced frame features