Behavior recognition method and device, equipment and storage medium

By fusing bone sequence data and visual features, using graph neural networks and cross-attention mechanisms to build multimodal behavioral features, the problems of inaccurate behavior recognition and high consumption of computing resources in the prior art are solved, and more efficient and accurate behavioral recognition is achieved.

CN120148103APending Publication Date: 2025-06-13GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146179.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing behavior recognition methods are difficult to effectively and accurately realize behavior recognition, especially to effectively distinguish behaviors with the same motion pattern. The video-based method requires high network size and training time, and the convergence situation is not ideal.

Method used

By obtaining the bone sequence data and visual features in the video frame, using graph neural network and cross-attention mechanism, fuse motion and visual features, construct multimodal behavioral features and perform behavior recognition.

Benefits of technology

It improves the accuracy and effectiveness of behavior recognition, can effectively distinguish behaviors with the same motion pattern, and reduces the requirements for network size and training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148103A_ABST
    Figure CN120148103A_ABST
Patent Text Reader

Abstract

The invention discloses a behavior recognition method and device, equipment and a storage medium. The method comprises the following steps: acquiring a plurality of video frames in a to-be-processed video; acquiring skeleton sequence data of to-be-identified objects in the plurality of video frames by using a preset attitude estimation network; according to the skeleton sequence data and a preset articulation point prior topological connection structure corresponding to different behaviors, using a graph neural network to extract motion features of the to-be-identified object from the plurality of video frames; extracting visual features of the to-be-identified object and an article interacted with the to-be-identified object from the plurality of video frames; and fusing the motion feature and the visual feature into a multi-modal behavior feature, and performing behavior recognition on the to-be-recognized object according to the multi-modal behavior feature to obtain a behavior recognition result. According to the invention, the accuracy and effectiveness of behavior recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular, to a behavior recognition method, device, electronic device, and computer-readable storage medium. Background Art

[0002] As a basic problem in computer vision, behavior recognition has broad application prospects, such as intelligent monitoring, somatosensory games for human-computer interaction, video retrieval, etc., thus attracting extensive attention in the industry.

[0003] Existing behavior recognition methods usually recognize based on bone sequences or based on videos. Among them, the method of recognizing behavior based on bone sequences recognizes from the motion patterns of behaviors. However, it cannot effectively distinguish behaviors with the same motion pattern. For example, "drinking water" and "eating", these two behaviors have the same motion pattern. If more samples are added for learning, overfitting is likely to occur; the method of recognizing behavior based on videos requires visual feature modeling in both time and space simultaneously. Therefore, this method has high requirements for network scale and training time, and the convergence situation is not ideal. Therefore, existing behavior recognition methods are difficult to effectively and accurately achieve behavior recognition. Summary of the Invention

[0004] The present invention provides a behavior recognition method, device, equipment, and storage medium to solve the technical problem that existing behavior recognition methods are difficult to effectively and accurately achieve behavior recognition.

[0005] To solve the above technical problem, in the first aspect of an embodiment of the present invention, a behavior recognition method is provided, including:

[0006] Obtain a plurality of video frames in a video to be processed;

[0007] Use a preset pose estimation network to obtain the bone sequence data of the object to be recognized in the plurality of video frames;

[0008] According to the bone sequence data and the preset prior topological connection structure of joint points corresponding to different behaviors, use a graph neural network to extract the motion features of the object to be recognized from the plurality of video frames;

[0009] Extract the visual features of the object to be recognized and the item with which it interacts from the plurality of video frames;

[0010] Fuse the motion features and the visual features into multi-modal behavior features, and perform behavior recognition on the object to be recognized according to the multi-modal behavior features to obtain a behavior recognition result.

[0011] As a preferred solution, the method specifically obtains the prior topological connection structure of the joints corresponding to different behaviors through the following steps:

[0012] Construct a number of sentences for characterizing the correlation between joints and behaviors;

[0013] Use a pre-trained natural language network to encode the number of sentences to obtain a number of sentence embedding vectors; wherein, the number of sentence embedding vectors is V×N, V represents the number of joints, and N represents the number of behaviors;

[0014] Perform average pooling on the number of sentence embedding vectors along the dimension for characterizing the number of behaviors to obtain the representation vector corresponding to each joint;

[0015] Calculate the cosine similarity of the representation vector corresponding to each joint to obtain a similarity matrix for characterizing the prior topological connection structure of the joints corresponding to different behaviors.

[0016] As a preferred solution, using the graph neural network to extract the motion features of the object to be recognized from a number of the video frames according to the bone sequence data and the prior topological connection structure of the joints corresponding to preset different behaviors specifically includes:

[0017] According to the bone sequence data, obtain the position time series data of each joint in a number of the video frames, and construct a bone spatio-temporal graph according to the position time series data;

[0018] According to the position time series data of each joint, the similarity matrix, a preset number of initial joint adjacency matrices, and the convolutional layer weight parameters of the graph neural network, use the graph neural network to extract joint features from the bone spatio-temporal graph;

[0019] Perform average pooling on the extracted joint features in the time dimension to obtain the motion features of the object to be recognized.

[0020] As a preferred solution, extracting the visual features of the object to be recognized and the items with which it interacts from a number of the video frames specifically includes:

[0021] Use a region proposal network to generate the bounding box of the object to be recognized in the video frame and the candidate boxes of each item other than the object to be recognized;

[0022] Calculate the intersection over union (IoU) between each candidate box and the bounding box;

[0023] Determine a number of interaction regions according to the IoU and a preset IoU threshold;

[0024] Perform region pooling cutting on the interaction region according to the bounding box of the interaction region to obtain the visual features of the object to be recognized and the items it interacts with; wherein, the bounding box of the interaction region is composed of the bounding box of the object to be recognized within the interaction region and the candidate boxes of the items.

[0025] As a preferred solution, calculating the intersection over union (IoU) between each of the candidate boxes and the bounding box specifically includes:

[0026] Determine the union area of each candidate box and the bounding box according to each candidate box and the bounding box, and determine the minimum box area between the candidate box and the bounding box;

[0027] Obtain the intersection over union (IoU) between each candidate box and the bounding box according to the ratio between the union area and the minimum box area.

[0028] As a preferred solution, fusing the motion feature and the visual feature into a multi-modal behavior feature and performing behavior recognition on the object to be recognized according to the multi-modal behavior feature to obtain a behavior recognition result specifically includes:

[0029] Take the motion feature as the query vector, take the visual feature as the key vector and value vector, and fuse the motion feature and the visual feature into the multi-modal behavior feature based on the cross-attention mechanism;

[0030] Use a preset classification head to map the multi-modal behavior feature to a behavior probability distribution, and determine the behavior recognition result based on the behavior probability distribution.

[0031] As a preferred solution, before obtaining a plurality of video frames in the video to be processed, the method further includes:

[0032] Create an image cache queue, a preprocessing data cache queue, parallel camera data input threads, image preprocessing threads, and a behavior recognition main thread;

[0033] Wherein, the camera data input thread is used to input a plurality of the obtained video frames into the image cache queue;

[0034] The image preprocessing thread is used to obtain the video frame from the head of the image cache queue, and obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the obtained video frame; input the obtained skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item into the preprocessing data cache queue;

[0035] The behavior recognition main thread is used to obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the head of the preprocessed data cache queue, and execute the behavior recognition process to obtain the behavior recognition result.

[0036] In a second aspect of the embodiments of the present invention, a behavior recognition device is provided, including:

[0037] A video processing module, configured to obtain a plurality of video frames in a video to be processed;

[0038] A skeleton sequence data acquisition module, configured to use a preset pose estimation network to obtain the skeleton sequence data of the object to be recognized in each of the video frames;

[0039] A motion feature extraction module, configured to extract the motion features of the object to be recognized from a plurality of the video frames by using a graph neural network according to the skeleton sequence data and a preset prior topological connection structure of joint points corresponding to different behaviors;

[0040] A visual feature extraction module, configured to extract the visual features of the object to be recognized and the items with which it interacts from a plurality of the video frames;

[0041] A behavior recognition module, configured to fuse the motion features and the visual features into multimodal behavior features, and perform behavior recognition on the object to be recognized according to the multimodal behavior features to obtain a behavior recognition result.

[0042] In a third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the behavior recognition method according to any one of the first aspect is implemented.

[0043] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, where the computer-readable storage medium includes a stored computer program, and when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the behavior recognition method according to any one of the first aspect.

[0044] Compared with the prior art, the beneficial effects of the embodiments of the present invention are that by simultaneously fusing motion features and visual features to construct multimodal behavior features and perform behavior recognition, it is possible to use the visual features for characterizing the human interaction relationship to make up for the deficiency that the motion features extracted based on the skeleton sequence data cannot distinguish behaviors with the same motion pattern, thereby improving the accuracy and effectiveness of behavior recognition.

[0045] In addition, during the motion feature extraction process, the prior topological connection structure of the joint points corresponding to different behaviors is additionally introduced, which strengthens the effect of the graph neural network on motion feature extraction, thereby further improving the accuracy of behavior recognition. Description of the Drawings

[0046] Figure 1 is a schematic flowchart of the behavior recognition method in an embodiment of the present invention;

[0047] Figure 2 is a system architecture diagram of the behavior recognition process in an embodiment of the present invention;

[0048] Figure 3 is a system architecture diagram of multi-modal behavior feature fusion in an embodiment of the present invention;

[0049] Figure 4 is a schematic diagram of concurrent threads in an embodiment of the present invention;

[0050] Figure 5 is a schematic structural diagram of the behavior recognition device in an embodiment of the present invention. Detailed Embodiments

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0052] Please refer to Figure 1 , the first aspect of the embodiment of the present invention provides a behavior recognition method, including the following steps S1 to S5:

[0053] Step S1, obtaining a plurality of video frames in the video to be processed;

[0054] Step S2, using a preset pose estimation network to obtain the skeletal sequence data of the object to be recognized in the plurality of video frames;

[0055] Step S3, according to the skeletal sequence data and the preset prior topological connection structure of the joint points corresponding to different behaviors, using a graph neural network to extract the motion features of the object to be recognized from the plurality of video frames;

[0056] Step S4, extracting the visual features of the object to be recognized and the items it interacts with from the plurality of video frames;

[0057] Step S5: Fuse the motion feature and the visual feature into a multi-modal behavior feature, and perform behavior recognition on the object to be recognized according to the multi-modal behavior feature to obtain a behavior recognition result.

[0058] Specifically, as Figure 2 shown, it is the system architecture diagram of the behavior recognition process in the embodiment of the present invention. In order to effectively obtain the skeletal sequence data of the object to be recognized (person or animal), in this embodiment, several video frames in the video to be processed are first obtained, and then the position data of each joint point of the object to be recognized in each video frame is obtained through the pose estimation network, so as to form the skeletal sequence data of the object to be recognized in the time dimension. It can be understood that the skeletal sequence data symbolizes the position movement of each joint point of the object to be recognized within the time period of several video frames. Further, in this embodiment, the graph neural network is used to extract the motion feature of the object to be recognized from several video frames based on the obtained skeletal sequence data. Among them, since the natural joint point topology is not completely applicable to the nodes with close relationships in human behavior. For example, the behavior of "walking" requires the cooperation of hand joints and leg joints, and their cooperation correlation should be very strong, while the hand and foot are not directly connected in the natural joint point topology. And the joints that originally had a stronger connection relationship are not connected, which makes the graph neural network need to pass through other nodes to establish the relationship between the two, and at this time, losses will occur, which will cause smoothness problems, making the features of different behaviors indistinguishable. Therefore, in this embodiment, an additional joint point prior topological connection structure corresponding to different behaviors is introduced to assist the graph neural network in extracting action features.

[0059] Furthermore, since the features encoded by the network pre-trained through the image classification task tend to be more towards the appearance features of classification, and many types of behaviors cannot be recognized separately from the visual appearance of the action object, in this embodiment, the visual features of the object to be recognized and the items it interacts with are extracted from several video frames, so as to be able to characterize the visual relationship between the object to be recognized and the interactive items.

[0060] Furthermore, on the premise of not increasing a large amount of network scale, the number of parameters, and the long inference time, in this embodiment, the motion feature and the visual feature are fused into a multi-modal behavior feature and behavior recognition is performed, and finally a behavior recognition result is obtained.

[0061] The behavior recognition method provided by the embodiment of the present invention constructs a multi-modal behavior feature by simultaneously fusing the motion feature and the visual feature and performs behavior recognition, and can use the visual feature used to characterize the human interaction relationship to make up for the deficiency that the motion feature extracted based on the skeletal sequence data cannot distinguish behaviors with the same motion pattern, so as to improve the accuracy and effectiveness of behavior recognition.

[0062] In addition, during the motion feature extraction process, the prior topological connection structure of joint points corresponding to different behaviors is additionally introduced, which enhances the motion feature extraction effect of the graph neural network, thereby further improving the accuracy of behavior recognition.

[0063] As a preferred solution, the method specifically obtains the prior topological connection structure of joint points corresponding to different behaviors through the following steps:

[0064] Construct a number of sentences for characterizing the correlation between joint points and behaviors;

[0065] Use a pre-trained natural language network to encode the number of sentences to obtain a number of sentence embedding vectors; where the number of sentence embedding vectors is V×N, V represents the number of joint points, and N represents the number of behaviors;

[0066] Perform average pooling on the number of sentence embedding vectors along the dimension for characterizing the number of behaviors to obtain the representation vector corresponding to each joint point;

[0067] Calculate the cosine similarity of the representation vector corresponding to each joint point to obtain a similarity matrix for characterizing the prior topological connection structure of joint points corresponding to different behaviors.

[0068] Specifically, in this embodiment, a sentence template is first constructed, such as "{keypoint} function in {action}", and then V joint points and N behaviors are placed in the corresponding positions in the sentence template to form a number of sentences for characterizing the correlation between joint points and behaviors. Further, considering the characteristics of the pre-trained natural language network for encoding sentences, which has good properties semantically, making the similarity of the features of similar sentences higher and the similarity of the features of distant sentences lower. Therefore, in this embodiment, each sentence is input into the pre-trained natural language network for encoding to obtain V×N sentence embedding vectors with a dimension of D. Among them, each joint point corresponds to a vector of dimension (N, D). Perform average pooling on the number of sentence embedding vectors along the dimension for characterizing the number of behaviors to obtain the representation vector corresponding to each joint point with a dimension of D. Further, calculate the cosine similarity between the representation vectors corresponding to all joint points to obtain a similarity matrix of dimension (V, V), which contains the connection relationship between joint points related to behaviors.

[0069] As a preferred solution, the use of the graph neural network to extract the motion features of the object to be recognized from a number of video frames according to the bone sequence data and the preset prior topological connection structure of joint points corresponding to different behaviors specifically includes:

[0070] According to the bone sequence data, obtain the position time series data of each joint point in a plurality of the video frames, and construct a bone spatio-temporal graph according to the position time series data;

[0071] According to the position time series data of each joint point, the similarity matrix, a plurality of preset initial adjacency matrices of joint points, and the convolutional layer weight parameters of the graph neural network, use the graph neural network to extract joint point features from the bone spatio-temporal graph;

[0072] Perform average pooling on the extracted joint point features in the time dimension to obtain the motion features of the object to be recognized.

[0073] Specifically, in this embodiment, first based on the bone sequence data, the position time series data of each joint point is obtained. Since there is a natural connection relationship between joint points, the human skeleton can be abstracted as a graph, and combined with the time dimension to form a bone spatio-temporal graph. The value of each joint point is the two-dimensional position of the current joint point in the video frame. By encoding graphs with different attributes (different values of joint points) but the same structure, the motion features of the object to be recognized can be extracted. This embodiment encodes the bone spatio-temporal graph based on the graph neural network. The encoding of the graph neural network for the graph is based on the topological structure of the graph, that is, the connection relationship between joint points, for message passing. Thus, for the feature extraction of the current joint point, in addition to combining the value of this joint point, the values of the joint points connected to it are combined into the current joint point through the weights of the connection edges. Therefore, when using the graph neural network to encode the bone spatio-temporal graph, the connection relationship between joint points is very important, and related nodes should have a stronger connection relationship to assist the graph neural network in encoding. Based on the fact that the graph neural network has adaptively learned several initial adjacency matrices of joint points during the training process, in order to break through the problem of the smoothness of the consistent node feature regions caused by the need for the graph neural network to stack depths to capture the relationship between joint points with multi-hop connections, this embodiment incorporates the obtained similarity matrix into one of the initial adjacency matrices. Thus, the process of the graph neural network extracting joint point features from the bone spatio-temporal graph is shown in the following expression:

[0074]

[0075] Among them, X out represents the extracted joint point features, A i represents the i-th initial adjacency matrix, p represents the number of initial adjacency matrices of joint points, and the (p + 1)-th initial adjacency matrix is the similarity matrix, X in represents the position time series data of the current joint point, and W represents the convolutional layer weight parameters of the graph neural network.

[0076] It can be understood that the initial adjacency matrix A in the above expressioni That is the weight of the connection edges between each joint point. One row of elements represents the adjacent relationship between the current joint point and other joint points. The sum of the elements in one row is 1, and the non - existent connection edge is 0. This is a "hard connection". In this embodiment, as long as there is a weight, it represents a "soft connection" relationship, and there is no need to discretize it into connected or unconnected states.

[0077] As a preferred solution, extracting the visual features of the object to be recognized and the items it interacts with from the several video frames specifically includes:

[0078] Using a region proposal network to generate the bounding box of the object to be recognized in the video frame and the candidate boxes of each item other than the object to be recognized;

[0079] Calculating the intersection - over - union area between each of the candidate boxes and the bounding box;

[0080] Determining several interaction regions according to the intersection - over - union area and a preset intersection - over - union area threshold;

[0081] Performing region pooling cutting on the interaction regions according to the bounding boxes of the interaction regions to obtain the visual features of the object to be recognized and the items it interacts with; wherein, the bounding box of the interaction region is composed of the bounding box of the object to be recognized within the interaction region and the candidate box of the item.

[0082] Specifically, the embodiment of the present invention proposes to process visual features based on a region proposal network and region of interest pooling, adding visual features centered on the object to be recognized to the network to assist the network in modeling and reasoning the visual features of behavior recognition. First, a region proposal network generates the bounding box of the object to be recognized in the video frame and the candidate boxes of each item other than the object to be recognized, and calculates the intersection - over - union area between the candidate boxes and the bounding box to indicate the possibility of interaction between the object to be recognized and the item. After selecting the corresponding interacting items based on the intersection - over - union area threshold (such as set to 0.25), the interaction region is composed of the region within the bounding box of the object to be recognized and the region within the candidate box of the corresponding interacting item. It can be understood that when the intersection - over - union area between a certain candidate box and the bounding box is greater than the set intersection - over - union area threshold, it is determined that the item corresponding to the candidate box is an interacting item. Further, according to the bounding box of the interaction region, performing region pooling cutting on the interaction region can obtain the visual features of the object to be recognized and the items it interacts with, helping the network to exclude image noise, focus on the region of interest, and further assisting the network in reasoning about human behavior. Exemplarily, the features of each interaction region are linearly pooled to a size of (7,7) in the region pooling, and after flattening, the final visual features are obtained.

[0083] As a preferred solution, calculating the intersection over union (IoU) between each of the candidate boxes and the bounding box specifically includes:

[0084] Based on each of the candidate boxes and the bounding box, determine the union area of each of the candidate boxes and the bounding box, and determine the area of the smallest box between the candidate box and the bounding box;

[0085] Based on the ratio between the union area and the area of the smallest box, obtain the intersection over union (IoU) between each of the candidate boxes and the bounding box.

[0086] Specifically, for the common calculation of the intersection over union (IoU), in order to filter out similar candidate boxes, the denominator is the union area of two candidate boxes. After experimentation in this embodiment, it is found that using the union area of the candidate box and the bounding box as the numerator and the area of the smallest box between the candidate box and the bounding box as the denominator, the calculated intersection over union (IoU) can better reflect the interaction possibility between the object to be recognized and the item.

[0087] As a preferred solution, fusing the motion feature and the visual feature into a multi-modal behavior feature, and performing behavior recognition on the object to be recognized according to the multi-modal behavior feature to obtain a behavior recognition result specifically includes:

[0088] Use the motion feature as the query vector, use the visual feature as the key vector and the value vector, and fuse the motion feature and the visual feature into the multi-modal behavior feature based on the cross-attention mechanism;

[0089] Use a preset classification head to map the multi-modal behavior feature to a behavior probability distribution, and determine the behavior recognition result based on the behavior probability distribution.

[0090] Specifically, as Figure 3 shown, in this embodiment, the motion feature and the visual feature are first mapped through a 1×1 convolutional layer, and the dimensions are unified to C. Then, using the motion feature as the query vector (Query), the visual feature as the key vector (Key) and the value vector (Value), the motion feature and the visual feature are fused into a multi-modal behavior feature based on the cross-attention mechanism. The expression is as follows:

[0091]

[0092] Among them, X out represents the multi-modal behavior feature, Q represents the query vector, K represents the key vector, V represents the value vector, and C represents the dimension of the key vector. The dimension of the attention matrix is (1,M), and the dimension of the output multi-modal behavior feature is (1,C). Finally, use a preset classification head to map the multi-modal behavior feature to a behavior probability distribution, and the behavior with the highest probability is used as the behavior recognition result.

[0093] As a preferred solution, before obtaining several video frames in the video to be processed, the method further includes:

[0094] Create an image cache queue, a preprocessing data cache queue, parallel camera data input threads, image preprocessing threads, and a behavior recognition main thread;

[0095] Among them, the camera data input thread is used to input the obtained several video frames into the image cache queue;

[0096] The image preprocessing thread is used to obtain the video frames from the head of the image cache queue, and obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the obtained video frames; input the obtained skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item into the preprocessing data cache queue;

[0097] The behavior recognition main thread is used to obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the head of the preprocessing data cache queue, and execute the behavior recognition process to obtain the behavior recognition result.

[0098] Specifically, as Figure 4 shown, it is a schematic diagram of concurrent threads in an embodiment of the present invention. In order to execute the current behavior recognition process, the video frames need to go through preprocessing of object detection and pose estimation to obtain candidate boxes of items, bounding boxes of objects to be recognized, and skeleton point data as inputs. In this embodiment, the yolov8 network is used to implement object detection and pose estimation. This embodiment is mainly aimed at the fact that the input of the behavior recognition process is a time series. From the frame input of the camera to the detection of the behavior, 32 frames of images need to be waited for. At this time, if a serial execution system architecture is used, the system needs to wait for 32 frames of images, which takes about 1 second under a 30fps camera. It takes about 0.02 seconds and 0.3 seconds to execute the yolov8 algorithm and the behavior recognition process respectively. At this time, the latency of the system for behavior recognition is greatly increased, which is 1 + 0.02 + 0.3 = 1.32 seconds. In order to enhance the real-time performance of the system, the embodiment of the present invention uses a multi-threaded architecture to build an algorithm, namely a camera data input thread, an image preprocessing thread, and a behavior recognition main thread. The three threads execute concurrently and communicate using a queue with a fixed window size. This queue with a fixed window size plays a role in communication between threads on the one hand, and also serves as a cache to provide the behavior recognition process with 32-frame inputs with less pause.

[0099] Specifically, both the image cache queue and the preprocessed data cache queue are queues that satisfy the "first in, first out" structure. The image cache queue is used to store the video frames of the camera, and the preprocessed data cache queue is used to store the skeleton sequence data, the bounding boxes of the objects to be recognized, and the candidate boxes of each item. Since the speed of the camera is 30fps and the inference speed of the yolov8 network on the GPU reaches 70fps, which is faster than the input speed of the camera, the images can be processed in a timely manner. The behavior recognition main thread executes the behavior recognition process and fetches 32 frames from the head of the preprocessed data cache queue as input each time. Additionally, according to the calculation of the algorithm's time consumption, each time the behavior recognition process is executed, it takes 0.3 seconds. At this time, the camera inputs approximately 10 frames of images. Therefore, each time the behavior recognition process fetches data for execution, only 10 frames at the head of the queue are popped, and the remaining 22 frames are combined with the latest frame as historical frames and used as the next input. Since when the capacity of the preprocessed data cache queue is full, the old data at the head of the queue is discarded and the new data is placed at the end of the queue, the data obtained by the behavior recognition process each time is the latest 32 frames for calculation, and there will be no delay accumulation caused by processing historical data due to a slowdown in speed, which would otherwise result in an exponential increase in delay.

[0100] Please refer to Figure 5 , the second aspect of the embodiments of the present invention provides a behavior recognition device, including:

[0101] A video processing module 101, configured to obtain a plurality of video frames in a video to be processed;

[0102] A skeleton sequence data acquisition module 102, configured to use a preset pose estimation network to obtain the skeleton sequence data of the object to be recognized in each of the video frames;

[0103] A motion feature extraction module 103, configured to extract the motion features of the object to be recognized from a plurality of the video frames by using a graph neural network according to the skeleton sequence data and a preset prior topological connection structure of joint points corresponding to different behaviors;

[0104] A visual feature extraction module 104, configured to extract the visual features of the object to be recognized and the items with which it interacts from a plurality of the video frames;

[0105] A behavior recognition module 105, configured to fuse the motion features and the visual features into multimodal behavior features, and perform behavior recognition on the object to be recognized according to the multimodal behavior features to obtain a behavior recognition result.

[0106] As a preferred solution, the device further includes a prior topological connection structure acquisition module for joint points, configured to:

[0107] Construct a plurality of sentences for characterizing the correlation between joint points and behaviors;

[0108] Encode several of the sentences using a pre-trained natural language network to obtain several sentence embedding vectors; where the number of the sentence embedding vectors is V×N, V represents the number of joint points, and N represents the number of behaviors;

[0109] Perform average pooling on several of the sentence embedding vectors along the dimension representing the number of behaviors to obtain a representation vector corresponding to each of the joint points;

[0110] Calculate the cosine similarity of the representation vector corresponding to each of the joint points to obtain a similarity matrix for characterizing the prior topological connection structure of the joint points corresponding to different behaviors.

[0111] As a preferred solution, the motion feature extraction module 103 is configured to extract the motion features of the object to be recognized from several of the video frames using a graph neural network according to the skeletal sequence data and the prior topological connection structure of the joint points corresponding to preset different behaviors, specifically including:

[0112] According to the skeletal sequence data, obtain the position time series data of each of the joint points in several of the video frames, and construct a skeletal spatio-temporal graph according to the position time series data;

[0113] According to the position time series data of each of the joint points, the similarity matrix, a preset plurality of joint point initialization adjacency matrices, and the convolutional layer weight parameters of the graph neural network, use the graph neural network to perform joint point feature extraction on the skeletal spatio-temporal graph;

[0114] Perform average pooling on the extracted joint point features in the time dimension to obtain the motion features of the object to be recognized.

[0115] As a preferred solution, the visual feature extraction module 104 is configured to extract the visual features of the object to be recognized and the items with which it interacts from several of the video frames, specifically including:

[0116] Use a region proposal network to generate a bounding box of the object to be recognized in the video frame and candidate boxes of each item other than the object to be recognized;

[0117] Calculate the intersection over union of the areas between each of the candidate boxes and the bounding box;

[0118] Determine several interaction regions according to the intersection over union and a preset intersection over union threshold;

[0119] Perform region pooling cutting on the interaction area according to the bounding box of the interaction area to obtain the visual features of the object to be recognized and the items it interacts with; wherein, the bounding box of the interaction area is composed of the bounding box of the object to be recognized within the interaction area and the candidate boxes of the items.

[0120] As a preferred solution, the visual feature extraction module 104 is used to calculate the intersection over union (IoU) between each candidate box and the bounding box, specifically including:

[0121] Determine the union area of each candidate box and the bounding box according to each candidate box and the bounding box, and determine the minimum box area between the candidate box and the bounding box;

[0122] Obtain the intersection over union between each candidate box and the bounding box according to the ratio between the union area and the minimum box area.

[0123] As a preferred solution, the behavior recognition module 105 is used to fuse the motion features and the visual features into multi-modal behavior features, and perform behavior recognition on the object to be recognized according to the multi-modal behavior features to obtain a behavior recognition result, specifically including:

[0124] Use the motion features as query vectors, use the visual features as key vectors and value vectors, and fuse the motion features and the visual features into the multi-modal behavior features based on the cross-attention mechanism;

[0125] Use a preset classification head to map the multi-modal behavior features to a behavior probability distribution, and determine the behavior recognition result based on the behavior probability distribution.

[0126] As a preferred solution, the device further includes a parallel thread and cache queue construction module for:

[0127] Create an image cache queue, a preprocessing data cache queue, parallel camera data input threads, image preprocessing threads, and a behavior recognition main thread;

[0128] Among them, the camera data input thread is used to input a plurality of the video frames obtained into the image cache queue;

[0129] The image preprocessing thread is used to obtain the video frames from the head of the image cache queue, and obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the obtained video frames; input the obtained skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item into the preprocessing data cache queue;

[0130] The behavior recognition main thread is used to obtain the skeleton sequence data, the bounding box of the object to be recognized, and the candidate boxes of each item from the head of the preprocessing data cache queue, and execute the process of behavior recognition to obtain the behavior recognition result.

[0131] The behavior recognition device provided by the embodiment of the present invention constructs multi-modal behavior features by simultaneously fusing motion features and visual features for behavior recognition, and can use visual features for characterizing the interaction relationship between people to make up for the deficiency that the motion features extracted based on the skeleton sequence data cannot distinguish behaviors with the same motion pattern, so as to improve the accuracy and effectiveness of behavior recognition.

[0132] In addition, a prior topological connection structure of joint points corresponding to different behaviors is additionally introduced in the process of motion feature extraction, which strengthens the extraction effect of the graph neural network on motion features, thereby further improving the accuracy of behavior recognition.

[0133] A third aspect of the embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the behavior recognition method described in any embodiment of the first aspect is implemented.

[0134] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.

[0135] The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device, and does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, buses, etc.

[0136] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and circuits.

[0137] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the electronic device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0138] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the behavior recognition method described in any embodiment of the first aspect.

[0139] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0140] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A behavior recognition method, characterized in that: include: Obtain several video frames in the video to be processed; Using a preset posture estimation network to obtain skeleton sequence data of the object to be identified in a plurality of video frames; According to the skeleton sequence data and the prior topological connection structure of the joint points corresponding to the preset different behaviors, the motion features of the object to be identified are extracted from the plurality of video frames using a graph neural network; Extracting visual features of the object to be identified and the item it interacts with from the plurality of video frames; The motion feature and the visual feature are fused into a multimodal behavior feature, and behavior recognition is performed on the object to be recognized based on the multimodal behavior feature to obtain a behavior recognition result.

2. The behavior recognition method according to claim 1, characterized in that: The method specifically obtains the priori topological connection structure of joint points corresponding to different behaviors through the following steps: Construct several sentences to represent the relationship between joint points and behaviors; Encoding a number of the sentences using a pre-trained natural language network to obtain a number of sentence embedding vectors; wherein the number of the sentence embedding vectors is V×N, where V represents the number of joint points and N represents the number of behaviors; Average pooling is performed on a plurality of the sentence embedding vectors along the dimension used to represent the number of behaviors to obtain a representation vector corresponding to each of the joint points; The cosine similarity calculation is performed on the characterization vector corresponding to each of the joint points to obtain a similarity matrix of the prior topological connection structure of the joint points corresponding to different behaviors.

3. The behavior recognition method according to claim 2, characterized in that: The method of extracting the motion features of the object to be identified from a plurality of video frames using a graph neural network according to the skeleton sequence data and the prior topological connection structure of the joint points corresponding to the preset different behaviors specifically includes: According to the skeleton sequence data, obtaining the position time series data of each joint point in a plurality of the video frames, and constructing a skeleton spatiotemporal graph according to the position time series data; According to the position time series data of each joint point, the similarity matrix, several preset joint point initialization adjacency matrices and the convolution layer weight parameters of the graph neural network, the graph neural network is used to extract joint point features from the skeleton spatiotemporal graph; The extracted joint point features are average pooled in the time dimension to obtain the motion features of the object to be identified.

4. The behavior recognition method according to claim 1, characterized in that: The extracting visual features of the object to be identified and the item with which it interacts from the plurality of video frames specifically includes: Generate a bounding box of the object to be identified in the video frame and candidate boxes of various items other than the object to be identified using a region candidate network; Calculating the area intersection-union ratio between each of the candidate boxes and the bounding box; Determining a plurality of interaction areas according to the area intersection-and-union ratio and a preset area intersection-and-union ratio threshold; According to the bounding box of the interaction area, the interaction area is subjected to regional pooling cutting to obtain visual features of the object to be identified and the item with which it interacts; wherein the bounding box of the interaction area is composed of a bounding box of the object to be identified and a candidate box of the item in the interaction area.

5. The behavior recognition method according to claim 4, characterized in that: The calculating the area intersection-and-union ratio between each of the candidate boxes and the bounding box specifically includes: Determine, according to each of the candidate boxes and the bounding box, a union area of ​​each of the candidate boxes and the bounding box, and determine a minimum box area between the candidate boxes and the bounding box; According to the ratio between the union area and the minimum box area, the area intersection-union ratio between each of the candidate boxes and the bounding box is obtained.

6. The behavior recognition method according to claim 1, characterized in that: The step of fusing the motion feature and the visual feature into a multimodal behavior feature, and performing behavior recognition on the object to be recognized according to the multimodal behavior feature to obtain a behavior recognition result specifically includes: Taking the motion feature as a query vector and the visual feature as a key vector and a value vector, the motion feature and the visual feature are fused into the multimodal behavior feature based on a cross attention mechanism; The multimodal behavior features are mapped into behavior probability distribution using a preset classification head, and the behavior recognition result is determined based on the behavior probability distribution.

7. The behavior recognition method according to any one of claims 4, characterized in that: Before obtaining a plurality of video frames in the video to be processed, the method further includes: Create image cache queue, pre-processing data cache queue, parallel camera data input thread, image pre-processing thread and behavior recognition main thread; Wherein, the camera data input thread is used to input the acquired several video frames into the image cache queue; The image preprocessing thread is used to obtain the video frame from the head of the image cache queue, and obtain the skeleton sequence data, the bounding box of the object to be identified, and the candidate box of each of the items from the obtained video frame; input the obtained skeleton sequence data, the bounding box of the object to be identified, and the candidate box of each of the items into the preprocessing data cache queue; The behavior recognition main thread is used to obtain the skeleton sequence data, the bounding box of the object to be recognized and the candidate box of each of the items from the head of the pre-processing data cache queue, and execute the behavior recognition process to obtain the behavior recognition result.

8. A behavior recognition device, characterized in that: include: A video processing module, used for obtaining a number of video frames in a video to be processed; A skeleton sequence data acquisition module, used to acquire skeleton sequence data of the object to be identified in each of the video frames using a preset posture estimation network; A motion feature extraction module, used to extract the motion features of the object to be identified from a plurality of the video frames using a graph neural network according to the skeleton sequence data and a priori topological connection structure of joint points corresponding to different preset behaviors; A visual feature extraction module, used to extract visual features of the object to be identified and the item it interacts with from the plurality of video frames; The behavior recognition module is used to fuse the motion feature and the visual feature into a multimodal behavior feature, and perform behavior recognition on the object to be recognized based on the multimodal behavior feature to obtain a behavior recognition result.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the behavior recognition method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the behavior recognition method according to any one of claims 1 to 7.