Three-dimensional human body behavior recognition method and device, terminal and medium

By extracting and processing the spatiotemporal features of four-dimensional point cloud data in the state space model, the problem of temporal dependence in the prior art is solved, and more efficient and accurate behavior recognition is achieved.

CN120088870AActive Publication Date: 2025-06-03PEKING UNIV SHENZHEN GRADUATE SCHOOL

Patent Information

Application Number
CN202510588159.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-03
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the prior art, the space-time dependence of four-dimensional point cloud data is difficult to efficiently capture, resulting in low accuracy and efficiency of behavior recognition results.

Method used

By entering the four-dimensional point cloud data to be analyzed into the trained state space model, subsequences of different time scales are extracted, and they are sorted and spliced ​​into an ordered spatiotemporal sequence, the central point features are determined, and action recognition is performed in combination with low-order spatiotemporal features.

Benefits of technology

This method can effectively capture complex space-time dependencies, reduce the computational complexity, and improve the accuracy and efficiency of behavior recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088870A_ABST
    Figure CN120088870A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a three-dimensional human body behavior recognition method and device, a terminal and a medium, and the method comprises the steps: inputting four-dimensional point cloud data into a trained state space model, and extracting subsequences of different time scales in the four-dimensional point cloud data; sorting each frame of three-dimensional point cloud in each sub-sequence to obtain an ordered space sequence; splicing the ordered space sequence according to a time sequence to obtain a spliced ordered space-time sequence corresponding to each subsequence; determining a center point feature corresponding to each sub-sequence according to the spliced ordered space-time sequence; and obtaining low-order spatial-temporal characteristics, and obtaining an action recognition result based on the low-order spatial-temporal characteristics and the central point characteristics of each sub-sequence. According to the method, space and time information is taken into consideration, a complex space-time dependency relationship can be captured, and the calculation complexity is reduced through the trained state space model, so that the accuracy and efficiency of a behavior recognition result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a three-dimensional human behavior recognition method, device, terminal and medium. Background Art

[0002] Four-dimensional point cloud videos can capture both the dynamic geometric information in three-dimensional space and the motion features that change over time, and thus can be applied to tasks such as action recognition, human pose estimation, environmental modeling, and intelligent interaction. Compared with traditional RGB videos and depth images, four-dimensional point cloud videos have higher robustness under low-light or perspective change conditions, and are particularly suitable for human behavior analysis in complex and dynamic environments.

[0003] However, in the prior art, architectures with quadratic complexity are difficult to efficiently capture the spatio-temporal dependencies of four-dimensional point clouds, while the selective state space model with linear complexity is a unidirectional recursive structure, which limits the application effect in four-dimensional point clouds with spatio-temporal disorder, and further leads to low accuracy and efficiency of behavior recognition results.

[0004] Therefore, there are deficiencies in the prior art and it needs to be improved and developed. Summary of the Invention

[0005] The present application provides a three-dimensional human behavior recognition method, device, terminal and medium to solve the technical problem of low accuracy and efficiency of behavior recognition results in related technologies.

[0006] To achieve the above object, the present application adopts the following technical solutions: A three-dimensional human behavior recognition method, wherein the method includes: Input the four-dimensional point cloud data to be analyzed into a trained state space model, and extract subsequences with different time scales in the four-dimensional point cloud data; Sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; Concatenate the ordered spatial sequences corresponding to all frames of three-dimensional point cloud in each subsequence in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence; Determine the center point feature corresponding to each subsequence according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence; Obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain the action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence.

[0007] In an embodiment of the present application, determining the center point feature corresponding to each subsequence according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence includes: Construct a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences; Perform normalization processing and feature fusion on the point features within each of the spatio-temporal neighborhood graphs to obtain the center point features corresponding to each subsequence.

[0008] In one embodiment of the present application, constructing a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences includes: Use the K-nearest neighbor method and the spatio-temporal embedding method to construct a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences.

[0009] In one embodiment of the present application, obtaining an action recognition result based on the low-order spatio-temporal features and the center point features of each of the subsequences includes: Input the low-order spatio-temporal features into a pose encoder to obtain predicted skeleton key points; Input the skeleton key points into a pose decoder to obtain high-dimensional geometric features; Fuse the high-dimensional geometric features and the center point features of each of the subsequences to obtain an action recognition result.

[0010] In one embodiment of the present application, the training steps of the state space model include: Obtain a training data set, where the training data set includes: four-dimensional point cloud training data and corresponding action labels; Input the four-dimensional point cloud training data into an initial state space model to extract training subsequences with different time scales in the four-dimensional point cloud training data; Sort each frame of three-dimensional point cloud in each of the training subsequences to obtain ordered space sequence training data corresponding to each frame of three-dimensional point cloud; Concatenate the ordered space sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence in chronological order to obtain spliced ordered spatio-temporal sequence training data corresponding to each training subsequence; Determine the center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence; Obtain low-order spatio-temporal feature training data extracted from the four-dimensional point cloud training data, and perform training based on the low-order spatio-temporal feature training data, the center point feature training data of each of the training subsequences, and the action labels to obtain a trained state space model.

[0011] In one embodiment of the present application, determining the center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence includes: Construct a spatio-temporal neighborhood training graph for each center point on each piece of the spliced ordered spatio-temporal sequence training data; Perform normalization processing and feature fusion on the point features within each of the spatio-temporal neighborhood training graphs to obtain the center point feature training data corresponding to each training subsequence.

[0012] In an embodiment of the present application, the training dataset further includes: true skeleton key points corresponding to four-dimensional point cloud training data; Train based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action label to obtain a trained state space model, including: Input the low-order spatio-temporal feature training data into a pose encoder to obtain predicted skeleton key points; Calculate the mean squared error loss between the predicted skeleton key points and the true skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder; Input the predicted skeleton key points into a pose decoder to obtain high-dimensional geometric feature training data; Fuse the high-dimensional geometric feature training data and the center point feature training data of each training subsequence, and train with the action label as the ground truth to obtain a trained state space model.

[0013] The present application also provides a three-dimensional human behavior recognition device, wherein the device includes: An extraction module, configured to input the four-dimensional point cloud data to be analyzed into the trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; A sorting module, configured to sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; A splicing module, configured to splice the ordered spatial sequences corresponding to all frames of three-dimensional point cloud in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; A determination module, configured to determine the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence; An identification module, configured to obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence.

[0014] The present application also provides a terminal, which includes: a memory, a processor, and a three-dimensional human behavior recognition program stored on the memory and executable on the processor. When the three-dimensional human behavior recognition program is executed by the processor, the steps of the three-dimensional human behavior recognition method described above are implemented.

[0015] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the three-dimensional human behavior recognition method as described above.

[0016] Advantages of the present invention: The method according to an embodiment of the present invention inputs the four-dimensional point cloud data to be analyzed into a trained state space model, and extracts subsequences of different time scales in the four-dimensional point cloud data; sorts each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; splices the ordered spatial sequences corresponding to all frames of three-dimensional point cloud in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; determines a center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence; obtains low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtains an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence, taking into account both spatial and temporal information, capable of capturing complex spatio-temporal dependence relationships, and reducing the computational complexity through the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of a preferred embodiment of the three-dimensional human behavior recognition method in the present invention.

[0018] Figure 2 is a principle block diagram of the present invention from the input of four-dimensional point cloud data to the output of action recognition result.

[0019] Figure 3 is a test result of the state space model of the present invention on a test set.

[0020] Figure 4 is a schematic diagram of attention visualization for three-dimensional human behavior recognition in the present invention.

[0021] Figure 5 is a functional principle block diagram of a preferred embodiment of the three-dimensional human behavior recognition device in the present invention.

[0022] Figure 6 is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] The existing technologies have the following key drawbacks and limitations in four-dimensional point cloud video analysis: First, it is unable to effectively capture spatio-temporal dependencies. Existing methods based on convolutional neural networks (CNNs) cannot effectively handle long-term spatio-temporal dependencies in four-dimensional point clouds. Although convolutional neural networks can extract local geometric features, their ability in temporal sequence modeling is weak. Especially when dealing with long sequences, it is easy to lose temporal information.

[0025] Second, it consumes a large amount of computing resources. Transformer-based models, although able to capture long-term dependencies, have a sharp increase in their computational complexity and memory consumption as the length of the input sequence increases, which limits their application in high-dimensional and long-sequence point cloud data. This leads to higher hardware requirements and computing resources, thus restricting the efficiency in practical applications.

[0026] Third, the problem of spatio-temporal disorder. Existing state space models (SSMs), although able to effectively process spatial data, due to the lack of effective time series modeling, fail to fully utilize spatio-temporal correlations, resulting in poor performance in the spatio-temporal joint modeling task of four-dimensional point cloud videos.

[0027] Fourth, the lack of robustness against incomplete and noisy data. Existing four-dimensional point cloud analysis methods usually assume that the input data is complete and noise-free. However, in practical applications, due to hardware limitations, point cloud data is often incomplete or contains noise. The existing technologies have weak processing capabilities for such incomplete and noisy data, resulting in poor robustness of the model when facing sparse and noisy datasets.

[0028] In view of the above deficiencies of the existing technologies, this application solves the problems of the existing methods being unable to effectively capture spatio-temporal dependencies, high computational complexity, difficult spatio-temporal disorder modeling, and lack of robustness. Specifically, this application can effectively handle spatio-temporal dependencies in four-dimensional point cloud videos through joint spatio-temporal serialization and structured modeling, and greatly reduces the consumption of computing resources by using state space models, improving the efficiency and accuracy of the model in long-sequence modeling.

[0029] The 3D human behavior recognition method, device, terminal, and medium according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problem in the related art mentioned in the above background art that the spatial point cloud data and time information of the four-dimensional point cloud data are often disordered, resulting in a low accuracy of the behavior recognition result, the present application provides a 3D human behavior recognition method. In this method, the to-be-analyzed four-dimensional point cloud data is input into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; each frame of the three-dimensional point cloud in each of the subsequences is sorted to obtain an ordered spatial sequence corresponding to each frame of the three-dimensional point cloud; the ordered spatial sequences corresponding to all frames of the three-dimensional point cloud in each subsequence are concatenated in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence; the central point feature corresponding to each subsequence is determined according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence; the low-order spatio-temporal features extracted from the to-be-analyzed four-dimensional point cloud data are obtained, and an action recognition result is obtained based on the low-order spatio-temporal features and the central point features of each subsequence. By sorting each frame of the three-dimensional point cloud to obtain an ordered spatial sequence, the present application takes into account both spatial and time information, can capture complex spatio-temporal dependence relationships, and reduces the computational complexity through the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result.

[0030] Please refer to Figure 1 , the 3D human behavior recognition method described in the embodiments of the present invention includes the following steps: Step S100: Input the to-be-analyzed four-dimensional point cloud data into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data.

[0031] The four-dimensional point cloud data of the present application can be a four-dimensional point cloud video, and the trained state space model provided by the present application is used for the efficient analysis of four-dimensional point cloud data and human action recognition. In one embodiment, the state space model includes: a hierarchical ordered sequencer, a cross-temporal serialization module, a spatio-temporal structure aggregation layer, and a pose-aware feature optimization module. In the embodiments of the present application, the disordered four-dimensional point cloud is converted into an ordered sequence through cross-temporal serialization, and then the state space model is used to efficiently capture spatio-temporal dependence relationships. At the same time, the spatio-temporal structure aggregation layer and the hierarchical ordered sequencer are introduced to further optimize feature extraction and multi-scale spatio-temporal modeling. In addition, the pose-aware feature optimization module enhances the robustness of the model when processing sparse, incomplete, and high-noise data sets by introducing a pose estimation branch. Therefore, the present application can efficiently and accurately capture complex spatio-temporal dependence relationships in four-dimensional point cloud data while maintaining a linear computational complexity, not only improving the accuracy and efficiency of action recognition, but also being superior to traditional convolutional neural networks and Transformer architectures in terms of running time and memory usage.

[0032] As Figure 2As shown Figure 2 The four-dimensional point cloud data input therein includes the four-dimensional point cloud corresponding to the moment. First, the four-dimensional point cloud data is processed using 4D convolution to facilitate the processing of the four-dimensional point cloud data by the hierarchical ordered sequencer and the pose perception feature optimization module. Among them, represents subsequences of different time scales.

[0033] Specifically, the hierarchical ordered sequencer can balance the high-frequency and low-frequency temporal variations in the four-dimensional point cloud, and expand the receptive field of the model through multi-scale temporal downsampling.

[0034] Specifically, the four-dimensional point cloud data is expressed as: ; where represents the four-dimensional point cloud data, R represents the set of real numbers (Real numbers), and T×N×3 represents a feature with T time steps, N points, and 3 channels for each point. The low-order spatio-temporal features of the four-dimensional point cloud data are expressed as: . Where represents the low-order spatio-temporal features, and T×N×C represents a feature with T time steps, N points, and C channels for each point.

[0035] The hierarchical ordered sequencer adopts a downsampling strategy with an exponential step size to extract subsequences of different time scales, and the formula is as follows: ; ; where represents a slice vector for X, represents a slice vector for , tS is the subscript for traversing values. For example, if the total number of frames is 24 frames, S is the sampling level. If S = 2, then t ranges from 1 to 12, so tS is 2, 4, 6, 8,..., 24. n traverses from 1 to N, c traverses from 1 to the coordinate dimension number 3, and f traverses from 1 to the feature dimension number C. The step size is , representing different time scales.

[0036] In the embodiment of the present application, by using the hierarchical ordered sequencer for multi-scale downsampling, the capture of high-frequency details and low-frequency structures is balanced, and the receptive field of the model is expanded, making it perform better in processing long-term actions.

[0037] As Figure 1 shown, the three-dimensional human behavior recognition method further includes the following steps: Step S200, sorting each frame of the three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of the three-dimensional point cloud.

[0038] Step S300: Concatenate the ordered spatial sequences corresponding to the 3D point clouds of all frames in each subsequence in chronological order to obtain the concatenated ordered spatio-temporal sequence corresponding to each subsequence.

[0039] In the embodiment of the present application, a cross-temporal serialization module is also provided. The cross-temporal serialization module converts the unordered four-dimensional point cloud data into an ordered sequence to meet the one-way modeling requirements of the state space model (SSM). The Hilbert curve is used to sort each frame of the 3D point cloud in the subsequence, maintaining local continuity in space and reducing the distance difference between adjacent points in the sequence. Moreover, the point clouds of each frame are serialized in chronological order to ensure the coherence of temporal information.

[0040] As Figure 1 shown, the 3D human behavior recognition method further includes the following steps: Step S400: Determine the center point feature corresponding to each subsequence according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence.

[0041] In the embodiment of the present application, step S400 specifically includes: Step S410: Construct a spatio-temporal neighborhood graph for each center point on each of the concatenated ordered spatio-temporal sequences; Step S420: Perform normalization processing and feature fusion on the point features within each of the spatio-temporal neighborhood graphs to obtain the center point feature corresponding to each subsequence.

[0042] Specifically, concatenate the ordered spatial sequences of all frames in chronological order to form an overall spatio-temporally ordered sequence and , where , represents the coordinates of the center point, L represents the sequence length, represents the center point feature. Since the input limit of the state space model is three-dimensional, including batch size, length, and number of points, it is necessary to multiply T and N and combine them into one L dimension, and the L dimension contains all the points within T frames .

[0043] The present application is provided with a spatio-temporal structure aggregation layer, and the spatio-temporal structure aggregation layer is used to construct a spatio-temporal neighborhood graph for each center point on each of the concatenated ordered spatio-temporal sequences. Among them, there are multiple center points on the concatenated ordered spatio-temporal sequence, and the center points are obtained by farthest point sampling from the input points. After normalizing the point features within each of the spatio-temporal neighborhood graphs, feature fusion is performed through a multi-layer perceptron (MLP) to generate an updated center point feature. The specific formula is as follows: ; ; ; ; 。

[0044] Among them, represents the center point feature, that is, it refers to , represents the coordinates of the center point, that is, it refers to . KNN represents the K-Nearest Neighbors method. Figure 2 The coordinates of the center point and the points within the neighborhood (abbreviated as neighboring points) in are all expressed as (x, y, z, t). represents the difference between the neighboring point after spatio-temporal embedding and the center point on the spatial x-axis. represents the difference between the neighboring point after spatio-temporal embedding and the center point on the spatial y-axis. represents the difference between the neighboring point after spatio-temporal embedding and the center point on the spatial z-axis. represents the neighboring point feature. represents the intermediate result after normalizing the difference between the neighboring point feature and the center point feature. represents a very small constant, usually used to avoid division by zero errors or numerical stability. represents the updated neighboring point feature. represents the natural logarithm. represents the updated center point feature obtained after feature fusion. K represents the number of neighboring points obtained after the KNN algorithm. i represents traversing i times from the 1st to the Kth neighboring point, and j represents traversing j times from the 1st to the Kth neighboring point. represents a multi-layer perceptron. represents the feature of the i-th point traversed in represents the feature of the j-th point traversed in

[0045] The embodiments of the present application can extract and integrate the local spatio-temporal features of the point cloud, and update the center point feature by constructing a spatio-temporal neighborhood graph.

[0046] In one embodiment of the present application, the step S410 is specifically: constructing a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences by using the K-Nearest Neighbors method and the spatio-temporal embedding method.

[0047] Specifically, the embodiment of the present application uses an extended K-Nearest Neighbor (KNN) method and combines a spatio-temporal embedding method to construct a spatio-temporal neighborhood graph for each center point.

[0048] As Figure 1 shown, the three-dimensional human behavior recognition method further includes the following steps: Step S500: Obtain low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence.

[0049] In the embodiment of the present application, step S500 specifically includes: Step S510: Input the low-order spatio-temporal features into a pose encoder to obtain predicted skeleton key points; Step S520: Input the skeleton key points into a pose decoder to obtain high-dimensional geometric features; Step S530: Fuse the high-dimensional geometric features and the center point features of each subsequence to obtain an action recognition result.

[0050] Specifically, the embodiment of the present application obtains low-order spatio-temporal features through a shared point 4D convolution, inputs the low-order spatio-temporal features into a trained pose encoder to obtain predicted skeleton key points, then uses a pose decoder to extract high-dimensional geometric features, and after pooling, fuses them with the center point features of each subsequence to obtain an action recognition result.

[0051] The embodiment of the present application uses a pose-aware feature optimization module for auxiliary learning of pose estimation, which improves the model's perception ability of the human body's geometric structure and motion pattern and enhances the recognition accuracy.

[0052] In an embodiment of the present application, the training steps of the state space model include: Obtain a training data set, where the training data set includes: four-dimensional point cloud training data and corresponding action labels; Input the four-dimensional point cloud training data into an initial state space model to extract training subsequences with different time scales in the four-dimensional point cloud training data; Sort each frame of the three-dimensional point cloud in each training subsequence to obtain ordered spatial sequence training data corresponding to each frame of the three-dimensional point cloud; Concatenate the ordered spatial sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence in chronological order to obtain concatenated ordered spatio-temporal sequence training data corresponding to each training subsequence; Determine the center point feature training data corresponding to each training subsequence according to the concatenated ordered spatio-temporal sequence training data corresponding to each training subsequence; Obtain the low-order spatio-temporal feature training data extracted from the four-dimensional point cloud training data, and train based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action label to obtain a trained state space model.

[0053] In the embodiment of the present application, through cross-temporal serialization, the four-dimensional point cloud data is effectively serialized, taking into account both spatial and temporal information, enabling the state space model to capture complex spatio-temporal dependencies in a unidirectional modeling framework. This unified modeling improves the model's understanding and recognition ability of dynamic actions. Compared with traditional convolutional neural networks and Transformer architectures, the present application reduces the running time and memory usage, especially when processing long-sequence four-dimensional point clouds, and can improve the computational efficiency. The present application is applicable to resource-constrained application scenarios. Specifically, it is not only applicable to human action recognition, but also can be extended to multiple fields such as robot navigation, autonomous driving, and intelligent monitoring, improving the system's action recognition and response ability in complex environments.

[0054] The state space model provided by the present application effectively solves the problems of high computational complexity and difficulty in capturing spatio-temporal dependencies in four-dimensional point cloud data analysis. The state space model of the present application not only outperforms traditional methods in terms of computational efficiency and memory usage, but also improves the robustness and recognition accuracy in complex environments through the pose perception mechanism. As Figure 3 and Figure 4 shown, Figure 3 are the test results of the state space model of the present application on the test set, Figure 4 is the attention visualization schematic diagram of three-dimensional human behavior recognition in the present application.

[0055] In an embodiment of the present application, determining the center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence includes: Construct a spatio-temporal neighborhood training graph for each center point on each spliced ordered spatio-temporal sequence training data; Perform normalization processing and feature fusion on the point features in each spatio-temporal neighborhood training graph to obtain the center point feature training data corresponding to each training subsequence.

[0056] In the embodiment of the present application, by constructing a spatio-temporal neighborhood training graph, efficient aggregation of local features is realized. At the same time, the state space model is responsible for capturing long-range dependencies, ensuring that the model does not sacrifice the understanding of global spatio-temporal relationships while maintaining linear complexity, thereby realizing efficient local feature extraction and global dependency capture.

[0057] In an embodiment of the present application, the training data set further includes: the true skeleton key points corresponding to the four-dimensional point cloud training data; Training is performed based on the low-order spatiotemporal feature training data, the center point feature training data of each of the training subsequences, and the action label to obtain a trained state space model, including: Inputting the low-order spatiotemporal feature training data into a posture encoder to obtain predicted skeleton key points; Calculating the mean square error loss between the predicted skeleton key points and the real skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder; Inputting the predicted skeleton key points into a posture decoder to obtain high-dimensional geometric feature training data; The high-dimensional geometric feature training data and the center point feature training data of each training subsequence are fused, and training is performed using the action label as the true value to obtain a trained state space model.

[0058] Specifically, in the training phase, the embodiment of the present application uses a convolution-based gesture encoder to map the low-order spatiotemporal feature training data to , where kp is the number of skeleton key points, which is the information in the training dataset, such as the 20 key points in the MSR Action3D dataset. The mean square error (MSE) loss between the predicted skeleton key points and the true skeleton key points is calculated to guide model learning.

[0059] The embodiments of the present application enhance the model's ability to learn the human skeleton structure and motion pattern and improve its robustness on sparse, incomplete and high-noise data sets by introducing a posture estimation task.

[0060] The embodiment of the present application utilizes a posture perception feature optimization module to perform auxiliary learning for posture estimation, thereby improving the model's ability to perceive the geometric structure and motion patterns of the human body, and enhancing the robustness and recognition accuracy on sparse, incomplete, and noisy data sets.

[0061] In addition, in terms of hardware, to improve the computational efficiency of the model, this application can be implemented on a GPU (Graphics Processing Unit), leveraging its parallel computing capabilities to accelerate the operations of CTS (Cross-Temporal Serialization), STSAL (Spatio-Temporal Structure Aggregation Layer), and SSM (State Space Model). In terms of software, this application can be implemented in mainstream deep learning frameworks (such as PyTorch), using rich APIs (Application Programming Interfaces) and optimization tools to simplify the model development and training process. Optimization techniques such as Mixed Precision Training and Gradient Accumulation can also be adopted to improve the efficiency and stability of model training. In terms of data processing, this application can also perform filtering on four-dimensional point cloud data to remove noise points and outliers, improving the input quality and recognition accuracy of the model. In addition to pose estimation, other auxiliary tasks (such as point cloud classification and segmentation) can be introduced to further improve the feature expression ability of the model through multi-task learning. The model provided by this application also has compatibility and scalability, being able to be compatible with existing four-dimensional point cloud processing architectures and can be integrated into existing systems as a plug-in module to enhance their spatio-temporal modeling capabilities. In this way, this application can flexibly adapt to different application requirements and technical environments, achieving efficient and accurate four-dimensional point cloud data analysis and human action recognition.

[0062] In one embodiment, as Figure 5 shown, based on the above three-dimensional human behavior recognition method, the present invention also correspondingly provides a three-dimensional human behavior recognition device, including: An extraction module 100, configured to input the four-dimensional point cloud data to be analyzed into a trained state space model, and extract subsequences of different time scales in the four-dimensional point cloud data; A sorting module 200, configured to sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; A splicing module 300, configured to splice the ordered spatial sequences corresponding to all frames of three-dimensional point cloud in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; A determination module 400, configured to determine the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence; A recognition module 500, configured to obtain low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each of the subsequences.

[0063] It should be noted that the foregoing explanation of the embodiments of the three-dimensional human behavior recognition method also applies to the three-dimensional human behavior recognition device of this embodiment, and will not be elaborated here.

[0064] The present invention discloses a three-dimensional human behavior recognition device. By inputting the four-dimensional point cloud data to be analyzed into a trained state space model, subsequences of different time scales in the four-dimensional point cloud data are extracted; each frame of the three-dimensional point cloud in each of the subsequences is sorted to obtain an ordered spatial sequence corresponding to each frame of the three-dimensional point cloud; the ordered spatial sequences corresponding to all frames of the three-dimensional point cloud in each subsequence are concatenated in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence; the central point feature corresponding to each subsequence is determined according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence; the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed are obtained, and an action recognition result is obtained based on the low-order spatio-temporal features and the central point features of each subsequence. In this application, by sorting each frame of the three-dimensional point cloud to obtain an ordered spatial sequence, both spatial and temporal information is taken into account, complex spatio-temporal dependence relationships can be captured, and the computational complexity is reduced by the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result.

[0065] Figure 6 It is a schematic structural diagram of a terminal provided in an embodiment of this application. The terminal may include: A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0066] When the processor 502 executes the program, it implements the three-dimensional human behavior recognition method provided in the above embodiment.

[0067] Further, the terminal further includes: A communication interface 503 for communication between the memory 501 and the processor 502.

[0068] The memory 501 is used to store a computer program executable on the processor 502.

[0069] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0070] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected via a bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in illustration, only one line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0071] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can complete communication with each other through an internal interface.

[0072] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0073] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned three-dimensional human behavior recognition method is implemented.

[0074] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0075] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0076] Any process or method description represented in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present application pertain.

[0077] The logic and / or steps represented in a flowchart or described otherwise herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, since the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0078] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0079] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0080] In addition, in each embodiment of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0081] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A three-dimensional human behavior recognition method, characterized in that: The method comprises: Inputting the four-dimensional point cloud data to be analyzed into the trained state space model, and extracting subsequences of different time scales in the four-dimensional point cloud data; Sort each frame of the three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of the three-dimensional point cloud; The ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence are spliced ​​in time order to obtain a spliced ​​ordered spatiotemporal sequence corresponding to each subsequence; Determine the center point feature corresponding to each subsequence according to the spliced ​​ordered spatiotemporal sequence corresponding to each subsequence; Low-order spatiotemporal features extracted from the four-dimensional point cloud data to be analyzed are obtained, and action recognition results are obtained based on the low-order spatiotemporal features and the center point features of each of the subsequences.

2. The three-dimensional human behavior recognition method according to claim 1, characterized in that: The center point feature corresponding to each subsequence is determined according to the spliced ​​ordered spatiotemporal sequence corresponding to each subsequence, including: Constructing a spatiotemporal neighborhood graph of each center point on each of the spliced ​​ordered spatiotemporal sequences; The point features in each of the spatiotemporal neighborhood graphs are normalized and feature fused to obtain the center point features corresponding to each subsequence.

3. The three-dimensional human behavior recognition method according to claim 2, characterized in that: Constructing a spatiotemporal neighborhood graph of each center point on each of the spliced ​​ordered spatiotemporal sequences, including: The K nearest neighbor method and the spatiotemporal embedding method are used to construct a spatiotemporal neighborhood graph of each center point on each of the spliced ​​ordered spatiotemporal sequences.

4. The three-dimensional human behavior recognition method according to claim 1, characterized in that: Obtaining action recognition results based on the low-order spatiotemporal features and the center point features of each of the subsequences includes: Inputting the low-order spatiotemporal features into a posture encoder to obtain predicted skeleton key points; Inputting the skeleton key points into a posture decoder to obtain high-dimensional geometric features; The high-dimensional geometric features and the center point features of each of the subsequences are fused to obtain an action recognition result.

5. The three-dimensional human behavior recognition method according to claim 1, characterized in that: The training steps of the state space model include: Acquire a training data set, wherein the training data set includes: four-dimensional point cloud training data and corresponding action labels; Inputting four-dimensional point cloud training data into an initial state space model, and extracting training subsequences of different time scales from the four-dimensional point cloud training data; Sorting each frame of three-dimensional point cloud in each of the training subsequences to obtain ordered spatial sequence training data corresponding to each frame of three-dimensional point cloud; The ordered spatial sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence are spliced ​​in time order to obtain the spliced ​​ordered spatiotemporal sequence training data corresponding to each training subsequence; Determine the center point feature training data corresponding to each training subsequence according to the spliced ​​ordered spatiotemporal sequence training data corresponding to each training subsequence; Obtain low-order spatiotemporal feature training data extracted from the four-dimensional point cloud training data, perform training based on the low-order spatiotemporal feature training data, the center point feature training data of each of the training subsequences, and the action label, and obtain a trained state space model.

6. The three-dimensional human behavior recognition method according to claim 5, characterized in that: Determine the center point feature training data corresponding to each training subsequence according to the spliced ​​ordered spatiotemporal sequence training data corresponding to each training subsequence, including: Constructing a spatiotemporal neighborhood training graph for each center point on each of the spliced ​​ordered spatiotemporal sequence training data; The point features in each of the spatiotemporal neighborhood training graphs are normalized and feature fused to obtain center point feature training data corresponding to each training subsequence.

7. The three-dimensional human behavior recognition method according to claim 5, characterized in that: The training data set also includes: real skeleton key points corresponding to the four-dimensional point cloud training data; Training is performed based on the low-order spatiotemporal feature training data, the center point feature training data of each of the training subsequences, and the action label to obtain a trained state space model, including: Inputting the low-order spatiotemporal feature training data into a posture encoder to obtain predicted skeleton key points; Calculating the mean square error loss between the predicted skeleton key points and the real skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder; Inputting the predicted skeleton key points into a posture decoder to obtain high-dimensional geometric feature training data; The high-dimensional geometric feature training data and the center point feature training data of each training subsequence are fused, and training is performed using the action label as the true value to obtain a trained state space model.

8. A three-dimensional human behavior recognition device, characterized in that: The device comprises: An extraction module, used for inputting the four-dimensional point cloud data to be analyzed into the trained state space model, and extracting subsequences of different time scales in the four-dimensional point cloud data; A sorting module, used for sorting each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; A splicing module is used to splice the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in time order to obtain a spliced ​​ordered spatiotemporal sequence corresponding to each subsequence; A determination module, used to determine the center point feature corresponding to each subsequence according to the spliced ​​ordered spatiotemporal sequence corresponding to each subsequence; The recognition module is used to obtain low-order spatiotemporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain action recognition results based on the low-order spatiotemporal features and the center point features of each subsequence.

9. A terminal, characterized in that: include: A memory, a processor, and a three-dimensional human behavior recognition program stored in the memory and executable on the processor, wherein the three-dimensional human behavior recognition program, when executed by the processor, implements the steps of the three-dimensional human behavior recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the three-dimensional human behavior recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, device and equipment based on Transform model and storage medium

    CN116721207A

Cited By

  • Personnel identification and behavior detection method and device and security robot

    CN120388411A