A 3D human behavior recognition method, device, terminal and medium

By processing four-dimensional point cloud data in the state space model, extracting and sorting subsequences, and building a spatiotemporal neighborhood map, the accuracy and efficiency of four-dimensional point cloud videos in behavior recognition are solved, and efficient action recognition and robustness are achieved.

CN120088870BActive Publication Date: 2025-07-11PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510588159.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-11
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

In the prior art, four-dimensional point cloud videos have low accuracy and efficiency in behavior recognition, difficult to effectively capture space-time dependence, high computational complexity, difficult to model space-time disordered modeling and lack robustness.

Method used

By inputting the four-dimensional point cloud data into the trained state space model, subsequences of different time scales are extracted, sorted and spliced into an ordered spatiotemporal sequence, space-time neighborhood map is constructed, low-order spatiotemporal and central point features are obtained, and action recognition is combined with the posture encoder and decoder.

Benefits of technology

It improves the accuracy and efficiency of behavior recognition, reduces the computational complexity, enhances the robustness of the model under sparse and noisy data, and is suitable for resource-constrained application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088870B_ABST
    Figure CN120088870B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and particularly to a three-dimensional human behavior recognition method, device, terminal and medium. The method includes inputting four-dimensional point cloud data into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; sorting each frame of three-dimensional point cloud in each subsequence to obtain an ordered space sequence; splicing the ordered space sequences in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; determining the central point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence; obtaining low-order spatio-temporal features, and obtaining an action recognition result based on the low-order spatio-temporal features and the central point features of each subsequence. This application takes into account both spatial and temporal information, can capture complex spatio-temporal dependence relationships, and reduces the computational complexity through the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a three-dimensional human behavior recognition method, device, terminal and medium. Background Technique

[0002] Four-dimensional point cloud videos can simultaneously capture the dynamic geometric information in three-dimensional space and the motion characteristics that change over time, and thus can be applied to tasks such as action recognition, human pose estimation, environmental modeling, and intelligent interaction. Compared with traditional RGB videos and depth images, four-dimensional point cloud videos have higher robustness under low-light or perspective change conditions, and are particularly suitable for human behavior analysis in complex and dynamic environments.

[0003] However, in the prior art, architectures with quadratic complexity are difficult to efficiently capture the spatio-temporal dependencies of four-dimensional point clouds, while the selective state space model with linear complexity is a unidirectional recursive structure, which limits the application effect in four-dimensional point clouds with spatio-temporal disorder, and further leads to low accuracy and efficiency of behavior recognition results.

[0004] Therefore, there are defects in the prior art and it needs to be improved and developed. Summary of the Invention

[0005] The present application provides a three-dimensional human behavior recognition method, device, terminal and medium to solve the technical problem of low accuracy and efficiency of behavior recognition results in related technologies.

[0006] To achieve the above object, the present application adopts the following technical solutions:

[0007] A three-dimensional human behavior recognition method, wherein the method includes:

[0008] Input the four-dimensional point cloud data to be analyzed into a trained state space model, and extract subsequences with different time scales in the four-dimensional point cloud data;

[0009] Sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud;

[0010] Concatenate the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence;

[0011] Determine the center point feature corresponding to each subsequence according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence;

[0012] Obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain the action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence.

[0013] In one embodiment of the present application, determining the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence includes:

[0014] Constructing a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences;

[0015] Performing normalization processing and feature fusion on the point features within each of the spatio-temporal neighborhood graphs to obtain the center point feature corresponding to each subsequence.

[0016] In one embodiment of the present application, constructing a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences includes:

[0017] Using the K-nearest neighbor method and the spatio-temporal embedding method to construct a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences.

[0018] In one embodiment of the present application, obtaining an action recognition result based on the low-order spatio-temporal feature and the center point features of each of the subsequences includes:

[0019] Inputting the low-order spatio-temporal feature into a pose encoder to obtain predicted skeleton key points;

[0020] Inputting the skeleton key points into a pose decoder to obtain high-dimensional geometric features;

[0021] Fusing the high-dimensional geometric features and the center point features of each of the subsequences to obtain an action recognition result.

[0022] In one embodiment of the present application, the training steps of the state space model include:

[0023] Obtaining a training data set, where the training data set includes: four-dimensional point cloud training data and corresponding action labels;

[0024] Inputting the four-dimensional point cloud training data into an initial state space model to extract training subsequences with different time scales in the four-dimensional point cloud training data;

[0025] Sorting each frame of three-dimensional point cloud in each of the training subsequences to obtain ordered space sequence training data corresponding to each frame of three-dimensional point cloud;

[0026] Splicing the ordered space sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence in chronological order to obtain spliced ordered spatio-temporal sequence training data corresponding to each training subsequence;

[0027] Determining center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence;

[0028] Obtain the low-order spatio-temporal feature training data extracted from the four-dimensional point cloud training data, and train based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action label to obtain a trained state space model.

[0029] In one embodiment of the present application, determining the center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence includes:

[0030] Construct a spatio-temporal neighborhood training graph for each center point on each spliced ordered spatio-temporal sequence training data;

[0031] Perform normalization processing and feature fusion on the point features within each spatio-temporal neighborhood training graph to obtain the center point feature training data corresponding to each training subsequence.

[0032] In one embodiment of the present application, the training data set further includes: the true skeleton key points corresponding to the four-dimensional point cloud training data;

[0033] Training based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action label to obtain a trained state space model includes:

[0034] Input the low-order spatio-temporal feature training data into a pose encoder to obtain predicted skeleton key points;

[0035] Calculate the mean square error loss between the predicted skeleton key points and the true skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder;

[0036] Input the predicted skeleton key points into a pose decoder to obtain high-dimensional geometric feature training data;

[0037] Fuse the high-dimensional geometric feature training data and the center point feature training data of each training subsequence, and train with the action label as the true value to obtain a trained state space model.

[0038] The present application also provides a three-dimensional human behavior recognition device, wherein the device includes:

[0039] An extraction module, configured to input the four-dimensional point cloud data to be analyzed into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data;

[0040] A sorting module, configured to sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered space sequence corresponding to each frame of three-dimensional point cloud;

[0041] A splicing module, configured to splice the ordered space sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence;

[0042] A determination module, configured to determine the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence;

[0043] An identification module, configured to obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each of the subsequences.

[0044] The present application further provides a terminal, which includes: a memory, a processor, and a three-dimensional human behavior recognition program stored on the memory and executable on the processor. When the three-dimensional human behavior recognition program is executed by the processor, the steps of the three-dimensional human behavior recognition method described above are implemented.

[0045] The present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the three-dimensional human behavior recognition method described above.

[0046] Advantages of the present invention: The method of the embodiment of the present invention inputs the four-dimensional point cloud data to be analyzed into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; sorts each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered space sequence corresponding to each frame of three-dimensional point cloud; splices the ordered space sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; determines the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence; obtains the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtains an action recognition result based on the low-order spatio-temporal features and the center point features of each of the subsequences. It takes into account both spatial and temporal information, can capture complex spatio-temporal dependence relationships, and reduces the computational complexity through the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result. Description of the Drawings

[0047] Figure 1 is a flowchart of a preferred embodiment of the three-dimensional human behavior recognition method in the present invention.

[0048] Figure 2 is a principle block diagram from the input of the four-dimensional point cloud data to the output of the action recognition result in the present invention.

[0049] Figure 3It is the test result of the state space model of the present invention on the test set.

[0050] Figure 4 It is the attention visualization schematic diagram of three-dimensional human behavior recognition in the present invention.

[0051] Figure 5 It is the functional principle block diagram of the preferred embodiment of the three-dimensional human behavior recognition device in the present invention.

[0052] Figure 6 It is the functional principle block diagram of the preferred embodiment of the terminal in the present invention. Detailed implementation manners

[0053] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] The prior art has the following key drawbacks and limitations in four-dimensional point cloud video analysis:

[0055] First, it is unable to effectively capture spatio-temporal dependencies. Existing methods based on convolutional neural networks (CNNs) are unable to effectively handle long-term spatio-temporal dependencies in four-dimensional point clouds. Although convolutional neural networks can extract local geometric features, their ability in temporal sequence modeling is weak, especially when dealing with long sequences, and it is easy to lose temporal information.

[0056] Second, it consumes a large amount of computing resources. Transformer-based models, although capable of capturing long-term dependencies, have a sharp increase in their computational complexity and memory consumption as the length of the input sequence increases, which limits their application in high-dimensional and long-sequence point cloud data. This results in higher hardware requirements and computing resources, thus limiting the efficiency in practical applications.

[0057] Third, the spatio-temporal disorder problem. Existing state space models (SSMs) can effectively process spatial data, but due to the lack of effective time series modeling, they fail to fully utilize spatio-temporal correlations, resulting in poor performance in the spatio-temporal joint modeling task of four-dimensional point cloud videos.

[0058] Fourth, the lack of robustness against incomplete and noisy data. Existing four-dimensional point cloud analysis methods usually assume that the input data is complete and noise-free. However, in practical applications, due to hardware limitations, point cloud data is often incomplete or contains noise. The prior art has weak processing capabilities for such incomplete and noisy data, resulting in poor robustness of the model when facing sparse and noisy datasets.

[0059] In view of the deficiencies of the above prior art, the present application solves the problems that existing methods cannot effectively capture spatio-temporal dependencies, have high computational complexity, are difficult to model spatio-temporal disorder, and lack robustness. Specifically, the present application can effectively process spatio-temporal dependencies in four-dimensional point cloud videos through joint spatio-temporal serialization and structured modeling, and greatly reduces the consumption of computing resources by using a state space model, improving the efficiency and accuracy of the model in long sequence modeling.

[0060] The three-dimensional human behavior recognition method, device, terminal and medium of the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problem in the related art that the spatial point cloud data and time information of four-dimensional point cloud data are often disordered, resulting in low accuracy of behavior recognition results, the present application provides a three-dimensional human behavior recognition method. In this method, the four-dimensional point cloud data to be analyzed is input into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; each frame of three-dimensional point cloud in each of the subsequences is sorted to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence are concatenated in time order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence; the central point feature corresponding to each subsequence is determined according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence; the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed are obtained, and an action recognition result is obtained based on the low-order spatio-temporal features and the central point features of each subsequence. By sorting each frame of three-dimensional point cloud to obtain an ordered spatial sequence, the present application takes into account both spatial and time information, can capture complex spatio-temporal dependency relationships, and reduces the computational complexity through a trained state space model, thereby improving the accuracy and efficiency of behavior recognition results.

[0061] Please refer to Figure 1 , the three-dimensional human behavior recognition method described in the embodiments of the present invention includes the following steps:

[0062] Step S100: Input the four-dimensional point cloud data to be analyzed into a trained state space model, and extract subsequences of different time scales in the four-dimensional point cloud data.

[0063] The four-dimensional point cloud data of this application can be a four-dimensional point cloud video, and the trained state space model provided by this application is used for the efficient analysis of four-dimensional point cloud data and human action recognition. In one embodiment, the state space model includes: a hierarchical ordered sequencer, a cross-temporal serialization module, a spatio-temporal structure aggregation layer, and a pose-aware feature optimization module. In the embodiment of this application, the disordered four-dimensional point cloud is converted into an ordered sequence through cross-temporal serialization, and then the state space model is used to efficiently capture spatio-temporal dependencies. At the same time, the spatio-temporal structure aggregation layer and the hierarchical ordered sequencer are introduced to further optimize feature extraction and multi-scale spatio-temporal modeling. In addition, the pose-aware feature optimization module enhances the robustness of the model when processing sparse, incomplete, and highly noisy data sets by introducing a pose estimation branch. Therefore, this application can efficiently and accurately capture complex spatio-temporal dependencies in four-dimensional point cloud data while maintaining a linear computational complexity, not only improving the accuracy and efficiency of action recognition, but also outperforming traditional convolutional neural networks and Transformer architectures in terms of running time and memory usage.

[0064] As Figure 2 shown, Figure 2 the input four-dimensional point cloud data includes the four-dimensional point cloud corresponding to the moment. First, the four-dimensional point cloud data is processed using 4D point convolution to facilitate the processing of the four-dimensional point cloud data by the hierarchical ordered sequencer and the pose-aware feature optimization module. Among them, represents subsequences of different time scales.

[0065] Specifically, the hierarchical ordered sequencer can balance the high-frequency and low-frequency temporal variations in the four-dimensional point cloud and expand the receptive field of the model through multi-scale temporal downsampling.

[0066] Specifically, the four-dimensional point cloud data is represented as: ; where represents the four-dimensional point cloud data, R represents the set of real numbers (Real numbers), and T×N×3 represents a feature with T time steps, N points, and 3 channels for each point. The low-order spatio-temporal features of the four-dimensional point cloud data are represented as: . Where represents the low-order spatio-temporal features, and T×N×C represents a feature with T time steps, N points, and C channels for each point.

[0067] The hierarchical ordered sequencer adopts a downsampling strategy with an exponential step size to extract subsequences of different time scales. The formula is as follows:

[0068] ;

[0069] ;

[0070] Among them, represents a slice vector for X, represents for a slice vector of, tS is the subscript for traversing values. For example, if the total number of frames is 24 frames and S is the sampling level, if S = 2, then t ranges from 1 to 12, so tS is 2, 4, 6, 8, …, 24. n traverses from 1 to N, c traverses from 1 to the coordinate dimension number 3, and f traverses from 1 to the feature dimension number C. The step size is , representing different time scales.

[0071] In the embodiments of the present application, by using a hierarchical ordered sequencer for multi-scale downsampling, the capture of high-frequency details and low-frequency structures is balanced, the receptive field of the model is expanded, and it performs better when processing long-time sequence actions.

[0072] As Figure 1 shown, the three-dimensional human behavior recognition method further includes the following steps:

[0073] Step S200: Sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud.

[0074] Step S300: Concatenate the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence.

[0075] In the embodiments of the present application, a cross-temporal serialization module is also provided. The cross-temporal serialization module converts unordered four-dimensional point cloud data into an ordered sequence to meet the one-way modeling requirements of the state space model (SSM). The Hilbert curve is used to sort each frame of three-dimensional point cloud in the subsequence, maintaining local continuity in space and reducing the distance difference between adjacent points in the sequence. Moreover, the point clouds of each frame are serialized in chronological order to ensure the coherence of temporal information.

[0076] As Figure 1 shown, the three-dimensional human behavior recognition method further includes the following steps:

[0077] Step S400: Determine the center point feature corresponding to each subsequence according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence.

[0078] In the embodiments of the present application, step S400 specifically includes:

[0079] Step S410: Construct a spatio-temporal neighborhood graph for each center point on each of the concatenated ordered spatio-temporal sequences;

[0080] Step S420: Normalize the point features in each of the spatio-temporal neighborhood graphs and perform feature fusion to obtain the center point features corresponding to each subsequence.

[0081] Specifically, the ordered spatial sequences of all frames are concatenated in chronological order to form an overall spatio-temporal ordered sequence and , where , represents the coordinates of the center point, L represents the sequence length, represents the center point feature. Since the input limit of the state space model is three-dimensional, including batch size, length, and number of points, it is necessary to multiply T and N and combine them into one L dimension, and the L dimension contains all the points within T frames .

[0082] This application is provided with a spatio-temporal structure aggregation layer, and the spatio-temporal structure aggregation layer is used to construct the spatio-temporal neighborhood graph of each center point on each of the concatenated ordered spatio-temporal sequences. Among them, there are multiple center points on the concatenated ordered spatio-temporal sequence, and the center points are obtained by farthest point sampling from the input points. After normalizing the point features in each of the spatio-temporal neighborhood graphs, feature fusion is performed through a multi-layer perceptron (MLP) to generate the updated center point features. The specific formula is as follows:

[0083] ;

[0084] ;

[0085] ;

[0086] ;

[0087] .

[0088] Among them, represents the center point feature, that is, it refers to , represents the coordinates of the center point, that is, it refers to . KNN represents the K-nearest neighbor method. Figure 2 The coordinates of the center point and the points in the neighborhood (abbreviated as neighboring points) in are both expressed as (x, y, z, t), represents the difference between the neighboring point after spatio-temporal embedding and the center point on the x-axis in space, represents the difference between the neighboring point after spatio-temporal embedding and the center point on the y-axis in space, Represents the difference in time t between the neighboring points and the central point after spatio-temporal embedding. Represents the neighboring point features. Represents the intermediate result after normalizing the difference between the neighboring point features and the central point features. Represents a very small constant, usually used to avoid division by zero errors or numerical stability. Represents the updated neighboring point features. Represents the natural logarithm. Represents the central point features updated after feature fusion. K represents the number of neighboring points obtained after the KNN algorithm. i represents traversing i times from the 1st to the Kth neighboring points, and j represents traversing j times from the 1st to the Kth neighboring points. Represents a multi-layer perceptron. Represents The features of the ith point traversed in Represents The features of the jth point traversed in

[0089] The embodiments of the present application can extract and integrate the local spatio-temporal features of the point cloud, and update the central point features by constructing a spatio-temporal neighborhood graph.

[0090] In an embodiment of the present application, the step S410 is specifically: constructing a spatio-temporal neighborhood graph for each central point on each of the spliced ordered spatio-temporal sequences by using the K-nearest neighbor method and the spatio-temporal embedding method.

[0091] Specifically, the embodiments of the present application use an extended K-nearest neighbor (KNN) method, combined with the spatio-temporal embedding method, to construct a spatio-temporal neighborhood graph for each central point.

[0092] As Figure 1 shown, the three-dimensional human behavior recognition method further includes the following steps:

[0093] Step S500, obtaining the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtaining an action recognition result based on the low-order spatio-temporal features and the central point features of each of the subsequences.

[0094] In the embodiments of the present application, the step S500 specifically includes:

[0095] Step S510, inputting the low-order spatio-temporal features into a pose encoder to obtain predicted skeleton key points;

[0096] Step S520, inputting the skeleton key points into a pose decoder to obtain high-dimensional geometric features;

[0097] Step S530: Fuse the high-dimensional geometric features and the center point features of each of the subsequences to obtain an action recognition result.

[0098] Specifically, in the embodiment of the present application, low-order spatio-temporal features are obtained through shared point 4D convolution, the low-order spatio-temporal features are input into a trained pose encoder to obtain predicted skeleton key points, and then a pose decoder is used to extract high-dimensional geometric features. After pooling, they are fused with the center point features of each subsequence to obtain an action recognition result.

[0099] In the embodiment of the present application, an auxiliary learning of pose estimation is performed using a pose-aware feature optimization module, which improves the model's perception ability of the human body's geometric structure and motion pattern and improves the recognition accuracy.

[0100] In an embodiment of the present application, the training steps of the state space model include:

[0101] Obtain a training data set, where the training data set includes: four-dimensional point cloud training data and corresponding action labels;

[0102] Input the four-dimensional point cloud training data into an initial state space model to extract training subsequences with different time scales in the four-dimensional point cloud training data;

[0103] Sort each frame of three-dimensional point cloud in each of the training subsequences to obtain ordered spatial sequence training data corresponding to each frame of three-dimensional point cloud;

[0104] Concatenate the ordered spatial sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence in chronological order to obtain concatenated ordered spatio-temporal sequence training data corresponding to each training subsequence;

[0105] Determine the center point feature training data corresponding to each training subsequence according to the concatenated ordered spatio-temporal sequence training data corresponding to each training subsequence;

[0106] Obtain low-order spatio-temporal feature training data extracted from the four-dimensional point cloud training data, and perform training based on the low-order spatio-temporal feature training data, the center point feature training data of each of the training subsequences, and the action labels to obtain a trained state space model.

[0107] In the embodiments of the present application, through cross-temporal serialization, four-dimensional point cloud data is effectively serialized, taking into account both spatial and temporal information, enabling the state space model to capture complex spatio-temporal dependencies in a unidirectional modeling framework. This unified modeling improves the model's understanding and recognition capabilities of dynamic actions. Compared with traditional convolutional neural networks and Transformer architectures, the present application reduces the running time and memory usage, especially when processing long-sequence four-dimensional point clouds, and can improve the computing efficiency. The present application is applicable to resource-constrained application scenarios. Specifically, it is not only applicable to human action recognition, but also can be extended to multiple fields such as robot navigation, autonomous driving, and intelligent monitoring, improving the system's action recognition and response capabilities in complex environments.

[0108] The state space model provided by the present application effectively solves the problems of high computational complexity and difficulty in capturing spatio-temporal dependencies in four-dimensional point cloud data analysis. The state space model of the present application not only outperforms traditional methods in terms of computational efficiency and memory usage, but also improves the robustness and recognition accuracy in complex environments through the pose perception mechanism. As Figure 3 and Figure 4 shown, Figure 3 are the test results of the state space model of the present application on the test set, Figure 4 is the attention visualization schematic diagram of three-dimensional human behavior recognition in the present application.

[0109] In an embodiment of the present application, determining the center point feature training data corresponding to each training subsequence according to the spliced ordered spatio-temporal sequence training data corresponding to each training subsequence includes:

[0110] Constructing a spatio-temporal neighborhood training graph for each center point on each of the spliced ordered spatio-temporal sequence training data;

[0111] Normalizing and fusing the point features within each of the spatio-temporal neighborhood training graphs to obtain the center point feature training data corresponding to each training subsequence.

[0112] In the embodiments of the present application, by constructing a spatio-temporal neighborhood training graph, efficient aggregation of local features is realized. At the same time, the state space model is responsible for capturing long-range dependencies, ensuring that the model does not sacrifice the understanding of global spatio-temporal relationships while maintaining linear complexity, thereby realizing efficient local feature extraction and global dependency capture.

[0113] In an embodiment of the present application, the training data set further includes: the true skeleton key points corresponding to the four-dimensional point cloud training data;

[0114] Training based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action labels to obtain a trained state space model, including:

[0115] Inputting the low-order spatiotemporal feature training data into a posture encoder to obtain predicted skeleton key points;

[0116] Calculating the mean square error loss between the predicted skeleton key points and the real skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder;

[0117] Inputting the predicted skeleton key points into a posture decoder to obtain high-dimensional geometric feature training data;

[0118] The high-dimensional geometric feature training data and the center point feature training data of each training subsequence are fused, and training is performed using the action label as the true value to obtain a trained state space model.

[0119] Specifically, in the training phase, the embodiment of the present application uses a convolution-based gesture encoder to map the low-order spatiotemporal feature training data to , where kp is the number of skeleton key points, which is the information in the training dataset, such as the 20 key points in the MSR Action3D dataset. The mean square error (MSE) loss between the predicted skeleton key points and the true skeleton key points is calculated to guide model learning.

[0120] The embodiments of the present application enhance the model's ability to learn the human skeleton structure and motion pattern and improve its robustness on sparse, incomplete and high-noise data sets by introducing a posture estimation task.

[0121] The embodiment of the present application utilizes a posture perception feature optimization module to perform auxiliary learning for posture estimation, thereby improving the model's ability to perceive the geometric structure and motion patterns of the human body, and enhancing the robustness and recognition accuracy on sparse, incomplete, and noisy data sets.

[0122] In addition, in terms of hardware, to improve the computational efficiency of the model, this application can be implemented on a GPU (Graphics Processing Unit), using the parallel computing ability to accelerate the operations of CTS (Cross-Temporal Serialization), STSAL (Spatio-Temporal Structure Aggregation Layer), and SSM (State Space Model). In terms of software, this application can be implemented in mainstream deep learning frameworks (such as PyTorch), using rich APIs (Application Programming Interfaces) and optimization tools to simplify the model development and training process. Optimization techniques such as Mixed Precision Training and Gradient Accumulation can also be adopted to improve the efficiency and stability of model training. In terms of data processing, this application can also filter four-dimensional point cloud data to remove noise points and outliers, improving the input quality and recognition accuracy of the model. In addition to pose estimation, other auxiliary tasks (such as point cloud classification and segmentation) can be introduced to further improve the feature expression ability of the model through multi-task learning. The model provided by this application also has compatibility and scalability, can be compatible with existing four-dimensional point cloud processing architectures, and can be integrated into existing systems as a plug-in module to improve their spatio-temporal modeling ability. In this way, this application can flexibly adapt to different application requirements and technical environments to achieve efficient and accurate four-dimensional point cloud data analysis and human action recognition.

[0123] In one embodiment, as Figure 5 shown, based on the above three-dimensional human behavior recognition method, the present invention also correspondingly provides a three-dimensional human behavior recognition device, including:

[0124] An extraction module 100, configured to input the four-dimensional point cloud data to be analyzed into a trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data;

[0125] A sorting module 200, configured to sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered space sequence corresponding to each frame of three-dimensional point cloud;

[0126] A splicing module 300, configured to splice the ordered space sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence;

[0127] A determination module 400, configured to determine the center point feature corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence;

[0128] An identification module 500, configured to obtain low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence.

[0129] It should be noted that the foregoing explanation of the embodiments of the three-dimensional human behavior recognition method is also applicable to the three-dimensional human behavior recognition device of this embodiment, and will not be repeated here.

[0130] The present invention discloses a three-dimensional human behavior recognition device. By inputting the four-dimensional point cloud data to be analyzed into a trained state space model, subsequences of different time scales in the four-dimensional point cloud data are extracted; each frame of three-dimensional point cloud in each of the subsequences is sorted to obtain an ordered space sequence corresponding to each frame of three-dimensional point cloud; the ordered space sequences corresponding to all frames of three-dimensional point cloud in each subsequence are concatenated in chronological order to obtain a concatenated ordered spatio-temporal sequence corresponding to each subsequence; the central point feature corresponding to each subsequence is determined according to the concatenated ordered spatio-temporal sequence corresponding to each subsequence; the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed are obtained, and an action recognition result is obtained based on the low-order spatio-temporal features and the central point features of each subsequence. In this application, by sorting each frame of three-dimensional point cloud to obtain an ordered space sequence, both spatial and temporal information is taken into account, complex spatio-temporal dependence relationships can be captured, and the computational complexity is reduced through the trained state space model, thereby improving the accuracy and efficiency of the behavior recognition result.

[0131] Figure 6 It is a schematic structural diagram of a terminal provided by an embodiment of this application. The terminal may include:

[0132] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.

[0133] When the processor 502 executes the program, it implements the three-dimensional human behavior recognition method provided in the foregoing embodiment.

[0134] Further, the terminal further includes:

[0135] A communication interface 503 for communication between the memory 501 and the processor 502.

[0136] The memory 501 is used to store a computer program executable on the processor 502.

[0137] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0138] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0139] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.

[0140] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0141] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above three-dimensional human behavior recognition method is implemented.

[0142] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0143] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0144] Any process or method description represented in a flowchart or otherwise described herein may be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application belong.

[0145] Logic and / or steps represented in a flowchart or otherwise described herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0146] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0147] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0148] In addition, in each embodiment of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0149] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A three-dimensional human behavior recognition method, characterized in that, The method includes: Input the four-dimensional point cloud data to be analyzed into the trained state space model, and extract subsequences with different time scales in the four-dimensional point cloud data; Sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; Stitch the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a stitched ordered spatio-temporal sequence corresponding to each subsequence; Determine the center point feature corresponding to each subsequence according to the stitched ordered spatio-temporal sequence corresponding to each subsequence; Obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain the action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence; Determine the center point feature corresponding to each subsequence according to the stitched ordered spatio-temporal sequence corresponding to each subsequence, including: Construct a spatio-temporal neighborhood graph for each center point on each of the stitched ordered spatio-temporal sequences; Perform normalization processing and feature fusion on the point features in each of the spatio-temporal neighborhood graphs to obtain the center point feature corresponding to each subsequence; Construct a spatio-temporal neighborhood graph for each center point on each of the stitched ordered spatio-temporal sequences, including: Use the K-nearest neighbor method and the spatio-temporal embedding method to construct a spatio-temporal neighborhood graph for each center point on each of the stitched ordered spatio-temporal sequences; Obtain the action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence, including: Input the low-order spatio-temporal features into the pose encoder to obtain the predicted skeleton key points; Input the skeleton key points into the pose decoder to obtain high-dimensional geometric features; Fuse the high-dimensional geometric features and the center point features of each subsequence to obtain the action recognition result.

2. The three-dimensional human behavior recognition method according to claim 1, characterized in that The training steps of the state space model include: Obtain a training data set, where the training data set includes: four-dimensional point cloud training data and corresponding action labels; Input the four-dimensional point cloud training data into the initial state space model, and extract training subsequences with different time scales in the four-dimensional point cloud training data; Sort each frame of three-dimensional point cloud in each of the training subsequences to obtain an ordered spatial sequence training data corresponding to each frame of three-dimensional point cloud; Stitch the ordered spatial sequence training data corresponding to the three-dimensional point clouds of all frames in each training subsequence in chronological order to obtain a stitched ordered spatio-temporal sequence training data corresponding to each training subsequence; Determine the center point feature training data corresponding to each training subsequence according to the stitched ordered spatio-temporal sequence training data corresponding to each training subsequence; Obtain the low-order spatio-temporal feature training data extracted from the four-dimensional point cloud training data, and perform training based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action labels to obtain the trained state space model.

3. The 3D human behavior recognition method according to claim 2, characterized in that, Determine the center point feature training data corresponding to each training subsequence according to the stitched ordered spatio-temporal sequence training data corresponding to each training subsequence, including: Construct a spatio-temporal neighborhood training graph for each center point on each of the stitched ordered spatio-temporal sequence training data; Normalize the point features in each of the spatio-temporal neighborhood training graphs and perform feature fusion to obtain the center point feature training data corresponding to each training subsequence.

4. The three-dimensional human behavior recognition method according to claim 2, wherein The training dataset further includes: the true skeleton key points corresponding to the four-dimensional point cloud training data; Train based on the low-order spatio-temporal feature training data, the center point feature training data of each training subsequence, and the action labels to obtain a trained state space model, including: Input the low-order spatio-temporal feature training data into the pose encoder to obtain predicted skeleton key points; Calculate the mean squared error loss between the predicted skeleton key points and the true skeleton key points corresponding to the four-dimensional point cloud training data to train the pose encoder; Input the predicted skeleton key points into the pose decoder to obtain high-dimensional geometric feature training data; Fuse the high-dimensional geometric feature training data and the center point feature training data of each training subsequence, and train with the action labels as the ground truth to obtain a trained state space model.

5. A three-dimensional human behavior recognition device, characterized in that, The device includes: An extraction module, configured to input the four-dimensional point cloud data to be analyzed into the trained state space model to extract subsequences of different time scales in the four-dimensional point cloud data; A sorting module, configured to sort each frame of three-dimensional point cloud in each of the subsequences to obtain an ordered spatial sequence corresponding to each frame of three-dimensional point cloud; A splicing module, configured to splice the ordered spatial sequences corresponding to the three-dimensional point clouds of all frames in each subsequence in chronological order to obtain a spliced ordered spatio-temporal sequence corresponding to each subsequence; A determination module, configured to determine the center point features corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence; An identification module, configured to obtain the low-order spatio-temporal features extracted from the four-dimensional point cloud data to be analyzed, and obtain an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence; Determining the center point features corresponding to each subsequence according to the spliced ordered spatio-temporal sequence corresponding to each subsequence includes: Construct a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences; Normalize the point features in each of the spatio-temporal neighborhood graphs and perform feature fusion to obtain the center point features corresponding to each subsequence; Constructing a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences includes: Using the K-nearest neighbor method and the spatio-temporal embedding method to construct a spatio-temporal neighborhood graph for each center point on each of the spliced ordered spatio-temporal sequences; Obtaining an action recognition result based on the low-order spatio-temporal features and the center point features of each subsequence includes: Input the low-order spatio-temporal features into the pose encoder to obtain predicted skeleton key points; Input the skeleton key points into the pose decoder to obtain high-dimensional geometric features; Fuse the high-dimensional geometric features and the center point features of each subsequence to obtain an action recognition result.

6. A terminal, characterized in that, Includes: A memory, a processor, and a three-dimensional human body behavior recognition program stored on the memory and executable on the processor. When the three-dimensional human body behavior recognition program is executed by the processor, it implements the steps of the three-dimensional human body behavior recognition method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the three-dimensional human body behavior recognition method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method, device and equipment based on Transform model and storage medium

    CN116721207A