A Behavior Recognition Method for Spatiotemporal Feature Network of Cross-Modal 3D Point Cloud Sequences
By introducing channels and spatial attention layers into the point cloud sequence, and designing spatiotemporal modeling and information injection modules, the problems of local feature extraction and spatiotemporal information loss in point cloud learning are solved, and more efficient behavior recognition effect is achieved.
Patent Information
- Application Number
- CN202210652520.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-09
AI Technical Summary
The prior art has problems in point cloud learning with strong local feature extraction capabilities but loss of space-time information, especially in non-uniform point cloud scenarios, and insufficient spatio-temporal structure information of depth graph sequences.
By introducing channel attention and spatial attention layers into the point cloud sequence, the spatiotemporal modeling module and spatiotemporal information injection module are designed, the spatiotemporal dimension information representation is enhanced, and the spatiotemporal feature network of cross-modal three-dimensional point cloud sequences is constructed, and the temporal and spatial feature information is injected to compensate for information losses.
Effectively integrating multi-scale human movement characteristics and space-time characteristics improves the utilization of the spatial structure and temporal change laws of behavior data, and improves the accuracy and robustness of behavior recognition.
Smart Images

Figure CN114973418B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and particularly to a method for behavior recognition of a cross-modal three-dimensional point cloud sequence spatio-temporal feature network. Background Art
[0002] With the continuous development of computer vision, behavior recognition has shown its wide application prospects and research value in many fields such as video surveillance and human-computer interaction; using depth map sequences for human behavior recognition is an important research field in machine vision and artificial intelligence. Although the widely used depth map sequences can provide depth information, they have a large amount of data redundancy and a large loss of spatio-temporal structure information of behavior data. The emergence of point clouds makes up for the disadvantages of depth map data. Point clouds are discrete point sets distributed in three-dimensional space, and they have unique advantages in expressing complex scenes and the shapes of objects. However, due to the irregular and disordered nature of point cloud distributions, it is not easy to apply deep learning to point clouds.
[0003] Currently, point cloud learning can be divided into volume-based methods and point-based methods:
[0004] (1) Volume-based methods: Volume-based methods usually voxelize point clouds into three-dimensional grids, and then apply three-dimensional convolutional neural networks to spatial representations for classification.
[0005] (2) Point-based methods: Point-based methods are directly executed on the original point clouds. The core idea of PointNet is to use a set of multi-layer perceptrons to abstract each point to learn its corresponding spatial encoding, and then aggregate all individual point features through a symmetric function to obtain a global point cloud feature. However, PointNet lacks the extraction and processing of local features, and the point clouds in real-world scenarios are often of different densities, while PointNet is trained based on uniformly sampled point clouds, resulting in a decrease in its accuracy in actual scenario point clouds;
[0006] Therefore, a hierarchical network PointNet++ is proposed in the prior art. The feature extraction of point sets consists of three parts, namely a sampling layer, a grouping layer, and a point-net-based learning layer. These three layers form an abstraction layer. PointNet++ consists of several sets of abstraction levels. PointNet++ gradually learns features using local region information through the hierarchical structure of several abstraction layers, and the network structure is more effective and robust; although PointNet++ can well extract local features through abstraction operations and gradually increase the receptive field, while performing abstraction operations, the farthest point sampling (FPS) will also reduce the number of outer contour points, which inevitably loses the spatio-temporal information of the original point cloud data. Summary of the Invention
[0007] In view of the deficiencies of depth maps, while retaining the powerful local feature extraction ability of PointNet++, the present invention makes up for the lost spatio-temporal feature information, and adds channel attention and spatial attention layers in the abstraction operation; and designs a spatio-temporal modeling module and a spatio-temporal information injection module. Channel attention and spatial attention are added to the spatio-temporal modeling module to enhance the ability of the spatio-temporal modeling module to capture important features, and then time and space feature information is injected into the feature sequence through the spatio-temporal information injection module to strengthen the information representation in the spatio-temporal dimension and make up for the information loss caused by FPS.
[0008] The technical solution adopted by the present invention is: a method for recognizing human behaviors of a cross-modal three-dimensional point cloud sequence spatio-temporal feature network includes the following steps:
[0009] S1. Collect human depth data, and cross-modally convert the depth map into a point cloud sequence through coordinate transformation;
[0010] S2. Input each frame of the point cloud sequence into a cross-modal three-dimensional point cloud sequence spatio-temporal feature network composed of a spatio-temporal modeling module and a spatio-temporal information injection module to obtain a feature vector sequence with temporal information and a spatial structure information feature vector sequence, splice them as the input of the fully connected layer, and perform human behavior recognition through a classifier.
[0011] Further, the spatio-temporal modeling module is composed of two abstraction operation layers, two groups of multi-layer perceptrons and a max pooling layer.
[0012] Further, the abstraction operation layer is composed of a sampling layer, a grouping layer, channel attention, spatial attention and a PointNet layer. The point cloud sequence is input into the abstraction operation layer, and the input is (T, n m , d + c i ) - dimensional;
[0013] The sampling layer uses farthest point sampling (FPS) to select n m points from the point set as centroids;
[0014] The grouping layer takes an n m-1 ×(d + c m-1 ) - dimensional point set and a set of centroid coordinates of size n m ×d as inputs, and the output is n m groups of point clusters of size n m ×k m ×(d + c m-1 ); where each group corresponds to a local area, and k m represents the number of local points within the neighborhood of the centroid point, and all points within the radius range are found through the ball radius query method, and k m is set as the upper limit within this radius range;
[0015] The input to the channel attention and spatial attention layers is a local region of n m ×k m ×(d + c m-1 )-dimensional points; m where n is the number of points in the local region;
[0016] First, the coordinates of the points in the local region are converted into a local coordinate system relative to the centroid point; second, the distance between each local point and the centroid is used as an additional 1D point feature; then, the feature-interaction attention mechanism is used to optimize the fusion effect of different features, and its expression is as follows:
[0017]
[0018] where represents the coordinates of the i-th point in the j-th region of the t-th point cloud frame, and are the centroid point coordinates and the point feature corresponding to respectively, is the and Euclidean distance between, A is the attention mechanism, and the coordinate and feature scores corresponding to each point are (3 + 1 + c m-1 )-dimensional. The attention scores in A are shared by all local points in all point cloud frames, and ⊙ are concatenation operation and dot product operation, is the region feature of the j-th region of the t-th point cloud frame after passing through the channel-spatial attention layer.
[0019] The channel attention module simultaneously uses the point cloud features after average pooling and max pooling, then feeds them into a multi-layer perceptron with shared weights in sequence, and finally merges the output feature vectors;
[0020] Spatial attention obtains one feature map each through max pooling and average pooling, then concatenates them into a 2D feature map, and then feeds it into a standard 7X7 convolution for parameter learning, and finally obtains a 1D weight feature map;
[0021] In the Pointnet layer, it consists of a group of mlp and a max pooling operation. The max pooling operation is used to combine the abstract features of all local points to generate a representation of the local region. Finally, the coordinates of the centroid point and its local region representation are concatenated into an abstract feature vector sequence of the centroid point
[0022] Finally, the spatio-temporal information of the entire point cloud frame is characterized by a group of multi-layer perceptrons and max pooling layers.
[0023] Furthermore, the spatio-temporal information injection module includes: a temporal information injection module and a spatial information injection module; by inputting the point cloud sequence of each frame and outputting the static appearance spatio-temporal feature vector of the corresponding frame to characterize the spatio-temporal structure information, the temporal information and spatial scale information are added to the static appearance spatio-temporal features of all frames through the spatio-temporal information injection module.
[0024] Furthermore, the temporal information injection module first encodes the time information of human actions, using a time position embedding layer, a shared MLPS layer, and a hierarchical pyramid max pooling layer. The time position embedding layer injects time position information using the order of the feature vector sequence. The shared MLPS layer performs a set of MLPS on each independent feature vector to extract the spatio-temporal information of each point cloud frame. The hierarchical pyramid max pooling layer extracts the sequence spatial information at multiple time scales.
[0025] Furthermore, the time position embedding layer uses sine and cosine functions with different frequencies as time position encoding:
[0026]
[0027]
[0028] where d sout represents the dimension of the feature vector, t is the time position, and h is the dimension position; the feature vector is updated by adding the position encoding as follows:
[0029]
[0030] where is the new feature vector after time position embedding; then, a new sequence of feature vectors
[0031] After passing through the time position embedding layer, the sequential information is simply embedded into the spatial information sequence. To further extract the spatio-temporal information, a set of MLPS is applied to each feature vector, and the formula is:
[0032]
[0033] where represents the feature vector updated using the MLP operation. Then, an updated sequence of feature vectors
[0034] The hierarchical pyramid max pooling layer (Two-MAX) is used to aggregate multiple feature vectors, and the vector sequence Perform multiple temporal partitions on an equal number of point cloud frames, and then perform max pooling operations on each partition to generate corresponding descriptors; use a hierarchical pyramid max pooling strategy with three partitions and two layers of pyramids; finally, concatenate the descriptors of all temporal partitions to form the sequence-level feature E of human behavior.
[0035] Furthermore, the sequence of feature vectors of temporal information includes: extracting the human action region-level feature M and the human action frame-level feature N, concatenating E, M, and N and outputting the temporal feature P; among them, the extraction formulas for the region-level feature M and the frame-level feature N are as follows:
[0036]
[0037]
[0038] Among them, is the abstract feature generated through the second set of abstract operations.
[0039] Furthermore, after temporal information injection in the spatio-temporal information injection module, a sequence of three-dimensional vector relationships with temporal information is generated through clustering. The sequence of three-dimensional vector relationships and the same group of random tensors jointly enter the inter-point attention mechanism module, learn the structural relationships between points in the point cloud data through the inter-point attention mechanism, and generate an inter-point relationship matrix representing the spatial structural relationships of the point cloud data;
[0040] Furthermore, the inter-point attention mechanism consists of a group of MLPs and softmax, and generates the inter-point relationship matrix. The formula is:
[0041] F s =MAX{MLP(R,E)} (8)
[0042] Among them, F s represents the sequence of spatial feature vectors that focus on representing spatial structural information, R represents the random tensor, and E is the sequence-level feature of human behavior;
[0043] Combine the inter-point relationships with each point in the point cloud sequence data to generate the spatial structural information feature F f The formula is:
[0044]
[0045] Among them, is the sequence of three-dimensional vector relationships generated after temporal information injection.
[0046] Furthermore, concatenate the sequence of feature vectors P with temporal information and the sequence of feature vectors F with spatial structural information f as the input Q of the fully connected layer, and then perform human action recognition through the classifier. The formula is as follows:
[0047]
[0048] Advantages of the present invention:
[0049] 1. A spatio-temporal information injection module is constructed to address the limitations of PointNet++, injecting dynamic temporal information into static point cloud sequences;
[0050] 2. Multi-scale human motion feature data and spatio-temporal feature data are fused, enabling full utilization of the spatial structure information and temporal variation patterns of behavior data;
[0051] 3. A cross-modal spatio-temporal feature network for 3D point cloud sequences is proposed, converting depth information across modalities into point cloud frame sequences to capture complex spatio-temporal structures and compensating for the deficiencies of depth map data. Brief Description of the Drawings
[0052] Figure 1 is a structural block diagram of the behavior recognition method of the cross-modal spatio-temporal feature network for 3D point cloud sequences of the present invention;
[0053] Figure 2 is a schematic diagram of the spatio-temporal modeling module of the present invention;
[0054] Figure 3 is a schematic diagram of the spatio-temporal information injection module of the present invention;
[0055] Figure 4 is the difference in recognition rate for different data inputs of the present invention;
[0056] Figure 5 is the difference in recognition rate for different feature fusion methods of the present invention. Detailed Embodiments
[0057] The present invention will be further described below in conjunction with the drawings and embodiments. This figure is a simplified schematic diagram, only illustrating the basic structure of the present invention in a schematic manner, and thus only showing the components related to the present invention.
[0058] To evaluate the effectiveness of the method of the present invention, experiments were conducted on a public dataset based on depth maps; large public datasets can provide more extensive training data for the model, making the model stronger; to verify the robustness of the method of the present invention, classic small datasets were also used in the selection of datasets. Therefore, experiments were conducted on several datasets with significantly different scales: MSR Action3D and NTU-RGB+D.
[0059] The present invention is based on the PyTorch framework, where the Python version is 3.7.0 and the PyTorch version is 1.10.1; the hardware platform for this experiment is a desktop computer, where the motherboard is MSI B460M MORTAR, the CPU is Intel i7 10700 with a main frequency of 2.9 GHz, the memory is 16 GB, the operating system is Windows 10 Professional, and the GPU resource is NVIDIA Tesla V100 with a video memory of 32 GB; the software tools used in the experiment are PyCharm and Anaconda3.
[0060] The MSR Action3D dataset records human action sequences, including a total of 20 action types and 10 subjects. Each subject performs each action 2 or 3 times, resulting in a total of 567 depth map sequences with a resolution of 640x240. The data is recorded using a depth sensor similar to the Kinect device. The action sequences of subjects numbered 1, 3, 5, 7, and 9 are used as the training set, and the rest are used as the test set.
[0061] NTU RGB+D belongs to a large dataset and provides more challenging action samples and more modal information; the NTU RGB+D dataset includes a total of 56,880 action samples, and uses acquisition devices at fixed different positions to provide four types of data for each sample captured from three perspectives: RGB, depth maps, skeleton sequences, and infrared radiation videos; it includes 60 action categories, and each action category is completed 1-2 times by 40 subjects; different from the acquisition method of MSR Action3D, NTU RGB+D provides 17 types of multi-view and multi-modal data collected by the acquisition device at different horizontal heights and distances; first, the cameras at different perspectives in the acquisition device are divided into two groups, where the 37,920 action sequences captured by cameras numbered 2 and 3 are used as the training set, and the 18,960 action sequences captured by camera numbered 1 are used as the test set.
[0062] As Figure 1 shown, the behavior recognition method of the cross-modal three-dimensional point cloud sequence spatio-temporal feature network includes the following steps:
[0063] Figure 1 The network structure consists of a spatio-temporal modeling module and a spatio-temporal information injection module. In the spatio-temporal modeling module, the point cloud set of each frame is input, and the static appearance spatio-temporal feature vector corresponding to the frame is output to represent the spatio-temporal structure information; the spatio-temporal information injection module adds temporal information and spatial scale information to the static appearance spatio-temporal features of all frames; then, the multi-scale human motion feature data and spatio-temporal feature data are effectively fused and the fully connected neural network is used for action classification and recognition.
[0064] S1. Collect the human body depth map data, and cross - modality convert the depth map into a point cloud sequence through coordinate transformation;
[0065] The point cloud sequence representing the point cloud framework of T frames, represents the unordered point set of the t - th frame point cloud framework, and n is the number of points;
[0066] S2. Input each frame of the point cloud sequence into a cross - modality 3D point cloud sequence spatio - temporal feature network composed of a spatio - temporal modeling module and a spatio - temporal information injection module to obtain a sequence of feature vectors with temporal information and a sequence of feature vectors of spatial structure information, perform a simple splicing as the input of the fully - connected layer, and perform human behavior recognition. The spatio - temporal information injection module includes: a temporal information injection module and a spatial information injection module;
[0067] Such as Figure 2 is the spatio - temporal modeling module. Further, the spatio - temporal modeling module consists of two abstract operation layers, two groups of multi - layer perceptrons and a max - pooling layer;
[0068] Further, the abstract operation layer consists of a sampling layer, a grouping layer, channel attention, spatial attention and a Pointnet layer. The point cloud sequence is input into the abstract operation layer, and the input is (T, n m , d + c i ) - dimensional, where d is set to 3 corresponding to the three - dimensional coordinates (x, y, z) of each point, and c i represents the c i -dimensional point feature, and c1 is set to 0;
[0069] In the sampling layer, use farthest point sampling (FPS) to select n m points from the point set as the centroids.
[0070] In the grouping layer, take the n m-1 ×(d + c m-1 ) - dimensional point set and a set of centroid coordinates of size n m ×d as the input, and the output is n m groups of size n m ×k m ×(d + c m-1 ) point clusters; where each group corresponds to a local area, and k m represents the number of local points within the neighborhood of the centroid point, and all points within the radius range are found through the ball - radius query method, and k m is set as the upper limit within this radius range.
[0071] In the channel attention layer and the spatial attention layer, channel attention and spatial attention are used to learn attention weights along the two dimensions of channels and space, adaptively adjust the point cloud features, obtain important features, compress unimportant features, and represent the temporal information and spatial structure of the static appearance of each frame of human behavior. The inputs of the channel attention layer and the spatial attention layer are local regions of n m ×k m ×(d + c m-1 ) - dimensional with n m points;
[0072] First, convert the coordinates of the points in the local region into a local coordinate system relative to the centroid point; second, use the distance between each local point and the centroid as a 1 - dimensional additional point feature to alleviate the impact of rotational motion on action recognition; then, use the feature - inter - attention mechanism to optimize the fusion effect of different features, and its manifestation is as follows:
[0073]
[0074] Among them, represents the coordinates of the i - th point in the j - th region of the t - th point cloud frame, and are the centroid point coordinates and the point feature corresponding to respectively, is and the Euclidean distance between, A is the attention mechanism, the coordinates and feature scores corresponding to each point are (3 + 1 + c m-1 ) - dimensional, and the attention scores in A are shared by all local points in all point cloud frames, and ⊙ are concatenation operation and dot - product operation, is the region feature of the j - th region of the t - th point cloud frame after passing through the channel and spatial attention layers.
[0075] Channel attention: The channel attention module simultaneously uses the point cloud features after average pooling and max pooling, then sends them into a multi - layer perceptron with shared weights in sequence, and finally merges the output feature vectors. To effectively calculate the channel attention, it is necessary to compress the spatial dimension of the input feature map. For the aggregation of spatial information, the commonly used method is average pooling. In addition, max pooling can collect more important clues between difficult - to - distinguish objects to obtain more detailed channel attention. Therefore, the features of average pooling and max pooling are used simultaneously.
[0076] Spatial attention: Spatial attention mainly focuses on which part has rich effective information, which is a complement to channel attention. One feature map is obtained through max pooling and average pooling respectively, and then they are concatenated into a 2D feature map, which is then fed into a standard 7X7 convolution for parameter learning. Finally, a 1D weight feature map is obtained, which encodes the positions that need to be attended to or suppressed. From a spatial perspective, channel attention is global while spatial attention is local.
[0077] The Pointnet layer consists of a group of MLPs and a max pooling operation. The max pooling operation is used to combine the abstract features of all local points to generate a representation of the local region. Finally, the coordinates of the centroid point and its local region representation are concatenated into a sequence of abstract feature vectors of the centroid point. where f t is the input point cloud S for each frame t and the corresponding output features for each frame.
[0078] Finally, the spatio-temporal information of the entire point cloud framework is characterized by a group of multi-layer perceptrons (MLPs) and max pooling layers.
[0079] Furthermore, the input point cloud sequence for each frame outputs a static appearance spatio-temporal feature vector corresponding to the frame to characterize the spatio-temporal structure information. The spatio-temporal information injection module adds temporal information and spatial scale information to the static appearance spatio-temporal features of all frames;
[0080] The spatio-temporal information injection module injects additional spatio-temporal structure information into the point cloud sequence, including temporal information injection and spatial information injection:
[0081] such as Figure 1 and 3 shown in the temporal information injection module. First, the time information of human actions is encoded using a time position embedding layer, a shared MLPS layer, and a hierarchical pyramid max pooling layer. The time position embedding layer injects time position information using the order of the feature vector sequence. The shared MLPS layer performs a group of MLPS on each independent feature vector to extract the spatio-temporal information of each point cloud framework. The hierarchical pyramid max pooling layer extracts sequence spatial information at multiple time scales.
[0082] The time position embedding layer uses sine and cosine functions with different frequencies as time position encoding:
[0083]
[0084]
[0085] where d soutrepresents the dimension of the feature vector, t is the time position, and h is the dimension position; the feature vector is updated by adding position encoding as follows:
[0086]
[0087] where, is the new feature vector after time position embedding; then, a new sequence of feature vectors
[0088] After passing through the time position embedding layer, the sequential information is simply embedded into the spatial information sequence. To further extract spatio-temporal information, a set of MLPs is applied to each feature vector, and the formula is:
[0089]
[0090] where, represents the feature vector updated using the MLP operation. Then, an updated sequence of feature vectors
[0091] The hierarchical pyramid max pooling layer (Two-MAX) is used to aggregate multiple feature vectors. To capture sub-actions within the point cloud sequence and encode more discriminative motion information, a hierarchical pyramid max pooling strategy is proposed: the sequence of feature vectors is divided into multiple time partitions for an equal number of point cloud frames, and then a max pooling operation is performed on each partition to generate the corresponding descriptor; in this embodiment, a hierarchical pyramid max pooling strategy with three partitions and two layers of pyramid is used; finally, the descriptors from all time partitions are simply concatenated to form the sequence-level feature E of human behavior.
[0092] To obtain more sufficient spatio-temporal information of human motion, human action features are integrated from different stages. For this purpose, district-level feature M and frame-level feature N are extracted, and the extraction methods are as follows:
[0093]
[0094]
[0095] where, is the abstract feature generated through the second set of abstraction operations; then, E, M, and N are simply concatenated as the temporal feature P.
[0096] Further, as Figure 3Then, spatial information injection is performed. After temporal information injection in the spatio-temporal information injection module, a three-dimensional vector relationship sequence with temporal information is generated through clustering. The three-dimensional vector relationship sequence and a group of random tensors jointly enter the inter-point attention mechanism module. Through the inter-point attention mechanism, the structural relationship between the point cloud data points is learned, and an inter-point relationship matrix representing the spatial structure relationship of the point cloud data is generated.
[0097] Random tensors can better perform deep learning on point clouds, enabling the network to autonomously learn a relationship matrix more suitable for representing the data spatial structure. In this embodiment, a set of tensors with a set size but random data is used. A tensor is a powerful method for representing direction and space. Through tensors, not only can the spatial structure information of the data be better represented, but also the running speed of the network can be accelerated.
[0098] The inter-point attention mechanism consists of a group of MLPs and softmax. The MLP can well learn the spatio-temporal information of the more key points in the point cloud data, and then through the softmax layer, it is converted into weight coefficients, that is, a relationship matrix that can use random tensors to represent the spatial structure relationship between points is generated. Its expression form is as follows:
[0099] F s =MAX{MLP(R , E)} (8)
[0100] Among them, F s represents the generated spatial feature vector sequence (spatio-temporal feature 1) that focuses on representing spatial structure information, R represents the random tensor, and E is the sequence-level feature of human behavior.
[0101] Furthermore, in order to combine the inter-point relationship with each point of the point cloud sequence data to generate the spatial structure information feature F f The formula is:
[0102]
[0103] Among them, is the three-dimensional vector relationship sequence generated after temporal information injection, and it is subjected to max pooling operation to be abstracted as spatio-temporal feature 2. Thus, the inter-point relationship is combined with each point to generate the spatial structure information feature F f ;
[0104] Then, the feature vector sequence P with temporal information and the spatial structure information feature vector sequence F f are simply concatenated as the input Q of the fully connected layer, and then human action recognition is performed through the classifier. The formula is as follows:
[0105]
[0106] The experimental process is as follows:
[0107] Sample 512 points from the point cloud set as the point cloud framework. First, randomly select 2048 points from the point cloud set, and then use the PFS algorithm to select 512 points from the 2048 points. In the spatio-temporal modeling module, perform two set abstraction operations on each point cloud framework to model the spatio-temporal structure. In the first set abstraction operation, select 128 centroids to determine point groups, set the group radius to 0.06, and set the number of points in each point group to 48; in the second set abstraction operation, select 32 centroids to determine point groups, set the group radius to 0.1, and set the number of points in each point group to 16. As shown in Table 1, before extracting the spatial structure information, first use clustering to generate a three-dimensional vector relationship sequence, set the clustering radius to 20. When extracting the spatial structure information, set the random tensor size to (8, 64, 64), and set dropout to 0.5. In order to prevent overfitting caused by the excessive size of the NTU RGB+d120 dataset, set dropout to 0.8 when testing the NTU RGB+d120 dataset. As shown in Table 2, adopt the same data augmentation strategy in 3DV-PointNet++ for the training data, including random rotation around the Y and X axes, jittering, and random point dropout; use Adam as the optimizer, start with a learning rate of 0.001, and decay at a rate of 0.5 every 10 epochs.
[0108] Table 1 Experimental settings for spatio-temporal modeling
[0109]
[0110] Table 2 Experimental settings for spatio-temporal information injection
[0111]
[0112] Since the data used in this experiment is a cross-modal point cloud sequence, although the spatial modeling has been performed through the spatio-temporal modeling module to make it have spatio-temporal structural features, the loss of some spatial structure information is inevitably caused by the disruption of the order of the point cloud sequence. Therefore, use the spatio-temporal information injection module to supplement the feature sequence; in order to explore which data is more conducive to the extraction of spatial information and the effects of different data extraction methods on the recognition rate, different experiments are carried out for comparison to find the most suitable experimental method.
[0113] First, experiments were conducted using the MSR Action3D small dataset. Two different types of data were used as the input to the spatio-temporal information injection module. One type of 3D point cloud data was the original 3D point cloud data, that is, the 3D point cloud data before abstraction operations; the other type of 3D point cloud data was the 3D vector relationship sequence generated through clustering after position encoding and spatio-temporal modeling (hereinafter referred to as the original data and relationship data). After that, multiple experiments were carried out and the final experimental results were recorded, as shown in Table 3. The data comparison graph is as Figure 4 , and the feature injection comparison graph is as Figure 5 ;
[0114] Table 3 Experimental process on MSR-ACTION3D
[0115]
[0116] It can be seen from Figure 4 that when using the original data as the input to the spatio-temporal information injection module, the highest recognition rate can reach 93.75%, but the lowest recognition rate is only 89.71%; when using the relationship data as the input to the spatio-temporal information injection module, the highest recognition rate reaches 93.01% and the lowest recognition rate is 90.81%. After comparing with the original results without injecting spatio-temporal features, it can be concluded that using the original data has a better effect but poor stability, while using the relationship data has a better effect and better stability;
[0117] It can be seen from Figure 5 that if only spatio-temporal feature 1 is injected for feature extraction, the final accuracy rate is only 86.76%. If only spatio-temporal feature 2 is injected for feature extraction, the accuracy rate is 91.18%. By comparing with other experiments, it can be known that using spatio-temporal feature 1 or spatio-temporal feature 2 alone to supplement the loss of spatio-temporal structure information has an even worse effect than not injecting spatio-temporal information. Among them, the effect of only injecting spatio-temporal feature 1 is even about 5% lower. This is because using only any one of spatio-temporal feature 1 or 2 cannot combine the relationship between points; while by aggregating spatio-temporal feature 1 and 2, a feature vector representing the spatio-temporal structure between each point of the point cloud is formed, compensating for the spatio-temporal structure information lost by using the point cloud sequence. Experiments also prove the rationality of the network model structure and theory.
[0118] After obtaining the results using the MSR Action3D small dataset, experiments on the NTU RGB+d120 and NTU RGB+d60 large datasets were started. First, the original data was used as the input to the spatio-temporal information injection module for experiments on the NTU large dataset, but the effect was not good; then the relationship data was used as the input to the spatio-temporal information injection module, and the results were recorded as shown in Table 4;
[0119] Table 4 Experimental process on NTU RGB+d60 / 120
[0120]
[0121] As can be seen from the results in Table 4, the accuracy of the network after spatio-temporal information injection has increased by up to 0.22%, and it also has a high accuracy on the NTURGB+d120 large dataset, which directly proves the rationality and feasibility of spatio-temporal information injection; through the experiment on the NTU RGB+d60 large dataset, it can be concluded that the method of the present invention has a higher accuracy for the classification of human behavior recognition.
[0122] NTU RGB+d60 dataset: First, compare the method of the present invention with the state-of-the-art method on the NTU RGB+d60 dataset. The NTU RGB+d60 dataset is a large-scale indoor human activity dataset. As shown in Table 5, the accuracy of the method of the present invention has reached 97.8%, and the method of the present invention shows better performance compared with other methods on the NTU RGB+d60 dataset.
[0123] Table 5 Action recognition accuracy on NTU RGB+D60
[0124]
[0125]
[0126] NTU RGB+d120 dataset: Then compare the method of the present invention with the state-of-the-art method on the NTU RGB+d120 dataset; the NTU RGB+d120 dataset is the largest dataset for 3D action recognition. Compared with the NTU RGB+d60 dataset, it is more challenging to perform three-dimensional human action recognition on the NTU RGB+d120 dataset; as shown in Table 6, the accuracy of the method of the present invention is 95.3%, and on the NTU RGB+d120 dataset, the method of the present invention shows better performance compared with other methods on the NTU RGB+d120 dataset.
[0127] Table 6 Action recognition accuracy on NTU RGB+D120
[0128]
[0129] MSR Action3D dataset: In order to comprehensively evaluate the method of the present invention, a comparative experiment was carried out on the small MSR Action3D dataset. In order to alleviate the overfitting problem on the small-scale dataset, the batch size was set to 8, and other parameter settings were the same as those on the two large datasets. Table 7 shows the recognition accuracies of different methods, and the method of the present invention shows better performance compared with other methods on the MSR Action3D dataset.
[0130] Table 7 Action recognition accuracy on MSR-ACTION3D
[0131]
[0132] Enlightened by the above ideal embodiments of the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A behavior recognition method for a spatio-temporal feature network of cross-modal three-dimensional point cloud sequences, characterized in that, Including the following steps: Collect human depth data, and cross - modality convert the depth map into a point cloud sequence through coordinate transformation; Input each frame of the point cloud sequence into a cross - modality three - dimensional point cloud sequence spatio - temporal feature network composed of a spatio - temporal modeling module and a spatio - temporal information injection module to obtain a sequence of feature vectors with temporal information and a sequence of feature vectors with spatial structure information, splice them as the input of the fully connected layer, and perform human behavior recognition through a classifier; The spatio - temporal modeling module consists of two abstract operation layers, two groups of multi - layer perceptrons and a max - pooling layer; The abstract operation layer consists of a sampling layer, a grouping layer, a channel attention, a spatial attention and a Pointnet layer; The sampling layer selects n m points from the point set as centroids using FPS; The grouping layer takes an n m-1 × (d + c m-1 ) - dimensional point set and a set of centroid coordinates of size n m × d as inputs, and outputs n m groups of point clusters of size n m × k m × (d + c m-1 ); The channel attention and spatial attention layers convert the coordinates of points in the local area into a local coordinate system relative to the centroid point; take the distance between each local point and the centroid as a 1 - D additional point feature; use the feature - inter attention mechanism to optimize the fusion effect of different features, and the formula is as follows: Among them, represents the coordinates of the i-th point in the j-th region of the t-th point cloud frame, and are respectively the centroid point coordinates and point pair features corresponding to , is and the Euclidean distance between, A is the attention mechanism, and the coordinate and feature scores corresponding to each point are (3 + 1 + c m-1 ) - dimensional. The attention scores in A are shared by all local points in all point cloud frames. and ⊙ are the concatenation operation and the dot - product operation, is the region feature of the j-th region of the t-th point cloud frame after passing through the channel - spatial attention layer; The Pointnet layer consists of a group of mlp and a max - pooling operation; The spatio - temporal information injection module includes: a temporal information injection module and a spatial information injection module; The timing information injection module encodes the time information of human actions, using a temporal positional embedding layer, a shared MLPS layer, and a hierarchical pyramid max pooling layer; after passing through the temporal positional embedding layer, the sequential information is embedded into the spatial information sequence; a set of MLPSs are applied to each feature vector; an updated feature vector sequence is generated; Two-MAX is used to aggregate multiple feature vectors, and the vector sequence performs multiple temporal partitions on an equal number of point cloud frames, and then performs a max pooling operation on each partition to generate corresponding descriptors; uses a hierarchical pyramid max pooling strategy with a three-partition two-layer pyramid; concatenates the descriptors of all temporal partitions to form a sequence-level feature E of human behavior; The time - position embedding layer uses sine and cosine functions with different frequencies as time - position encoding, and the formula is as follows: where d sout represents the dimension of the feature vector, t is the time position, and h is the dimension position; the feature vector is updated by adding positional encoding as follows: Among them, is the new feature vector after time position embedding; Then, a new sequence of feature vectors is obtained 2. The behavior recognition method of the cross-modal three-dimensional point cloud sequence spatio-temporal feature network according to claim 1, wherein, The sequence of feature vectors of the temporal information includes: extracting the human action region - level feature M, the human action frame - level feature N and the sequence - level feature E of human behavior, connecting E, M and N and outputting the temporal feature P; among them, the extraction formulas of the region - level feature M and the frame - level feature N are as follows: Among them, is an abstract feature generated by a second set of abstract operations.
3. The behavior recognition method of the cross-modal three-dimensional point cloud sequence spatio-temporal feature network according to claim 2, characterized in that: The spatial information injection module generates a three - dimensional vector relationship sequence with temporal information through clustering after temporal information injection. The three - dimensional vector relationship sequence and a group of random tensors jointly enter the point - to - point attention mechanism module, learn the structural relationship between points of the point cloud data through the point - to - point attention mechanism, and generate a point - to - point relationship matrix representing the spatial structure relationship of the point cloud data.
4. The behavior recognition method of the cross-modal three-dimensional point cloud sequence spatio-temporal feature network according to claim 3, wherein: The point - to - point attention mechanism consists of a group of MLPS and softmax, and generates a point - to - point relationship matrix, and the formula is: F s = MAX{MLP(R, E)} (8) Among them, F s represents the generated sequence of spatial feature vectors focusing on representing spatial structure information, R represents a random tensor, and E is the sequence-level feature of human behavior; Combining the relationship between points with each point of the point cloud sequence data to generate the spatial structure information feature F f The formula is as follows: Among them, is a three-dimensional vector relationship sequence generated after injecting timing information.
Citation Information
Patent Citations
Depth video human body behavior recognition method based on three-dimensional space sequential modeling
CN110852182A
Multi-modal target detection method and system suitable for modal deficiency
CN114359586A