Multi-scale space-time frequency behavior identification method based on point cloud sequence

By constructing multi-scale spatiotemporal feature extraction and spatiotemporal frequency enhancement modules, the problem of insufficient processing of low-level semantic and high-level semantic information in deep networks when processing point cloud sequences is solved, and higher behavior recognition accuracy is achieved.

CN120340136APending Publication Date: 2025-07-18HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510594491.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When processing point cloud sequences, existing deep neural networks are difficult to fully and meticulously process the differences in low-level semantic information and high-level semantic information in human behavior data, resulting in limitations in understanding and characterizing human behavior data.

Method used

A multi-scale spatiotemporal feature extraction module and spatiotemporal frequency enhancement module are built. Through multi-scale spatiotemporal convolution and spatiotemporal frequency enhancement convolution, combined with discrete cosine transformation and attention mechanism, the features of point cloud sequences are extracted and weighted, and pooled operations are carried out to retain key feature dimensions for classification recognition.

Benefits of technology

The accuracy of point cloud sequence behavior recognition has been significantly improved, especially on the NTU-RGB+D 60, NTU-RGB+D 120 and MSR-Action3D datasets, reaching 97.2%, 89.2% and 93.5%, respectively, which is better than the existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340136A_ABST
    Figure CN120340136A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, and relates to the technical field of behavior recognition, and the method comprises the steps: introducing two modules, namely a multi-scale spatio-temporal feature extraction module and a spatio-temporal feature frequency enhancement module; the multi-scale spatio-temporal feature extraction module is used for decomposing an original point cloud sequence in a spatial dimension and a time dimension, and enriching low-level semantic information of the spatial dimension by adopting multi-scale spatio-temporal convolution; and the space-time frequency enhancement module adaptively learns channel information of high-level semantics based on frequency information of a time dimension, so that the discrimination ability of behavior space-time characteristics is improved. Experimental results show that compared with a conventional mainstream point cloud-based method, the method is remarkably improved, and the recognition precision is also superior to that of most methods in other directions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of behavior recognition, and particularly to a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence. Background Art

[0002] Point cloud data has the characteristics of rich information and easy processing, and has received extensive attention in the field of three-dimensional understanding. Compared with two-dimensional images, point clouds are unordered point sets and are difficult to process with traditional methods. Recently, many deep learning-based methods, such as PointNet and PointNet++, have been proposed to directly process raw point clouds through MLP structures, avoiding the complexity of data conversion and information loss.

[0003] Dynamic three-dimensional data is closer to the real world than static data. Therefore, compared with static point cloud data, point cloud sequences provide richer temporal and spatial information. The difficulty in processing point cloud sequences lies not only in the addition of the time dimension, but also in the difficulty of capturing the movement of objects. The PointRNN model predicts the movement trajectory of points by combining the coordinates and time information of points, but faces challenges in computing resources and accuracy. The MeteorNet model learns the features of each point by aggregating the spatio-temporal neighborhood information of point cloud sequences to achieve a wider range of information aggregation.

[0004] However, the current deep neural networks fail to differentially process the low-level semantic information and high-level semantic information in human behavior data, and the processing of features is not sufficient and detailed enough, which leads to limitations in the deep network's understanding and representation of human behavior data. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, including the following steps: S1. Construct a multi-scale spatio-temporal feature extraction module to perform multi-scale spatio-temporal convolution on the input point cloud sequence to obtain spatio-temporal features; S2. Construct a spatio-temporal frequency enhancement module to perform spatio-temporal frequency enhancement convolution on the output features of the multi-scale spatio-temporal feature extraction module and perform weighted processing on the semantic information; S3. Perform a pooling operation on the output features of the spatio-temporal frequency enhancement module to retain the feature dimensions; S4. Send the final features obtained by the pooling process to a fully connected classifier to recognize human actions.

[0006] The further limited technical solution of the present invention is: Further, step S1 specifically includes the following sub-steps: S1.1. Select several time anchor frames in the input point cloud sequence; S1.2. Select spatial anchor points in the anchor frame through farthest point sampling; S1.3. Map the anchor points in the anchor frame to adjacent frames, and use these original and mapped anchor points to construct the central axis of the spatio-temporal axis body; S1.4. For the spatial neighborhood of each anchor point in each frame, adopt a multi-scale method to perform spatial convolution; perform temporal convolution on the spatio-temporal axis bodies generated at different scales to obtain their respective spatio-temporal features; S1.5. Fuse the obtained spatio-temporal features to obtain new features representing the dynamic changes in the local area of the point cloud sequence.

[0007] For a multi-scale spatio-temporal frequency behavior recognition method based on point cloud sequence as described above, in step S1.1, the point cloud sequence is [P1,F1], [P2,F2],..., [P L ,F L ; where, P t ∈R 3×N , F t ∈R C×N represent the point coordinates and features in the t-th frame of the point cloud sequence, L represents the number of frames, and N and C respectively represent the number of points and the number of channels.

[0008] For a multi-scale spatio-temporal frequency behavior recognition method based on point cloud sequence as described above, in step S1.1, select temporal anchor frames based on the temporal kernel size and temporal stride; the temporal kernel represents the length of the spatio-temporal axis body, and the temporal stride represents the inter-frame distance between two adjacent anchor frames.

[0009] For a multi-scale spatio-temporal frequency behavior recognition method based on point cloud sequence as described above, in step S1.4, decouple the temporal dimension and spatial dimension of the point cloud sequence, first model the unordered spatial structure to obtain the spatial local structure of the point cloud in 3D space; then perform temporal convolution to obtain the spatio-temporal feature representation of the point cloud sequence; The multi-scale spatio-temporal convolution formula is as follows: where, (x,y,z)∈P t and (δ x ,δ y ,δ z ) represent displacement, r i ∈R represents the spatial search radius, represents the spatial convolution kernel, represents the temporal convolution kernel, l represents the size of the temporal convolution kernel, C m represents the dimension of the intermediate feature, M t represents the sequence feature after spatial convolution, represents the sequence feature after temporal convolution.

[0010] As described above, in a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, in step S1.4, relative coordinates are used in space to generate a convolution kernel, the kernel is converted into a displacement function, and the updated spatial convolution formula is as follows: Where, represents a function of (δ x , δ y , δ z ), parameterized by θ, and different spatial convolution kernels are generated according to different displacements ; To further reduce the computational complexity, the function f is decomposed into f((δ x , δ y , δ z ); θ) = θ d ∙(δ x , δ y , δ z ) T ∙1∙⊙θ s , where θ = [θ d , θ s , represents the displacement transformation kernel, represents the shared kernel, 1 = (1,..., 1) ∈ R 1×C represents the broadcast kernel, and ⊙ represents the element-wise product.

[0011] As described above, in a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, step S2 specifically includes the following sub-steps: S2.1. After the point cloud sequence undergoes multi-scale spatio-temporal convolution operations, discrete cosine transform is applied in the time dimension to obtain frequency information, resulting in multi-frequency vectors; S2.2. Apply the attention mechanism to the multi-frequency vectors across the entire frequency channel to obtain weighted multi-frequency features.

[0012] As described above, in a multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, in step S2.1, discrete cosine transform is used for frequency domain feature decomposition, and the formula for discrete cosine transform is as follows: Where, f ∈ R H represents the spectrum of the discrete cosine transform, h ∈ {0, 1,..., H - 1}; x ∈ R H represents the input data, H represents the length of the input component, and the cosine function represents the weight; To obtain all frequency components, the inverse transform of the discrete cosine transform formula is performed, and the formula is as follows: where \(i\in\{0,1,\cdots,H - 1\}\); To fuse the frequency components, the input \(X\) is divided into \(n\) parts along the channel dimension, and the corresponding discrete cosine transform frequency components are assigned to each part. The formula is as follows: where \(i\in\{0,1,\cdots,n - 1\}\), GAP represents the average pooling operation in the spatial dimension \(w\) direction, and Freq i \(\in\mathbb{R}\) C' represents the compressed \(C'\)-dimensional vector; the entire compressed vector is obtained by concatenation: where Freq \(\in\mathbb{R}\) c represents the obtained multi-frequency vector.

[0013] In a multi-scale spatio-temporal frequency behavior recognition method based on point cloud sequences as described above, in step S2.2, the obtained multi-frequency vector is applied with an attention mechanism in the frequency channel to weight the entire feature channel. The formula is as follows: where freq_att represents the multi-frequency vector after weighting processing.

[0014] In a multi-scale spatio-temporal frequency behavior recognition method based on point cloud sequences as described above, in step S3, average pooling and max pooling are respectively used for spatial pooling and temporal pooling to eliminate the spatial and temporal dimensions and only retain the feature dimension for classification processing.

[0015] The beneficial effects of the present invention are as follows: In the present invention, a multi-scale spatio-temporal frequency behavior recognition network is proposed for recognizing three-dimensional behaviors. The network mainly includes two modules, namely, a multi-scale spatio-temporal feature extraction module and a spatio-temporal feature frequency enhancement module; the multi-scale spatio-temporal feature extraction module decomposes the original point cloud sequence in the spatial dimension and the temporal dimension, and adopts a multi-scale strategy to comprehensively enrich the low-level semantic information in the spatial dimension; the spatio-temporal feature frequency enhancement module adaptively learns the channel information of high-level semantics based on the frequency information in the temporal dimension, thereby enhancing the discriminability of the behavior spatio-temporal features; the model achieves a cross-view accuracy of 97.2% on the NTU-RGB+D 60 dataset, a cross-subject accuracy of 89.2% on the NTU-RGB+D 120 dataset, and an accuracy of 93.5% on the MSR-Action3D dataset, which is significantly improved compared with the previous mainstream point cloud-based methods, and the recognition accuracy is also better than most methods in other directions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1Schematic diagram of the overall process of the present invention; Figure 2 Schematic diagram of the structure of the multi-scale spatio-temporal feature extraction module in the embodiment of the present invention; Figure 3 Schematic diagram of the structure of the spatio-temporal feature frequency enhancement module in the embodiment of the present invention. Detailed implementation manners

[0017] A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence provided in this embodiment, as Figure 1 shown, includes the following steps: S1. Construct a multi-scale spatio-temporal feature extraction module, perform multi-scale spatio-temporal convolution on the input point cloud sequence, and obtain spatio-temporal features.

[0018] Step S1 specifically includes the following sub-steps: S1.1. Select several time anchor frames in the input point cloud sequence. The input point cloud sequence is [P1, F1], [P2, F2],..., [P L , F L ; where P t ∈ R 3×N , F t ∈ R C×N represent the point coordinates and features in the t-th frame of the point cloud sequence, L represents the number of frames, and N and C respectively represent the number of points and the number of channels; For the input point cloud sequence, select time anchor frames based on the time kernel size and time stride; where the time kernel represents the length of the spatio-temporal axis body, and the time stride represents the frame distance between two adjacent anchor frames.

[0019] S1.2. Select spatial anchor points in the anchor frames by farthest point sampling.

[0020] S1.3. Map the anchor points in the anchor frames to the adjacent frames, and use these original and mapped anchor points to construct the central axis of the spatio-temporal axis body.

[0021] S1.4. As Figure 2 shown, since the point cloud sequence is irregular and disordered in space and ordered in time, to prevent the influence of the point cloud irregularity on time modeling, for the input point cloud sequence, decouple the time dimension and the space dimension, first model the disordered spatial structure to obtain the spatial local structure of the point cloud in 3D space; then perform time convolution to obtain the spatio-temporal feature representation of the point cloud sequence.

[0022] The multi-scale spatio-temporal convolution formula is as follows: where (x, y, z) ∈ P t and (δ x, δ y , δ z ) represents displacement, r i ∈R represents the spatial search radius, represents the spatial convolution kernel, represents the temporal convolution kernel, l represents the size of the temporal convolution kernel, C m represents the dimension of the intermediate feature, M t represents the sequence feature after spatial convolution, represents the sequence feature after temporal convolution.

[0023] Generate a convolution kernel using relative coordinates in space, convert the kernel to a displacement function, and the updated spatial convolution formula is as follows: Among them, represents the function of (δ x , δ y , δ z ), parameterized by θ, and different spatial convolution kernels are generated according to different displacements .

[0024] To further reduce the computational complexity, decompose the function f into f((δ x , δ y , δ z ); θ) = θ d ∙(δ x , δ y , δ z ) T ∙1∙⊙θ s , where θ = [θ d , θ s , represents the displacement transformation kernel, represents the shared kernel, 1 = (1,..., 1) ∈ R 1×C represents the broadcast kernel, and ⊙ represents the element-wise product.

[0025] S1.5. Fuse the obtained spatio-temporal features to obtain new features representing the dynamic changes within the local region of the point cloud sequence.

[0026] S2. Construct a spatio-temporal frequency enhancement module, as Figure 3 shown, perform spatio-temporal frequency enhancement convolution on the output features of the multi-scale spatio-temporal feature extraction module, and weight the high-level semantic information.

[0027] Step S2 specifically includes the following sub-steps: S2.1. After the point cloud sequence undergoes multi-scale spatio-temporal convolution operations, the feature dimension is increased. The discrete cosine transform (DCT) is used to extract the frequency information in the time dimension. First, DCT is used for frequency-domain feature decomposition, and the formula of DCT is as follows: where \(f\in\mathbb{R}\) H represents the spectrum of the discrete cosine transform, \(h\in\{0,1,\ldots,H - 1\}\); \(x\in\mathbb{R}\) H represents the input data, \(H\) represents the length of the input component, and the cosine function represents the weight.

[0028] To obtain all the frequency components, the inverse transform of the DCT formula is performed, and the formula is as follows: where \(i\in\{0,1,\ldots,H - 1\}\).

[0029] To fuse the frequency components, the input \(X\) is divided into \(n\) parts along the channel dimension, and each part is assigned the corresponding DCT frequency component. The formula is as follows: where \(i\in\{0,1,\ldots,n - 1\}\), GAP represents the average pooling operation in the spatial dimension \(w\) direction, and Freq i \(\in\mathbb{R}\) C' represents the compressed \(C'\)-dimensional vector; the entire compressed vector is obtained by concatenation: where Freq\(\in\mathbb{R}\) c represents the obtained multi-frequency vector.

[0030] S2.2. Apply the attention mechanism to the obtained multi-frequency vector in the frequency channel to perform weighted processing on the entire feature channel. The formula is as follows: where freq_att represents the multi-frequency vector after weighted processing.

[0031] S3. Apply average pooling and max pooling to the output features of the spatio-temporal frequency enhancement module in the spatial dimension and time dimension respectively to eliminate the spatial and time dimensions and only retain the feature dimension for classification processing.

[0032] S4. Send the final features obtained by the pooling process to a fully connected classifier to recognize human actions.

[0033] The method of this embodiment designs a multi-scale spatio-temporal frequency behavior recognition network, which is used to differentially process the low-level semantic information and high-level semantic information of point cloud data, effectively improving the network's understanding and recognition ability of human behavior data; it achieves a cross-view accuracy of 97.2% on the NTU-RGB+D 60 dataset, a cross-subject accuracy of 89.2% on the NTU-RGB+D 120 dataset, and an accuracy of 93.5% on the MSR-Action3D dataset, significantly superior to the methods of the prior art.

[0034] In addition to the above embodiments, the present invention may also have other embodiments. All technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.

Claims

1. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence, characterized in that: It includes the following steps: S1. Construct a multi-scale spatio-temporal feature extraction module, perform multi-scale spatio-temporal convolution on the input point cloud sequence, and obtain spatio-temporal features; S2. Construct a spatio-temporal frequency enhancement module, perform spatio-temporal frequency enhancement convolution on the output features of the multi-scale spatio-temporal feature extraction module, and perform weighted processing on the semantic information; S3. Perform a pooling operation on the output features of the spatio-temporal frequency enhancement module to retain the feature dimension; S4. Send the final features obtained by the pooling process to a fully connected classifier to recognize human actions.

2. The multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 1, wherein: The specific steps of step S1 include the following sub-steps: S1.

1. Select several time anchor frames in the input point cloud sequence; S1.

2. Select spatial anchor points in the anchor frames by farthest point sampling; S1.

3. Map the anchor points in the anchor frames to adjacent frames, and use these original and mapped anchor points to construct the central axis of the spatio-temporal axis body; S1.

4. For the spatial neighborhood of each anchor point in each frame, adopt a multi-scale method to perform spatial convolution; perform temporal convolution on the spatio-temporal axis bodies generated at different scales to obtain their respective spatio-temporal features; S1.

5. Fuse the obtained spatio-temporal features to obtain new features representing the dynamic changes in the local area of the point cloud sequence.

3. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 2, characterized in that: In the step S1.1, the point cloud sequence is [P1, F1], [P2, F2],..., [P L , F L ; where P t ∈ R 3×N , F t ∈ R C×N represents the point coordinates and features in the t-th frame of the point cloud sequence, L represents the number of frames, and N and C represent the number of points and the number of channels respectively.

4. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 3, characterized in that: In step S1.1, select time anchor frames based on the time kernel size and time stride; The time kernel represents the length of the spatio-temporal axis body, and the time stride represents the frame distance between two adjacent anchor frames.

5. The multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 2, characterized in that: In step S1.4, decouple the time dimension and the spatial dimension of the point cloud sequence. First, model the unordered spatial structure to obtain the spatial local structure of the point cloud in 3D space; then perform temporal convolution to obtain the spatio-temporal feature representation of the point cloud sequence; The multi-scale spatio-temporal convolution formula is as follows: where \((x, y, z)\in P\) t and \((\delta\) x , \(\delta\) y , \(\delta\) z ) represents displacement, \(r\) i \(\in R\) represents the spatial search radius, represents the spatial convolution kernel, represents the temporal convolution kernel, \(l\) represents the size of the temporal convolution kernel, \(C\) m represents the dimension of the intermediate feature, \(M\) t represents the sequence feature after spatial convolution, represents the sequence feature after temporal convolution.

6. The multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 5, wherein: In step S1.4, use relative coordinates in space to generate a convolution kernel, convert the kernel into a displacement function, and the updated spatial convolution formula is as follows: Among them, represents the function of (δ x , δ y , δ z ), parameterized by θ, and generates different spatial convolution kernels according to different displacements ; To further reduce the computational complexity, the function f is decomposed into f((δ x ,δ y ,δ z );θ)=θ d ∙(δ x ,δ y ,δ z ) T ∙1∙⊙θ s , where θ = [θ d ,θ s , represents the displacement transformation kernel, represents the shared kernel, 1 = (1,..., 1) ∈ R 1×C represents the broadcast kernel, and ⊙ represents the element-wise product.

7. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 1, characterized in that: The specific steps of step S2 include the following sub-steps: S2.

1. After the point cloud sequence undergoes multi-scale spatio-temporal convolution operations, apply the discrete cosine transform in the time dimension to obtain frequency information and get multi-frequency vectors; S2.

2. Apply the attention mechanism to the multi-frequency vectors in the entire frequency channel to obtain weighted multi-frequency features.

8. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 7, characterized in that: In step S2.1, use the discrete cosine transform for frequency domain feature decomposition, and the formula of the discrete cosine transform is as follows: where \(f\in\mathbb{R}\) H represents the spectrum of the discrete cosine transform, \(h\in\{0,1,\ldots,H - 1\}\); \(x\in\mathbb{R}\) H represents the input data, \(H\) represents the length of the input component, and the cosine function represents the weight; To obtain all frequency components, perform the inverse transform on the discrete cosine transform formula, and the formula is as follows: where, i ∈ {0, 1,..., H - 1}; To fuse the frequency components, divide the input X into n parts according to the channel dimension, and assign corresponding discrete cosine transform frequency components to each part, and the formula is as follows: where \(i\in\{0,1,\cdots,n - 1\}\), GAP represents the average pooling operation in the spatial dimension \(w\) direction, and Freq i \(\in\mathbb{R}\) C' represents the compressed \(C'\)-dimensional vector; the entire compressed vector is obtained by concatenation: where Freq ∈ R c represents the obtained multi-frequency vector.

9. The multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 7, characterized in that: In step S2.2, apply the attention mechanism to the obtained multi-frequency vectors in the frequency channel to perform weighted processing on the entire feature channel, and the formula representation is as follows: where, freq_att represents the multi-frequency vector after weighted processing.

10. A multi-scale spatio-temporal frequency behavior recognition method based on a point cloud sequence according to claim 1, characterized in that: In step S3, use average pooling and max pooling for spatial pooling and temporal pooling respectively to eliminate the spatial and time dimensions, and only retain the feature dimension for classification processing.