Multi-scale feature enhanced point cloud data dynamic gesture recognition method and system

By decoupling point cloud data into spatial and dynamic flows, and combining attention mechanisms and residual networks, extracting and integrating multi-scale features of gestures, the problem of low gesture recognition accuracy in the prior art is solved, and higher recognition accuracy and robustness are achieved.

CN120183044APending Publication Date: 2025-06-20CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326666.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing gesture recognition technology has low accuracy when processing complex spatiotemporal data, making it difficult to capture the overall motion pattern and subtle posture changes of the hand at the same time, resulting in poor recognition effect.

Method used

The dynamic gesture recognition method of point cloud data is adopted for multi-scale feature enhancement. By decoupling point cloud data into spatial flow and dynamic flow, combining attention mechanism and residual network, macroscopic spatial information and microscopic action features are extracted, and these features are integrated through a bidirectional feature fusion mechanism.

Benefits of technology

It improves the accuracy and robustness of gesture recognition, can handle viewing angle changes, lighting conditions and background interference more effectively, and enhances the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183044A_ABST
    Figure CN120183044A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale feature enhanced point cloud data dynamic gesture recognition method and system, and relates to the technical field of gesture recognition. According to the method, multi-scale information of spatial features and dynamic features is combined, so that the model can capture dynamic changes of a global macroscopic spatial structure and microscopic details at the same time, and complex information in gesture recognition can be better understood and processed by introducing a multi-scale feature enhancement network, so that the recognition precision is improved; according to the method, the point cloud data are decoupled into the spatial stream and the dynamic stream, global shape information and local action features are respectively caught, redundant information can be effectively reduced, the characterization capability of the features can be improved, the spatial distribution characteristic expression of the point cloud data can be further improved by combining the BPS technology and the residual network, and the method has the advantages of being high in practicability and the like. The feature extraction of the point cloud is more accurate; and a key point detection module based on an attention mechanism is designed in the dynamic path, so that the model can concentrate on important local features in the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gesture recognition, and specifically to a multi-scale feature enhanced point cloud data dynamic gesture recognition method and system. Background Art

[0002] Gesture recognition technology, as the core means to achieve human-computer interaction, captures and analyzes hand movements. It has shown broad application prospects in fields such as virtual reality, intelligent device control, and game entertainment. Traditional methods mainly rely on continuous image frame analysis or joint position tracking, and there are problems such as low accuracy and poor gesture recognition effect in gesture recognition in the scenario of continuous gesture actions. Gesture recognition technology based on 3D point cloud has the advantages of being able to effectively process complex spatio-temporal data, accurately capture the hand surface and its geometric structure, and provide rich spatial information.

[0003] The idea of the multi-scale feature enhanced network method is a feature extraction framework based on a two-stream architecture. Among them, the dynamic feature path designs a key point detection module based on the attention mechanism, dynamically extracts fine geometric features from the original point cloud, and combines the key point cloud data with the original point cloud network to analyze the dynamic features of the hand. The other spatial feature path effectively combines the ResNet50 network with the residual base point set (BPS) for encoding, focusing on capturing the overall motion pattern of the hand. By introducing the bidirectional feature fusion technology, the macroscopic spatial information and microscopic motion features are effectively integrated, thereby realizing effective modeling in the time dimension and effectively improving the accuracy of gesture recognition.

[0004] Currently, gesture recognition based on graphics and skeletons has relatively high accuracy, but it also faces challenges in geometric feature extraction: at the macroscopic level, the overall displacement and motion trajectory of the hand constitute a large-scale spatial pattern; at the microscopic level, the subtle posture changes of the fingers determine the semantic connotation of the gesture. This multi-scale characteristic requires the ability to handle large-scale motion patterns and capture fine action details simultaneously, increasing the complexity of technical implementation. Therefore, how to ensure high timeliness while accurately extracting the feature information of dynamic hands and accurately recognizing gesture information is a problem that needs to be solved by those skilled in the art.

[0005] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a multi-scale feature enhanced point cloud data dynamic gesture recognition method and system. The point cloud data not only helps to understand the overall posture of the hand, but also can accurately track the subtle changes of the fingers, and has stronger robustness to perspective changes, lighting conditions, and background interference, thereby greatly improving the recognition accuracy and system stability, so as to solve the technical problems proposed in the background art.

[0007] To achieve the above object, the present invention provides the following technical solutions: A multi-scale feature enhanced point cloud data dynamic gesture recognition method, which at least includes the following steps:

[0008] S1: Convert the depth map into a 3D point cloud sequence, decouple the input 3D point cloud sequence into a spatial stream and a dynamic stream, and focus on feature information at different scales;

[0009] S2: Design a key point detection module based on the attention mechanism in the dynamic path to enhance the model's ability to understand the internal detailed structure of the point cloud. Designing the key point detection module can effectively capture the long-range dependencies in the point cloud data and obtain the dynamic features of the hand. That is, a key point detection module based on the self-attention mechanism is designed in the local stream, which is more conducive to the extraction of local features;

[0010] S3: Effectively combine the residual network (ResNet50 network) with the basic point set (BPS) for encoding in the spatial path, combine the basic point set technology to process the point cloud data, and use the residual network structure for efficient feature extraction;

[0011] S4: Propose a bidirectional feature fusion mechanism to effectively integrate the macroscopic spatial information and microscopic motion features, utilize the complementary and associated characteristics of the spatial features and dynamic features, and avoid the problems of insufficient fusion and feature conflict caused by simple element-level addition, thereby realizing dynamic gesture recognition.

[0012] Further, S1 at least includes the following steps:

[0013] S1.1: Select 32 key frames from the original gesture video sequence F = {f1, f2,..., f N}, for each key frame f i Read the depth image and crop the hand area. In use, median filtering is used to reduce noise, and the point cloud production function is used to generate point cloud data. The point cloud production function is used to generate point cloud data. Convert the two-dimensional depth image into a three-dimensional point cloud representation, so as to more intuitively display the shape of the hand and its positional relationship in space. Then convert the point cloud data into the XYZ coordinate system, and the conversion formula is:

[0014] Z = d uv

[0015]

[0016] where u and v are the horizontal and vertical pixel coordinates of the target point on the depth image respectively; Z is the actual distance of the point relative to the camera plane, that is, the depth; f x and f y are the focal lengths of the camera in the x-axis and y-axis directions respectively; c x and c yare the coordinates of the center point of the image plane, also known as the principal point;

[0017] S1.2: Decouple into spatial streams. The spatial stream (glob) emphasizes the overall structure and shape features of the entire gesture sequence and performs centering by removing the average position of the entire sequence;

[0018] S1.3: Decouple into dynamic streams (loc). The dynamic stream focuses on local features within each frame and performs centering by removing the respective average position of each frame to capture details and changes in actions within the frame.

[0019] Furthermore, S2 includes at least the following steps:

[0020] S2.1: Extract local features and apply batch normalization and attention modules;

[0021] Extract local features from the input point cloud data. Each layer is followed by a batch normalization layer to accelerate the training process and stabilize the gradients, and an attention module is applied after the last convolution;

[0022] Map the input features to a common space through a linear transformation, and split the mapped features into multiple heads, with each head responsible for a part of the information;

[0023] S2.2: Calculate attention scores using scaled dot-product attention and apply the softmax function for normalization;

[0024] Finally, perform a weighted sum of the values according to the attention scores and output a fully connected layer through an additional linear transformation for further processing the output of the attention layer and finally outputting the embedding vector;

[0025] The attention mechanism module Attention enhances the data query feature representation through the interaction between the query Q, key K, and value V;

[0026] The query Q is obtained by applying a linear transformation to the features after global max pooling:

[0027] Q = fc q (max(x, dim = 2))

[0028] where max(x, dim = 2) represents global max pooling along the point dimension;

[0029] The key K and value V are directly obtained by applying a linear transformation to the features x after transposition:

[0030] K = fc k (x T )

[0031] V = fc v (xT )

[0032] Among them, x T is the transpose of the feature matrix, and the key K and value V are obtained through linear transformation;

[0033] S2.3: For each head, the attention weight Ah is calculated as follows:

[0034] Q h = Q (h) , K h = K (h) , V h = V (h)

[0035]

[0036] Among them, the total attention dimension D is divided into H heads, and the dimension of each head is D' = D / H;

[0037] S2.4: Weighted summation and output calculation:

[0038] The output O of each head h is calculated according to the following formula:

[0039] O h = fc o (Q h + A h V h )

[0040] The final output O is the concatenation result of the outputs of all heads: O = Concat(O1,..., O H ).

[0041] Furthermore, S3 for efficient feature extraction includes at least the following steps:

[0042] Using the method of the basic point set to convert the input point cloud into a set of predefined basic points distance feature representation;

[0043] By calculating the minimum distance from each input point to each point in the basic point set, the spatial distribution characteristics of the point cloud are captured, thereby effectively reducing the dimension of the original point cloud data, and extracting the distance to the closest basic point as the feature vector;

[0044] Then, feature enhancement is performed through a series of residual blocks, while retaining the key spatial information.

[0045] Furthermore, by combining the residual network and the basic point set (BPS) as an alternative representation form, the spatial feature expression ability of the original point cloud is enhanced;

[0046] Calculate the squared pairwise Euclidean distance between two point clouds, which is calculated as follows:

[0047] dist[i,j] = ||x[i,:] - y[j,:]|| 2

[0048] Where, given two matrices x and y, where x is an N×d matrix and y is an M×d matrix, the return value is an N×M matrix. Both N and M represent the number of rows of the matrix, that is, the number of samples contained in the matrix, and d represents the number of columns of the matrix, that is, the number of feature dimensions of each sample or point;

[0049] Each element dist[i,j] represents the squared Euclidean distance between x[i,:] and y[j,:];

[0050] The BPS technology can effectively capture the spatial distribution characteristics of the point cloud by calculating the minimum distance between the input point cloud and a set of predefined basis points, while the residual network ensures the effective training of the deep model;

[0051] Calculate the pairwise distance between the point cloud and the basis point set, and the correlation calculation formula is:

[0052] dist = pairwise_distances(x, basis)

[0053]

[0054] Combines the BPS technology to process point cloud data and uses the residual network structure for global feature extraction.

[0055] Furthermore, S4 at least includes the following steps:

[0056] Through two independent spatial dynamic stream outputs, the output spatial features and dynamic feature data are divided into two parts F1 and F2, and each feature data is processed by an independent bidirectional LSTM, as shown in the following formula:

[0057] H1 = BiLSTM(F1)

[0058] H2 = BiLSTM(F2)

[0059] Where, BiLSTM represents the bidirectional long short-term memory network;

[0060] Interpolation combines global features and local features, comprehensively captures the temporal dynamic characteristics of the point cloud data, and integrates this information through the feature fusion layer to generate a richer feature representation;

[0061] Feature fusion: Through the feature fusion layer, the outputs H1 and H2 of two LSTMs are fused to form a unified feature representation, as shown in the following formula:

[0062] F fused =[H1; H2]

[0063] To comprehensively capture the temporal dynamic characteristics of the point cloud data, interpolation processing is further performed on the fused features. The interpolation operation can balance between global features and local features and integrate different temporal information in the data after feature fusion. The specific interpolation calculation formula is:

[0064] F interpolated =αF fused [:,:,:2H]+(1-α)F fused [:,:,2H:]

[0065] where: α is the interpolation weight.

[0066] Furthermore, the design of combining the encoder-decoder architecture and the self-attention mechanism in the local flow is as follows:

[0067] The encoder can effectively capture the long-range dependencies in the point cloud through the multi-layer self-attention mechanism. The decoder uses the "partition query" vector to divide the point cloud into multiple meaningful subsets and generates a consistent partition structure through cross-attention;

[0068] Through the "partition exchange" operation, the problem of geometric distortion is successfully avoided, and the efficient circulation of information of similar point clouds is promoted;

[0069] Finally, through the spatial max pooling operation, the information of all points is aggregated into a comprehensive feature descriptor, represented in the form of a single vector, providing a compact and rich feature set for subsequent processing and enhancing the 3D point cloud data.

[0070] A system for dynamic gesture recognition of multi-scale feature-enhanced point cloud data at least includes an acquisition module, a spatial feature extraction module, a dynamic feature extraction module, a bidirectional feature fusion module, and a temporal modeling and classification module

[0071] The acquisition module is used to select key frames from the gesture video sequence to generate point cloud data;

[0072] The spatial feature extraction module is used to capture the extensive spatial connections in the point cloud data;

[0073] The dynamic feature extraction module is used to identify the detailed geometric features of hand movements;

[0074] The bidirectional feature fusion module is to optimize the fusion mechanism between spatial features and dynamic features;

[0075] The time modeling and classification module combines the powerful time series modeling ability of LSTM to efficiently capture time dynamic information for gesture recognition and classification.

[0076] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0077] 1. The method of the present invention combines multi-scale information of spatial features and dynamic features, enabling the model to simultaneously capture the global macroscopic spatial structure and the dynamic changes of microscopic details. By introducing a multi-scale feature enhancement network, it can better understand and process complex information in gesture recognition, thereby improving the recognition accuracy.

[0078] 2. By decoupling point cloud data into a spatial flow and a dynamic flow, the present invention focuses on capturing global shape information and local action features respectively, which can effectively reduce redundant information and enhance the feature representation ability. Combining the BPS technology and the residual network can further enhance the expression of the spatial distribution characteristics of point cloud data, making the feature extraction of point cloud more accurate.

[0079] 3. The present invention designs a key point detection module based on the attention mechanism in the dynamic path, enabling the model to focus on important local features in the point cloud, enhancing the understanding ability of point cloud details, which helps to handle subtle changes in gesture actions and improve the robustness of recognition.

[0080] 4. The present invention introduces a bidirectional feature fusion mechanism, which can effectively integrate spatial features and dynamic features, avoiding feature conflicts and information loss caused by simple element-level addition. By processing the spatial flow and the dynamic flow through a bidirectional LSTM module, the effect of feature fusion is further optimized, ensuring more sufficient fusion of spatial and dynamic information.

[0081] 5. By combining with the LSTM network for time series modeling, the present invention can effectively capture the time dynamic information in point cloud data, thereby enhancing the modeling ability of the time series dependence relationship in dynamic gesture recognition. In this way, the model can consider time context information when processing gesture actions and provide more accurate classification results.

[0082] 6. Reducing computational complexity and redundant information, the present invention uses the BPS technology to reduce the dimension of the original point cloud data. By calculating the minimum distance from each input point to the basic point set, the computational amount is effectively reduced, and the feature representation is further enhanced through the structure of the residual network, making the feature extraction more efficient.

[0083] 7. By combining the partition query vector and the cross-attention mechanism, the present invention effectively avoids the common geometric distortion problems in the process of point cloud data processing, promotes the efficient circulation of similar point cloud information, and ensures the accuracy and integrity of the geometric structure of point cloud data.

[0084] 8. The present invention pools the information of all points into a comprehensive feature descriptor through spatial max pooling operation, making the final output feature have a more compact and informative representation, which is conducive to subsequent processing and classification, and enhances the performance of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0086] Figure 1 is a schematic flow chart of the whole of the present invention;

[0087] Figure 2 is a gesture point cloud map generated by the present invention;

[0088] Figure 3 is a structural diagram of the gesture recognition system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0090] Embodiment 1:

[0091] Please refer to Figure 1 , a multi-scale feature enhanced point cloud data dynamic gesture recognition method, at least including the following steps:

[0092] S1: Convert the depth map into a 3D point cloud sequence, decouple the input 3D point cloud sequence into a spatial stream and a dynamic stream, and focus on feature information at different scales;

[0093] S2: Design a key point detection module based on the attention mechanism in the dynamic path to enhance the model's understanding ability of the internal detailed structure of the point cloud. Designing the key point detection module can effectively capture the long-distance dependence relationship in the point cloud data and obtain the dynamic features of the hand. That is, a key point detection module based on the self-attention mechanism is designed in the local stream, which is more conducive to the extraction of local features;

[0094] S3: Effectively combine the residual network (ResNet50 network) with the basic point set (BPS) for encoding in the spatial path, combine the basic point set technology to process the point cloud data, and use the residual network structure for efficient feature extraction;

[0095] S4: A two-way feature fusion mechanism is proposed, which effectively integrates macroscopic spatial information and microscopic motion features. By utilizing the complementary and correlative characteristics of spatial features and dynamic features, it avoids the problems of insufficient fusion and feature conflict caused by simple element-level addition, thereby achieving dynamic gesture recognition.

[0096] As Figure 2 shown, S1 first selects key frames and then generates point clouds, including at least the following steps:

[0097] S1.1: Select 32 key frames from the original gesture video sequence F = {f1, f2,..., f N}. For each key frame f i read the depth image and crop the hand region. During use, median filtering is used to reduce noise, and the point cloud production function is used to generate point cloud data. The two-dimensional depth image is converted into a three-dimensional point cloud representation to more intuitively display the shape of the hand and its positional relationship in space. Then, the point cloud data is converted into the XYZ coordinate system, and the conversion formula is:

[0098] Z = d uv

[0099]

[0100] where u and v are the horizontal and vertical pixel coordinates of the target point on the depth image respectively; Z is the actual distance of the point relative to the camera plane, i.e., the depth; f x and f y are the focal lengths of the camera in the x-axis and y-axis directions respectively; c x and c y are the coordinates of the center point of the image plane, also known as the principal point;

[0101] S1.2: Decouple it into a spatial flow. The spatial flow (glob) emphasizes the overall structure and shape features of the entire gesture sequence, and performs centering processing by removing the average position of the entire sequence;

[0102] Among them, for the local point cloud, the correlation calculation formula is:

[0103]

[0104] where pcd t,n represents the point cloud data (position or coordinate) at the t-th frame (time point) and the n-th point. The point cloud data is usually composed of a set of three-dimensional points, representing the surface of an object or the environment at a certain time point or in space;

[0105] Calculate the average position of the n-th point in all point clouds. At each frame t, the coordinates of all N points are averaged to obtain the average position of the entire point cloud;

[0106] loc t,n represents the point cloud data after centering processing, which is obtained by subtracting the average position of all points in the frame from the coordinates of each point, that is, "decentralizing" by removing the translational information of the point cloud, making the point cloud align with its average position, so as to eliminate the global translational transformation and better capture the shape and local features;

[0107] S1.3: Decouple into the dynamic flow (loc), and the dynamic flow focuses on the local features within each frame. Centering is performed by removing the respective average position of each frame to capture the changes in details and actions within the frame.

[0108] Among them, for the global point cloud, the correlation calculation formula is:

[0109]

[0110] Among them, represents the average of the coordinates of all time frames and point clouds to obtain the global average position of the entire dataset. Specifically, this is to calculate the average position of all point clouds over all frames (t = 1 to t = T) and all points (n = 1 to n = N);

[0111] glob t,n represents the point cloud data after centering processing, removing the average position of the entire dataset. This is to perform decentralization processing on the point cloud in each frame to eliminate the global translational offset. Through this processing, the point cloud data pays more attention to local changes rather than being affected by translational transformations.

[0112] S2 includes at least the following steps:

[0113] S2.1: Extract local features and apply batch normalization and attention modules;

[0114] Extract local features from the input point cloud data. Each layer is followed by a batch normalization layer to accelerate the training process and stabilize the gradients, and an attention module is applied after the last convolution;

[0115] Map the input features to a common space through a linear transformation, and split these mapped features into multiple heads, with each head responsible for a part of the information;

[0116] S2.2: Calculate attention scores using scaled dot-product attention and apply the softmax function for normalization;

[0117] Finally, the values are weighted and summed according to the attention scores, and a fully connected output is obtained through an additional linear transformation, which is used to further process the output of the attention layer and finally output the embedding vector;

[0118] The attention mechanism module Attention enhances the data query feature representation through the interaction between the query Q, the key K, and the value V;

[0119] The query Q is obtained by applying a linear transformation to the features after global max pooling:

[0120] Q = fc q (max(x, dim = 2))

[0121] where max(x, dim = 2) represents global max pooling along the point dimension;

[0122] The key K and the value V are directly obtained by applying a linear transformation to the transposed features x:

[0123] K = fc k (x T )

[0124] V = fc v (x T )

[0125] where x T is the transpose of the feature matrix, and the key K and the value V are obtained through linear transformation;

[0126] S2.3: For each head, the attention weight Ah is calculated as follows:

[0127] Q h = Q (h) , K h = K (h) , V h = V (h)

[0128]

[0129] where the total attention dimension D is divided into H heads, and the dimension of each head is D' = D / H;

[0130] S2.4: Weighted sum and output calculation:

[0131] The output O h of each head is calculated according to the following formula:

[0132] O h = fc o (Q h + A h V h )

[0133] The final output O is the concatenation of all head outputs: O = Concat(O1,..., O H ).

[0134] Performing efficient feature extraction in S3 includes at least the following steps:

[0135] Using the method of the base point set to convert the input point cloud into a set of predefined base points for distance feature representation;

[0136] By calculating the minimum distance from each input point to each point in the base point set, the spatial distribution characteristics of the point cloud are captured, effectively reducing the dimension of the original point cloud data, and extracting the distance to the closest base point as the feature vector;

[0137] Then, feature enhancement is performed through a series of residual blocks while retaining the key spatial information.

[0138] Combining the residual network and the base point set (BPS), an alternative representation form, enhances the spatial feature expression ability of the original point cloud;

[0139] Calculating the squared pairwise Euclidean distance between two point clouds is calculated as follows:

[0140] dist[i,j] = ||x · [i,:] - y[j,:]|| 2

[0141] where, given two matrices x and y, where x is an N×d matrix and y is an M×d matrix, the return value is an N×M matrix, N and M both represent the number of rows of the matrix, that is, the number of samples contained in the matrix, and d represents the number of columns of the matrix, that is, the number of feature dimensions of each sample or point;

[0142] where each element dist[i,j] represents the squared Euclidean distance between x[i,:] and y[j,:];

[0143] The BPS technology can effectively capture the spatial distribution characteristics of the point cloud by calculating the minimum distance between the input point cloud and a set of predefined base points, while the residual network ensures the effective training of the deep model;

[0144] Calculating the pairwise distance between the point cloud and the base point set, the correlation calculation formula is:

[0145] dist = pairwise_distances(x, basis)

[0146]

[0147] Combines the BPS technology to process point cloud data and uses the residual network structure for global feature extraction.

[0148] S4 includes at least the following steps:

[0149] Through two independent spatial dynamic stream outputs, the output spatial features and dynamic feature data are divided into two parts F1 and F2, and each feature data is processed by an independent bidirectional LSTM respectively, as shown in the following formula:

[0150] H1 = BiLSTM(F1)

[0151] H2 = BiLSTM(F2)

[0152] Where, BiLSTM represents the bidirectional long short-term memory network;

[0153] Interpolation combines global features and local features, comprehensively captures the temporal dynamic characteristics of point cloud data, and integrates this information through a feature fusion layer to generate a richer feature representation;

[0154] Feature fusion: Through the feature fusion layer, the outputs H1 and H2 of the two LSTMs are fused to form a unified feature representation, as shown in the following formula:

[0155] F fused = [H1; H2]

[0156] In order to comprehensively capture the temporal dynamic characteristics of point cloud data, further interpolation processing is performed on the fused features. The interpolation operation can balance between global features and local features and integrate different temporal information in the data after feature fusion. The specific interpolation calculation formula is:

[0157] F interpolated = αF fused [:,:,:2H] + (1 - α)F fused [:,:,2H:]

[0158] Where: α is the interpolation weight.

[0159] The design of combining the encoder-decoder architecture and the self-attention mechanism in the local flow is as follows:

[0160] The encoder can effectively capture the long-range dependencies in the point cloud through a multi-layer self-attention mechanism. The decoder uses the "partition query" vector to divide the point cloud into multiple meaningful subsets and generates a coordinated partition structure through cross-attention;

[0161] Through the "partition exchange" operation, the problem of geometric distortion is successfully avoided and the efficient circulation of similar point cloud information is promoted;

[0162] Finally, through the spatial max pooling operation, the information of all points is aggregated into a comprehensive feature descriptor, which is represented in the form of a single vector, providing a compact and rich feature set for subsequent processing and enhancing the 3D point cloud data.

[0163] Based on the above content, the following verification is proposed:

[0164] The accuracy rate is verified by comparison on the open-source dataset SCHREC17. Table 1 shows the comparison results of the accuracy rates of the method of the present invention and other 4 advanced methods under different attributes.

[0165] As can be seen from Table 1, the multi-scale feature enhancement network method for dynamic gesture recognition based on point cloud data described in the present invention has an accuracy rate of 97.2 for 14-class classification and 95.6 for 28-class classification on the public dataset SCHREC17. And compared with other advanced recognition methods, point cloud data is usually better than skeleton data. This is because skeleton data mainly provides a sparse representation of the key joints of the hand, while point cloud can completely cover the hand surface and capture more details, thus achieving higher accuracy and better results in gesture recognition and having strong competitiveness.

[0166] Table 1 Accuracy Rates of 14G and 28G on the SCHREC17 Dataset

[0167]

[0168] Example 2:

[0169] Based on the above method, this embodiment proposes a specific dynamic gesture recognition system based on point cloud data, as Figure 3 shown, which at least includes an acquisition module, a spatial feature extraction module, a dynamic feature extraction module, a bidirectional feature fusion module, and a temporal modeling and classification module

[0170] The acquisition module is used to select key frames from the gesture video sequence and generate point cloud data;

[0171] The spatial feature extraction module is used to capture the extensive spatial connections in the point cloud data;

[0172] The dynamic feature extraction module is used to identify the detailed geometric features of hand movements;

[0173] The bidirectional feature fusion module is to optimize the fusion mechanism between spatial features and dynamic features;

[0174] The temporal modeling and classification module combines the powerful temporal modeling ability of LSTM to efficiently capture temporal dynamic information and perform gesture recognition and classification.

[0175] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data, characterized in that: At least the following steps are included: S1: Convert the depth map into a 3D point cloud sequence, decouple the input 3D point cloud sequence into spatial flow and dynamic flow, and focus on feature information of different scales; S2: Design a key point detection module based on the attention mechanism in the dynamic path to enhance the model's ability to understand the detailed structure of the point cloud. The key point detection module can effectively capture the long-distance dependencies in the point cloud data and obtain the dynamic features of the hand. That is, a key point detection module based on the self-attention mechanism is designed in the local flow, which is more conducive to the extraction of local features. S3: The residual network is effectively combined with the basic point set for encoding in the spatial path, the basic point set technology is combined to process the point cloud data, and the residual network structure is used for efficient feature extraction; S4: A bidirectional feature fusion mechanism is proposed, which effectively integrates macroscopic spatial information and microscopic motion features. It utilizes the complementary and correlation characteristics of spatial features and dynamic features to avoid the problems of insufficient fusion and feature conflict caused by simple element-level addition, thereby realizing dynamic gesture recognition.

2. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 1, characterized in that: S1 includes at least the following steps: S1.1: From the original gesture video sequence F = {f1,f2,...,f N }Select 32 key frames, and for each key frame f i Read the depth image and crop the hand area, use median filtering to reduce noise, use the point cloud production function to generate point cloud data, use the point cloud production function to generate point cloud data, convert the two-dimensional depth image into a three-dimensional point cloud representation, so as to more intuitively show the shape of the hand and its position in space, and then convert the point cloud data into the XYZ coordinate system. The conversion formula is: Z=d uv Where u and v are the horizontal and vertical pixel coordinates of the target point on the depth image, respectively; Z is the actual distance of the point relative to the camera plane, i.e., the depth; f x and f y are the focal lengths of the camera in the x-axis and y-axis directions respectively; c x and c y are the coordinates of the center point of the image plane, also called the principal point; S1.2: decoupling into spatial streams, which emphasizes the overall structure and shape characteristics of the entire gesture sequence and performs central processing by removing the average position of the entire sequence; S1.3: Decoupled into dynamic flow, the dynamic flow focuses on the local features within each frame and is centralized by removing the average position of each frame to capture changes in details and movements within the frame.

3. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 1, characterized in that: S2 includes at least the following steps: S2.1: Extract local features and apply batch normalization and attention modules; Local features will be extracted from the input point cloud data. Each layer is followed by a batch normalization layer to speed up the training process and stabilize the gradient. An attention module is applied after the last convolution layer. Map the input features to a common space through linear transformation, and split these mapped features into multiple heads, each head is responsible for a part of the information; S2.2: Calculate the attention score using scaled dot product attention and apply the softmax function for normalization; Finally, the values ​​are weighted and summed according to the attention score and fully connected through an additional linear transformation output for further processing the output of the attention layer and finally outputting the embedding vector; The attention mechanism module Attention enhances the data query feature representation through the interaction between query Q, key K and value V; The query Q is obtained by applying a linear transformation to the features after global maximum pooling: Q=fc q (max(x,dim=2)) Among them, max(x, dim=2) means global maximum pooling along the point dimension; The key K and value V are directly obtained from the feature x by transposing it and applying a linear transformation: K=fc k (x T ) V=fc v (x T ) Among them, x T It is the transpose of the feature matrix, and the key K and value V are obtained through linear transformation; S2.3: For each head, the attention weight Ah is calculated as follows: Q h =Q (h) ,K h =K (h) ,V h =V (h) The total dimension D of attention is divided into H heads, and the dimension of each head is D′=D / H; S2.4: Weighted summation and output calculation: The output of each head O h The calculation of is as follows: About h =fc o (Q h +A h In h ) The final output O is the concatenation of all the header outputs: O = Concat(O1,...,O H ).

4. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 1, characterized in that: S3 performs efficient feature extraction by at least the following steps: Using the basic point set method, the input point cloud Convert to a set of predefined base points The distance feature representation of The spatial distribution characteristics of the point cloud are captured by calculating the minimum distance from each input point to each point in the basic point set, thereby effectively reducing the dimension of the original point cloud data and extracting the distance to the closest basic point as the feature vector; Feature enhancement is then performed through a series of residual blocks while retaining key spatial information.

5. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 4, characterized in that: Combining the residual network with the base point set as an alternative representation form, the spatial feature expression capability of the original point cloud is enhanced; Calculate the squared pairwise Euclidean distance between two point clouds as follows: dist[i,j]=||x · [i,:]-y[j,:]|| 2 Given two matrices x and y, where x is an N×d matrix and y is an M×d matrix, the return value is an N×M matrix, where N and M represent the number of rows in the matrix, i.e. the number of samples contained in the matrix, and d represents the number of columns in the matrix, i.e. the number of feature dimensions of each sample or point; Where each element dist[i,j] represents the squared Euclidean distance between x[i,:] and y[j,:]; BPS technology can effectively capture the spatial distribution characteristics of point clouds by calculating the minimum distance between the input point cloud and a set of predefined basic points, while the residual network ensures the effective training of the deep model; Calculate the pairwise distance between the point cloud and the base point set. The correlation calculation formula is: dist=pairwise_distances(x,basis) The BPS technology is combined to process point cloud data, and the residual network structure is used to extract global features.

6. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 1, characterized in that: S4 at least includes the following steps: Through two independent spatial dynamic stream outputs, the output spatial feature and dynamic feature data are divided into two parts F1 and F2. Each feature data is processed by an independent bidirectional LSTM, see the following formula: H1=BiLSTM(F1) H2=BiLSTM(F2) Among them, BiLSTM represents a bidirectional long short-term memory network; Interpolation combines global features and local features to comprehensively capture the temporal dynamic characteristics of point cloud data, and integrates this information through the feature fusion layer to generate richer feature representations; Feature fusion: Through the feature fusion layer, the outputs H1 and H2 of the two LSTMs are fused to form a unified feature representation, see the following formula: F fused =[H1;H2] In order to comprehensively capture the temporal dynamic characteristics of point cloud data, the fused features are further interpolated. The interpolation operation can balance the global features and local features, and integrate different temporal information in the data after feature fusion. The specific interpolation calculation formula is: F interpolated =αF fused [:,:,:2H]+(1-α)F fused [:,:,2H:] Where: α is the interpolation weight.

7. A method for dynamic gesture recognition using multi-scale feature-enhanced point cloud data according to claim 1, characterized in that: The design of combining the encoder-decoder architecture and the self-attention mechanism in the local stream is as follows: The encoder can effectively capture long-distance dependencies in the point cloud through a multi-layer self-attention mechanism, while the decoder uses the "partition query" vector to divide the point cloud into multiple meaningful subsets and generates a coordinated partition structure through cross-attention; Through the "partition exchange" operation, the problem of geometric distortion is successfully avoided and the efficient circulation of similar point cloud information is promoted; Finally, the information of all points is aggregated into a comprehensive feature descriptor through the spatial maximum pooling operation and represented in the form of a single vector, which provides a compact and rich feature set for subsequent processing and enhances the 3D point cloud data.

8. A system for dynamic gesture recognition using multi-scale feature enhanced point cloud data, used in a method for dynamic gesture recognition using multi-scale feature enhanced point cloud data according to any one of claims 1 to 7, characterized in that: At least includes acquisition module, spatial feature extraction module, dynamic feature extraction module, bidirectional feature fusion module and time modeling and classification module The acquisition module is used to select key frames from the gesture video sequence and generate point cloud data; The spatial feature extraction module is used to capture the extensive spatial connections in point cloud data; The dynamic feature extraction module is used to identify the detailed geometric features of hand movements; The bidirectional feature fusion module is a fusion mechanism to optimize spatial features and dynamic features; The temporal modeling and classification module combines the powerful temporal modeling capabilities of LSTM to efficiently capture temporal dynamic information and perform gesture recognition and classification.