Gesture recognition method and device and storage medium
By employing multi-view point cloud fusion and feature processing methods, the problems of external interference and hand self-occlusion in gesture recognition technology are solved, achieving low-latency and high-precision gesture recognition, which is suitable for gesture interaction in virtual reality and augmented reality technologies.
Patent Information
- Application Number
- CN202511105268.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing gesture recognition technologies are easily affected by external factors, resulting in poor timeliness and low recognition accuracy. Furthermore, traditional single-view acquisition is prone to hand self-occlusion problems, making it difficult to simultaneously meet the requirements of low latency, high accuracy, low equipment cost, and high reliability.
By acquiring depth image data from different acquisition perspectives through multiple preset sensors, performing point cloud fusion processing, and combining feature extraction and feature fusion, comprehensive acquisition of gesture data and accurate reconstruction of complete 3D structure are achieved.
It achieves low-latency, high-precision gesture recognition, overcomes the problem of hand self-occlusion, improves the continuity and reliability of recognition, and adapts to the flexibility of different application scenarios.
Smart Images

Figure CN120997880A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, and in particular to a gesture recognition method, device and storage medium. BACKGROUND
[0002] With the development of virtual reality and augmented reality technology, gesture capture as an important technology of human-computer interaction has been widely used in film and television production, industrial control, medical training and other fields.
[0003] At present, the capture of dynamic gestures mainly includes three categories: wearable sensing, optical vision capture and marker-based motion capture.
[0004] However, these technical solutions are easily disturbed by external factors, which leads to the problem that it is difficult to simultaneously meet the requirements of low delay, high precision, low cost, high reliability and high mobile portability of gesture recognition. SUMMARY
[0005] The present application provides a gesture recognition method, device and storage medium to solve the problem that related technologies are easily disturbed by external factors, leading to poor timeliness and low recognition accuracy.
[0006] In a first aspect, the present application provides a gesture recognition method, comprising:
[0007] obtaining a plurality of depth image data of a gesture to be recognized; wherein the plurality of depth image data is derived from different collection angles;
[0008] performing first preprocessing on the depth image data to obtain corresponding first preprocessed point cloud data;
[0009] performing point cloud fusion processing according to the first preprocessed point cloud data to obtain fused point cloud data;
[0010] performing second preprocessing on the fused point cloud data to obtain second preprocessed point cloud data;
[0011] performing feature extraction processing and feature fusion processing according to the second preprocessed point cloud data to obtain a recognition result of the gesture to be recognized.
[0012] In a possible implementation, the obtaining of the plurality of depth image data of the gesture to be recognized comprises: collecting, by a plurality of preset sensors, the plurality of depth image data of the gesture to be recognized; wherein the plurality of preset sensors are used to cover a preset space around the gesture to be recognized.
[0013] In a possible implementation, the preset sensors include a first preset sensor, a second preset sensor, and a third preset sensor; correspondingly, the collecting, by the plurality of preset sensors, of the plurality of depth image data of the to-be-recognized gesture includes: collecting, by the first preset sensor, top-view depth image data of the to-be-recognized gesture; collecting, by the second preset sensor, left-lower-view depth image data of the to-be-recognized gesture; and collecting, by the third preset sensor, right-lower-view depth image data of the to-be-recognized gesture.
[0014] In a possible implementation, the performing, according to the second preprocessed point cloud data, of feature extraction processing and feature fusion processing to obtain a recognition result of the to-be-recognized gesture includes: obtaining a trained point cloud deep learning network and a trained fully connected fusion network; the trained point cloud deep learning network is trained by using second preprocessed point cloud data samples and corresponding point cloud feature labels, and the trained fully connected fusion network is trained by using the point cloud feature labels and corresponding feature space labels; the feature extraction processing and the feature fusion processing are performed on the second preprocessed point cloud data according to the trained point cloud deep learning network and the trained fully connected fusion network, to obtain a feature space of the to-be-recognized gesture; and the recognition result of the to-be-recognized gesture is generated according to the feature space.
[0015] In a possible implementation, before the obtaining of the trained point cloud deep learning network and the trained fully connected fusion network, the method further includes: obtaining an initial point cloud deep learning network and an initial fully connected fusion network; configuring a dynamic weight-based loss function for the initial point cloud deep learning network and the initial fully connected fusion network; and training the initial point cloud deep learning network and the initial fully connected fusion network by using the second preprocessed point cloud data samples, the point cloud feature labels, and the feature space labels according to the dynamic weight-based loss function, to obtain the trained point cloud deep learning network and the trained fully connected fusion network.
[0016] In a possible implementation, the first pre-processing of the depth image data to obtain corresponding first preprocessed point cloud data includes: performing dimension coordinate conversion processing on the depth image data to obtain corresponding three-dimensional point cloud data; performing coordinate system unification processing on the three-dimensional point cloud data to obtain corresponding unified point cloud data; and performing time sequence alignment processing on the unified point cloud data to obtain the first preprocessed point cloud data.
[0017] In a possible implementation, the point cloud fusion processing is performed on the first pre-processed point cloud data to obtain fused point cloud data, including: performing outlier filtering processing on the first pre-processed point cloud data to obtain filtered point cloud data; and performing point cloud fusion processing on the filtered point cloud data based on a least square adaptive fusion strategy to obtain the fused point cloud data.
[0018] In a possible implementation, the second pre-processing is performed on the fused point cloud data to obtain second pre-processed point cloud data, including: performing normalization processing and rigid transformation processing on the fused point cloud data to obtain the second pre-processed point cloud data.
[0019] In a second aspect, the present application provides a gesture recognition device, including: a memory, a processor;
[0020] The memory stores computer execution instructions.
[0021] The processor executes the computer execution instructions stored in the memory, so that the processor executes the first aspect and / or various possible implementations of the first aspect.
[0022] In a third aspect, the present application provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to implement the first aspect and / or various possible implementations of the first aspect.
[0023] The gesture recognition method, device and storage medium provided by the present application can realize comprehensive collection of gesture data through depth image data from different collection angles, overcome the inevitable hand self-occlusion problem in traditional single-angle collection, accurately restore the complete three-dimensional structure information of the hand through point cloud fusion processing, obtain the recognition result of the gesture to be recognized through feature extraction and feature fusion processing, and realize low-delay and high-precision gesture recognition. BRIEF DESCRIPTION OF DRAWINGS
[0024] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.
[0025] Figure 1 An application scenario diagram of the gesture recognition method provided by the embodiment of the present application is shown.
[0026] Figure 2 A flowchart of the gesture recognition method provided by the embodiment of the present application is shown.
[0027] Figure 3A schematic diagram of a preset sensor spatial distribution provided for an embodiment of the present application;
[0028] Figure 4 A schematic diagram of a gesture recognition system provided for an embodiment of the present application;
[0029] Figure 5 A schematic diagram of fused point cloud data and gesture recognition results provided for an embodiment of the present application;
[0030] Figure 6 A schematic diagram of a structure of a gesture recognition device provided for an embodiment of the present application;
[0031] Figure 7 A schematic diagram of a structure of a gesture recognition device provided for an embodiment of the present application.
[0032] Through the above-described drawings, the explicit embodiments of the present application have been shown, and will be described in more detail hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application by any means, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0033] Exemplary embodiments will be described in detail herein with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.
[0034] The optical vision capture scheme adopts a depth camera or a binocular vision system to realize non-contact gesture recognition through a three-dimensional reconstruction algorithm. The implementation includes integrating a depth camera or a binocular camera in a virtual reality head-mounted display to collect hand images, and using a deep learning algorithm for key point detection and dynamic hand modeling. However, the optical vision capture has the problems of field of view angle limitation and recognition blind area caused by single view imaging, and is prone to hand self-occlusion problems, resulting in gesture recognition failure or inaccuracy. At the same time, real-time operation requires high computing resources, which not only increases the hardware cost, but also causes problems such as increased signal processing delay and shortened device battery life. In addition, this scheme has strict requirements on environmental lighting conditions and is easily disturbed by external factors, resulting in low gesture recognition accuracy.
[0035] The gesture recognition method provided in the application can realize comprehensive collection of gesture data through depth image data from different collection angles, overcome the inevitable hand self-occlusion problem in traditional single-angle collection, accurately restore the complete three-dimensional structure information of the hand through point cloud fusion processing, obtain the recognition result of the gesture to be recognized by combining feature extraction and feature fusion processing, and realize low-delay and high-precision gesture recognition.
[0036] Figure 1 The application scenario of the gesture recognition method provided in the embodiment of the application is shown in FIG. 1, which is a computer device including a receiving device 101, a processor 102 and a display device 103. Figure 1
[0037] It can be understood that the structure shown in the embodiment of the application does not constitute a specific limitation on the gesture recognition method. In other feasible embodiments of the application, the above architecture can include more or fewer components than the diagram, or combine certain components, or split certain components, or different component arrangement, which can be determined according to the actual application scenario, and is not limited herein. Figure 1 The components shown in the diagram can be realized in hardware, software, or a combination of software and hardware.
[0038] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can obtain multiple depth image data of the gesture to be recognized.
[0039] The processor 102 can process the multiple depth image data of the gesture to be recognized to determine the recognition result of the gesture to be recognized.
[0040] The display device 103 can be used to display the recognition result of the gesture to be recognized and the like.
[0041] The display device can also be a touch display screen, which is used to receive user instructions while displaying the above content to realize interaction with the user.
[0042] It should be understood that the above processor can be realized by reading instructions in the memory and executing the instructions by the processor, or by a chip circuit.
[0043] In addition, the network architecture and business scenario described in the embodiments of the application are used to more clearly illustrate the technical solutions of the embodiments of the application, and do not constitute a limitation on the technical solutions provided by the embodiments of the application. Those skilled in the art can know that, with the evolution of network architecture and the appearance of new business scenarios, the technical solutions provided by the embodiments of the application are also applicable to similar technical problems.
[0044] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes can not be described again in some examples. The embodiments of the present application will be described below with reference to the drawings.
[0045] Figure 2 A flowchart of a gesture recognition method provided for an embodiment of the present application is shown. The execution subject of the embodiment can be a computer device as shown in Figure 1 , which is not particularly limited here. As shown in Figure 2 , the method comprises:
[0046] S201: Obtain a plurality of depth image data of a gesture to be recognized; wherein the plurality of depth image data is derived from different collection angles.
[0047] Specifically, a plurality of depth image data of a gesture to be recognized is collected by a plurality of preset sensors; wherein the plurality of preset sensors are uniformly distributed in a ring shape in space, for covering a preset space around the gesture to be recognized.
[0048] Wherein, the preset space can completely cover all areas of the gesture to be recognized.
[0049] Wherein, the preset sensors include a first preset sensor, a second preset sensor and a third preset sensor.
[0050] Wherein, the first preset sensor, the second preset sensor and the third preset sensor are fixedly connected on a bracelet through a connecting device, and the bracelet is used for wearing on the wrist. The first preset sensor, the second preset sensor and the third preset sensor are uniformly distributed in a ring shape in space, and are kept relatively fixed with the forearm and the hand through the connecting device. As shown in Figure 3 , one of the preset sensors is located directly above the hand, for obtaining overhead angle information such as palm opening and closing, fingertip movement, etc.; the other two preset sensors are arranged at a preset angle below the left and right sides of the hand respectively, and are kept perpendicular to the wrist center axis, for collecting key information such as finger bending which is difficult to capture by the top sensor. Wherein, the preset angle can be 30 degrees.
[0051] Through the optimized layout strategy of the plurality of preset sensors being uniformly distributed in a ring shape in space, the inevitable hand self-occlusion problem in traditional single-view collection is overcome, and the integrity and accuracy of the hand action data are ensured.
[0052] Wherein, the three preset sensors form a triangular collection network.
[0053] The three preset sensors form a triangular acquisition network, realizing multi-dimensional synchronous acquisition of hand motion information, and the point cloud data collected by multiple sensors has high complementarity and can restore complete three-dimensional structure information of the hand.
[0054] It should be noted that the number and position of the preset sensors are not limited in the present application, and can be flexibly adjusted according to the accuracy requirements of the actual application scene.
[0055] The structure of the plurality of preset sensors uniformly distributed in a ring in the present application has excellent fault tolerance performance. Even if individual sensors are temporarily blocked or environmentally disturbed, the entire system can still rely on the data of other sensors to maintain stable operation, ensuring the continuity and reliability of gesture recognition. It provides a flexible way to optimize performance and makes it possible to customize deployment in different application scenarios. By increasing the number of sensors, more detailed hand motion information can be obtained, thereby further improving the accuracy and reliability of gesture recognition.
[0056] Specifically, the top view depth image data of the gesture to be recognized is collected by the first preset sensor. The left lower view depth image data of the gesture to be recognized is collected by the second preset sensor. The right lower view depth image data of the gesture to be recognized is collected by the third preset sensor.
[0057] The preset sensor can be a laser radar sensor.
[0058] S202: performing first preprocessing on the depth image data to obtain corresponding first preprocessed point cloud data.
[0059] Specifically, S202 specifically includes S2021-S2023:
[0060] S2021: performing dimension coordinate conversion processing on the depth image data to obtain corresponding three-dimensional point cloud data.
[0061] Specifically, the depth image data is processed by a depth map conversion module according to the intrinsic matrix of the preset sensor and the mapping of perspective projection, to obtain corresponding three-dimensional point cloud data.
[0062] Each three-dimensional point cloud data coordinate in the point cloud
[0063]
[0064] In the formula, denotes the depth image collected by the kth sensor (k = 1, 2, 3) at time t, and its size is h x w; (u, v) denotes the coordinates of the pixel in the depth image;h denotes the horizontal field of view angle of the kth sensor; θ v denotes the vertical field of view angle of the kth sensor.
[0065] S2022: Coordinate system unification processing is performed on the three-dimensional point cloud data to obtain corresponding unified post-point cloud data.
[0066] Specifically, there are three preset sensors, which are respectively at the top, left lower part and right lower part of the hand, and a unified world coordinate system is established with the top sensor (k = 1) as the reference sensor, and the point cloud data of other sensors is mapped into the coordinate system through rigid transformation.
[0067] wherein the calculation formula of the rigid transformation is:
[0068]
[0069] In the formula, denotes the point cloud transformation of the kth sensor coordinate system to the point cloud coordinate of the reference sensor coordinate system; denotes the point cloud coordinate collected by the kth sensor at time t, and the coordinate system is the coordinate system of the kth sensor itself; denotes the rotation matrix from the kth sensor coordinate system k to the reference sensor coordinate system 1, which satisfies the orthogonality denotes the translation vector, which describes the position of the origin of the coordinate system k in the coordinate system 1; denotes a full 1 column vector, which is used to broadcast the translation vector to all points in the point cloud data.
[0070] S2023: Time sequence alignment processing is performed on the unified post-point cloud data to obtain first preprocessed point cloud data.
[0071] Specifically, interpolation is performed based on the timestamp and motion estimation, and the unified post-point cloud data at different times is aligned to a preset reference time to obtain the first preprocessed point cloud data.
[0072] Optionally, an independent data cache queue is configured for each preset sensor, and the latest point cloud data of the data cache queue is extracted in real time.
[0073] Quasi-synchronous sampling is realized through time window constraint, which can ensure the consistency of the point cloud data in the time dimension and lay a foundation for subsequent spatial fusion.
[0074] S203: Point cloud fusion processing is performed on the first preprocessed point cloud data to obtain fused point cloud data.
[0075] Specifically, S203 specifically includes S2031-S2032:
[0076] S2031: Perform outlier filtering processing on the first pre-processed point cloud data to obtain filtered point cloud data.
[0077] Specifically, for each point of the first pre-processed point cloud data, the average distance of the point to its k-nearest neighbor points is calculated to obtain the average distance of all points; according to the average distance of all points, the statistical characteristics of the average distance are calculated; according to the statistical characteristics, a filtering threshold is obtained; if the average distance of a certain point to its k-nearest neighbor points is greater than the filtering threshold, the certain point is filtered.
[0078] The calculation formula of the average distance of the point to its k-nearest neighbor points is:
[0079]
[0080] In the formula, p i represents the three-dimensional coordinates of the i-th point in the first pre-processed point cloud data; represents the k-nearest neighbor point set of the point p i represents the average distance of p i to ; p j represents the j-th point in the k-nearest neighbor point set .
[0081] The calculation formula of the average distance of all points is:
[0082]
[0083] In the formula, N represents the total number of points in the point cloud data; μ represents the global average distance, that is, the average value of the average distance d i of all points to k-nearest neighbor points.
[0084] The calculation formula of the statistical characteristics is:
[0085]
[0086] In the formula, σ represents the standard deviation of the distance.
[0087] The calculation formula of the filtering threshold is:
[0088] a = μ + ασ
[0089] In the formula, a represents the filtering threshold; α represents the threshold parameter.
[0090] S2032: According to the filtered point cloud data, perform point cloud fusion processing based on a least squares adaptive fusion strategy to obtain fused point cloud data.
[0091] Specifically, an observation equation is established; a covariance matrix is determined according to error characteristics of each sensor; a minimum error sum of squares is obtained according to the covariance matrix and the observation equation; and an optimal fusion estimation is obtained according to the minimum error sum of squares, so as to obtain fused point cloud data.
[0092] The formula of the observation equation is:
[0093] z k =H k x+v k ,k=1,2,3
[0094] In the formula, z k is an observation value of the kth sensor, x is a real spatial position, H k is an observation matrix of the kth sensor, and v k is observation noise.
[0095] The calculation formula of the minimum error sum of squares is:
[0096]
[0097] In the formula, J(x) represents the minimum error sum of squares; represents an inverse matrix of the covariance matrix of the kth sensor.
[0098] The calculation formula of the optimal fusion estimation is:
[0099]
[0100] In the formula, P t represents the fused point cloud data.
[0101] S204: The fused point cloud data is subjected to a second preprocessing, so as to obtain second preprocessed point cloud data.
[0102] Specifically, the fused point cloud data is subjected to normalization processing and rigid transformation processing, so as to obtain the second preprocessed point cloud data.
[0103] Specifically, the maximum value and the minimum value in each coordinate axis direction of the fused point cloud data are calculated; the normalized point cloud data is obtained through normalization according to the maximum value and the minimum value; and the second preprocessed point cloud data is obtained through rigid transformation processing on the normalized point cloud data.
[0104] The calculation formula of the maximum value and the minimum value is:
[0105] x max =max i=1,…,N x i ,x min =min i=1,…,Nx i ,y max =max i=1,…,N y i
[0106] y min =min i=1,…,N y i ,z max =max i=1,…,N z i ,z min =min i=1,…,N z i
[0107] In the formula, x max , x min respectively represent the maximum and minimum values of the fused point cloud data in the x coordinate axis direction; y max , y min respectively represent the maximum and minimum values of the fused point cloud data in the y coordinate axis direction; z max , z min respectively represent the maximum and minimum values of the fused point cloud data in the z coordinate axis direction; x i , y i , z i respectively represent the coordinate values of the fused point cloud data in each coordinate axis direction; and N represents the total number of points in the point cloud data.
[0108] Each point in the normalized point cloud may be represented as:
[0109]
[0110] The calculation formula for rigid transformation processing of the normalized point cloud data is:
[0111]
[0112] In the formula, X t represents the second preprocessed point cloud data: is a rotation matrix, is a translation vector, is a full one vector. R t and t t are obtained by performing principal component analysis on the point cloud .
[0113] S205: According to the second preprocessed point cloud data, feature extraction processing and feature fusion processing are performed to obtain the recognition result of the to-be-recognized gesture.
[0114] Specifically, S205 specifically includes S2051-S2053:
[0115] S2051: Obtain a trained point cloud deep learning network and a trained fully connected fusion network; wherein the trained point cloud deep learning network is trained by the second pre-processed point cloud data sample and the corresponding point cloud feature label, and the trained fully connected fusion network is trained by the point cloud feature label and the corresponding feature space label.
[0116] S2052: According to the trained point cloud deep learning network and the trained fully connected fusion network, the second pre-processed point cloud data is subjected to feature extraction processing and feature fusion processing to obtain a feature space of a gesture to be recognized.
[0117] Optionally, the second pre-processed point cloud data is subjected to feature extraction processing to obtain a point cloud feature of a current time; a gesture recognition result of a previous time is obtained; a hot start mechanism is used to pre-process data of the point cloud feature of the current time and the gesture recognition result of the previous time to obtain a time series context feature; the time series context feature is subjected to feature fusion processing to obtain the feature space of the gesture to be recognized.
[0118] The calculation formula of the feature extraction processing is:
[0119]
[0120] In the formula, z represents a feature extraction network; represents a parameter of the network; z t represents the point cloud feature of the current t time extracted.
[0121] The calculation formula of the time series context feature is:
[0122]
[0123] In the formula, h t is a time series context feature containing historical information, and ψ is a feature construction parameter of the hot start mechanism, is the gesture recognition result of the previous time.
[0124] The calculation formula of the feature fusion processing is:
[0125]
[0126] In the formula, z represents a feature space of a gesture to be recognized at a current t time; f θ represents a fully connected fusion network; and θ represents a parameter of the network.
[0127] The continuity and accuracy of gesture recognition are improved through feature fusion based on timing information.
[0128] S2053: generating a recognition result of the gesture to be recognized according to the feature space.
[0129] Specifically, the feature space is inversely transformed to generate the recognition result of the gesture to be recognized.
[0130] The formula for inversely transforming the feature space is:
[0131]
[0132] In the formula, denotes an inverse transformation function of coordinate transformation; denotes an inverse matrix of a rotation matrix.
[0133] The gesture recognition method provided by the embodiments of the present application realizes comprehensive acquisition of gesture data from depth image data from different collection perspectives, overcomes the inevitable hand self-occlusion problem in traditional single-perspective collection, and through point cloud fusion processing, can accurately restore complete three-dimensional structure information of the hand, and in combination with feature extraction and feature fusion processing, obtains a recognition result of the gesture to be recognized. Low-delay and high-precision gesture recognition is realized.
[0134] In an embodiment of the present application, on the basis of the above-mentioned embodiments, before step S2051, the process of training the point cloud deep learning network and the fully connected fusion network is further included, which is described in detail as follows:
[0135] S2054: obtaining an initial point cloud deep learning network and an initial fully connected fusion network.
[0136] S2055: configuring a loss function based on dynamic weights for the initial point cloud deep learning network and the initial fully connected fusion network.
[0137] The formula for the loss function based on dynamic weights is:
[0138]
[0139] In the formula, f θ denotes a fully connected fusion network; and ψ is a feature construction parameter of a warm start mechanism; denotes a feature extraction network; denotes a true value of the gesture at the current t moment; ||·|| denotes a certain norm metric in the gesture space; the superscript (i) indicates that the sample comes from the i th sample sequence in the training set; T i denotes the length of the i th sample sequence; and N denotes the total number of points in the point cloud data. It is an adaptive function based on the magnitude of gesture changes at adjacent time points. The input is two n-dimensional vectors, and the output is a real number; n represents the dimension of the gesture output vector.
[0140] It should be noted that when a drastic change in gesture is detected (i.e., and If the difference is significant, the weight of that sample will be automatically increased.
[0141] S2056: Based on the loss function based on dynamic weights, the initial point cloud deep learning network and the initial fully connected fusion network are trained using the second preprocessed point cloud data samples, point cloud feature labels, and feature space labels to obtain the trained point cloud deep learning network and the trained fully connected fusion network.
[0142] The gesture recognition method provided in this application improves the recognition accuracy of fast gestures and enhances the recognition performance of high-dynamic gestures by combining a loss function based on dynamic weights during the training process of the point cloud deep learning network and the fully connected fusion network.
[0143] Figure 4 This is a schematic diagram of a gesture recognition system provided in an embodiment of this application. Figure 4 As shown, the gesture recognition system includes a LiDAR sensor, a depth map conversion module, a point cloud fusion module, a gesture recognition module, and an adaptive gesture dynamic weight adjustment module.
[0144] The lidar sensors, used to acquire depth image data and send it to the depth conversion module, include a top lidar sensor, a lower left lidar sensor, and a lower right lidar sensor. The top lidar sensor acquires depth image data from the top viewpoint. The lower left lidar sensor acquires depth image data from the lower left viewpoint. The lower right lidar sensor acquires depth image data from the lower right viewpoint.
[0145] The depth map conversion module is used to convert depth image data from different acquisition perspectives into point cloud data from different acquisition perspectives, and then send the point cloud data from different acquisition perspectives to the point cloud fusion module.
[0146] The point cloud fusion module is configured to perform coordinate system unification and fusion processing on the point cloud data collected from different perspectives to obtain fused point cloud data, and send the fused point cloud data to the gesture recognition module. The point cloud fusion module comprises a time synchronization module, a point cloud registration module, a coordinate transformation module, and a noise filtering and fusion module. The time synchronization module is configured to synchronize the point cloud data collected from different perspectives in time. The point cloud registration module is configured to register the coordinate systems of the point cloud data collected from different perspectives after time synchronization. The coordinate transformation module is configured to convert the point cloud data collected from different perspectives after coordinate system registration to a reference coordinate system. The noise filtering and fusion module is configured to identify and remove outliers in the point cloud data after coordinate system unification, and fuse the point cloud data collected from different perspectives.
[0147] The gesture recognition module is configured to perform gesture recognition based on the fused point cloud data at the current moment and the gesture recognition preliminary result at the previous moment to obtain a gesture recognition preliminary result, and send the gesture recognition preliminary result to the adaptive gesture dynamic weight adjustment module. The gesture recognition module comprises a preprocessing module, a feature extraction module, a hot start module, and a fully connected network. The preprocessing module is configured to perform normalization and rigid transformation processing on the fused point cloud data at the current moment, send the processed point cloud data at the current moment to the feature extraction module, and obtain and send the gesture recognition preliminary result at the previous moment to the hot start module. The feature extraction module is configured to perform feature extraction on the processed point cloud data at the current moment to obtain point cloud features at the current moment, and send the point cloud features at the current moment to the hot start module. The hot start module is configured to perform data preprocessing on the gesture recognition preliminary result at the previous moment and the point cloud features at the current moment to obtain time sequence context features, and send the time sequence context features to the fully connected network. The fully connected network is configured to perform feature fusion processing on the time sequence context features to obtain the gesture recognition preliminary result, and send the gesture recognition preliminary result to the adaptive gesture dynamic weight adjustment module and the preprocessing module.
[0148] The adaptive gesture dynamic weight adjustment module is configured to perform dynamic weight adjustment on the gesture recognition preliminary result to obtain a gesture recognition result.
[0149] Figure 5 The gesture recognition result and the fused point cloud data provided by the embodiments of the present application are shown in the schematic diagram.
[0150] The gesture recognition method provided by the embodiments of the present application realizes comprehensive collection of gesture data through depth image data from different collection perspectives, overcomes the inevitable hand self-occlusion problem in traditional single-perspective collection, accurately restores the complete three-dimensional structure information of the hand through point cloud fusion processing, and obtains the recognition result of the gesture to be recognized through feature extraction and feature fusion processing. The gesture recognition method realizes low-latency and high-precision gesture recognition.
[0151] Figure 6 A structural schematic diagram of a gesture recognition device provided for an embodiment of the present application is shown in FIG. 6. As shown in FIG. 6, the gesture recognition device 60 provided by the embodiment includes an acquisition module 601, a processing module 602, a fusion module 603, and a recognition module 604. Figure 6
[0152] The acquisition module 601 is configured to acquire a plurality of depth image data of a gesture to be recognized, wherein the plurality of depth image data is derived from different collection perspectives.
[0153] The processing module 602 is configured to perform first preprocessing on the depth image data to obtain corresponding first preprocessed point cloud data.
[0154] The fusion module 603 is configured to perform point cloud fusion processing according to the first preprocessed point cloud data to obtain fused point cloud data.
[0155] The processing module 602 is further configured to perform second preprocessing on the fused point cloud data to obtain second preprocessed point cloud data.
[0156] The recognition module 604 is configured to perform feature extraction processing and feature fusion processing according to the second preprocessed point cloud data to obtain a recognition result of the gesture to be recognized.
[0157] In a possible implementation, the acquisition module 601 is specifically configured to collect the plurality of depth image data of the gesture to be recognized by a plurality of preset sensors, wherein the plurality of preset sensors are configured to cover a preset space around the gesture to be recognized.
[0158] In a possible implementation, the preset sensors include a first preset sensor, a second preset sensor, and a third preset sensor. Correspondingly, the acquisition module 601 is specifically configured to collect top perspective depth image data of the gesture to be recognized by the first preset sensor, collect left lower perspective depth image data of the gesture to be recognized by the second preset sensor, and collect right lower perspective depth image data of the gesture to be recognized by the third preset sensor.
[0159] In a possible implementation, the recognition module 604, specifically configured to: obtain a trained point cloud deep learning network and a trained fully connected fusion network; wherein the trained point cloud deep learning network is trained by the second preprocessed point cloud data sample and the corresponding point cloud feature label, and the trained fully connected fusion network is trained by the point cloud feature label and the corresponding feature space label; perform feature extraction processing and feature fusion processing on the second preprocessed point cloud data according to the trained point cloud deep learning network and the trained fully connected fusion network, to obtain a feature space of the gesture to be recognized; and generate a recognition result of the gesture to be recognized according to the feature space.
[0160] In a possible implementation, the gesture recognition apparatus 60 further includes:
[0161] a training module configured to: obtain an initial point cloud deep learning network and an initial fully connected fusion network; configure a loss function based on dynamic weights for the initial point cloud deep learning network and the initial fully connected fusion network; and train the initial point cloud deep learning network and the initial fully connected fusion network by using the second preprocessed point cloud data sample, the point cloud feature label, and the feature space label according to the loss function based on dynamic weights, to obtain the trained point cloud deep learning network and the trained fully connected fusion network.
[0162] In a possible implementation, the processing module 602, specifically configured to: perform dimension coordinate conversion processing on the depth image data, to obtain corresponding three-dimensional point cloud data; perform coordinate system unification processing on the three-dimensional point cloud data, to obtain corresponding unified point cloud data; and perform time sequence alignment processing on the unified point cloud data, to obtain the first preprocessed point cloud data.
[0163] In a possible implementation, the fusion module 603, specifically configured to: perform outlier filtering processing on the first preprocessed point cloud data, to obtain filtered point cloud data; and perform point cloud fusion processing on the filtered point cloud data based on a least square adaptive fusion strategy, to obtain the fused point cloud data.
[0164] In a possible implementation, the processing module 602, specifically configured to: perform normalization processing and rigid transformation processing on the fused point cloud data, to obtain the second preprocessed point cloud data.
[0165] The gesture recognition apparatus provided in this embodiment can perform the method provided in the method embodiments, and has similar implementation principles and technical effects, which will not be described here in detail.
[0166] Figure 7 A structural schematic diagram of a gesture recognition device is provided in the embodiments of the present application. As shown in the figure, the gesture recognition device 70 provided in the embodiments includes at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. The processor 701, the memory 702 and the communication component 703 are connected through a bus 704. Figure 7
[0167] In the implementation process, the at least one processor 701 executes the computer execution instructions stored in the memory 702, so that the at least one processor 701 executes the above-mentioned method.
[0168] The specific implementation process of the processor 701 can refer to the method embodiments described above, which has similar implementation principles and technical effects, and will not be described here in detail.
[0169] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC) and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like. The steps of the method disclosed in the application can be directly embodied as the execution of the hardware processor, or the execution of the combination of the hardware and software modules in the processor.
[0170] The memory can include a random access memory (RAM), and can also include a non-volatile memory (NVM), for example, at least one disk memory.
[0171] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus or an extended industry standard architecture (EISA) bus and the like. The bus can be divided into an address bus, a data bus, a control bus and the like. For the convenience of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.
[0172] The application further provides a computer readable storage medium, and the computer readable storage medium stores computer execution instructions.
[0173] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method.
[0174] The readable storage medium can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0175] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.
[0176] The division of units is only a logical function division, and in actual implementation, there can be another division mode, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0177] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.
[0178] In addition, the functional units in each embodiment of the application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0179] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0180] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction-related hardware. The aforementioned program can be stored in a computer readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: a ROM, a RAM, a magnetic disk or an optical disk, and various media that can store program codes.
[0181] Finally, it should be noted that: those skilled in the art will easily derive other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known or customary technical means in the art that are not disclosed in the present application, and is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is only limited by the appended claims.
Claims
1. A gesture recognition method, characterized in that, include: Acquire multiple depth image data of the gesture to be recognized; wherein the multiple depth image data are obtained from different acquisition perspectives; The depth image data is subjected to a first preprocessing to obtain the corresponding first preprocessed point cloud data; Based on the first preprocessed point cloud data, point cloud fusion processing is performed to obtain fused point cloud data; The fused point cloud data is subjected to a second preprocessing to obtain the second preprocessed point cloud data; Based on the second preprocessed point cloud data, feature extraction and feature fusion are performed to obtain the recognition result of the gesture to be recognized.
2. The method according to claim 1, characterized in that, The acquisition of multiple depth image data of the gesture to be recognized includes: Multiple depth image data of the gesture to be recognized are acquired through multiple preset sensors; wherein, the multiple preset sensors are used to cover a preset space around the gesture to be recognized.
3. The method according to claim 2, characterized in that, The preset sensors include a first preset sensor, a second preset sensor, and a third preset sensor; Accordingly, the process of acquiring multiple depth image data of the gesture to be recognized through multiple preset sensors includes: The top-view depth image data of the gesture to be recognized is acquired through the first preset sensor; The lower left perspective depth image data of the gesture to be recognized is acquired through the second preset sensor. The third preset sensor acquires depth image data of the lower right view of the gesture to be recognized.
4. The method according to any one of claims 1 to 3, characterized in that, The step of performing feature extraction and feature fusion processing based on the second preprocessed point cloud data to obtain the recognition result of the gesture to be recognized includes: A trained point cloud deep learning network and a trained fully connected fusion network are obtained; wherein, the trained point cloud deep learning network is trained using point cloud data samples after the second preprocessing and the corresponding point cloud feature labels, and the trained fully connected fusion network is trained using the point cloud feature labels and the corresponding feature space labels. Based on the trained point cloud deep learning network and the trained fully connected fusion network, feature extraction and feature fusion processing are performed on the second preprocessed point cloud data to obtain the feature space of the gesture to be recognized. Based on the feature space, the recognition result of the gesture to be recognized is generated.
5. The method according to claim 4, characterized in that, Before obtaining the trained point cloud deep learning network and the trained fully connected fusion network, the following steps are also included: Obtain the initial point cloud deep learning network and the initial fully connected fusion network; For the initial point cloud deep learning network and the initial fully connected fusion network, configure a loss function based on dynamic weights; Based on the loss function based on dynamic weights, the initial point cloud deep learning network and the initial fully connected fusion network are trained using the second preprocessed point cloud data samples, the point cloud feature labels, and the feature space labels, so as to obtain the trained point cloud deep learning network and the trained fully connected fusion network.
6. The method according to any one of claims 1 to 3, characterized in that, The first preprocessing of the depth image data to obtain corresponding preprocessed point cloud data includes: The depth image data is subjected to dimensional coordinate transformation to obtain the corresponding three-dimensional point cloud data; The three-dimensional point cloud data is processed using a coordinate system to obtain the corresponding unified point cloud data; Based on the unified point cloud data, temporal alignment processing is performed to obtain the first preprocessed point cloud data.
7. The method according to any one of claims 1 to 3, characterized in that, The step of performing point cloud fusion processing based on the first preprocessed point cloud data to obtain fused point cloud data includes: Outlier filtering is performed on the first preprocessed point cloud data to obtain filtered point cloud data. Based on the filtered point cloud data, point cloud fusion processing is performed using an adaptive fusion strategy based on least squares to obtain the fused point cloud data.
8. The method according to any one of claims 1 to 3, characterized in that, The second preprocessing of the fused point cloud data to obtain second preprocessed point cloud data includes: The fused point cloud data is then normalized and subjected to rigid transformation to obtain the second preprocessed point cloud data.
9. A gesture recognition device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
Citation Information
Patent Citations
A method and apparatus for recognizing static gestures
CN109002811A
Gesture recognition method and system based on depth camera and contour extraction
CN116311492A
Human-computer interaction method based on gesture recognition and interaction system thereof
CN118247850A
Gesture recognition method and device based on multi-modal information fusion and storage medium
CN118334742A
Dynamic gesture recognition method based on point cloud
CN118629087A