A gesture recognition method and system based on lightweight computing and storage

By building a lightweight millimeter-wave radar gesture recognition system, the CNN-LSTM hybrid neural network model is used to achieve low latency and high accuracy gesture recognition on embedded devices, solving the problems of high computing complexity, large storage requirements and privacy leakage, and is suitable for consumer electronics and Internet of Things scenarios with limited resources.

CN120318920BActive Publication Date: 2025-08-19SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510804930.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-19
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing millimeter-wave radar gesture recognition technology has high computing complexity, high storage requirements, high power consumption and high risk of privacy data leakage on embedded devices, making it difficult to deploy on a large scale in consumer electronics and IoT scenarios with resource-constrained.

Method used

Using lightweight computing and storage gesture recognition methods, data is collected through millimeter wave radar to generate point clouds, feature vector matrix is ​​constructed, CNN-LSTM hybrid neural network model is converted into a lightweight neural network model and deployed on a millimeter wave radar chip with Arm Cortex-M core for identification.

Benefits of technology

It realizes low latency and high accuracy gesture recognition on resource-constrained devices, reduces computing complexity and storage requirements, and enhances system independence and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318920B_ABST
    Figure CN120318920B_ABST
Patent Text Reader

Abstract

This application discloses a gesture recognition method and system based on lightweight computing and storage, belonging to the field of gesture recognition technology. The method includes: using a millimeter radar to collect millimeter-wave radar raw data and generate point cloud data; based on the point cloud data, extracting eigenvectors and constructing an eigenvector matrix to capture dynamic changes in gestures; based on the eigenvector matrix, collecting millimeter-wave radar data of different gestures and manually annotating them to obtain a millimeter-wave radar gesture dataset; constructing a CNN-LSTM hybrid neural network model and training it based on the gesture dataset to obtain a Keras model; converting the Keras model into a lightweight neural network model suitable for embedded devices; and deploying the lightweight neural network model to a millimeter-wave radar chip with an Arm Cortex-M core for gesture recognition. This method reduces computational complexity and storage requirements, and enhances system independence and data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gesture recognition technology, and in particular to a gesture recognition method and system based on lightweight computing and storage. Background Art

[0002] With the rapid development of intelligent hardware technology, gesture recognition, as a key method of human-computer interaction, is becoming increasingly valuable in smart devices. Traditional gesture recognition technology typically relies on cameras or infrared sensors to collect information. While significant progress has been made, its performance is susceptible to changes in ambient lighting and detection range, and it requires significant computing resources to process high-dimensional sensor data. In contrast, millimeter-wave radar, with its all-weather capability, non-contact detection characteristics, and excellent privacy protection, has demonstrated unique technological value in areas such as biometric monitoring and automotive radar.

[0003] In recent years, the deep integration of millimeter-wave radar and artificial intelligence has opened up new avenues for gesture recognition. Current mainstream solutions generally employ deep neural network architectures deployed in the cloud or on a host computer. While these solutions maintain high recognition accuracy, they face three challenges in practical implementation: First, massive parameter models place significant pressure on the storage space of embedded devices. For example, on-chip storage in a typical millimeter-wave radar processing platform often struggles to accommodate neural networks with over a million parameters. Second, the computational intensity of complex models far exceeds the computing power limits of edge devices, making real-time performance difficult to guarantee. Third, while some cloud-edge collaborative architectures alleviate local computing pressure, they introduce significant communication latency, resulting in increased system power consumption and the risk of data privacy leakage. These technical bottlenecks severely restrict the application of millimeter-wave radar gesture recognition technology in resource-constrained scenarios such as consumer electronics and the Internet of Things. Existing solutions, in particular, face obstacles to large-scale deployment in areas such as wearables and smart homes, where low power consumption and real-time response are stringent requirements. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present application provides a gesture recognition method and system based on lightweight computing and storage, which solves the problems of computing dependence on the cloud and host computer, high computing power and storage requirements, high system power consumption, and privacy leakage risks.

[0005] In order to achieve the above-mentioned invention objectives, the technical solutions adopted in this application are:

[0006] First aspect:

[0007] This application provides a gesture recognition method based on lightweight computing and storage, including:

[0008] S1: Use the millimeter radar to collect millimeter-wave radar raw data and generate millimeter-wave radar point cloud data based on the radar data link;

[0009] S2: extracting feature vectors of the millimeter-wave radar data based on the millimeter-wave radar point cloud data, and constructing a feature vector matrix of the millimeter-wave radar data to capture dynamic changes in gestures;

[0010] S3: Based on the eigenvector matrix of the millimeter-wave radar data, collect millimeter-wave radar data of different gestures, and manually annotate the millimeter-wave radar data of the different gestures using an annotation tool to obtain a millimeter-wave radar gesture dataset;

[0011] S4: Build a CNN-LSTM hybrid neural network model;

[0012] S5: Inputting the millimeter-wave radar gesture dataset into the CNN-LSTM hybrid neural network model for training to obtain a Keras model;

[0013] S6: Convert the Keras model into a lightweight neural network model suitable for the embedded end;

[0014] S7: Deploy the lightweight neural network model on a millimeter-wave radar chip with an Arm Cortex-M core, and perform gesture recognition based on the millimeter-wave radar chip.

[0015] Furthermore, the millimeter-wave radar point cloud data in S1 includes spatial coordinates, speed, and signal-to-noise ratio information.

[0016] Furthermore, the S2 includes:

[0017] S201: Record the total number of point clouds at the current moment, and respectively record the number of point clouds with radial velocities greater than 0 and less than 0 at the current moment: 、 、 ;

[0018] S202: Calculate the average signal-to-noise ratio, weighted average distance, and weighted average speed of all point cloud data based on the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 ;

[0019] S203: Calculate the weighted average distance and weighted average velocity of the point clouds corresponding to radial velocities greater than 0 and radial velocities less than 0, respectively, based on the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 、 ;

[0020] S204: Calculate the weighted centroid coordinates at the centroid of the millimeter-wave radar point cloud data according to the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 ;

[0021] S205: Arrange the number of point clouds, average signal-to-noise ratio, weighted average distance, weighted average velocity, and weighted centroid coordinates into a column vector to form a feature vector of the millimeter-wave radar data:

[0022]

[0023] in, is the total number of point clouds in a frame of data, is the number of point clouds with radial velocity values greater than 0 in the total number of point clouds, is the number of point clouds with radial velocity values less than 0 in the total number of point clouds, is the average signal-to-noise ratio of all point clouds, is the weighted average distance of all point clouds, is the weighted average velocity of all point clouds, is the weighted average distance of the point cloud corresponding to the radial velocity greater than 0, is the weighted average velocity of the point cloud corresponding to the velocity greater than 0, is the weighted average distance of the point cloud corresponding to the radial velocity less than 0, is the weighted average velocity of the point cloud corresponding to the radial velocity being less than 0, 、 、 is the Cartesian space coordinate of the weighted centroid, for The feature vector of the millimeter-wave radar data at the moment;

[0024] S206: Based on the eigenvectors of the millimeter-wave radar data, construct an eigenvector matrix of the millimeter-wave radar data by using a sliding time window technology.

[0025] Furthermore, the step S206 of constructing a feature vector matrix of millimeter wave radar data by using a sliding time window technique includes:

[0026] Using the sliding time window technique, the feature vectors of continuously collected millimeter-wave radar data are arranged in chronological order. The 13-dimensional feature vector of each frame is used as a column of the matrix. The features of 10 consecutive frames are arranged to form a 13×10 feature matrix:

[0027]

[0028] in, Indicates the end time of the current sliding time window, Indicates the starting time of the current sliding time window.

[0029] Furthermore, the calculation method of the average signal-to-noise ratio, weighted average distance, weighted average speed and weighted center of mass coordinates at the center of mass includes:

[0030] The formula for calculating the average signal-to-noise ratio is:

[0031]

[0032] in, is the number of points in the point cloud in a frame of data, The first The signal-to-noise ratio of the point cloud;

[0033] The calculation formula of weighted average distance is:

[0034]

[0035] in, The first The Euclidean distance from the point cloud to the origin, The first The Cartesian space coordinates of the point cloud points, is the weighted average distance of all point clouds in a frame of data;

[0036] The formula for calculating the weighted average speed is:

[0037]

[0038] in, The first The radial velocity of the point cloud relative to the millimeter wave radar sensor, is the weighted average velocity of all point clouds in a frame of data;

[0039] The weighted centroid coordinate calculation formula at the centroid is:

[0040]

[0041]

[0042]

[0043] in, 、 、 The first The Cartesian space coordinates of the centroid of the point cloud.

[0044] Furthermore, the CNN-LSTM hybrid neural network model includes:

[0045] Input module, used to receive a feature matrix with a shape of [10, 13];

[0046] The preliminary feature extraction module converts the feature matrix with an input shape of [10, 13] into a preliminary feature matrix with an output shape of [4, 13] through a one-dimensional convolution operation;

[0047] The feature processing and enhancement module includes two submodules, each of which consists of a one-dimensional convolutional layer, a batch normalization layer, and a first activation layer. The one-dimensional convolutional layer uses 24 convolution kernels of size 3, a step size of 1, and a linear activation function to convert the preliminary feature matrix [4, 13] into a feature matrix with an output shape of [4, 24]. The batch normalization layer standardizes the output of the one-dimensional convolutional layer to an output shape of [4, 24]. The first activation layer introduces nonlinearity through the ReLU activation function to convert the output of the batch normalization layer into a nonlinear output with an output shape of [4, 24].

[0048] A time series feature extraction module includes a long short-term memory layer, which uses a hyperbolic tangent function and a sigmoid function as activation functions and a recurrent activation function to capture long-term dependencies in sequence data;

[0049] The classification decision module includes a fully connected layer and a second activation layer, wherein the fully connected layer is configured to have an output size of 6, uses a linear activation function, and has an output shape of 6 to generate a score for each category; the second activation layer uses a Softmax function for activation to convert the output of the fully connected layer into a probability distribution for gesture classification.

[0050] Furthermore, in S5, the millimeter-wave radar gesture dataset is input into the CNN-LSTM hybrid neural network model for training to obtain a Keras model, including:

[0051] The CNN-LSTM hybrid neural network model is trained using the TensorFlow machine learning framework to obtain a Keras model.

[0052] Furthermore, the S6 converts the Keras model into a lightweight neural network model suitable for the embedded end, specifically including:

[0053] A1: Use the model conversion tool to convert the Keras model into a lightweight TensorFlow Lite model.

[0054] A2: Optimize the TensorFlow Lite model to meet the performance and storage requirements of the embedded device.

[0055] A3: Use the TensorFlow Lite Micro framework to convert the optimized TensorFlow Lite model into a lightweight neural network model.

[0056] Second aspect:

[0057] This application provides a gesture recognition system based on lightweight computing and storage, including:

[0058] A millimeter-wave radar data acquisition module, data processing module, and gesture recognition module integrated on the same millimeter-wave radar chip with an Arm Cortex-M core;

[0059] The millimeter wave radar data acquisition module includes a millimeter wave radar transceiver system for collecting millimeter wave radar raw data;

[0060] The data processing module is used to preprocess and extract features from the millimeter-wave radar raw data to generate millimeter-wave radar gesture data that can be used for training and recognition;

[0061] A model training and conversion module, comprising a model training unit and a model conversion unit, wherein the model training unit uses the millimeter-wave radar gesture dataset of the data processing module to train a CNN-LSTM hybrid neural network model; the model conversion unit is used to convert the trained neural network model into a lightweight embedded model and optimize it to meet the computing and storage requirements of the embedded device;

[0062] The gesture recognition module includes a data input unit and a neural network inference unit, wherein the data input unit is used to receive the millimeter wave radar feature vector matrix data output by the data processing module; the neural network inference unit is used to perform neural network operations based on a lightweight embedded model and output the gesture type;

[0063] The judgment result display module includes a display unit for displaying the gesture recognition result in real time according to the recognition result output by the gesture recognition module, wherein the gesture recognition result includes the gesture type and the corresponding gesture action information.

[0064] Furthermore, the gesture recognition module is deployed on the Arm Cortex-M processor of the millimeter wave radar chip and runs on the processor.

[0065] The beneficial effects of this application are:

[0066] This application presents a gesture recognition method and system based on lightweight computation and storage. By utilizing an optimized eigenvector matrix construction method and a lightweight neural network model, they achieve low-latency, high-accuracy gesture recognition on resource-constrained embedded devices. Through model compression and optimization, computational complexity and storage requirements are significantly reduced, while enhancing system independence and data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0068] Figure 1 A flowchart of a gesture recognition method based on lightweight computing and storage is provided in an embodiment of the present application.

[0069] Figure 2 A schematic flow chart of an existing radar data link processing method provided in an embodiment of the present application.

[0070] Figure 3 A schematic diagram of the process of extracting a time series feature matrix from point cloud data provided in an embodiment of the present application.

[0071] Figure 4 This is a schematic diagram of the structure of a gesture recognition model built based on a CNN-LSTM hybrid neural network provided in an embodiment of the present application.

[0072] Figure 5 A schematic diagram of a gesture recognition system based on lightweight computing and storage provided in an embodiment of the present application. DETAILED DESCRIPTION

[0073] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0074] Example 1:

[0075] The present application embodiment provides a gesture recognition method based on lightweight computing and storage, which can be found in Figure 1 , Figure 1 A method flow chart of a gesture recognition method based on lightweight computing and storage provided in an embodiment of the present application includes:

[0076] S1: Use the millimeter radar to collect millimeter-wave radar raw data and generate millimeter-wave radar point cloud data based on the radar data link.

[0077] Furthermore, the millimeter-wave radar point cloud data in S1 includes spatial coordinates, speed, and signal-to-noise ratio information.

[0078] In one embodiment of the present application, Figure 2 As shown, Figure 2 A flowchart of an existing radar data link processing method provided in an embodiment of the present application. Generating millimeter-wave radar point cloud data based on a radar data link involves the following: the millimeter-wave radar transmits a frequency modulated continuous wave (FMCW) signal and receives the reflected wave. After frequency mixing and analog-to-digital conversion, the raw sampled data is obtained. The format of the raw sampled data is the number of sampling points × number of frames × number of antennas. Subsequently, a fast Fourier transform is performed on the sampling points and frame dimensions to generate a range-Doppler matrix, and static filtering is used to remove clutter. The angle information is calculated using the Direction of Arrival (DOA) algorithm, and the target is detected using the Constant False-Alarm Rate (CFAR) algorithm. Ultimately, a millimeter-wave radar point cloud containing spatial coordinates, velocity information, and signal-to-noise ratio information is generated. This processing flow is part of existing point cloud generation technology and will not be described in detail here.

[0079] Based on the above processing flow, millimeter-wave radar can output the target's multi-dimensional information in the form of a point cloud, providing a basis for subsequent target detection and recognition. Each millimeter-wave radar point cloud contains the following information:

[0080] (1) The spatial coordinates (Xn, Yn, Zn) of the point cloud relative to the millimeter-wave radar sensor as the origin, in meters (m);

[0081] (2) The radial velocity (Vn) of the point cloud relative to the millimeter-wave radar sensor as the origin, in meters per second (m / s). When Vn is greater than 0, it means that the point cloud is moving away from the millimeter-wave radar sensor; when Vn is less than 0, it means that the point cloud is moving closer to the millimeter-wave radar sensor.

[0082] (3) The signal-to-noise ratio (SNR) of the point cloud calculated from the peak value found by the constant false alarm rate algorithm and the noise, in decibels (dB).

[0083] S2: Based on the millimeter-wave radar point cloud data, extract the feature vector of the millimeter-wave radar data, and construct a feature vector matrix of the millimeter-wave radar data to capture dynamic changes in gestures.

[0084] In one embodiment of the present application, Figure 3 As shown, Figure 3 Schematic diagram of the process of extracting a time series feature matrix from point cloud data provided by an embodiment of the present application. Count and record the total number of point clouds at the current moment, and record the number of point clouds with radial velocities greater than 0 and less than 0 at the current moment: 、 、 ; According to the size of the signal-to-noise ratio, calculate the average signal-to-noise ratio, weighted average distance and weighted average speed of all point clouds: 、 、 ; According to the size of the signal-to-noise ratio, the weighted average distance and weighted average velocity of the point cloud corresponding to radial velocity greater than 0 and radial velocity less than 0 are calculated respectively: 、 、 、 ; Calculate the weighted centroid coordinates at the centroid according to the signal-to-noise ratio: 、 、 ; The above 13 data 、 、 、 、 、 、 、 、 、 、 、 、 Arranged into column vectors to form the feature vectors of millimeter wave radar data:

[0085]

[0086] in, (Numbers of Total Point, total number of point clouds) is the total number of point clouds in a frame of data, (Numbers of Centrifuge Point, total number of centrifugal points) is the number of point clouds with radial velocity values greater than 0 in the total number of point clouds. Numbers of Centripetal Point (Numbers of Centripetal Points) is the number of point clouds with radial velocity values less than 0 in the total number of point clouds. is the average signal-to-noise ratio of all point clouds, is the weighted average distance of all point clouds, is the weighted average velocity of all point clouds, is the weighted average distance of the point cloud corresponding to the radial velocity greater than 0, is the weighted average velocity of the point cloud corresponding to the velocity greater than 0, is the weighted average distance of the point cloud corresponding to the radial velocity less than 0, is the weighted average velocity of the point cloud corresponding to the radial velocity being less than 0, 、 、 is the Cartesian space coordinate of the weighted centroid, for The feature vector of the millimeter-wave radar data at time t.

[0087] In one embodiment of the present application, a method for calculating an average signal-to-noise ratio is as follows:

[0088]

[0089] in, is the number of points in the point cloud in a frame of data, The first The signal-to-noise ratio of a point cloud.

[0090] A method for calculating weighted average distance is as follows:

[0091]

[0092] in, The first The Euclidean distance from the point cloud to the origin, The first The Cartesian space coordinates of the point cloud points, It is the weighted average distance of all point clouds in a frame of data.

[0093] A method for calculating weighted average speed is as follows:

[0094]

[0095] in, The first The radial velocity of the point cloud relative to the millimeter wave radar sensor, It is the weighted average velocity of all point clouds in a frame of data.

[0096] The weighted centroid coordinate calculation formula at the centroid is:

[0097]

[0098]

[0099]

[0100] in, 、 、 The first The Cartesian space coordinates of the centroid of the point cloud.

[0101] Apply the above method to calculate the weighted average distance, calculate the weighted average distance of all point clouds, and record it as (Range of Total Point); Apply the above method to calculate the weighted average velocity, calculate the weighted average velocity of all point clouds, denoted as (Velocity of Total Point); Apply the above method to calculate the weighted average distance, and calculate the weighted average distance of the point cloud corresponding to the radial velocity greater than 0 (Range of Centrifuge Point); Apply the above method to calculate the weighted average distance and calculate the weighted average velocity of the point cloud corresponding to the radial velocity greater than 0 (Velocity of Centrifuge Point); Apply the above method to calculate the weighted average distance to calculate the weighted average distance of the point cloud corresponding to the radial velocity less than 0 (Range of Centripetal Point); Apply the above method to calculate the weighted average distance and calculate the weighted average velocity of the point cloud corresponding to the radial velocity less than 0 (Velocity of Centripetal Point).

[0102] In one embodiment of the present application, the 13 data obtained after calculating the point cloud in a frame are 、 、 、 、 、 、 、 、 、 、 、 、 , which constitute the feature vectors of the frame point cloud. These feature vectors jointly describe the spatial distribution, motion state and weighted center position of the frame point cloud.

[0103] When constructing the sliding window matrix, the window size is set to 10 frames, the sliding step is set to 1 frame, and the 13-dimensional feature vector of each frame is used as a column of the matrix. The features of 10 consecutive frames are arranged to form a 13×10 millimeter-wave radar time series feature matrix. The time point corresponding to the last column in the feature matrix is , the time point corresponding to the first column is ,Right now:

[0104]

[0105] in, Indicates the end time of the current sliding time window, Indicates the starting time of the current sliding time window.

[0106] As the sliding window advances with a step size of 1 frame, the data is updated one frame at a time and a new feature matrix is formed. This method can capture the dynamic characteristics of gestures in the temporal dimension and provide high-dimensional feature input for subsequent classification and recognition.

[0107] S3: Based on the eigenvector matrix of the millimeter-wave radar data, collect millimeter-wave radar data of different gestures, and manually annotate the millimeter-wave radar data of different gestures using an annotation tool to obtain a millimeter-wave radar gesture dataset.

[0108] In one embodiment of the present application, the dataset construction process is as follows: (1) Millimeter wave radar selection and configuration: The Texas Instruments IWRL6432BOOST millimeter wave radar device is selected, which has an operating frequency range of 57Ghz-61Ghz and can capture small changes in gestures with high resolution. The radar device is installed in a fixed position to cover the target gesture collection area to ensure that the gestures can be fully presented within the effective detection range of the radar; (2) Gesture sample collection: In this embodiment, five gestures are designed, including waving, grasping, rotating, pushing, and pulling, which can cover the common gesture types in daily interactions. At least 10 volunteers are invited to participate in data collection to ensure the diversity and representativeness of the data. Volunteers are required to perform according to the preset gestures within the radar detection area. Each gesture is repeated at least 20 times to ensure the richness of the data. Richness and consistency. During the acquisition process, the radar equipment records the echo signal of the gesture action in real time at a frequency of 10Hz to generate millimeter-wave radar data; (3) Radar data annotation: Use professional annotation tools to annotate the collected millimeter-wave radar data. The annotator accurately marks the start and end timestamps of each gesture action in the time series of the radar data according to the preset gesture categories (such as waving, grasping, rotating, pushing, and pulling), and annotates the corresponding gesture name. During the annotation process, the accuracy and consistency of the annotation results are ensured through multi-person annotation and cross-validation; (4) Dataset construction: The annotated dataset is divided into training set, validation set and test set, usually in a ratio of 70%, 15% and 15%. The training set is used for model training, the validation set is used for model parameter adjustment and selection, and the test set is used to evaluate the final performance of the model.

[0109] S4: Build a CNN-LSTM hybrid neural network model.

[0110] In one embodiment of the present application, Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a gesture recognition model based on a CNN-LSTM hybrid neural network provided in an embodiment of the present application. This CNN-LSTM hybrid neural network model combines a convolutional neural network (CNN) and a long short-term memory network (LSTM) to train the dataset obtained in step S3. The CNN-LSTM hybrid neural network model is composed as follows:

[0111] (1) Input module: Receives the millimeter-wave radar time series feature vector of shape [10, 13] as input data, where 10 represents the number of time steps and 13 represents the number of features per time step;

[0112] (2) Preliminary feature extraction module: Through one-dimensional convolution operation, 13 convolution kernels of size 3 are used, with a stride of 2, no padding, and a linear activation function to preserve the original distribution of features. The output shape of this module is [4, 13], which is used to extract preliminary features.

[0113] (3) Feature processing and enhancement module: It includes two submodules, each of which consists of a one-dimensional convolution layer, a batch normalization layer, and the first activation layer;

[0114] Among them, the one-dimensional convolution layer: uses 24 convolution kernels of size 3, a stride of 1, adopts the 'same' padding method, the activation function is linear, and the output shape is [4, 24]. The batch normalization layer: is used to standardize the output of the previous layer to improve the stability of the model, and the output shape is [4, 24]. The first activation layer: introduces nonlinearity through the ReLU activation function to enhance the expressiveness of the model, and the output shape is [4, 24].

[0115] (4) Time series feature extraction module: uses a long short-term memory layer (LSTM) with 12 units, uses the hyperbolic tangent function (tanh) as the activation function, the recurrent activation function is Sigmoid, and the output shape is [4, 12] to capture long-term dependencies in sequence data;

[0116] (5) Classification decision module: including the fully connected layer and the second activation layer;

[0117] Among them, the fully connected layer is configured with an output size of 6, uses a linear activation function, and has an output shape of 6 to generate scores for each category; the second activation layer uses the Softmax function for activation to convert the output of the fully connected layer into a probability distribution to facilitate classification decisions.

[0118] These layers together constitute the key part of the CNN-LSTM hybrid neural network model of this application, which is used to extract features from the input data, capture long-term dependencies in time series data through the LSTM layer, and finally generate the probability distribution results of gesture classification through the fully connected layer and the Softmax layer.

[0119] S5: Input the millimeter-wave radar gesture dataset into the CNN-LSTM hybrid neural network model for training to obtain a Keras model.

[0120] In one embodiment of the present application, after constructing the CNN-LSTM hybrid neural network model, the millimeter-wave radar gesture dataset obtained in step S3 is input into the neural network model obtained in step S4 for training. The TensorFlow machine learning framework is used for training during the training process, and a Keras model is obtained after training.

[0121] In step S3, the dataset of this embodiment is divided into a training set, a validation set, and a test set, with ratios of 70%, 15%, and 15%, respectively. On this basis, this embodiment collected 1,000 sets of gesture data, each set continuously collecting gesture signals for 10 seconds, and divided them according to the above ratios. After testing, the average gesture recognition accuracy of the neural network model of this embodiment on the test set was 87.5%, indicating that the model has good recognition performance and generalization ability.

[0122] S6: Convert the Keras model into a lightweight neural network model suitable for the embedded end.

[0123] In one embodiment of the present application, a trained Keras model is converted into model code suitable for embedded devices. This process involves converting the model into an optimized format to ensure that it can run efficiently in resource-constrained embedded systems. The converted model will be further optimized to meet the performance and storage requirements of the target device. The specific steps are as follows:

[0124] A1: Use TensorFlow Lite Converter to convert the Keras model to a TensorFlow Lite model. That is, the Keras model is converted into a .tflite format model file.

[0125] A2: Convert the TensorFlow Lite model to a TensorFlow Lite Micro (lightweight neural network model) model. This means converting the generated .tflite model file into a TensorFlow Lite Micro model suitable for embedded devices.

[0126] The specific steps are: Use TensorFlow Lite Converter to convert the .h5 model file into a .tflite model file. Then, use the tools provided by TensorFlow Lite Micro to convert the .tflite file into C language code. This tool extracts the model's parameters and structure and generates code that can run on embedded devices. Integrate the generated code into the embedded development environment and use the runtime library provided by the TensorFlow Lite Micro framework to complete the model deployment.

[0127] Through the above steps, this application realizes the efficient conversion of Keras models into TensorFlow Lite Micro models suitable for embedded devices, providing technical support for neural network applications on embedded devices.

[0128] S7: Deploy the lightweight neural network model on a millimeter-wave radar chip with an Arm Cortex-M core, and perform gesture recognition based on the millimeter-wave radar chip.

[0129] In one embodiment of the present application, the converted neural network model is deployed on a millimeter-wave radar chip with an Arm Cortex-M core and an operating frequency of 60 GHz.

[0130] This embodiment uses the Texas Instruments IWRL6432 millimeter-wave radar chip, which integrates an Arm Cortex-M core and performs raw data acquisition, data preprocessing, and computational processing of embedded neural network models. By integrating these functions into a single chip, it significantly reduces data transmission latency and energy consumption, improving overall system efficiency while also lowering hardware costs and system complexity, and enhancing system reliability and integration.

[0131] In one embodiment of the present application, a gesture tester faces the millimeter-wave radar transceiver system and randomly performs a series of predefined gestures. These gestures include but are not limited to waving, grasping, rotating, pushing, pulling, etc., covering common gesture types in daily interactions. After the millimeter-wave radar receives the data, it is processed by the radar data and the neural network model deployed on the embedded end, and then the result of gesture judgment at the current moment is output.

[0132] Example 2:

[0133] The embodiment of the present application provides a gesture recognition system based on lightweight computing and storage, which can be seen in Figure 5 , Figure 5A schematic diagram of a gesture recognition system based on lightweight computing and storage provided in an embodiment of the present application, comprising: a millimeter-wave radar data acquisition module, a data processing module, a gesture recognition module, a model training and conversion module, and a judgment result display module;

[0134] The millimeter-wave radar data acquisition module, data processing module and gesture recognition module are integrated on the same millimeter-wave radar chip with an Arm Cortex-M core, wherein the gesture recognition module is deployed on and runs on the Arm Cortex-M processor.

[0135] The millimeter wave radar data acquisition module includes a millimeter wave radar transceiver system for collecting millimeter wave radar raw data.

[0136] In one embodiment of the present application, the system constructed in this embodiment uses the Texas Instruments IWRL6432 millimeter-wave radar chip, which integrates an RF transmission front end, an analog-to-digital converter (ADC), a digital signal processing unit, and multiple functional modules, and can efficiently process raw data and generate point cloud data.

[0137] The data processing module is used to preprocess and extract features from the millimeter-wave radar raw data to generate millimeter-wave radar gesture data that can be used for training and recognition, thereby improving data availability and recognition accuracy.

[0138] The model training and conversion module includes a model training unit and a model conversion unit, wherein the model training unit uses the millimeter-wave radar gesture data set of the data processing module to train the CNN-LSTM hybrid neural network model; the model conversion unit is used to convert the trained neural network model into a lightweight embedded model and optimize it to adapt to the computing and storage requirements of the embedded device.

[0139] In one embodiment of the present application, in order to achieve efficient deployment of the neural network model in an embedded system, the original Keras model is converted and optimized. Through a series of technical means, including model compression, quantization and structural optimization, the size of the model is significantly reduced. Specifically, the volume of the optimized model is reduced by 96% compared to the original Keras model, and the final model size is 60KB. This optimized model can be stored in the Flash memory of the embedded system, which meets the requirements of the neural network model running on the embedded end. This significant size reduction not only improves the storage feasibility of the model on the embedded device, but also reduces the memory usage and energy consumption during the model loading and running process, further improving the overall operating efficiency of the system.

[0140] The gesture recognition module includes a data input unit and a neural network inference unit, wherein the data input unit is used to receive the millimeter-wave radar feature vector matrix data output by the data processing module; the neural network inference unit is used to perform neural network operations based on a lightweight embedded model and output the gesture type.

[0141] In one embodiment of the present application, the module is deployed on an Arm Cortex-M processor and integrated with other key modules on the same millimeter-wave radar chip to improve system integration and operational efficiency.

[0142] The judgment result display module includes a display unit for displaying the gesture recognition result in real time according to the recognition result output by the gesture recognition module, wherein the gesture recognition result includes the gesture type and the corresponding gesture action information.

[0143] In one embodiment of the present application, the judgment result display module uses an OLED screen to clearly and in real time display the gesture recognition rate and gesture name output by the millimeter-wave radar gesture recognition system, providing users with intuitive and convenient interactive feedback. The display unit can be in the form of a screen, indicator light, or other visual display device.

[0144] The present application provides a gesture recognition method and system based on lightweight computing and storage. By constructing a feature vector matrix of millimeter-wave radar data, an innovative balance is achieved between data compression and feature retention. The matrix can fully capture the spatiotemporal characteristics of gesture movement while retaining only a portion of the data. This efficient representation method enables the millimeter-wave radar gesture recognition system to run in real time on resource-constrained embedded platforms while avoiding the delays and privacy issues caused by cloud processing. At the same time, the constructed CNN-LSTM hybrid neural network model is optimized into a Keras model. The optimized model has a smaller model size and reduced computational complexity, allowing the model to run efficiently on embedded devices without relying on high-performance computing platforms or cloud resources, thereby significantly reducing system power consumption and deployment costs. Moreover, by converting the Keras model into a lightweight neural network model suitable for embedded devices, the model size is reduced, significantly reducing the computing power and storage requirements for deploying neural networks on the embedded side. In addition, deploying a lightweight neural network model on a millimeter-wave radar chip with an Arm Cortex-M core can quickly extract hand feature vectors and perform real-time inference through an optimized neural network, thereby effectively reducing recognition latency and ensuring high response speed and accurate recognition even in resource-constrained environments.

[0145] The present invention provides an efficient and reliable solution for human-computer interaction of smart devices, has wide applicability and good user experience.

[0146] It should be noted that those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of this application, and it should be understood that the scope of protection of this application is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in this application without departing from the essence of this application, and such variations and combinations are still within the scope of protection of this application.

Claims

1. A gesture recognition method based on lightweight computing and storage, characterized in that: include: S1: Use the millimeter radar to collect millimeter-wave radar raw data and generate millimeter-wave radar point cloud data based on the radar data link; S2: extracting feature vectors of the millimeter-wave radar data based on the millimeter-wave radar point cloud data, and constructing a feature vector matrix of the millimeter-wave radar data to capture dynamic changes in gestures; S3: Based on the eigenvector matrix of the millimeter-wave radar data, collect millimeter-wave radar data of different gestures, and manually annotate the millimeter-wave radar data of the different gestures using an annotation tool to obtain a millimeter-wave radar gesture dataset; S4: Build a CNN-LSTM hybrid neural network model; S5: Inputting the millimeter-wave radar gesture dataset into the CNN-LSTM hybrid neural network model for training to obtain a Keras model; S6: Convert the Keras model into a lightweight neural network model suitable for the embedded end; S7: deploying the lightweight neural network model on a millimeter-wave radar chip with an Arm Cortex-M core, and performing gesture recognition based on the millimeter-wave radar chip; In S6, the Keras model is converted into a lightweight neural network model suitable for the embedded end, specifically including: A1: Use the model conversion tool to convert the Keras model into a lightweight TensorFlow Lite model. A2: Optimize the TensorFlow Lite model to meet the performance and storage requirements of the embedded device. A3: Use the TensorFlow Lite Micro framework to convert the optimized TensorFlow Lite model into a lightweight neural network model.

2. The gesture recognition method based on lightweight computing and storage according to claim 1, characterized in that: The millimeter-wave radar point cloud data in S1 includes spatial coordinates, speed, and signal-to-noise ratio information.

3. The gesture recognition method based on lightweight computing and storage according to claim 2, characterized in that: The S2 includes: S201: Record the total number of point clouds at the current moment, and respectively record the number of point clouds with radial velocities greater than 0 and less than 0 at the current moment: 、 、 ; S202: Calculate the average signal-to-noise ratio, weighted average distance, and weighted average speed of all point cloud data based on the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 ; S203: Calculate the weighted average distance and weighted average velocity of the point clouds corresponding to radial velocities greater than 0 and radial velocities less than 0, respectively, based on the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 、 ; S204: Calculate the weighted centroid coordinates at the centroid of the millimeter-wave radar point cloud data according to the signal-to-noise ratio in the millimeter-wave radar point cloud data: 、 、 ; S205: Arrange the number of point clouds, average signal-to-noise ratio, weighted average distance, weighted average velocity, and weighted centroid coordinates into a column vector to form a feature vector of the millimeter-wave radar data: in, is the total number of point clouds in a frame of data, is the number of point clouds with radial velocity values greater than 0 in the total number of point clouds, is the number of point clouds with radial velocity values less than 0 in the total number of point clouds, is the average signal-to-noise ratio of all point clouds, is the weighted average distance of all point clouds, is the weighted average velocity of all point clouds, is the weighted average distance of the point cloud corresponding to the radial velocity greater than 0, is the weighted average velocity of the point cloud corresponding to the velocity greater than 0, is the weighted average distance of the point cloud corresponding to the radial velocity less than 0, is the weighted average velocity of the point cloud corresponding to the radial velocity being less than 0, 、 、 is the Cartesian space coordinate of the weighted centroid, for The feature vector of the millimeter-wave radar data at the moment; S206: Based on the eigenvectors of the millimeter-wave radar data, construct an eigenvector matrix of the millimeter-wave radar data by using a sliding time window technology.

4. The gesture recognition method based on lightweight computing and storage according to claim 3, characterized in that: The step S206 constructs the characteristic vector matrix of the millimeter wave radar data by using the sliding time window technology, including: Using the sliding time window technique, the feature vectors of continuously collected millimeter-wave radar data are arranged in chronological order. The 13-dimensional feature vector of each frame is used as a column of the matrix. The features of 10 consecutive frames are arranged to form a 13×10 feature matrix: in, Indicates the end time of the current sliding time window, Indicates the starting time of the current sliding time window.

5. The gesture recognition method based on lightweight computing and storage according to claim 3, characterized in that: The calculation method of the average signal-to-noise ratio, weighted average distance, weighted average speed and weighted center of mass coordinates at the center of mass includes: The formula for calculating the average signal-to-noise ratio is: in, is the number of points in the point cloud in a frame of data, The first The signal-to-noise ratio of the point cloud; The calculation formula of weighted average distance is: in, The first The Euclidean distance from the point cloud to the origin, The first The Cartesian space coordinates of the point cloud points, is the weighted average distance of all point clouds in a frame of data; The formula for calculating the weighted average speed is: in, The first The radial velocity of the point cloud relative to the millimeter wave radar sensor, is the weighted average velocity of all point clouds in a frame of data; The weighted centroid coordinate calculation formula at the centroid is: in, 、 、 The first The Cartesian space coordinates of the centroid of the point cloud.

6. The gesture recognition method based on lightweight computing and storage according to claim 1, characterized in that: The CNN-LSTM hybrid neural network model includes: Input module, used to receive a feature matrix with a shape of [10, 13]; The preliminary feature extraction module converts the feature matrix with an input shape of [10, 13] into a preliminary feature matrix with an output shape of [4, 13] through a one-dimensional convolution operation; The feature processing and enhancement module includes two submodules, each of which consists of a one-dimensional convolutional layer, a batch normalization layer, and a first activation layer. The one-dimensional convolutional layer uses 24 convolution kernels of size 3, a step size of 1, and a linear activation function to convert the preliminary feature matrix [4, 13] into a feature matrix with an output shape of [4, 24]. The batch normalization layer standardizes the output of the one-dimensional convolutional layer to an output shape of [4, 24]. The first activation layer introduces nonlinearity through the ReLU activation function to convert the output of the batch normalization layer into a nonlinear output with an output shape of [4, 24]. A time series feature extraction module includes a long short-term memory layer, which uses a hyperbolic tangent function and a sigmoid function as activation functions and a recurrent activation function to capture long-term dependencies in sequence data; The classification decision module includes a fully connected layer and a second activation layer, wherein the fully connected layer is configured to have an output size of 6, uses a linear activation function, and has an output shape of 6 to generate a score for each category; the second activation layer uses a Softmax function for activation to convert the output of the fully connected layer into a probability distribution for gesture classification.

7. The gesture recognition method based on lightweight computing and storage according to claim 1, characterized in that: In S5, the millimeter-wave radar gesture dataset is input into the CNN-LSTM hybrid neural network model for training to obtain a Keras model, including: The CNN-LSTM hybrid neural network model is trained using the TensorFlow machine learning framework to obtain a Keras model.

8. A system according to any one of claims 1 to 7, characterized in that: include: A millimeter-wave radar data acquisition module, data processing module, and gesture recognition module integrated on the same millimeter-wave radar chip with an Arm Cortex-M core; The millimeter wave radar data acquisition module includes a millimeter wave radar transceiver system for collecting millimeter wave radar raw data; The data processing module is used to preprocess and extract features from the millimeter-wave radar raw data to generate millimeter-wave radar gesture data that can be used for training and recognition; A model training and conversion module, comprising a model training unit and a model conversion unit, wherein the model training unit uses the millimeter-wave radar gesture dataset of the data processing module to train a CNN-LSTM hybrid neural network model; the model conversion unit is used to convert the trained neural network model into a lightweight embedded model and optimize it to meet the computing and storage requirements of the embedded device; The gesture recognition module includes a data input unit and a neural network inference unit, wherein the data input unit is used to receive the millimeter wave radar feature vector matrix data output by the data processing module; the neural network inference unit is used to perform neural network operations based on a lightweight embedded model and output the gesture type; The judgment result display module includes a display unit for displaying the gesture recognition result in real time according to the recognition result output by the gesture recognition module, wherein the gesture recognition result includes the gesture type and the corresponding gesture action information.

9. The system according to claim 8, characterized in that The gesture recognition module is deployed on the Arm Cortex-M processor of the millimeter wave radar chip and runs on the processor.

Citation Information

Patent Citations

  • Gesture recognition using 3D mm-wave radar

    CN110941331A

  • Multi-feature lightweight gesture recognition method based on millimeter wave radar

    CN117935368A