Adaptive data labeling method and system for millimeter wave radar gesture recognition

By using an adaptive data annotation method, and leveraging the maximum signal-to-noise ratio point and distance and velocity features, the problem of poor consistency of training samples in millimeter-wave radar gesture recognition was solved, achieving high-precision gesture recognition.

CN121542746BActive Publication Date: 2026-05-01POSSUMIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
POSSUMIC TECH CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing millimeter-wave radar gesture recognition technology, the training samples have poor consistency and are easily affected by noise, resulting in inaccurate recognition results. In particular, since the echo energy of human targets is stronger than that of gesture echoes, and the environmental noise is complex, the data labeling aliasing phenomenon is serious.

Method used

By controlling the radar to repeatedly collect hand gestures, the point with the highest signal-to-noise ratio is selected as the feature point. Combining distance, speed and signal-to-noise ratio features, adaptive data annotation of hand gestures is performed in the time dimension. The range of distance slices is divided and point cloud subset segmentation is performed to generate high-quality data samples.

Benefits of technology

It effectively filters out environmental noise and human target interference, improves the accuracy of gesture recognition, generates unbiased high-quality datasets, and enhances the robustness of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542746B_ABST
    Figure CN121542746B_ABST
Patent Text Reader

Abstract

An adaptive data annotation method and system for millimeter-wave radar gesture recognition includes: S100, collecting the same gesture M times within time L, and obtaining... Frame point cloud data; S200, select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, statistically analyze the distance between the feature points, and determine the estimated distance to the human target. S300, Estimated distance to human target Using a base point, N range slices are defined in the direction of the radar, and the point cloud data is divided into N point cloud subsets according to the N range slices; S400, the maximum velocity change rate of each point cloud subset is calculated, and the point cloud subset with the most significant change is selected as the data annotation point cloud subset; S500, the gesture action time annotation is performed by combining the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset; S600, based on the gesture action time annotation points and the N point cloud subsets, the annotated gesture action data samples are obtained. This invention improves the accuracy of gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

An Adaptive Data Labeling Method and System for Millimeter-Wave Radar Gesture Recognition Technical Field

[0001] This invention belongs to the field of radar detection technology, specifically relating to an adaptive data annotation method and system for millimeter-wave radar gesture recognition. Background Technology

[0002] Gestures can easily and intuitively express the intent of commands, making their recognition an indispensable research area in human-computer interaction. Gesture recognition applications based on millimeter-wave radar are increasingly popular due to their strong environmental adaptability, privacy protection, and non-contact nature. Currently, gesture recognition for millimeter-wave radar often involves extracting feature vectors from radar echo signals using neural networks, and then using deep learning techniques to classify and recognize these feature vectors. To ensure the accuracy and robustness of the application, training the neural network requires a large amount of accurately labeled radar gesture data as samples. However, current research focuses primarily on innovation in network design, with training samples often extracted entirely from the gesture command or labeled based on a single radar data feature. This results in poor sample consistency and susceptibility to noise interference.

[0003] In the process of annotating radar gesture action point cloud data, the human target is larger, and its echo energy is much stronger than the gesture echo energy. Furthermore, the acquisition environment is complex and subject to significant noise interference. Annotating data based on only a single feature is highly susceptible to training data aliasing, affecting recognition results. For example, some current radar gesture action annotation methods do not perform gesture target distance detection, directly using single feature information for data annotation in the time dimension, easily leading to aliasing between human targets and gesture targets. Other methods directly select the position with the maximum absolute velocity value as the gesture action distance, without considering the potential interference from the surrounding environment. Summary of the Invention

[0004] This invention provides an adaptive data annotation method and system for gesture recognition using millimeter-wave radar, aiming to improve the accuracy of gesture recognition. The technical solution for implementing this invention is as follows:

[0005] In a first aspect, the present invention provides an adaptive data annotation method for millimeter-wave radar gesture recognition, comprising:

[0006] S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain... Frame point cloud data;

[0007] S200. Select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, and statistically analyze the distance between the feature points to determine the estimated distance to the human target. ;

[0008] S300, Estimated distance to human target Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0009] S400. Calculate the maximum rate of change of velocity for each point cloud subset, select the point cloud subset with the most significant change as the data label point cloud subset, and use the center of its corresponding distance slice as the estimated distance of the gesture action. ;

[0010] S500. Combining the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset, perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0011] S600. Based on the time markers of the gesture actions, the N point cloud subsets are segmented to obtain millimeter-wave radar gesture action data samples with labeled multi-distance slice ranges.

[0012] As a preferred technical solution, before step S100, it is first determined whether the human target within a preset range in front of the radar is in a stable state. If so, step S100 is executed; otherwise, step S100 is postponed. The method for determining the stable state includes:

[0013] Set human body stability detection signal Record point cloud data, and take the maximum signal-to-noise ratio of each frame of point cloud data to form a feature vector. ,in Representing the The maximum signal-to-noise ratio of frame point cloud data, K is Preset size Updated in real time as time increases. Calculate its effective mean ,in , For neighbors The number of point cloud frames with a maximum non-zero signal-to-noise ratio, and based on... Calculate vectors If the L1 norm is less than a preset threshold, then It determines whether the human target is in a stable state or not, and otherwise determines whether the human target is in a non-stable state.

[0014] As a preferred technical solution, step S200, the method for determining the estimated distance to the human target includes: […]. The distance vector of feature points in a frame point cloud data is denoted as ,in This represents the distance value of the feature point in frame t. and the distance vector Statistical analysis was conducted, and the average effective distance was taken as the estimated distance to the human target. In this process, the effective distance for calculation is limited. Only when the signal-to-noise ratio of a feature point, i.e., the maximum signal-to-noise ratio of the corresponding frame's point cloud data, meets the threshold range is the corresponding distance value considered as an effective distance. , for The number of frames in the point cloud data whose maximum signal-to-noise ratio meets the threshold range.

[0015] As a preferred technical solution, in step S300, the estimated distance to the human target is used. Using the radar as a base point, the method for defining N range slices in the direction of the radar is as follows: using the estimated distance to the human target... Using the distance as the base point, divide the distance slices along the distance dimension. ,in , The distance is the width of the slice. The distance is the distance between the centers of adjacent distance slices, and there is an overlapping interval between adjacent distance slices; the point cloud data is divided into N point cloud subsets according to the range of the N distance slices.

[0016] As a preferred technical solution, step S400 specifically includes:

[0017] S410, Define the velocity expression function ,in For the first The maximum velocity value of frame point cloud data. Its absolute value, ε is the set speed threshold. Characterized by the first [unit] within the distance slice range Does the maximum speed of the frame point cloud data meet the judgment requirements?

[0018] S420. During the motion acquisition time, the maximum velocity of the point cloud subsets across all slice ranges is statistically analyzed, and a statistical function is defined. , Characterization in the The maximum rate of change of velocity of a subset of point clouds within a distance slice. The number of frames for a subset of the point cloud;

[0019] S430, Select The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. , , The distance to the center of adjacent slices is used to divide the slice range using step S300. A subset of the point cloud is used as a data annotation subset of the point cloud. This is the width of the slice.

[0020] As a preferred technical solution, step S500 specifically includes:

[0021] S510. The data annotation point cloud subset is segmented according to the time of the action command, to obtain M segments of gesture action point cloud subset data.

[0022] S520. Based on the speed of the action, the data annotation point cloud subset of each gesture action is filtered, and the maximum speed of each frame of the data annotation point cloud subset is selected to form the speed feature vector of the gesture action. The maximum value of the speed feature vector is selected as the candidate point for action time annotation.

[0023] S530. Based on the signal-to-noise ratio, which represents the magnitude of the action energy, the candidate points for action time annotation are screened to determine the gesture action time annotation point.

[0024] As a preferred technical solution, in step S530, the operation of filtering each segment of gesture action point cloud data based on the signal-to-noise ratio (SNR) representing the magnitude of action energy specifically includes: selecting the maximum SNR of each frame of point cloud subset to construct an SNR feature vector. Differential processing is performed on the signal-to-noise ratio eigenvectors. The gesture action time marker point must satisfy the corresponding signal-to-noise ratio difference vector value being greater than zero, and at least one neighboring point's difference vector value being greater than zero; the first velocity maximum point that satisfies the signal-to-noise ratio difference vector condition in chronological order is selected as the gesture action time marker point for that segment of the gesture.

[0025] As a preferred technical solution, the gesture action data samples in step S600 include positive samples. The method for obtaining the positive samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; segmenting the feature map according to the gesture action time annotation points, and incorporating the segmented map into the database as positive samples of the gesture action.

[0026] As a preferred technical solution, the gesture action data samples in step S600 include negative samples. The method for obtaining the negative samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; inferring multiple intervals as no-action intervals between adjacent gesture action time annotation points, and taking the starting point of the no-action interval as the negative sample annotation point; segmenting the feature map according to the negative sample annotation point, and incorporating the segmented map into the database as negative samples of gesture actions.

[0027] Secondly, the present invention provides an adaptive data annotation system for millimeter-wave radar gesture recognition, comprising:

[0028] The point cloud data acquisition module controls the radar to repeatedly collect the same hand gesture M times within a time period L, thus obtaining... Frame point cloud data;

[0029] The human target distance estimation module selects the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, calculates the distance between the feature points, and determines the estimated human target distance. ;

[0030] The point cloud subset acquisition module uses the estimated distance to the human target. Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0031] The gesture distance estimation module calculates the maximum rate of change of velocity for each point cloud subset, selects the point cloud subset with the most significant change as the data annotation subset, and uses the center of its corresponding distance slice as the estimated gesture distance value. ;

[0032] The gesture action time annotation module combines the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset to perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0033] The data sample acquisition module segments N point cloud subsets based on the time markers of the gesture actions to obtain millimeter-wave radar gesture action data samples with multiple distance slice ranges.

[0034] The adaptive data annotation method and system for millimeter-wave radar gesture recognition provided by this invention have the following advantages: by using the three-dimensional radar data features of distance, speed, and signal-to-noise ratio, environmental noise is filtered out and human target interference is shielded, and adaptive data annotation of gesture action data is achieved to obtain unbiased, high-quality millimeter-wave radar gesture action dataset, which can greatly improve the accuracy of gesture recognition. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 is a flowchart of an adaptive data annotation method for millimeter-wave radar gesture recognition provided in an embodiment of the present invention.

[0037] Figure 2 is a flowchart of obtaining the gesture distance estimate in the adaptive data annotation method for millimeter-wave radar gesture recognition provided in an embodiment of the present invention.

[0038] Figure 3 is a flowchart of the gesture action time annotation method for millimeter-wave radar gesture recognition provided in an embodiment of the present invention.

[0039] Figure 4 is a block diagram of the adaptive data annotation system for millimeter-wave radar gesture recognition provided in an embodiment of the present invention. Detailed Implementation

[0040] To make the technical solution of the present invention clearer and its technical advantages more apparent, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of the present invention.

[0041] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0042] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0043] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0044] It is understandable that radar receives echoes and processes them to generate point cloud data. Since human targets are large and their echo energy is strong, they are very likely to interfere with the echo signals of hand gestures. Therefore, distance estimation is essentially about refining the distance dimension range. This measure can extract the distance range where the changes in hand gestures are most significant to the greatest extent. Subsequent action time labeling within this range can effectively avoid inaccurate labeling caused by background noise and human interference.

[0045] The specific embodiments of the present invention provide an adaptive data annotation method and system for millimeter-wave radar gesture recognition. In order to improve the accuracy of gesture recognition, the main approaches are as follows: First, determining the range of motion of the gesture (distance detection) and distinguishing between human targets and gesture targets is the first step in whether the data annotation can be successful; Second, based on determining the distance of the gesture, a detection algorithm combining signal-to-noise ratio information and velocity information is used to annotate the gesture data stream in the time dimension.

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0047] Referring to Figure 1, as a basic implementation method, the adaptive data annotation method for millimeter-wave radar gesture recognition provided in this embodiment includes:

[0048] S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain... Frame point cloud data;

[0049] S200. Select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, and statistically analyze the distance between the feature points to determine the estimated distance to the human target. ;

[0050] S300, Estimated distance to human target Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0051] S400. Calculate the maximum rate of change of velocity for each point cloud subset, select the point cloud subset with the most significant change as the data label point cloud subset, and use the center of its corresponding distance slice as the estimated distance of the gesture action. ;

[0052] S500. Combining the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset, perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0053] S600. Based on the time markers of the gesture actions, the N point cloud subsets are segmented to obtain millimeter-wave radar gesture action data samples with labeled multi-distance slice ranges.

[0054] When the millimeter-wave radar collects hand gestures from a human target, it repeats each hand gesture M times according to the command signal, and records the total collection time of the gesture as time L. The total time L is then roughly divided into M segments using the time points when the command signal occurs. The time period of the gesture action is recorded as follows: , Positive integers; gesture annotation involves selecting the time point with the most significant action feature within each time period as the action annotation point.

[0055] Taking the FMCW millimeter-wave radar as an example, its transmission carrier frequency is 60GHz. It collects hand gestures made by targets within a range of 0.5m to 2m in front of the radar. The radar frame period is set according to the complexity of the action. Generally, a frame period of 50ms can accurately reflect commonly defined hand gestures, such as waving to the left (right) or waving forward (backward). The point cloud feature map size is set to size=[number_chirp, length_capture], where number_chirp is the number of radar chirps and length_capture is the number of frames captured for the detected action. Generally, size=[64,32] can more completely reflect the action characteristics.

[0056] Furthermore, to achieve optimal data acquisition results when using millimeter-wave radar to collect gesture data, a relatively static environment is generally required, meaning the target object must be in a stable state. Therefore, the target (human body) needs to remain within a preset range in front of the radar for a period of time before data acquisition. Thus, as a preferred implementation, the adaptive data annotation method for millimeter-wave radar gesture recognition first determines whether the human target within the preset range in front of the radar is in a stable state before step S100. If so, step S100 is executed; otherwise, step S100 is postponed.

[0057] Specifically, the method for determining the stable state includes: setting a human body stability detection signal. Record point cloud data, and take the maximum signal-to-noise ratio of each frame of point cloud data to form a feature vector. ,in Representing the Maximum signal-to-noise ratio of frame point cloud data, Preset size , Updated in real time as time increases. Calculate its effective mean ,in , For neighbors The number of point cloud frames with a maximum non-zero signal-to-noise ratio, and based on... Calculate vectors If the L1 norm is less than a preset threshold, then It was determined that the human target was in a stable state, and subsequent data collection could proceed.

[0058] After receiving the target stabilization signal feedback, the acquisition system begins to issue motion acquisition commands, using a timer to record the motion time range corresponding to the gesture commands, such as the first... The time period of the gesture action is recorded as follows: , A positive integer, after collecting a preset number of times M, the collection stops and the timing stops, and the distance to the human target is estimated from the data within the operation time L.

[0059] In step S200, the method for determining the estimated distance to the human target includes: […]. The distance vector of feature points in a frame point cloud data is denoted as ,in This represents the distance value of the feature point in frame t. and the distance vector Statistical analysis was conducted, and the average effective distance was taken as the estimated distance to the human target. Furthermore, the effective distances involved in the calculation can be limited. Only when the signal-to-noise ratio of a feature point, i.e., the maximum signal-to-noise ratio of the corresponding frame's point cloud data, meets a threshold range, will its corresponding distance value be considered as an effective distance. , for The number of frames in the point cloud data whose maximum signal-to-noise ratio meets the threshold range.

[0060] In step S300, the estimated distance to the human target is used. Using the radar as a base point, the method for defining N range slices in the direction of the radar is as follows: using the estimated distance to the human target... Using the distance as the base point, divide the distance slices along the distance dimension. ,in , The distance is the width of the slice. It's generally best to set it to half an arm's length, for example, set... , The distance between the centers of adjacent distance slices is typically set to the radar range resolution; furthermore, there are overlapping regions between adjacent distance slices.

[0061] Referring to Figure 2, step S400 specifically includes:

[0062] S410, Define the velocity expression function ,in For the first Maximum frame rate Its absolute value, ε is the set speed threshold. Characterized within the range of the slice Does the maximum speed of the frame point cloud data meet the judgment requirements?

[0063] S420. Within the motion acquisition time L, the maximum velocity of each point cloud subset is statistically analyzed, and a statistical function is defined. , Characterization in the The maximum rate of change of velocity of a subset of point clouds within a distance slice. The number of frames for a subset of the point cloud;

[0064] S430, Select The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. , , The distance to the center of adjacent slices is used to divide the slice range using step S300. A subset of the point cloud is used as a data annotation subset of the point cloud.

[0065] Next, the data of the labeled point cloud subset is labeled along the time dimension. Because many environmental noises and human body shaking effects have been removed, the point cloud data (including signal-to-noise ratio, velocity, etc.) shows obvious fluctuations in the time dimension with the gesture movement, which can adaptively and accurately realize the action segmentation.

[0066] Referring to Figure 3, step S500 specifically includes:

[0067] S510. Segment the data annotation point cloud subset according to the time of the action command, and obtain M segments of gesture action point cloud subset data.

[0068] S520. Based on the speed of the action, the data annotation point cloud subset of each gesture action is filtered, and the maximum value of the speed feature vector is selected as the candidate point for action time annotation. Specifically, the maximum speed of each frame of the data annotation point cloud subset selected in step S430 is selected to form the speed feature vector of the gesture action, and the maximum value of the speed feature vector is selected as the candidate point for action time annotation.

[0069] S530. Based on the signal-to-noise ratio, which represents the magnitude of the action energy, the candidate points for action time annotation are screened to determine the gesture action time annotation point.

[0070] In step S520, the above-mentioned data annotation point cloud subset is taken, and for each segment of data from the M collection actions, the action start point is annotated: within the effective operation time of each gesture action recording, the maximum velocity of each frame of point cloud data is selected to form a velocity feature vector. ,in Representing the The maximum velocity value of the frame point cloud subset, selected The maximum points form the candidate point set.

[0071] In step S530, the operation of filtering the point cloud subset data of each gesture action based on the signal-to-noise ratio (SNR) representing the magnitude of the action energy specifically includes: the SNR of the point cloud subset data shows an increasing trend at the start of the action, therefore, the maximum SNR of each frame's point cloud subset is selected to form the SNR feature vector. Differential processing is performed on the signal-to-noise ratio eigenvectors. For each action time marker, the corresponding signal-to-noise ratio (SNR) difference vector value must be greater than zero, and at least one neighboring point must also have an SNR difference vector value greater than zero. The first velocity maximum point satisfying the SNR difference vector condition in chronological order is selected as the action time marker for that gesture segment. This operation is repeated M times to obtain all action time markers.

[0072] In step S600, the gesture action data samples include positive samples. The method for obtaining positive samples includes: splicing feature information from the point cloud subset data of N distance slice ranges in step S300 along the time dimension to generate a feature map; segmenting the feature map according to the action time markers, and incorporating the segmented map into the database as positive samples of the gesture action. The point cloud data features include distance, velocity, orientation angle, pitch angle, and signal-to-noise ratio, which are spliced ​​along the time dimension to form multiple feature maps, such as generating a range map (RTM), a Doppler map (DTM), and an angle map (ATM). Furthermore, the map width is the radar chirp count, and the length is the number of acquisition frames. Since the action markers are located at the start of the action, and gesture actions usually have a clearer intention within the initial time range, covering the first half of the action with acquisition frames is sufficient to meet the recognition requirements. After determining the accurate action time markers, processing point cloud data feature maps at different distance ranges can cover appropriate noise and human interference, obtaining feature information over a wider area to expand the database.

[0073] In step S600, with the determination of the action time markers, negative samples in the database can also be accurately collected. That is, the gesture action data samples include negative samples. The method for obtaining negative samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; between adjacent action time markers, using prior conditions to infer that multiple intervals within a reasonable time range are no-action intervals, and taking the starting point of the no-action interval as the negative sample marker; based on the negative sample markers, segmenting the feature map, and incorporating the segmented map into the database as negative samples of gesture actions.

[0074] In this way, multiple training sets containing noise and interference can be generated using the point cloud subsets under different distance slices, thereby augmenting the training data and enhancing the robustness of the trained neural network.

[0075] The present invention also provides an adaptive data annotation system for millimeter-wave radar gesture recognition, which executes the adaptive data annotation method for millimeter-wave radar gesture recognition described above.

[0076] Specifically, referring to Figure 4, the adaptive data annotation system for millimeter-wave radar gesture recognition includes:

[0077] The point cloud data acquisition module controls the radar to repeatedly collect the same hand gesture M times within a time period L, thus obtaining... Frame point cloud data;

[0078] The human target distance estimation module selects the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, calculates the distance between the feature points, and determines the estimated human target distance. ;

[0079] The point cloud subset acquisition module uses the estimated distance to the human target. Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0080] The gesture distance estimation module calculates the maximum rate of change of velocity for each point cloud subset, selects the point cloud subset with the most significant change as the data label point cloud subset, and uses the center of the corresponding distance slice as the gesture distance estimate.

[0081] The gesture action time annotation module combines the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset to perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0082] The data sample acquisition module segments N point cloud subsets based on the time markers of the gesture actions to obtain millimeter-wave radar gesture action data samples with multiple distance slice ranges.

[0083] The beneficial effects of the technical solution of the present invention are as follows: the present invention uses the three-dimensional radar data features of distance, speed and signal-to-noise ratio to filter environmental noise and shield human target interference, and achieves adaptive data annotation of gesture action data to obtain unbiased high-quality millimeter-wave radar gesture action dataset, which can greatly improve the accuracy of gesture recognition.

[0084] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. An adaptive data annotation method for millimeter-wave radar gesture recognition, characterized in that, include: S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain... Frame point cloud data; S200, select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, statistically analyze the distance between the feature points, and determine the estimated distance to the human target. S300, Estimated distance to human target Using the base point as the reference point, N range slices are defined along the range dimension towards the direction where the radar is located. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals; S400, calculate the maximum velocity change rate of each point cloud subset, and select the point cloud subset with the most significant change as the data labeling point cloud subset; During the motion capture time, a statistical function is used. The maximum speed of the point cloud subsets across all slice ranges was statistically analyzed, and selected... The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. S500. Based on the speed representing the speed of the action and the signal-to-noise ratio representing the magnitude of the action energy, the gesture action time is marked in the time dimension and the gesture action time mark points are recorded; S600. Based on the gesture action time mark points, the N point cloud subsets are segmented to obtain the millimeter-wave radar gesture action data samples with marked multi-distance slice range.

2. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, Before step S100, it is first determined whether the human target within a preset range in front of the radar is in a stable state. If so, step S100 is executed; otherwise, step S100 is postponed. The method for determining the stable state includes: setting a human stability detection signal. Record point cloud data, and take the maximum signal-to-noise ratio of each frame of point cloud data to form a feature vector. ,in Representing the The maximum signal-to-noise ratio of frame point cloud data, K is Preset size, increasing over time. Real-time updates Calculate its effective mean ,in , For neighbors The number of point cloud frames with a maximum non-zero signal-to-noise ratio, and based on... Calculate vectors If the L1 norm is less than a preset threshold, then It determines whether the human target is in a stable state or not, and otherwise determines whether the human target is in a non-stable state.

3. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, In step S200, the method for determining the estimated distance to the human target includes: […]. The distance vector of feature points in a frame point cloud data is denoted as ,in This represents the distance value of the feature point in frame t. and the distance vector Statistical analysis was conducted, and the average effective distance was taken as the estimated distance to the human target. In this process, the effective distances used in the calculation are limited; only when the signal-to-noise ratio of a feature point meets a threshold range are its corresponding distance values ​​considered as effective distances. , for The number of frames in the point cloud data whose maximum signal-to-noise ratio meets the threshold range.

4. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, In step S300, the estimated distance to the human target is used. Using the radar as a base point, the method for defining N range slices in the direction of the radar is as follows: using the estimated distance to the human target... Using the radar as the base point, range slices are divided along the range dimension towards the direction where the radar is located. ,in , The distance is the width of the slice. The distance is the distance between the centers of adjacent distance slices, and there is an overlapping interval between adjacent distance slices; the point cloud data is divided into N point cloud subsets according to the range of the N distance slices.

5. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, Step S400 specifically includes: S410, defining the velocity expression function. ,in For the first The maximum velocity value of frame point cloud data. Its absolute value, where ε is the set speed threshold. Characterized by the first [unit] within the distance slice range Does the maximum speed of the frame point cloud data meet the judgment requirements? S420. During the action acquisition time, the maximum speed of the point cloud subsets of all slice ranges is statistically analyzed, and a statistical function is defined. , Characterization in the The maximum rate of change of velocity of a subset of point clouds within a distance slice. The number of frames for a subset of the point cloud; S430, select The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. , , The distance to the center of adjacent slices is used to divide the slice range using step S300. A subset of the point cloud is used as a data annotation subset of the point cloud. This is the width of the slice.

6. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, Step S500 specifically includes: S510, segmenting the data annotation point cloud subset according to the time of the action command, obtaining M segments of gesture action point cloud subset data; S520, filtering the data annotation point cloud subset of each gesture action according to the speed characterization of the action speed, selecting the maximum speed of each frame of the data annotation point cloud subset to form the speed feature vector of the gesture action segment, and filtering out the maximum value of the speed feature vector as the action time annotation candidate point; S530, filtering the action time annotation candidate points according to the signal-to-noise ratio characterization of the action energy, and determining the gesture action time annotation point.

7. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 6, characterized in that, In step S530, the operation of filtering each subset of gesture action point cloud data based on the signal-to-noise ratio (SNR) representing the magnitude of action energy specifically includes: selecting the maximum SNR of each frame's point cloud subset to construct an SNR feature vector. Differential processing is performed on the signal-to-noise ratio eigenvectors. The gesture action time marker point must satisfy the corresponding signal-to-noise ratio difference vector value being greater than zero, and at least one neighboring point's difference vector value being greater than zero; the first velocity maximum point that satisfies the signal-to-noise ratio difference vector condition in chronological order is selected as the gesture action time marker point for that segment of the gesture.

8. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, The gesture action data samples mentioned in step S600 include positive samples. The method for obtaining the positive samples includes: splicing feature information on the point cloud subset data of N distance slice ranges in step S300 in the time dimension to generate a feature map; segmenting the feature map according to the gesture action time annotation points, and incorporating the segmented map into the database as positive samples of the gesture action.

9. The adaptive data annotation method for millimeter-wave radar gesture recognition according to claim 1, characterized in that, The gesture action data samples mentioned in step S600 include negative samples. The method for obtaining negative samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; inferring multiple intervals as no-action intervals between adjacent gesture action time annotation points, and taking the starting point of the no-action interval as the negative sample annotation point; segmenting the feature map according to the negative sample annotation point, and incorporating the segmented map into the database as negative samples of gesture actions.

10. An adaptive data annotation system for millimeter-wave radar gesture recognition, characterized in that, include: The point cloud data acquisition module controls the radar to repeatedly collect the same hand gesture M times within a time period L, thus obtaining... Frame point cloud data; The human target distance estimation module selects the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, calculates the distance between the feature points, and determines the estimated human target distance. The point cloud subset acquisition module uses the estimated distance to the human target. Using the radar as the base point, N range slices are defined along the range dimension towards the radar's location. The point cloud data is then divided into N point cloud subsets according to these N range slices, with overlapping intervals between adjacent range slices. The gesture action distance estimation module calculates the maximum velocity change rate of each point cloud subset and selects the point cloud subset with the most significant change as the data annotation point cloud subset. During the motion capture time, a statistical function is used. The maximum speed of the point cloud subsets across all slice ranges was statistically analyzed, and selected... The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. The gesture action time annotation module uses speed to represent the speed of the action and signal-to-noise ratio to represent the magnitude of the action energy. It annotates the gesture action time in the time dimension and records the gesture action time annotation points. The data sample acquisition module segments N point cloud subsets based on the gesture action time markers to obtain millimeter-wave radar gesture action data samples with multiple distance slice ranges.

Citation Information

Patent Citations

  • Gesture recognition method based on millimeter wave radar point cloud data

    CN118035881A

  • Gesture recognition method and system based on lightweight calculation and storage

    CN120318920A