Adaptive data labeling method and system for millimeter wave radar gesture recognition

By using an adaptive data annotation method, gesture action data is annotated in the time dimension using the point of maximum signal-to-noise ratio and the rate of change of velocity. This solves the problems of poor data consistency and noise interference in existing technologies and achieves high-precision gesture recognition.

CN121542746AActive Publication Date: 2026-02-17POSSUMIC TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610072338.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-02-17
Estimated Expiration
2046-01-20

AI Technical Summary

Technical Problem

In existing millimeter-wave radar gesture recognition technology, the consistency of gesture action data annotation is poor and it is easily affected by noise interference, resulting in inaccurate recognition results.

Method used

By controlling the radar to repeatedly collect hand gestures, selecting the point with the highest signal-to-noise ratio as the feature point, dividing the range into distance slices, and combining the rate of change of velocity and the signal-to-noise ratio to perform data annotation in the time dimension, environmental noise and human interference are eliminated to generate high-quality data samples.

Benefits of technology

It improves the accuracy of gesture recognition, generates an unbiased millimeter-wave radar gesture dataset, and enhances the training robustness of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542746A_ABST
    Figure CN121542746A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive data labeling method and system for millimeter wave radar gesture recognition, and the method comprises the steps: S100, collecting the same gesture action for M times within the time L, and obtaining frame point cloud data; s200, selecting a point with the maximum signal-to-noise ratio in each frame of point cloud data as a feature point, performing statistics on the distance of the feature point, and determining a human body target distance estimation value; s300, taking the human body target distance estimation value as a base point, delimiting N distance slice ranges in the direction where the radar is located, and dividing the point cloud data into N point cloud subsets according to the N distance slices; s400, calculating the maximum speed change rate of each point cloud subset, and selecting the point cloud subset with the most significant change as a data labeling point cloud subset; s500, in combination with the speed features and the signal-to-noise ratios of the data labeling point cloud subsets, gesture action time labeling is carried out; and S600, according to the gesture action time marking points and the N point cloud subsets, obtaining marked gesture action data samples. According to the invention, the accuracy of gesture recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of radar detection technology, specifically relating to an adaptive data annotation method and system for millimeter-wave radar gesture recognition. Background Technology

[0002] Gestures can easily and intuitively express the intent of commands, making their recognition an indispensable research area in human-computer interaction. Gesture recognition applications based on millimeter-wave radar are increasingly popular due to their strong environmental adaptability, privacy protection, and non-contact nature. Currently, gesture recognition for millimeter-wave radar often involves extracting feature vectors from radar echo signals using neural networks, and then using deep learning techniques to classify and recognize these feature vectors. To ensure the accuracy and robustness of the application, training the neural network requires a large amount of accurately labeled radar gesture data as samples. However, current research focuses primarily on innovation in network design, with training samples often extracted entirely from the gesture command or labeled based on a single radar data feature. This results in poor sample consistency and susceptibility to noise interference.

[0003] In the process of annotating radar gesture action point cloud data, the human target is larger, and its echo energy is much stronger than the gesture echo energy. Furthermore, the acquisition environment is complex and subject to significant noise interference. Annotating data based on only a single feature is highly susceptible to training data aliasing, affecting recognition results. For example, some current radar gesture action annotation methods do not perform gesture target distance detection, directly using single feature information for data annotation in the time dimension, easily leading to aliasing between human targets and gesture targets. Other methods directly select the position with the maximum absolute velocity value as the gesture action distance, without considering the potential interference from the surrounding environment. Summary of the Invention

[0004] This invention provides an adaptive data annotation method and system for gesture recognition using millimeter-wave radar, aiming to improve the accuracy of gesture recognition. The technical solution for implementing this invention is as follows:

[0005] In a first aspect, the present invention provides an adaptive data annotation method for millimeter-wave radar gesture recognition, comprising:

[0006] S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain Frame point cloud data;

[0007] S200. Select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, and statistically analyze the distance between the feature points to determine the estimated distance to the human target. ;

[0008] S300, Estimated distance to human target Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0009] S400. Calculate the maximum rate of change of velocity for each point cloud subset, select the point cloud subset with the most significant change as the data label point cloud subset, and use the center of its corresponding distance slice as the estimated distance of the gesture action. ;

[0010] S500. Combining the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset, perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0011] S600. Based on the time markers of the gesture actions, the N point cloud subsets are segmented to obtain millimeter-wave radar gesture action data samples with labeled multi-distance slice ranges.

[0012] As a preferred technical solution, before step S100, it is first determined whether the human target within a preset range in front of the radar is in a stable state. If so, step S100 is executed; otherwise, step S100 is postponed. The method for determining the stable state includes:

[0013] Set human body stability detection signal Record point cloud data, and take the maximum signal-to-noise ratio of each frame of point cloud data to form a feature vector. ,in Representing the The maximum signal-to-noise ratio of frame point cloud data, K is Preset size Updated in real time as time increases. Calculate its effective mean ,in , For neighbors The number of point cloud frames with a maximum non-zero signal-to-noise ratio, and based on... Calculate vectors If the L1 norm is less than a preset threshold, then It determines whether the human target is in a stable state or not, and determines whether the human target is in a non-stable state.

[0014] As a preferred technical solution, step S200, the method for determining the estimated distance to the human target includes: […]. The distance vector of feature points in a frame point cloud data is denoted as ,in This represents the distance value of the feature point in frame t. and the distance vector Statistical analysis was conducted, and the average effective distance was taken as the estimated distance to the human target. In this process, the effective distance for calculation is limited. Only when the signal-to-noise ratio of a feature point, i.e., the maximum signal-to-noise ratio of the corresponding frame's point cloud data, meets the threshold range is the corresponding distance value considered as an effective distance. , for The number of frames in the point cloud data whose maximum signal-to-noise ratio meets the threshold range.

[0015] As a preferred technical solution, in step S300, the estimated distance to the human target is used. Using the radar as a base point, the method for defining N range slices in the direction of the radar is as follows: using the estimated distance to the human target... Using the base point, divide the distance slices along the distance dimension. ,in , The distance is the width of the slice. The distance is the distance between the centers of adjacent distance slices, and there is an overlapping interval between adjacent distance slices; the point cloud data is divided into N point cloud subsets according to the range of the N distance slices.

[0016] As a preferred technical solution, step S400 specifically includes:

[0017] S410, Define the velocity expression function ,in For the first The maximum velocity value of frame point cloud data. Its absolute value, ε is the set speed threshold. Characterized by the first [unit] within the distance slice range Does the maximum speed of the frame point cloud data meet the judgment requirements?

[0018] S420. During the motion acquisition time, the maximum velocity of the point cloud subsets of all slice ranges is statistically analyzed, and a statistical function is defined. , Characterization in the The maximum rate of change of velocity of a subset of point clouds within a distance slice. The number of frames for a subset of the point cloud;

[0019] S430, Select The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. , , The distance to the center of adjacent slices is used to divide the slice range using step S300. A subset of the point cloud is used as a data annotation subset of the point cloud. This is the width of the slice.

[0020] As a preferred technical solution, step S500 specifically includes:

[0021] S510. The data annotation point cloud subset is segmented according to the time of the action command, to obtain M segments of gesture action point cloud subset data.

[0022] S520. Based on the speed of the action, the data annotation point cloud subset of each gesture action is filtered, and the maximum speed of each frame of the data annotation point cloud subset is selected to form the speed feature vector of the gesture action. The maximum value of the speed feature vector is selected as the candidate point for action time annotation.

[0023] S530. Based on the signal-to-noise ratio, which represents the magnitude of the action energy, the candidate points for action time annotation are screened to determine the gesture action time annotation point.

[0024] As a preferred technical solution, in step S530, the operation of filtering each segment of gesture action point cloud data based on the signal-to-noise ratio (SNR) representing the magnitude of action energy specifically includes: selecting the maximum SNR of each frame of point cloud subset to construct an SNR feature vector. Differential processing is performed on the signal-to-noise ratio eigenvectors. The gesture action time marker point must satisfy the corresponding signal-to-noise ratio difference vector value being greater than zero, and at least one neighboring point's difference vector value being greater than zero; the first velocity maximum point that satisfies the signal-to-noise ratio difference vector condition in chronological order is selected as the gesture action time marker point for that segment of the gesture.

[0025] As a preferred technical solution, the gesture action data samples in step S600 include positive samples. The method for obtaining the positive samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; segmenting the feature map according to the gesture action time annotation points, and incorporating the segmented map into the database as positive samples of the gesture action.

[0026] As a preferred technical solution, the gesture action data samples in step S600 include negative samples. The method for obtaining the negative samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; inferring multiple intervals as no-action intervals between adjacent gesture action time annotation points, and taking the starting point of the no-action interval as the negative sample annotation point; segmenting the feature map according to the negative sample annotation point, and incorporating the segmented map into the database as negative samples of gesture actions.

[0027] Secondly, the present invention provides an adaptive data annotation system for millimeter-wave radar gesture recognition, comprising:

[0028] The point cloud data acquisition module controls the radar to repeatedly collect the same hand gesture M times within a time period L, thus obtaining... Frame point cloud data;

[0029] The human target distance estimation module selects the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, calculates the distance between the feature points, and determines the estimated human target distance. ;

[0030] The point cloud subset acquisition module uses the estimated distance to the human target. Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0031] The gesture distance estimation module calculates the maximum rate of change of velocity for each point cloud subset, selects the point cloud subset with the most significant change as the data annotation subset, and uses the center of its corresponding distance slice as the estimated gesture distance value. ;

[0032] The gesture action time annotation module combines the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset to perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0033] The data sample acquisition module segments N point cloud subsets based on the time markers of the gesture actions to obtain millimeter-wave radar gesture action data samples with multiple distance slice ranges.

[0034] The adaptive data annotation method and system for millimeter-wave radar gesture recognition provided by this invention have the following advantages: by using the three-dimensional radar data features of distance, speed, and signal-to-noise ratio, environmental noise is filtered out and human target interference is shielded, and adaptive data annotation of gesture action data is achieved to obtain unbiased, high-quality millimeter-wave radar gesture action dataset, which can greatly improve the accuracy of gesture recognition. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of an adaptive data annotation method for millimeter-wave radar gesture recognition provided in an embodiment of the present invention.

[0037] Figure 2 This is a flowchart illustrating the acquisition of gesture distance estimation in the adaptive data annotation method for millimeter-wave radar gesture recognition provided in this embodiment of the invention.

[0038] Figure 3 This is a flowchart of the gesture action time annotation method for millimeter-wave radar gesture recognition provided in the embodiments of the present invention.

[0039] Figure 4 This is a block diagram of the adaptive data annotation system for millimeter-wave radar gesture recognition provided in an embodiment of the present invention. Detailed Implementation

[0040] To make the technical solution of the present invention clearer and its technical advantages more apparent, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of the present invention.

[0041] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.

[0042] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0043] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0044] It is understandable that radar receives echoes and processes them to generate point cloud data. Since human targets are large and their echo energy is strong, they are very likely to interfere with the echo signals of hand gestures. Therefore, distance estimation is essentially about refining the distance dimension range. This measure can extract the distance range where the changes in hand gestures are most significant to the greatest extent. Subsequent action time labeling within this range can effectively avoid inaccurate labeling caused by background noise and human interference.

[0045] The specific embodiments of the present invention provide an adaptive data annotation method and system for millimeter-wave radar gesture recognition. In order to improve the accuracy of gesture recognition, the main approaches are as follows: First, determining the range of motion of the gesture (distance detection) and distinguishing between human targets and gesture targets is the first step in whether the data annotation can be successful; Second, based on determining the distance of the gesture, a detection algorithm combining signal-to-noise ratio information and velocity information is used to annotate the gesture data stream in the time dimension.

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0047] See Figure 1 As shown, as a basic implementation method, the adaptive data annotation method for millimeter-wave radar gesture recognition provided in this embodiment includes:

[0048] S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain Frame point cloud data;

[0049] S200. Select the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, and statistically analyze the distance between the feature points to determine the estimated distance to the human target. ;

[0050] S300, Estimated distance to human target Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0051] S400. Calculate the maximum rate of change of velocity for each point cloud subset, select the point cloud subset with the most significant change as the data label point cloud subset, and use the center of its corresponding distance slice as the estimated distance of the gesture action. ;

[0052] S500. Combining the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset, perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0053] S600. Based on the time markers of the gesture actions, the N point cloud subsets are segmented to obtain millimeter-wave radar gesture action data samples with labeled multi-distance slice ranges.

[0054] When the millimeter-wave radar collects hand gestures from a human target, it repeats each hand gesture M times according to the command signal, and records the total collection time of the gesture as time L. The total time L is then roughly divided into M segments using the time points when the command signal occurs. The time period of the gesture action is recorded as follows: , Positive integers; gesture annotation involves selecting the time point with the most significant action feature within each time period as the action annotation point.

[0055] Taking the FMCW millimeter-wave radar as an example, its transmission carrier frequency is 60GHz. It collects hand gestures made by targets within a range of 0.5m to 2m in front of the radar. The radar frame period is set according to the complexity of the action. Generally, a frame period of 50ms can accurately reflect commonly defined hand gestures, such as waving to the left (right) or waving forward (backward). The point cloud feature map size is set to size=[number_chirp, length_capture], where number_chirp is the number of radar chirps and length_capture is the number of frames captured for the detected action. Generally, size=[64,32] can more completely reflect the action characteristics.

[0056] Furthermore, to achieve optimal data acquisition results when using millimeter-wave radar to collect gesture data, a relatively static environment is generally required, meaning the target object must be in a stable state. Therefore, the target (human body) needs to remain within a preset range in front of the radar for a period of time before data acquisition. Thus, as a preferred implementation, the adaptive data annotation method for millimeter-wave radar gesture recognition first determines whether the human target within the preset range in front of the radar is in a stable state before step S100. If so, step S100 is executed; otherwise, step S100 is postponed.

[0057] Specifically, the method for determining the stable state includes: setting a human body stability detection signal. Record point cloud data, and take the maximum signal-to-noise ratio of each frame of point cloud data to form a feature vector. ,in Representing the Maximum signal-to-noise ratio of frame point cloud data, Preset size , Updated in real time as time increases. Calculate its effective mean ,in , For neighbors The number of point cloud frames with a maximum non-zero signal-to-noise ratio, and based on... Calculate vectors If the L1 norm is less than a preset threshold, then It was determined that the human target was in a stable state and that subsequent data collection could proceed.

[0058] After receiving the target stabilization signal feedback, the acquisition system begins to issue motion acquisition commands, using a timer to record the motion time range corresponding to the gesture commands, such as the first... The time period of the gesture action is recorded as follows: , A positive integer, after collecting a preset number of times M, the collection stops and the timing stops, and the distance to the human target is estimated from the data within the operation time L.

[0059] In step S200, the method for determining the estimated distance to the human target includes: […]. The distance vector of feature points in a frame point cloud data is denoted as ,in This represents the distance value of the feature point in frame t. and the distance vector Statistical analysis was conducted, and the average effective distance was taken as the estimated distance to the human target. Furthermore, the effective distances involved in the calculation can be limited. Only when the signal-to-noise ratio of a feature point, i.e., the maximum signal-to-noise ratio of the corresponding frame's point cloud data, meets a threshold range, will its corresponding distance value be considered as an effective distance. , for The number of frames in the point cloud data whose maximum signal-to-noise ratio meets the threshold range.

[0060] In step S300, the estimated distance to the human target is used. Using the radar as a base point, the method for defining N range slices in the direction of the radar is as follows: using the estimated distance to the human target... Using the base point, divide the distance slices along the distance dimension. ,in , The distance is the width of the slice. It's generally best to set it to half an arm's length, for example, set... , The distance between the centers of adjacent distance slices is typically set to the radar range resolution; furthermore, there are overlapping regions between adjacent distance slices.

[0061] Combination Figure 2 As shown, step S400 specifically includes:

[0062] S410, Define the velocity expression function ,in For the first Maximum frame rate Its absolute value, ε is the set speed threshold. Characterized within the range of the slice Does the maximum speed of the frame point cloud data meet the judgment requirements?

[0063] S420. Within the motion acquisition time L, the maximum velocity of each point cloud subset is statistically analyzed, and a statistical function is defined. , Characterization in the The maximum rate of change of velocity of a subset of point clouds within a distance slice. The number of frames for a subset of the point cloud;

[0064] S430, Select The distance from the center of the slice with the maximum value is used as the estimated distance for the gesture. , , The distance to the center of adjacent slices is used to divide the slice range using step S300. A subset of the point cloud is used as a data annotation subset of the point cloud.

[0065] Next, the data of the labeled point cloud subset is labeled along the time dimension. Because many environmental noises and human body shaking effects have been removed, the point cloud data (including signal-to-noise ratio, velocity, etc.) shows obvious fluctuations in the time dimension with the gesture movement, which can adaptively and accurately realize the action segmentation.

[0066] Combination Figure 3 As shown, step S500 specifically includes:

[0067] S510. Segment the data annotation point cloud subset according to the time of the action command, and obtain M segments of gesture action point cloud subset data.

[0068] S520. Based on the speed of the action, the data annotation point cloud subset of each gesture action is filtered, and the maximum value of the speed feature vector is selected as the candidate point for action time annotation. Specifically, the maximum speed of each frame of the data annotation point cloud subset selected in step S430 is selected to form the speed feature vector of the gesture action, and the maximum value of the speed feature vector is selected as the candidate point for action time annotation.

[0069] S530. Based on the signal-to-noise ratio, which represents the magnitude of the action energy, the candidate points for action time annotation are screened to determine the gesture action time annotation point.

[0070] In step S520, the above-mentioned data annotation point cloud subset is taken, and for each segment of data from the M collection actions, the action start point is annotated: within the effective operation time of each gesture action recording, the maximum velocity of each frame of point cloud data is selected to form a velocity feature vector. ,in Representing the The maximum velocity value of the frame point cloud subset, selected The maximum points form the candidate point set.

[0071] In step S530, the operation of filtering the point cloud subset data of each gesture action based on the signal-to-noise ratio (SNR) representing the magnitude of the action energy specifically includes: the SNR of the point cloud subset data shows an increasing trend at the start of the action, therefore, the maximum SNR of each frame's point cloud subset is selected to form the SNR feature vector. Differential processing is performed on the signal-to-noise ratio eigenvectors. For each action time marker, the corresponding signal-to-noise ratio (SNR) difference vector value must be greater than zero, and at least one neighboring point must also have an SNR difference vector value greater than zero. The first velocity maximum point satisfying the SNR difference vector condition in chronological order is selected as the action time marker for that gesture segment. This operation is repeated M times to obtain all action time markers.

[0072] In step S600, the gesture action data samples include positive samples. The method for obtaining positive samples includes: splicing feature information from the point cloud subset data of N distance slice ranges in step S300 along the time dimension to generate a feature map; segmenting the feature map according to the action time markers, and incorporating the segmented map into the database as positive samples of the gesture action. The point cloud data features include distance, velocity, orientation angle, pitch angle, and signal-to-noise ratio, which are spliced ​​along the time dimension to form multiple feature maps, such as generating a range map (RTM), a Doppler map (DTM), and an angle map (ATM). Furthermore, the map width is the radar chirp count, and the length is the number of acquisition frames. Since the action markers are located at the start of the action, and gesture actions usually have a clearer intention within the initial time range, covering the first half of the action with acquisition frames is sufficient to meet the recognition requirements. After determining the accurate action time markers, processing point cloud data feature maps at different distance ranges can cover appropriate noise and human interference, obtaining feature information over a wider area to expand the database.

[0073] In step S600, with the determination of the action time markers, negative samples in the database can also be accurately collected. That is, the gesture action data samples include negative samples. The method for obtaining negative samples includes: splicing feature information in the time dimension of the point cloud subset data of N distance slice ranges in step S300 to generate a feature map; between adjacent action time markers, using prior conditions to infer that multiple intervals within a reasonable time range are no-action intervals, and taking the starting point of the no-action interval as the negative sample marker; based on the negative sample markers, segmenting the feature map, and incorporating the segmented map into the database as negative samples of gesture actions.

[0074] In this way, multiple training sets containing noise and interference can be generated using the point cloud subsets under different distance slices, thereby augmenting the training data and enhancing the robustness of the trained neural network.

[0075] The present invention also provides an adaptive data annotation system for millimeter-wave radar gesture recognition, which executes the adaptive data annotation method for millimeter-wave radar gesture recognition described above.

[0076] Specifically, see Figure 4 As shown, the adaptive data annotation system for millimeter-wave radar gesture recognition includes:

[0077] The point cloud data acquisition module controls the radar to repeatedly collect the same hand gesture M times within a time period L, thus obtaining... Frame point cloud data;

[0078] The human target distance estimation module selects the point with the highest signal-to-noise ratio in each frame of point cloud data as the feature point, calculates the distance between the feature points, and determines the estimated human target distance. ;

[0079] The point cloud subset acquisition module uses the estimated distance to the human target. Using the radar as the base point, N range slices are defined in the direction of the radar. The point cloud data is divided into N point cloud subsets according to the N range slices, where adjacent range slices have overlapping intervals.

[0080] The gesture distance estimation module calculates the maximum rate of change of velocity for each point cloud subset, selects the point cloud subset with the most significant change as the data label point cloud subset, and uses the center of the corresponding distance slice as the gesture distance estimate.

[0081] The gesture action time annotation module combines the velocity characteristics and signal-to-noise ratio of the data annotation point cloud subset to perform gesture action time annotation in the time dimension and record the gesture action time annotation points.

[0082] The data sample acquisition module segments N point cloud subsets based on the time markers of the gesture actions to obtain millimeter-wave radar gesture action data samples with multiple distance slice ranges.

[0083] The beneficial effects of the technical solution of the present invention are as follows: the present invention uses the three-dimensional radar data features of distance, speed and signal-to-noise ratio to filter environmental noise and shield human target interference, and achieves adaptive data annotation of gesture action data to obtain unbiased high-quality millimeter-wave radar gesture action dataset, which can greatly improve the accuracy of gesture recognition.

[0084] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. An adaptive data labeling method for millimeter wave radar gesture recognition, characterized in that, The method comprises the following steps: S100, control the radar to repeatedly collect the same hand gesture M times within time L, and obtain Frame point cloud data; S200, select the point with the maximum signal-to-noise ratio in each frame of point cloud data as a feature point, and determine a human target distance estimation value by counting the distance of the feature point ; S300, estimate the distance of the human target As the base point, N distance slice ranges are drawn in the direction where the radar is located, and the point cloud data is divided into N point cloud subsets according to the N distance slice ranges, wherein adjacent distance slice ranges have an overlapping interval; S400, calculate the maximum speed change rate of each point cloud subset, select the point cloud subset with the most significant change as the data labeling point cloud subset, and the distance slice center corresponding to the data labeling point cloud subset is taken as the gesture action distance estimation value ; S500, in combination with the speed feature and signal-to-noise ratio of the data-labeled point cloud subset, gesture action time labeling is performed in the time dimension, and a gesture action time labeling point is recorded; S600, according to the gesture action time labeling point, N point cloud subsets are segmented to obtain a millimeter wave radar gesture action data sample with a labeled multi-distance slice range.

2. The adaptive data labeling method for mm-wave radar gesture recognition of claim 1, wherein, Before step S100, it is judged whether a human target in a preset range in front of the radar is in a stable state, if yes, step S100 is executed, otherwise, step S100 is temporarily suspended; the method for judging the stable state comprises: Setting human body stability detection signal , record point cloud data, take the maximum signal-to-noise ratio of each frame of point cloud data to constitute a feature vector , wherein represents the first frame of point cloud data, K is a preset size, increases with time, ; real-time update , calculate its effective mean , wherein , is the number of point cloud frames with non-zero maximum signal-to-noise ratio of adjacent frame point cloud data, and the L1 norm of the vector is calculated according to , if it is less than a preset threshold, then , and it is determined that the human body target is in a stable state, otherwise it is determined that the human body target is in a non-stable state.

3. The adaptive data labeling method for mm-wave radar gesture recognition of claim 1, wherein, In step S200, the method for determining the human target distance estimation value comprises: The distance vector of the feature point of the frame point cloud data is denoted as , wherein represents the distance value of the feature point of the t-th frame, and the effective distance mean is taken as the human target distance estimation value , wherein the effective distance participating in the calculation is limited, and only the distance value corresponding to the feature point with the signal-to-noise ratio satisfying the threshold range is taken as the effective distance , , is the frame number of the frame point cloud data with the maximum signal-to-noise ratio satisfying the threshold range.

4. The adaptive data labeling method for mm-wave radar gesture recognition of claim 1, wherein, In step S300, the human body target distance estimation value is taken as the base point, and N distance slice ranges are defined in the direction in which the radar is located is taken as the base point, and distance slices are divided in the direction in which the radar is located along the distance dimension , wherein , is the width of the distance slice is the distance between the centers of adjacent distance slices, and there is an overlapping interval between the adjacent distance slices The point cloud data is divided into N point cloud subsets according to the N distance slice ranges.

5. The adaptive data labeling method for mm-wave radar gesture recognition of claim 1, wherein, Step S400 specifically comprises: S410, define a speed expression function wherein is the first maximum speed value of the frame point cloud data, is the absolute value thereof, and ε is a set speed threshold value, characterizes whether the maximum speed of the first frame point cloud data within the distance slice range meets a determination requirement; S420, in the action collection time, the maximum speed of the point cloud subset in all slice ranges is counted, and a statistical function is defined , characterizing the maximum speed change rate of the point cloud subset in the first distance slice range, is the frame number of the point cloud subset; S430, selecting The center of the distance slice with the maximum value is taken as the gesture action distance estimation value , , is the distance between adjacent distance slice centers, and the distance slice range is selected in step S300 The point cloud subset is taken as the data annotation point cloud subset is the width of the distance slice.

6. The adaptive data labeling method for mm-wave radar gesture recognition of claim 1, wherein, Step S500 specifically comprises: S510, the data-labeled point cloud subset is segmented according to the time of the action instruction, M gesture action point cloud subset data are obtained; S520, according to the speed feature vector, the maximum speed of each frame of the data-labeled point cloud subset is selected as the speed feature vector of the gesture action, and the maximum value of the speed feature vector is selected as the action time labeling candidate point; S530, according to the signal-to-noise ratio, the action time labeling candidate point is selected, and the gesture action time labeling point is determined.

7. The adaptive data labeling method for mm-wave radar gesture recognition of claim 6, wherein, In step S530, according to the size of the signal-to-noise ratio representing the action energy, the operation of screening each segment of gesture action point cloud subset data, specifically includes: selecting the maximum signal-to-noise ratio of each frame of point cloud subset to form a signal-to-noise ratio feature vector , the signal-to-noise ratio feature vector is subjected to difference processing , the gesture action time annotation point must satisfy that the corresponding signal-to-noise ratio difference vector value is greater than zero, and the difference vector value of at least one adjacent point is greater than zero; selecting the first speed maximum value point satisfying the signal-to-noise ratio difference vector condition in time sequence as the gesture action time annotation point of the segment gesture.

8. The adaptive data labeling method for millimeter wave radar gesture recognition of claim 1, wherein, The gesture action data sample in step S600 comprises a positive sample, and the method for obtaining the positive sample comprises: feature information of the point cloud subset data in the N distance slice ranges in step S300 is spliced in the time dimension to generate a feature graph; according to the gesture action time labeling point, the feature graph is cut, and the cut feature graph is put into a database as a gesture action positive sample.

9. The adaptive data labeling method for millimeter wave radar gesture recognition of claim 1, wherein, The gesture action data sample in step S600 comprises a negative sample, and the method for obtaining the negative sample comprises: feature information of the point cloud subset data in the N distance slice ranges in step S300 is spliced in the time dimension to generate a feature graph; between adjacent gesture action time labeling points, a plurality of intervals are inferred as non-action intervals, and the start point of the non-action interval is taken as a negative sample labeling point; according to the negative sample labeling point, the feature graph is cut, and the cut feature graph is put into a database as a gesture action negative sample.

10. An adaptive data labeling system for millimeter wave radar gesture recognition, comprising: The method comprises the following steps: The point cloud data acquisition module controls the radar to repeatedly collect the same gesture action M times within a time L to obtain frame point cloud data; The human body target distance estimation module selects a point with the maximum signal-to-noise ratio in each frame of point cloud data as a feature point, and determines a human body target distance estimation value by counting the distance of the feature point ; The point cloud subset obtaining module obtains a distance estimation value of the human body target The point cloud data is divided into N point cloud subsets according to the N distance slice ranges, and there is an overlapping interval between adjacent distance slice ranges. The gesture action distance estimation module calculates the maximum speed change rate of each point cloud subset, selects the point cloud subset with the most significant change as the data labeling point cloud subset, and takes the distance slice center corresponding to the data labeling point cloud subset as the gesture action distance estimation value ; A gesture action time labeling module, in combination with the speed feature and signal-to-noise ratio of the data-labeled point cloud subset, gesture action time labeling is performed in the time dimension, and a gesture action time labeling point is recorded; A data sample acquisition module, according to the gesture action time labeling point, N point cloud subsets are segmented to obtain a millimeter wave radar gesture action data sample with a multi-distance slice range.

Citation Information

Patent Citations

  • Millimeter wave radar sensing gesture recognition method based on few sample learning

    CN114708663A

  • Gesture recognition method based on millimeter wave radar point cloud data

    CN118035881A

  • Gesture recognition method and system based on lightweight calculation and storage

    CN120318920A

  • Video and text driven millimeter wave radar gesture data generation method

    CN121167307A