A multi-person interference gesture recognition method for a millimeter wave radar

CN122821631APending Publication Date: 2026-09-25ANHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611024518.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]现阶段毫米波雷达手势识别技术已广泛应用于智能家居、车载智能控制等领域,但现有常规识别方案在复杂多人实景环境中仍存在多项难以规避的技术缺陷

Benefits of technology

本发明在数据源预处理环节通过多级FFT变换与杂波抑制完成原始回波数据的标准化处理,拆解距离、速度、角度核心物理维度信息,既完整保留手部手势微动对应的精细化回波特征,又滤除环境静态杂波带来的无效噪声,从源头提升原始特征的数据信噪比,减少底层信号瑕疵对后续识别流程的干扰,夯实全流程精准识别的数据基础。依托空间距离约束、节点与边特征联动的动态权重计算逻辑,能够自适应区分目标手部小幅局部运动与周边人员走动、摆臂带来的大范围干扰运动,实现手势信号与干扰信号的数据解耦,有效破解多人同场环境下多目标回波混叠造成的识别错乱难题,大幅提升复杂居家、车载等多干扰实景下的识别鲁棒性,摆脱传统雷达手势识别仅能在单人纯净环境稳定工作的局限性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821631A_ABST
    Figure CN122821631A_ABST
Patent Text Reader

Abstract

The application discloses a multi-person interference gesture recognition method of a millimeter wave radar and relates to the technical field of gesture recognition, and the technical solution points of the application comprise the following steps: collecting original echo data of the millimeter wave radar, and obtaining radar feature representation containing distance, speed and angle dimensions through preprocessing of the original echo data; constructing a reliability-aware dynamic radar apparent motion field based on the radar feature representation, decoupling target gesture motion and non-target personnel interference through a dynamic weight evaluation mechanism, and extracting a gesture motion sequence; inputting the gesture motion sequence into a lightweight space-time feature extraction network, determining context dependency of gesture actions, and obtaining high-dimensional space-time features; and inputting the high-dimensional space-time features into a classifier to output gesture category recognition results, so that the determination accuracy of gesture classification results is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gesture recognition technology, and more specifically, to a method for recognizing multi-person interference gestures using millimeter-wave radar. Background Technology

[0002] Currently, millimeter-wave radar gesture recognition technology has been widely used in smart homes, vehicle intelligent control and other fields. However, existing conventional recognition solutions still have several technical defects that are difficult to avoid in complex multi-person real-world environments.

[0003] Traditional radar data processing mostly performs basic range and Doppler transformations on the raw echoes, lacking systematic three-dimensional feature normalization. It relies solely on single-dimensional data for feature analysis, failing to comprehensively integrate the three types of physical information: distance, velocity, and azimuth. Static wall clutter and environmental electromagnetic noise easily remain in the feature data, resulting in a low signal-to-noise ratio at the lower level, thus limiting the improvement of recognition accuracy from the data source level. Furthermore, it lacks targeted optimization for interference conditions such as people walking by, close-range arm swings, and simultaneous activities by multiple people. It cannot effectively distinguish between the small, subtle hand movements of the target user and the large-scale movements of unrelated personnel. In real-world scenarios with multiple users, the interference echoes and gesture echoes overlap, easily leading to target features being covered by interference information and misidentification due to target drift. This results in a significant drop in recognition accuracy in multi-person environments, making it unsuitable for deployment scenarios with frequent personnel activity, such as homes and cockpits.

[0004] Existing feature extraction networks generally adopt standard Transformer or deep CNN structures with large parameters. The number of model parameters and computational power consumption are relatively high. Even if the recognition accuracy is excellent, it is impossible to achieve local real-time deployment on embedded radar terminals with limited computing power and storage space. Most solutions can only rely on cloud computing power to complete the recognition calculation, thus posing risks of data transmission delay and privacy leakage. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method for recognizing multi-person interference gestures using millimeter-wave radar.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for recognizing multiple interference gestures using millimeter-wave radar, comprising the following steps: The raw echo data of millimeter-wave radar is collected, and the raw echo data is preprocessed to obtain radar feature representations containing range, velocity, and angle dimensions; Based on radar feature representation, a reliability-aware dynamic radar apparent motion field is constructed. The target gesture motion and non-target personnel interference are decoupled through a dynamic weight evaluation mechanism to extract the gesture motion sequence. The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features. Input high-dimensional spatiotemporal features into the classifier and output gesture category recognition results.

[0007] Preferably, the raw echo data is preprocessed to obtain radar feature representations including range, velocity, and angle dimensions, specifically as follows: The raw echo data is processed sequentially by range FFT and Doppler FFT, and static clutter is suppressed by background removal or mean elimination. Then, it is processed by antenna dimension angle FFT or beamforming to obtain radar feature representation containing range, velocity and angle dimensions.

[0008] Preferably, it further includes: Based on the energy distribution, range variation, velocity variation and motion continuity represented by radar features, a rough motion area of ​​the target user is extracted to generate a target area guidance map; The target area guidance map is used as an additional channel and spliced ​​with the radar feature representation to form an enhanced input feature.

[0009] Preferably, the dynamic weight evaluation mechanism is as follows: The enhanced input features are fed into a lightweight feature extraction network, which automatically learns the feature weights of the target region and non-target regions, enhances the response of the target gesture region, and reduces the feature contribution of non-target motion regions. Temporal modeling is performed on the radar feature representation of consecutive multiple frames to extract short-term local motion features of gestures. Interference areas are identified based on the spatial range, duration, and consistency with the target gesture of non-target motion, and a decoupled gesture motion sequence is generated by dynamically suppressing weight output.

[0010] Preferably, the gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features, specifically: The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network. The lightweight spatiotemporal feature extraction network adopts a structure of lightweight Transformer combined with bidirectional gated recurrent units. At the same time, it adopts lightweight convolution, compact temporal modeling and efficient feature fusion strategies to compress the number of model parameters and computational load, and focuses computational resources on the target gesture-related region to obtain high-dimensional spatiotemporal features.

[0011] Preferably, the high-dimensional spatiotemporal features are input into the classifier to output the gesture category recognition result, specifically as follows: High-dimensional spatiotemporal features are input into a classifier consisting of a fully connected layer and a Softmax classifier, and the probability distribution of various gestures is output to obtain the final gesture category recognition result.

[0012] Preferably, it also includes anti-interference optimization during the training phase, collecting multi-scene samples such as single person without interference, surrounding people stationary, surrounding people walking, surrounding people swinging their arms, and multiple people with dynamic interference, adding background noise samples for training, setting loss penalty terms and Top-2 softmax difference detection thresholds, and filtering low-confidence predictions.

[0013] Preferably, the feature weights in the dynamic weight evaluation mechanism are calculated using the following formula: ; in, Let N be the output feature of node i, N be the neighborhood set of node j, and D be the output feature of node i. ij Let X be the distance between nodes i and j, r be the neighborhood radius, MLP be a shared multilayer perceptron, and X be the distance between nodes i and j. j E is the input feature of node j. ij Let be the edge characteristics from node j to node i.

[0014] Preferably, the point cloud data output by the millimeter-wave radar includes a sequence number, timestamp, point number, three-dimensional coordinates, distance, velocity, Doppler cells, azimuth angle, and reflection intensity fields.

[0015] Preferably, the radar feature representation includes input forms of point clouds and time-series images.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention standardizes the raw echo data through multi-level FFT transformation and clutter suppression in the data source preprocessing stage, deconstructing core physical dimensions such as distance, velocity, and angle. This process not only fully preserves the refined echo features corresponding to subtle hand gestures but also filters out invalid noise from static environmental clutter. This improves the signal-to-noise ratio of the original features from the source, reduces interference from underlying signal defects in subsequent recognition processes, and solidifies the data foundation for accurate recognition throughout the entire process. Based on the dynamic weight calculation logic of spatial distance constraints and node-edge feature linkage, it can adaptively distinguish between small local movements of the target hand and large-scale interference movements caused by surrounding people walking and swinging their arms. This achieves data decoupling between gesture signals and interference signals, effectively solving the problem of recognition errors caused by multi-target echo aliasing in multi-person environments. It significantly improves the robustness of recognition in complex real-world scenarios with multiple interferences, such as at home and in vehicles, overcoming the limitation of traditional radar gesture recognition which can only work stably in a single, clean environment.

[0017] The lightweight spatiotemporal feature extraction network integrates a lightweight Transformer and a bidirectional gated recurrent unit, along with optimization strategies such as lightweight convolution and compact temporal modeling. While fully mining the temporal contextual dependencies of gestures and extracting high-dimensional spatiotemporal features, it significantly reduces the number of model parameters and computational overhead, balancing feature extraction completeness and operational efficiency. It can be adapted to embedded terminal devices with limited computing resources, enabling low-power real-time inference deployment on the local end. It meets the needs of miniaturization and low-cost deployment of smart home hardware, ensuring the accuracy of gesture classification results while avoiding misjudgments under extreme interference through low-confidence prediction filtering. It balances anti-interference ability, recognition accuracy, and engineering practicality, and is suitable for diverse application scenarios such as vehicle control and smart home appliance gesture control. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of an embodiment of the present invention. Detailed Implementation

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0021] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0022] Reference Figure 1 As shown.

[0023] The embodiments further illustrate the multi-person interference gesture recognition method for millimeter-wave radar proposed in this invention.

[0024] A method for recognizing multiple interference gestures using millimeter-wave radar, comprising the following steps: The raw echo data of millimeter-wave radar is collected, and the raw echo data is preprocessed to obtain radar feature representations containing range, velocity, and angle dimensions; Based on radar feature representation, a reliability-aware dynamic radar apparent motion field is constructed. The target gesture motion and non-target personnel interference are decoupled through a dynamic weight evaluation mechanism to extract the gesture motion sequence. The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features. Input high-dimensional spatiotemporal features into the classifier and output gesture category recognition results.

[0025] The raw echo data is preprocessed to obtain radar feature representations that include range, velocity, and angle dimensions, specifically: The raw echo data is processed sequentially by range FFT and Doppler FFT, and static clutter is suppressed by background removal or mean elimination. Then, it is processed by antenna dimension angle FFT or beamforming to obtain radar feature representation containing range, velocity and angle dimensions.

[0026] Range-based FFT relies on the frequency domain ranging principle of Fourier transform to convert the time-domain echo sampling signal received by the radar into the range frequency domain space. The difference frequency between the frequency-modulated continuous wave signal emitted by the radar and the target reflected echo has a fixed linear correspondence with the radial distance of the target relative to the radar. By extracting the amplitude information corresponding to each frequency point, range-based FFT calculates the radial distance values ​​between all reflecting targets in space and the radar antenna, completing the decomposition and representation of the physical information in the range dimension, thereby generating range-dimensional basic data.

[0027] Doppler FFT is performed to calculate the radial velocity of the target. The same spatial target will generate a continuous echo sampling sequence within multiple consecutive frequency sweep cycles of the radar. Doppler FFT performs Fourier transform on the sampling data of different frequency sweep cycles under the same range gate. Utilizing the Doppler frequency shift characteristics of the echo brought by the moving target, the magnitude of the frequency offset of the frequency domain result directly corresponds to the radial movement rate of the target relative to the radar. The positive or negative frequency offset can distinguish whether the target is moving towards or away from the radar. This operation is used to achieve complete extraction of velocity dimension features. After two-step transformation of range FFT and Doppler FFT, the data has been transformed into a range-Doppler two-dimensional data matrix.

[0028] After the two-dimensional matrix is ​​formed, a static clutter suppression process is introduced. In nature, stationary objects such as walls, furniture, and the ground continuously reflect radar echoes, forming fixed static clutter. This type of clutter signal has stable amplitude and minimal temporal variation, which can easily mask the effective echo signals of micro-moving targets such as hands. Background removal algorithms or time-domain mean elimination algorithms are selected to filter out clutter. Mean elimination is achieved by calculating the mean of multiple frames of time-series data at the same Doppler position and subtracting the mean component from the original data as a static clutter component. Background removal is achieved by pre-collecting blank radar background data in a target-free scenario and performing a difference operation between the real-time collected data and the background reference data. Both methods can remove fixed static interference signals, retain the effective echo data of moving targets such as human hands to the greatest extent, and improve the signal-to-noise ratio of the features.

[0029] After completing the clutter suppression and purification data, the angle dimension is calculated based on the multi-channel sampling data of the radar array antenna. The processing methods are divided into two categories: antenna dimension angle FFT operation and digital beamforming. Angle FFT utilizes the spatial phase difference generated when different array elements of the array antenna receive the echo of the same target, and uses Fourier transform to convert the spatial phase difference into the elevation angle or azimuth angle information of the target relative to the radar normal. Digital beamforming applies weighting coefficients to the multi-channel echo data of the array, superimposes signal gain in the specified spatial angle direction, suppresses sidelobe interference, and directionally obtains the target echo amplitude information at the corresponding angle. Both processing modes can complete the solution of the physical parameters of the angle dimension.

[0030] Through the aforementioned step-by-step signal processing across the entire chain, the original radar echo data, which originally only contained temporal amplitude information, was decomposed, converted, and purified into standardized radar feature data that integrates three physical dimensions: range, speed, and spatial orientation. This three-dimensional feature fully encompasses three key types of information: the target's spatial location, speed of movement, and spatial orientation. It can comprehensively depict the spatial motion details of the target, such as hand gestures, and successfully completes the transformation of the original radar signal into a structured three-dimensional feature representation, providing complete underlying feature support for target classification, interference removal, and lightweight network feature extraction.

[0031] Also includes: Based on the energy distribution, range variation, velocity variation and motion continuity represented by radar features, a rough motion area of ​​the target user is extracted to generate a target area guidance map; The target area guidance map is used as an additional channel and spliced ​​with the radar feature representation to form an enhanced input feature.

[0032] In the target area guidance map generation stage, the energy distribution, distance variation, velocity variation, and motion continuity features in radar characteristics are used to roughly delineate the target user's movement area. The energy distribution feature originates from the cumulative amplitude of radar echo signals at various spatial locations. The effective reflected echoes from the target user's hands and body exhibit stable and concentrated energy accumulation characteristics compared to environmental clutter and interference echoes from irrelevant passersby. Based on the energy distribution, candidate areas with concentrated energy within the radar detection range are delineated, eliminating large areas of low-energy invalid background regions. The distance variation feature characterizes the radial distance fluctuation pattern corresponding to the same spatial point within continuous sampling frames. When the target user makes a gesture, the hand exhibits small-range, continuous distance fluctuations in the near-range, while stationary obstacles show no distance change, and the distance variation range and amplitude of distant moving interfering persons are different. There is a clear distinction between near-field and long-range hand gesture targets. The effective target range can be narrowed down from the candidate region by measuring the temporal fluctuation amplitude of the distance. The velocity change feature is determined by the radial velocity temporal data obtained by Doppler calculation. The hand micro-movements corresponding to the hand gestures show a low-speed, small-amplitude frequency change pattern. Irrelevant walking personnel have a wide range of high-speed speed change characteristics. The velocity of static objects always remains at zero. By using the difference in the amplitude and frequency of velocity changes, interference pixels that do not match the motion attributes can be eliminated. Motion continuity constrains the target area at the multi-frame temporal correlation level. The motion area of ​​the real user's gesture has the temporal characteristics of spatial position continuity and motion trajectory continuity in consecutive radar sampling frames. Instantaneous sporadic clutter noise and short-term passing scattered interference signals cannot meet the constraint condition of continuous motion across frames. This feature is used to eliminate instantaneous invalid isolated areas.

[0033] After four features work together and are filtered through multiple levels, a rough motion area corresponding to the target user is identified. Based on the regional location information, a target area guidance map with a format and size matching the original radar features is generated. The position of the target motion area in the guidance map is assigned a high-weight label, and non-target interference areas are assigned a low-weight label, intuitively distinguishing between effective target space and ineffective interference space. After the guidance map is constructed, the feature stitching and enhancement stage begins. The generated target area guidance map is treated as a new feature channel and fused with the original radar feature representation containing distance, speed, and angle information in the channel dimension through tensor stitching. The original radar features fully retain the basic physical information of target distance, motion rate, and spatial orientation. The guidance map channel is supplemented with the effectiveness weight information of the spatial region. Finally, the combination forms an enhanced input feature that integrates basic physical features and prior information of regional weights. This enhanced feature can automatically focus on the effective target area and weaken the feature interference caused by irrelevant interference areas during network operation, relying on the weight information of the guidance channel. It not only preserves the original multi-dimensional physical details of the radar but also realizes the prior guidance constraint of the target area, providing high-quality input data with spatial focusing attributes for the anti-interference recognition model.

[0034] The dynamic weight evaluation mechanism is as follows: The enhanced input features are fed into a lightweight feature extraction network, which automatically learns the feature weights of the target region and non-target regions, enhances the response of the target gesture region, and reduces the feature contribution of non-target motion regions. Temporal modeling is performed on the radar feature representation of consecutive multiple frames to extract short-term local motion features of gestures. Interference areas are identified based on the spatial range, duration, and consistency with the target gesture of non-target motion, and a decoupled gesture motion sequence is generated by dynamically suppressing weight output.

[0035] The lightweight feature extraction network integrates the original radar 3D information with the enhanced input features of the target guidance channel. Leveraging its own convolutional and parameter structures, and adapting to the low-computing conditions of embedded devices, the lightweight network autonomously completes the adaptive allocation of feature weights across the entire region. Combined with prior spatial information carried by the target region guidance map, it autonomously distinguishes between the target gesture area and irrelevant non-target motion areas within the image. For the target gesture area, the feature weight coefficients are gradually increased. This weight increase amplifies the amplitude of distance, speed, and angle-related feature responses corresponding to the gesture, enhancing the representation of subtle hand movements in the feature data. For non-target motion areas such as people walking or minor environmental vibrations, the network automatically reduces the feature weight ratio at the corresponding location, spatially reducing the contribution of interfering features in subsequent calculations. This completes feature selection and gain adjustment in the single-frame spatial dimension, achieving the optimization effect of highlighting target features and weakening interfering features in the spatial domain.

[0036] After optimizing the spatial features of a single frame, the process moves to temporal modeling of continuous multi-frame data. This stage uses multi-frame radar feature data stacked in time as the processing object, and uses temporal modeling algorithms to mine the correlation of motion changes between frames, extracting short-term local motion features of specific hand gestures. Hand gestures are characterized by a motion space concentrated in a small area of ​​the hand, short duration, continuous motion trajectory, and fixed local deformation patterns. In contrast, non-target interference motions in the scene are mostly characterized by a wider activity space, irregular motion duration, and significantly different motion trajectories and patterns of change from the target hand gestures. The algorithm sequentially performs analysis on each motion block based on spatial range, duration, and motion consistency. The attribute discrimination process accurately identifies interference blocks that do not belong to the gesture action from all motion areas. For the identified interference areas, a dynamic suppression weight that changes in real time with the scene is generated. Using this dynamic weight, a weighted suppression operation is performed on the multi-frame temporal features to gradually remove the interference motion components mixed in with the gesture data. Finally, the interference signal and the valid gesture signal are decoupled, and the gesture motion sequence that retains only the pure hand action information is output. After a two-layer optimization process of spatial weighting and temporal dynamic suppression, the final gesture motion sequence not only removes irrelevant clutter interference in the single-frame space, but also filters out motion aliasing caused by people walking around in the temporal dimension, providing high-quality temporal feature data with a very low interference ratio for gesture classification and recognition.

[0037] The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features, specifically: The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network. The lightweight spatiotemporal feature extraction network adopts a structure of lightweight Transformer combined with bidirectional gated recurrent units. At the same time, it adopts lightweight convolution, compact temporal modeling and efficient feature fusion strategies to compress the number of model parameters and computational cost, and focuses computational resources on the target gesture-related region to obtain high-dimensional spatiotemporal features.

[0038] The lightweight spatiotemporal feature extraction network relies on a composite architecture of a lightweight Transformer and a bidirectional gated recurrent unit to achieve spatiotemporal feature modeling. The lightweight Transformer architecture is mainly responsible for capturing the long-distance contextual dependencies of gesture spatial points across the entire temporal span. Relying on an improved lightweight self-attention mechanism, it establishes a correlation mapping between hand spatial position, distance value, and velocity changes at different times while reducing redundant parameters, thus depicting the overall evolution logic of the gesture from the initial action to the final action. The bidirectional gated recurrent unit performs refined extraction of local temporal features from continuous gesture sequences in both forward and reverse temporal directions. The forward traversal records the details of action evolution along the time sequence of the gesture action, while the reverse traversal traces back the changes in the early stages of the action from the end node of the gesture. The bidirectional operations complement each other, fully capturing subtle hand movement changes between short-term adjacent frames and improving local temporal correlation information.

[0039] The system employs lightweight convolution, compact temporal modeling, and efficient feature fusion to achieve model slimming and computational optimization. Lightweight convolution reduces redundant weights and floating-point operations in spatial convolution by splitting the convolution kernel and pruning sparse channels, compressing the computational overhead of spatial dimensions while preserving the spatial texture and distance-velocity distribution of the gesture region. Compact temporal modeling abandons the traditional redundant full temporal traversal method, performing modeling operations only on the selected effective gesture temporal segments, eliminating the unnecessary computational consumption caused by invalid blank frames and interference residual frames. The efficient feature fusion strategy performs hierarchical weighted fusion of long-range temporal features output by the Transformer and local short-term features output by the bidirectional gated recurrent unit, selecting and retaining effective feature components strongly related to gesture actions, discarding redundant and overlapping features, and further reducing feature dimensions and parameter count.

[0040] The entire network achieves significant compression in terms of both model parameter quantity and overall computational load, making it adaptable to the limited computing power environment of embedded devices. During the computation process, limited computing resources are concentrated on the effective area corresponding to the target gesture, reducing invalid computation on the remaining irrelevant background area. Based on the composite network structure, the contextual dependency logic of the gesture action across frames is mined. Finally, spatial morphology and temporal change information are integrated to generate high-dimensional spatiotemporal features containing refined spatial details and complete temporal evolution rules. These features comprehensively include the dynamic change information of distance, speed, and spatial position at each stage of the gesture.

[0041] The high-dimensional spatiotemporal features are input into the classifier, and the output gesture category recognition result is as follows: High-dimensional spatiotemporal features are input into a classifier consisting of a fully connected layer and a Softmax classifier, and the probability distribution of various gestures is output to obtain the final gesture category recognition result.

[0042] High-dimensional spatiotemporal features encompass the full range of characteristics of gestures, including spatial orientation, radial distance, speed, and temporal evolution. The data presents a high-dimensional, multi-dimensional fused tensor form, which cannot be directly used for classification calculations. Therefore, the feature data is input into a fully connected layer for feature mapping and dimensionality regularization. The fully connected layer, relying on the weight and bias parameters set within the layer, performs global correlation and weighting operations on the dispersed high-dimensional spatiotemporal features, aggregates and compresses redundant feature components, and uniformly maps irregular multi-dimensional features to a feature dimension space that matches the preset number of gesture categories. During the mapping process, the unique feature representation of each gesture is enhanced, the feature differences between different gestures are amplified, and the feature dispersion deviation within the same type of gesture is reduced, completing the format transformation from abstract spatiotemporal features to pre-classification feature vectors.

[0043] The feature vector, after being normalized by the fully connected layer, is then fed into the Softmax classifier for probability normalization. The Softmax classifier performs an exponential transformation and overall normalization on each feature value in the input feature vector, converting the originally unconstrained feature scores into probability values ​​between zero and one. The sum of the probability values ​​corresponding to all gesture categories is fixed at one, thus generating a probability distribution for all preset gesture categories. This probability distribution directly reflects the confidence level of the current input gesture belonging to each gesture category. After generating the complete probability distribution data, the category with the highest probability value is selected as the final classification category for the current gesture, thus obtaining the gesture category recognition result.

[0044] It also includes anti-interference optimization during the training phase, collecting multi-scene samples such as single person without interference, surrounding people stationary, surrounding people walking, surrounding people swinging their arms, and multiple people with dynamic interference, adding background noise samples for training, setting loss penalty terms and Top-2 softmax difference detection thresholds, and filtering low-confidence predictions.

[0045] In the dataset construction phase, five types of differentiated samples were collected in a gradient according to the interference level of actual application scenarios. Among them, the single-person interference-free sample served as the baseline sample, corresponding to the ideal test environment where only the target user makes gestures and there are no other people or extra movement interferences in the space, which is used to solidify the model's basic learning ability for various standard gesture ontological features; the surrounding people stationary sample retained the target gesture, while other people in the scene remained stationary. This type of sample introduced human static echo clutter, training the model's ability to distinguish effective gesture features from static clutter background; the surrounding people walking sample simulated the interference scenario of irregular pacing and movement of people in the daily environment. The walking brought a large-scale continuous Doppler velocity clutter, which can train the model to distinguish the feature differences between small hand gestures and large-scale human movement; the surrounding people arm swing sample... This study targets interference from unconscious arm raising or hand swinging by bystanders. Such interference is highly similar to the target gesture in terms of distance and speed characteristics, which can specifically improve the model's ability to distinguish similar motion interference. Multi-person dynamic interference samples are superimposed with composite interference conditions of multiple people walking and swinging their arms at the same time, recreating the extremely complex working conditions of multi-target signal superposition, comprehensively enhancing the model's anti-multi-interference performance. In addition to these five types of motion-related samples, background noise samples are added to include irregular static noise data such as walls, furniture, and environmental electromagnetic clutter, filling in the noise feature samples in blank scenes, allowing the model to fully learn the radar feature distribution patterns corresponding to various interference sources. Based on rich multi-condition samples, supervised training is completed, prompting the network to autonomously summarize the feature patterns of various interferences when learning gesture features, and completing the underlying optimization of anti-interference capability from the data source level.

[0046] In terms of model loss control, an additional dedicated loss penalty term is added during the training process. When the model encounters a sample with interference and misclassifies the class or shifts the confidence level, the penalty term increases the corresponding loss value, which in turn drives the network parameters to iteratively optimize. This tightens the feature discrimination boundary of the model, reduces the feature confusion caused by the incorporation of interference features, and constrains the model from overly favoring the feature patterns of clean samples, thus balancing the recognition logic of clean samples and interference samples.

[0047] In the output result filtering stage, low-confidence prediction filtering is achieved by relying on the Top-2 softmax difference detection threshold. The best class probability ranked first and the second best class probability ranked second are extracted from the probability distribution of each category output by Softmax. The difference between the two probabilities is calculated and compared with the pre-set detection threshold. When the difference is less than the set threshold, it means that the model cannot form an effective distinction between the two types of gestures and the confidence of the current prediction result is too low. The prediction is determined to be an invalid recognition result and is filtered out. Only the prediction result with a difference higher than the threshold is retained as a valid recognition output. This is to filter out unreliable judgments caused by excessive interference and blurred features. Finally, the model anti-interference optimization system is improved from the entire chain of training samples, loss constraints, and post-processing screening.

[0048] The feature weights in the dynamic weight evaluation mechanism are calculated using the following formula: ; in, Let N be the output feature of node i, N be the neighborhood set of node j, and D be the output feature of node i. ij Let X be the distance between nodes i and j, r be the neighborhood radius, MLP be a shared multilayer perceptron, and X be the distance between nodes i and j. j E is the input feature of node j. ij Let be the edge characteristics from node j to node i.

[0049] represents the optimized feature output by node i after dynamic weighting, which integrates the optimal association information within the node's effective neighborhood and is a representation of the result after dynamic weighting adjustment; N delineates the set of all available neighboring candidate nodes j globally, providing candidate data sources for neighborhood selection; D ij This is used to quantify the physical distance between node i and its neighboring candidate node j in the radar spatial dimension, corresponding to the spatial distance information of the radar point cloud. r is a pre-set neighborhood radius threshold. During the calculation, a distance filtering constraint is first executed to remove irrelevant nodes that are too far away and exceed the set spatial range, thus narrowing the neighborhood calculation range at the spatial level and filtering out interference nodes that are too far away and have no effective motion association, achieving coarse-grained neighborhood limitation in the spatial dimension. For the compliant neighboring node j that is retained after filtering, two basic input parameters are extracted and fed into a shared multilayer perceptron (MLP) for feature mapping calculation, where X... j E represents the native input features carried by the neighboring node j itself, corresponding to the original attribute information of that point in the radar data, such as distance, radial velocity, and echo energy, and carrying the target's own physical characteristics; ijThese are the edge features connecting node j to node i, used to characterize the association attributes between the two nodes, including the velocity difference, angle deviation, and difference in motion continuity between the two points, supplementing the coupling constraints between nodes. The shared multilayer perceptron (MLP) adopts a structure design with shared parameters across the entire network, and for each paired X... j With E ij A unified nonlinear feature transformation and weight mapping is performed. Relying on a multi-layer network to autonomously learn the intrinsic correlation between node and edge features, differentiated feature weights are automatically generated for different neighborhood nodes, completing the nonlinear purification and weighted optimization of the original features. After all compliant neighborhood nodes have undergone MLP feature transformation, a maximum value selection operation is performed on all output results, selecting the feature result with the best numerical value from all processed neighborhood features and assigning it to... Using this as the output feature after node i is updated, the operation logic of taking the maximum value can adaptively retain the feature component with the strongest correlation and the richest effective information in the neighborhood and automatically weaken the inferior feature contribution brought by invalid interference nodes in the neighborhood, thus fully implementing the core logic of dynamic weight evaluation.

[0050] The entire calculation process progresses step by step from spatial distance screening, node and edge dual feature input, shared network nonlinear weighting, and optimal feature selection. It dynamically adjusts feature weights based on the spatial distribution of radar point clouds and node association attributes. During the hand gesture point cloud feature processing, it automatically focuses on the target hand-related nodes and suppresses the feature weights of clutter and interference point clouds, providing standardized node feature data with dynamic weight optimization for anti-interference feature extraction.

[0051] The point cloud data output by millimeter-wave radar includes sequence number, timestamp, point number, three-dimensional coordinates, range, velocity, Doppler cells, azimuth angle, and reflection intensity fields.

[0052] Radar feature representation includes input forms of point clouds and time-series images.

[0053] The sequence number is used to mark the acquisition frame number corresponding to a single frame of radar data. This field distinguishes point cloud groups acquired under different radar sweep cycles, enabling time-series grouping and aggregation of multi-frame data. It serves as the basic index identifier for stitching together continuous time sequences. The timestamp records the absolute moment each radar reflection point is received and acquired. The timestamp allows for precise alignment of the time coordinates of different sampling points, providing a precise time-series reference for cross-frame motion continuity analysis and short-term gesture time-series modeling. The point number assigns a unique code to each spatial reflection point within the same frame, distinguishing each independent spatial point in the single-frame point cloud. This facilitates single-point filtering, neighborhood association, and data tracing of the point cloud. The three-dimensional coordinates define the X, Y, and Z three-dimensional spatial positions of the reflection point within the radar detection space, intuitively representing the three-dimensional spatial distribution of target reflection points. This is a core spatial parameter for constructing point cloud topology and defining spatial neighborhood ranges. The distance field directly represents the radial straight-line distance between the corresponding reflecting target and the radar antenna, calculated from the radar difference frequency signal via distance F. The FT calculation is used to define the distribution of targets at near and far distances, and is also a key basis for screening close-range hand gesture targets and eliminating long-range interference clutter. The velocity field is obtained by converting Doppler frequency offset to obtain the radial movement rate of the target relative to the radar. Positive and negative values ​​can also distinguish the direction of movement of the target as it approaches or moves away from the radar, and can intuitively distinguish the difference in movement rate caused by micro-movement of the hand and large-scale walking by others. The Doppler unit corresponds to the frequency domain index position after the velocity FFT calculation, which is used to anchor the velocity discrete interval to which the point cloud belongs, and assists the algorithm to batch collect point clusters in the same velocity interval, realizing the grouping of point clouds with similar motion attributes. The azimuth angle represents the horizontal deflection angle of the reflecting target relative to the radar normal direction. Combined with distance and three-dimensional coordinates, it improves the target's planar azimuth information and supports the data analysis of spatial region division and beam dimension. The reflection intensity corresponds to the magnitude of the target echo signal, which is determined by the target material and surface area. The reflection intensity of different objects such as hands, walls, and clothing is significantly different. This field can be used to distinguish effective targets from environmental clutter from the energy dimension.

[0054] After outputting multi-field point cloud data, the radar feature representation is split into parallel input formats of point cloud and temporal image. The point cloud format fully retains the discrete spatial point information of the above nine fields, adapting to algorithm modules that rely on single-point topological association, such as graph neural networks and neighborhood weight calculation. It can flexibly extract single-point distance, velocity, and edge association features to complete dynamic weight calculation. The temporal image format maps multiple consecutive frames of point cloud into regular two-dimensional image data according to the distance, velocity, and orientation dimensions, transforming discrete point information into pixelated feature arrangement. It is adapted to network structures for gridded data, such as lightweight convolution and Transformer. The two input formats match the data input requirements of different algorithm branches, taking into account the dual advantages of retaining the original information of discrete points and convenient operation of regular images, realizing the standardized conversion of radar bottom-level data into multi-path feature input for subsequent algorithms.

[0055] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0056] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for recognizing multiple-person interference gestures using millimeter-wave radar, characterized in that, Includes the following steps: The raw echo data of millimeter-wave radar is collected, and the raw echo data is preprocessed to obtain radar feature representations containing range, velocity, and angle dimensions; Based on radar feature representation, a reliability-aware dynamic radar apparent motion field is constructed. The target gesture motion and non-target personnel interference are decoupled through a dynamic weight evaluation mechanism to extract the gesture motion sequence. The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features. Input high-dimensional spatiotemporal features into the classifier and output gesture category recognition results.

2. The method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 1, characterized in that, The raw echo data is preprocessed to obtain radar feature representations that include range, velocity, and angle dimensions, specifically: The raw echo data is processed sequentially by range FFT and Doppler FFT, and static clutter is suppressed by background removal or mean elimination. Then, it is processed by antenna dimension angle FFT or beamforming to obtain radar feature representation containing range, velocity and angle dimensions.

3. The method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 2, characterized in that, Also includes: Based on the energy distribution, range variation, velocity variation and motion continuity represented by radar features, a rough motion area of ​​the target user is extracted to generate a target area guidance map; The target area guidance map is used as an additional channel and spliced ​​with the radar feature representation to form an enhanced input feature.

4. The method for recognizing multiple interference gestures using millimeter-wave radar according to claim 3, characterized in that, The dynamic weight evaluation mechanism is as follows: The enhanced input features are fed into a lightweight feature extraction network, which automatically learns the feature weights of the target region and non-target regions, enhances the response of the target gesture region, and reduces the feature contribution of non-target motion regions. Temporal modeling is performed on the radar feature representation of consecutive multiple frames to extract short-term local motion features of gestures. Interference areas are identified based on the spatial range, duration, and consistency with the target gesture of non-target motion, and a decoupled gesture motion sequence is generated by dynamically suppressing weight output.

5. The method for recognizing multiple interference gestures using millimeter-wave radar according to claim 4, characterized in that, The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network to determine the contextual dependencies of the gesture actions and obtain high-dimensional spatiotemporal features, specifically: The gesture motion sequence is input into a lightweight spatiotemporal feature extraction network. The lightweight spatiotemporal feature extraction network adopts a structure of lightweight Transformer combined with bidirectional gated recurrent units. At the same time, it adopts lightweight convolution, compact temporal modeling and efficient feature fusion strategies to compress the number of model parameters and computational load, and focuses computational resources on the target gesture-related region to obtain high-dimensional spatiotemporal features.

6. The method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 5, characterized in that, The high-dimensional spatiotemporal features are input into the classifier, and the output gesture category recognition result is as follows: High-dimensional spatiotemporal features are input into a classifier consisting of a fully connected layer and a Softmax classifier, and the probability distribution of various gestures is output to obtain the final gesture category recognition result.

7. The method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 6, characterized in that, It also includes anti-interference optimization during the training phase, collecting multi-scene samples such as single person without interference, surrounding people stationary, surrounding people walking, surrounding people swinging their arms, and multiple people with dynamic interference, adding background noise samples for training, setting loss penalty terms and Top-2 softmax difference detection thresholds, and filtering low-confidence predictions.

8. A method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 7, characterized in that, The feature weights in the dynamic weight evaluation mechanism are calculated using the following formula: ; in, Let N be the output feature of node i, N be the neighborhood set of node j, and D be the output feature of node i. ij Let be the distance between nodes i and j, r be the neighborhood radius, MLP be a shared multilayer perceptron, and X be the distance between nodes i and j. j E is the input feature of node j. ij Let be the edge characteristics from node j to node i.

9. A method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 8, characterized in that, The point cloud data output by the millimeter-wave radar includes sequence number, timestamp, point number, three-dimensional coordinates, distance, velocity, Doppler cells, azimuth angle, and reflection intensity fields.

10. A method for recognizing multiple-person interference gestures using millimeter-wave radar according to claim 9, characterized in that, The radar feature representation includes input forms of point clouds and time-series images.