Radar-based human skeleton estimation method, system and product

By generating a four-dimensional point cloud of the target and dividing it into human body topology point cloud blocks, the accuracy and robustness issues of radar human body attitude estimation are solved by using a human skeleton estimation network, achieving stable and high-precision human body attitude prediction in all weather conditions.

CN121817844AActive Publication Date: 2026-04-10SHENZHEN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing radar-based human pose estimation methods have shortcomings in terms of accuracy and robustness. In particular, millimeter-wave radar point clouds are sparse and noisy, making it difficult to construct implicit human topological priors, and they are affected by illumination, occlusion and environmental changes.

Method used

By performing target detection on radar echo data to generate a four-dimensional point cloud of the target, the human body topology point cloud blocks are divided based on the similarity of velocity direction, and the human skeleton estimation network is used to estimate the human skeleton, thus constructing a priori human body topology structure and enhancing the model's ability to model the spatial relationships of key parts of the human body.

Benefits of technology

It achieves stable human posture estimation around the clock, improves the accuracy and robustness of posture prediction, is suitable for places with high security and privacy requirements, and adapts to diverse daily human behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817844A_ABST
    Figure CN121817844A_ABST
Patent Text Reader

Abstract

The invention provides a radar-based human skeleton estimation method, system and product, and the method comprises the steps: carrying out the target detection of collected radar echo data, and generating a target four-dimensional point cloud containing distance information, speed information and angle information based on a target detection result; dividing the target four-dimensional point cloud into a plurality of human body topology point cloud blocks based on the velocity direction similarity of each point in the target four-dimensional point cloud; and inputting the human body topology point cloud block into a human body skeleton estimation network for human body skeleton estimation to obtain a human body skeleton estimation result. According to the human skeleton estimation method provided by the invention, human topology priori can be constructed from sparse millimeter wave radar point clouds, point cloud block features sensed by a structure are extracted, the modeling capability of a model for the spatial relationship of key parts of a human body is enhanced, diversified and natural daily human behaviors in a non-inductive monitoring scene can be adapted, and the human skeleton estimation accuracy is improved. The human body skeleton estimation with higher generalization ability is realized, and the accuracy and robustness of human body posture prediction can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of radar, in particular to a human skeleton estimation method, system and product based on radar. BACKGROUND

[0002] With the increasing demand for daily health monitoring, abnormal behavior recognition and fall warning, through continuous and natural perception and analysis of human posture and action in a home environment, early identification of risks such as declining mobility and abnormal gait can be achieved, and timely warning and auxiliary response can be provided in the event of an emergency. However, due to the cost of professional care and the increasing number of people living alone, there is an urgent need for a long-term monitoring technology that can balance privacy protection while ensuring safety. Among them, human skeleton estimation, as an important basic technology for describing human posture and action, has important application value in behavior understanding, health assessment and safety monitoring scenarios.

[0003] There are many types of sensors currently in use, mainly including: wearable devices, which infer the human skeleton by wearing sensors on specific parts of the human body and collecting inertial measurement unit data, but are susceptible to calibration errors and drift due to the need for long-term user wear; visual and infrared sensors, which capture human appearance images or thermal distribution information and perform key point detection and skeleton reconstruction, but are highly sensitive to light, occlusion and environmental conditions, and pose privacy risks; Wi-Fi radio frequency devices, which estimate the human skeleton by mapping wireless channel information into human spatial structure information, but are susceptible to bandwidth and spatial resolution limitations inherent to the communication system, and are susceptible to multipath effects and environmental changes; millimeter wave radar, which transmits millimeter waves and receives their reflected signals, uses distance, Doppler and angle information to obtain human scatter point distribution, and has advantages such as non-contact, anti-occlusion and privacy-friendly. However, millimeter wave radar point clouds are naturally sparse, noisy and have discontinuous shapes in local regions, and lack explicit geometric and topological connection relationships between body parts, making it difficult to construct implicit human topological priors without explicit skeleton labels, and to use point cloud blocks to model the association between different joint regions of the human topological structure. SUMMARY

[0004] The purpose of the present application is to provide a human skeleton estimation method, system and product based on radar, which can at least solve the problem of poor accuracy and robustness of the human posture estimation method based on radar in the related art.

[0005] To solve the above technical problems, the first aspect of the embodiments of the present application provides a human skeleton estimation method based on radar, comprising: detecting the collected radar echo data, and generating a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result; the target four-dimensional point cloud is divided into a plurality of human topological point cloud blocks based on the similarity of the velocity directions of the points in the target four-dimensional point cloud; The human topological point cloud blocks are input into a human skeleton estimation network for human skeleton estimation, to obtain a human skeleton estimation result.

[0006] The second aspect of the embodiments of the present application provides a human skeleton estimation system based on a radar, comprising: A point cloud generation module is configured to perform target detection on the collected radar echo data, and generate a target four-dimensional point cloud containing distance information, velocity information and angle information based on the target detection result. A point cloud division module is configured to divide the target four-dimensional point cloud into a plurality of human topological point cloud blocks based on the similarity of the velocity directions of the points in the target four-dimensional point cloud. A skeleton estimation module is configured to input the human topological point cloud blocks into a human skeleton estimation network for human skeleton estimation, to obtain a human skeleton estimation result.

[0007] The third aspect of the present application provides a radar, comprising a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory, and when the processor executes the computer program, each step of the human skeleton estimation method described in the first aspect of the embodiments of the present application is implemented.

[0008] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, each step of the human skeleton estimation method described in the first aspect of the embodiments of the present application is implemented.

[0009] From the above, the embodiment of the present application first performs target detection on the collected radar echo data, and generates a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result, then divides the target four-dimensional point cloud into multiple human body topology point cloud blocks based on the speed direction similarity of each point in the target four-dimensional point cloud, and finally inputs the human body topology point cloud block into a human body skeleton estimation network for human body skeleton estimation to obtain a human body skeleton estimation result. The human body skeleton estimation method proposed in the present application has the privacy-friendly feature compared with the human body skeleton estimation method of other sensors, and is not affected by light and weather conditions, can realize all-weather stable operation, and is suitable for places with high safety and privacy requirements; in addition, the method can construct human body topology prior from sparse millimeter wave radar point cloud, extract structure-aware point cloud block features, enhance the modeling ability of the model for the spatial relationship of human body key parts, further, compared with the existing method which relies more on the single behavior type and limited data scale of the data set, the method can adapt to the diversified and natural daily human behaviors in the non-contact monitoring scene, realize human body skeleton estimation with more generalization ability, and thus can effectively improve the accuracy and robustness of human body posture prediction.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the related art or the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the related art or the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and not all embodiments. Those skilled in the art can also obtain other drawings according to these drawings without creating any creative labor.

[0012] Figure 1 A flowchart of the human body skeleton estimation method provided by the embodiment of the present application; Figure 2 A real skeleton graph collected by a depth camera when a human body performs a leg kicking motion in the human body skeleton estimation method provided by the embodiment of the present application; Figure 3 A three-dimensional point cloud clustering result graph when a human body performs a leg kicking motion in the human body skeleton estimation method provided by the embodiment of the present application; Figure 4 A real skeleton graph collected by a depth camera when a human body performs a double-hand lifting action in the human body skeleton estimation method provided by the embodiment of the present application; Figure 5A three-dimensional point cloud clustering result diagram of a human body performing a double-hand lifting action in a human skeleton estimation method provided by an embodiment of the present application; Figure 6 A loss constraint concept schematic diagram in a human skeleton estimation method provided by an embodiment of the present application; Figure 7 A human skeleton estimation network framework schematic diagram in a human skeleton estimation method provided by an embodiment of the present application; Figure 8 A refinement process schematic diagram of a human skeleton estimation method provided by an embodiment of the present application; Figure 9 A program module schematic diagram of a human skeleton estimation system provided by an embodiment of the present application; Figure 10 A module block diagram of a radar provided by an embodiment of the present application; Figure 11 A module block diagram of a computer-readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION

[0013] In order to make the objectives, technical solutions and advantages of the present application more obvious and easy to understand, the present application will be described clearly and completely below in conjunction with the embodiments of the present application and the accompanying drawings, in which the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. It should be understood that the embodiments of the present application described below are only used to explain the present application and do not limit the present application, that is, based on the various embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application. In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0014] With the increasing demand for daily health monitoring, abnormal behavior recognition and fall warning, through continuous and natural perception and analysis of human posture and action in a home environment, early identification of risks such as declining mobility and abnormal gait can be achieved, and timely warning and auxiliary response can be provided in the event of an emergency. However, due to the limitation of professional nursing costs and the increasing number of people living alone, there is an urgent need for a long-term monitoring technology that can ensure safety while protecting privacy. Among them, human skeleton estimation as an important basic technology for describing human posture and action has important application value in behavior understanding, health assessment and safety monitoring scenarios.

[0015] Currently, there are many types of sensors in use, mainly including: wearable devices, which collect inertial measurement unit data by wearing sensors on specific parts of the human body to infer the human skeleton, but are prone to calibration errors and drift due to long-term wear by users; visual and infrared sensors, which capture human appearance images or heat distribution information and perform key point detection and skeleton reconstruction, but are highly sensitive to light, occlusion and environmental conditions, and pose privacy risks; Wi-Fi radio frequency devices, which map wireless channel information to human spatial structure information to estimate the human skeleton, but are subject to bandwidth and spatial resolution limitations inherent to the communication system, and are susceptible to multipath effects and environmental changes; millimeter wave radar, which transmits millimeter waves and receives their reflected signals, uses distance, Doppler and angle information to obtain human scattering point distribution, and has advantages such as non-contact, anti-occlusion and privacy-friendly. However, millimeter wave radar point clouds are naturally sparse, noisy and have discontinuous local region shapes, and lack explicit geometric and topological connection relationships between body parts, making it difficult to construct implicit human topological priors without explicit skeleton labels, and use point cloud blocks to model the association between different joint regions of the human topological structure.

[0016] In summary, the present application provides a human skeleton estimation method, specifically, please refer to Figure 1 , Figure 1 The flowchart of the radar-based human skeleton estimation method provided by the embodiments of the present application is shown in the figure. The human skeleton estimation method includes the following steps 101-103.

[0017] Step 101, target detection is performed on the collected radar echo data, and a target four-dimensional point cloud containing distance information, velocity information and angle information is generated based on the target detection result.

[0018] First, radar echo data is collected and target detection is performed, and effective signal components related to human targets are selected, and a target four-dimensional point cloud is generated based on the target detection result. The target four-dimensional point cloud contains at least distance information, velocity information and angle information, and can fully depict the position and motion characteristics of the human target in space.

[0019] In some embodiments of the present application, before performing target detection on the collected radar echo data, it further includes: performing signal preprocessing on the collected radar echo signal to obtain a multi-dimensional echo signal containing distance, velocity and angle information.

[0020] Specifically, a multi-channel millimeter-wave radar continuously transmits electromagnetic waves into the monitoring space and receives echo signals reflected from objects. After signal preprocessing such as signal amplification, mixing, and analog-to-digital conversion, a multi-dimensional echo signal containing information in the range, velocity, and angle dimensions is obtained. Subsequently, power spectrum analysis is performed on the multi-dimensional echo signal, and target detection algorithms such as constant false alarm rate are used to determine the existence of the target and parameters such as range and velocity. For the effective target points detected, information such as angle and distance is obtained.

[0021] Furthermore, in some embodiments of this application, the acquired radar echo signal is preprocessed to obtain a multidimensional echo signal containing range, velocity, and angle dimensions. This includes: preprocessing the acquired raw radar echo signal to obtain a first discrete echo signal. ;in, For fast time dimension, it means the first... One sampling point, Let be the slow time dimension, indicating the th A linear frequency modulated continuous wave signal Let be the antenna dimension, and let represent the th . The received signals from each channel are analyzed; based on the first discrete echo signal, a fast Fourier transform is performed in the fast time dimension to obtain the second discrete echo signal. The second discrete echo signal is a range-slow time-antenna dimension signal. This indicates range cell sampling; range-dimensional clutter suppression and coherent accumulation are performed based on the second discrete echo signal to obtain the initial multidimensional echo signal. The final multidimensional echo signal is obtained by performing a Fast Fourier Transform (FFT) on the initial multidimensional echo signal in the slow time dimension. Among them, the multidimensional echo signal is a range-Doppler-antenna dimensional signal. , indicating Doppler unit sampling.

[0022] Specifically, firstly, the radar continuously transmits electromagnetic wave signals into space. After being reflected by objects, the signals are received by the radar receiver. Then, after being sampled by a signal amplifier, mixer, and ADC, a discrete echo signal containing the distance, velocity, and angle dimensions is obtained, which is the first discrete echo signal.

[0023] The first discrete echo signal can be represented as: ,in, For fast time dimension, it means the first... One sampling point; Let be the antenna dimension, and let represent the th . The received signals of each channel.

[0024] Then based on the first discrete echo signal, first perform fast Fourier transform in the fast time dimension to convert the signal from the time domain to the frequency domain to obtain a second discrete echo signal. The second discrete echo signal, that is, the distance-slow time dimension-antenna dimension signal, can be represented as , wherein , represents a distance unit sample.

[0025] Further, the distance dimension clutter suppression and coherent accumulation are performed on the second discrete echo signal, and the environmental noise and non-target echo interference are filtered out through a sliding average algorithm or a time domain accumulation method to obtain a multi-dimensional signal after clutter suppression and coherent accumulation, that is, an initial multi-dimensional echo signal, which can be represented as .

[0026] Finally, the initial multi-dimensional echo signal is further subjected to fast Fourier transform in the slow time dimension to further extract target speed information to obtain a final multi-dimensional echo signal, that is, a distance-Doppler-antenna dimension signal, which can be represented as , wherein , represents a Doppler unit sample.

[0027] It can be understood that the multi-channel millimeter wave radar selected by the embodiments of the present application can be a one-transmit multi-receive radar or a multi-transmit multi-receive radar, and the present application does not make specific limitations thereto.

[0028] In the embodiments of the present application, target detection is performed on the collected radar echo data, and a target four-dimensional point cloud containing distance information, speed information and angle information is generated based on the target detection result, including: performing fast Fourier transform processing on the collected single-frame radar echo data to obtain short-time scale point cloud information; performing non-uniform Fourier transform processing on the collected multi-frame radar echo data to obtain long-time scale point cloud information; performing target detection based on the short-time scale point cloud information and the long-time scale point cloud information, and generating a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result.

[0029] Specifically, considering that the motion states of each bone joint of the human body not only differ at the same time, but also present obvious dynamic change characteristics at different time scales. If the same observation time window and coherent integration strategy are used for point cloud generation for fast motion joints and slow motion joints, the Doppler information of high-speed motion joints may be broadened or the characteristics of low-speed motion joints may be submerged, thereby causing the loss of part of the key joint information. Therefore, the present application introduces a multi-time scale point cloud generation mechanism.

[0030] Firstly, the short-time scale point cloud generation operation is performed. For the radar echo data collected in a single frame, after the three-dimensional data cube signal is obtained by down-conversion and analog-to-digital conversion, the feature capture of the human body fast motion joint (such as arm swinging, leg kicking, etc.) is focused on, and the fast Fourier transform (FFT) is directly performed on the Doppler dimension of the single-frame data. This processing method can quickly extract point cloud information reflecting the instantaneous motion characteristics of the target, avoid the Doppler information broadening of the high-speed motion joint due to the long observation time window, and finally obtain short-time scale point cloud information. The core advantage lies in the accurate capture of instantaneous dynamic characteristics.

[0031] Secondly, the long-time scale point cloud generation operation is performed. For the radar echo data collected in continuous multiple frames, considering that the features of the human body slow motion joint (such as trunk swinging, lower limbs when walking slowly) are easily submerged by noise, the Doppler resolution needs to be improved to enhance the motion continuity description. Therefore, after the Doppler dimension of the continuous multi-frame data is fused, the non-uniform fast Fourier transform (NUFFT) is used for signal analysis. This transform can adapt to the non-uniform distribution characteristics of multi-frame data in the Doppler dimension, effectively improve the feature recognition of low-speed motion targets, and finally obtain long-time scale point cloud information.

[0032] Finally, based on the above two types of point cloud information, target detection is performed and target four-dimensional point cloud is generated. The short-time scale point cloud information focuses on instantaneous motion characteristics, and the long-time scale point cloud information emphasizes high-resolution steady-state characteristics. The two complement each other to form a comprehensive representation of target features. By jointly performing target detection on the two types of point cloud information, the effective signal components related to the human body target are selected, and environmental noise and clutter interference are removed. Finally, the target four-dimensional point cloud containing radial distance, velocity information and angle information is obtained. This target four-dimensional point cloud can accurately depict the dynamic changes of the human body fast motion joint and clearly present the spatial position of the slow motion joint. By fusing the four-dimensional point clouds generated at different time scales, a multi-time scale point cloud representation that takes into account both fast and slow motion characteristics can be obtained.

[0033] In the embodiment of the present application, target detection is performed based on short-time-scale point cloud information and long-time-scale point cloud information, and a target four-dimensional point cloud containing distance information, speed information and angle information is generated based on the target detection result, including: performing target detection on the short-time-scale point cloud information and the long-time-scale point cloud information based on a constant false alarm rate algorithm to obtain a target detection result; wherein the target detection result includes effective target points; extracting distance-Doppler indexes of the effective target points; wherein the distance-Doppler indexes include radial distance information and radial speed information of the effective target points; performing beamforming on the short-time-scale point cloud information and the long-time-scale point cloud information corresponding to the distance-Doppler indexes to extract angle indexes of the effective target points; wherein the angle indexes include azimuth angle information and elevation angle information of the effective target points; establishing a spatial coordinate system based on a radar installation mode, and obtaining four-dimensional coordinates of the effective target points based on the distance-Doppler indexes and the angle indexes to generate a target four-dimensional point cloud containing distance information, speed information and angle information.

[0034] First, an effective target point screening operation is performed. For the generated short-time-scale point cloud and long-time-scale point cloud, a constant false alarm rate (CFAR) detection algorithm is used for target detection. This algorithm dynamically generates a detection threshold by statistically analyzing the noise power of the local neighborhood in the point cloud data, and determines that the points with signal power higher than the threshold are effective target points, and the points with signal power lower than the threshold are environmental noise or clutter and are removed. Through this operation, the set of effective target points related to human targets can be screened from the two types of time-scale point clouds.

[0035] Second, the distance-Doppler indexes of the effective target points are extracted. The distance-Doppler index is a key parameter representing the core motion and position information of the effective target points, and its extraction is based on the preprocessing results of the two types of point clouds: after the short-time-scale point cloud and the long-time-scale point cloud are processed by distance dimension FFT and Doppler dimension FFT / NUFFT, they have formed a feature spectrum containing distance and speed information. The radial distance information and radial speed information corresponding to each effective target point are directly extracted from this spectrum, i.e., the construction of the distance-Doppler index is completed.

[0036] Then, angle estimation is performed based on the range-Doppler index. For each valid target point, according to its range-Doppler index, the corresponding antenna channel dimension signal is extracted from the three-dimensional data cube signal of the short time scale point cloud and the long time scale point cloud; the antenna channel signal is processed by using a digital beamforming (DBF) algorithm, the direction of arrival of the valid target point is estimated by weighted synthesis of multi-antenna receiving signals, and then the angle index containing azimuth angle information and elevation angle information is obtained. The angle index cooperates with the range-Doppler index to completely depict the spatial position and motion state of the valid target point.

[0037] Finally, a spatial coordinate system is established and a target four-dimensional point cloud is generated. According to the installation mode of the radar (such as side-mounted on a support or a wall), a three-dimensional spatial coordinate system suitable for the monitoring scene is established; the range-Doppler index (radial distance, radial velocity) and the angle index (azimuth angle, elevation angle) of each valid target point are substituted into the coordinate system, and the four-dimensional coordinates of the valid target point are calculated through coordinate conversion; the four-dimensional coordinates of all valid target points are integrated to generate a target four-dimensional point cloud containing distance information, velocity information and angle information. The target four-dimensional point cloud has the advantages of the two types of time scale point clouds, and through accurate index extraction and angle estimation, the accuracy of each dimension information is ensured.

[0038] Specifically, in some embodiments of the present application, when target detection is performed, that is, target detection is actually performed based on the multi-dimensional echo signal, including: obtaining the Doppler power spectrum of the multi-dimensional echo signal; performing peak extraction on the Doppler power spectrum based on a constant false alarm rate algorithm, and performing target detection on the extracted peak points to obtain a target detection result; wherein the target detection result includes valid target points.

[0039] Specifically, the multi-dimensional echo signal contains distance dimension , velocity dimension and antenna dimension information, and the range-Doppler-antenna dimension signal The power spectrum of the signal can be used to obtain the Doppler power spectrum. By analyzing the Doppler power spectrum of the signal, the presence of a target within the radar detection range and the target's distance and velocity information can be directly reflected. Then, the peak values ​​in the power spectrum are extracted using constant false alarm rate (CFAR) algorithms. CFAR algorithms specifically include Cell-Averaging Constant False Alarm Rate (CA-CFAR) and Ordered Statistics Constant False Alarm Rate (OS-CFAR). By setting a detection threshold that matches the intensity of environmental clutter, peak points corresponding to the true target are selected, noise and interference are eliminated, and valid target points are obtained, thereby determining the existence of the target and its preliminary position and velocity parameters.

[0040] Specifically, the first step is to extract the distance-Doppler of the effective target points. An index is generated, corresponding to the range and velocity characteristics of the target in the multidimensional echo signal. Then, the index corresponding to the above index is retrieved. The antenna-dimensional signal is used for beamforming, and the peak power spectrum index after beamforming is taken as the angle index corresponding to the effective points. Thus, the distance, velocity, angle, and energy information of each effective point of the target are known, enabling precise target location. Commonly used beamforming methods include Digital Beamforming (DBF), Multiple Signal Classification (MUSIC), and Minimum Variance Distortionless Response (MVDR) algorithm (also known as the Capon algorithm). Finally, a spatial coordinate system is established according to the millimeter-wave radar installation method (the spatial coordinate system varies depending on the installation method; common installation methods include side mounting at 45°, side mounting, top mounting, etc.). In the spatial coordinate system, the coordinates of all effective points of the target, i.e., the coordinates of all point cloud imaging points, can be obtained through the distance and angle information. All coordinate points constitute the target point cloud, intuitively representing the spatial distribution of the human body within the radar detection range.

[0041] Step 102: Based on the similarity of the velocity directions of each point in the target four-dimensional point cloud, divide the target four-dimensional point cloud into multiple human body topology point cloud blocks.

[0042] Specifically, based on the similarity of the speed direction of each point in the target four-dimensional point cloud, the target four-dimensional point cloud is topologically structured and divided to obtain multiple human topological point cloud blocks. During human motion, the scattering points of the same limb part have consistent speed direction characteristics, and the scattering points of different limb parts have significant differences in speed direction. This method utilizes the kinematic characteristics of human body, measures the similarity of the speed direction of each point, adaptively divides the discrete sparse point cloud into point cloud blocks with implicit topological correlation, and thus constructs the human topological structure prior, thereby solving the technical problem of lack of explicit body part connection relationship of sparse point cloud.

[0043] In the embodiment of the present application, based on the speed direction similarity of each point in the target four-dimensional point cloud, the target four-dimensional point cloud is divided into multiple human topological point cloud blocks, including: performing unit normalization processing on the radial velocity information of the effective target points to obtain normalized velocity vectors; constructing a distance metric based on the cosine similarity of each normalized velocity vector, and performing clustering processing on the effective target points according to the distance metric to obtain a clustering processing result; and dividing the effective target points based on the clustering processing result to obtain multiple human topological point cloud blocks.

[0044] Specifically, for the four-dimensional point cloud data generated in multiple time scales, the present application further introduces a human topological structure modeling mechanism. Inspired by the kinematic characteristics of human body, the direction consistency of the velocity vectors of each scattering point in the radar point cloud is considered, and the human point cloud is topologically structured and divided by measuring the similarity between the speed directions. Specifically, the cosine similarity is used as a measure of the similarity of the speed direction, and the human dynamic point cloud is adaptively divided into five main topological regions, respectively corresponding to the main trunk, left upper limb, right upper limb, left lower limb and right lower limb. For example, as shown in FIG. 6, it is a real skeleton graph collected by a depth camera when a human body performs a leg kicking motion, as shown in FIG. 7, it is a three-dimensional point cloud clustering result graph when a human body performs a leg kicking motion, as shown in FIG. 8, it is a real skeleton graph collected by a depth camera when a human body performs a double-hand lifting action, as shown in FIG. 9, it is a three-dimensional point cloud clustering result graph when a human body performs a double-hand lifting action. Figure 2 Figure 3 Figure 4 Figure 5

[0045] ​​​​In detail, first, unit normalization processing of radial velocity information is performed. For the effective target points screened, the radial velocity information of each point is extracted and a velocity vector is constructed, and unit normalization operation is performed on the velocity vector. The core purpose of this operation is to eliminate the difference in velocity amplitude (i.e., the speed of movement) between different effective target points, and only to retain the velocity direction information. When the human body moves, the movement directions of the effective target points of the same limb part (such as the left upper limb and the main trunk) are naturally consistent, while the movement directions of the effective target points of different limb parts are significantly different. After normalization processing, this core feature can be highlighted, and a direction basis for subsequent similarity judgment is provided. Finally, a normalized velocity vector with a constant modulus of 1 and consistent with the original velocity vector in direction is obtained.

[0046] Secondly, a distance metric is constructed based on cosine similarity and clustering processing is performed. For the normalized velocity vectors of all effective target points, the cosine similarity between any two vectors is calculated. The similarity is used to quantify the degree of coincidence of the velocity directions of two effective target points. The closer the similarity is to 1, the more consistent the velocity directions of the two points are, and the greater the probability that they belong to the same limb part. The closer the similarity is to 0 or -1, the greater the difference in velocity direction, and the greater the probability that they belong to different limb parts. In order to adapt to the distance minimization clustering logic of the clustering algorithm, the cosine similarity is converted into a distance metric, and then the K-means clustering algorithm is used to cluster all effective target points, so that the effective target points with consistent velocity directions are classified into the same cluster, and the clustering processing result is obtained.

[0047] Finally, the human topological point cloud block is divided based on the clustering processing result. After clustering processing, each cluster corresponds to a human body part with implicit topological association (such as a main trunk cluster and a left upper limb cluster), and each cluster is taken as an independent human topological point cloud block to complete the structured division of the target four-dimensional point cloud.

[0048] It can be understood that, in order to avoid the limitation of the model expression ability caused by manually setting fixed body part labels, the present application only gives each point cloud cluster an independent coding identifier, without hard constraint on the specific human body part corresponding to the point cloud cluster, and the potential topological semantic relationship is adaptively learned by the subsequent network during the training process.

[0049] In addition, in some embodiments of the present application, in order to unify the data into a fixed size tensor to adapt to the network input, for the human topological point cloud block with a point number less than the preset standard, a two-point linear interpolation operation is performed according to the proportion of the original point number of each cluster, and the point number is padded to the preset point number N, so as to ensure that all human topological point cloud blocks are fixed size tensors.

[0050] Step 103, inputting the human topological point cloud block into a human skeleton estimation network to perform human skeleton estimation, and obtaining a human skeleton estimation result.

[0051] Specifically, the divided human topological point cloud block is input to a human skeleton estimation network trained in advance, and the structural features of the topological point cloud block are learned and modeled through the network, and finally the human skeleton estimation result is output. The human skeleton estimation network can fully mine the local topological information and global pose correlation contained in the point cloud block, improve the ability to describe the spatial relationship of the key joints of the human body, and thus ensure the accuracy and robustness of the skeleton estimation.

[0052] In the embodiment of the present application, the human skeleton estimation network comprises a point cloud block topology modeling module, a multi-level feature fusion module and a skeleton regression module connected in sequence; the point cloud block topology modeling module is used for dynamically modeling the input human topological point cloud block, and outputting point-level features with enhanced topological semantics; the multi-level feature fusion module is used for adaptively fusing the point-level features with enhanced topological semantics, original point cloud geometric features and global context features, and outputting point cloud fusion features; and the skeleton regression module is used for mapping the point cloud fusion features to human skeleton joint coordinates to obtain the human skeleton estimation result.

[0053] Firstly, the point cloud block topology modeling module performs dynamic topology modeling. The input of this module is the divided human topological point cloud block. Although these point cloud blocks have implicit human part association, they still lack explicit topological structure feature expression. The module mines the spatial geometric relationship and semantic association between points and clusters and between clusters, dynamically models the discrete point cloud block, strengthens the structured association between points and clusters and between clusters, and finally outputs point-level features with enhanced topological semantics. This feature not only retains the local geometric information of individual points, but also integrates the global topological constraints of human parts, solving the technical difficulty of sparse point cloud in describing human structure association.

[0054] Secondly, the multi-level feature fusion module performs adaptive feature fusion. This module receives three types of input features: first, the point-level features with enhanced topological semantics output by the point cloud block topology modeling module; second, the original point cloud geometric features directly extracted from the human topological point cloud block, reflecting the spatial position attributes of points; and third, the pose encoding global context features obtained through global pooling, reflecting the overall pose trend of the human body. Since the three types of features correspond to different semantic levels: local topology, original geometry and global pose, the module balances the contribution of each type of feature through an adaptive fusion mechanism, avoids the limitations of single feature expression, and finally outputs point cloud fusion features with more comprehensive information and stronger representation ability.

[0055] Finally, the skeleton regression module performs human skeleton joint coordinate mapping. The module takes the point cloud fusion features output by the multi-level feature fusion module as input, and through nonlinear mapping and regression calculation, the high-dimensional fusion features are converted into three-dimensional coordinates of human skeleton joints. The role of this module is to establish the mapping relationship between the fusion features and the spatial position of the human joints, learn the corresponding rules of the features and the joint coordinates through network training, and finally output the skeleton estimation result that can accurately reflect the human pose.

[0056] In the embodiment of the present application, the point cloud block topology modeling module comprises: a geometric perception soft clustering unit, configured to classify the points in the human body topology point cloud block to each cluster in a soft assignment manner, and aggregate to obtain cluster-level features; a cluster-level spatial geometric perception attention unit, configured to enhance the cluster-level features through an attention mechanism based on the spatial geometric relationship between clusters; and a feature diffusion unit, configured to diffuse the enhanced cluster-level features back to the point level based on the spatial proximity of points and clusters, to generate point-level features with enhanced topology semantics.

[0057] Specifically, due to the characteristics of strong sparsity, many noise points and unstable human scattering structure of the millimeter wave radar point cloud, the human body cannot always be reliably divided into a fixed number of local topology regions with clear boundaries in a single frame of point cloud. If hard assignment is used to cluster the point cloud, it is easy to cause local structure breakage or misassignment, thereby affecting the stability of subsequent feature modeling.

[0058] Therefore, the present application proposes a geometric perception soft clustering method. The method does not make discrete decisions on the ownership relationship between point cloud samples and clusters, but instead introduces a learnable adaptive parameter σ, assigns a continuous soft ownership weight to each point based on the spatial geometric distance, to achieve flexible modeling of the local structure of the human body.

[0059] In detail, the method comprises: (1) Initial input The initial input consists of three parts: a point-level feature matrix extracted based on multi-scale hollow convolution , a corresponding set of three-dimensional spatial coordinates , and a point-level initial cluster number obtained by process 2, which is used for subsequent topology structure modeling and feature propagation.

[0060] (2) Cluster center initialization According to the initial coarse assignment cluster number, the cluster center is initialized, and the geometric center of the cth cluster is obtained: , wherein, is an indicator function, which is used to represent whether the nth point is assigned to the cth point cloud cluster. When the point n belongs to the cluster c, the value is 1, otherwise it is 0.

[0061] (3) Distance-based soft assignment weight calculation Avoid the instability brought by hard assignment, introduce a learnable scale parameter Calculate soft assignment weight according to the spatial distance between points and cluster centers: , where, is the standard deviation dynamically predicted by the geometric features in each cluster through a small MLP stacked with several layers, introducing an adaptive It can make the model flexibly adjust the strictness of clustering according to the actual distribution of point cloud, thereby enhancing the robustness to noise and non-rigid deformation.

[0062] (4) Cluster-level feature aggregation: Based on the soft assignment weight, the point-level features are weighted and aggregated to construct the cluster-level feature representation: , And update the cluster center simultaneously: .

[0063] Secondly, the cluster-level spatial geometry-aware attention unit performs cluster-level feature enhancement, which takes the cluster-level features output by the geometry-aware soft clustering unit and the updated cluster center coordinates as input. The core is to mine the human topological relationship between clusters. In detail, it includes: (1) Geometric relationship modeling between clusters For any pair of clusters , based on the spatial distance between clusters, introduce a geometric affinity weight to characterize the spatial proximity relationship between different topological parts of the human body: , where, is used to map the Euclidean distance between clusters to the geometric affinity weight in the interval [0, 1], and clusters with closer distances have greater affinity weights.

[0064] (2) Introduce geometry prior attention weight construction Under the multi-head attention mechanism, in the hth attention head, first perform linear mapping on the feature vector of the first cluster to obtain the transformed cluster feature: , Then, the source cluster feature, the target cluster feature and their spatial direction information are spliced to construct a joint representation for attention modeling: , Based on the above joint representation, the hth cluster obtains the attention weight for the first cluster through a learnable attention mapping function: Unnormalized geometric-enhanced attention weighting score of a cluster: , wherein, is the feature mapping matrix of the th attention head; is the learnable attention weight vector; geometric affinity weight As a subtractive modulation term, it suppresses the inter-cluster relationship that does not conform to the human spatial structure. By explicitly introducing a spatial distance-based geometric prior constraint in the attention scoring stage, the error information transmission between distant and weakly associated clusters can be effectively reduced, and the stable modeling ability of the local topology of the human body can be enhanced.

[0065] On this basis, the attention score is normalized to obtain the attention weight under the th attention head: , (3) Cluster-level feature update based on multi-head attention Based on the attention weight, the adjacent cluster features are weighted aggregated to obtain the output of the th attention head: , wherein C represents the C topological point cloud clusters of the human body. Then, the outputs of the attention heads are spliced and mapped back to the original feature dimension to obtain the cluster-level topology-enhanced feature: , wherein, is the output projection matrix, and LN represents the layer normalization operation.

[0066] Further, to avoid the loss of point-level details caused by cluster-level modeling, the present application further reversely diffuses the cluster-level semantic information to the point-level space. This mechanism takes the spatial distance between the point and the geometric center of each cluster as the basis, and measures the influence degree of different clusters on the point through soft weight, so as to realize the weighted fusion of multi-cluster information. In detail, it includes: (1) Point-cluster spatial weight calculation: For any point and cluster , the diffusion weight thereof is calculated according to the Euclidean distance in three-dimensional space, and is defined as: , wherein, represents the spatial coordinates of the nth point, represents the geometric center of the th cluster; is a scale control parameter, which is the same as the The parameters are kept consistent to ensure the consistency and stability of spatial weight calculation.

[0067] (2) Cluster-level features diffuse to point-level features Based on the above weights, the cluster-level features are diffused to the point-level space to obtain the enhanced point features: , in, Indicates the first Aggregation features of enhanced topological clusters The total number of clusters. Through the cluster-point feature diffusion process described above, the point-level representation not only preserves local geometric information but also explicitly integrates global semantic constraints from different topological parts of the human body, thereby improving the ability to model sparse point clouds and complex human poses.

[0068] Furthermore, in this embodiment, to fully integrate feature information from different levels and semantic sources, this application proposes a multi-branch feature adaptive fusion method based on global awareness gating, applied to a multi-level feature fusion module. This method adaptively learns the relative importance of the three features by globally statistically analyzing their overall distribution characteristics, and accordingly completes weighted feature fusion. Details: In terms of point features, the method has three input branches: 1. Pointwise geometric features encoded by multi-scale dilated convolution; 2. Local semantic features containing human topology from the diffusion module; 3. Pose-encoded global context embedding features obtained through global pooling.

[0069] First, the global awareness gating weights are generated, and global average pooling is performed on the three features to construct a channel-level global description vector. Subsequently, through a gating mapping consisting of two layers of fully connected networks, and then... Normalization generates fusion weights for the three features: , in, This represents a lightweight mapping function connected by non-linear activation functions, used to characterize the relative importance of different feature branches in the current sample.

[0070] Then, based on the above fusion weights, the three point-level features are weighted and summed to obtain the fused point-level feature representation: , This fusion process achieves dynamic balance of multi-level features while maintaining the point-level spatial distribution. The fused point-level features are further enhanced through linear mapping, layer normalization, and nonlinear activation operations, thereby improving the stability of feature representation and nonlinear modeling capabilities.

[0071] In the embodiment of the present application, before the human body topology point cloud block is input to the human body skeleton estimation network for human body skeleton estimation, it further includes: training the original human body skeleton estimation network based on a joint loss function to obtain an optimized human body skeleton estimation network; wherein the joint loss function includes a first loss term and a second loss term, the first loss term is used to constrain the human joint coordinate error, and the second loss term is used to constrain the human body part coordinate error.

[0072] Specifically, a joint loss function is constructed. The joint loss function is composed of a first loss term and a second loss term, which respectively constrain the human joint coordinate error and the part coordinate error. The first loss term (human joint coordinate error constraint) focuses on the coordinate prediction accuracy difference between the key joints and the non-key joints of the human body, and gives higher weight to important joints (such as right shoulder, right wrist, left shoulder, left wrist, etc. which are active in motion and key to pose representation) and relatively lower weight to non-important joints. By calculating the Euclidean distance error between the predicted joint coordinates and the real joint coordinates, the accuracy of the joint-level position prediction is constrained to ensure the estimation accuracy of the key joints; the second loss term (human body part coordinate error constraint) considers that the absolute coordinates of the joints in the data set are variable and difficult to accurately constrain. The human body is divided into five parts: main trunk, left upper limb, right upper limb, left lower limb and right lower limb. The corresponding reference root node is selected for each part, the relative position error of each joint in the part and the root node is calculated, and the normalized weight is given to each part according to the difference in motion amplitude and pose stability of different parts, so as to accurately constrain the topology structure of each part of the human body and avoid local part pose distortion. In detail, as shown in Figure 6 First, in order to depict the contribution difference of different human joints in pose estimation, the application introduces a joint importance weighting mechanism in the joint-level error calculation, which gives higher weight to key joints, and normalizes the weight of all joints to ensure the stability of the overall error scale. The joint weighted MPJPE is defined as: , where a higher weight is assigned to important joints , represents the frame index, F is the total number of human skeleton joints, represents the set of important joints (right shoulder, right wrist, left shoulder and left wrist) and is the non-important joint. The predicted position and the real position of the important joint are represented as and , and the predicted position and the real position of the non-important joint are represented as and .

[0073] ​Secondly, due to the large variation of joint positions in the data set, it is difficult to accurately estimate the absolute joint position. The present application divides the human body into five parts: the main trunk, the left upper limb, the right upper limb, the left lower limb and the right lower limb, and selects the corresponding reference root node, and constructs the relative geometric constraint of the joint nodes in the part. The relative geometric loss of a single human body part is defined as: , wherein, represents the joint set of the part, represents the j-th joint of the part, represents the reference root joint of the part. Here, the Manhattan distance is used to calculate the distance. Considering the differences in pose stability and motion amplitude of different human body parts, the present application introduces a normalized part weight to the relative geometric loss of each part to achieve balanced constraints of the whole body structure and local dynamics. The multi-part weighted geometric consistency loss is defined as:

[0074] , , wherein, is the weight coefficient of the corresponding human body part, considering the fast change of the joint speed of the upper limb, the weight of the upper limb part can be set higher than that of the trunk and the lower limb.

[0075] Finally, the joint weighted regression error and the multi-part relative geometric constraint are jointly optimized, and the overall loss function is represented as: .

[0076] The human skeleton estimation network optimized by the joint loss function can not only accurately constrain the coordinate error of a single joint, but also guarantee the topological structure consistency of each part of the human body, effectively improving the modeling ability and estimation robustness of complex human poses, providing a reliable model foundation for subsequent accurate skeleton estimation of input human topological point cloud blocks.

[0077] Human skeleton estimation belongs to the regression task of pose parameters. In the model training stage, first, a synchronous data acquisition system of millimeter wave radar and visual sensor is constructed, and through time alignment and space calibration, consistent acquisition of multi-source sensing data is realized. Secondly, the human skeleton joint coordinates extracted by the visual sensor are used as the supervision label information, and four-dimensional millimeter wave radar multi-time scale point clouds are generated, and are divided into multiple human topological blocks according to the human structure relationship, and the corresponding structure and motion features are extracted as the input of the multi-level feature adaptive fusion human skeleton estimation network.

[0078] ​In the training process, the human skeleton joint coordinates are taken as the supervision signal of the pose regression task to guide the network to learn the mapping relationship between the radar point cloud features and the human skeleton structure, thereby improving the accuracy and stability of human skeleton estimation. After the training is completed, the human skeleton estimation model is deployed in the actual application system to realize real-time human skeleton estimation based on the millimeter wave radar point cloud data.

[0079] The application also provides a multi-level feature adaptive fusion human skeleton estimation network framework, as shown in Figure 7 The radar module is first subjected to preliminary processing to obtain a normalized human topology point cloud block from the echo signal, as shown in Figure 7 Then, the normalized human topology point cloud block is input into the multi-level feature adaptive fusion human skeleton estimation network framework for processing. Specifically, the normalized human topology point cloud block is input into a local feature extraction and topology structure modeling module, a geometric feature extraction module and a global feature extraction module.

[0080] The local feature extraction and topology structure modeling module is the core innovative part of the network, and is executed in three steps of geometric perception soft clustering, cluster-level spatial geometric perception attention modeling and cluster-point feature diffusion based on spatial proximity. The input of the geometric perception soft clustering unit is the point-level feature matrix (F) of the topology point cloud block, the three-dimensional coordinate set (P) and the initial cluster number. The soft assignment weight of the point and the cluster is calculated through the learnable parameter σ, the cluster-level features are aggregated and the cluster center is updated to avoid structure breakage caused by hard assignment. The cluster-level spatial geometric perception attention unit first obtains the geometric affinity weight based on the Euclidean distance between clusters through MLP mapping, then constructs a joint representation containing spatial direction information to generate a geometric constraint attention weight. The adjacent cluster features are aggregated through multi-head attention to output enhanced cluster-level topology features to strengthen the spatial correlation of each part of the human body. The cluster-point feature diffusion unit diffuses the enhanced cluster-level features to the point level. The diffusion weight is calculated based on the spatial proximity between the point and the cluster center, and the multi-cluster semantic information is fused to each point to generate point-level features with enhanced topology semantics, which not only retains the point-level details but also integrates the global topology constraint. For fine extraction of local geometric features of the point cloud, multi-scale dilated convolution is used to perform convolution operation on the original three-dimensional coordinates of the topology point cloud block to extract different local geometric features and output point-level geometric features, which makes up for the possible loss of point-level details in cluster-level learning and complements the enhanced topology semantic point-level features of the cluster-level path. The geometric feature extraction module preliminarily integrates the topology semantic features output by the cluster-level path and the local geometric features output by the point-level path, and unifies the feature dimensions through linear mapping and layer normalization. The global feature extraction module inputs the local features integrated by the cluster-level learning and the point-level learning. The local features are first mapped into global pose semantic vectors (such as overall human body inclination and left-right inclination trend) through a pose encoding module, and then the global context features are obtained through splicing and pooling operations.

[0081] The adaptive gating fusion module of multi-level features first performs global average pooling on the three types of features to construct a channel-level global description vector; then generates fusion weights through two fully connected networks, and after normalization, the three types of features are weighted and summed according to the weights to obtain the point cloud fusion features.

[0082] The loss function module constrains the loss function from the joint accuracy and the part structure, to ensure that the output skeleton conforms to the joint position accuracy and meets the human body topological structure consistency, and completely matches the demand for skeleton estimation accuracy and robustness in home monitoring and other scenes. The residual regression module inputs the fusion features into the fully connected layer of the residual structure, and through nonlinear mapping, the high-dimensional features are converted into the three-dimensional coordinates of the human skeleton joints, and the final human skeleton estimation result is output.

[0083] As can be seen from the above, the embodiment of the present application first performs target detection on the collected radar echo data, and generates a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result, then divides the target four-dimensional point cloud into multiple human topological point cloud blocks based on the similarity of the speed direction of each point in the target four-dimensional point cloud, and finally inputs the human topological point cloud block into the human skeleton estimation network for human skeleton estimation to obtain the human skeleton estimation result. The present application proposes to construct human topological point cloud blocks based on the similarity of the speed direction of the point cloud, proposes an implicit human topological prior modeling method for millimeter wave radar, realizes the extraction of structured point cloud block features from sparse point clouds, and designs a multi-level feature adaptive fusion skeleton estimation framework, proposes dynamic modeling of the topological structure of the point cloud block, and realizes full-scale, multi-level fusion from fine-grained local joints to overall posture topology, significantly improving the stability and robustness of posture prediction. And compared with the human skeleton estimation method of other sensors, the human skeleton estimation method proposed in the present application has the characteristics of privacy-friendly, and is not affected by light and weather conditions, and can realize stable operation all day long, and is suitable for places with high requirements for safety and privacy; in addition, the method can construct human topological priors from sparse millimeter wave radar point clouds, extract structure-aware point cloud block features, and enhance the modeling ability of the model for the spatial relationship of key parts of the human body. Further, compared with the existing method which relies more on a single behavior type and a limited data set, the present method can adapt to diversified and natural daily human behaviors in unobtrusive monitoring scenarios, realize human skeleton estimation with more generalization ability, and thus effectively improve the accuracy and robustness of human posture prediction.

[0084] It should be understood that the size of the serial number of each step in the embodiment does not mean the order of execution of the steps, and the execution order of each step should be determined according to its function and inherent logic, and should not constitute the only limitation on the implementation process of the embodiment of the present application.

[0085] In summary, the detailed process of the human skeleton estimation method in the embodiment of the present application can be referred to in the followingFigure 8 , specifically: Step 801, performing fast Fourier transform processing on the collected single-frame radar echo data to obtain short-time scale point cloud information; Step 802, performing non-uniform Fourier transform processing on the collected multi-frame radar echo data to obtain long-time scale point cloud information; Step 803, performing target detection on the short-time scale point cloud information and the long-time scale point cloud information based on a constant false alarm rate algorithm to obtain a target detection result; Step 804, extracting a range-Doppler index of an effective target point; wherein the range-Doppler index includes radial distance information and radial velocity information of the effective target point; Step 805, performing beamforming on the short-time scale point cloud information and the long-time scale point cloud information corresponding to the range-Doppler index to extract an angle index of the effective target point; Step 806, establishing a spatial coordinate system based on a radar installation mode, and obtaining a four-dimensional coordinate of the effective target point based on the range-Doppler index and the angle index to generate a target four-dimensional point cloud including distance information, velocity information, and angle information; Step 807, performing unit normalization processing on the radial velocity information of the effective target point to obtain a normalized velocity vector; Step 808, constructing a distance metric based on the cosine similarity of each normalized velocity vector, and performing clustering processing on the effective target point according to the distance metric to obtain a clustering processing result; Step 809, dividing the effective target point based on the clustering processing result to obtain multiple human body topological point cloud blocks; Step 810, inputting the human body topological point cloud block into a human body skeleton estimation network for human body skeleton estimation to obtain a human body skeleton estimation result.

[0086] For more detailed processes of each of steps 801-810, please refer to the related part of the foregoing description, which will not be repeated here.

[0087] Please refer to Figure 9 , Figure 9 The radar-based human body skeleton estimation system provided by the embodiments of the present application. The system can be used to implement the human body skeleton estimation method related by the embodiments of the present application. The radar-based human body skeleton estimation system mainly includes: The point cloud generation module 901 is configured to perform target detection on the collected radar echo data, and generate a target four-dimensional point cloud including distance information, velocity information, and angle information based on the target detection result; The point cloud division module 902 is configured to divide the target four-dimensional point cloud into multiple human body topological point cloud blocks based on the velocity direction similarity of each point in the target four-dimensional point cloud. The skeleton estimation module 903 is configured to input the human body topological point cloud block to a human body skeleton estimation network for human body skeleton estimation, to obtain a human body skeleton estimation result.

[0088] In some embodiments of the present embodiment, when the point cloud generation module 901 performs target detection on the collected radar echo data and generates a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result, the point cloud generation module 901 is configured to: perform fast Fourier transform processing on the collected single-frame radar echo data to obtain short-time-scale point cloud information; perform non-uniform Fourier transform processing on the collected multi-frame radar echo data to obtain long-time-scale point cloud information; perform target detection based on the short-time-scale point cloud information and the long-time-scale point cloud information, and generate a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result.

[0089] Further, in some embodiments of the present embodiment, when the point cloud generation module 901 performs target detection based on the short-time-scale point cloud information and the long-time-scale point cloud information, and generates a target four-dimensional point cloud containing distance information, speed information and angle information based on the target detection result, the point cloud generation module 901 is specifically configured to: perform target detection on the short-time-scale point cloud information and the long-time-scale point cloud information based on a constant false alarm rate algorithm to obtain a target detection result; wherein the target detection result includes effective target points; extract distance-Doppler indexes of the effective target points; wherein the distance-Doppler indexes include radial distance information and radial speed information of the effective target points; perform beamforming on the short-time-scale point cloud information and the long-time-scale point cloud information corresponding to the distance-Doppler indexes to extract angle indexes of the effective target points; wherein the angle indexes include azimuth angle information and pitch angle information of the effective target points; establish a spatial coordinate system based on a radar installation mode, and obtain four-dimensional coordinates of the effective target points based on the distance-Doppler indexes and the angle indexes to generate a target four-dimensional point cloud containing distance information, speed information and angle information.

[0090] In some embodiments of the present embodiment, when the point cloud division module 902 divides the target four-dimensional point cloud into multiple human body topological point cloud blocks based on the speed direction similarity of each point in the target four-dimensional point cloud, the point cloud division module 902 is configured to: perform unit normalization processing on the radial speed information of the effective target points to obtain normalized speed vectors; construct a distance metric based on the cosine similarity of each normalized speed vector, and perform clustering processing on the effective target points according to the distance metric to obtain a clustering processing result; divide the effective target points based on the clustering processing result to obtain the multiple human body topological point cloud blocks.

[0091] In some embodiments of the present embodiment, the human skeleton estimation network in the human skeleton estimation system comprises, connected in sequence, a point cloud block topology modeling module, a multi-level feature fusion module, and a skeleton regression module: the point cloud block topology modeling module is configured to model the input human topology point cloud block in a dynamic topology structure, and output point-level features with enhanced topology semantics; the multi-level feature fusion module is configured to adaptively fuse the point-level features with enhanced topology semantics, original point cloud geometric features, and global context features, and output point cloud fusion features; and the skeleton regression module is configured to map the point cloud fusion features to human skeleton joint coordinates, to obtain a human skeleton estimation result.

[0092] In some embodiments of the present embodiment, the point cloud block topology modeling module in the human skeleton estimation system comprises: a geometric perception soft clustering unit, configured to classify points in the human topology point cloud block in a soft assignment manner to each cluster, and aggregate to obtain cluster-level features; a cluster-level spatial geometry perception attention unit, configured to enhance the cluster-level features through an attention mechanism based on spatial geometric relationships between clusters; and a feature diffusion unit, configured to diffuse the enhanced cluster-level features back to the point level based on the spatial proximity of points and clusters, to generate point-level features with enhanced topology semantics.

[0093] In some embodiments of the present embodiment, before the skeleton estimation module 903 performs the function of inputting the human topology point cloud block to the human skeleton estimation network for human skeleton estimation, the skeleton estimation module 903 is further configured to train the original human skeleton estimation network based on a joint loss function to obtain an optimized human skeleton estimation network; wherein the joint loss function comprises a first loss term and a second loss term, the first loss term is used to constrain human joint coordinate errors, and the second loss term is used to constrain human part coordinate errors.

[0094] In detail, each module in the radar-based human skeleton estimation system provided by the embodiments of the present application uses the same technical means as the human skeleton estimation method in the above Figure 1 , and can produce the same technical effects, which will not be described here.

[0095] Please refer to Figure 10 , Figure 10 for the module block diagram of the radar provided by the embodiments of the present application.

[0096] As Figure 10As shown in the figure, the embodiment of the present application further provides a radar, which can be used to implement the human skeleton estimation method in the foregoing embodiment, and the radar comprises a memory 1001, at least one processor 1002, a signal generator 1003 and a signal receiver 1004; the memory 1001 is used to store at least one program, and when the at least one program is executed by the at least one processor 1002, the at least one processor 1002 executes the human skeleton estimation method provided by the embodiment of the present application.

[0097] As shown in the figure, the embodiment of the present application further provides a radar, which can be used to implement the human skeleton estimation method in the foregoing embodiment, and the radar comprises a memory 1001, at least one processor 1002, a signal generator 1003 and a signal receiver 1004; the memory 1001 is used to store at least one program, and when the at least one program is executed by the at least one processor 1002, the at least one processor 1002 executes the human skeleton estimation method provided by the embodiment of the present application. Figure 11 , Figure 11 A module block diagram of the computer readable storage medium provided by the embodiment of the present application is shown in the figure.

[0098] As Figure 11 shown in the figure, the embodiment of the present application further provides a computer readable storage medium 1100, and the computer readable storage medium 1100 stores executable instructions 1110, and the executable instructions 1110 are executed to execute the human skeleton estimation method provided by the embodiment of the present application.

[0099] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), an internal memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of storage medium known in the technical field.

[0100] In the embodiments described above, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as Digital Video Disk, DVD), or semiconductor media (such as Solid State Disk) and the like.

[0101] It should be noted that each of the embodiments in the present application is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between each embodiment can be referred to each other. For product class embodiments, since they are similar to method class embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the description of the method class embodiment.

[0102] It should also be noted that in the present application, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or equipment including the element.

[0103] The above description of disclosed embodiments allows a skilled person to implement or use the present content. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined can be applied to other embodiments without departing from the spirit or scope of the content. Thus, the content is not to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed.

Claims

1. A radar-based method for estimating human skeletons, characterized in that, include: Target detection is performed on the collected radar echo data, and a four-dimensional point cloud of the target containing distance, velocity and angle information is generated based on the target detection results. Based on the similarity of the velocity directions of each point in the target four-dimensional point cloud, the target four-dimensional point cloud is divided into multiple human body topology point cloud blocks; The human body topology point cloud blocks are input into the human skeleton estimation network to perform human skeleton estimation and obtain the human skeleton estimation result.

2. The human skeleton estimation method according to claim 1, characterized in that, The process of detecting targets from the acquired radar echo data and generating a four-dimensional point cloud of the target containing range, velocity, and angle information based on the target detection results includes: Fast Fourier Transform is performed on the collected single-frame radar echo data to obtain short-timescale point cloud information. Non-uniform Fourier transform processing is performed on the collected multi-frame radar echo data to obtain long-term point cloud information; Target detection is performed based on the short-timescale point cloud information and the long-timescale point cloud information, and a four-dimensional point cloud of the target containing distance information, velocity information and angle information is generated based on the target detection results.

3. The human skeleton estimation method according to claim 2, characterized in that, The step of performing target detection based on the short-timescale point cloud information and the long-timescale point cloud information, and generating a four-dimensional point cloud of the target containing distance, velocity, and angle information based on the target detection results, includes: Target detection is performed on the short-timescale point cloud information and the long-timescale point cloud information based on the constant false alarm rate algorithm to obtain target detection results; wherein, the target detection results include valid target points; Extract the range-Doppler index of the effective target point; wherein, the range-Doppler index includes the radial distance information and radial velocity information of the effective target point; Beamforming is performed based on the short-timescale point cloud information and the long-timescale point cloud information corresponding to the range-Doppler index to extract the angle index of the effective target point; wherein, the angle index includes the azimuth and elevation angle information of the effective target point; A spatial coordinate system is established based on the radar installation method, and the four-dimensional coordinates of the effective target point are obtained based on the range-Doppler index and the angle index, generating a target four-dimensional point cloud containing range information, velocity information and angle information.

4. The human skeleton estimation method according to claim 3, characterized in that, Based on the similarity of velocity directions of each point in the target four-dimensional point cloud, the target four-dimensional point cloud is divided into multiple human body topological point cloud blocks, including: The radial velocity information of the effective target points is normalized to obtain a normalized velocity vector. A distance metric is constructed based on the cosine similarity of each normalized velocity vector, and the effective target points are clustered according to the distance metric to obtain the clustering result. Based on the clustering results, the effective target points are divided to obtain multiple human body topological point cloud blocks.

5. The human skeleton estimation method according to claim 1, characterized in that, The human skeleton estimation network includes a point cloud block topology modeling module, a multi-level feature fusion module, and a skeleton regression module connected in sequence. The point cloud block topology modeling module is used to perform dynamic topology modeling on the input human body topology point cloud block and output point-level features with enhanced topological semantics. The multi-level feature fusion module is used to adaptively fuse the point-level features with enhanced topological semantics, the original point cloud geometric features, and the global context features to output point cloud fusion features. The skeleton regression module is used to map the point cloud fusion features into human skeleton joint coordinates to obtain human skeleton estimation results.

6. The human skeleton estimation method according to claim 5, characterized in that, The point cloud block topology modeling module includes: The geometric perception soft clustering unit is used to classify the points in the human body topological point cloud block into clusters in a soft allocation manner and aggregate them to obtain cluster-level features; A cluster-level spatial geometry perception attention unit is used to enhance the cluster-level features based on the spatial geometric relationships between clusters through an attention mechanism; The feature diffusion unit is used to diffuse the enhanced cluster-level features back to the point level based on the spatial proximity between the point and the cluster, thereby generating the point-level features with enhanced topological semantics.

7. The human skeleton estimation method according to claim 5, characterized in that, Before inputting the human body topology point cloud blocks into the human skeleton estimation network for human skeleton estimation, the method further includes: The original human skeleton estimation network is trained based on the joint loss function to obtain the optimized human skeleton estimation network. The joint loss function includes a first loss term and a second loss term. The first loss term is used to constrain the coordinate error of human joints, and the second loss term is used to constrain the coordinate error of various parts of the human body.

8. A radar-based human skeleton estimation system, characterized in that, include: The point cloud generation module is used to perform target detection on the collected radar echo data and generate a four-dimensional point cloud of the target containing distance, velocity and angle information based on the target detection results. The point cloud partitioning module is used to divide the target four-dimensional point cloud into multiple human body topology point cloud blocks based on the similarity of the velocity directions of each point in the target four-dimensional point cloud. The skeleton estimation module is used to input the human body topology point cloud blocks into the human body skeleton estimation network to perform human body skeleton estimation and obtain the human body skeleton estimation result.

9. A radar, characterized in that, Includes memory and processor, of which: The processor is used to execute computer programs stored in the memory; When the processor executes the computer program, it implements the steps in the human skeleton estimation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the human skeleton estimation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-target identity recognition and behavior monitoring method based on millimeter wave radar

    CN118334736A

  • Millimeter wave radar motion evaluation method, system and equipment based on human skeleton model

    CN120392055A

  • Skeleton identification method and device, equipment, storage medium and computer program product

    CN121214479A

  • Fall detection method and system based on millimeter wave radar fused with human body posture

    CN121489438A

  • Millimeter wave radar fine-grained human body pose sensing method based on multi-dimensional feature extraction

    CN121541194A

Cited By

  • Knee bending angle measurement method, electronic equipment and storage medium

    CN122030954A

  • A knee flexion angle measurement method, electronic device and storage medium

    CN122030954B