Lightweight AI Personnel Perception Method and System Based on Millimeter-Wave Radar
By performing phase compensation and amplitude calibration on the millimeter-wave radar echo signal, combined with bidirectional joint filtering and adaptive point cloud feature extraction, a lightweight personnel perception method is realized, solving the problems of high computing complexity and high error detection rate in the prior art. It is suitable for resource-constrained devices, improving detection accuracy and real-timeness.
Patent Information
- Application Number
- CN202510672610.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing personnel perception methods based on millimeter wave radar are large in computing and high in power consumption, making them difficult to adapt to dynamic changes in different scenarios, and have high error detection rates and high computational complexity in low signal-to-noise ratio environments, which are difficult to meet real-time requirements, especially in resource-constrained edge devices.
The millimeter wave echo signal is processed by phase compensation and amplitude calibration methods, and three-dimensional spatial point cloud data is constructed in combination with bidirectional joint filtering. By calculating the center position of the point cloud cluster and the adaptive sampling radius, the personnel position is identified using the dual threshold judgment module to generate the thermal map output perceptual results.
It reduces the computational complexity, improves the accuracy and real-timeness of personnel detection, enhances the adaptability and robustness in complex environments, is suitable for embedded devices with limited resources, and supports smart home and security monitoring scenarios.
Smart Images

Figure CN120195652B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to personnel perception technology, and particularly to a lightweight AI personnel perception method and system based on millimeter-wave radar. Background Art
[0002] Existing personnel perception methods based on millimeter-wave radar mainly rely on large-scale neural networks for point cloud processing and target recognition. Although this method ensures a certain level of accuracy, it has a large computational load and high device power consumption, making it unsuitable for edge devices or embedded systems with limited computing resources. In addition, traditional point cloud processing methods are vulnerable to environmental noise interference in complex scenarios, resulting in insufficient robustness in personnel recognition.
[0003] Current millimeter-wave radar personnel perception methods usually use fixed thresholds for target detection. However, fixed thresholds are difficult to adapt to the dynamic changes of different scenarios, resulting in a high false detection rate in low signal-to-noise ratio environments. In addition, existing technologies lack effective adaptive processing methods for point cloud data, leading to a decrease in the detection accuracy of non-rigid targets (such as walking or posture-changing personnel). At the same time, in terms of personnel position tracking, existing methods mainly rely on multi-frame fusion or deep learning modeling, but these methods have a high computational complexity and are difficult to meet the real-time requirements, restricting the deployment of millimeter-wave radar in edge computing devices.
[0004] Therefore, there is an urgent need for a lightweight AI personnel perception method and system based on millimeter-wave radar to improve personnel detection accuracy, reduce the computational burden, and be applicable to resource-constrained terminal devices. Summary of the Invention
[0005] Embodiments of the present invention provide a lightweight AI personnel perception method and system based on millimeter-wave radar, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] A lightweight AI personnel perception method based on millimeter-wave radar is provided, including:
[0008] Collecting millimeter-wave echo signals in a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals in both the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data based on the filtered echo data;
[0009] Calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix based on the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as the reference, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and pose parameters based on the extracted point cloud features;
[0010] Input the motion vector and pose parameters into the dual-threshold decision module, calculate the speed threshold value and pose threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there are people in the target area, generate a personnel position heat map based on the center position of the point cloud cluster, and output the personnel perception result based on the personnel position heat map.
[0011] In an alternative embodiment,
[0012] Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibrated signal, perform two-way joint filtering processing on the calibrated signal in the distance dimension and angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data, including:
[0013] Collect the millimeter-wave echo signal of the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signal, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signal with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain a calibrated signal;
[0014] Perform wavelet packet decomposition on the calibrated signal to obtain multi-scale signal components, extract the energy entropy feature vector of the multi-scale signal components, calculate the adaptive reconstruction weight of each scale signal component based on the energy entropy feature vector, multiply the adaptive reconstruction weight by the corresponding scale signal component and superimpose them to obtain an enhanced signal;
[0015] Extract local maximum points as peak points on the time-frequency energy spectrum of the enhanced signal, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, use the threshold surface to extract the distance-dimensional target echo, construct an orthogonal projection operator with the covariance matrix eigenvector of the distance-dimensional target echo, and project the calibrated signal onto the null space of the orthogonal projection operator to obtain the angle-dimensional target echo;
[0016] Feature extraction is performed on the range-dimensional target echo and the angle-dimensional target echo to obtain a multi-level feature map. The multi-level feature map is input into an attention fusion network to generate a target scattering probability map. Filtered echo data is extracted according to the maximum value position of the target scattering probability map;
[0017] The filtered echo data is mapped to an initial point cloud through polar coordinates. The local density gradient of the initial point cloud is calculated. A diffusion equation of a point cloud growth model is established according to the local density gradient, and the diffusion equation is solved to obtain an enhanced point cloud;
[0018] The principal curvature and normal vector of the enhanced point cloud are calculated, and a spectral clustering is performed using the principal curvature and normal vector to construct an affinity matrix, obtaining three-dimensional spatial point cloud data.
[0019] In an alternative embodiment,
[0020] Feature extraction is performed on the range-dimensional target echo and the angle-dimensional target echo to obtain a multi-level feature map. The multi-level feature map is input into an attention fusion network to generate a target scattering probability map. Filtered echo data is extracted according to the maximum value position of the target scattering probability map, including:
[0021] The range-dimensional target echo is input into a phase-sensitive complex-valued filter bank to obtain a first feature layer. The conditional entropy of the first feature layer is calculated to obtain the redundancy between feature layers. The information gain rate of the first feature layer is calculated to obtain the feature discrimination degree. Based on the redundancy between feature layers and the feature discrimination degree, the time-scale parameter and phase parameter of the phase-sensitive complex-valued filter bank are optimized to obtain a first multi-level feature map;
[0022] The angle-dimensional target echo is input into a direction-sensitive filter bank. The local region direction consistency of the angle-dimensional target echo is calculated to obtain a direction response map. The main direction and secondary direction are extracted based on the direction response map. The main direction and secondary direction are used as the reference directions of the direction-sensitive filter bank. The feature response on the main direction and secondary direction is calculated to obtain the feature response difference. Based on the feature response difference, the angle resolution parameter of the direction-sensitive filter bank is optimized to obtain a second multi-level feature map;
[0023] The mutual information matrix is calculated for the first multi-level feature map and the second multi-level feature map. The cross-attention weight is constructed based on the principal component vector of the mutual information matrix. The channel attention weight is obtained based on the channel statistics of the first multi-level feature map and the second multi-level feature map. The cross-attention weight and the channel attention weight are combined to obtain a fusion attention weight;
[0024] Calculate the similarity pattern of local features based on the fused attention weights to obtain the feature sampling offset. Resample the first multi-level feature map and the second multi-level feature map according to the feature sampling offset to obtain resampled features. Calculate the local density distribution matrix and the density jump matrix of the resampled features, and generate a target scattering probability map based on the local density distribution matrix and the density jump matrix;
[0025] Determine the scattering center position coordinates based on the maximum position of the target scattering probability map, construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, and perform a convolution operation on the adaptive kernel function with the target echo in the distance dimension and the target echo in the angle dimension to obtain the filtered echo data.
[0026] In an alternative embodiment,
[0027] Calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and the attitude parameters based on the extracted point cloud features, including:
[0028] Construct a radial density distribution matrix of the three-dimensional spatial point cloud data, calculate the adaptive search radius of each point based on the gradient change of the radial density distribution matrix, extract adjacent point sets within the adaptive search radius, and calculate the distance difference and the angle difference between point pairs in the adjacent point sets;
[0029] Construct a two-layer weight modulation function. The first-layer weight coefficient of the two-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Input the distance difference and the angle difference into the two-layer weight modulation function to obtain weight features, and construct a spatial distribution matrix based on the weight features;
[0030] Perform eigen-decomposition on the spatial distribution matrix to obtain the main eigen-subspace and the secondary eigen-subspace, map the main eigen-subspace and the secondary eigen-subspace to a topological feature network, calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and use the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters;
[0031] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculate the density difference between any pair of points within the point cloud clustering cluster, determine the density increasing direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum. Determine the position of the local density maximum as the center position of the point cloud cluster;
[0032] Construct a deformation kernel function with the center position of the point cloud cluster as the reference point, calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameter of the kernel function, adjust the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, construct a feature distribution tensor within the range of the adaptive sampling radius, extract the main feature components of the feature distribution tensor, and combine the main feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud;
[0033] Construct a feature propagation network. Input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain the motion vectors of the feature points, and perform manifold constraint optimization on the motion vectors of the feature points to obtain the pose parameters of the target object.
[0034] In an alternative embodiment,
[0035] Construct a feature propagation network. Input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain the motion vectors of the feature points, and performing manifold constraint optimization on the motion vectors of the feature points to obtain the pose parameters of the target object includes:
[0036] Obtain the dynamic features of the point cloud in consecutive time frames, extract the local shape descriptors and global descriptors representing the regional distribution from the dynamic features of the point cloud; perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features; divide the enhanced dynamic features into the current frame feature set and the subsequent frame feature set according to the time frames;
[0037] Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distances and feature similarities between feature points at each sampling level; perform dynamic grouping on the feature points based on the geometric distances and feature similarities to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure;
[0038] Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal variation; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fusion feature points;
[0039] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fusion feature points;
[0040] Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints; combine the temporal constraint and the cross-scale feature point matching constraints into an objective function;
[0041] Optimize the objective function sequentially at multiple sampling levels to obtain the feature point matching relationship; based on the feature point matching relationship, pair the three-dimensional space coordinates of the matching feature points under their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector;
[0042] Construct a manifold space representation of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and take the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and take the convergence result as the pose parameter of the target object.
[0043] In an alternative embodiment,
[0044] Input the motion vector and the pose parameter into a dual-threshold decision module, calculate the speed threshold value and the pose threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there are people in the target area, and a personnel position heat map is generated based on the center position of the point cloud cluster. The personnel perception result output based on the personnel position heat map includes:
[0045] Sample the motion vector within a sliding time window of a specified length, and construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vector according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a temporal change model for the pose parameter, extract the mean and standard deviation of the pose parameter based on the temporal change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value;
[0046] Compare the motion vector with a velocity threshold value, and compare the attitude parameter with an attitude threshold value; when the motion vector is greater than the velocity threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area; use the region growing algorithm to perform point cloud segmentation on the target area to obtain a target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain a filtered point cloud cluster;
[0047] Calculate the centroid of the filtered point cloud cluster to obtain the center position coordinates of the point cloud cluster, and map the center position coordinates of the point cloud cluster to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinates; use a Gaussian kernel function with an adaptive bandwidth to perform kernel density estimation on the two-dimensional ground plane coordinates to obtain a spatial density distribution; calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud cluster, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain an initial heat map;
[0048] Construct a recursive filter to perform temporal filtering on the initial heat map to obtain a filtered heat map; calculate the fusion weight based on the similarity between the filtered heat map and the heat map of the previous moment; apply the fusion weight to the filtered heat map and the heat map of the previous moment to obtain a heat map of the personnel position; apply the local maximum suppression algorithm to the heat map of the personnel position to extract the extreme points of the heat map; use the extreme points of the heat map as candidate points for the personnel position;
[0049] Perform inter-frame matching on the candidate points for the personnel position to obtain an initial personnel trajectory, and use a Kalman filter to smooth and predict the initial personnel trajectory to obtain a personnel motion trajectory; construct a multi-feature fusion model by combining the motion vector, attitude parameter, and historical information of the personnel motion trajectory; use the multi-feature fusion model to calculate the confidence score of each candidate point for the personnel position; set a dynamic threshold to screen the confidence scores, and output the candidate points for the personnel position that meet the dynamic threshold and their corresponding personnel motion trajectories as the personnel perception result.
[0050] In an alternative embodiment,
[0051] Performing inter-frame matching on the candidate points for the personnel position to obtain an initial personnel trajectory, and using a Kalman filter to smooth and predict the initial personnel trajectory to obtain a personnel motion trajectory includes:
[0052] Obtain the candidate points for the personnel position in consecutive time frames; calculate the Euclidean distance between the candidate points for the personnel position in adjacent time frames to obtain a distance feature; calculate the motion speed based on the historical position information of the candidate points for the personnel position, and calculate the speed difference between the candidate points for the personnel position in adjacent time frames based on the motion speed to obtain a speed feature; calculate the motion direction based on the motion speed, and calculate the direction difference between the candidate points for the personnel position in adjacent time frames based on the motion direction to obtain a direction feature;
[0053] Construct the distance feature, speed feature, and direction feature into a matching cost matrix; use the candidate points of the personnel position in the current time frame as the target point set, and use the candidate points of the personnel position in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain matching point pairs; connect the matching point pairs in chronological order to obtain the initial personnel trajectory;
[0054] Form a state vector from the position coordinates, motion speed, and acceleration in the initial personnel trajectory; construct a state transition equation according to the state vector; construct an observation equation based on the position coordinates of the initial personnel trajectory; calculate the prediction error between the predicted value of the state transition equation and the actual observation value; update the process noise covariance according to the prediction error;
[0055] Predict the state vector using the state transition equation and process noise covariance to obtain the state prediction value and prediction covariance; calculate the Kalman gain according to the state prediction value, prediction covariance, and observation equation; correct the state prediction value using the Kalman gain to obtain the state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value;
[0056] Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract the historical motion pattern within the prediction window length; use the historical motion pattern to repair the trajectory break points;
[0057] Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than the preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain the optimized personnel motion trajectory.
[0058] In the second aspect of the embodiments of the present invention,
[0059] Provide a lightweight AI personnel perception system based on a millimeter-wave radar, including:
[0060] A first unit for collecting millimeter-wave echo signals in a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data according to the filtered echo data;
[0061] A second unit is configured to calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as a reference, extract continuously changing point cloud features within the adaptive sampling radius, and calculate a motion vector and pose parameters based on the extracted point cloud features;
[0062] A third unit is configured to input the motion vector and the pose parameters into a dual-threshold decision module, calculate a speed threshold value and a pose threshold value respectively through the dual-threshold decision module, determine that there is a person in the target area when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, generate a personnel position heat map according to the center position of the point cloud cluster, and output a personnel perception result based on the personnel position heat map.
[0063] In a third aspect of the embodiments of the present invention,
[0064] There is provided an electronic device, including:
[0065] A processor;
[0066] A memory for storing instructions executable by the processor;
[0067] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0068] In a fourth aspect of the embodiments of the present invention,
[0069] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0070] In this embodiment, by performing phase compensation and amplitude calibration on the millimeter-wave echo signal, the accuracy of the point cloud data is improved, making subsequent target detection more reliable. The bidirectional joint filtering method is adopted to optimize the signal in the range dimension and the angle dimension, reduce environmental noise interference, and effectively improve the stability of detection. At the same time, by constructing a spatial distribution matrix and an adaptive sampling radius, accurate identification of point cloud clusters is achieved, thereby improving the accuracy of dynamic personnel detection and avoiding the misjudgment problem caused by the fixed threshold method. It has lightweight computing capabilities and adopts an adaptive point cloud feature extraction method to reduce the computational complexity while ensuring the detection accuracy, and is suitable for embedded and edge computing devices. Based on the point cloud features, the motion vector and attitude parameters are calculated, and combined with the dual-threshold decision mechanism, efficient identification of the presence of personnel is achieved, which can effectively distinguish the static background from the real target and improve the reliability of detection. In addition, using the personnel position heat map, the intuitive visualization of the personnel distribution in the space can be realized, providing accurate perception capabilities for intelligent security, unmanned systems, and smart homes. While ensuring the detection accuracy, the computational burden is reduced, the adaptability to complex environments is enhanced, and the real-time performance is improved, providing strong support for the wide application of millimeter-wave radars in the field of intelligent perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 FIG. is a schematic flow chart of the lightweight AI personnel perception method based on millimeter-wave radar according to an embodiment of the present invention;
[0072] Figure 2 FIG. is a scatter plot of the feature extraction performance according to an embodiment of the present invention;
[0073] Figure 3 FIG. is a comparison effect diagram of the point cloud cluster segmentation based on the sequential filtering of the features according to an embodiment of the present invention;
[0074] Figure 4 FIG. is a schematic structural diagram of the lightweight AI personnel perception system based on millimeter-wave radar according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0076] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0077] Figure 1 This is a schematic flowchart of the lightweight AI personnel perception method based on millimeter-wave radar according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0078] Collect the millimeter-wave echo signals in the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, perform two-way joint filtering processing on the calibrated signals according to the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data based on the filtered echo data;
[0079] Calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix based on the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features;
[0080] Input the motion vector and attitude parameters into a dual-threshold decision module, calculate the speed threshold value and the attitude threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area, and a personnel position heat map is generated according to the center position of the point cloud cluster, and a personnel perception result is output based on the personnel position heat map.
[0081] In an alternative embodiment, collecting the millimeter-wave echo signals in the target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals according to the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data based on the filtered echo data includes:
[0082] Collect the millimeter-wave echo signals in the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signals, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signals with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain calibrated signals;
[0083] Perform wavelet packet decomposition on the calibrated signals to obtain multi-scale signal components, extract the energy entropy feature vectors of the multi-scale signal components, calculate the adaptive reconstruction weights of each scale signal component based on the energy entropy feature vectors, multiply the adaptive reconstruction weights by the corresponding scale signal components and superimpose them to obtain an enhanced signal;
[0084] Extract local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, extract range-dimension target echoes using the threshold surface, construct an orthogonal projection operator with the covariance matrix eigenvectors of the range-dimension target echoes, and project the calibration signal onto the null space of the orthogonal projection operator to obtain azimuth-dimension target echoes;
[0085] Perform feature extraction on the range-dimension target echoes and azimuth-dimension target echoes to obtain a multi-level feature map, input the multi-level feature map into an attention fusion network to generate a target scattering probability map, and extract filtered echo data based on the maximum position of the target scattering probability map;
[0086] Obtain an initial point cloud by performing polar coordinate mapping on the filtered echo data, calculate the local density gradient of the initial point cloud, establish a diffusion equation for the point cloud growth model based on the local density gradient, and solve the diffusion equation to obtain an enhanced point cloud;
[0087] Calculate the principal curvature and normal vector of the enhanced point cloud, construct an affinity matrix using the principal curvature and normal vector for spectral clustering, and obtain 3D spatial point cloud data.
[0088] A method for constructing 3D spatial point cloud data based on millimeter-wave radar echo signals, which is used to accurately perceive the 3D information of the target area. This method performs multi-level processing on millimeter-wave echo signals, including phase compensation and amplitude calibration, bidirectional joint filtering, as well as point cloud enhancement and clustering, and finally generates high-quality 3D point cloud data.
[0089] Exemplarily, first, collect the millimeter-wave echo signals of the target area. For example, use a 77GHz millimeter-wave radar to collect the echo signals of the target area with a sampling period of 10ms.
[0090] Then, perform phase compensation and amplitude calibration on the collected millimeter-wave echo signals. Calculate the phase differences between adjacent antennas. For example, calculate the phase difference between the first antenna and the second antenna, the phase difference between the second antenna and the third antenna, and so on. Construct a dynamic self-correction coefficient matrix with these phase differences. At the same time, calculate the amplitude ratios between adjacent antennas and map them into a non-linear compensation coefficient curve. For example, use the method of polynomial fitting to fit the amplitude ratio and the corresponding compensation coefficient into a curve. Perform matrix multiplication operations on the collected millimeter-wave echo signals with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain the calibrated signal. Assume that the millimeter-wave radar has 4 receiving antennas, then the dynamic self-correction coefficient matrix is a 4x4 matrix, and the non-linear compensation coefficient curve is a piecewise function.
[0091] Next, perform wavelet packet decomposition on the calibrated signal. For example, use the db4 wavelet for 4-layer decomposition to obtain signal components at 16 scales. Extract the energy entropy feature vectors of the signal components at each scale. For example, calculate the energy entropy values of the signal components at each scale and form a 16-dimensional feature vector. Calculate the adaptive reconstruction weights of the signal components at each scale based on the energy entropy feature vectors. For example, assign a weight to each scale according to the magnitude of the energy entropy value, and the larger the energy entropy value, the larger the weight. Multiply the adaptive reconstruction weights by the corresponding scale signal components and sum them up to obtain the enhanced signal.
[0092] Extract the local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points. For example, set a threshold and consider the points exceeding the threshold as peak points. Calculate the spatial clustering distribution characteristics of the peak points. For example, calculate the distances between the peak points and perform clustering based on the distances. Generate a threshold surface based on the spatial clustering distribution characteristics. For example, use the interpolation method to connect the clustering centers to form a surface. Use the threshold surface to extract the range-dimensional target echo. Construct an orthogonal projection operator with the eigenvectors of the covariance matrix of the range-dimensional target echo. Project the calibrated signal onto the null space of the orthogonal projection operator to obtain the angle-dimensional target echo.
[0093] Extract features from the range-dimensional target echo and the angle-dimensional target echo to obtain a multi-level feature map. For example, extract the peak features of the range-dimensional target echo and the energy distribution features of the angle-dimensional target echo. Input the multi-level feature map into the attention fusion network to generate the target scattering probability map. For example, use a two-layer fully connected network for fusion. Extract the filtered echo data according to the maximum position of the target scattering probability map.
[0094] Obtain the initial point cloud by polar coordinate mapping of the filtered echo data. For example, convert the range and angle information of the echo data into three-dimensional coordinates. Calculate the local density gradient of the initial point cloud. For example, calculate the number of points around each point and calculate its gradient. Establish the diffusion equation of the point cloud growth model based on the local density gradient and solve the diffusion equation to obtain the enhanced point cloud.
[0095] Calculate the principal curvature and normal vector of the enhanced point cloud. Use the principal curvature and normal vector to construct an affinity matrix for spectral clustering to obtain the final three-dimensional spatial point cloud data. For example, use the k-means algorithm for clustering to finally obtain the three-dimensional point cloud data of the target.
[0096] In this embodiment, through multi-level signal processing and point cloud enhancement techniques, the influence of noise and clutter is effectively suppressed, the accuracy and integrity of the point cloud are improved, and the constructed point cloud data is clearer and more accurate. By adopting an adaptive phase compensation and amplitude calibration method, it can effectively cope with signal changes in different environments, improving the robustness and anti-interference ability of the system. By combining techniques such as wavelet packet decomposition, attention mechanism, and spectral clustering, efficient processing of millimeter-wave radar echo signals and point cloud construction are realized, reducing the computational complexity and improving the processing efficiency.
[0097] In an alternative embodiment, feature extraction is performed on the range-dimension target echo and the angle-dimension target echo to obtain a multi-level feature map, and the multi-level feature map is input into an attention fusion network to generate a target scattering probability map. Extracting the filtered echo data according to the maximum value position of the target scattering probability map includes:
[0098] Input the range-dimension target echo into a phase-sensitive complex-valued filter bank to obtain a first feature layer, calculate the conditional entropy of the first feature layer to obtain the redundancy between feature layers, calculate the information gain rate of the first feature layer to obtain the feature discrimination degree, and optimize the time-scale parameter and phase parameter of the phase-sensitive complex-valued filter bank based on the redundancy between feature layers and the feature discrimination degree to obtain a first multi-level feature map;
[0099] Input the angle-dimension target echo into a direction-sensitive filter bank, calculate the local region direction consistency of the angle-dimension target echo to obtain a direction response map, extract the main direction and the secondary direction based on the direction response map, use the main direction and the secondary direction as the reference directions of the direction-sensitive filter bank, calculate the feature response differences in the main direction and the secondary direction, and optimize the angle resolution parameter of the direction-sensitive filter bank based on the feature response differences to obtain a second multi-level feature map;
[0100] Calculate the mutual information matrix for the first multi-level feature map and the second multi-level feature map, construct cross-attention weights based on the principal component vectors of the mutual information matrix, obtain channel attention weights based on the channel statistics of the first multi-level feature map and the second multi-level feature map, and combine the cross-attention weights and the channel attention weights to obtain fusion attention weights;
[0101] Calculate the similarity pattern of local features based on the fusion attention weights to obtain a feature sampling offset, resample the first multi-level feature map and the second multi-level feature map according to the feature sampling offset to obtain resampled features, calculate the local density distribution matrix and density jump matrix of the resampled features, and generate a target scattering probability map based on the local density distribution matrix and the density jump matrix;
[0102] Determine the scattering center position coordinates based on the maximum position of the target scattering probability map, construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, and perform a convolution operation on the adaptive kernel function with the range-dimensional target echo and the angle-dimensional target echo to obtain the filtered echo data.
[0103] Exemplarily, first obtain the range-dimensional echo and the angle-dimensional echo data of the target. Assume that the range-dimensional echo is a vector containing 128 sampling points, and the angle-dimensional echo is a 16x16 matrix, where each element represents the echo intensity at different angles.
[0104] Next, perform feature extraction on the range-dimensional target echo. Input the range-dimensional echo into a set of phase-sensitive complex-valued filters, such as using 8 Gabor filters with different time scales and phase parameters. Each filter performs a convolution operation on the echo to obtain 8 different feature maps, constituting the first feature layer. Calculate the conditional entropy of the first feature layer, such as using the Shannon entropy formula, to measure the redundancy between feature layers. At the same time, calculate the information gain rate of the first feature layer to evaluate the distinguishability of features. According to the conditional entropy and the information gain rate, adjust the time scale and phase parameters of the Gabor filters, such as using the gradient descent method to minimize the conditional entropy and maximize the information gain rate. Filter the range-dimensional echo again through the optimized filters to obtain the optimized first multi-level feature map. Assume that the optimized first multi-level feature map still contains 8 feature maps, and the size of each feature map is 128.
[0105] Then, perform feature extraction on the angle-dimensional target echo. Calculate the local region direction consistency of the angle-dimensional target echo, such as using the Sobel operator to calculate the gradient to obtain the direction response map. Assume that the size of the direction response map is 16x16. Extract the main direction and the secondary direction according to the direction response map, such as selecting the direction with the largest gradient magnitude as the main direction and the direction orthogonal to it as the secondary direction. Use the main direction and the secondary direction as the reference directions of the direction-sensitive filter bank, such as using 8 Gabor filters with different angle resolution parameters, where 4 filters are aligned with the main direction and the other 4 filters are aligned with the secondary direction. Calculate the feature responses in the main direction and the secondary direction, and calculate the difference between them. According to the feature response difference, such as using the gradient descent method to minimize the feature response difference, optimize the angle resolution parameters of the direction-sensitive filter bank. Filter the angle-dimensional echo with the optimized filters to obtain the second multi-level feature map. Assume that the second multi-level feature map contains 8 feature maps, and the size of each feature map is 16x16.
[0106] Next, feature fusion is performed. Calculate the mutual information matrix between the first multi-level feature map and the second multi-level feature map, for example, using a histogram-based mutual information estimation method. Perform principal component analysis on the mutual information matrix to obtain the principal component vectors, and use them as cross-attention weights. At the same time, calculate the channel statistics of the first multi-level feature map and the second multi-level feature map, such as the mean and standard deviation, and use them to construct channel attention weights, for example, using the Sigmoid function to normalize the statistics. Combine the cross-attention weights and the channel attention weights, for example, by weighted averaging, to obtain the fused attention weights.
[0107] Based on the fused attention weights, calculate the similarity pattern of the local features, for example, calculate the cosine similarity between the feature vectors to obtain the feature sampling offset. According to the feature sampling offset, resample the first multi-level feature map and the second multi-level feature map, for example, using bilinear interpolation. Calculate the local density distribution matrix and the density jump matrix of the resampled features, for example, using kernel density estimation and differential operations. Based on the local density distribution matrix and the density jump matrix, for example, multiply the two to generate the target scattering probability map.
[0108] Finally, echo filtering is performed. According to the maximum position of the target scattering probability map, determine the coordinates of the scattering center position. Construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, for example, using a Gaussian kernel function, whose parameters are determined by the eigenvalues at the scattering center position coordinates. Convolve the adaptive kernel function with the range-dimensional target echo and the angle-dimensional target echo to obtain the filtered echo data.
[0109] In the prior art, during the process of target echo processing, a single feature extraction method or a method with fixed filtering parameters is usually adopted, resulting in inaccurate characterization of the target scattering characteristics, being easily affected by noise interference, and affecting the accuracy of target recognition. In this application, by constructing a phase-sensitive complex-valued filter bank and a direction-sensitive filter bank, feature extraction is respectively performed on the target echoes in the range dimension and the angle dimension, and the filtering parameters are optimized based on conditional entropy, information gain rate, and direction consistency, making the feature extraction more accurate. Meanwhile, cross-attention and channel-attention weights are introduced to achieve deep fusion of multi-level feature mappings, fully utilize the information complementarity between different features, and improve the expression ability of target scattering features. During the generation process of the target scattering probability map, by calculating the local density distribution matrix and the density jump matrix, the distinguishability of the target scattering features is enhanced, making the target positioning more accurate. In addition, this application adopts an adaptive kernel function that dynamically adjusts according to the position of the scattering center to filter the echo data, enabling the filtering method to adapt to the changes in the target scattering characteristics, reducing the influence of noise, and improving the robustness of target detection. Compared with the prior art, this application can extract and fuse target scattering features more accurately, maintain a high detection accuracy in complex environments, and enhance the adaptability and robustness of the system.
[0110] In an optional implementation manner, calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as a reference, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and the attitude parameters based on the extracted point cloud features, including:
[0111] Construct a radial density distribution matrix of the three-dimensional spatial point cloud data, calculate the adaptive search radius of each point based on the gradient change of the radial density distribution matrix, extract adjacent point sets within the adaptive search radius, and calculate the distance difference and the angle difference between point pairs in the adjacent point sets;
[0112] Construct a two-layer weight modulation function, where the first-layer weight coefficient of the two-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Input the distance difference and the angle difference into the two-layer weight modulation function to obtain a weight feature, and construct a spatial distribution matrix based on the weight feature;
[0113] Perform eigen-decomposition on the spatial distribution matrix to obtain a principal eigen-subspace and a secondary eigen-subspace, map the principal eigen-subspace and the secondary eigen-subspace to a topological feature network, calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and use the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters;
[0114] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculate the density difference between any pair of points within the point cloud clustering cluster, determine the density increasing direction based on the density difference, and obtain the local density maximum position by iteratively updating the search position until convergence. Determine the local density maximum position as the point cloud cluster center position;
[0115] Construct a deformation kernel function with the point cloud cluster center position as the reference point, calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameter of the kernel function, adjust the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, construct a feature distribution tensor within the range of the adaptive sampling radius, extract the main feature components of the feature distribution tensor, and combine the main feature components with the multi-dimensional feature vector to obtain the point cloud dynamic feature;
[0116] Construct a feature propagation network, input the point cloud dynamic feature into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the feature point motion vector, and perform manifold constraint optimization on the feature point motion vector to obtain the pose parameters of the target object.
[0117] Exemplarily, first obtain the three-dimensional point cloud data and construct a radial density distribution matrix. Calculate the average distance from each point to other points within a certain range around it, and use this as the radial density value of the point. Store the radial density values of all points in a matrix to form the radial density distribution matrix. For example, in a point cloud data containing 1000 points, the average distance from each point to its 10 nearest neighbor points can be calculated, and these 1000 average distance values can be stored in a 1000x1 matrix, which is the radial density distribution matrix.
[0118] Calculate the adaptive search radius of each point according to the gradient change of the radial density distribution matrix. The greater the gradient change, the smaller the search radius, and vice versa. For example, if the radial density value of a certain point differs greatly from the radial density values of its surrounding points, the search radius of this point will be smaller; conversely, if the radial density value of this point differs little from the radial density values of its surrounding points, the search radius of this point will be larger.
[0119] Extract the adjacent point set within the adaptive search radius, and calculate the distance difference and angle difference between point pairs in the adjacent point set. Assume that the adaptive search radius of a certain point is 0.1, then the adjacent point set of this point is all points whose distance from it is less than 0.1. For any two points in the adjacent point set, calculate the Euclidean distance between them and the angle formed by the line connecting them to the center point to obtain the distance difference and angle difference.
[0120] Construct a two-layer weight modulation function. The weight coefficients of the first layer are determined by the change rate of the radial density distribution matrix. The greater the change rate, the greater the weight coefficient. The weight coefficients of the second layer are determined by the distribution consistency of the angular differences. The more consistent the distribution, the greater the weight coefficient. For example, if the points around a point are concentrated in one direction, the distribution consistency of the angular differences is relatively high, and the corresponding weight coefficient is also relatively large. Input the distance difference and the angular difference into the two-layer weight modulation function to obtain the weight features.
[0121] Construct a spatial distribution matrix based on the weight features. Store the weight features of all point pairs in a matrix to form the spatial distribution matrix. For example, in a point cloud data containing 1000 points, the dimension of the spatial distribution matrix is 1000x1000, where each element represents the weight feature between the corresponding two points.
[0122] Perform eigen-decomposition on the spatial distribution matrix to obtain the principal eigen-subspace and the secondary eigen-subspace. Map the principal eigen-subspace and the secondary eigen-subspace to the topological feature network.
[0123] Calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector. Use the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters.
[0124] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function. Calculate the density difference between any point pairs within the point cloud clustering cluster, determine the density rising direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum, and determine the position of the local density maximum as the center position of the point cloud cluster.
[0125] Taking the center position of the point cloud cluster as the reference point, construct a deformation kernel function. Calculate the local curvature distribution within the neighborhood of the reference point, and map the local curvature distribution to the deformation parameter of the kernel function. Adjust the shape of the kernel function according to the deformation parameter to obtain the adaptive sampling radius. Construct a feature distribution tensor within the range of the adaptive sampling radius, extract the principal feature components of the feature distribution tensor, and combine the principal feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud.
[0126] Construct a feature propagation network and input the dynamic features of the point cloud into the feature propagation network. Identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain the motion vectors of the feature points. Perform manifold constraint optimization on the motion vectors of the feature points to obtain the attitude parameters of the target object.
[0127] In this embodiment, through the adaptive search radius and the double-layer weight modulation function, the point cloud features can be extracted more accurately, thereby improving the accuracy of subsequent motion estimation and attitude parameter calculation. By using the density gradient optimization function to determine the center position of the point cloud cluster, the influence of noise can be effectively avoided and the robustness of the algorithm can be enhanced. Through the feature propagation network and the manifold constraint optimization, the attitude parameters of the target object can be accurately estimated, and high accuracy can be maintained even when the target object performs complex motions.
[0128] In an alternative embodiment, a feature propagation network is constructed. The dynamic features of the point cloud are input into the feature propagation network to identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain the feature point motion vectors, and perform manifold constraint optimization on the feature point motion vectors to obtain the attitude parameters of the target object, including:
[0129] Obtain the dynamic features of the point cloud in consecutive time frames, and extract local shape descriptors and global descriptors representing the regional distribution from the dynamic features of the point cloud; perform adaptive weighted fusion on the local shape descriptors and the global descriptors to obtain enhanced dynamic features; divide the enhanced dynamic features into a current frame feature set and a subsequent frame feature set according to the time frames;
[0130] Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distances and feature similarities between feature points at each sampling level; perform dynamic grouping on the feature points based on the geometric distances and feature similarities to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure;
[0131] Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal change; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fusion feature points;
[0132] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fusion feature points;
[0133] Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints; combine the temporal constraint and the cross-scale feature point matching constraints into an objective function;
[0134] Optimize the objective function successively at multiple sampling levels to obtain the feature point matching relationship; based on the feature point matching relationship, correspond and pair the three-dimensional space coordinates of the matching feature points under their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector.
[0135] Construct a manifold space expression of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and use the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and use the convergence result as the pose parameter of the target object.
[0136] A target pose estimation method based on a feature propagation network for identifying the target pose from dynamic point cloud data. This method constructs a feature propagation network to capture the spatio-temporal features of the point cloud and combines manifold constraint optimization to achieve accurate target pose estimation.
[0137] First, obtain the dynamic features of the point cloud in consecutive time frames. For example, in the first-frame point cloud, each point has three-dimensional coordinates (x, y, z) and a reflection intensity value. By calculating the geometric relationship and feature differences between each point and its neighboring points, extract local shape descriptors, such as local curvature and normal vector. At the same time, statistically analyze the overall distribution of the point cloud and extract global descriptors, such as the centroid and bounding box size of the point cloud. Then, perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features. For example, according to the discriminability of the local descriptors and the stability of the global descriptors, assign different weights and linearly combine them to obtain enhanced features.
[0138] Next, divide the enhanced dynamic features into a current-frame feature set and a subsequent-frame feature set according to the time frame. For example, take the enhanced features of the first frame as the current-frame feature set and the enhanced features of the second frame as the subsequent-frame feature set. Perform multi-level downsampling on the enhanced dynamic features. For example, use the voxel grid downsampling method to divide the point cloud into multiple voxels and select a representative point within each voxel. At each sampling level, calculate the geometric distance between feature points, such as the Euclidean distance, and the feature similarity, such as the cosine similarity. Dynamically group the feature points based on the geometric distance and feature similarity. For example, use the DBSCAN clustering algorithm to group points that are close in distance and similar in features into the same group. Establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure. For example, connect a feature point group at a coarser level to all its subgroups at a finer level.
[0139] Based on the constructed hierarchical feature point structure, calculate the association strength of feature points in the dimensions of spatial position, feature expression, and temporal variation. For example, the spatial position association strength can be calculated from the Euclidean distance between feature points, the feature expression association strength can be calculated from the cosine similarity between feature vectors, and the temporal variation association strength can be calculated from the displacement magnitude of feature points between consecutive frames. Convert these association strengths into feature transfer weights. For example, use the softmax function to normalize the association strength between 0 and 1. Utilize the feature transfer weights to perform information interaction among feature points at different sampling levels to generate multi-scale fusion feature points. For example, weighted transfer the feature information at a coarser level to a finer level, and weighted transfer the feature information at a finer level to a coarser level.
[0140] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points. For example, use a Kalman filter to predict the position of a feature point in the next frame based on its motion trajectory in the past few frames. Calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint. For example, calculate the Euclidean distance between the predicted position and the actual position. Calculate the feature similarity matrix for the multi-scale fusion feature points. For example, calculate the cosine similarity between each pair of feature points. Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints. For example, transfer the feature similarity matrix at a coarser level to a finer level to guide the feature point matching at the finer level. Combine the temporal constraint and the cross-scale feature point matching constraint into an objective function. For example, use the weighted sum of the temporal constraint and the cross-scale matching constraint as the objective function.
[0141] Optimize the objective function successively at multiple sampling levels to obtain the feature point matching relationship. For example, use the iterative closest point algorithm to minimize the objective function and find the best feature point matching relationship. Based on the feature point matching relationship, pair the three-dimensional spatial coordinates of the matching feature points under their respective time frames. Calculate the difference in Euclidean distance of the paired coordinates to obtain the feature point displacement. Divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector. Suppose there are two matching points, with coordinates (1, 2, 3) in the first frame and coordinates (2, 3, 4) in the second frame, and the time interval is 0.1 seconds. Then the feature point displacement is (1, 1, 1), and the feature point motion vector is (10, 10, 10).
[0142] Construct a manifold space representation of the feature point motion vectors. For example, use Lie algebra to represent rotational motion. Calculate the orthogonal projection of the feature point motion vectors onto the manifold space. For example, project the motion vectors onto the rotation plane. Take the difference between the orthogonal projection and the original motion vectors as the projection error. Iteratively optimize the weighted sum of squares of the projection error until convergence. For example, use the gradient descent method to minimize the projection error. Take the convergence result as the pose parameters of the target object. For example, convert the final rotation matrix into Euler angles or quaternion to represent the pose.
[0143] In this embodiment, through multi-scale feature fusion and manifold constraint optimization, the spatio-temporal features and motion laws of the point cloud are effectively captured, thereby improving the accuracy of pose estimation. It has strong robustness to noise and occlusion and can stably estimate the target pose in complex scenarios. The design of the hierarchical feature point structure and the feature propagation network reduces the computational complexity and improves the efficiency of pose estimation.
[0144] In an alternative embodiment, input the motion vectors and pose parameters into a dual threshold decision module. Calculate the speed threshold value and the pose threshold value respectively through the dual threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area, and a person position heat map is generated based on the center position of the point cloud cluster. The output of the person perception result based on the person position heat map includes:
[0145] Sample the motion vectors within a sliding time window of a specified length, and construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vectors according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a time series change model for the pose parameters, extract the mean and standard deviation of the pose parameters based on the time series change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value;
[0146] Compare the motion vector with the speed threshold value, and compare the pose parameter with the pose threshold value; when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area; use the region growing algorithm to segment the point cloud in the target area to obtain the target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain the filtered point cloud cluster;
[0147] Calculate the centroid of the filtered point cloud clusters to obtain the center position coordinates of the point cloud clusters, and map the center position coordinates of the point cloud clusters to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinates; perform kernel density estimation on the two-dimensional ground plane coordinates using a Gaussian kernel function with an adaptive bandwidth to obtain the spatial density distribution; calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud clusters, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map;
[0148] Construct a recursive filter to perform temporal filtering on the initial heat map to obtain the filtered heat map; calculate the fusion weight based on the similarity between the filtered heat map and the heat map of the previous moment; apply the fusion weight to the filtered heat map and the heat map of the previous moment to obtain the personnel position heat map; apply the local maximum suppression algorithm on the personnel position heat map to extract the extreme points of the heat map; use the extreme points of the heat map as the candidate personnel positions;
[0149] Perform inter-frame matching on the candidate personnel positions to obtain the initial personnel trajectory, and use a Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory; construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the personnel movement trajectory; use the multi-feature fusion model to calculate the confidence score of each candidate personnel position; set a dynamic threshold to screen the confidence scores, and output the candidate personnel positions that meet the dynamic threshold and their corresponding personnel movement trajectories as the personnel perception results.
[0150] A personnel perception method aims to accurately identify and track personnel from three-dimensional point cloud data. This method uses motion information, pose features, and point cloud processing techniques to generate a personnel position heat map, and combines multi-feature fusion and trajectory prediction to achieve reliable perception of personnel.
[0151] First, in the preprocessing stage, perform denoising and filtering on the input three-dimensional point cloud data to remove the influence of environmental noise and stray points, so as to improve the accuracy of subsequent processing steps. For example, use a statistical filter to remove outliers and use a bilateral filter to smooth the point cloud data while retaining edge information.
[0152] Next, perform motion estimation on the preprocessed point cloud data. Within a sliding time window of a specified length (e.g., 1 second), sample the motion vectors of each point. Assume that 10 motion vectors are sampled within a time window, with magnitudes of 0.1 m / s, 0.2 m / s, ..., 1.0 m / s respectively. Based on these sampled data, construct a motion vector distribution model, such as a Gaussian distribution model. According to this model, calculate the statistical mean (0.55 m / s in this example) and standard deviation (e.g., 0.28 m / s) of the motion vectors. Determine the weighted combination of the statistical mean and standard deviation (e.g., mean weight is 0.7, standard deviation weight is 0.3) as the speed threshold value, such as 0.47 m / s.
[0153] Meanwhile, establish a temporal variation model for the pose parameters. For example, assume the pose parameter is the human height. Within a time window, sample the human height and establish a time series model. Based on this model, extract the mean and standard deviation of the pose parameters. Similar to the calculation method of the speed threshold value, determine the weighted combination of the mean and standard deviation as the pose threshold value. For example, assume the mean is 1.75 m, the standard deviation is 0.1 m, and the pose threshold value after weighted combination is 1.72 m.
[0154] Then, compare the motion vector of each point with the speed threshold value, and compare the pose parameter with the pose threshold value. When the motion vector of a point is greater than the speed threshold value (e.g., 0.5 m / s is greater than 0.47 m / s) and the pose parameter is greater than the pose threshold value (e.g., 1.8 m is greater than 1.72 m), then determine that this point belongs to the target area, indicating that there may be a person.
[0155] Next, use the region growing algorithm to segment the point cloud of the target area to obtain the target point cloud cluster. Perform spatial filtering and noise removal on the target point cloud cluster to obtain the filtered point cloud cluster. Calculate the centroid of the filtered point cloud cluster to obtain the coordinates of the center position of the point cloud cluster. Map the coordinates of the center position of the point cloud cluster to the ground plane coordinate system through the projection matrix to obtain the two-dimensional ground plane coordinates. For example, assume the three-dimensional coordinates of the center of the point cloud cluster are (1, 2, 3), and map it to the ground plane coordinate system through the projection matrix to obtain the two-dimensional coordinates (1, 2).
[0156] Use a Gaussian kernel function with an adaptive bandwidth to perform kernel density estimation on the two-dimensional ground plane coordinates to obtain the spatial density distribution. Calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud cluster. Perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map.
[0157] Construct a recursive filter to perform temporal filtering on the initial heatmap to obtain a filtered heatmap. Calculate the fusion weight based on the similarity between the filtered heatmap and the heatmap of the previous moment. Apply the fusion weight to the filtered heatmap and the heatmap of the previous moment to obtain the personnel position heatmap. Apply the local maximum suppression algorithm to the personnel position heatmap to extract the extreme points of the heatmap. Use the extreme points of the heatmap as candidate points for personnel positions.
[0158] Perform inter-frame matching on the candidate points for personnel positions to obtain the initial personnel trajectory. Use the Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory. Construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the personnel movement trajectory. Use the multi-feature fusion model to calculate the confidence score for each candidate point for personnel positions. Set a dynamic threshold to screen the confidence scores, and output the candidate points for personnel positions that meet the dynamic threshold and their corresponding personnel movement trajectories as the personnel perception result.
[0159] In this embodiment, by comprehensively using motion information, pose features, and point cloud processing technology, personnel can be effectively distinguished from other objects, and the false detection rate can be reduced. By using temporal filtering and multi-feature fusion technology, complex scenarios such as noise and occlusion can be effectively processed, and the stability of personnel perception can be improved. By generating the personnel position heatmap and combining the Kalman filter for trajectory prediction, precise positioning and continuous tracking of personnel can be achieved.
[0160] In an alternative embodiment, performing inter-frame matching on the candidate points for personnel positions to obtain the initial personnel trajectory, and using the Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory includes:
[0161] Obtain the candidate points for personnel positions in consecutive time frames; calculate the Euclidean distance between the candidate points for personnel positions in adjacent time frames to obtain the distance feature; calculate the motion speed based on the historical position information of the candidate points for personnel positions, and calculate the speed difference between the candidate points for personnel positions in adjacent time frames based on the motion speed to obtain the speed feature; calculate the motion direction based on the motion speed, and calculate the direction difference between the candidate points for personnel positions in adjacent time frames based on the motion direction to obtain the direction feature;
[0162] Construct a matching cost matrix from the distance feature, speed feature, and direction feature; use the candidate points for personnel positions in the current time frame as the target point set and the candidate points for personnel positions in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain matching point pairs; connect the matching point pairs in chronological order to obtain the initial personnel trajectory;
[0163] Construct a state vector from the position coordinates, movement speed, and acceleration in the initial trajectory of the person; construct a state transition equation based on the state vector; construct an observation equation based on the position coordinates of the initial trajectory of the person; calculate the prediction error between the predicted value and the actual observed value of the state transition equation; update the process noise covariance according to the prediction error;
[0164] Use the state transition equation and the process noise covariance to predict the state vector to obtain a state prediction value and a prediction covariance; calculate the Kalman gain according to the state prediction value, the prediction covariance, and the observation equation; use the Kalman gain to correct the state prediction value to obtain a state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value;
[0165] Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract the historical motion pattern within the prediction window length; use the historical motion pattern to repair the trajectory break points;
[0166] Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than the preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain an optimized person movement trajectory.
[0167] A method for generating a person movement trajectory based on multi-feature matching and Kalman filtering, which can effectively handle complex scenarios such as occlusion and noise, and generate a smooth, continuous, and accurate person movement trajectory.
[0168] First, obtain the candidate points of the person's position in continuous time frames. For example, use an object detection algorithm to detect the person object in each frame image of the video, obtain the bounding box coordinates of each person, and use the center point of the bounding box as the candidate point of the person's position. Assume that 3 people are detected in the t-th frame, and their position coordinates are (10, 20), (30, 40), and (50, 60) respectively.
[0169] Next, calculate the Euclidean distance, speed difference, and direction difference between the candidate points of the personnel positions in adjacent time frames to construct a multi-feature matching cost matrix. For example, assume that the position of a certain person in the (t - 1)-th frame is (8, 18), and calculate its Euclidean distances from the three candidate points (10, 20), (30, 40), and (50, 60) in the t-th frame, which are 2.83, 22.83, and 42.83 respectively. At the same time, calculate the movement speed and direction based on the historical position information. Assume that the speed of this person in the (t - 1)-th frame is (2, 2) and the direction is 45 degrees, then calculate the speed difference and direction difference between it and the three candidate points in the t-th frame. Combine the distance, speed difference, and direction difference with weights to construct the matching cost matrix.
[0170] Then, construct a bipartite graph based on the matching cost matrix and use the Hungarian algorithm for minimum cost matching to obtain the matching point pairs. For example, assume that the cost matrix indicates that the matching cost between (8, 18) and (10, 20) is the smallest, then connect them to form a matching point pair. Connect all the matching point pairs in chronological order to obtain the initial trajectory of the personnel.
[0171] Next, form a state vector with the position coordinates, movement speed, and acceleration in the initial trajectory of the personnel, and construct a Kalman filter. According to the kinematic model, construct a state transition equation to describe the variation law of the state vector over time. At the same time, construct an observation equation to describe the relationship between the observed value (i.e., the personnel position coordinates) and the state vector. Use the Kalman filter to smooth and predict the initial trajectory. For example, calculate the prediction error between the predicted value of the state transition equation and the actual observed value based on the initial trajectory, and update the process noise covariance accordingly. Then, use the updated state transition equation and process noise covariance to predict the state vector to obtain the state prediction value and prediction covariance. Finally, correct the state prediction value using the Kalman gain to obtain the state estimate value, and generate a smooth trajectory based on the position coordinates in the state estimate value.
[0172] Detect the continuity of the smooth trajectory and identify the trajectory break points. Statistically analyze the effective trajectory length before the trajectory break points and determine the prediction window length based on this length. Extract the historical motion pattern within the prediction window length and use this pattern to repair the trajectory break points. For example, if a person's trajectory breaks at the 10th frame and the effective trajectory length before that is 9 frames, then the prediction window length can be set to 9 frames, extract the motion pattern of this person in the previous 9 frames, and use this pattern to predict their position after the 10th frame to repair the trajectory break.
[0173] Calculate the position similarity and velocity similarity between the repaired trajectories, and use the weighted combination of the two as the trajectory similarity. When the trajectory similarity is greater than the preset threshold, calculate the reliability weight according to the number of valid observations of the trajectory, and use this weight to merge the similar trajectories to obtain the optimized personnel movement trajectory. For example, if two repaired trajectories are very similar in terms of position and velocity, they can be merged into one trajectory.
[0174] Existing methods for generating personnel movement trajectories mainly rely on single feature matching, making it difficult to ensure the continuity and accuracy of trajectories in complex scenarios. They are vulnerable to occlusion and noise interference, resulting in trajectory breaks or incorrect matches. In addition, traditional trajectory repair mostly uses simple interpolation, lacking the utilization of individual movement patterns, with low prediction accuracy. Trajectory merging also lacks effective similarity metrics, affecting the quality of the final trajectory. This application uses a multi-feature matching method to comprehensively construct a matching cost matrix based on distance, velocity, and direction features, and uses the Hungarian algorithm to achieve optimal matching, improving the accuracy of frame-by-frame matching and reducing incorrect associations. At the same time, a Kalman filter is introduced to construct a state transition equation and an observation equation to smooth and predict the trajectory, making the trajectory more stable and continuous and reducing noise interference. In terms of trajectory repair, this application detects trajectory breakpoints, extracts historical movement patterns, and completes the trajectory within the prediction window, making the repaired trajectory more in line with the actual movement trend. For trajectory optimization, the position similarity and velocity similarity are comprehensively calculated, and a reliability weight is introduced for trajectory merging to avoid incorrect merging problems caused by a single feature. Compared with the prior art, this application has been improved in terms of matching accuracy, trajectory smoothing, trajectory repair, and trajectory optimization, making the finally generated personnel movement trajectory smoother, more accurate, and more coherent, capable of effectively dealing with complex scenarios such as occlusion and noise, and improving the robustness and reliability of trajectory generation.
[0175] Figure 2 This is the scatter plot of the feature extraction performance for the embodiments of the present invention. As Figure 2 shown, the scatter plot of the feature extraction performance shows the distribution of different methods in terms of feature dimension and computational complexity. This technical solution (circular markers) achieves a feature extraction accuracy of 0.95 at 128 feature dimensions, with a computational complexity of O(nlog n ); the traditional PCA method (square markers) has an accuracy of 0.82 at the same feature dimension, with a complexity of O(n²); the ICA algorithm (triangle markers) has an accuracy of 0.87, and its complexity is between the two. By introducing a multi-level feature mapping and an attention mechanism, this technical solution significantly improves the accuracy of feature extraction while maintaining a low computational complexity, especially showing a better clustering effect in the high-dimensional feature space, and the discrimination between features has been improved by about 15%. The distribution density of the scatter points also indicates that the feature extraction results of this technical solution are more stable, with the variance reduced by about 25%.
[0176] Figure 3 This is a comparative effect diagram of point cloud cluster segmentation based on temporal filtering for the features of the embodiments of the present invention. As Figure 3 shown, the comparative diagram of point cloud cluster segmentation shows the performance of different segmentation methods in 100 consecutive frames. This technical solution (circular marker) achieves a segmentation accuracy rate of 95.2% through temporal filtering and region growing algorithm, with a noise rate of 2.8% and an average processing time of 15 ms / frame. The segmentation accuracy rate of the traditional region growing algorithm (square marker) is 82.5%, the noise rate is 8.6%, and the processing time is 25 ms / frame. The segmentation accuracy rate of the Euclidean clustering method (triangle marker) is 85.3%, the noise rate is 7.2%, and the processing time is 22 ms / frame. In a scenario with a personnel density of 3 people per square meter, the missed detection rate of this solution is only 1.5%, significantly lower than 4.8% and 4.2% of the traditional methods. By introducing temporal information and adaptive parameter adjustment, this solution has achieved obvious improvements in both the accuracy and real-time performance of point cloud segmentation.
[0177] Figure 4 This is a schematic structural diagram of a lightweight AI personnel perception system based on a millimeter-wave radar for the embodiments of the present invention. As Figure 4 shown, the system includes:
[0178] A first unit for collecting millimeter-wave echo signals in a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals according to distance dimension and angle dimension to obtain filtered echo data, and constructing three-dimensional space point cloud data based on the filtered echo data;
[0179] A second unit for calculating the distance difference and angle difference between adjacent points in the three-dimensional space point cloud data, constructing a spatial distribution matrix based on the distance difference and angle difference, calculating the center position of the point cloud cluster based on the spatial distribution matrix, establishing an adaptive sampling radius with the center position of the point cloud cluster as a reference, extracting continuously changing point cloud features within the adaptive sampling radius, and calculating motion vectors and attitude parameters based on the extracted point cloud features;
[0180] A third unit for inputting the motion vectors and attitude parameters into a dual-threshold decision module, calculating a speed threshold value and an attitude threshold value respectively through the dual-threshold decision module, determining that there are personnel in the target area when the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, generating a personnel position heat map based on the center position of the point cloud cluster, and outputting a personnel perception result based on the personnel position heat map.
[0181] In the third aspect of the embodiments of the present invention,
[0182] Provided is an electronic device, comprising:
[0183] a processor;
[0184] a memory for storing instructions executable by the processor;
[0185] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0186] In a fourth aspect of the embodiments of the present invention,
[0187] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0188] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0189] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A lightweight AI personnel perception method based on millimeter-wave radar, characterized in that, Including: Collecting millimeter-wave echo signals of a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data according to the filtered echo data; Constructing a radial density distribution matrix of the three-dimensional spatial point cloud data, calculating an adaptive search radius for each point based on the gradient change of the radial density distribution matrix, extracting an adjacent point set within the adaptive search radius, and calculating the distance difference and angle difference between point pairs in the adjacent point set; Constructing a double-layer weight modulation function, where the first-layer weight coefficient of the double-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Inputting the distance difference and angle difference into the double-layer weight modulation function to obtain weight features, and constructing a spatial distribution matrix based on the weight features; Performing eigen-decomposition on the spatial distribution matrix to obtain a principal eigen-subspace and a secondary eigen-subspace, mapping the principal eigen-subspace and the secondary eigen-subspace to a topological feature network, calculating the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and using the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters; Constructing a density gradient optimization function within each point cloud clustering cluster, inputting the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculating the density difference between any point pairs within the point cloud clustering cluster, determining the density increasing direction based on the density difference, and iteratively updating the search position until convergence to obtain the position of the local density maximum, and determining the position of the local density maximum as the center position of the point cloud cluster; Constructing a deformation kernel function with the center position of the point cloud cluster as a reference point, calculating the local curvature distribution within the neighborhood of the reference point, mapping the local curvature distribution to the deformation parameter of the kernel function, adjusting the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, constructing a feature distribution tensor within the adaptive sampling radius range, extracting the principal feature components of the feature distribution tensor, and combining the principal feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud; Constructing a feature propagation network, inputting the dynamic features of the point cloud into the feature propagation network, identifying the spatial positions of corresponding feature points in consecutive time frames, calculating the displacement between the corresponding feature points to obtain a feature point motion vector, performing manifold constraint optimization on the feature point motion vector to obtain the attitude parameters of the target object; inputting the motion vector and the attitude parameters into a double-threshold decision module, calculating a speed threshold value and an attitude threshold value respectively through the double-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area, and a personnel position heat map is generated according to the center position of the point cloud cluster, and a personnel perception result is output based on the personnel position heat map.
2. The method according to claim 1, characterized in that, Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibration signal, perform two-way joint filtering processing on the calibration signal in the distance dimension and the angle dimension to obtain the filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data, including: Collect the millimeter-wave echo signal of the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signal, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signal with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain a calibration signal; Perform wavelet packet decomposition on the calibration signal to obtain multi-scale signal components, extract the energy entropy feature vector of the multi-scale signal components, calculate the adaptive reconstruction weight of each scale signal component based on the energy entropy feature vector, multiply the adaptive reconstruction weight by the corresponding scale signal component and sum them up to obtain an enhanced signal; Extract local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, use the threshold surface to extract the distance-dimensional target echo, construct an orthogonal projection operator with the covariance matrix eigenvector of the distance-dimensional target echo, and project the calibration signal into the null space of the orthogonal projection operator to obtain the angle-dimensional target echo; Extract multi-level feature maps by performing feature extraction on the distance-dimensional target echo and the angle-dimensional target echo, input the multi-level feature maps into an attention fusion network to generate a target scattering probability map, and extract the filtered echo data according to the maximum value position of the target scattering probability map; Obtain the initial point cloud by performing polar coordinate mapping on the filtered echo data, calculate the local density gradient of the initial point cloud, establish a diffusion equation of the point cloud growth model based on the local density gradient, and solve the diffusion equation to obtain the enhanced point cloud; Calculate the principal curvature and normal vector of the enhanced point cloud, construct an affinity matrix using the principal curvature and normal vector for spectral clustering, and obtain three-dimensional spatial point cloud data.
3. The method according to claim 2, wherein Extract multi-level feature maps by performing feature extraction on the distance-dimensional target echo and the angle-dimensional target echo, input the multi-level feature maps into an attention fusion network to generate a target scattering probability map, and extract the filtered echo data according to the maximum value position of the target scattering probability map, including: Input the distance-dimensional target echo into a phase-sensitive complex-valued filter bank to obtain a first feature layer, calculate the conditional entropy of the first feature layer to obtain the redundancy between feature layers, calculate the information gain rate of the first feature layer to obtain the feature discriminability, and optimize the time scale parameter and phase parameter of the phase-sensitive complex-valued filter bank based on the redundancy between feature layers and the feature discriminability to obtain a first multi-level feature map; Input the angle-dimensional target echo into the direction-sensitive filter bank, calculate the local region direction consistency of the angle-dimensional target echo to obtain a direction response map, extract the main direction and the secondary direction based on the direction response map, use the main direction and the secondary direction as the reference directions of the direction-sensitive filter bank, calculate the feature responses in the main direction and the secondary direction to obtain a feature response difference, and optimize the angle resolution parameter of the direction-sensitive filter bank based on the feature response difference to obtain a second multi-level feature map; Calculate the mutual information matrix for the first multi-level feature map and the second multi-level feature map, construct the cross-attention weight based on the principal component vector of the mutual information matrix, obtain the channel attention weight based on the channel statistics of the first multi-level feature map and the second multi-level feature map, and combine the cross-attention weight and the channel attention weight to obtain a fusion attention weight; Calculate the similarity pattern of the local features based on the fusion attention weight to obtain a feature sampling offset, resample the first multi-level feature map and the second multi-level feature map according to the feature sampling offset to obtain resampled features, calculate the local density distribution matrix and the density jump matrix of the resampled features, and generate a target scattering probability map based on the local density distribution matrix and the density jump matrix; Determine the coordinates of the scattering center position based on the maximum value position of the target scattering probability map, construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, and perform a convolution operation on the adaptive kernel function with the range-dimensional target echo and the angle-dimensional target echo to obtain filtered echo data.
4. The method according to claim 1, wherein Construct a feature propagation network, input the point cloud dynamic features into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain feature point motion vectors, and perform manifold constraint optimization on the feature point motion vectors to obtain the pose parameters of the target object, including: Obtain the point cloud dynamic features of consecutive time frames, extract the local shape descriptor and the global descriptor representing the regional distribution for the point cloud dynamic features; perform adaptive weighted fusion on the local shape descriptor and the global descriptor to obtain enhanced dynamic features; divide the enhanced dynamic features into a current frame feature set and a subsequent frame feature set according to the time frame; Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distances and feature similarities between feature points at each sampling level; perform dynamic grouping on the feature points based on the geometric distances and feature similarities to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure; Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal change; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fusion feature points; Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fusion feature points; Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints; combine the temporal constraint and the cross-scale feature point matching constraints into an objective function; Optimize the objective function sequentially at multiple sampling levels to obtain the feature point matching relationship; based on the feature point matching relationship, correspond and pair the three-dimensional space coordinates of the matching feature points under their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector; Construct a manifold space representation of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and take the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and take the convergence result as the pose parameter of the target object.
5. The method according to claim 1, wherein Input the motion vector and the pose parameter into a dual-threshold decision module, and calculate the speed threshold value and the pose threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area, and a person position heat map is generated according to the center position of the point cloud cluster. Based on the person position heat map, the person perception result is output, including: Sample the motion vector within a sliding time window of a specified length, and construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vector according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a temporal change model for the pose parameter, extract the mean and standard deviation of the pose parameter based on the temporal change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value; Compare the motion vector with the speed threshold value, and compare the pose parameter with the pose threshold value; when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area; use the region growing algorithm to segment the point cloud of the target area to obtain the target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain the filtered point cloud cluster; Calculate the centroid of the filtered point cloud cluster to obtain the center position coordinate of the point cloud cluster, map the center position coordinate of the point cloud cluster to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinate; perform kernel density estimation on the two-dimensional ground plane coordinate using a Gaussian kernel function with an adaptive bandwidth to obtain the spatial density distribution; calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud cluster, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map; Construct a recursive filter to perform temporal filtering on the initial heatmap to obtain a filtered heatmap; calculate the fusion weight based on the similarity between the filtered heatmap and the heatmap of the previous moment; apply the fusion weight to the filtered heatmap and the heatmap of the previous moment to obtain a personnel position heatmap; apply the local maximum suppression algorithm to the personnel position heatmap to extract the extreme points of the heatmap; use the extreme points of the heatmap as candidate points for personnel positions; Perform inter-frame matching on the candidate points for personnel positions to obtain an initial personnel trajectory, and use a Kalman filter to smooth and predict the initial personnel trajectory to obtain a personnel movement trajectory; construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the personnel movement trajectory; calculate the confidence score for each candidate point for personnel positions using the multi-feature fusion model; set a dynamic threshold to screen the confidence scores, and output the candidate points for personnel positions that meet the dynamic threshold and their corresponding personnel movement trajectories as the personnel perception result.
6. The method according to claim 5, wherein Performing inter-frame matching on the candidate points for personnel positions to obtain an initial personnel trajectory, and using a Kalman filter to smooth and predict the initial personnel trajectory to obtain a personnel movement trajectory includes: Obtain the candidate points for personnel positions in consecutive time frames; calculate the Euclidean distance between the candidate points for personnel positions in adjacent time frames to obtain a distance feature; calculate the motion speed based on the historical position information of the candidate points for personnel positions, and calculate the speed difference between the candidate points for personnel positions in adjacent time frames based on the motion speed to obtain a speed feature; calculate the motion direction based on the motion speed, and calculate the direction difference between the candidate points for personnel positions in adjacent time frames based on the motion direction to obtain a direction feature; Construct a matching cost matrix from the distance feature, speed feature, and direction feature; use the candidate points for personnel positions in the current time frame as the target point set and the candidate points for personnel positions in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain matching point pairs; connect the matching point pairs in chronological order to obtain an initial personnel trajectory; Form a state vector from the position coordinates, motion speed, and acceleration in the initial personnel trajectory; construct a state transition equation based on the state vector; construct an observation equation based on the position coordinates of the initial personnel trajectory; calculate the prediction error between the predicted value and the actual observed value of the state transition equation; update the process noise covariance according to the prediction error; Use the state transition equation and the process noise covariance to predict the state vector to obtain a state prediction value and a prediction covariance; calculate the Kalman gain based on the state prediction value, prediction covariance, and observation equation; correct the state prediction value using the Kalman gain to obtain a state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value; Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract historical motion patterns within the prediction window length; use the historical motion patterns to repair the trajectory break points; Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than a preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain an optimized personnel motion trajectory.
7. A lightweight AI personnel perception system based on millimeter-wave radar, used to implement the method described in any one of the foregoing claims 1-6, characterized in that, Comprising: A first unit, configured to collect millimeter-wave echo signals of a target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, perform two-way joint filtering processing on the calibrated signals according to the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data; A second unit, configured to calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate a motion vector and attitude parameters based on the extracted point cloud features; A third unit, configured to input the motion vector and the attitude parameters into a dual-threshold decision module, calculate a speed threshold value and an attitude threshold value respectively through the dual-threshold decision module, when the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, determine that there are personnel in the target area, generate a personnel position heat map according to the center position of the point cloud cluster, and output a personnel perception result based on the personnel position heat map.
8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Fall posture recognition method and system based on millimeter wave radar point cloud
CN114942434A
Multi-person 5D radar falling detection method and system based on artificial intelligence algorithm
CN117079416A