Lightweight AI personnel sensing method and system based on millimeter wave radar

By performing phase compensation and amplitude calibration on the millimeter wave echo signal, combined with bidirectional joint filtering processing and adaptive sampling radius technology, three-dimensional spatial point cloud data are constructed and personnel are identified, which solves the problems of large amount of calculation and insufficient robustness in the existing technology, and achieves efficient and reliable personnel perception effects.

CN120195652AActive Publication Date: 2025-06-24DEXIAOBAO HEALTH TECHNOLOGY (CHANGZHOU) CO LTD

Patent Information

Application Number
CN202510672610.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing personnel perception method based on millimeter wave radar is large in computing and high in equipment power consumption, and is not suitable for edge devices. It has insufficient recognition robustness in complex scenarios, high error detection rate, and lacks effective adaptive processing methods for point cloud data.

Method used

The lightweight AI personnel perception method is adopted to perform phase compensation and amplitude calibration of the millimeter wave echo signal, combined with bidirectional joint filtering processing, three-dimensional spatial point cloud data is constructed, and the center of the point cloud cluster is identified through the spatial distribution matrix and adaptive sampling radius, continuously changing point cloud features are extracted, motion vectors and pose parameters are calculated, and personnel existence is determined using the dual threshold judgment module.

Benefits of technology

It improves personnel detection accuracy, reduces computing burden, is suitable for terminal devices with limited resources, enhances adaptability and real-timeness to complex environments, and improves detection reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120195652A_ABST
    Figure CN120195652A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight AI personnel perception method and system based on a millimeter wave radar, and relates to the technical field of personnel perception, and the method comprises the steps: carrying out the phase compensation, amplitude calibration and bidirectional combined filtering processing of a millimeter wave echo signal, and constructing three-dimensional space point cloud data. Then, constructing a spatial distribution matrix based on the point cloud data, calculating the center position of a point cloud cluster, establishing an adaptive sampling radius to extract point cloud features, and calculating a motion vector and attitude parameters; and finally, inputting the motion vector and the attitude parameter into a double threshold judgment module, judging whether a person exists in the target area, generating a person position thermodynamic diagram based on the center position of the point cloud cluster, and outputting a person perception result. According to the method, the calculation complexity is reduced, the real-time performance and accuracy of personnel perception are improved, and the method is suitable for smart home, security monitoring and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to personnel perception technology, and particularly to a lightweight AI personnel perception method and system based on millimeter-wave radar. Background Art

[0002] Existing personnel perception methods based on millimeter-wave radar mainly rely on large-scale neural networks for point cloud processing and target recognition. Although this method ensures a certain degree of accuracy, it has a large amount of calculation and high device power consumption, and is not suitable for edge devices or embedded systems with limited computing resources. In addition, traditional point cloud processing methods are vulnerable to environmental noise interference in complex scenarios, resulting in insufficient robustness of personnel recognition.

[0003] Current millimeter-wave radar personnel perception methods usually use fixed thresholds for target detection, but fixed thresholds are difficult to adapt to the dynamic changes of different scenarios, resulting in a high false detection rate in low signal-to-noise ratio environments. In addition, existing technologies lack effective adaptive processing methods for point cloud data, resulting in a decrease in the detection accuracy of non-rigid targets (such as people walking or changing postures). At the same time, in terms of personnel position tracking, existing methods mainly rely on multi-frame fusion or deep learning modeling, but these methods have a high computational complexity and are difficult to meet the real-time requirements, which limits the deployment of millimeter-wave radar in edge computing devices.

[0004] Therefore, there is an urgent need for a lightweight AI personnel perception method and system based on millimeter-wave radar to improve personnel detection accuracy, reduce the computational burden, and be applicable to resource-constrained terminal devices. Summary of the Invention

[0005] Embodiments of the present invention provide a lightweight AI personnel perception method and system based on millimeter-wave radar, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention,

[0007] A lightweight AI personnel perception method based on millimeter-wave radar is provided, including:

[0008] Collecting millimeter-wave echo signals in a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional space point cloud data according to the filtered echo data;

[0009] Calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix based on the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as the reference, extract continuously varying point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features;

[0010] Input the motion vector and attitude parameters into the dual-threshold decision module, calculate the speed threshold value and attitude threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area, generate a heat map of the personnel position based on the center position of the point cloud cluster, and output the personnel perception result based on the heat map of the personnel position.

[0011] In an alternative embodiment,

[0012] Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibrated signal, perform two-way joint filtering processing on the calibrated signal in the distance dimension and angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data, including:

[0013] Collect the millimeter-wave echo signal of the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signal, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signal with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain a calibrated signal;

[0014] Perform wavelet packet decomposition on the calibrated signal to obtain multi-scale signal components, extract the energy entropy feature vector of the multi-scale signal components, calculate the adaptive reconstruction weight of each scale signal component based on the energy entropy feature vector, multiply the adaptive reconstruction weight by the corresponding scale signal component and sum them to obtain an enhanced signal;

[0015] Extract local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, use the threshold surface to extract the distance-dimensional target echo, construct an orthogonal projection operator with the eigenvector of the covariance matrix of the distance-dimensional target echo, and project the calibrated signal onto the null space of the orthogonal projection operator to obtain the angle-dimensional target echo;

[0016] Feature extraction is performed on the range - dimension target echo and the angle - dimension target echo to obtain a multi - level feature map. The multi - level feature map is input into an attention fusion network to generate a target scattering probability map, and filtered echo data is extracted according to the maximum value position of the target scattering probability map;

[0017] The filtered echo data is mapped to an initial point cloud through polar coordinate mapping. The local density gradient of the initial point cloud is calculated, and a diffusion equation of a point cloud growth model is established based on the local density gradient. The diffusion equation is solved to obtain an enhanced point cloud;

[0018] The principal curvature and normal vector of the enhanced point cloud are calculated, and a spectral clustering is performed using the principal curvature and normal vector to construct an affinity matrix, obtaining three - dimensional space point cloud data.

[0019] In an alternative embodiment,

[0020] Feature extraction is performed on the range - dimension target echo and the angle - dimension target echo to obtain a multi - level feature map. The multi - level feature map is input into an attention fusion network to generate a target scattering probability map, and filtering the echo data according to the maximum value position of the target scattering probability map includes:

[0021] The range - dimension target echo is input into a phase - sensitive complex - valued filter bank to obtain a first feature layer. The conditional entropy of the first feature layer is calculated to obtain the redundancy between feature layers, and the information gain rate of the first feature layer is calculated to obtain the feature discrimination degree. Based on the redundancy between feature layers and the feature discrimination degree, the time - scale parameter and phase parameter of the phase - sensitive complex - valued filter bank are optimized to obtain a first multi - level feature map;

[0022] The angle - dimension target echo is input into a direction - sensitive filter bank. The local region direction consistency of the angle - dimension target echo is calculated to obtain a direction response map. The main direction and secondary direction are extracted based on the direction response map, and the main direction and secondary direction are used as the reference directions of the direction - sensitive filter bank. The feature response differences in the main direction and secondary direction are calculated, and the angle resolution parameter of the direction - sensitive filter bank is optimized based on the feature response differences to obtain a second multi - level feature map;

[0023] The mutual information matrix is calculated for the first multi - level feature map and the second multi - level feature map. The cross - attention weights are constructed based on the principal component vectors of the mutual information matrix, and the channel attention weights are obtained based on the channel statistics of the first multi - level feature map and the second multi - level feature map. The cross - attention weights and channel attention weights are combined to obtain fusion attention weights;

[0024] Calculate the similarity pattern of local features based on the fused attention weights to obtain a feature sampling offset. Resample the first multi-level feature map and the second multi-level feature map according to the feature sampling offset to obtain resampled features. Calculate the local density distribution matrix and the density jump matrix of the resampled features, and generate a target scattering probability map based on the local density distribution matrix and the density jump matrix;

[0025] Determine the scattering center position coordinates based on the maximum value position of the target scattering probability map, construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, and perform a convolution operation on the adaptive kernel function with the target echo in the distance dimension and the target echo in the angle dimension to obtain the filtered echo data.

[0026] In an alternative embodiment,

[0027] Calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data. Construct a spatial distribution matrix based on the distance difference and the angle difference. Calculate the center position of the point cloud cluster based on the spatial distribution matrix. Establish an adaptive sampling radius based on the center position of the point cloud cluster. Extract continuously varying point cloud features within the adaptive sampling radius, and calculate the motion vector and the pose parameters based on the extracted point cloud features, including:

[0028] Construct a radial density distribution matrix of the three-dimensional spatial point cloud data. Calculate the adaptive search radius of each point based on the gradient change of the radial density distribution matrix. Extract adjacent point sets within the adaptive search radius, and calculate the distance difference and the angle difference between point pairs in the adjacent point sets;

[0029] Construct a two-layer weight modulation function. The first-layer weight coefficient of the two-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Input the distance difference and the angle difference into the two-layer weight modulation function to obtain weight features, and construct a spatial distribution matrix based on the weight features;

[0030] Perform eigen-decomposition on the spatial distribution matrix to obtain the principal eigen-subspace and the secondary eigen-subspace. Map the principal eigen-subspace and the secondary eigen-subspace to a topological feature network. Calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and use the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters;

[0031] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculate the density difference between any pair of points within the point cloud clustering cluster, determine the density increasing direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum. Determine the position of the local density maximum as the center position of the point cloud cluster;

[0032] Construct a deformation kernel function with the center position of the point cloud cluster as the reference point, calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameter of the kernel function, adjust the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, construct a feature distribution tensor within the range of the adaptive sampling radius, extract the principal feature components of the feature distribution tensor, and combine the principal feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud;

[0033] Construct a feature propagation network. Input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vector of the feature points, and perform manifold constraint optimization on the motion vector of the feature points to obtain the pose parameters of the target object.

[0034] In an alternative embodiment,

[0035] Construct a feature propagation network. Input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vector of the feature points, and performing manifold constraint optimization on the motion vector of the feature points to obtain the pose parameters of the target object includes:

[0036] Obtain the dynamic features of the point cloud in consecutive time frames, extract local shape descriptors and global descriptors representing regional distributions from the dynamic features of the point cloud; perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features; divide the enhanced dynamic features into a current frame feature set and a subsequent frame feature set according to the time frame;

[0037] Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distance and feature similarity between feature points at each sampling level; perform dynamic grouping on the feature points based on the geometric distance and feature similarity to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure;

[0038] Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal variation; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fusion feature points;

[0039] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fusion feature points;

[0040] Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints; combine the temporal constraint and the cross-scale feature point matching constraints into an objective function;

[0041] Optimize the objective function successively at multiple sampling levels to obtain the feature point matching relationship; based on the feature point matching relationship, pair the three-dimensional spatial coordinates of the matching feature points under their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector;

[0042] Construct a manifold space representation of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and take the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and take the convergence result as the pose parameter of the target object.

[0043] In an alternative embodiment,

[0044] Input the motion vector and the pose parameter into a dual-threshold decision module, calculate the speed threshold value and the pose threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area, and a person position heat map is generated according to the center position of the point cloud cluster. The person perception result output based on the person position heat map includes:

[0045] Sample the motion vector within a sliding time window of a specified length, and construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vector according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a temporal change model for the pose parameter, extract the mean and standard deviation of the pose parameter based on the temporal change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value;

[0046] Compare the motion vector with a speed threshold value, and compare the attitude parameter with an attitude threshold value; when the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area; use the region growing algorithm to segment the point cloud of the target area to obtain a target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain a filtered point cloud cluster;

[0047] Calculate the centroid of the filtered point cloud cluster to obtain the center position coordinates of the point cloud cluster, and map the center position coordinates of the point cloud cluster to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinates; use a Gaussian kernel function with an adaptive bandwidth to perform kernel density estimation on the two-dimensional ground plane coordinates to obtain a spatial density distribution; calculate a confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud cluster, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain an initial heat map;

[0048] Construct a recursive filter to perform temporal filtering on the initial heat map to obtain a filtered heat map; calculate a fusion weight based on the similarity between the filtered heat map and the heat map of the previous moment; apply the fusion weight to the filtered heat map and the heat map of the previous moment to obtain a heat map of the person's position; apply a local maximum suppression algorithm to the heat map of the person's position to extract the extreme points of the heat map; use the extreme points of the heat map as candidate points for the person's position;

[0049] Perform inter-frame matching on the candidate points for the person's position to obtain an initial trajectory of the person, and use a Kalman filter to smooth and predict the initial trajectory of the person to obtain the movement trajectory of the person; construct a multi-feature fusion model by combining the motion vector, attitude parameter, and historical information of the person's movement trajectory; use the multi-feature fusion model to calculate the confidence score of each candidate point for the person's position; set a dynamic threshold to screen the confidence scores, and output the candidate points for the person's position that meet the dynamic threshold and their corresponding person's movement trajectories as the person perception result.

[0050] In an alternative embodiment,

[0051] Performing inter-frame matching on the candidate points for the person's position to obtain an initial trajectory of the person, and using a Kalman filter to smooth and predict the initial trajectory of the person to obtain the movement trajectory of the person includes:

[0052] Obtain the candidate points for the person's position in consecutive time frames; calculate the Euclidean distance between the candidate points for the person's position in adjacent time frames to obtain a distance feature; calculate the movement speed based on the historical position information of the candidate points for the person's position, and calculate the speed difference between the candidate points for the person's position in adjacent time frames based on the movement speed to obtain a speed feature; calculate the movement direction based on the movement speed, and calculate the direction difference between the candidate points for the person's position in adjacent time frames based on the movement direction to obtain a direction feature;

[0053] Construct the distance feature, speed feature, and direction feature into a matching cost matrix; use the candidate points of the personnel position in the current time frame as the target point set, and use the candidate points of the personnel position in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain matching point pairs; connect the matching point pairs in chronological order to obtain the initial trajectory of the personnel.

[0054] Form a state vector from the position coordinates, motion speed, and acceleration in the initial trajectory of the personnel; construct a state transition equation according to the state vector; construct an observation equation based on the position coordinates of the initial trajectory of the personnel; calculate the prediction error between the predicted value of the state transition equation and the actual observation value; update the process noise covariance according to the prediction error.

[0055] Predict the state vector using the state transition equation and the process noise covariance to obtain a state prediction value and a prediction covariance; calculate the Kalman gain according to the state prediction value, the prediction covariance, and the observation equation; correct the state prediction value using the Kalman gain to obtain a state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value.

[0056] Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract the historical motion pattern within the prediction window length; use the historical motion pattern to repair the trajectory break points.

[0057] Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than a preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain an optimized personnel motion trajectory.

[0058] In the second aspect of the embodiments of the present invention,

[0059] Provide a lightweight AI personnel perception system based on a millimeter-wave radar, including:

[0060] A first unit for collecting millimeter-wave echo signals in a target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, performing two-way joint filtering processing on the calibrated signals in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data according to the filtered echo data.

[0061] A second unit, configured to calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as a reference, extract continuously changing point cloud features within the adaptive sampling radius, and calculate a motion vector and pose parameters based on the extracted point cloud features;

[0062] A third unit, configured to input the motion vector and the pose parameters into a dual-threshold decision module, calculate a speed threshold value and a pose threshold value respectively through the dual-threshold decision module, determine that there is a person in the target area when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, generate a personnel position heat map according to the center position of the point cloud cluster, and output a personnel perception result based on the personnel position heat map.

[0063] In a third aspect of the embodiments of the present invention,

[0064] There is provided an electronic device, including:

[0065] A processor;

[0066] A memory for storing instructions executable by the processor;

[0067] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0068] In a fourth aspect of the embodiments of the present invention,

[0069] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0070] In this embodiment, by performing phase compensation and amplitude calibration on the millimeter-wave echo signal, the accuracy of the point cloud data is improved, making subsequent target detection more reliable. The bidirectional joint filtering method is adopted to optimize the signal in the range dimension and the angle dimension, reduce environmental noise interference, and effectively improve the stability of detection. At the same time, by constructing a spatial distribution matrix and an adaptive sampling radius, accurate identification of point cloud clusters is achieved, thereby improving the accuracy of dynamic personnel detection and avoiding the misjudgment problem caused by the fixed threshold method. It has lightweight computing capabilities and adopts an adaptive point cloud feature extraction method to reduce the computational complexity while ensuring the detection accuracy, and is suitable for embedded and edge computing devices. Based on the point cloud features, the motion vector and attitude parameters are calculated, and combined with the dual-threshold decision mechanism, efficient identification of the presence of personnel is achieved, which can effectively distinguish the static background from the real target and improve the reliability of detection. In addition, using the personnel position heat map, the intuitive visualization of the personnel distribution in the space can be realized, providing accurate perception capabilities for intelligent security, unmanned systems, and smart homes. While ensuring the detection accuracy, the computational burden is reduced, the adaptability to complex environments is enhanced, and the real-time performance is improved, providing strong support for the wide application of millimeter-wave radars in the field of intelligent perception. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 FIG. is a schematic flowchart of the lightweight AI personnel perception method based on millimeter-wave radar according to an embodiment of the present invention;

[0072] Figure 2 FIG. is a scatter plot of the feature extraction performance according to an embodiment of the present invention;

[0073] Figure 3 FIG. is a comparison effect diagram of the point cloud cluster segmentation based on the time series filtering of the features according to an embodiment of the present invention;

[0074] Figure 4 FIG. is a schematic structural diagram of the lightweight AI personnel perception system based on millimeter-wave radar according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0076] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0077] Figure 1 This is a schematic flowchart of the lightweight AI personnel perception method based on millimeter-wave radar according to an embodiment of the present invention. As Figure 1 shown, the method includes:

[0078] Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibration signal, perform two-way joint filtering processing on the calibration signal in the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data;

[0079] Calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features;

[0080] Input the motion vector and attitude parameters into the dual-threshold decision module, calculate the speed threshold value and the attitude threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area, and a personnel position heat map is generated according to the center position of the point cloud cluster, and a personnel perception result is output based on the personnel position heat map.

[0081] In an alternative embodiment, collecting the millimeter-wave echo signal of the target area, performing phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibration signal, performing two-way joint filtering processing on the calibration signal in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data according to the filtered echo data includes:

[0082] Collect the millimeter-wave echo signal of the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signal, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signal with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain a calibration signal;

[0083] Perform wavelet packet decomposition on the calibration signal to obtain multi-scale signal components, extract the energy entropy feature vector of the multi-scale signal components, calculate the adaptive reconstruction weight of each scale signal component based on the energy entropy feature vector, multiply the adaptive reconstruction weight by the corresponding scale signal component and superimpose them to obtain an enhanced signal;

[0084] Extract local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, extract range-dimension target echoes using the threshold surface, construct an orthogonal projection operator with the covariance matrix eigenvectors of the range-dimension target echoes, and project the calibration signal onto the null space of the orthogonal projection operator to obtain azimuth-dimension target echoes;

[0085] Extract multi-level feature maps from the range-dimension target echoes and azimuth-dimension target echoes, input the multi-level feature maps into an attention fusion network to generate a target scattering probability map, and extract filtered echo data based on the maximum position of the target scattering probability map;

[0086] Obtain an initial point cloud by polar coordinate mapping of the filtered echo data, calculate the local density gradient of the initial point cloud, establish a diffusion equation for the point cloud growth model based on the local density gradient, and solve the diffusion equation to obtain an enhanced point cloud;

[0087] Calculate the principal curvature and normal vector of the enhanced point cloud, construct an affinity matrix using the principal curvature and normal vector for spectral clustering, and obtain 3D spatial point cloud data.

[0088] A method for constructing 3D spatial point cloud data based on millimeter-wave radar echo signals, which is used to accurately perceive the 3D information of the target area. This method performs multi-level processing on millimeter-wave echo signals, including phase compensation and amplitude calibration, bidirectional joint filtering, as well as point cloud enhancement and clustering, and finally generates high-quality 3D point cloud data.

[0089] Exemplarily, first, collect millimeter-wave echo signals of the target area. For example, use a 77GHz millimeter-wave radar to collect echo signals of the target area with a sampling period of 10ms.

[0090] Then, perform phase compensation and amplitude calibration on the collected millimeter-wave echo signals. Calculate the phase differences between adjacent antennas. For example, calculate the phase difference between the first antenna and the second antenna, the phase difference between the second antenna and the third antenna, and so on. Construct a dynamic self-correction coefficient matrix with these phase differences. At the same time, calculate the amplitude ratios between adjacent antennas and map them into a non-linear compensation coefficient curve. For example, use the method of polynomial fitting to fit the amplitude ratio and the corresponding compensation coefficient into a curve. Perform matrix multiplication operations on the collected millimeter-wave echo signals with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain the calibrated signal. Assume that the millimeter-wave radar has 4 receiving antennas, then the dynamic self-correction coefficient matrix is a 4x4 matrix, and the non-linear compensation coefficient curve is a piecewise function.

[0091] Next, perform wavelet packet decomposition on the calibrated signal. For example, use the db4 wavelet for 4-layer decomposition to obtain signal components at 16 scales. Extract the energy entropy feature vectors of the signal components at each scale. For example, calculate the energy entropy values of the signal components at each scale and form a 16-dimensional feature vector. Calculate the adaptive reconstruction weights of the signal components at each scale based on the energy entropy feature vectors. For example, assign a weight to each scale according to the magnitude of the energy entropy value. The larger the energy entropy value, the larger the weight. Multiply the adaptive reconstruction weights by the corresponding scale signal components and sum them up to obtain the enhanced signal.

[0092] Extract the local maximum points as peak points on the time-frequency energy spectrum of the enhanced signal. For example, set a threshold and consider the points exceeding the threshold as peak points. Calculate the spatial clustering distribution characteristics of the peak points. For example, calculate the distances between the peak points and perform clustering based on the distances. Generate a threshold surface based on the spatial clustering distribution characteristics. For example, use the interpolation method to connect the clustering centers to form a surface. Use the threshold surface to extract the range-dimensional target echo. Construct an orthogonal projection operator with the eigenvector of the covariance matrix of the range-dimensional target echo. Project the calibrated signal onto the null space of the orthogonal projection operator to obtain the angle-dimensional target echo.

[0093] Perform feature extraction on the range-dimensional target echo and the angle-dimensional target echo to obtain a multi-level feature map. For example, extract the peak features of the range-dimensional target echo and the energy distribution features of the angle-dimensional target echo. Input the multi-level feature map into an attention fusion network to generate a target scattering probability map. For example, use a two-layer fully connected network for fusion. Extract the filtered echo data according to the maximum value position of the target scattering probability map.

[0094] Obtain the initial point cloud by performing polar coordinate mapping on the filtered echo data. For example, convert the range and angle information of the echo data into three-dimensional coordinates. Calculate the local density gradient of the initial point cloud. For example, calculate the number of points around each point and calculate its gradient. Establish a diffusion equation for the point cloud growth model based on the local density gradient and solve the diffusion equation to obtain the enhanced point cloud.

[0095] Calculate the principal curvature and normal vector of the enhanced point cloud. Construct an affinity matrix using the principal curvature and normal vector for spectral clustering to obtain the final three-dimensional spatial point cloud data. For example, use the k-means algorithm for clustering to finally obtain the three-dimensional point cloud data of the target.

[0096] In this embodiment, through multi-level signal processing and point cloud enhancement techniques, the influence of noise and clutter is effectively suppressed, the accuracy and integrity of the point cloud are improved, and the constructed point cloud data is clearer and more accurate. By adopting an adaptive phase compensation and amplitude calibration method, it can effectively cope with signal changes in different environments, improving the robustness and anti-interference ability of the system. By combining techniques such as wavelet packet decomposition, attention mechanism, and spectral clustering, the efficient processing of millimeter-wave radar echo signals and point cloud construction are realized, reducing the computational complexity and improving the processing efficiency.

[0097] In an alternative embodiment, feature extraction is performed on the range-dimension target echo and the angle-dimension target echo to obtain a multi-level feature map, and the multi-level feature map is input into an attention fusion network to generate a target scattering probability map. Extracting the filtered echo data according to the maximum position of the target scattering probability map includes:

[0098] Input the range-dimension target echo into a phase-sensitive complex-valued filter bank to obtain a first feature layer, calculate the conditional entropy of the first feature layer to obtain the redundancy between feature layers, calculate the information gain rate of the first feature layer to obtain the feature discrimination degree, and optimize the time-scale parameter and phase parameter of the phase-sensitive complex-valued filter bank based on the redundancy between feature layers and the feature discrimination degree to obtain a first multi-level feature map;

[0099] Input the angle-dimension target echo into a direction-sensitive filter bank, calculate the local region direction consistency of the angle-dimension target echo to obtain a direction response map, extract the main direction and the secondary direction based on the direction response map, use the main direction and the secondary direction as the reference directions of the direction-sensitive filter bank, calculate the feature response differences in the main direction and the secondary direction, and optimize the angle resolution parameter of the direction-sensitive filter bank based on the feature response differences to obtain a second multi-level feature map;

[0100] Calculate the mutual information matrix for the first multi-level feature map and the second multi-level feature map, construct cross-attention weights based on the principal component vectors of the mutual information matrix, obtain channel attention weights based on the channel statistics of the first multi-level feature map and the second multi-level feature map, and combine the cross-attention weights and the channel attention weights to obtain fusion attention weights;

[0101] Calculate the similarity pattern of local features based on the fusion attention weights to obtain a feature sampling offset, resample the first multi-level feature map and the second multi-level feature map according to the feature sampling offset to obtain resampled features, calculate the local density distribution matrix and the density jump matrix of the resampled features, and generate a target scattering probability map based on the local density distribution matrix and the density jump matrix;

[0102] Determine the scattering center position coordinates based on the maximum value position of the target scattering probability map, construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, and perform a convolution operation on the adaptive kernel function with the range dimension target echo and the angle dimension target echo to obtain the filtered echo data.

[0103] Exemplarily, first obtain the range dimension echo and the angle dimension echo data of the target. Assume that the range dimension echo is a vector containing 128 sampling points, and the angle dimension echo is a 16x16 matrix, where each element represents the echo intensity at different angles.

[0104] Next, perform feature extraction on the range dimension target echo. Input the range dimension echo into a set of phase-sensitive complex-valued filters, such as using 8 Gabor filters with different time scales and phase parameters. Each filter performs a convolution operation on the echo to obtain 8 different feature maps, forming the first feature layer. Calculate the conditional entropy of the first feature layer, such as using the Shannon entropy formula, to measure the redundancy between feature layers. At the same time, calculate the information gain rate of the first feature layer to evaluate the distinguishability of features. According to the conditional entropy and the information gain rate, adjust the time scale and phase parameters of the Gabor filters, such as using the gradient descent method to minimize the conditional entropy and maximize the information gain rate. Filter the range dimension echo again through the optimized filters to obtain the optimized first multi-level feature map. Assume that the optimized first multi-level feature map still contains 8 feature maps, and the size of each feature map is 128.

[0105] Then, perform feature extraction on the angle dimension target echo. Calculate the local region direction consistency of the angle dimension target echo, such as using the Sobel operator to calculate the gradient to obtain the direction response map. Assume that the size of the direction response map is 16x16. Extract the main direction and the secondary direction from the direction response map, such as selecting the direction with the largest gradient magnitude as the main direction and the direction orthogonal to it as the secondary direction. Use the main direction and the secondary direction as the reference directions of the direction-sensitive filter bank, such as using 8 Gabor filters with different angle resolution parameters, where 4 filters are aligned with the main direction and the other 4 filters are aligned with the secondary direction. Calculate the feature responses in the main direction and the secondary direction, and calculate the difference between them. According to the feature response difference, such as using the gradient descent method to minimize the feature response difference, optimize the angle resolution parameters of the direction-sensitive filter bank. Filter the angle dimension echo with the optimized filters to obtain the second multi-level feature map. Assume that the second multi-level feature map contains 8 feature maps, and the size of each feature map is 16x16.

[0106] Next, feature fusion is performed. Calculate the mutual information matrix between the first multi-level feature map and the second multi-level feature map, for example, using a histogram-based mutual information estimation method. Perform principal component analysis on the mutual information matrix to obtain the principal component vector, and use it as the cross-attention weight. At the same time, calculate the channel statistics of the first multi-level feature map and the second multi-level feature map, such as the mean and standard deviation, and use them to construct the channel attention weight, for example, using the Sigmoid function to normalize the statistics. Combine the cross-attention weight and the channel attention weight, for example, by weighted averaging, to obtain the fused attention weight.

[0107] Based on the fused attention weight, calculate the similarity pattern of the local features, for example, calculate the cosine similarity between the feature vectors to obtain the feature sampling offset. According to the feature sampling offset, resample the first multi-level feature map and the second multi-level feature map, for example, using bilinear interpolation. Calculate the local density distribution matrix and the density jump matrix of the resampled features, for example, using kernel density estimation and differential operations. Based on the local density distribution matrix and the density jump matrix, for example, multiply the two to generate the target scattering probability map.

[0108] Finally, echo filtering is performed. According to the maximum position of the target scattering probability map, determine the coordinates of the scattering center position. Construct an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates, for example, using a Gaussian kernel function, whose parameters are determined by the eigenvalue at the scattering center position coordinates. Convolve the adaptive kernel function with the range-dimensional target echo and the angle-dimensional target echo to obtain the filtered echo data.

[0109] In the prior art, during the process of target echo processing, a single feature extraction method or a method with fixed filtering parameters is usually adopted, resulting in inaccurate characterization of the target scattering characteristics, being easily interfered by noise, and affecting the accuracy of target recognition. In this application, a phase-sensitive complex-valued filter bank and a direction-sensitive filter bank are constructed to extract features from the target echoes in the range dimension and the angle dimension respectively, and the filtering parameters are optimized based on conditional entropy, information gain rate, and direction consistency to make the feature extraction more accurate. At the same time, cross-attention and channel-attention weights are introduced to achieve deep fusion of multi-level feature maps, fully utilize the information complementarity between different features, and improve the expression ability of target scattering features. During the generation process of the target scattering probability map, by calculating the local density distribution matrix and the density jump matrix, the distinguishability of the target scattering features is enhanced, making the target positioning more accurate. In addition, this application adopts an adaptive kernel function that dynamically adjusts according to the position of the scattering center to filter the echo data, enabling the filtering method to adapt to the changes in the target scattering characteristics, reducing the influence of noise, and improving the robustness of target detection. Compared with the prior art, this application can extract and fuse target scattering features more accurately, maintain a high detection accuracy in complex environments, and enhance the adaptability and robustness of the system.

[0110] In an optional implementation manner, calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius with the center position of the point cloud cluster as the benchmark, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features, including:

[0111] Construct a radial density distribution matrix of the three-dimensional spatial point cloud data, calculate the adaptive search radius of each point based on the gradient change of the radial density distribution matrix, extract adjacent point sets within the adaptive search radius, and calculate the distance difference and angle difference between point pairs in the adjacent point sets;

[0112] Construct a double-layer weight modulation function, where the first-layer weight coefficient of the double-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Input the distance difference and angle difference into the double-layer weight modulation function to obtain weight features, and construct a spatial distribution matrix based on the weight features;

[0113] Perform eigen-decomposition on the spatial distribution matrix to obtain the main eigen-subspace and the secondary eigen-subspace, map the main eigen-subspace and the secondary eigen-subspace to a topological feature network, calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and use the multi-dimensional feature vector to divide the three-dimensional spatial point cloud data into multiple point cloud clustering clusters;

[0114] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculate the density difference between any pair of points within the point cloud clustering cluster, determine the density increasing direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum, and determine the position of the local density maximum as the center position of the point cloud cluster;

[0115] Construct a deformation kernel function with the center position of the point cloud cluster as the reference point, calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameter of the kernel function, adjust the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, construct a feature distribution tensor within the range of the adaptive sampling radius, extract the main feature components of the feature distribution tensor, and combine the main feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud;

[0116] Construct a feature propagation network, input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vector of the feature points, and perform manifold constraint optimization on the motion vector of the feature points to obtain the pose parameters of the target object.

[0117] Exemplarily, first obtain the three-dimensional point cloud data and construct a radial density distribution matrix. Calculate the average distance from each point to other points within a certain range around it, and use this as the radial density value of the point. Store the radial density values of all points in a matrix to form a radial density distribution matrix. For example, in a point cloud data containing 1000 points, the average distance from each point to its 10 nearest neighbor points can be calculated, and these 1000 average distance values are stored in a 1000x1 matrix, which is the radial density distribution matrix.

[0118] Calculate the adaptive search radius of each point according to the gradient change of the radial density distribution matrix. The greater the gradient change, the smaller the search radius, and vice versa. For example, if the radial density value of a certain point differs greatly from the radial density values of its surrounding points, the search radius of this point will be smaller; conversely, if the radial density value of this point differs little from the radial density values of its surrounding points, the search radius of this point will be larger.

[0119] Extract the adjacent point set within the adaptive search radius, and calculate the distance difference and angle difference between point pairs in the adjacent point set. Assume that the adaptive search radius of a certain point is 0.1, then the adjacent point set of this point is all points whose distance from it is less than 0.1. For any two points in the adjacent point set, calculate the Euclidean distance between them and the included angle formed by their connection lines with the center point to obtain the distance difference and angle difference.

[0120] Construct a two - layer weight modulation function. The weight coefficients of the first layer are determined by the change rate of the radial density distribution matrix. The larger the change rate, the larger the weight coefficient. The weight coefficients of the second layer are determined by the distribution consistency of the angular differences. The more consistent the distribution, the larger the weight coefficient. For example, if the points around a point are concentrated in one direction, the distribution consistency of the angular differences is relatively high, and the corresponding weight coefficient is also large. Input the distance difference and the angular difference into the two - layer weight modulation function to obtain weight features.

[0121] Construct a spatial distribution matrix based on the weight features. Store the weight features of all point pairs in a matrix to form a spatial distribution matrix. For example, in a point cloud data containing 1000 points, the dimension of the spatial distribution matrix is 1000x1000, where each element represents the weight feature between the corresponding two points.

[0122] Perform eigen - decomposition on the spatial distribution matrix to obtain the principal eigen - subspace and the secondary eigen - subspace. Map the principal eigen - subspace and the secondary eigen - subspace to the topological feature network.

[0123] Calculate the hierarchical entropy of the topological feature network to obtain a multi - dimensional feature vector. Use the multi - dimensional feature vector to divide the 3D spatial point cloud data into multiple point cloud clustering clusters.

[0124] Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi - dimensional feature vector and the radial density distribution matrix into the density gradient optimization function. Calculate the density difference between any point pairs within the point cloud clustering cluster, determine the density increasing direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum, and determine the position of the local density maximum as the center position of the point cloud cluster.

[0125] Taking the center position of the point cloud cluster as a reference point, construct a deformation kernel function. Calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameters of the kernel function. Adjust the shape of the kernel function according to the deformation parameters to obtain an adaptive sampling radius. Construct a feature distribution tensor within the range of the adaptive sampling radius, extract the principal feature components of the feature distribution tensor, and combine the principal feature components with the multi - dimensional feature vector to obtain the dynamic features of the point cloud.

[0126] Construct a feature propagation network and input the dynamic features of the point cloud into the feature propagation network. Identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vectors of the feature points. Perform manifold - constraint optimization on the motion vectors of the feature points to obtain the attitude parameters of the target object.

[0127] In this embodiment, through the adaptive search radius and the double-layer weight modulation function, the point cloud features can be extracted more accurately, thereby improving the accuracy of subsequent motion estimation and attitude parameter calculation. By using the density gradient optimization function to determine the center position of the point cloud cluster, the influence of noise can be effectively avoided, and the robustness of the algorithm can be enhanced. Through the feature propagation network and the manifold constraint optimization, the attitude parameters of the target object can be accurately estimated, and high accuracy can be maintained even when the target object performs complex motions.

[0128] In an alternative embodiment, a feature propagation network is constructed. The dynamic features of the point cloud are input into the feature propagation network to identify the spatial positions of the corresponding feature points in consecutive time frames, calculate the displacements between the corresponding feature points to obtain the feature point motion vectors, and perform manifold constraint optimization on the feature point motion vectors to obtain the attitude parameters of the target object, including:

[0129] Obtain the dynamic features of the point cloud in consecutive time frames, and extract local shape descriptors and global descriptors representing the regional distribution from the dynamic features of the point cloud; perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features; divide the enhanced dynamic features into a current frame feature set and a subsequent frame feature set according to the time frame;

[0130] Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distances and feature similarities between feature points at each sampling level; perform dynamic grouping on the feature points based on the geometric distances and feature similarities to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure;

[0131] Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal variation; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fusion feature points;

[0132] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fusion feature points;

[0133] Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints; combine the temporal constraint and the cross-scale feature point matching constraints into an objective function;

[0134] Optimize the objective function successively at multiple sampling levels to obtain the feature point matching relationship; based on the feature point matching relationship, pair the three-dimensional spatial coordinates of the matching feature points under their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector.

[0135] Construct a manifold space representation of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and take the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and take the convergence result as the pose parameter of the target object.

[0136] A target pose estimation method based on a feature propagation network for identifying the target pose from dynamic point cloud data. This method constructs a feature propagation network to capture the spatio-temporal features of the point cloud and combines manifold constraint optimization to achieve accurate target pose estimation.

[0137] First, obtain the dynamic features of the point cloud in consecutive time frames. For example, in the first-frame point cloud, each point has three-dimensional coordinates (x, y, z) and a reflection intensity value. By calculating the geometric relationship and feature difference between each point and its neighboring points, extract local shape descriptors, such as local curvature and normal vector. At the same time, statistically analyze the overall distribution of the point cloud and extract global descriptors, such as the centroid and bounding box size of the point cloud. Then, perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features. For example, according to the discriminability of the local descriptors and the stability of the global descriptors, assign different weights and linearly combine them to obtain the enhanced features.

[0138] Next, divide the enhanced dynamic features into the current-frame feature set and the subsequent-frame feature set according to the time frame. For example, take the enhanced features of the first frame as the current-frame feature set and the enhanced features of the second frame as the subsequent-frame feature set. Perform multi-level downsampling on the enhanced dynamic features. For example, use the voxel grid downsampling method to divide the point cloud into multiple voxels and select a representative point within each voxel. At each sampling level, calculate the geometric distance between feature points, such as the Euclidean distance, and the feature similarity, such as the cosine similarity. Dynamically group the feature points based on the geometric distance and feature similarity. For example, use the DBSCAN clustering algorithm to group points that are close in distance and similar in features into the same group. Establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure. For example, connect a feature point group at a coarser level to all its subgroups at a finer level.

[0139] Based on the constructed hierarchical feature point structure, calculate the association strength of feature points in the dimensions of spatial position, feature expression, and temporal variation. For example, the spatial position association strength can be calculated by the Euclidean distance between feature points, the feature expression association strength can be calculated by the cosine similarity between feature vectors, and the temporal variation association strength can be calculated by the displacement magnitude of feature points between consecutive frames. Convert these association strengths into feature transfer weights. For example, use the softmax function to normalize the association strength between 0 and 1. Utilize the feature transfer weights to perform information interaction among feature points at different sampling levels to generate multi-scale fusion feature points. For example, weight and transfer the feature information at a coarser level to a finer level, and weight and transfer the feature information at a finer level to a coarser level.

[0140] Combine the multi-scale fusion feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points. For example, use the Kalman filter to predict the position of a feature point in the next frame based on its motion trajectory in the past few frames. Calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint. For example, calculate the Euclidean distance between the predicted position and the actual position. Calculate the feature similarity matrix for the multi-scale fusion feature points. For example, calculate the cosine similarity between each pair of feature points. Transfer the feature similarity matrix between different sampling levels to establish cross-scale feature point matching constraints. For example, transfer the feature similarity matrix at a coarser level to a finer level to guide the feature point matching at the finer level. Combine the temporal constraint and the cross-scale feature point matching constraint into an objective function. For example, use the weighted sum of the temporal constraint and the cross-scale matching constraint as the objective function.

[0141] Optimize the objective function successively at multiple sampling levels to obtain the feature point matching relationship. For example, use the iterative closest point algorithm to minimize the objective function and find the best feature point matching relationship. Based on the feature point matching relationship, pair the three-dimensional spatial coordinates of the matching feature points under their respective time frames. Calculate the difference in Euclidean distance of the paired coordinates to obtain the feature point displacement. Divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector. Suppose there are two matching points, with coordinates (1, 2, 3) in the first frame and (2, 3, 4) in the second frame, and the time interval is 0.1 seconds. Then the feature point displacement is (1, 1, 1), and the feature point motion vector is (10, 10, 10).

[0142] Construct a manifold space representation of the feature point motion vectors. For example, use Lie algebra to represent rotational motion. Calculate the orthogonal projection of the feature point motion vectors onto the manifold space. For example, project the motion vectors onto the rotation plane. Take the difference between the orthogonal projection and the original motion vectors as the projection error. Iteratively optimize the weighted sum of squares of the projection error until convergence. For example, use the gradient descent method to minimize the projection error. Take the convergence result as the pose parameters of the target object. For example, convert the final rotation matrix into Euler angles or quaternion to represent the pose.

[0143] In this embodiment, through multi-scale feature fusion and manifold constraint optimization, the spatio-temporal features and motion laws of the point cloud are effectively captured, thereby improving the accuracy of pose estimation. It has strong robustness to noise and occlusion and can stably estimate the target pose in complex scenarios. The design of the hierarchical feature point structure and the feature propagation network reduces the computational complexity and improves the efficiency of pose estimation.

[0144] In an alternative embodiment, input the motion vectors and pose parameters into a dual threshold decision module. Calculate the speed threshold value and the pose threshold value respectively through the dual threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area, and a person position heat map is generated based on the center position of the point cloud cluster. The output of the person perception result based on the person position heat map includes:

[0145] Sample the motion vectors within a sliding time window of a specified length, and construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vectors according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a time series change model for the pose parameters, extract the mean and standard deviation of the pose parameters based on the time series change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value;

[0146] Compare the motion vector with the speed threshold value, and compare the pose parameter with the pose threshold value; when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, it is determined that there is a person in the target area; use the region growing algorithm to segment the point cloud of the target area to obtain the target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain the filtered point cloud cluster;

[0147] Calculate the centroid of the filtered point cloud clusters to obtain the center position coordinates of the point cloud clusters, and map the center position coordinates of the point cloud clusters to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinates; perform kernel density estimation on the two-dimensional ground plane coordinates using a Gaussian kernel function with an adaptive bandwidth to obtain the spatial density distribution; calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud clusters, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map;

[0148] Construct a recursive filter to perform temporal filtering on the initial heat map to obtain the filtered heat map; calculate the fusion weight based on the similarity between the filtered heat map and the heat map of the previous moment; apply the fusion weight to the filtered heat map and the heat map of the previous moment to obtain the personnel position heat map; apply the local maximum suppression algorithm to the personnel position heat map to extract the extreme points of the heat map; use the extreme points of the heat map as candidate personnel positions;

[0149] Perform inter-frame matching on the candidate personnel positions to obtain the initial personnel trajectory, and use a Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory; construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the personnel movement trajectory; use the multi-feature fusion model to calculate the confidence score of each candidate personnel position; set a dynamic threshold to screen the confidence scores, and output the candidate personnel positions that meet the dynamic threshold and their corresponding personnel movement trajectories as the personnel perception results.

[0150] A personnel perception method aims to accurately identify and track personnel from three-dimensional point cloud data. This method uses motion information, pose features, and point cloud processing techniques to generate a personnel position heat map, and combines multi-feature fusion and trajectory prediction to achieve reliable perception of personnel.

[0151] First, in the preprocessing stage, perform denoising and filtering on the input three-dimensional point cloud data to remove the influence of environmental noise and stray points, so as to improve the accuracy of subsequent processing steps. For example, use a statistical filter to remove outliers and a bilateral filter to smooth the point cloud data while retaining edge information.

[0152] Next, perform motion estimation on the preprocessed point cloud data. Within a sliding time window of a specified length (e.g., 1 second), sample the motion vectors of each point. Suppose 10 motion vectors are sampled within a time window, with magnitudes of 0.1 m / s, 0.2 m / s, ..., 1.0 m / s respectively. Based on these sampled data, construct a motion vector distribution model, such as a Gaussian distribution model. According to this model, calculate the statistical mean (0.55 m / s in this example) and standard deviation (e.g., 0.28 m / s) of the motion vectors. Determine the weighted combination of the statistical mean and standard deviation (e.g., mean weight is 0.7, standard deviation weight is 0.3) as the velocity threshold value, such as 0.47 m / s.

[0153] Meanwhile, establish a time series change model for the pose parameters. For example, assume the pose parameter is the human height. Within a time window, sample the human height and establish a time series model. Based on this model, extract the mean and standard deviation of the pose parameters. Similar to the calculation method of the velocity threshold value, determine the weighted combination of the mean and standard deviation as the pose threshold value. For example, assume the mean is 1.75 m, the standard deviation is 0.1 m, and the pose threshold value after weighted combination is 1.72 m.

[0154] Then, compare the motion vector of each point with the velocity threshold value, and compare the pose parameter with the pose threshold value. When the motion vector of a point is greater than the velocity threshold value (e.g., 0.5 m / s is greater than 0.47 m / s) and the pose parameter is greater than the pose threshold value (e.g., 1.8 m is greater than 1.72 m), then determine that this point belongs to the target area, indicating that there may be a person.

[0155] Next, use the region growing algorithm to segment the point cloud of the target area to obtain the target point cloud cluster. Perform spatial filtering and noise removal on the target point cloud cluster to obtain the filtered point cloud cluster. Calculate the centroid of the filtered point cloud cluster to obtain the coordinates of the center position of the point cloud cluster. Map the coordinates of the center position of the point cloud cluster to the ground plane coordinate system through the projection matrix to obtain the two-dimensional ground plane coordinates. For example, assume the three-dimensional coordinates of the center of the point cloud cluster are (1, 2, 3), and map it to the ground plane coordinate system through the projection matrix to obtain the two-dimensional coordinates (1, 2).

[0156] Use a Gaussian kernel function with an adaptive bandwidth to perform kernel density estimation on the two-dimensional ground plane coordinates to obtain the spatial density distribution. Calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud cluster. Perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map.

[0157] Construct a recursive filter to perform temporal filtering on the initial heatmap to obtain a filtered heatmap. Calculate the fusion weight based on the similarity between the filtered heatmap and the heatmap of the previous moment. Apply the fusion weight to the filtered heatmap and the heatmap of the previous moment to obtain the personnel position heatmap. Apply the local maximum suppression algorithm on the personnel position heatmap to extract the extreme points of the heatmap. Use the extreme points of the heatmap as candidate points for personnel positions.

[0158] Perform inter-frame matching on the candidate points for personnel positions to obtain the initial personnel trajectory. Use the Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory. Construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the personnel movement trajectory. Use the multi-feature fusion model to calculate the confidence score for each candidate point for personnel positions. Set a dynamic threshold to screen the confidence scores, and output the candidate points for personnel positions that meet the dynamic threshold and their corresponding personnel movement trajectories as the personnel perception result.

[0159] In this embodiment, by comprehensively using motion information, pose features, and point cloud processing technology, personnel can be effectively distinguished from other objects, reducing the false detection rate. By using temporal filtering and multi-feature fusion technology, complex scenarios such as noise and occlusion can be effectively processed, improving the stability of personnel perception. By generating the personnel position heatmap and combining the Kalman filter for trajectory prediction, accurate positioning and continuous tracking of personnel can be achieved.

[0160] In an alternative implementation, performing inter-frame matching on the candidate points for personnel positions to obtain the initial personnel trajectory, and using the Kalman filter to smooth and predict the initial personnel trajectory to obtain the personnel movement trajectory includes:

[0161] Obtain the candidate points for personnel positions in consecutive time frames; calculate the Euclidean distance between the candidate points for personnel positions in adjacent time frames to obtain the distance feature; calculate the motion speed based on the historical position information of the candidate points for personnel positions, and calculate the speed difference between the candidate points for personnel positions in adjacent time frames based on the motion speed to obtain the speed feature; calculate the motion direction based on the motion speed, and calculate the direction difference between the candidate points for personnel positions in adjacent time frames based on the motion direction to obtain the direction feature;

[0162] Construct a matching cost matrix from the distance feature, speed feature, and direction feature; use the candidate points for personnel positions in the current time frame as the target point set, and use the candidate points for personnel positions in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain the matching point pairs; connect the matching point pairs in chronological order to obtain the initial personnel trajectory;

[0163] Form a state vector from the position coordinates, motion speed, and acceleration in the initial trajectory of the person; construct a state transition equation based on the state vector; construct an observation equation based on the position coordinates of the initial trajectory of the person; calculate the prediction error between the predicted value of the state transition equation and the actual observed value; update the process noise covariance according to the prediction error;

[0164] Use the state transition equation and the process noise covariance to predict the state vector to obtain a state prediction value and a prediction covariance; calculate the Kalman gain according to the state prediction value, the prediction covariance, and the observation equation; use the Kalman gain to correct the state prediction value to obtain a state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value;

[0165] Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract the historical motion pattern within the prediction window length; use the historical motion pattern to repair the trajectory break points;

[0166] Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than a preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain an optimized person motion trajectory.

[0167] A method for generating a person motion trajectory based on multi-feature matching and Kalman filtering, which can effectively handle complex scenarios such as occlusion and noise, and generate a smooth, continuous, and accurate person motion trajectory.

[0168] First, obtain the candidate points of the person's position in consecutive time frames. For example, use an object detection algorithm to detect the person target in each frame image of the video, obtain the bounding box coordinates of each person, and use the center point of the bounding box as the candidate point of the person's position. Assume that 3 people are detected in the t-th frame, and their position coordinates are (10, 20), (30, 40), and (50, 60) respectively.

[0169] Next, calculate the Euclidean distance, speed difference, and direction difference between the candidate points of the person's position in adjacent time frames to construct a multi-feature matching cost matrix. For example, assume that the position of a person in the (t-1)-th frame is (8, 18), and calculate its Euclidean distances from the three candidate points (10, 20), (30, 40), and (50, 60) in the t-th frame, which are 2.83, 22.83, and 42.83 respectively. At the same time, calculate the movement speed and direction based on the historical position information. Assume that the speed of this person in the (t-1)-th frame is (2, 2) and the direction is 45 degrees, then calculate the speed difference and direction difference between it and the three candidate points in the t-th frame. Combine the distance, speed difference, and direction difference with weights to construct the matching cost matrix.

[0170] Then, construct a bipartite graph based on the matching cost matrix and use the Hungarian algorithm to perform minimum-cost matching to obtain the matching point pairs. For example, assume that the cost matrix indicates that the matching cost between (8, 18) and (10, 20) is the smallest, then connect them to form a matching point pair. Connect all the matching point pairs in chronological order to obtain the initial trajectory of the person.

[0171] Next, form a state vector from the position coordinates, movement speed, and acceleration in the initial trajectory of the person and construct a Kalman filter. According to the kinematic model, construct a state transition equation to describe the variation law of the state vector over time. At the same time, construct an observation equation to describe the relationship between the observed value (i.e., the person's position coordinates) and the state vector. Use the Kalman filter to smooth and predict the initial trajectory. For example, calculate the prediction error between the predicted value of the state transition equation and the actual observed value based on the initial trajectory, and update the process noise covariance accordingly. Then, use the updated state transition equation and process noise covariance to predict the state vector to obtain the state prediction value and prediction covariance. Finally, correct the state prediction value using the Kalman gain to obtain the state estimate value, and generate a smooth trajectory based on the position coordinates in the state estimate value.

[0172] Detect the continuity of the smooth trajectory and identify the trajectory break points. Statistically calculate the effective trajectory length before the trajectory break point and determine the prediction window length based on this length. Extract the historical motion pattern within the prediction window length and use this pattern to repair the trajectory break point. For example, if a person's trajectory breaks at the 10th frame and the effective trajectory length before that is 9 frames, then the prediction window length can be set to 9 frames, extract the motion pattern of this person in the previous 9 frames, and use this pattern to predict their position after the 10th frame to repair the trajectory break.

[0173] Calculate the position similarity and velocity similarity between the repaired trajectories, and use the weighted combination of the two as the trajectory similarity. When the trajectory similarity is greater than the preset threshold, calculate the reliability weight according to the number of valid observations of the trajectory, and use this weight to merge the similar trajectories to obtain the optimized personnel movement trajectory. For example, if two repaired trajectories are very similar in terms of position and velocity, they can be merged into one trajectory.

[0174] Existing methods for generating personnel movement trajectories mainly rely on single feature matching, which is difficult to ensure the continuity and accuracy of trajectories in complex scenarios and is vulnerable to occlusion and noise interference, resulting in trajectory breaks or incorrect matches. In addition, traditional trajectory repair mostly uses simple interpolation, lacking the utilization of individual movement patterns, with low prediction accuracy, and there is also a lack of effective similarity measurement for trajectory merging, affecting the quality of the final trajectory. In this application, through a multi-feature matching method, a matching cost matrix is comprehensively constructed using distance, velocity, and direction features, and the Hungarian algorithm is used to achieve optimal matching, improving the accuracy of inter-frame matching and reducing incorrect associations. At the same time, a Kalman filter is introduced to construct a state transition equation and an observation equation to smooth and predict the trajectory, making the trajectory more stable and continuous and reducing noise interference. In terms of trajectory repair, this application detects trajectory breakpoints, extracts historical movement patterns, and completes the trajectory within the prediction window, making the repaired trajectory more in line with the actual movement trend. For trajectory optimization, the position similarity and velocity similarity are comprehensively calculated, and a reliability weight is introduced for trajectory merging to avoid incorrect merging problems caused by single features. Compared with the prior art, this application has been improved in terms of matching accuracy, trajectory smoothing, trajectory repair, and trajectory optimization, making the finally generated personnel movement trajectory smoother, more accurate, and more coherent, capable of effectively dealing with complex scenarios such as occlusion and noise, and improving the robustness and reliability of trajectory generation.

[0175] Figure 2 This is a scatter plot of the feature extraction performance for the embodiments of the present invention. As Figure 2 shown, the scatter plot of the feature extraction performance shows the distribution of different methods in terms of feature dimension and computational complexity. This technical solution (circular markers) achieves a feature extraction accuracy of 0.95 at 128 feature dimensions, with a computational complexity of O(nlog n ); the traditional PCA method (square markers) has an accuracy of 0.82 at the same feature dimension, with a complexity of O(n²); the ICA algorithm (triangular markers) has an accuracy of 0.87, and its complexity is between the two. By introducing a multi-level feature mapping and an attention mechanism, this technical solution significantly improves the accuracy of feature extraction while maintaining a low computational complexity, especially showing a better clustering effect in the high-dimensional feature space, and the discrimination between features has increased by about 15%. The distribution density of the scatter points also indicates that the feature extraction results of this technical solution are more stable, and the variance has decreased by about 25%.

[0176] Figure 3 This is the comparison effect diagram of point cloud cluster segmentation based on temporal filtering for the features of the embodiments of the present invention. As Figure 3 shown, the comparison diagram of point cloud cluster segmentation shows the performance of different segmentation methods in 100 consecutive frames. The technical solution (circular marker) achieves a segmentation accuracy of 95.2% through temporal filtering and region growing algorithm, with a noise rate of 2.8% and an average processing time of 15 ms / frame. The segmentation accuracy of the traditional region growing algorithm (square marker) is 82.5%, the noise rate is 8.6%, and the processing time is 25 ms / frame. The segmentation accuracy of the Euclidean clustering method (triangle marker) is 85.3%, the noise rate is 7.2%, and the processing time is 22 ms / frame. In a scenario with a personnel density of 3 people per square meter, the missed detection rate of this solution is only 1.5%, significantly lower than 4.8% and 4.2% of the traditional methods. By introducing temporal information and adaptive parameter adjustment, this solution has achieved obvious improvements in both the accuracy and real-time performance of point cloud segmentation.

[0177] Figure 4 This is the structural schematic diagram of a lightweight AI personnel perception system based on a millimeter-wave radar for the embodiments of the present invention. As Figure 4 shown, the system includes:

[0178] The first unit is used to collect the millimeter-wave echo signals of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, perform two-way joint filtering processing on the calibrated signals according to the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional space point cloud data based on the filtered echo data;

[0179] The second unit is used to calculate the distance difference and angle difference between adjacent points in the three-dimensional space point cloud data, construct a spatial distribution matrix according to the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features;

[0180] The third unit is used to input the motion vector and attitude parameters into a dual-threshold decision module, calculate the speed threshold value and attitude threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are personnel in the target area, generate a personnel position heat map according to the center position of the point cloud cluster, and output the personnel perception result based on the personnel position heat map.

[0181] In the third aspect of the embodiments of the present invention,

[0182] Provided is an electronic device, comprising:

[0183] a processor;

[0184] a memory for storing instructions executable by the processor;

[0185] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0186] In a fourth aspect of the embodiments of the present invention,

[0187] provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0188] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for implementing various aspects of the present invention are loaded.

[0189] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lightweight AI personnel perception method based on millimeter-wave radar, characterized in that, Including: Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibration signal, perform two-way joint filtering processing on the calibration signal in the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data; Calculate the distance difference and angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate the motion vector and attitude parameters based on the extracted point cloud features; Input the motion vector and attitude parameters into a dual-threshold decision module, calculate the speed threshold value and attitude threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the attitude parameter is greater than the attitude threshold value, it is determined that there are people in the target area, and a personnel position heat map is generated according to the center position of the point cloud cluster, and a personnel perception result is output based on the personnel position heat map.

2. The method according to claim 1, wherein Collect the millimeter-wave echo signal of the target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signal to obtain a calibration signal, perform two-way joint filtering processing on the calibration signal in the distance dimension and the angle dimension to obtain filtered echo data, and constructing three-dimensional spatial point cloud data according to the filtered echo data includes: Collect the millimeter-wave echo signal of the target area, calculate the phase difference and amplitude ratio between adjacent antennas of the millimeter-wave echo signal, map the phase difference to a dynamic self-correction coefficient matrix, map the amplitude ratio to a non-linear compensation coefficient curve, and perform matrix multiplication operations on the millimeter-wave echo signal with the dynamic self-correction coefficient matrix and the non-linear compensation coefficient curve respectively to obtain a calibration signal; Perform wavelet packet decomposition on the calibration signal to obtain multi-scale signal components, extract the energy entropy feature vector of the multi-scale signal components, calculate the adaptive reconstruction weight of each scale signal component based on the energy entropy feature vector, multiply the adaptive reconstruction weight by the corresponding scale signal component and sum them to obtain an enhanced signal; Extract local maximum points on the time-frequency energy spectrum of the enhanced signal as peak points, calculate the spatial clustering distribution characteristics of the peak points, generate a threshold surface based on the spatial clustering distribution characteristics, use the threshold surface to extract the distance-dimensional target echo, construct an orthogonal projection operator with the covariance matrix eigenvector of the distance-dimensional target echo, and project the calibration signal onto the null space of the orthogonal projection operator to obtain the angle-dimensional target echo; Extract multi-level feature maps from the distance-dimensional target echo and the angle-dimensional target echo, input the multi-level feature maps into an attention fusion network to generate a target scattering probability map, and extract the filtered echo data according to the maximum value position of the target scattering probability map; The filtered echo data is mapped through polar coordinates to obtain an initial point cloud, the local density gradient of the initial point cloud is calculated, a diffusion equation of a point cloud growth model is established according to the local density gradient, and the diffusion equation is solved to obtain an enhanced point cloud; The principal curvature and normal vector of the enhanced point cloud are calculated, and an affinity matrix is constructed using the principal curvature and normal vector for spectral clustering to obtain three-dimensional space point cloud data.

3. The method according to claim 2, wherein Feature extraction is performed on the range-dimensional target echo and the angle-dimensional target echo to obtain a multi-level feature map. The multi-level feature map is input into an attention fusion network to generate a target scattering probability map. Extracting the filtered echo data according to the maximum value position of the target scattering probability map includes: The range-dimensional target echo is input into a phase-sensitive complex-valued filter bank to obtain a first feature layer. The conditional entropy of the first feature layer is calculated to obtain the redundancy between feature layers, and the information gain rate of the first feature layer is calculated to obtain the feature discrimination degree. Based on the redundancy between feature layers and the feature discrimination degree, the time scale parameter and phase parameter of the phase-sensitive complex-valued filter bank are optimized to obtain a first multi-level feature map; The angle-dimensional target echo is input into a direction-sensitive filter bank. The local region direction consistency of the angle-dimensional target echo is calculated to obtain a direction response map. The main direction and secondary direction are extracted based on the direction response map. The main direction and secondary direction are used as the reference directions of the direction-sensitive filter bank, and the feature response differences in the main direction and secondary direction are calculated. Based on the feature response differences, the angle resolution parameter of the direction-sensitive filter bank is optimized to obtain a second multi-level feature map; The mutual information matrix is calculated for the first multi-level feature map and the second multi-level feature map. The cross-attention weight is constructed based on the principal component vector of the mutual information matrix, and the channel attention weight is obtained based on the channel statistics of the first multi-level feature map and the second multi-level feature map. The cross-attention weight and the channel attention weight are combined to obtain a fusion attention weight; Based on the fusion attention weight, the similarity pattern of local features is calculated to obtain a feature sampling offset. The first multi-level feature map and the second multi-level feature map are resampled according to the feature sampling offset to obtain resampled features. The local density distribution matrix and density jump matrix of the resampled features are calculated, and a target scattering probability map is generated based on the local density distribution matrix and the density jump matrix; Based on the maximum value position of the target scattering probability map, the scattering center position coordinates are determined, an adaptive kernel function that dynamically adjusts with the local features at the scattering center position coordinates is constructed, and the adaptive kernel function is convolved with the range-dimensional target echo and the angle-dimensional target echo to obtain filtered echo data.

4. The method according to claim 1, wherein The distance difference and angle difference between adjacent points in the three-dimensional space point cloud data are calculated. A spatial distribution matrix is constructed according to the distance difference and angle difference. The point cloud cluster center position is calculated based on the spatial distribution matrix. An adaptive sampling radius is established with the point cloud cluster center position as the reference. Continuously varying point cloud features are extracted within the adaptive sampling radius, and the motion vector and attitude parameters are calculated based on the extracted point cloud features, including: Construct a radial density distribution matrix of 3D spatial point cloud data, calculate the adaptive search radius of each point based on the gradient change of the radial density distribution matrix, extract adjacent point sets within the adaptive search radius, and calculate the distance difference and angle difference between point pairs in the adjacent point sets; Construct a double-layer weight modulation function. The first-layer weight coefficient of the double-layer weight modulation function is determined by the change rate of the radial density distribution matrix, and the second-layer weight coefficient is determined by the distribution consistency of the angle difference. Input the distance difference and angle difference into the double-layer weight modulation function to obtain weight features, and construct a spatial distribution matrix based on the weight features; Perform eigen-decomposition on the spatial distribution matrix to obtain the principal eigen-subspace and the secondary eigen-subspace. Map the principal eigen-subspace and the secondary eigen-subspace to a topological feature network, calculate the hierarchical entropy of the topological feature network to obtain a multi-dimensional feature vector, and use the multi-dimensional feature vector to divide the 3D spatial point cloud data into multiple point cloud clustering clusters; Construct a density gradient optimization function within each point cloud clustering cluster. Input the multi-dimensional feature vector and the radial density distribution matrix into the density gradient optimization function, calculate the density difference between any point pairs within the point cloud clustering cluster, determine the density rising direction based on the density difference, and iteratively update the search position until convergence to obtain the position of the local density maximum, and determine the position of the local density maximum as the center position of the point cloud cluster; Construct a deformation kernel function with the center position of the point cloud cluster as the reference point, calculate the local curvature distribution within the neighborhood of the reference point, map the local curvature distribution to the deformation parameter of the kernel function, adjust the shape of the kernel function according to the deformation parameter to obtain an adaptive sampling radius, construct a feature distribution tensor within the adaptive sampling radius range, extract the principal feature components of the feature distribution tensor, and combine the principal feature components with the multi-dimensional feature vector to obtain the dynamic features of the point cloud; Construct a feature propagation network, input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vectors of the feature points, and perform manifold constraint optimization on the motion vectors of the feature points to obtain the pose parameters of the target object.

5. The method according to claim 4, wherein Construct a feature propagation network, input the dynamic features of the point cloud into the feature propagation network, identify the spatial positions of corresponding feature points in consecutive time frames, calculate the displacement between the corresponding feature points to obtain the motion vectors of the feature points, and perform manifold constraint optimization on the motion vectors of the feature points to obtain the pose parameters of the target object, including: Obtain the dynamic features of the point cloud in consecutive time frames, extract local shape descriptors and global descriptors representing regional distributions from the dynamic features of the point cloud; perform adaptive weighted fusion on the local shape descriptors and global descriptors to obtain enhanced dynamic features; divide the enhanced dynamic features into a current frame feature set and a subsequent frame feature set according to time frames; Perform multi-level downsampling on the enhanced dynamic features, calculate the geometric distance and feature similarity between feature points at each sampling level; dynamically group the feature points based on the geometric distance and feature similarity to obtain multi-level feature point groups; establish a transfer relationship between the feature point groups at adjacent sampling levels to generate a hierarchical feature point structure; Based on the hierarchical feature point structure, calculate the association strength of the feature points in the dimensions of spatial position, feature expression, and temporal change; convert the association strength into a feature transfer weight; use the feature transfer weight to perform information interaction on the feature points at different sampling levels to generate multi-scale fused feature points; Combine the multi-scale fused feature points with the historical motion sequence of the feature points to predict the motion trend of the feature points; calculate the deviation value between the predicted motion trend and the actual feature point distribution as the temporal constraint; calculate the feature similarity matrix for the multi-scale fused feature points; Transfer the feature similarity matrix between different sampling levels to establish a cross-scale feature point matching constraint; combine the temporal constraint and the cross-scale feature point matching constraint into an objective function; Optimize the objective function sequentially at multiple sampling levels to obtain a feature point matching relationship; based on the feature point matching relationship, correspond and pair the three-dimensional spatial coordinates of the matching feature points in their respective time frames; calculate the Euclidean distance difference of the paired coordinates to obtain the feature point displacement; divide the feature point displacement by the corresponding time interval to obtain the feature point motion vector; Construct a manifold space expression of the feature point motion vector, calculate the orthogonal projection of the feature point motion vector onto the manifold space, and use the difference between the orthogonal projection and the original motion vector as the projection error; iteratively optimize the weighted sum of squares of the projection error until convergence, and use the convergence result as the pose parameter of the target object.

6. The method according to claim 1, wherein Input the motion vector and the pose parameter into a dual-threshold decision module, calculate the speed threshold value and the pose threshold value respectively through the dual-threshold decision module. When the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, determine that there is a person in the target area, generate a personnel position heat map based on the center position of the point cloud cluster, and output the personnel perception result based on the personnel position heat map, including: Sample the motion vector within a sliding time window of a specified length, construct a motion vector distribution model based on the sampled data; calculate the statistical mean and standard deviation of the motion vector according to the motion vector distribution model, and determine the weighted combination of the statistical mean and the standard deviation as the speed threshold value; establish a temporal change model for the pose parameter, extract the mean and standard deviation of the pose parameter based on the temporal change model, and determine the weighted combination of the mean and the standard deviation as the pose threshold value; Compare the motion vector with the speed threshold value, and compare the pose parameter with the pose threshold value; when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, determine that there is a person in the target area; use the region growing algorithm to segment the point cloud in the target area to obtain a target point cloud cluster; perform spatial filtering and noise removal on the target point cloud cluster to obtain a filtered point cloud cluster; Calculate the centroid of the filtered point cloud clusters to obtain the center position coordinates of the point cloud clusters, and map the center position coordinates of the point cloud clusters to the ground plane coordinate system through a projection matrix to obtain the two-dimensional ground plane coordinates; perform kernel density estimation on the two-dimensional ground plane coordinates using a Gaussian kernel function with an adaptive bandwidth to obtain the spatial density distribution; calculate the confidence weight based on the spatial distribution characteristics and temporal consistency of the filtered point cloud clusters, and perform a convolution operation on the confidence weight and the spatial density distribution to obtain the initial heat map; Construct a recursive filter to perform temporal filtering on the initial heat map to obtain the filtered heat map; calculate the fusion weight based on the similarity between the filtered heat map and the heat map of the previous moment; apply the fusion weight to the filtered heat map and the heat map of the previous moment to obtain the person position heat map; apply the local maximum suppression algorithm to the person position heat map to extract the extreme points of the heat map; use the extreme points of the heat map as the person position candidate points; Perform inter-frame matching on the person position candidate points to obtain the initial person trajectory, and use a Kalman filter to smooth and predict the initial person trajectory to obtain the person movement trajectory; construct a multi-feature fusion model by combining the motion vector, pose parameters, and historical information of the person movement trajectory; calculate the confidence score of each person position candidate point using the multi-feature fusion model; set a dynamic threshold to screen the confidence scores, and output the person position candidate points that meet the dynamic threshold and their corresponding person movement trajectories as the person perception results.

7. The method according to claim 6, wherein Performing inter-frame matching on the person position candidate points to obtain the initial person trajectory, and using a Kalman filter to smooth and predict the initial person trajectory to obtain the person movement trajectory includes: Obtain the person position candidate points in consecutive time frames; calculate the Euclidean distance between the person position candidate points in adjacent time frames to obtain the distance feature; calculate the movement speed based on the historical position information of the person position candidate points, and calculate the speed difference between the person position candidate points in adjacent time frames based on the movement speed to obtain the speed feature; calculate the movement direction based on the movement speed, and calculate the direction difference between the person position candidate points in adjacent time frames based on the movement direction to obtain the direction feature; Construct a matching cost matrix from the distance feature, speed feature, and direction feature; use the person position candidate points in the current time frame as the target point set and the person position candidate points in the previous time frame as the source point set; construct a bipartite graph based on the matching cost matrix; perform minimum cost matching on the bipartite graph using the Hungarian algorithm to obtain the matching point pairs; connect the matching point pairs in chronological order to obtain the initial person trajectory; Form a state vector from the position coordinates, movement speed, and acceleration in the initial person trajectory; construct a state transition equation based on the state vector; construct an observation equation based on the position coordinates of the initial person trajectory; calculate the prediction error between the predicted value and the actual observation value of the state transition equation; update the process noise covariance according to the prediction error; Predict the state vector using the state transition equation and the process noise covariance to obtain a state prediction value and a prediction covariance; calculate the Kalman gain based on the state prediction value, the prediction covariance, and the observation equation; correct the state prediction value using the Kalman gain to obtain a state estimate value; generate a smooth trajectory based on the position coordinates in the state estimate value; Detect the continuity of the smooth trajectory to obtain trajectory break points; count the effective trajectory length before the trajectory break points; determine the prediction window length according to the effective trajectory length; extract the historical motion pattern within the prediction window length; use the historical motion pattern to repair the trajectory break points; Calculate the position distance between the repaired trajectories to obtain the position similarity; calculate the speed difference between the repaired trajectories to obtain the speed similarity; use the weighted combination of the position similarity and the speed similarity as the trajectory similarity; when the trajectory similarity is greater than a preset threshold, calculate the reliability weight according to the effective observation number of the trajectory; use the reliability weight to merge the similar trajectories to obtain an optimized personnel motion trajectory.

8. A lightweight AI personnel perception system based on millimeter-wave radar for implementing the method described in any one of the preceding claims 1-7, characterized in that, Comprising: A first unit, configured to collect millimeter-wave echo signals in a target area, perform phase compensation and amplitude calibration on the millimeter-wave echo signals to obtain calibrated signals, perform two-way joint filtering processing on the calibrated signals according to the distance dimension and the angle dimension to obtain filtered echo data, and construct three-dimensional spatial point cloud data according to the filtered echo data; A second unit, configured to calculate the distance difference and the angle difference between adjacent points in the three-dimensional spatial point cloud data, construct a spatial distribution matrix according to the distance difference and the angle difference, calculate the center position of the point cloud cluster based on the spatial distribution matrix, establish an adaptive sampling radius based on the center position of the point cloud cluster, extract continuously changing point cloud features within the adaptive sampling radius, and calculate a motion vector and pose parameters based on the extracted point cloud features; A third unit, configured to input the motion vector and the pose parameters into a dual-threshold decision module, calculate a speed threshold value and a pose threshold value respectively through the dual-threshold decision module, when the motion vector is greater than the speed threshold value and the pose parameter is greater than the pose threshold value, determine that there are personnel in the target area, generate a personnel position heat map according to the center position of the point cloud cluster, and output a personnel perception result based on the personnel position heat map.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Fall posture recognition method and system based on millimeter wave radar point cloud

    CN114942434A

  • Multi-person 5D radar falling detection method and system based on artificial intelligence algorithm

    CN117079416A

  • Human body tumble identification method based on millimeter wave radar perception

    CN119291630A

Cited By

  • Tunnel lining leakage infrared-millimeter wave fusion detection method and system

    CN120491045A

  • Real-time dynamic trajectory tracking method and system for millimeter wave radar gesture recognition

    CN120802203A

  • Point cloud data processing method, system and equipment of humanoid robot and medium

    CN120997630A

  • Millimeter wave radar personnel perception method based on time-frequency domain and deep CNN

    CN121165058A

  • Radar intelligent detection and tracking method based on target two-dimensional image

    CN121477186A