Key point recognition apparatus and method based on wireless radar signals
By using wireless radar signals and neural network technology, high-accuracy key point detection of various actions is achieved, solving the problem of limited application scenarios of existing radar detection technology and providing an easy-to-operate and privacy-protected solution.
Patent Information
- Application Number
- CN202110607536.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2041-06-01
AI Technical Summary
Existing radar-based key point detection technologies can only detect a small number of specific actions, limiting their application scenarios, and video technology performs poorly under privacy and environmental conditions.
A key point detection method based on wireless radar signals is adopted. Point cloud data is obtained by radar sensing objects, and feature extraction and cascading are performed. The fusion feature extraction and key point detection model of neural network is used to achieve accurate detection of various actions.
It achieves high-accuracy detection of various actions, requires few computing resources, is easy to operate, and has strong noise resistance and high privacy protection.
Smart Images

Figure CN115436894B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of radar detection technology. Background Technology
[0002] During the detection of human movements, key points of the human body can be detected, such as the head, neck, arms, feet, and waist. Human key point detection has a wide range of applications and is a key technology for applications such as smart homes, health monitoring, and behavior understanding.
[0003] Currently, video-based human keypoint detection technology is widely used. However, video severely infringes on privacy and cannot be applied in private situations. In addition, video-based keypoint detection is greatly affected by the environment (such as occlusion, lighting, smoke, etc.) and cannot function in dark or occluded scenes; it is also greatly affected by clothing, posture, and viewing angle.
[0004] Radar detects objects (such as the human body) via wireless signals, without exposing privacy, and is not dependent on external conditions such as lighting. It can also function normally in partially obscured environments. Therefore, radar-based key point detection can compensate for the shortcomings of video technology.
[0005] It should be noted that the above introduction to the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of this application and for the convenience of those skilled in the art to understand them. It should not be assumed that the above technical solutions are known to those skilled in the art simply because these solutions have been described in the background section of this application. Summary of the Invention
[0006] However, the inventors discovered that conventional radar-based key point detection can only detect a small number of specific actions, which greatly limits its application scenarios.
[0007] To address at least one of the aforementioned technical problems, embodiments of this application provide a key point detection device and method based on wireless radar signals. This method detects key points of objects (e.g., the human body) based on radar point clouds, without limiting the type of action, requiring minimal computational resources, and achieving high detection accuracy.
[0008] According to one aspect of the embodiments of this application, a key point detection device based on wireless radar signals is provided, comprising:
[0009] The sensing unit uses radar to sense objects and obtain point cloud data.
[0010] The feature extraction unit extracts features from the point cloud data obtained over a period of time, and obtains the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud.
[0011] A cascade unit that cascades the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data;
[0012] The feature detection unit utilizes a neural network-based fusion feature extraction model to detect the cascaded feature data to obtain fused feature information; and
[0013] The key point detection unit uses a neural network-based key point detection model to detect the fused feature information and output the key point data of the object.
[0014] According to another aspect of the embodiments of this application, a key point detection method based on wireless radar signals is provided, comprising:
[0015] Point cloud data is obtained by sensing objects using radar.
[0016] Feature extraction is performed on point cloud data acquired over a period of time to obtain the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud;
[0017] The first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data are concatenated;
[0018] A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information; and
[0019] The fused feature information is detected using a key point detection model based on a neural network to output the key point data of the object.
[0020] One of the beneficial effects of this application's embodiments is that: feature extraction is performed on point cloud data obtained over a period of time to obtain first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflected point cloud; a neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information; and a neural network-based key point detection model is used to detect the fused feature information to output the object's key point data. Therefore, detecting key points of objects (e.g., the human body) based on radar point clouds can be performed without limiting the action category, requires less computational resources, and has high detection accuracy; furthermore, it is easy to implement, simple to operate, has strong noise resistance, and provides high privacy protection.
[0021] Referring to the following description and accompanying drawings, specific implementation methods of the embodiments of this application are disclosed in detail, indicating how the principles of the embodiments of this application can be adopted. It should be understood that the implementation methods of this application are not limited in scope. Within the spirit and scope of the appended claims, the implementation methods of this application include many changes, modifications, and equivalents. Attached Figure Description
[0022] The accompanying drawings, which form part of the specification, are used to provide a further understanding of the embodiments of this application and illustrate the implementation methods of this application, together with the textual description, to explain the principles of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other implementation methods based on these drawings without creative effort. In the drawings:
[0023] Figure 1 This is a schematic diagram of a key point detection method based on wireless radar signals according to an embodiment of this application;
[0024] Figure 2 This is an example diagram of key human body points according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the reflection point distance value according to an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of multiple network models according to embodiments of this application;
[0027] Figure 5 This is a schematic diagram of the feature data of an embodiment of this application;
[0028] Figure 6 This is a schematic diagram of spatial feature data transformation according to an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of a neural network-based spatial transformation model according to an embodiment of this application;
[0030] Figure 8 This is another schematic diagram of several network models in the embodiments of this application;
[0031] Figure 9 This is a schematic diagram of a neural network-based fusion feature extraction model according to an embodiment of this application;
[0032] Figure 10 This is a schematic diagram of a key point detection model based on a neural network according to an embodiment of this application;
[0033] Figure 11 This is a schematic diagram of a key point identification device based on wireless radar signals according to an embodiment of this application.
[0034] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0035] Referring to the accompanying drawings, the foregoing and other features of the embodiments of this application will become apparent from the following description. Specific embodiments of this application are specifically disclosed in the description and drawings, illustrating partial implementations in which the principles of the embodiments of this application can be adopted. It should be understood that this application is not limited to the described embodiments; rather, the embodiments of this application include all modifications, variations, and equivalents falling within the scope of the appended claims.
[0036] In the embodiments of this application, the terms "first," "second," etc., are used to distinguish different elements by name, but do not indicate the spatial arrangement or chronological order of these elements, and these elements should not be limited by these terms. The term "and / or" includes any one or more of the terms listed in association and all combinations thereof. The terms "comprising," "including," "having," etc., refer to the presence of the stated features, elements, components, or assemblies, but do not exclude the presence or addition of one or more other features, elements, components, or assemblies.
[0037] In the embodiments of this application, the singular forms "a," "the," etc., including the plural forms, should be broadly understood as "a kind" or "a class" rather than limited to the meaning of "an." Furthermore, the term "the" should be understood to include both the singular and plural forms, unless the context explicitly indicates otherwise. Additionally, the term "according to" should be understood as "at least partially based on…," and the term "based on" should be understood as "at least partially based on…," unless the context explicitly indicates otherwise.
[0038] Features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or substituted for features in other embodiments. The term "comprising / including" as used herein means the presence of a feature, integral, step, or component, but does not exclude the presence or addition of one or more other features, integrals, steps, or components.
[0039] In this embodiment, the radar can be a millimeter-wave (mmWave) radar, but is not limited to this. The radar transmits electromagnetic waves through a transmitting antenna, and after reflection from different objects, receives the corresponding reflected waves (which can be called radar echo information). By analyzing the radar echo information, information such as the object's distance from the radar and its radial velocity can be effectively extracted, which can meet the needs of many application scenarios.
[0040] In the embodiments of this application, the object being detected can be a person of various ages, such as an elderly person, a child, or an elderly person and / or caregiver, or a child and / or guardian. This application is not limited to these; the object being detected can also be an animal with living characteristics, or a robot without living characteristics, etc. The following explanation uses the human body as an example.
[0041] First aspect of the embodiments
[0042] This application provides a key point detection method based on wireless radar signals. Figure 1 This is a schematic diagram of a key point detection method based on wireless radar signals according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes:
[0043] 101. Use radar to sense objects and obtain point cloud data;
[0044] 102. Feature extraction is performed on the point cloud data obtained over a period of time to obtain the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud;
[0045] 103. The first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and time distribution feature data are concatenated.
[0046] 104. A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information; and
[0047] 105. A key point detection model based on neural networks is used to detect fused feature information in order to output key point data of the object.
[0048] It is worth noting that the above appendix Figure 1 The embodiments of this application have only been illustrated schematically, and the application is not limited thereto. For example, the execution order between various operations can be appropriately adjusted, and other operations can be added or some operations can be removed. Those skilled in the art can make appropriate modifications based on the above description, and are not limited to the above-described embodiments. Figure 1 The records.
[0049] In some embodiments, radar senses the external space via wireless signals. The point cloud data output by the radar includes information such as the distance, velocity, and position of objects detected by the radar in the space. The radar periodically transmits wireless signals for detection, and the point cloud information is also output periodically. The point cloud information output by the radar in one detection operation can be called one frame of point cloud data.
[0050] In this embodiment, N consecutive frames of radar point cloud data are taken as input. After preprocessing of the radar point cloud data and calculation by the motion detection model, key point information of the human body is output to represent the human's motion. The N frames of radar point cloud data can be used... This indicates that the frame number i is arranged in chronological order; the larger the frame number, the later the corresponding point cloud data appeared. Therefore, P N This indicates the latest radar point cloud data.
[0051] In some embodiments, a frame of point cloud data P from the radar consists of several points, P = {p j Let p, 1≤j≤n}, where n is the number of points contained in the point cloud of that frame, and p j This is the j-th point. In point cloud data, a point is represented by p, where p = (s, v, p, x, y, z), where s is the frame number, v is the Doppler velocity relative to the radar, p is the signal strength of that point, and (x, y, z) are its spatial coordinates.
[0052] Therefore, the raw spatial feature data obtained directly from the radar output signal can be used as the first spatial feature data of the reflection point cloud, the Doppler velocity v obtained directly from the radar output signal can be used as the Doppler velocity feature data of the reflection point cloud, and the signal strength p obtained directly from the radar output signal can be used as the reflection energy feature data of the reflection point cloud.
[0053] Furthermore, the raw spatial feature data, Doppler velocity, and signal strength obtained from the radar output signal can be processed, and the processed data can be used as the first spatial feature data, Doppler velocity feature data, and reflection energy feature data of the reflection point cloud. For example, the raw spatial coordinate values (x0, y0, z0) of the radar output signal can be translated, rotated, etc., and the transformed values (x1, y1, z1) of the processed spatial coordinates can be used as the first spatial feature data of the reflection point cloud, etc.; this application is not limited to this.
[0054] In some embodiments, key human body points correspond to major joints or organs of the human body, such as the nose, shoulder, elbow, wrist, and hip. This application does not limit the selection of key human body points; different key human body points can be selected according to the specific application requirements.
[0055] Figure 2 This is an example diagram of key human body points according to an embodiment of this application. For example... Figure 2 As shown, key point information of the human body can refer to the relative positional relationship of the joints or organs corresponding to the key points, denoted by H = {(x k ,y k ,z k ),1≤k≤m} means that (x k ,yk ,z k Let (x, y) be the coordinates of the k-th keypoint. When only 2D positional information is available, one dimension of the keypoint coordinates can be set to 0; for example, z can be set to 0, so that only (x, y) contains valid positional information.
[0056] The above provides an illustrative description of radar point cloud data, but this application is not limited thereto. Furthermore, density distribution characteristic data and temporal distribution characteristic data can also be obtained from radar point cloud data.
[0057] The distribution of radar point clouds is random, and the density of radar point clouds varies at different locations during different movements. For example, when the upper limb of the target moves with a large amplitude, the radar point cloud is more concentrated at the location of the upper limb, while the point cloud distribution at other locations is relatively sparse.
[0058] In some embodiments, obtaining density distribution feature data of a reflection point cloud includes: calculating the distance value between each reflection point and other reflection points; calculating the probability corresponding to the distance value using a probability density function; and using the probability as the density distribution feature data.
[0059] Figure 3 This is a schematic diagram of the reflection point distance values according to an embodiment of this application, using the average value as an example, but this application is not limited thereto. Figure 3 As shown, for N*3 reflection points, N*N distances between the reflection points can be calculated, thus obtaining the average of N*1 distances. Therefore, the density d of a single reflection point can be obtained. i (i = 1, ..., N).
[0060] For example, the probability density function is:
[0061]
[0062] d i The distance value is (e.g., mean); μ is a preset distance value parameter, and σ is the variance value;
[0063] In some embodiments, different variance values correspond to different probabilities to obtain multiple density distribution feature data.
[0064] For example, μ = 0 and σ = 0.1 can obtain one density distribution feature; correspondingly, σ = 0.2 can obtain another density distribution feature; σ = 0.4 can obtain another density distribution feature; and σ = 0.6 can obtain another density distribution feature. Thus, four different density distribution features can be obtained for each reflection point, forming N*4 dimensional density distribution feature data, which can be used as input features for different channels in the neural network structure.
[0065] In some embodiments, obtaining temporal distribution feature data of reflection point clouds includes: calculating the time value corresponding to each reflection point; calculating the probability corresponding to the time value using a time density function; and using the probability as the temporal distribution feature data.
[0066] For example, the frame number of each reflection point can be directly obtained from the radar output signal (e.g., the time interval of the radar output signal is 0.1s, i.e., 1 frame / 0.1s). The frame numbers (1, 2, ...) of the radar reflection points can be used to calculate the distribution characteristics of the point cloud frame numbers, thereby obtaining the time series distribution characteristics of the radar reflection point cloud.
[0067] For example, the time density function is:
[0068]
[0069] f i The time value is μ; μ is a preset time value parameter, and σ is the variance value.
[0070] In some embodiments, different variance values correspond to different probabilities to obtain multiple time distribution feature data.
[0071] For example, f i The frame number (1, 2, ...) of the reflection point represents the temporal continuity information of the point cloud. μ = 0, σ = 0.5 yields one temporal distribution feature; correspondingly, σ = 1.0 yields another; σ = 1.5 yields yet another; and σ = 2.0 yields yet another. Thus, four different temporal distribution features can be obtained for each reflection point, forming N*4 dimensional temporal distribution feature data, which is then used as input features for different channels in the neural network structure.
[0072] The above provides an illustrative description of the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data. The embodiments of this application utilize these feature data to identify key points.
[0073] Figure 4 This is a schematic diagram of multiple network models according to embodiments of this application, such as... Figure 4 As shown, the first spatial feature data (e.g., N*3 dimensions), Doppler velocity feature data (e.g., N*1 dimensions), reflection energy feature data (e.g., N*1 dimensions), density distribution feature data (e.g., N*4 dimensions), and time distribution feature data (e.g., N*4 dimensions), totaling N*13 dimensions, can be used as input data for the fusion feature extraction model 401.
[0074] like Figure 4As shown, the feature extraction model 401 outputs feature data, which serves as the input data for the keypoint detection model 402. Figure 4 As shown, the key point detection model 402 can output key point information, thereby realizing the key point recognition of an object.
[0075] Therefore, detecting key points of objects (such as the human body) based on radar point clouds can be performed regardless of the action category, requiring less computational resources and achieving high detection accuracy. Furthermore, by performing deep analysis of feature data through multiple network models, it is possible to distinguish the feature differences between different actions, further accurately identifying key point information for different actions.
[0076] In some embodiments, the first spatial feature data (original spatial feature data) can be transformed using a neural network-based spatial transformation model to obtain normalized second spatial feature data; and the normalized second spatial feature data can be concatenated with the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and time distribution feature data, and the concatenated feature data can be input into the fusion feature extraction model.
[0077] Figure 5 This is a schematic diagram of the feature data in an embodiment of this application. For example... Figure 5 As shown, feature extraction can be performed on point cloud data obtained over a period of time (as shown in 501) to obtain first spatial feature data (original spatial feature data, as shown in 502), Doppler velocity feature data (as shown in 503), reflection energy feature data (as shown in 504), density distribution feature data (as shown in 505), and temporal distribution feature data (as shown in 506).
[0078] like Figure 5 As shown, the first spatial feature data (original spatial feature data) can also be transformed using a neural network-based spatial transformation model to obtain normalized second spatial feature data (as shown in Figure 507). The normalized second spatial feature data can be concatenated with the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and time distribution feature data.
[0079] like Figure 5 As shown, the first spatial feature data (original spatial feature data), along with the Doppler velocity feature data, reflection energy feature data, density distribution feature data, and time distribution feature data, can be referred to as input 1, and the normalized second spatial feature data can be referred to as input 2. Both of these can be input into the fusion feature extraction model 508.
[0080] Figure 6This is a schematic diagram of spatial feature data transformation according to an embodiment of this application, as shown below. Figure 6 As shown, the N*3 dimensional first spatial feature data (as shown in 601) can be extended to obtain N*M dimensional extended spatial feature data (as shown in 602); pooling operations (e.g., max pooling) can be performed on the N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data (as shown in 603); feature transformation can be performed on the 1*M dimensional global spatial feature data to obtain 1*9 dimensional feature data, which is then transformed into 3*3 dimensional spatial feature data (as shown in 604); and matrix multiplication can be performed between the N*3 dimensional first spatial feature data and the 3*3 dimensional spatial feature data to obtain the normalized second spatial feature data (as shown in 605); where N is a positive integer greater than 1 and M is a positive integer greater than 3.
[0081] Figure 7 This is a schematic diagram of a spatial transformation model based on a neural network according to an embodiment of this application, as shown below. Figure 7 As shown, for example, the 200*3 dimensional first spatial feature data can be extended using a one-dimensional convolutional neural network (as shown in 701) to obtain 200*256 dimensional extended spatial feature data; max pooling operation can be performed on the 200*256 dimensional extended spatial feature data (as shown in 702) to obtain 1*256 dimensional global spatial feature data; the 1*256 dimensional global spatial feature data can be transformed using a fully connected neural network (as shown in 703) to obtain 1*64 dimensional feature data, which is then transformed again into 1*9 dimensional spatial feature data using a fully connected network, and then the 1*9 dimensional feature data can be transformed into 3*3 dimensional spatial feature data (as shown in 704); and the 200*3 dimensional first spatial feature data and the 3*3 dimensional spatial feature data can be matrix multiplied (as shown in 705) to obtain the normalized 200*3 dimensional second spatial feature data.
[0082] Figure 8 This is another schematic diagram of several network models in the embodiments of this application, such as... Figure 8 As shown, the first spatial feature data (e.g., N*3 dimensions), Doppler velocity feature data (e.g., N*1 dimensions), reflection energy feature data (e.g., N*1 dimensions), density distribution feature data (e.g., N*4 dimensions), and time distribution feature data (e.g., N*4 dimensions), totaling N*13 dimensions, can be used as input 1 to the fusion feature extraction model 801.
[0083] like Figure 8As shown, the first spatial feature data (original spatial feature data) can also be transformed using a neural network-based spatial transformation model 802 to obtain normalized second spatial feature data (e.g., N*3 dimensions), and the normalized second spatial feature data can be used as input 2 of the fusion feature extraction model 801.
[0084] like Figure 8 As shown, the feature extraction model 801 outputs feature data, which serves as the input data for the keypoint detection model 803. Figure 8 As shown, the key point detection model 803 can output key point information, thereby realizing the key point recognition of an object.
[0085] Therefore, detecting key points of objects (such as the human body) based on radar point clouds can be done regardless of the action category, requires less computational resources, and further improves detection accuracy. Furthermore, by performing deep analysis of feature data through multiple network models, the feature differences between different actions can be distinguished, further accurately identifying key point information for different actions.
[0086] In some embodiments, a neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information, including: expanding the N*L-dimensional cascaded feature data to obtain N*K-dimensional expanded feature data; performing pooling operations on the N*K-dimensional expanded feature data to obtain 1*K-dimensional global feature data; and performing feature transformation on the 1*K-dimensional global feature data to obtain 1*L... 2 The N*L dimensional feature data is transformed into L*L dimensional feature data; and the N*L dimensional concatenated feature data is multiplied by the L*L dimensional feature data to obtain N*L dimensional fused feature information; where N, L and K are all positive integers greater than 1, and L is less than K.
[0087] Figure 9 This is a schematic diagram of a neural network-based fusion feature extraction model according to an embodiment of this application, as shown below. Figure 9 As shown, for example, the first spatial feature data (200*3D), the Doppler velocity feature data (200*1D), the reflection energy feature data (200*1D), the density distribution feature data (200*4D), and the time distribution feature data (200*4D) can be used as input 1 of 200*13D, and the normalized second spatial feature data (200*3D) can be used as input 2 of 200*3D. Input 1 and input 2 are concatenated (as shown in 901) to form fusion feature information of 200*16D, which is used as the input data of the fusion feature extraction model.
[0088] like Figure 9As shown, the 200*16 dimensional fused feature data can be expanded using a one-dimensional convolutional neural network (as shown in 902) to obtain 200*64 dimensional feature data (Feature Data 1), and then 200*256 dimensional feature data (as shown in 903). Max pooling is then performed on the 200*256 dimensional feature data (as shown in 904) to obtain 1*256 dimensional feature data. The 1*256 dimensional feature data is then transformed using a fully connected neural network (as shown in 905) to obtain 1*64 dimensional feature data, which is then further transformed into 1*4096 dimensional feature data using a fully connected network. Finally, the 1*4096 dimensional feature data is transformed into 64*64 dimensional feature data (Feature Data 2). 2) (as shown in 906); and perform matrix multiplication of the 200*64 dimensional feature data 1 and the 64*64 dimensional feature data 2 (as shown in 907) to obtain 200*64 dimensional feature data (fusion feature information, as shown in 908).
[0089] In some embodiments, a key point detection model based on a neural network is used to detect fused feature information to output key point data of an object, including: extending the N*L-dimensional fused feature information to obtain N*J-dimensional extended feature data; performing pooling operations on the N*J-dimensional extended feature data to obtain 1*J-dimensional global feature data; and performing feature transformation on the 1*J-dimensional global feature data to obtain 1*P-dimensional key point position information; wherein N, L, and J are all positive integers greater than 1, and L is less than J.
[0090] Figure 10 This is a schematic diagram of a key point detection model based on a neural network according to an embodiment of this application, as shown below. Figure 10 As shown, for example, the 200*64-dimensional feature data output by the fusion feature extraction model can be used as the input data for the key point detection model.
[0091] like Figure 10 As shown, 200*64 dimensional feature data can be expanded using a one-dimensional convolutional neural network (as shown in 1001) to obtain 200*128 dimensional feature data, and then 200*256 dimensional feature data (as shown in 1002); max pooling operation is performed on the 200*256 dimensional feature data (as shown in 1003) to obtain 1*256 dimensional feature data; feature transformation is performed on the 1*256 dimensional feature data through a fully connected neural network (as shown in 1004) to obtain 1*64 dimensional feature data, and then transformed again into 1*34 dimensional feature data using a fully connected network (as shown in 1005), and then the key point information is output.
[0092] In the embodiments of this application, simulation experiments show that using N*13 dimensional feature data and multiple network models for feature extraction and detection can accurately detect key point information, improving detection accuracy compared to using the original N*6 dimensional data. Furthermore, using N*13 dimensional feature data and multiple network models for feature extraction and detection can further improve detection accuracy.
[0093] The above provides an illustrative description of the recognition method and model. In this embodiment, one or more sets of optimal parameters can be obtained through supervised training; these parameters are then applied to the detection model to perform calculations on the input radar point cloud data to obtain the corresponding human key point information. This embodiment does not limit the specific training of the model; for example, SGD (Stochastic Gradient Descent) optimization, Adam (Adaptive Moment Estimation) optimization, etc., can be used.
[0094] The above description only covers the steps or processes related to this application, but this application is not limited thereto. The action detection method may also include other steps or processes; for details of these steps or processes, please refer to the prior art. Furthermore, the above description only uses some structural examples of action detection models to illustrate the embodiments of this application, but this application is not limited to these structures, and appropriate modifications can be made to these structures. All such modifications should be included within the scope of the embodiments of this application.
[0095] The above embodiments are merely illustrative examples of embodiments of this application, but this application is not limited thereto, and appropriate modifications can be made based on the above embodiments. For example, the above embodiments can be used alone, or one or more of the above embodiments can be combined.
[0096] As can be seen from the above embodiments, feature extraction is performed on point cloud data obtained over a period of time to obtain first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflected point cloud. A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information. A neural network-based keypoint detection model is then used to detect the fused feature information to output the object's keypoint data. Therefore, detecting keypoints of objects (e.g., the human body) based on radar point clouds is not limited to any action category, requires less computational resources, and has high detection accuracy. Furthermore, it is easy to implement, simple to operate, has strong noise resistance, and offers high privacy protection.
[0097] Second aspect of the embodiments
[0098] This application provides a key point identification device based on wireless radar signals, and the same content as the first aspect of the embodiment will not be repeated.
[0099] Figure 11 This is a schematic diagram of a key point identification device based on wireless radar signals according to an embodiment of this application, as shown below. Figure 11 As shown, the key point identification device 1100 based on wireless radar signals includes:
[0100] The sensing unit 1101 uses radar to sense objects and obtain point cloud data.
[0101] The feature extraction unit 1102 performs feature extraction on point cloud data obtained over a period of time to obtain the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud.
[0102] The cascade unit 1103 cascades the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the time distribution feature data.
[0103] Feature detection unit 1104 uses a neural network-based fusion feature extraction model to detect the cascaded feature data to obtain fused feature information; and
[0104] The key point detection unit 1105 uses a key point detection model based on a neural network to detect the fused feature information and output the key point data of the object.
[0105] In some embodiments, such as Figure 11 As shown, the key point identification device 1100 based on wireless radar signals may further include:
[0106] The feature transformation unit 1106 uses a spatial transformation model based on a neural network to transform the first spatial feature data to obtain the normalized second spatial feature data.
[0107] The cascade unit 1103 can also be used to cascade the normalized second spatial feature data with the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the time distribution feature data, and input the cascaded feature data into the fusion feature extraction model.
[0108] In some embodiments, the feature transformation unit 1106 is used for:
[0109] Feature expansion is performed on the N*3 dimensional first spatial feature data to obtain N*M dimensional expanded spatial feature data;
[0110] Pooling is performed on N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data.
[0111] Feature transformation is performed on 1*M dimensional global spatial feature data to obtain 1*9 dimensional feature data, which is then transformed into 3*3 dimensional spatial feature data; and
[0112] The N*3 dimensional first spatial feature data is multiplied by the 3*3 dimensional spatial feature data to obtain the normalized second spatial feature data.
[0113] Where N is a positive integer greater than 1, and M is a positive integer greater than 3.
[0114] In some embodiments, the feature detection unit 1104 is used for:
[0115] Feature expansion is performed on N*L dimensional cascaded feature data to obtain N*K dimensional extended feature data;
[0116] Pooling is performed on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data;
[0117] Perform feature transformation on 1*K dimensional global feature data to obtain 1*L dimensional data. 2 Transform the 1D feature data into L*L dimensional feature data; and
[0118] The N*L dimensional cascaded feature data is multiplied by the L*L dimensional feature data to obtain N*L dimensional fused feature information.
[0119] Where N, L, and K are all positive integers greater than 1, and L is less than K.
[0120] In some embodiments, the feature extraction unit 1102 obtains the density distribution feature data of the reflection point cloud, including:
[0121] Calculate the distance between each reflection point and other reflection points;
[0122] The probability corresponding to the distance value is calculated using a probability density function, and the probability is used as the density distribution feature data.
[0123] In some embodiments, the probability density function is:
[0124]
[0125] d i The distance value is denoted by μ; μ is a preset distance value parameter, and σ is the variance value.
[0126] Different variance values correspond to different probabilities, thereby obtaining multiple density distribution feature data.
[0127] In some embodiments, the feature extraction unit 1102 obtains the temporal distribution feature data of the reflection point cloud, including:
[0128] Calculate the time value corresponding to each reflection point;
[0129] The probability corresponding to the time value is calculated using a time density function, and the probability is used as the time distribution feature data.
[0130] In some embodiments, the time density function is:
[0131]
[0132] f i The time value is μ; μ is a preset time value parameter, and σ is the variance value.
[0133] Different variance values correspond to different probabilities, thereby obtaining multiple time distribution feature data.
[0134] In some embodiments, the key point detection unit 1105 is used for:
[0135] Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data;
[0136] Pooling is performed on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data;
[0137] Feature transformation is performed on 1*J-dimensional global feature data to obtain 1*P-dimensional key point location information;
[0138] Where N, L, J and P are all positive integers greater than 1, and L is less than J and P is less than J.
[0139] It is worth noting that the above description only covers the components or modules relevant to this application, but this application is not limited thereto. The key point identification device 1100 based on wireless radar signals may also include other components or modules, and for details regarding these components or modules, please refer to relevant technologies.
[0140] For the sake of simplicity, Figure 11 The diagram only exemplifies the connection relationships or signal flow between various components or modules; however, those skilled in the art should understand that various related technologies, such as bus connections, can be employed. The aforementioned components or modules can be implemented using hardware facilities such as processors and memory; this application does not limit the scope of the embodiments.
[0141] The above embodiments are merely illustrative examples of embodiments of this application, but this application is not limited thereto, and appropriate modifications can be made based on the above embodiments. For example, the above embodiments can be used alone, or one or more of the above embodiments can be combined.
[0142] As can be seen from the above embodiments, feature extraction is performed on point cloud data obtained over a period of time to obtain first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflected point cloud. A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information. A neural network-based keypoint detection model is then used to detect the fused feature information to output the object's keypoint data. Therefore, detecting keypoints of objects (e.g., the human body) based on radar point clouds is not limited to any action category, requires less computational resources, and has high detection accuracy. Furthermore, it is easy to implement, simple to operate, has strong noise resistance, and offers high privacy protection.
[0143] Third aspect of the embodiments
[0144] This application provides an electronic device including a key point identification device 1100 based on wireless radar signals as described in the second aspect of the embodiment, the contents of which are incorporated herein by reference. This electronic device may be, for example, a computer, server, workstation, laptop computer, smartphone, etc.; however, this application is not limited thereto.
[0145] Figure 12 This is a schematic diagram of an electronic device according to an embodiment of this application. For example... Figure 12 As shown, the electronic device 1200 may include a processor (e.g., a central processing unit, CPU) 1210 and a memory 1220; the memory 1220 is coupled to the central processing unit 1210. The memory 1220 can store various data; in addition, it also stores an information processing program 1221, and executes the program 1221 under the control of the processor 1210.
[0146] In some embodiments, the functionality of the key point recognition device 1100 based on wireless radar signals is integrated into the processor 1210. The processor 1210 is configured to implement the key point recognition method based on wireless radar signals as described in the embodiments of the first aspect.
[0147] In some embodiments, the key point recognition device 1100 based on wireless radar signals is configured separately from the processor 1210. For example, the key point recognition device 1100 based on wireless radar signals can be configured as a chip connected to the processor 1210, and the function of the key point recognition device 1100 based on wireless radar signals can be realized through the control of the processor 1210.
[0148] For example, processor 1210 is configured to perform the following control: perceive an object using radar to obtain point cloud data; extract features from the point cloud data obtained over a period of time to obtain first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflected point cloud; cascade the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the temporal distribution feature data; detect the cascaded feature data using a neural network-based fusion feature extraction model to obtain fused feature information; and detect the fused feature information using a neural network-based keypoint detection model to output keypoint data of the object.
[0149] In addition, such as Figure 12 As shown, the electronic device 1200 may further include: an input / output (I / O) device 1230 and a display 1240, etc.; the functions of the above components are similar to those in the prior art, and will not be described in detail here. It is worth noting that the electronic device 1200 is not necessarily required to include... Figure 12 All components shown; in addition, the electronic device 1200 may also include Figure 12 For components not shown, please refer to relevant technologies.
[0150] This application also provides a computer-readable program, wherein when the program is executed in an electronic device, the program causes the computer in the electronic device to perform the key point identification method based on wireless radar signals as described in the first aspect embodiment.
[0151] This application also provides a storage medium storing a computer-readable program, wherein the computer-readable program causes a computer in an electronic device to perform the key point identification method based on wireless radar signals as described in the first aspect embodiment.
[0152] The apparatus and methods described above in this application can be implemented in hardware or in combination with software. This application relates to a computer-readable program that, when executed by a logic component, enables the logic component to implement the apparatus or components described above, or to implement the various methods or steps described above. This application also relates to storage media for storing the above programs, such as hard disks, magnetic disks, optical disks, DVDs, flash memory, etc.
[0153] The methods / apparatus described in conjunction with the embodiments of this application can be directly embodied in hardware, software modules executed by a processor, or a combination of both. For example, one or more and / or combinations of one or more functional block diagrams shown in the figures can correspond to various software modules in a computer program flow, or to various hardware modules. These software modules can correspond to the various steps shown in the figures, respectively. These hardware modules can be implemented, for example, using a field-programmable gate array (FPGA) to embed these software modules.
[0154] The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. A storage medium can be coupled to the processor, enabling the processor to read information from and write information to the storage medium; or the storage medium can be an integral part of the processor. The processor and storage medium can reside in an ASIC. The software module can be stored in the memory of a mobile terminal or in a memory card that can be inserted into the mobile terminal. For example, if the device (such as a mobile terminal) uses a high-capacity MEGA-SIM card or a high-capacity flash memory device, the software module can be stored in the MEGA-SIM card or the high-capacity flash memory device.
[0155] One or more and / or one or more combinations of functional blocks described in the accompanying drawings can be implemented as a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any suitable combination thereof for performing the functions described herein. One or more and / or one or more combinations of functional blocks described in the accompanying drawings can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in communication with a DSP, or any other such configuration.
[0156] The present application has been described above with reference to specific embodiments. However, those skilled in the art should understand that these descriptions are exemplary and not intended to limit the scope of protection of the present application. Those skilled in the art can make various modifications and variations to the present application based on the principles thereof, and these modifications and variations are also within the scope of the present application.
[0157] Regarding the implementation methods including the above embodiments, the following notes are also disclosed:
[0158] Appendix 1. A key point identification method based on wireless radar signals, comprising:
[0159] Point cloud data is obtained by sensing objects using radar.
[0160] Feature extraction is performed on point cloud data acquired over a period of time to obtain the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud;
[0161] The first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data are concatenated;
[0162] A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information; and
[0163] The fused feature information is detected using a key point detection model based on a neural network to output the key point data of the object.
[0164] Appendix 2. The method according to Appendix 1, wherein the method further comprises:
[0165] The first spatial feature data is transformed using a neural network-based spatial transformation model to obtain normalized second spatial feature data; and
[0166] The normalized second spatial feature data is concatenated with the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the time distribution feature data, and the concatenated feature data is input into the fusion feature extraction model.
[0167] Appendix 3. The method according to Appendix 1 or 2, wherein transforming the first spatial feature data using a neural network-based spatial transformation model includes:
[0168] Feature expansion is performed on the N*3 dimensional first spatial feature data to obtain N*M dimensional expanded spatial feature data;
[0169] Pooling is performed on N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data.
[0170] Feature transformation is performed on 1*M dimensional global spatial feature data to obtain 1*9 dimensional feature data, which is then transformed into 3*3 dimensional spatial feature data; and
[0171] The N*3 dimensional first spatial feature data is multiplied by the 3*3 dimensional spatial feature data to obtain the normalized second spatial feature data.
[0172] Where N is a positive integer greater than 1, and M is a positive integer greater than 3.
[0173] Appendix 4. The method according to any one of Appendices 1 to 3, wherein detecting the cascaded feature data using a neural network-based fusion feature extraction model includes:
[0174] Feature expansion is performed on N*L dimensional cascaded feature data to obtain N*K dimensional extended feature data;
[0175] Pooling is performed on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data;
[0176] Perform feature transformation on 1*K dimensional global feature data to obtain 1*L dimensional data. 2 Transform the 1D feature data into L*L dimensional feature data; and
[0177] The N*L dimensional cascaded feature data is multiplied by the L*L dimensional feature data to obtain N*L dimensional fused feature information.
[0178] Where N, L, and K are all positive integers greater than 1, and L is less than K.
[0179] Appendix 5. The method according to any one of Appendices 1 to 4, wherein obtaining the density distribution characteristic data of the reflection point cloud includes:
[0180] Calculate the distance between each reflection point and other reflection points;
[0181] The probability corresponding to the distance value is calculated using a probability density function, and the probability is used as the density distribution feature data.
[0182] Appendix 6. According to the method described in Appendix 5, the probability density function is:
[0183]
[0184] d i σ is the distance value; μ is a preset distance value parameter, and σ is the variance value.
[0185] Note 7. According to the method described in Note 6, different variance values correspond to different probabilities to obtain multiple density distribution feature data.
[0186] Appendix 8. The method according to any one of Appendices 1 to 7, wherein obtaining the temporal distribution characteristic data of the reflection point cloud includes:
[0187] Calculate the time value corresponding to each reflection point;
[0188] The probability corresponding to the time value is calculated using a time density function, and the probability is used as the time distribution feature data.
[0189] Appendix 9. According to the method described in Appendix 8, the time density function is:
[0190]
[0191] f i The time value is denoted by μ; μ is a preset time value parameter; and σ is the variance value.
[0192] Note 10. According to the method described in Note 9, different variance values correspond to different probabilities to obtain multiple time distribution feature data.
[0193] Appendix 11. The method according to any one of Appendices 1 to 10, wherein detecting the fused feature information using a neural network-based keypoint detection model includes:
[0194] Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data;
[0195] Pooling is performed on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data;
[0196] Feature transformation is performed on 1*J-dimensional global feature data to obtain 1*P-dimensional key point location information;
[0197] Where N, L, J and P are all positive integers greater than 1, and L is less than J and P is less than J.
[0198] Appendix 12. An electronic device comprising a memory and a processor, the memory storing a computer program and the processor being configured to execute the computer program to implement the key point identification method based on wireless radar signals as described in any one of Appendices 1 to 11.
[0199] Note 13. A storage medium storing a computer-readable program, wherein the computer-readable program causes a computer in an electronic device to perform a key point identification method based on wireless radar signals as described in any one of Notes 1 to 11.
Claims
1. A key point identification device based on wireless radar signals, characterized in that, The device includes: The sensing unit uses radar to sense objects and obtain point cloud data. The feature extraction unit extracts features from the point cloud data obtained over a period of time, and obtains the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud. A cascade unit that cascades the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; The feature detection unit utilizes a neural network-based fusion feature extraction model to detect the cascaded feature data to obtain fused feature information; and The key point detection unit uses a neural network-based key point detection model to detect the fused feature information and outputs the key point data of the object. The device further includes: The feature transformation unit uses a neural network-based spatial transformation model to transform the first spatial feature data to obtain the normalized second spatial feature data. The cascade unit is also used to cascade the normalized second spatial feature data with the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the time distribution feature data, and input the cascaded feature data into the fusion feature extraction model; The feature extraction unit obtains the density distribution feature data of the reflection point cloud, including: Calculate the distance between each reflection point and other reflection points; The probability corresponding to the distance value is calculated using a probability density function, and the probability is used as the density distribution feature data. The probability density function is: d i The distance value is μ; μ is a preset distance value parameter; σ is the variance value; where different variance values correspond to different probabilities, so as to obtain multiple density distribution feature data. The feature extraction unit obtains the temporal distribution feature data of the reflection point cloud, including: Calculate the time value corresponding to each reflection point; The probability corresponding to the time value is calculated using a time density function, and the probability is used as the time distribution feature data. The time density function is: f i The time value is denoted by μ, which is a preset time value parameter, and σ is the variance value. Different variance values correspond to different probabilities, so as to obtain multiple time distribution feature data.
2. The apparatus according to claim 1, wherein, The feature transformation unit is used for: Feature expansion is performed on the N*3 dimensional first spatial feature data to obtain N*M dimensional expanded spatial feature data; Pooling is performed on N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data. Feature transformation is performed on 1*M dimensional global spatial feature data to obtain 1*9 dimensional feature data, which is then transformed into 3*3 dimensional spatial feature data; and The N*3 dimensional first spatial feature data is multiplied by the 3*3 dimensional spatial feature data to obtain the normalized second spatial feature data. Where N is a positive integer greater than 1, and M is a positive integer greater than 3.
3. The apparatus according to claim 1, wherein, The feature detection unit is used for: Feature expansion is performed on N*L dimensional cascaded feature data to obtain N*K dimensional extended feature data; Pooling is performed on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data; Perform feature transformation on 1*K dimensional global feature data to obtain 1*L dimensional data. 2 Transform the 1D feature data into L*L dimensional feature data; and The N*L dimensional cascaded feature data is multiplied by the L*L dimensional feature data to obtain N*L dimensional fused feature information. Where N, L, and K are all positive integers greater than 1, and L is less than K.
4. The apparatus according to claim 1, wherein, The key point detection unit is used for: Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data; Pooling is performed on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data; Feature transformation is performed on 1*J-dimensional global feature data to obtain 1*P-dimensional key point location information; Where N, L, J and P are all positive integers greater than 1, and L is less than J and P is less than J.
5. A key point identification method based on wireless radar signals, characterized in that, The method includes: Point cloud data is obtained by sensing objects using radar. Feature extraction is performed on point cloud data acquired over a period of time to obtain the first spatial feature data, Doppler velocity feature data, reflection energy feature data, density distribution feature data, and temporal distribution feature data of the reflection point cloud; The first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data are concatenated; A neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fused feature information; and The fused feature information is detected using a key point detection model based on a neural network to output the key point data of the object; The method further includes: The first spatial feature data is transformed using a neural network-based spatial transformation model to obtain the normalized second spatial feature data. The normalized second spatial feature data is concatenated with the first spatial feature data, the Doppler velocity feature data, the reflection energy feature data, the density distribution feature data, and the time distribution feature data, and the concatenated feature data is input into the fusion feature extraction model; The acquisition of the density distribution feature data of the reflection point cloud includes: Calculate the distance between each reflection point and other reflection points; The probability corresponding to the distance value is calculated using a probability density function, and the probability is used as the density distribution feature data. The probability density function is: d i The distance value is μ; μ is a preset distance value parameter; σ is the variance value; where different variance values correspond to different probabilities, so as to obtain multiple density distribution feature data. The acquisition of the temporal distribution feature data of the reflection point cloud includes: Calculate the time value corresponding to each reflection point; The probability corresponding to the time value is calculated using a time density function, and the probability is used as the time distribution feature data. The time density function is: f i The time value is denoted by μ, which is a preset time value parameter, and σ is the variance value. Different variance values correspond to different probabilities, so as to obtain multiple time distribution feature data.
Citation Information
Patent Citations
Obstacle detection system and method based on depth information
CN110070570A
Tool and method for annotating a human pose in 3D point cloud data
CN111695402A