Apparatus and method for keypoint recognition based on radio radar signals
The keypoint detection method using wireless radar signals addresses the limitations of conventional radar by performing multi-feature extraction and neural network-based detection, achieving high-accuracy and computationally efficient keypoint recognition across various actions with enhanced privacy protection.
Patent Information
- Application Number
- JP2022073909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-01
- Filing Date
- 2022-04-27
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-04-27
AI Technical Summary
Conventional radar-based keypoint detection is limited to a small number of specific actions, restricting its application scenarios and is computationally intensive.
A keypoint detection apparatus and method using wireless radar signals that perform feature extraction on point cloud data, combining spatial, Doppler velocity, reflected energy, density distribution, and time distribution features using neural networks to detect keypoints without action category limitations, requiring fewer computational resources.
Enables high-accuracy keypoint detection of objects, such as human bodies, with improved computational efficiency and privacy protection, resistant to noise and environmental variations.
Smart Images

Figure 0007790267000009 
Figure 0007790267000010 
Figure 0007790267000011
Abstract
Description
[Technical Field]
[0001] The present invention relates to the technical field of radar detection. [Background technology]
[0002] In the process of detecting human body movements, it is possible to detect key points (also called feature points) on the human body, such as the head, neck, arms, legs, waist, etc. Human body key point detection has a wide range of application scenarios and is an important technology for applications such as smart homes, health monitoring, and behavior understanding.
[0003] Currently, video-based human body keypoint detection technology is widely used. However, video can invade privacy and cannot be used in private situations. In addition, video-based keypoint detection is heavily affected by the environment (e.g., blocking, lighting, haze, etc.), making it ineffective in scenes with no light or blocking. Furthermore, it can be significantly affected by clothing, posture, viewpoint, etc.
[0004] In contrast, radar detects objects (e.g., human bodies) using radio signals, which does not expose privacy, is independent of external conditions such as lighting, and can work well in partially obstructed scenes. Therefore, radar-based keypoint detection can compensate for the shortcomings of video-based detection. Summary of the Invention [Problem to be solved by the invention]
[0005] However, the inventors have discovered that conventional radar-based keypoint detection can only detect a small number of specific actions (movements or movements), significantly limiting its application scenarios.
[0006] To solve the above technical problems, an embodiment of the present invention provides a keypoint detection apparatus and method based on wireless radar signals, which can not only detect keypoints of an object (e.g., a human body) based on radar point clouds, but also is not limited to motion categories, requires less computational resources, and has high detection accuracy. [Means for solving the problem]
[0007] According to one aspect of an embodiment of the present invention, there is provided a keypoint detection device based on a radio radar signal, which includes: A detection unit that detects (senses) objects using radar and obtains point cloud data; a feature extraction unit that performs feature extraction on the point cloud data acquired within a predetermined period of time to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; a cascade connection unit for cascading the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; a feature detection unit that uses a neural network-based fusion feature extraction model to detect the cascaded feature data and obtain fusion feature information; and A keypoint detection unit is included for detecting the fused feature information using a keypoint detection model based on a neural network, and outputting keypoint data of the object.
[0008] According to another aspect of an embodiment of the present invention, there is provided a keypoint detection method based on a wireless radar signal, which includes: Detect (sense) objects using radar and obtain point cloud data; Perform feature extraction on the point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; cascading the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; Detecting the cascaded feature data using a neural network-based fusion feature extraction model to obtain fusion feature information; and The method includes detecting the fused feature information using a neural network-based keypoint detection model and outputting keypoint data of the object. [Effects of the Invention]
[0009] The advantageous effects of the embodiments of the present invention are as follows: feature extraction is performed on point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; a neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fusion feature information; and a neural network-based keypoint detection model is used to detect the fusion feature information and output keypoint data of the object, thereby enabling keypoints of an object (e.g., a human body) to be detected based on the radar point cloud, and is not limited to action categories, requires few computing resources, and has high detection accuracy; furthermore, it is easy to implement and operate, has high noise resistance, and has excellent privacy protection. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 illustrates a wireless radar signal-based keypoint detection method according to an embodiment of the present invention. [Figure 2]FIG. 2 is a diagram showing an example of key points of a human body in an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram showing the distance between reflection points in an embodiment of the present invention. [Figure 4] FIG. 1 illustrates a plurality of network models in an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram showing feature data in the embodiment of the present invention. [Figure 6] FIG. 10 is a diagram illustrating the transformation of spatial feature data in an embodiment of the present invention. [Figure 7] FIG. 1 illustrates a neural network-based spatial transformation model in an embodiment of the present invention. [Figure 8] FIG. 10 is another diagram illustrating multiple network models in an embodiment of the present invention. [Figure 9] FIG. 1 illustrates a neural network-based feature fusion model in an embodiment of the present invention. [Figure 10] FIG. 1 illustrates a neural network-based keypoint detection model in an embodiment of the present invention. [Figure 11] FIG. 1 illustrates a wireless radar signal based keypoint recognition device in an embodiment of the present invention. [Figure 12] 1 is a diagram illustrating an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. However, these embodiments are merely illustrative and are not intended to limit the scope of the present invention.
[0012] In an embodiment of the present invention, the radar may be, but is not limited to, a millimeter wave (mm Wave) radar. The radar transmits electromagnetic waves through a transmitting antenna, and receives corresponding reflected waves (also called radar echo information) after the electromagnetic waves are reflected by different objects. By analyzing the radar echo information, information such as the distance (position) from the object to the radar and the radial movement speed can be effectively extracted, which can meet the needs of many application scenarios.
[0013] In an embodiment of the present invention, the detection target object may be people of various ages, such as elderly people and children, elderly people and / or nursing staff, or children and / or nursing staff. The present invention is not limited thereto, and the detection target object may also be a living animal or an inanimate robot. The following description will be given taking the human body as an example.
[0014] <Example of the first aspect> In an embodiment of the present invention, a keypoint detection method based on a wireless radar signal is provided. Figure 1 is a diagram showing the keypoint detection method based on a wireless radar signal in an embodiment of the present invention. As shown in Figure 1, the method includes the following steps:
[0015] 101: Detect (sense) objects using radar and obtain point cloud data; 102: Perform feature extraction on the point cloud data acquired within a predetermined period of time to acquire first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; 103: Cascade-connecting first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data; 104: Obtain fusion feature information by detecting the cascaded feature data using a neural network-based fusion feature extraction model; and 105: Output keypoint data of the object by detecting fused feature information using a keypoint detection model based on a neural network.
[0016] Note that the above-described FIG. 1 is merely for illustrative purposes of an embodiment of the present invention, and the present invention is not limited thereto. For example, the execution order of each operation (step) can be appropriately adjusted, or some operations can be added or removed. Those skilled in the art can make appropriate changes and modifications to the above content without being limited to the description of FIG. 1.
[0017] In some embodiments, a radar detects (senses) an external space using a radio signal, and the point cloud data output by the radar includes distance, speed, and position information of objects in the space detected by the radar. The radar periodically transmits a radio signal for detection, and the point cloud information is also periodically output. The point cloud information output by the radar for one detection may be referred to as one frame of point cloud data.
[0018] In an embodiment of the present invention, N frames of continuous radar point cloud data are input, and after pre-processing of the radar point cloud data and calculation of a motion detection model, the key point information of the human body is output, which is used to represent the motion of the human body. The radar point cloud data of N frames is expressed as P={P i , 1≦i≦N}, where the frame number (sequence number) i is arranged in chronological order, and the larger the frame number, the later the corresponding point cloud data appears. N represents the latest radar point cloud data.
[0019] In some embodiments, the point cloud data P of one frame of radar consists of multiple points, and P={p j , 1≦j≦n}, where n is the number of points contained in the point cloud of the frame, and p jis the j-th point. A point in the point cloud data is represented by p, where p = (s,v,p,x,y,z), where s is the frame number, v is the Doppler velocity relative to the radar, p is the signal strength of the point, and (x,y,z) are its coordinates in space.
[0020] Therefore, the original spatial feature data obtained by the radar output signal can be directly used as the first spatial feature data of the reflection point cloud, the Doppler velocity v obtained by the radar output signal can be used as the Doppler velocity feature data of the reflection point cloud, and the signal intensity p obtained by the radar output signal can be used as the reflection energy feature data of the reflection point cloud.
[0021] Alternatively, the original spatial feature data, Doppler velocity, and signal intensity obtained from the radar output signal can be processed, and the processed data can be used as the first spatial feature data, Doppler velocity feature data, and reflected energy feature data of the reflection point cloud, respectively. For example, the original spatial coordinate values (x0, y0, z0) of the radar output signal can be subjected to processing such as translation and rotation, and the transformed spatial coordinate values (x1, y1, z1) obtained after processing can be used as the first spatial feature data of the reflection point cloud, but the present invention is not limited to this.
[0022] In some embodiments, the human body key points correspond to major joints or organs of the human body, such as the nose, shoulders, elbows, wrists, waist, etc. The present invention is not limited to the selection of human body key points, and other key points may be selected according to the needs of a specific application.
[0023] FIG. 2 is a diagram showing an example of human body key points in an embodiment of the present invention. As shown in FIG. 2, human body key point information may refer to the relative positional relationship of the joints and organs of the human body corresponding to the key points, which can be expressed as H={(x k ,y k ,z k ), 1≦k≦m}, and (x k ,yk ,z k ) is the coordinate of the k-th keypoint. When there is only two-dimensional position information, one dimension of the keypoint coordinate can be set to 0, for example, z is set to 0, and then only (x, y) contains valid position information.
[0024] Although radar point cloud data has been described as an example, the present invention is not limited to this. Density distribution feature data and time distribution feature data can also be obtained based on radar point cloud data.
[0025] The distribution position of the radar point cloud is random, and the density of the radar point cloud at different positions varies with different movements. For example, when the movement range of the upper limbs of the detection target is relatively large, the radar point cloud is relatively concentrated at the position where the upper limbs are located, and the point cloud distribution at other positions is relatively sparse.
[0026] In some embodiments, obtaining density distribution feature data of the reflectance point cloud includes calculating a distance (value) between each reflectance point and other reflectance points; and calculating a probability corresponding to the distance using a probability density function, and setting the probability as the density distribution feature data.
[0027] FIG. 3 is a diagram showing the distance between reflection points in an embodiment of the present invention. Here, the average value is taken as an example, but the present invention is not limited to this. As shown in FIG. 3, for N*3 reflection points, by calculating N*N distances between the reflection points, it is possible to obtain the average value of N*1 distances. This allows the density d of single reflection points to be calculated. i (i=1,…,N) can be obtained.
[0028] For example, the probability density function is
number
[0029] In some embodiments, different variances correspond to different probabilities, thereby obtaining multiple density distribution feature data.
[0030] For example, when μ=0 and σ=0.1, one density distribution feature data can be obtained; correspondingly, when σ=0.2, another density distribution feature data can be obtained; when σ=0.4, another density distribution feature data can be obtained; and when σ=0.6, still another density distribution feature data can be obtained. Thus, four different density distribution features can be obtained for each reflection point, forming N*4-dimensional density distribution feature data, which can then be used as input features for different channels in a neural network configuration.
[0031] In some embodiments, obtaining time distribution feature data of the reflectance point cloud includes calculating a time (value) corresponding to each reflectance point; and calculating a probability corresponding to the time using a time density function, and taking the probability as the time distribution feature data.
[0032] For example, the frame number of each reflection point can be directly obtained from the radar output signal (for example, the time interval of the radar output signal is 0.1 s, i.e., 1 frame / 0.1 s). By calculating the distribution feature of the point cloud frame number using the frame number of the radar reflection point (1, 2, ...), the distribution feature of the time series of the radar reflection point cloud can be obtained.
[0033] For example, the time density function is
number
[0034] In some embodiments, different variances correspond to different probabilities, thereby providing multiple time distribution feature data.
[0035] For example, f i is the frame number (1, 2, ...) of the reflection point, representing the temporal continuity information of the point cloud. When μ = 0 and σ = 0.5, one time distribution feature data can be obtained; correspondingly, when σ = 1.0, another time distribution feature data can be obtained; when σ = 1.5, another time distribution feature data can be obtained; and when σ = 2.0, yet another time distribution feature data can be obtained. Thus, by obtaining four different time distribution features for each reflection point, N*4-dimensional time distribution feature data can be formed, which can then be used as input features for different channels in the neural network configuration.
[0036] The first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data have been described above as examples. An embodiment of the present invention can recognize keypoints based on these feature data.
[0037] 4 is a diagram illustrating a plurality of network models according to an embodiment of the present invention. As shown in FIG. 4, a total of N*13-dimensional feature data including first spatial feature data (e.g., N*3 dimensions), Doppler velocity feature data (e.g., N*1 dimension), reflected energy feature data (e.g., N*1 dimension), density distribution feature data (e.g., N*4 dimensions), and time distribution feature data (e.g., N*4 dimensions) can be used as input data for the fusion feature extraction model 401.
[0038] As shown in Fig. 4, the fusion feature extraction model 401 outputs feature data, which is used as input data for the keypoint detection model 402. As shown in Fig. 4, the keypoint detection model 402 outputs keypoint information, thereby realizing keypoint recognition of an object.
[0039] This allows the detection of keypoints of objects (e.g., human bodies) based on radar point clouds to be unrestricted by action category, requiring fewer computational resources and achieving high detection accuracy. Furthermore, by performing deep analysis of feature data using multiple network models, it is possible to distinguish between feature differences between different actions, resulting in more accurate recognition of keypoint information for different actions.
[0040] In some embodiments, the first spatial feature data (original spatial feature data) may be transformed using a spatial transformation model based on a neural network to obtain normalized second spatial feature data; and the normalized second spatial feature data may be cascaded with the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data, and the cascaded feature data may be input to the fusion feature extraction model.
[0041] Fig. 5 is a diagram showing feature data in an embodiment of the present invention. As shown in Fig. 5, feature extraction is performed on point cloud data acquired within a predetermined period (see 501), and first spatial feature data (original spatial feature data (see 502)), Doppler velocity feature data (see 503), reflected energy feature data (see 504), density distribution feature data (see 505), and time distribution feature data (see 506) can be acquired.
[0042] 5, the first spatial feature data (original spatial feature data) can be further transformed using a spatial transformation model based on a neural network to obtain normalized second spatial feature data (see 507). Then, the normalized second spatial feature data can be cascaded with the first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data.
[0043] As shown in FIG. 5 , the first spatial feature data (original spatial feature data), Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data can be jointly input as input 1 (input 1), and the normalized second spatial feature data as input 2 (input 2) to a fusion feature extraction model 508.
[0044] 6 is a diagram illustrating the transformation of spatial feature data in an embodiment of the present invention. As shown in FIG. 6, feature expansion is performed on N*3-dimensional first spatial feature data (see 601) to obtain N*M-dimensional expanded spatial feature data (see 602); pooling (e.g., max pooling) is performed on the N*M-dimensional expanded spatial feature data to obtain 1*M-dimensional global spatial feature data (see 603); feature transformation is performed on the 1*M-dimensional global spatial feature data to obtain 1*9-dimensional feature data, which is then transformed into 3*3-dimensional spatial feature data (see 604); and matrix multiplication of the N*3-dimensional first spatial feature data and the 3*3-dimensional spatial feature data is performed to obtain the normalized second spatial feature data (see 605), where N is a positive integer greater than 1 and M is a positive integer greater than 3.
[0045] 7 is a diagram illustrating a spatial transformation model based on a neural network in an embodiment of the present invention. As shown in FIG. 7, for example, feature expansion can be performed on 200*3-dimensional first spatial feature data using a one-dimensional convolutional neural network (see 701) to obtain 200*256-dimensional expanded spatial feature data; max pooling can be performed on the 200*256-dimensional expanded spatial feature data (see 702) to obtain 1*256-dimensional global spatial feature data; feature transformation can be performed on the 1*256-dimensional global spatial feature data using a fully connected neural network (see 703) to obtain 1*64-dimensional feature data, which can then be converted into 1*9-dimensional spatial feature data using a fully connected neural network again, and the 1*9-dimensional feature data can then be converted into 3*3-dimensional spatial feature data (see 704); and matrix multiplication can be performed on the 200*3-dimensional first spatial feature data and the 3*3-dimensional spatial feature data (see 705) to obtain normalized 200*3-dimensional second spatial feature data.
[0046] 8 is another diagram illustrating multiple network models in an embodiment of the present invention. As shown in FIG. 8, a total of N*13-dimensional feature data including first spatial feature data (e.g., N*3 dimensions), Doppler velocity feature data (e.g., N*1 dimension), reflected energy feature data (e.g., N*1 dimension), density distribution feature data (e.g., N*4 dimensions), and time distribution feature data (e.g., N*4 dimensions) can be used as input 1 of a fusion feature extraction model 801.
[0047] As shown in FIG. 8 , a spatial transformation model 802 based on a neural network is further used to transform the first spatial feature data (original spatial feature data) to obtain normalized second spatial feature data (e.g., N*3 dimensions), which can then be used as input 2 of the fusion feature extraction model 801.
[0048] As shown in Fig. 8, the fusion feature extraction model 801 outputs feature data, which is used as input data for the keypoint detection model 803. As shown in Fig. 8, the keypoint detection model 803 outputs keypoint information, thereby realizing recognition of the keypoints of an object.
[0049] This allows the detection of key points of an object (e.g., a human body) based on radar point clouds to be unrestricted by action category, requiring fewer computational resources and further improving detection accuracy. Furthermore, by performing deep analysis of feature data using multiple network models, it is possible to distinguish feature differences between different actions, resulting in more accurate recognition of key point information for different actions.
[0050] In some embodiments, using a neural network-based feature fusion extraction model to detect cascaded feature data and obtain fused feature information includes: performing feature expansion on the N*L-dimensional cascaded feature data to obtain N*K-dimensional expanded feature data; performing pooling on the N*K-dimensional expanded feature data to obtain 1*K-dimensional global feature data; and performing feature transformation on the 1*K-dimensional global feature data to obtain 1*L-dimensional global feature data. 2 The method includes: obtaining N*L-dimensional feature data, converting the obtained N*L-dimensional feature data into L*L-dimensional feature data; and performing matrix multiplication between the N*L-dimensional cascaded feature data and the L*L-dimensional feature data to obtain N*L-dimensional fused feature information, where N, L, and K are all positive integers greater than 1, and L is less than K.
[0051] 9 is a diagram illustrating a neural network-based fusion feature extraction model according to an embodiment of the present invention. As shown in FIG. 9, for example, the first spatial feature data (200*3 dimensions), the Doppler velocity feature data (200*1 dimension), the reflected energy feature data (200*1 dimension), the density distribution feature data (200*4 dimensions), and the time distribution feature data (200*4 dimensions) are used as a 200*13-dimensional input 1, and the normalized second spatial feature data (200*3 dimensions) is used as a 200*3-dimensional input 2. Then, input 1 and input 2 are cascaded (see 901) to form 200*16-dimensional fusion feature information, which is used as the input data for the fusion feature extraction model.
[0052] As shown in Figure 9, feature expansion is performed on the 200*16 dimensional fused feature data using a 1D convolutional neural network (see 902), obtaining 200*64 dimensional feature data (Feature 1), and then obtaining 200*256 dimensional feature data (see 903); max pooling is performed on the 200*256 dimensional feature data (see 904), obtaining 1*256 dimensional feature data; feature transformation is performed on the 1*256 dimensional feature data using a fully connected neural network (see 905), obtaining 1*64 dimensional feature data, and then converting the 1*4096 dimensional feature data into 64*64 dimensional feature data (Feature 2). 2)) (see 906); and then, matrix multiplication of the 200*64 dimensional feature data 1 and the 64*64 dimensional feature data 2 is performed (see 907), thereby obtaining 200*64 dimensional feature data (fused feature information (see 908)).
[0053] In some embodiments, using a neural network-based keypoint detection model to detect fusion feature information and output keypoint data of the object includes: performing feature expansion on the N*L-dimensional fusion feature information to obtain N*J-dimensional expanded feature data; performing a pooling operation on the N*J-dimensional expanded feature data to obtain 1*J-dimensional global feature data; and performing feature transformation on the 1*J-dimensional global feature data to obtain 1*P-dimensional keypoint location information, where N, L, and J are all positive integers greater than 1, and L is less than J.
[0054] 10 is a diagram illustrating a keypoint detection model based on a neural network in an embodiment of the present invention. As shown in FIG. 10, for example, the 200*64 dimensional feature data output by the fusion feature extraction model can be used as input data for the keypoint detection model.
[0055] As shown in FIG. 10, feature expansion is performed on 200*64 dimensional feature data using a one-dimensional convolutional neural network (see 1001), obtaining 200*128 dimensional feature data, and then obtaining 200*256 dimensional feature data (see 1002); max pooling is performed on the 200*256 dimensional feature data (see 1003), obtaining 1*256 dimensional feature data; feature transformation is performed on the 1*256 dimensional feature data using a fully connected neural network (see 1004), obtaining 1*64 dimensional feature data, and then converting it into 1*34 dimensional feature data using a fully connected neural network again (see 1005), after which keypoint information can be output.
[0056] In the embodiments of the present invention, the following has been found through simulation experiments: By using N*13-dimensional feature data and performing feature extraction and detection using multiple network models, keypoint information can be accurately detected, and the detection accuracy can be improved compared to when the original N*6-dimensional data is used. Furthermore, by using N*13-dimensional feature data and performing feature extraction and detection using multiple network models, the detection accuracy can be further improved.
[0057] The above is an exemplary description of a recognition method and model. In an embodiment of the present invention, a supervised training method is used to obtain one or more sets of optimal parameters; the parameters are then applied to a detection model, and calculations are performed on the input radar point cloud data to obtain corresponding human body keypoint information. The embodiment of the present invention does not limit the specific training method of the model; for example, Stochastic Gradient Descent (SGD) optimization, Adaptive Moment Estimation (Adam) optimization, etc. can be used.
[0058] Furthermore, although the above describes each step or process directly related to the present invention, the present invention is not limited thereto. The motion detection method may further include other steps or processes. For the specific content of these steps or processes, reference may be made to the prior art. Furthermore, although the above describes exemplary embodiments of the present invention using several configurations of the motion detection model as examples, the present invention is not limited to these configurations, and appropriate modifications can be made to these configurations, and all such modifications fall within the scope of the embodiments of the present invention.
[0059] The above-described embodiments are merely illustrative of the present invention, and the present invention is not limited thereto. Appropriate modifications can be made based on the above-described embodiments. For example, the above-described embodiments can be used alone, or one or more of the above-described embodiments can be used in combination.
[0060] As can be seen from the above embodiments, feature extraction is performed on point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; a neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fusion feature information; and a neural network-based keypoint detection model is used to detect the fusion feature information and output keypoint data of the object, thereby enabling keypoints of an object (e.g., a human body) to be detected based on the radar point cloud, and is not limited to action categories, requires few computing resources, and has high detection accuracy; furthermore, it is easy to implement and operate, has high noise resistance, and has excellent privacy protection.
[0061] <Example of the second aspect> An embodiment of the present invention provides a keypoint recognition device based on a radio radar signal, in which the same content as the embodiment of the first aspect is omitted.
[0062] 11 is a diagram illustrating a keypoint recognition device based on a wireless radar signal in an embodiment of the present invention. As shown in FIG. 11, a keypoint recognition device 1100 based on a wireless radar signal includes:
[0063] A detection unit 1101: detects an object by radar and obtains point cloud data; A feature extraction unit 1102: performs feature extraction on the point cloud data acquired within a predetermined period of time, and obtains first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; a cascade connection unit 1103: performing cascade connection on the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; A feature detection unit 1104: detects the cascaded feature data using a neural network-based feature fusion extraction model to obtain fusion feature information; and Keypoint detection unit 1105: detects the fused feature information using a neural network-based keypoint detection model, and outputs keypoint data of the object.
[0064] In some embodiments, as shown in FIG. 11, the wireless radar signal based keypoint recognition device 1100 may further include:
[0065] A feature transformation unit 1106: performs transformation on the first spatial feature data using a spatial transformation model based on a neural network to obtain a normalized second spatial feature data.
[0066] The cascade connection unit 1103 is further used to cascade the normalized second spatial feature data with the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data, and input the cascaded feature data into the fusion feature extraction model.
[0067] In some embodiments, the feature transformation unit 1106 comprises: Feature expansion is performed on the N*3-dimensional first spatial feature data to obtain N*M-dimensional expanded spatial feature data; Perform pooling operation on the N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data; Perform feature transformation on the 1*M-dimensional global spatial feature data to obtain 1*9-dimensional feature data, which is then transformed into 3*3-dimensional spatial feature data; and performing matrix multiplication of the N*3-dimensional first spatial feature data and the 3*3-dimensional spatial feature data to obtain the normalized second spatial feature data; In this, N is a positive integer greater than 1, and M is a positive integer greater than 3.
[0068] In some embodiments, the feature detection unit 1104 comprises: Perform feature expansion on the N*L dimensional cascaded feature data to obtain N*K dimensional expanded feature data; Perform pooling on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data; Feature transformation is performed on the 1*K dimensional global feature data, and 1*L 2 After obtaining the dimensional feature data, convert it into L*L dimensional feature data; and used for performing matrix multiplication of the N*L-dimensional cascaded feature data and the L*L-dimensional feature data to obtain N*L-dimensional fused feature information; Among them, N, L, and K are all positive integers greater than 1, and L is smaller than K.
[0069] In some embodiments, the feature extraction unit 1102 obtaining the density distribution feature data of the reflectance point cloud comprises: Calculate the distance between each reflecting point and other reflecting points; and The method includes calculating a probability corresponding to the distance using a probability density function, and setting the probability as the density distribution feature data.
[0070] In some embodiments, the probability density function is:
number
[0071] Wherein, different variances correspond to different probabilities, thereby obtaining a plurality of said density distribution feature data.
[0072] In some embodiments, the feature extraction unit 1102 obtaining the time distribution feature data of the reflectance point cloud comprises: Calculate the time corresponding to each reflection point; and Using the time density function The time and determining the probability as the time distribution feature data.
[0073] In some embodiments, the time density function is:
number
[0074] Wherein, different variance values correspond to different probabilities, thereby obtaining a plurality of the time distribution feature data.
[0075] In some embodiments, the keypoint detection unit 1105 comprises: Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data; Perform pooling on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data; and It is used to perform feature transformation on 1*J-dimensional global feature data to obtain 1*P-dimensional keypoint position information. Among them, N, L, J, and P are all positive integers greater than 1, and L is smaller than J, and P is smaller than J.
[0076] Although the components or modules directly related to the present invention have been described above, the present invention is not limited thereto. The keypoint recognition device 1100 based on wireless radar signals may further include other components or modules. For details of these components or modules, please refer to the related art.
[0077] 11 shows only the connection relationships or signal directions between the components or modules, but it should be understood by those skilled in the art that various related technologies, such as a bus-based connection, may also be used. Each of these components or modules may be realized by hardware, such as a processor, a memory, etc., and the embodiments of the present invention are not limited thereto.
[0078] The above-described embodiments are merely illustrative of the present invention, and the present invention is not limited thereto. Appropriate modifications can be made based on the above-described embodiments. For example, the above-described embodiments may be used alone, or one or more of the above-described embodiments may be used in combination.
[0079] As can be seen from the above embodiments, feature extraction is performed on point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; a neural network-based fusion feature extraction model is used to detect the cascaded feature data to obtain fusion feature information; and a neural network-based keypoint detection model is used to detect the fusion feature information and output keypoint data of the object, thereby enabling keypoints of an object (e.g., a human body) to be detected based on the radar point cloud, and is not limited to action categories, requires few computing resources, and has high detection accuracy; furthermore, it is easy to implement and operate, has high noise resistance, and has excellent privacy protection.
[0080] <Example of the third aspect> In an embodiment of the present invention, an electronic device is provided, which includes the keypoint recognition device 1100 based on a wireless radar signal as described in the embodiment of the second aspect, the contents of which are incorporated herein. The electronic device may be, for example, a computer, a server, a workstation, a laptop, a smartphone, etc., but the embodiment of the present invention is not limited thereto.
[0081] 12 is a diagram illustrating an electronic device according to an embodiment of the present invention. As shown in FIG. 12, the electronic device 1200 may include a processor (e.g., a central processing unit (CPU)) 1210 and a memory 1220, which is connected to the central processing unit 1210. The memory 1220 can store various data and can further store a program 1221 for information processing, and can execute the program 1221 under the control of the processor 1210.
[0082] In some embodiments, the functionality of the wireless radar signal-based keypoint recognition device 1100 can be integrated into the processor 1210, where the processor 1210 is configured to implement the method for wireless radar signal-based keypoint recognition described in the embodiments of the first aspect.
[0083] In some embodiments, the wireless radar signal based keypoint recognition device 1100 may be located separately from the processor 1210, for example, the wireless radar signal based keypoint recognition device 1100 may be configured as a chip connected to the processor 1210, and the functions of the wireless radar signal based keypoint recognition device 1100 may be realized under the control of the processor 1210.
[0084] For example, the processor 1210 may be configured to perform the following controls: detect an object using a radar to obtain point cloud data; perform feature extraction on the point cloud data obtained within a predetermined period of time to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; cascade the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; detect the cascaded feature data using a neural network-based fusion feature extraction model to obtain fusion feature information; and detect the fusion feature information using a neural network-based keypoint detection model to output keypoint data of the object.
[0085] 12, the electronic device 1200 may further include an input / output (I / O) device 1230, a display 1240, etc., and the functions of these components are similar to those of the prior art, so detailed descriptions thereof will be omitted here. Note that the electronic device 1200 does not need to include all of the components shown in FIG. 12. The electronic device 1200 may also include components not shown in FIG. 12, and reference can be made to related art for this information.
[0086] An embodiment of the present invention further provides a computer-readable program, which, when executed by an electronic device, causes a computer to perform the method for recognizing keypoints based on wireless radar signals described in the embodiment of the first aspect on the electronic device.
[0087] An embodiment of the present invention further provides a storage medium storing a computer-readable program, wherein the computer-readable program causes a computer to execute the method for recognizing keypoints based on wireless radar signals described in the embodiment of the first aspect in an electronic device.
[0088] The above-described apparatus and method may be realized by software or hardware, or by a combination of hardware and software. The present invention further relates to a computer-readable program as described below, which, when executed by a logic component, causes the logic component to realize the above-described apparatus or component, or to perform the above-described various methods or steps. The logic component may be, for example, an FPGA (Field Programmable Gate Array), a microprocessor, or a processing unit used in a computer. The present invention also relates to a storage medium, such as a hard disk, magnetic disk, optical hard disk, DVD, or flash memory, that stores the above-described program.
[0089] Furthermore, one or more combinations of the functional blocks illustrated in the figures and / or one or more combinations of the functional blocks may be implemented as a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic component, a discrete gate or transistor logic component, a discrete hardware assembly, or any other suitable combination for performing the functions described herein. Also, one or more combinations of the functional blocks illustrated in the figures and / or one or more combinations of the functional blocks may be further configured as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors communicatively coupled with a DSP, or any other configuration.
[0090] Furthermore, the following supplementary notes are disclosed regarding the above-mentioned embodiments.
[0091] (Appendix 1) A keypoint recognition method based on a radio radar signal, comprising: Radar detects objects and captures point cloud data; Perform feature extraction on the point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; cascading the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; Detecting the cascaded feature data using a neural network-based fusion feature extraction model to obtain fusion feature information; and detecting the fused feature information using a neural network-based keypoint detection model and outputting keypoint data of the object.
[0092] (Appendix 2) The method of claim 1, further comprising: Transforming the first spatial feature data using a spatial transformation model based on a neural network to obtain normalized second spatial feature data; and cascading the normalized second spatial feature data with the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data, and inputting the cascaded feature data into the fusion feature extraction model.
[0093] (Appendix 3) 10. The method according to claim 1 or 2, performing a transformation on the first spatial feature data using a neural network-based spatial transformation model; Feature expansion is performed on the N*3-dimensional first spatial feature data to obtain N*M-dimensional expanded spatial feature data; Perform pooling operation on the N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data; Perform feature transformation on the 1*M-dimensional global spatial feature data to obtain 1*9-dimensional feature data, which is then transformed into 3*3-dimensional spatial feature data; and performing matrix multiplication of the N*3-dimensional first spatial feature data and the 3*3-dimensional spatial feature data to obtain the normalized second spatial feature data; A method wherein N is a positive integer greater than 1 and M is a positive integer greater than 3.
[0094] (Appendix 4) 4. The method of any one of claims 1 to 3, comprising: Detecting cascaded feature data using a neural network-based fusion feature extraction model Perform feature expansion on the N*L dimensional cascaded feature data to obtain N*K dimensional expanded feature data; Perform pooling on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data; Feature transformation is performed on the 1*K dimensional global feature data, and 1*L 2 After obtaining the dimensional feature data, convert it into L*L dimensional feature data; and performing matrix multiplication of the N*L-dimensional cascaded feature data and the L*L-dimensional feature data to obtain N*L-dimensional fused feature information; A method wherein N, L and K are all positive integers greater than 1, and L is less than K.
[0095] (Appendix 5) 5. The method of any one of claims 1 to 4, comprising: Obtaining the density distribution feature data of the reflectance point cloud includes: Calculate the distance between the reflecting point and other reflecting points; and calculating a probability corresponding to the distance using a probability density function, and setting the probability as the density distribution feature data.
[0096] (Appendix 6) 6. The method of claim 5, The probability density function is
number
[0097] (Appendix 7) 7. The method of claim 6, The method, wherein different variances correspond to different probabilities, thereby obtaining a plurality of said density distribution feature data.
[0098] (Appendix 8) 8. The method of any one of claims 1 to 7, comprising: Obtaining the time distribution feature data of the reflectance point cloud includes: Calculate the time corresponding to each reflection point; and calculating a probability corresponding to the time using a time density function, and setting the probability as the time distribution feature data.
[0099] (Appendix 9) 9. The method of claim 8, The time density function is
number
[0100] (Appendix 10) 10. The method of claim 9, The method, wherein different variances correspond to different probabilities, thereby obtaining a plurality of said time distribution feature data.
[0101] (Appendix 11) 11. The method of any one of claims 1 to 10, comprising: Detecting the fused feature information using a neural network-based keypoint detection model includes: Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data; Perform pooling on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data; and Performing feature transformation on the 1*J-dimensional global feature data to obtain 1*P-dimensional keypoint location information; A method in which N, L, J, and P are all positive integers greater than 1, and L is less than J and P is less than J.
[0102] (Appendix 12) An electronic device, a memory and a processor; The storage device stores a computer program, 12. An electronic device, wherein the processor is configured to execute the computer program to implement the method for recognizing keypoints based on a radio radar signal according to any one of Supplementary Notes 1 to 11.
[0103] (Appendix 13) A storage medium storing a computer-readable program, A storage medium, the computer-readable program causing a computer to execute the method for recognizing keypoints based on radio radar signals according to any one of Supplementary Notes 1 to 11 in an electronic device.
[0104] Although the preferred embodiment of the present invention has been described above, the present invention is not limited to this embodiment, and any modification to the present invention falls within the technical scope of the present invention as long as it does not depart from the spirit of the present invention.
Claims
1. A keypoint recognition device based on radio radar signals, comprising: a detection unit that detects objects by radar and obtains point cloud data; a feature extraction unit that performs feature extraction on the point cloud data acquired within a predetermined period of time to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data and time distribution feature data of the reflection point cloud; a cascade connection unit for cascading the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; a feature detection unit that uses a neural network-based fusion feature extraction model to detect the cascaded feature data and obtain fusion feature information; and a keypoint detection unit for detecting the fused feature information using a neural network-based keypoint detection model and outputting keypoint data of the object;
2. 2. The keypoint recognition device according to claim 1, a feature transformation unit that performs transformation on the first spatial feature data using a spatial transformation model based on a neural network to obtain normalized second spatial feature data; the cascade connection unit further cascades the normalized second spatial feature data with the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data, and inputs the cascaded feature data into the fusion feature extraction model.
3. 3. The keypoint recognition device according to claim 2, The feature transformation unit Feature expansion is performed on the N*3-dimensional first spatial feature data to obtain N*M-dimensional expanded spatial feature data; Perform pooling on the N*M dimensional extended spatial feature data to obtain 1*M dimensional global spatial feature data; Perform feature transformation on the 1*M-dimensional global spatial feature data to obtain 1*9-dimensional feature data, which is then transformed into 3*3-dimensional spatial feature data; and performing matrix multiplication of the N*3-dimensional first spatial feature data and the 3*3-dimensional spatial feature data to obtain the normalized second spatial feature data; where N is a positive integer greater than 1 and M is a positive integer greater than 3.
4. 2. The keypoint recognition device according to claim 1, The feature detection unit Feature expansion is performed on the N*L dimensional cascaded feature data to obtain N*K dimensional expanded feature data; Perform pooling on the N*K dimensional extended feature data to obtain 1*K dimensional global feature data; Feature transformation is performed on the 1*K dimensional global feature data, and 1*L 2 After obtaining the dimensional feature data, convert it into L*L dimensional feature data; and Perform matrix multiplication of the N*L-dimensional cascaded feature data and the L*L-dimensional feature data to obtain N*L-dimensional fused feature information; Here, N, L, and K are all positive integers greater than 1, and L is less than K.
5. 2. The keypoint recognition device according to claim 1, The feature extraction unit obtaining the density distribution feature data of the reflectance point cloud includes: Calculating the distance between each reflecting point and other reflecting points; and A keypoint recognition device comprising: calculating a probability corresponding to the distance using a probability density function; and setting the probability as the density distribution feature data.
6. 6. The keypoint recognition device according to claim 5, The probability density function is [Equation 1] and where d i is the distance, μ is a predetermined distance parameter, and σ is the variance, The keypoint recognition device, wherein different distributions correspond to different probabilities, thereby obtaining a plurality of said density distribution feature data.
7. 2. The keypoint recognition device according to claim 1, The feature extraction unit obtaining the time distribution feature data of the reflectance point cloud includes: Calculating the time corresponding to each reflection point; and A keypoint recognition device comprising: calculating a probability corresponding to the time using a time density function; and setting the probability as the time-distributed feature data.
8. 8. The keypoint recognition device according to claim 7, The time density function is [Equation 2] and where f i is the time, μ is a predetermined time parameter, and σ is the variance, The keypoint recognition device, wherein different variances correspond to different probabilities, thereby obtaining a plurality of said time-distributed feature data.
9. 2. The keypoint recognition device according to claim 1, The keypoint detection unit Feature expansion is performed on the N*L dimensional fused feature information to obtain N*J dimensional expanded feature data; Perform a pooling operation on the N*J dimensional extended feature data to obtain 1*J dimensional global feature data; and Feature transformation is performed on the 1*J-dimensional global feature data to obtain 1*P-dimensional keypoint position information. Here, N, L, J, and P are all positive integers greater than 1, and L is less than J and P is less than J. A keypoint recognition device.
10. A keypoint recognition method based on a radio radar signal, comprising: Detect objects with radar and obtain point cloud data; Perform feature extraction on the point cloud data acquired within a predetermined period to obtain first spatial feature data, Doppler velocity feature data, reflected energy feature data, density distribution feature data, and time distribution feature data of the reflection point cloud; cascading the first spatial feature data, the Doppler velocity feature data, the reflected energy feature data, the density distribution feature data, and the time distribution feature data; Detecting the cascaded feature data using a neural network-based feature fusion extraction model to obtain fusion feature information; and a neural network-based keypoint detection model for detecting the fused feature information and outputting keypoint data of the object;
Citation Information
Patent Citations
Extracting device, method, and program
JP2015184061A
Lidar-based multi-person pose estimation
US20200167954A1