Neural network-based motion detection device and method
The neural network-based motion detection system enhances radar-based human body movement detection by processing point cloud data to improve accuracy and reduce computational resources.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2022-01-19
- Publication Date
- 2026-05-19
AI Technical Summary
Conventional radar-based methods for detecting human body movements are limited in application scenarios due to their inability to detect a wide range of motions and require significant computational resources.
A neural network-based motion detection system that processes radar point cloud data through preprocessing and feature extraction to identify key points of the human body, utilizing a neural network-based motion detection model to enhance detection accuracy without restricting motion classes.
Improves detection accuracy and reduces computational requirements by effectively identifying key points of the human body using radar point clouds.
Smart Images

Figure 0007861407000003 
Figure 0007861407000004 
Figure 0007861407000005
Abstract
Description
[Technical Field]
[0001] Embodiments of the present invention relate to the technical field of radar detection. [Background technology]
[0002] The process of detecting human body movements can identify key points of the human body, such as the head, neck, arms, legs, and waist. Detection of key points of the human body has a wide range of application scenarios and is an important technology in applications such as smart homes, health monitoring, and behavioral understanding. Currently, video-based human body key point detection technology is widely used. However, video seriously infringes on privacy and cannot be applied in private situations. Furthermore, video is highly susceptible to environmental influences and cannot function effectively in dark or obstructed scenarios.
[0003] Radar can detect objects (e.g., human bodies) using radio signals, without exposing privacy, without depending on lighting conditions, and can function effectively in partially obscured scenarios. Therefore, radar-based keypoint detection can compensate for the shortcomings of video technology.
[0004] The above-mentioned explanation of the technical background is intended to provide a clear and complete understanding of the proposed technical aspects of the present invention, and is written to be understandable to those skilled in the art. These proposed technical aspects are merely described as background art for the present invention and are not well known to those skilled in the art. [Overview of the project] [Problems that the invention aims to solve]
[0005] However, according to the inventors of this invention, conventional radar-based methods for detecting human body movements can only detect a small number of specific movements, severely limiting their application scenarios.
[0006] In view of at least one of the above technical problems, embodiments of the present invention provide a neural network-based motion detection device and method that can detect keypoints of an object (e.g., the human body) based on a radar point cloud, without limiting the motion class, reducing the required computational resources, and improving the accuracy of detection. [Means for solving the problem]
[0007] In one embodiment of the present invention, a neural network-based motion detection device is provided, comprising: a sensing unit that senses an object using radar and acquires point cloud data; a preprocessing unit that performs preprocessing on the point cloud data acquired within a predetermined period and acquires multidimensional feature data of a plurality of rearranged points; and a detection unit that performs motion detection on the feature data using a neural network-based motion detection model and outputs keypoint data of the object.
[0008] Another embodiment of the present invention provides a neural network-based motion detection method, comprising the steps of: sensing an object with radar and acquiring point cloud data; preprocessing the point cloud data acquired within a predetermined period to acquire multidimensional feature data of a plurality of rearranged points; and performing motion detection on the feature data using a neural network-based motion detection model and outputting keypoint data of the object.
[0009] The advantageous effects of the embodiments of the present invention are as follows. An object is sensed by a radar to obtain point cloud data, preprocessing is performed on the point cloud data obtained within a predetermined period, multi-dimensional feature data of a plurality of arranged points is obtained, and motion detection is performed on the feature data using a motion detection model based on a neural network, and keypoint data of the object is output. Thereby, by detecting the keypoints of an object (for example, a human body) based on radar point clouds, it is possible to improve the detection accuracy by reducing the necessary computational resources without restricting the motion class.
[0010] Specific embodiments of the present invention are disclosed in detail as shown in the following description and drawings, showing a manner in which the principles of the present invention can be adopted. Note that the embodiments of the present invention are not limited in scope. The embodiments of the present invention include various changes, modifications, and equivalents within the scope of the gist and content of the appended claims.
Brief Description of the Drawings
[0011] The drawings included herein are for understanding the embodiments of the present invention, form part of this specification, and are for exemplifying the embodiments of the present invention, and explain the principles of the present invention in conjunction with the written description. Note that the drawings described herein are merely for explaining the embodiments of the present invention, and those skilled in the art can easily obtain other drawings based on these drawings. [Figure 1] It is a schematic diagram of a method for motion detection based on a neural network according to an embodiment of the present invention. [Figure 2] It is a schematic diagram of one of the keypoints of a human body according to an embodiment of the present invention. [Figure 3] It is a schematic diagram of data preprocessing in an embodiment of the present invention. [Figure 4] It is a schematic diagram of feature expansion of point cloud data according to an embodiment of the present invention. [Figure 5] It is a schematic diagram of motion detection according to an embodiment of the present invention. [Figure 6] This is a schematic diagram of one of the neural network-based motion detection models according to an embodiment of the present invention. [Figure 7] This is a schematic diagram of one single-point feature extraction layer according to an embodiment of the present invention. [Figure 8] This is a schematic diagram of one local feature extraction layer according to an embodiment of the present invention. [Figure 9] This is a schematic diagram of one global feature extraction layer according to an embodiment of the present invention. [Figure 10] This is a schematic diagram of one feature composite layer according to an embodiment of the present invention. [Figure 11] This is a schematic diagram of one of the motion recognition layers according to an embodiment of the present invention. [Figure 12] This is a schematic diagram of one of the neural network-based motion detection devices according to an embodiment of the present invention. [Figure 13] This is a schematic diagram of an electronic device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0012] The above and other features of the present invention will become clearer from the drawings and the following description. The specification and drawings disclose specific embodiments of the present invention, i.e., some embodiments that conform to the principles of the present invention. However, the present invention is not limited to the embodiments described, and includes all modifications, variations, and equivalents of the claims.
[0013] In embodiments of the present invention, the terms "first" and "second" are used to distinguish different elements by name and do not imply a spatial arrangement or temporal order of these elements, and these elements are not limited to these terms. The term "and / or" includes any one or more of the enumerated terms and any combination thereof. The terms "include," "contain," and "have" mean the presence of the described features, elements, components or members, but do not exclude the presence or addition of one or more other features, elements, components or members.
[0014] In the embodiments of the present invention, the singular forms "one," "the," etc., include the plural forms and mean "one kind" or "one class," and are not limited to "one." Also, the term "the foregoing" includes both singular and plural forms unless explicitly indicated by the context. Also, unless explicitly indicated by the context, the term "correspondingly" means "at least partially in accordance with," and the term "based on" means "at least partially based on."
[0015] Features described and / or shown in one embodiment may be used in the same or similar manner in one or more other embodiments, in combination with features in other embodiments, or in place of features in other embodiments. The terms “inclusive” or “includes” mean the presence of the described features, elements, components or members, but do not exclude the presence or addition of one or more other features, elements, components or members.
[0016] In embodiments of the present invention, the radar may, but is not limited to, a millimeter-wave (mmWave) radar. The radar emits electromagnetic waves through a transmitting antenna, which are reflected by various objects, and the corresponding reflected waves (which may be referred to as radar echo wave information) are received. By analyzing the radar echo wave information, information such as the position of an object from the radar and its radial velocity can be effectively extracted, and this information can meet the needs of many application scenarios.
[0017] In embodiments of the present invention, the object to be detected may be a person of various ages, such as an elderly person, a child, an elderly person and / or nursing staff, or a child and / or guardian. The present invention is not limited to these, and the object to be detected may be an animal with vital signs, or a robot without vital signs. The following will be explained using the human body as an example.
[0018] <Example 1> An embodiment of the present invention provides a neural network-based motion detection method. Figure 1 is a schematic diagram of one neural network-based motion detection method according to an embodiment of the present invention. As shown in Figure 1, the method includes the following steps.
[0019] Step 101: Detect objects using radar and acquire point cloud data.
[0020] Step 102: Preprocess the point cloud data acquired within a specified period to obtain multidimensional feature data of the rearranged points.
[0021] Step 103: Perform motion detection on the feature data using a neural network-based motion detection model and output keypoint data for the object.
[0022] Figure 1 above is merely illustrative to illustrate an embodiment of the present invention, and the present invention is not limited thereto. For example, the execution order between each step may be adjusted as appropriate, other steps may be added, or some steps may be deleted. Those skilled in the art may make modifications based on the above description, and the invention is not limited to the description in Figure 1.
[0023] In some embodiments, radar senses the external space using radio signals, and the point cloud data output by the radar includes distance, velocity, and position information of objects detected in the space. The radar periodically emits radio signals to perform detection, and point cloud information is also output periodically. The point cloud information output each time the radar performs detection may be referred to as a frame of point cloud data.
[0024] In an embodiment of the present invention, continuous radar point cloud data of N frames is taken as input, and after preprocessing of the radar point cloud data and calculation of a motion detection model, keypoint information of the human body to represent human motion is output. The radar point cloud data of N frames (outside 1) It may be represented by TIFF0007861407000001.tif20127. The frame number i is sorted in chronological order, and the larger the frame number, the later the corresponding point cloud data appears. Therefore, P k represents the latest radar point cloud data.
[0025] In some embodiments, the point cloud data P of one frame of the radar consists of several points, and P = {p j , 1 ≤ j ≤ n}, where n is the number of points included in the point cloud of the frame, and p j is the j-th point. The points in the point cloud data are represented by p, and p = (s, v, p, x, y, z), where s is the frame number, v is the Doppler speed relative to the radar, p is the signal strength of the point, and (x, y, z) are its spatial coordinate values.
[0026] In some embodiments, the key points of the human body correspond to the main joints or organs of the human body, such as the nose, shoulders, elbows, wrists, and waist. The present invention does not limit the selection of the key points of the human body, and various key points of the human body may be selected according to the needs of specific applications.
[0027] FIG. 2 is a schematic diagram of one of the key points of the human body according to an embodiment of the present invention. As shown in FIG. 2, the key point information of the human body may mean the relative positional relationship of the joints or organs of the human body corresponding to the key points, and H = {(x k , y k , z k ), 1 ≤ k ≤ m}, and (x k , y k , z k ) are the coordinates of the k-th key point. When there is only two-dimensional position information, the dimension of the key point coordinates may be set to 0. For example, z may be set to 0, so that the valid position information is only included in (x, y).
[0028] The above is an exemplary description of the radar point cloud data, and the following will describe the data preprocessing.
[0029] Figure 3 is a schematic diagram of one data preprocessing step in an embodiment of the present invention. As shown in Figure 3, the preprocessing may include the following steps.
[0030] Step 301: Perform noise filtering on the point cloud data.
[0031] In some embodiments, points whose absolute value of Doppler velocity is less than a velocity threshold T1 may be filtered out (deleted) from the N-frame point cloud data acquired within a predetermined period.
[0032] For example, as mentioned above, the input data is N-frame radar point cloud data. (outside 2) It is represented as TIFF0007861407000002.tif20127, where P i This is the radar point cloud data for the i-th frame. When a person performs a specific action, a point cloud with a specific velocity is generated in the radar. Point clouds with low Doppler velocities do not play a significant role in predicting human actions. Therefore, in the noise filtering operation of the point cloud data, points with an absolute value of Doppler velocity smaller than the threshold T1 may be removed from the radar point cloud data.
[0033] Step 302: Merge the point cloud data.
[0034] In some embodiments, point cloud data from N frames may be merged in chronological order to obtain multiple point cloud data points that appear most recently.
[0035] For example, to solve the problem of insufficient radar point cloud data for a single frame, a point cloud data merging operation is performed to merge the radar point cloud data of the input N frames in descending order of frame number, and obtain a point cloud where the number of frames is less than or equal to the threshold T2. That is, the point cloud data merging operation selects the most recently appearing points with a frame number less than or equal to T2 from the point cloud data of N frames. The point cloud data of the merged frames is P c ={p jIt can also be expressed as {1 ≤ j ≤ C}, where C ≤ T2 is the number of points in the merged point cloud data.
[0036] Step 303: Perform feature enhancement on the point cloud data.
[0037] In some embodiments, feature augmentation may be performed on multiple point cloud data based on spatial position information, and multiple additional features may be obtained for each point cloud data. For example, the feature augmentation operation of point cloud data calculates the spatial characteristics of adjacent regions as additional features of the point cloud based on the spatial position information of the point cloud.
[0038] Figure 4 is a schematic diagram of one feature enhancement of point cloud data according to an embodiment of the present invention. As shown in Figure 4, the feature enhancement includes the following steps.
[0039] Step 401: Select a single point Pj in the point cloud data.
[0040] Step 402: For point Pj, calculate the Euclidean distance between point Pj and all other points.
[0041] For example, merged point cloud data P c For a specific point pj in the data, the point cloud data P merged with pj c Calculate the Euclidean distance between this point and all other points.
[0042] Step 403: Calculate the average value of the smallest D distances among the Euclidean distances.
[0043] For example, assuming D is 4, select the four smallest Euclidean distances from the multiple Euclidean distances calculated in step 402, and calculate the average value of these four Euclidean distances.
[0044] Step 404: Let the mean value be one additional feature of the point Pj, where D is a positive integer.
[0045] As a result, steps 403 and 404 above allow us to extend one feature of point Pj.
[0046] In some embodiments, M distinct values of D are selected for the point Pj, and N additional features corresponding to these M distinct values of D are obtained, where M is a positive integer.
[0047] As shown in Figure 4, feature augmentation may further include the following steps.
[0048] Step 405: Determine whether M D values have already been used. If YES, perform step 407; otherwise, perform step 406.
[0049] Step 406: Update the value of D and continue running Step 403.
[0050] For example, D has three different values (M=3), which are 4, 5, and 6. If D=4, steps 403 and 404 may be performed to obtain one additional feature, and then D may be updated to 5. If D=5, steps 403 and 404 may be performed to obtain another additional feature, and then D may be updated to 6. If D=6, steps 403 and 404 may be performed to obtain yet another additional feature. In this way, three (M=3) additional features can be obtained.
[0051] Step 407: Determine if there are any other points. If there are, perform Step 401 and continue feature augmentation for those other points. If there are no other points, terminate the feature augmentation process.
[0052] Figure 4 above is merely illustrative to illustrate an embodiment of the present invention, and the present invention is not limited thereto. For example, the execution order between each step may be adjusted as appropriate, other steps may be added, or some steps may be deleted. Those skilled in the art may make modifications based on the above description, and the invention is not limited to the description in Figure 4 above.
[0053] As shown in Figure 3, the pretreatment may further include the following steps.
[0054] Step 304: Normalize the features of the point cloud data.
[0055] In some embodiments, normalization operations may be performed on features other than Doppler velocity and height.
[0056] For example, the feature normalization operation for point cloud data involves performing the normalization operation on all features except the Doppler velocity v and height z of the feature-extended point cloud. Specifically, when normalizing feature A of a feature-extended point cloud for one frame, the mean value E of feature A for all points in the point cloud is first calculated. A and standard deviation V A Next, calculate the mean value E from the values of all the point cloud features A. A Subtracting the standard deviation V A Divide by . Feature A may be a point cloud feature (s,p,x,y) and an extended distance-related feature.
[0057] Step 305: Sort the point cloud data.
[0058] In some embodiments, multiple point cloud data are sorted according to the Doppler velocity. For example, the sorting operation of the point cloud data sorts all points in the point cloud data in descending (or ascending) order according to the absolute value of the Doppler velocity. However, the present invention is not limited to this, and for example, the data may be sorted according to signal strength or according to frame number.
[0059] Step 306: Interpolate the point cloud data.
[0060] In some embodiments, if the number of points in the point cloud data is less than a threshold T2, empty data is added so that the number of points becomes equal to the threshold T2, where T2 is a positive integer.
[0061] For example, the point cloud data interpolation operation adds empty data to the point cloud data so that the number of points in the point cloud is equal to T2. So-called empty data means that all feature dimensions of the point cloud data have a value of 0. This operation ensures that the point clouds entering the subsequent motion detection model have the same number of points and the data lengths are the same.
[0062] Figure 3 above is merely illustrative to illustrate an embodiment of the present invention, and the present invention is not limited thereto. For example, the execution order between each step may be adjusted as appropriate, other steps may be added, or some steps may be deleted. Those skilled in the art may modify the above content as appropriate, and are not limited to the description in Figure 3 above.
[0063] According to the point cloud data preprocessing operation described above, N frames of radar point cloud data are integrated into one frame of point cloud data, and the result is P'={p j It can also be expressed as ',1≦j≦T2}, p j '=(s j ',v j ,p j ',x j ',y j ',z j d j,1 ',d j,2 ',…,d j,E ') is the preprocessed point data. Here, v j and z j s are the Doppler frequency point and height value of the point, respectively. j 'and p j ' represents the normalized frame number and energy value of the point data, respectively, (x j ',y j ') is the coordinate value on the horizontal plane of the normalized point data, and (d j,1 ',d j,2 ',…,d j,E ') represents the normalized extended features of the point data, and E represents the number of extended features.
[0064] The preprocessed point cloud data P' consists of T2 points, each point having E+6 dimension features. Therefore, P' may represent a two-dimensional matrix of size (T2, E+6), where each row represents the data for one point in the point cloud, and each column represents the data for all points in the point cloud at a specific feature dimension.
[0065] The above provides an illustrative explanation of data preprocessing; the following explains motion detection.
[0066] Figure 5 is a schematic diagram of one example of motion detection according to the present invention. As shown in Figure 5, the motion detection includes the following steps.
[0067] Step 501: Perform convolution and nonlinear operations on a single point feature to obtain single-point feature data.
[0068] Step 502: Perform convolution and nonlinear operations on the features of multiple adjacent points to obtain local feature data.
[0069] Step 503: Perform convolution and nonlinear operations on all point features to obtain global feature data.
[0070] Step 504: Perform feature merging on the single-point feature data, the local feature data, and the global feature data.
[0071] Step 505: Perform calculations on the merged feature data and output the keypoint data.
[0072] Figure 5 above is merely illustrative to illustrate an embodiment of the present invention, and the present invention is not limited thereto. For example, the execution order between each step may be adjusted as appropriate, other steps may be added, or some steps may be deleted. Those skilled in the art may make modifications based on the above description, and the invention is not limited to the description in Figure 5 above.
[0073] In the embodiments of the present invention, the motion detection model is a neural network-based model, and after inputting the preprocessed point cloud data P' into the motion detection model, the keypoint information of the human body at the corresponding time point in the point cloud H={(x k ,y k ,z k ),1≦k≦m} can also be output, where (x k ,y k ,z k ) is the spatial coordinate of the k-th key point of the human body.
[0074] Figure 6 is a schematic diagram of one neural network-based motion detection model according to an embodiment of the present invention. As shown in Figure 6, the neural network-based motion detection model includes the following layers.
[0075] The single-point feature extraction layer 601 performs convolution and nonlinear operations on the features of a single point to obtain single-point feature data.
[0076] The local feature extraction layer 602 obtains local feature data by performing convolution and nonlinear operations on the features of multiple adjacent points.
[0077] The global feature extraction layer 603 performs convolution and nonlinear operations on the features of all points to obtain global feature data.
[0078] The feature merging layer 604 performs feature merging on the single-point feature data, the local feature data, and the global feature data.
[0079] The motion recognition layer 605 performs calculations on the merged feature data and outputs the keypoint data.
[0080] In some embodiments, the single-point feature extraction layer 601 performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on single-point features, and the size of the convolution kernel is 1 when performing convolution and nonlinear operations on single-point features.
[0081] Figure 7 is a schematic diagram of one single-point feature extraction layer according to an embodiment of the present invention. As shown in Figure 7, the first layer is a one-dimensional batch normalization layer (1D Batch Normalization Layer) 701, which normalizes the data in each input column (i.e., the data for each feature in the point cloud).
[0082] As shown in Figure 7, the second layer is a one-dimensional convolutional layer (1D Convolutional Layer) 702, which performs convolution and nonlinear operations on the data for all features at each point, where the convolution kernel size is 1, the step size is 1, and the number of output channels is N1.
[0083] The input to the single-point feature extraction layer 601 is pre-processed point cloud data P', which is a two-dimensional matrix with size (T2, E+6). After processing by the single-point feature extraction layer 601 shown in Figure 7, single-point feature data F1 is obtained, with size (T2, N1), and may be used as input to the subsequent local feature extraction layer 602. To obtain more complex single-point features, the structure shown in Figure 7 may be superimposed multiple times to construct a more complex single-point feature extraction layer 601.
[0084] In some embodiments, the local feature extraction layer 602 performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on the features of multiple adjacent points. When performing convolution and nonlinear operations on the features of multiple adjacent points, the size of the convolution kernel is an integer greater than 1.
[0085] Figure 8 is a schematic diagram of one local feature extraction layer according to an embodiment of the present invention. As shown in Figure 8, the first layer is a one-dimensional batch normalization layer 801, which normalizes the data in each input column (i.e., the data in each channel of F1). The second layer is a one-dimensional convolution layer 802, which performs convolution and nonlinear operations on K1 rows of F1 data, where the convolution kernel size is K1, the step size is 1, the number of output channels is N2, and K1 is an integer greater than 1.
[0086] After processing by the local feature extraction layer 602 shown in Figure 8, local feature data F2 of the point cloud is obtained, with size (T2, N2), and may be used as input for the subsequent global feature extraction layer 603. To obtain more complex local features of the point cloud, the structure shown in Figure 8 may be superimposed multiple times to construct a more complex local feature extraction layer 602.
[0087] In some embodiments, the global feature extraction layer 603 transposes the single-point feature data or local feature data before performing convolution and nonlinear operations on all point features, and then performs one-dimensional batch normalization on the transposed feature data. When performing convolution and nonlinear operations on all point features, the size of the convolution kernel is an integer greater than 0.
[0088] Figure 9 is a schematic diagram of one global feature extraction layer according to an embodiment of the present invention. As shown in Figure 9, the first layer performs transposition 901 on the input data (i.e., local feature data F2 of the point cloud) to obtain point cloud data of size (N2, T2). Here, each row represents the data in a specific channel of all neighboring point clouds in the point cloud, and each column represents the data in each channel of a group of neighboring point clouds in the point cloud.
[0089] As shown in Figure 9, the second layer is a one-dimensional batch normalization layer 902, which normalizes the data in each column output by the first layer (i.e., the data in each channel of the neighboring point cloud). The third layer is a one-dimensional convolutional layer 903, which performs convolution and nonlinear operations on the output data of the upper layer with K2 rows (i.e., all neighboring point cloud data in K2 channels), where the convolution kernel size is K2, the step size is 1, and the number of channels is N3. K2 is an integer greater than 0.
[0090] After processing by the global feature extraction layer 603 shown in Figure 9, global feature data F3 of the point cloud is obtained, with size (N2, N3). To obtain more complex global features of the point cloud, the second and third layers shown in Figure 9 may be superimposed multiple times to obtain a more complex global feature extraction layer 603.
[0091] In some embodiments, the feature merging layer 604 may perform pooling on the single-point feature data, the local feature data, and the global feature data before performing feature merging on the single-point feature data, the local feature data, and the global feature data.
[0092] In some embodiments, when the feature merging layer 604 performs feature merging on the single-point feature data, local feature data, and global feature data, it sums the pooled single-point feature data, local feature data, and global feature data, and expands the summed data into a one-dimensional vector.
[0093] Figure 10 is a schematic diagram of one feature merging layer according to an embodiment of the present invention. As shown in Figure 10, the feature merging layer 604 merges a single point feature F1, a local point cloud feature F2, and a global point cloud feature F3 to form a point cloud feature F for subsequent action recognition.
[0094] As shown in Figure 10, a single-point feature F1 is a data matrix of size (T2, N1), and after processing by the pooling layer 1001, some features are extracted to obtain a data matrix of size (T2, N4). A local feature F2 in the point cloud is a data matrix of size (T2, N2), and after processing by the pooling layer 1002, some features are extracted to obtain a data matrix of size (T2, N4). F3 is a data matrix of size (N2, N3), and after processing by the data transpose layer 1003 and the pooling layer 1004, some features are extracted to obtain a data matrix of size (T2, N4).
[0095] Next, as shown in Figure 10, the fifth layer 1005 sums all the pooled results and expands them into a one-dimensional vector to obtain a feature F for subsequent action recognition. F is a one-dimensional vector of size T2 × N4. Embodiments of the present invention are not limited to the specific algorithm used by the pooling layer, and may employ, for example, conventional max pooling or average pooling.
[0096] In some embodiments, the motion recognition layer 605 performs weighting and nonlinear calculations on the merged feature data using fully connected connections when performing calculations on the merged feature data and outputting the keypoint data.
[0097] Figure 11 is a schematic diagram of one motion recognition layer according to an embodiment of the present invention. As shown in Figure 11, the motion recognition layer 605 takes merged feature F as input and outputs calculated human body keypoint data H. As shown in Figure 11, the fully connected layer 1101 performs weighting and nonlinear calculations on the merged feature F and outputs human body keypoint data. To adapt to more complex situations, the structure shown in Figure 11 may be superimposed multiple times to obtain a multi-layer fully connected network for forming a more complex motion recognition layer 605.
[0098] The above provides an illustrative description of a motion detection method and motion detection model. The motion detection model includes numerous parameters. In embodiments of the present invention, one or more sets of optimal parameters may be obtained by a supervised training method, and these parameters may be applied to the motion detection model to perform calculations on input radar point cloud data and obtain corresponding keypoint information of the human body. Embodiments of the present invention are not limited to the specific training of the model, and may also use, for example, SGD (Stochastic Gradient Descent) optimization, Adam (Adaptive Moment Estimation) optimization, etc.
[0099] The above merely describes steps or processes related to the present invention, and the present invention is not limited thereto. The motion detection method may further include other steps or processes, and the specific details of these steps or processes may be referenced to the prior art. Furthermore, the above merely illustrates embodiments of the present invention using several structures of motion detection models as examples, and the present invention is not limited to these structures, and appropriate modifications may be made to these structures, and these modifications should be included within the scope of embodiments of the present invention.
[0100] The above embodiments are merely illustrative examples illustrating embodiments of the present invention, and the present invention is not limited thereto. Appropriate modifications may be made based on the various embodiments described above. For example, each of the above embodiments may be used individually, or one or more of the above embodiments may be used in combination.
[0101] According to this embodiment, an object is detected by radar and point cloud data is acquired. Preprocessing is performed on the point cloud data acquired within a predetermined period to obtain multidimensional feature data of multiple rearranged points. Motion detection is performed on the feature data using a motion detection model based on a neural network, and keypoint data of the object is output. As a result, by detecting keypoints of an object (e.g., the human body) based on radar point clouds, it is possible to reduce the required computational resources without limiting the motion class and to improve the accuracy of detection.
[0102] <Example 2> The embodiments of the present invention provide a motion detection device based on a neural network, and the description of the contents is the same as in Embodiment 1, so the explanation will be omitted.
[0103] Figure 12 is a schematic diagram of one neural network-based motion detection device according to an embodiment of the present invention. As shown in Figure 12, the neural network-based motion detection device 1200 includes the following parts.
[0104] The sensing unit 1201 detects objects using radar and acquires point cloud data.
[0105] The preprocessing unit 1202 performs preprocessing on the point cloud data acquired within a predetermined period to obtain multidimensional feature data of the rearranged points.
[0106] The detection unit 1203 performs motion detection on the feature data using a neural network-based motion detection model and outputs keypoint data for the object.
[0107] In some embodiments, the preprocessor 1202 filters out points from the N frames of point cloud data acquired within a predetermined period in which the absolute value of the Doppler velocity is less than the velocity threshold T1, and / or merges the N frames of point cloud data in chronological order to obtain the most recently appearing point cloud data. Here, both N and T1 are positive integers.
[0108] In some embodiments, the preprocessor 1202 performs feature augmentation on multiple point cloud data based on spatial position information, obtains multiple additional features for each point cloud data, and performs a normalization operation on features other than Doppler velocity and height.
[0109] In some embodiments, the feature extension includes, for a point P, calculating the Euclidean distance between point P and all other points, calculating the average of the smallest D distances among the Euclidean distances, and taking the average as one additional feature of point P, where D is a positive integer, and selecting M distinct values of D for point P, and obtaining M additional features corresponding to the M distinct values of D, where M is a positive integer.
[0110] In some embodiments, the preprocessor 1202 sorts a plurality of point cloud data according to the Doppler velocity, and if the number of points in the point cloud data is less than a number threshold T2, it adds empty data so that the number of points becomes equal to the number threshold T2, where T2 is a positive integer.
[0111] In some embodiments, as shown in Figure 6, the motion detection model based on the neural network includes the following layers:
[0112] The single-point feature extraction layer 601 performs convolution and nonlinear operations on the features of a single point to obtain single-point feature data.
[0113] The local feature extraction layer 602 obtains local feature data by performing convolution and nonlinear operations on the features of multiple adjacent points.
[0114] The global feature extraction layer 603 performs convolution and nonlinear operations on the features of all points to obtain global feature data.
[0115] The feature merging layer 604 performs feature merging on the single-point feature data, the local feature data, and the global feature data.
[0116] The motion recognition layer 605 performs calculations on the merged feature data and outputs the keypoint data.
[0117] In some embodiments, the single-point feature extraction layer 601 performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on single-point features, and the size of the convolution kernel is 1 when performing convolution and nonlinear operations on single-point features.
[0118] The local feature extraction layer 602 performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on the features of multiple adjacent points, and the size of the convolution kernel is an integer greater than 1 when performing convolution and nonlinear operations on the features of multiple adjacent points.
[0119] The global feature extraction layer 603 performs a transpose on single-point feature data or local feature data before performing convolution and nonlinear operations on all point features, then performs one-dimensional batch normalization on the transposed feature data, and when performing convolution and nonlinear operations on all point features, the size of the convolution kernel is an integer greater than 0.
[0120] In some embodiments, the feature merging layer 604 performs pooling on the single-point feature data, the local feature data, and the global feature data, before performing feature merging on the single-point feature data, the local feature data, and the global feature data.
[0121] In some embodiments, the feature merging layer 604 sums the pooled single-point feature data, local feature data, and global feature data, and expands the summed data into a one-dimensional vector.
[0122] In some embodiments, the motion recognition layer 605 uses fully connected layers to perform weighting and nonlinear operations on the merged feature data.
[0123] The above merely describes the various components or modules related to the present invention, and the present invention is not limited thereto. The neural network-based motion detection device 1200 may include other components or modules, and the specific details of these components or modules may be referenced from related technologies.
[0124] For simplicity, Figure 12 merely illustrates the connection relationships or signal directions between each component or module, and it will be apparent to those skilled in the art that various related techniques, such as bus connections, can be used. The various components or modules described above may also be implemented by hardware devices such as processors and memory, and the embodiments of the present invention are not limited thereto.
[0125] The above embodiments are merely illustrative examples illustrating embodiments of the present invention, and the present invention is not limited thereto. Appropriate modifications may be made based on the various embodiments described above. For example, each of the above embodiments may be used individually, or one or more of the above embodiments may be used in combination.
[0126] According to this embodiment, an object is detected by radar and point cloud data is acquired. Preprocessing is performed on the point cloud data acquired within a predetermined period to obtain multidimensional feature data of multiple rearranged points. Motion detection is performed on the feature data using a motion detection model based on a neural network, and keypoint data of the object is output. As a result, by detecting keypoints of an object (e.g., the human body) based on radar point clouds, it is possible to reduce the required computational resources without limiting the motion class and to improve the accuracy of detection.
[0127] <Example 3> An embodiment of the present invention provides an electronic device including a neural network-based motion detection device 1200 according to Embodiment 2, the contents of which are incorporated herein by reference. The electronic device may be, for example, a computer, server, workstation, laptop computer, smartphone, etc., but the embodiments of the present invention are not limited to these.
[0128] Figure 13 is a schematic diagram of an electronic device according to an embodiment of the present invention. As shown in Figure 13, the electronic device 1300 according to an embodiment of the present invention includes a processor (e.g., a central processing unit (CPU)) 1310 and a memory 1320. The memory 1320 is connected to the processor 1310. The memory 1320 may store various data and may further store an information processing program 1321. The program 1321 is executed under the control of the processor 1310.
[0129] In one embodiment, the functions of the neural network-based motion detection device 1200 may be integrated into the processor 1310. Here, the processor 1310 may be configured to implement the neural network-based motion detection method described in Example 1.
[0130] In another embodiment, the neural network-based motion detection device 1200 may be located with the processor 1310. For example, the neural network-based motion detection device 1200 may be a chip connected to the processor 1310 and configured to realize the functions of the neural network-based motion detection device 1200 through control of the processor 1310.
[0131] For example, the processor 1310 may be configured to detect an object using radar and acquire point cloud data, perform preprocessing on the point cloud data acquired within a predetermined period, acquire multidimensional feature data of a plurality of rearranged points, perform motion detection on the feature data using a motion detection model based on a neural network, and output keypoint data of the object.
[0132] As shown in Figure 13, the electronic device 1300 may further include an input / output (I / O) device 1330 and a display 1340, etc. The functions of these components are the same as in the prior art, and their explanation is omitted here. The electronic device 1300 does not have to include all the components shown in Figure 13. Furthermore, the electronic device 1300 may include components not shown in Figure 13, and may refer to the prior art.
[0133] An embodiment of the present invention provides a computer-readable program that, when the program is executed in an electronic device, causes a computer to execute the neural network-based motion detection method described in Embodiment 1 in the electronic device.
[0134] Embodiments of the present invention further provide a storage medium which stores a computer-readable program that causes a computer to execute the neural network-based motion detection method described in Embodiment 1 in an electronic device.
[0135] The above-described apparatus and method of the present invention may be implemented by hardware, or by combining hardware and software. The present invention relates to a computer-readable program, and when the program is executed by a logic unit, it can cause the logic unit to implement the above-described apparatus or configuration requirements, or to implement the above-described methods or steps. The present invention relates to a storage medium for storing the above-described program, such as a hard disk, magnetic disk, optical disk, DVD, flash memory, etc.
[0136] The methods / apparatus described with reference to embodiments of the present invention may be implemented in hardware, software modules executed by a processor, or a combination of both. For example, one or more functional block diagrams shown in the drawings, or one or more combinations of functional block diagrams, may correspond to each software module of a computer program flow, or to each hardware module. These software modules may correspond to each step shown in the drawings. These hardware modules may be implemented by hardwareizing these software modules, for example, using a field-programmable gate array (FPGA).
[0137] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, mobile hard disk, CD-ROM, or any other form of storage medium known to those skilled in the art. The storage medium may be connected to the processor so that the processor can read information from or write information to the storage medium, or the storage medium may be a component of the processor. The processor and the storage medium reside in an ASIC. The software module may be stored in the memory of the mobile terminal or on a memory card inserted into the mobile terminal. For example, if the device (e.g., a mobile terminal) uses a relatively large capacity MEGA-SIM card or a high-capacity flash memory device, the software module may be stored on the MEGA-SIM card or high-capacity flash memory device.
[0138] One or more functional blocks and / or one or more combinations of functional blocks shown in the drawings may be implemented by a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any suitable combination thereof for performing the functions described in the present invention. One or more functional blocks and / or one or more combinations of functional blocks shown in the drawings may be implemented, for example, by a combination of computing equipment, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, one or more microprocessors combined with DSP communication, or any other configuration.
[0139] Although the present invention has been described above with reference to specific embodiments, the above description is merely illustrative and does not limit the scope of protection of the present invention. Various modifications and changes can be made to the present invention as long as they do not deviate from the spirit and principles of the present invention, and these modifications and changes also fall within the scope of the present invention.
[0140] Furthermore, the following additional information is disclosed regarding embodiments including the above-described examples. (Note 1) A method for detecting motion based on a neural network, The steps include detecting an object using radar and acquiring point cloud data, The steps include: performing preprocessing on point cloud data acquired within a predetermined period to obtain multidimensional feature data of multiple rearranged points; A method comprising the steps of performing motion detection on the feature data using a motion detection model based on a neural network and outputting keypoint data of the object. (Note 2) The aforementioned pretreatment is The method according to Appendix 1, further comprising filtering out (deleting) points from the N-frame point cloud data acquired within the predetermined period in which the absolute value of the Doppler velocity is less than the velocity threshold T1. (Note 3) The aforementioned pretreatment is The method according to Appendix 1 or 2, further comprising merging N-frame point cloud data in chronological order to obtain the most recently occurring multiple point cloud data. (Note 4) The aforementioned pretreatment is This involves performing feature augmentation on multiple point cloud datasets based on spatial position information, and obtaining multiple additional features for each point cloud dataset. The method according to any one of the appendices 1 to 3, comprising performing a normalization operation on features other than Doppler velocity and height. (Note 5) The aforementioned feature extension is, For point P, calculate the Euclidean distance between point P and all other points, Calculate the average value of the minimum D distances among the aforementioned Euclidean distances, This includes making the aforementioned average value one additional feature of point P, Here, D is a positive integer, as described in Appendix 4. (Note 6) The aforementioned feature extension is, The method further includes selecting M distinct values of D for point P and obtaining M additional features corresponding to the M distinct values of D, Here, M is a positive integer, as described in Appendix 5. (Note 7) The aforementioned pretreatment is Sorting multiple point cloud data according to Doppler velocity, If the number of points in the point cloud data is less than a numerical threshold T2, the method further includes adding empty data so that the number of points becomes equal to the numerical threshold T2. Here, T2 is a positive integer, and the method is one of those described in Appendix 1 to 6. (Note 8) The aforementioned motion detection is, This involves obtaining single-point feature data by performing convolution and nonlinear operations on the features of a single point, This involves obtaining local feature data by performing convolution and nonlinear operations on the features of multiple adjacent points, This involves performing convolution and nonlinear operations on the features of all points to obtain global feature data, The single-point feature data, the local feature data, and the global feature data are subjected to feature merging. A method according to any one of the appendices 1 to 7, comprising performing calculations on the merged feature data to output the keypoint data. (Note 9) Before performing convolution and nonlinear operations on single-point features, the input data is subjected to one-dimensional batch normalization. The method described in Appendix 8, wherein the size of the convolution kernel is 1 when performing convolution and nonlinear operations on a single point feature. (Note 10) Before performing convolution and nonlinear operations on the features of multiple adjacent points, the input data is subjected to one-dimensional batch normalization. The method according to Appendix 8 or 9, wherein when performing convolution and nonlinear operations on features of multiple adjacent points, the size of the convolution kernel is an integer greater than 1. (Note 11) Before performing convolution and nonlinear operations on all point features, the single-point feature data or local feature data is transposed, and the transposed feature data is subjected to one-dimensional batch normalization. The method according to any one of the appendices 8 to 10, wherein the size of the convolution kernel is an integer greater than 0 when performing convolution and nonlinear operations on all point features. (Note 12) Before performing feature merging on the single-point feature data, the local feature data, and the global feature data, The steps include performing pooling on the single-point feature data, The steps include performing pooling on the local feature data, The method according to any one of the appendices 8 to 11, further comprising the step of transposing the global feature data and performing pooling. (Note 13) Performing feature merging on the single-point feature data, the local feature data, and the global feature data is: The pooling of the single-point feature data, the local feature data, and the global feature data is performed to sum them up. The method described in Appendix 12, which includes expanding the summed data into a one-dimensional vector. (Note 14) Performing calculations on the merged feature data to output the keypoint data is: The method according to any one of the appendices 8 to 13, comprising performing weighting and nonlinear operations on the merged feature data using a fully connected system. (Note 15) An electronic device comprising a memory in which a computer program is stored and a processor, wherein the processor executes the computer program to realize a neural network-based motion detection method described in any of appendices 1 to 14. (Note 16) A storage medium storing a computer-readable program, wherein the computer-readable program causes a computer to execute an action detection method based on a neural network as described in any of the appendices 1 to 14 on an electronic device.
Claims
1. A motion detection device based on a neural network, A sensing unit that detects objects using radar and acquires point cloud data, A preprocessing unit that performs preprocessing, including sorting, on point cloud data of multiple frames acquired within a predetermined period, and acquires multidimensional feature data of the sorted multiple points, wherein the multidimensional feature data includes Doppler velocity, A device comprising: a detection unit that uses a neural network-based motion detection model to perform calculations to obtain keypoint data for motion detection of an object from the feature data, and outputs the keypoint data of the object at a point in time corresponding to the feature data obtained by the preprocessing.
2. The aforementioned pre-processing unit, From the point cloud data of N frames acquired within the predetermined period, the absolute value of the Doppler velocity is the velocity threshold T. 1 Points smaller than this are filtered out, and / or, The point cloud data from N frames is merged in chronological order, and the most recently appearing multiple point cloud data points are obtained. Here, N and T 1 The apparatus according to claim 1, wherein all of are positive integers.
3. The aforementioned pre-processing unit, Feature augmentation is performed on multiple point cloud data based on spatial position information, and multiple additional features are obtained for each point cloud data. The apparatus according to claim 1, wherein a normalization operation is performed on features other than the Doppler velocity and height, and the Doppler velocity and height, as well as the other features on which the normalization operation has been performed, are used as the multidimensional feature data.
4. The aforementioned feature extension is, For point P, calculate the Euclidean distance between point P and all other points, Calculate the average value of the minimum D distances among the aforementioned Euclidean distances, This includes making the aforementioned average value one additional feature of point P, Here, D is a positive integer, Select M different values of D for point P, and obtain M additional features corresponding to the M different values of D. The apparatus according to claim 3, where M is a positive integer.
5. The aforementioned pre-processing unit, By sorting multiple point cloud data according to the Doppler velocity, The number of points in the aforementioned point cloud data is a threshold T 2 If the number of points is less than the number threshold T 2 Add free data so that it is equal to, Here, T 2 The apparatus according to claim 1, wherein is a positive integer.
6. The aforementioned neural network-based motion detection model is A single-point feature extraction layer that performs convolution and nonlinear operations on the features of a single point to obtain single-point feature data, A local feature extraction layer that performs convolution and nonlinear operations on the features of multiple adjacent points to obtain local feature data, A global feature extraction layer that performs convolution and nonlinear operations on the features of all points to obtain global feature data, A feature merging layer that performs feature merging on the single-point feature data, the local feature data, and the global feature data, The apparatus according to claim 1, further comprising: an action recognition layer that performs calculations on the merged feature data and outputs the keypoint data.
7. The single-point feature extraction layer performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on the single-point feature, and when performing convolution and nonlinear operations on the single-point feature, the size of the convolution kernel is 1. The local feature extraction layer performs one-dimensional batch normalization on the input data before performing convolution and nonlinear operations on the features of multiple adjacent points, and when performing convolution and nonlinear operations on the features of multiple adjacent points, the size of the convolution kernel is an integer greater than 1. The apparatus according to claim 6, wherein the global feature extraction layer transposes single-point feature data or local feature data before performing convolution and nonlinear operations on the features of all points, performs one-dimensional batch normalization on the transposed feature data, and when performing convolution and nonlinear operations on the features of all points, the size of the convolution kernel is an integer greater than 0.
8. The feature merging layer performs pooling on the single-point feature data, the local feature data, and the global feature data before performing feature merging on the single-point feature data, the local feature data, and the global feature data, and then performs transposition and pooling on the global feature data. The apparatus according to claim 6, wherein the feature merging layer sums the pooled single-point feature data, the local feature data, and the global feature data, and expands the summed data into a one-dimensional vector.
9. The apparatus according to claim 6, wherein the motion recognition layer uses fully connected devices to perform weighting and nonlinear calculations on the merged feature data.
10. A method for detecting motion based on a neural network, The steps include detecting an object using radar and acquiring point cloud data, A step of performing preprocessing, including sorting, on point cloud data of multiple frames acquired within a predetermined period, and obtaining multidimensional feature data of the sorted multiple points, wherein the multidimensional feature data includes Doppler velocity, A method comprising the steps of: using a neural network-based motion detection model to perform calculations to obtain keypoint data for motion detection of the object from the feature data, and outputting the keypoint data of the object at a point in time corresponding to the feature data obtained by the preprocessing.