Attitude detection methods and modules, millimeter-wave radar

By converting point cloud data received by millimeter-wave radar into a two-dimensional depth map and using the target model to identify the coordinates of key points, the problem of privacy infringement in visual detection is solved, and effective human posture detection is achieved while protecting privacy.

CN116482639BActive Publication Date: 2026-03-06NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing vision-based human pose recognition methods have privacy issues and are difficult to effectively detect human poses while protecting privacy.

Method used

Millimeter-wave radar is used to receive millimeter-wave point cloud data of the target object. This data is then converted into a two-dimensional depth map. The target model is used to identify the coordinates of key points to achieve attitude detection. A deep neural network is then used for attitude analysis to protect privacy information.

Benefits of technology

Effective human posture detection in low-light environments while protecting the privacy of the monitored individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116482639B_ABST
    Figure CN116482639B_ABST
Patent Text Reader

Abstract

This application provides a posture detection method, comprising: receiving millimeter-wave point cloud data of a target object; converting the millimeter-wave point cloud data into a two-dimensional depth map; inputting the two-dimensional depth map into a target model to obtain the coordinates of key points of the target object, wherein the target model is obtained through transfer learning based on a source model, and the source model contains knowledge for identifying key points; and obtaining the posture of the target object based on the key point coordinates. Furthermore, this application also provides a posture detection module and a millimeter-wave radar. The posture detection method provided by this application can detect the posture of a human body while protecting the privacy of the monitored person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of millimeter-wave radar technology, and in particular to an attitude detection method and its module, and a millimeter-wave radar. Background Technology

[0002] Human pose detection plays a crucial role in practical applications such as behavior recognition and fall detection. Classic vision-based human pose recognition methods involve inputting an image containing key human points into a neural network. The network outputs heatmaps of the detected joints, and connecting these heatmaps to represent the joints forms the key point information of the human body.

[0003] However, in practical applications, using visual detection of the human body raises privacy concerns. Summary of the Invention

[0004] In view of this, it is necessary to provide a posture detection method and its module, as well as a millimeter-wave radar, that can detect human posture while protecting the privacy of the monitored person.

[0005] In a first aspect, embodiments of this application provide a posture detection method, the posture detection method comprising:

[0006] Receive millimeter-wave point cloud data of the target object;

[0007] The millimeter-wave point cloud data is converted into a two-dimensional depth map;

[0008] The two-dimensional depth map is input into the target model to obtain the key point coordinates of the target object, wherein the target model is obtained through transfer learning from the source model, and the source model contains knowledge for identifying key points; and

[0009] The pose of the target object is obtained based on the coordinates of the key points.

[0010] Secondly, embodiments of this application provide an attitude detection module, the attitude detection module comprising:

[0011] Memory, used to store program instructions; and

[0012] A processor for executing the program instructions to implement the attitude detection method as described above.

[0013] Thirdly, embodiments of this application provide a millimeter-wave radar, which includes a main body and an attitude detection module as described above, wherein the attitude detection module is disposed on the main body.

[0014] The aforementioned attitude detection method and its module, along with millimeter-wave radar, trains a target model based on a source model through knowledge transfer. Since the source model contains knowledge about key point identification, the trained target model can also identify key points in the 2D depth map generated from millimeter-wave point cloud data, thus determining the target object's attitude. Mapping millimeter-wave point cloud data to a 2D depth map transforms point cloud data containing 3D spatial information into a 2D image format, enabling the target model to directly identify the 2D depth map based on learned knowledge. Using millimeter-wave radar as an attitude detection device, employing a deep neural network (i.e., the target model) for attitude detection, overcomes the problem of insufficient ambient light and significantly protects the privacy of the monitored individual. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0016] Figure 1 A flowchart of the attitude detection method provided in the embodiments of this application.

[0017] Figure 2 This is a sub-flowchart of the attitude detection method provided in the embodiments of this application.

[0018] Figure 3 This is a first sub-flowchart of the attitude detection method provided in the first embodiment of this application.

[0019] Figure 4 This is a second sub-flowchart of the attitude detection method provided in the first embodiment of this application.

[0020] Figure 5 This is a third sub-flowchart of the attitude detection method provided in the first embodiment of this application.

[0021] Figure 6 This is the fourth sub-flowchart of the attitude detection method provided in the first embodiment of this application.

[0022] Figure 7 This is a first sub-flowchart of the attitude detection method provided in the second embodiment of this application.

[0023] Figure 8 This is a second sub-flowchart of the attitude detection method provided in the second embodiment of this application.

[0024] Figure 9This is a schematic diagram illustrating an application scenario of the attitude detection method provided in the embodiments of this application.

[0025] Figure 10 for Figure 3 A schematic diagram of the data acquisition device shown.

[0026] Figure 11 for Figure 3 The image shows a comparison of data collected by the data acquisition device.

[0027] Figure 12 for Figure 3 The training flowchart for the target model is shown.

[0028] Figure 13 for Figure 3 The diagram shown illustrates the identification of key points of a target object using a target model.

[0029] Figure 14 for Figure 3 The bar chart shown is a graph illustrating the detection error of key points of the target object identified by the target model.

[0030] Figure 15 for Figure 4 The diagram shows the structure of the source model.

[0031] Figure 16 for Figure 6 The diagram shown illustrates how the target model learns knowledge from the source model.

[0032] Figure 17 for Figure 7 The image shows a comparison of the detection errors of the target model in identifying key points of the target object.

[0033] Figure 18 This is a schematic diagram of the internal structure of the attitude detection module provided in the embodiments of this application.

[0034] Figure 19 This is a schematic diagram of the internal structure of a millimeter-wave radar provided in an embodiment of this application.

[0035] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0037] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar planned objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data are interchangeable where appropriate; in other words, the described embodiments are implemented according to a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, may also include other content; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0038] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0039] Please refer to the following: Figure 1 and Figure 9 , Figure 1 This is a flowchart of the attitude detection method provided in the embodiments of this application. Figure 9 This is a schematic diagram illustrating an application scenario of the attitude detection method provided in an embodiment of this application. Figure 9 Taking the illustrated application scenario as an example, the millimeter-wave radar 20 is installed in an indoor environment to detect target objects 90 within that environment. The target objects 90 include, but are not limited to, human bodies and animals. In some feasible embodiments, the millimeter-wave radar 20 can also be installed in an outdoor environment to detect objects in that environment. An attitude detection method is applied to the millimeter-wave radar 20 to process the point cloud data formed by the echo signals received by the millimeter-wave radar 20, thereby obtaining the attitude of the target object 90.

[0040] The attitude detection method specifically includes the following steps.

[0041] Step S102: Receive millimeter-wave point cloud data of the target object.

[0042] The millimeter-wave radar 20 transmits a detection signal and receives the echo signal formed by the reflection of the detection signal from the target object 90. The millimeter-wave radar 20 generates millimeter-wave point cloud data based on the detection signal and the echo signal. In this embodiment, the millimeter-wave radar 20 includes a data processing module for generating millimeter-wave point cloud data and a detection module (not shown) for attitude detection. After generating the millimeter-wave point cloud data, the data processing module sends the millimeter-wave point cloud data to the detection module, and the detection module receives the millimeter-wave point cloud data of the target object 90.

[0043] Specifically, the data processing module transmits a detection signal within a conical range via a transmitting antenna and receives the echo signal reflected from the target object 90 via a receiving antenna. The detection signal is a frequency-modulated continuous wave (FMCW) millimeter wave with a frequency range of 30-300 GHz. The data processing module calculates the frequency difference between the detection signal and the echo signal, calculates the distance and angle between the target object 90 and the transmitting antenna based on the frequency difference, and then generates millimeter-wave point cloud data of the target object 90 based on the distance and angle.

[0044] Step S104: Convert the millimeter-wave point cloud data into a two-dimensional depth map.

[0045] The detection module converts the three-dimensional millimeter-wave point cloud data into two-dimensional data, namely a two-dimensional depth map.

[0046] The specific process of converting millimeter-wave point cloud data into a two-dimensional depth map will be described in detail below.

[0047] Step S106: Input the two-dimensional depth map into the target model to obtain the key point coordinates of the target object.

[0048] The detection module inputs a 2D depth map into the target model, which then outputs the coordinates of 90° key points of the target object. In this embodiment, the target model is obtained through transfer learning from a source model, which contains knowledge about key point identification. This key point identification knowledge from the source model can be used to annotate the key point locations in the input image. The image annotation results can be used to supervise the training of the target model, enabling it to learn the knowledge of key point identification and thus recognize the key point information, i.e., the key point coordinates, of the target object contained in the millimeter-wave radar point cloud data.

[0049] The specific process of training the target model based on the source model will be described in detail below.

[0050] Step S108: Obtain the pose of the target object based on the coordinates of the key points.

[0051] The detection module analyzes the coordinates of key points to obtain the 90° pose of the target object.

[0052] In the above embodiments, a target model is trained based on a source model through knowledge transfer. Since the source model contains knowledge about identifying key points, the trained target model can also identify key points in the two-dimensional depth map converted from millimeter-wave point cloud data, thereby obtaining the pose of the target object. Mapping millimeter-wave point cloud data to a two-dimensional depth map transforms point cloud data containing three-dimensional spatial information into a two-dimensional image format, enabling the target model to directly identify the two-dimensional depth map based on the learned knowledge. Using millimeter-wave radar as a pose detection device and employing a deep neural network, i.e., the target model, for pose detection can overcome the problem of insufficient ambient light and can also largely protect the privacy information of the monitored person.

[0053] Please refer to the following: Figure 2 This is a sub-flowchart of the attitude detection method provided in the embodiments of this application. Step S104 specifically includes the following steps.

[0054] Step S202: Obtain the three-dimensional coordinates of all points in the millimeter-wave point cloud data.

[0055] The detection module acquires the three-dimensional coordinates of all points in the millimeter-wave point cloud data. It can be understood that millimeter-wave point cloud data is three-dimensional point cloud data. Specifically, the detection module calculates the corresponding three-dimensional coordinates based on the distance and angle of each point in the millimeter-wave point cloud data. These three-dimensional coordinates include width, height, and depth values. In this implementation, each point in the millimeter-wave point cloud data is stored as a one-dimensional array, and the three-dimensional coordinates of a point can be represented as (x, y, z). Correspondingly, x represents the width value, y represents the height value, and z represents the depth value.

[0056] Step S204: Construct a two-dimensional coordinate system.

[0057] The detection module constructs a two-dimensional coordinate system. This two-dimensional coordinate system includes an abscissa and a ordinate.

[0058] Step S206: Map the width and height values ​​to the horizontal and vertical coordinates respectively, and convert the corresponding depth values ​​into pixel values ​​to obtain a two-dimensional depth map.

[0059] The detection module maps the width values ​​of all points to the x-coordinate of a two-dimensional coordinate system and the height values ​​to the y-coordinate, and converts the depth values ​​of the corresponding points into pixel values, thus forming a two-dimensional depth map corresponding to the millimeter-wave point cloud data. Specifically, the points in the millimeter-wave point cloud data correspond to the pixels in the two-dimensional depth map.

[0060] In this embodiment, the detection module maps millimeter-wave point cloud data to a two-dimensional depth map according to Formula 1. Specifically, Formula 1 is:

[0061]

[0062] Where u represents the x-coordinate of the pixel in the 2D depth map, v represents the y-coordinate of the pixel in the 2D depth map, and z represents the pixel value of the pixel in the 2D depth map. x f represents the focal length multiplied by the horizontal scaling factor between the three-dimensional and two-dimensional coordinate systems. y This represents the focal length multiplied by the vertical scaling factor between the three-dimensional and two-dimensional coordinate systems; c x This represents the lateral distance, c, from the origin of the 3D coordinate system to the top-left corner of the 2D depth map after mapping from the 3D coordinate system to the 2D coordinate system. y This represents the vertical distance from the origin of the 3D coordinate system to the top-left corner of the 2D depth map after mapping from the 3D coordinate system to the 2D coordinate system. K represents the intrinsic parameter matrix of the camera used to acquire the data, and P represents a point in the millimeter-wave point cloud data.

[0063] Furthermore, to meet visualization requirements, the detection module maps the pixel value of each pixel in the 2D depth map to a value between 0 and 255. In this embodiment, the detection module performs the mapping according to Formula 2. Specifically, Formula 2 is:

[0064]

[0065] Among them, z max =max(z|z∈Z). z max z represents the maximum pixel value among all pixels in the 2D depth map, and z represents the currently mapped pixel value. i This represents the mapped pixel value.

[0066] In the above embodiments, by mapping, millimeter-wave point cloud data containing any number of points can be unified into the same form to facilitate subsequent feature extraction and key point recognition.

[0067] Please refer to the following: Figure 3 This is a first sub-flowchart of the pose detection method provided in the first embodiment of this application. Components used for training the target model include, but are not limited to, a data acquisition module 30 and a data training module (not shown). Before executing step S106, the pose detection method further includes the following steps.

[0068] Step S302: Synchronously collect data from the same object from the same angle to obtain initial point cloud data and camera images that correspond one-to-one.

[0069] The data acquisition module 30 is used to acquire data, including initial point cloud data and camera images. In this embodiment, the data acquisition module 30 includes a color camera 31 and a millimeter-wave radar 32 (e.g., [missing information]). Figure 10(As shown). The color camera 31 is used to acquire video images, and the millimeter-wave radar 32 is used to acquire initial point cloud data. It is understood that the video images are color photographs, and the initial point cloud data is also millimeter-wave point cloud data.

[0070] Specifically, the color camera 31 and the acquisition millimeter-wave radar 32 are set together so that the acquisition millimeter-wave radar 32 and the color camera 31 can acquire objects from the same angle. Figure 10 In the image, the left side shows a millimeter-wave radar 32, and the right side shows a color camera 31. The data acquisition module 30 synchronously drives the millimeter-wave radar 32 and the color camera 31 to acquire data about the same object, thereby obtaining corresponding initial point cloud data and camera images. Each frame of initial point cloud data corresponds to at least one frame of camera image.

[0071] Step S304: Standardize the initial point cloud data to obtain an initial depth map.

[0072] The data acquisition module 30 or the data training module uses the camera imaging principle to map each frame of initial point cloud data onto a two-dimensional image, obtaining an initial depth map corresponding to each frame of initial point cloud data. This converts the initial point cloud data into a uniformly formatted depth map, facilitating subsequent feature extraction and model training. The process of standardizing the initial point cloud data into an initial depth map is essentially the same as the process of converting millimeter-wave point cloud data into a two-dimensional depth map, and will not be elaborated further here.

[0073] The captured video images are as follows Figure 11 As shown on the right, the initial depth map obtained after standardization is as follows: Figure 11 As shown on the left, the initial depth map displays the point cloud outlines of the head, arms, abdomen, legs, and other parts of the human body as it moves.

[0074] Step S306: Input the camera image into the source model to obtain the first output result.

[0075] The data training module inputs camera images into the source model to obtain the first output. Understandably, because the source model contains knowledge of keypoint recognition, it can identify the key points of the object's pose in the camera image when the image is input.

[0076] The specific process of inputting camera images into the source model to obtain the first output result will be described in detail below.

[0077] Step S308: Input the initial depth map into the initial model to obtain the second output result.

[0078] The data training module inputs the initial depth map into the initial model to obtain the second output result.

[0079] The specific process of inputting the initial depth map into the initial model to obtain the second output result will be described in detail below.

[0080] Step S310: Calculate the first loss function based on the first output result and the second output result.

[0081] The data training module calculates the first loss function based on the first output and the second output.

[0082] The specific process of calculating the first loss function based on the first and second output results will be described in detail below.

[0083] Step S312: Iteratively train the initial model according to the first loss function to obtain the target model.

[0084] The data training module iteratively trains the initial model using the first loss function to obtain the target model. In this embodiment, the process of the data training module iteratively training the initial model into the target model is basically the same as the process of iterative training of existing models, and will not be described in detail here.

[0085] The model training process includes data acquisition, data preprocessing, and knowledge transfer training (such as...). Figure 12 (As shown). Among them, data acquisition involves acquiring initial point cloud data and camera images; data preprocessing involves processing the initial point cloud data into an initial depth map; knowledge transfer training involves supervising the initial model based on the annotation results of the camera images obtained from the source model, so that the trained target model has the ability to recognize key point information in the depth map.

[0086] The target model trained accordingly demonstrates the following performance in detecting the 90-degree pose of the target object: Figure 13 As shown. Figure 13 The left side shows the pose detection result based on camera images, while the right side shows the pose detection result based on a 2D depth map formed from millimeter-wave point cloud data. The comparison shows that the target model can roughly detect the coordinates of 90° key points of the target object. The detection error of the target model for 90° pose detection of the target object is as follows: Figure 14 As shown in the figure, the target model performs well in 1 (head), 2 (chest), 3 (left shoulder), 4 (right shoulder), 9 (left hip), and 10 (right hip), with errors between 12 and 16 pixels; it performs slightly worse in 7 (left elbow), 8 (right elbow), 11 (left knee), and 12 (right knee), with errors between 22 and 24 pixels; and it performs the worst in 5 (left wrist), 6 (right wrist), 13 (left ankle), and 14 (right ankle), with errors between 32 and 33 pixels.

[0087] In the above embodiments, initial point cloud data and camera images of the same object at the same angle are acquired simultaneously, so that each frame of initial point cloud data has corresponding visual data, i.e., a camera image. The initial model is trained using transfer learning based on the source model, transferring the knowledge of recognizing human key points from the source model to the initial model, thereby enabling the trained target model to acquire the ability to detect key points.

[0088] Please refer to the following: Figure 4 This is the second sub-flowchart of the attitude detection method provided in the first embodiment of this application. Step S306 specifically includes the following steps.

[0089] Step S402: Input the camera image into the feature extraction network to obtain the camera feature map.

[0090] The data training module inputs camera images into the feature extraction network to obtain corresponding camera feature maps. In this embodiment, the feature extraction network is VGG-19.

[0091] Step S404: Input the camera feature map into the source model to obtain the first output result.

[0092] The data training module inputs the camera feature map into the source model to obtain the corresponding first output result. In this embodiment, the first output result includes a first joint confidence map and a first partial affinity field.

[0093] The structure of the source model is as follows Figure 15 As shown, the source model includes T P +T C There are several stages, including T. P One stage is used to predict the partial affinity field of the camera feature map, T C Each stage is used to predict the joint confidence of the camera feature map to obtain the joint confidence map.

[0094] Please refer to the following: Figure 5 This is the third sub-flowchart of the attitude detection method provided in the first embodiment of this application. Step S308 specifically includes the following steps.

[0095] Step S502: Input the initial depth map into the feature extraction network to obtain a depth feature map.

[0096] The data training module inputs the initial depth map into the feature extraction network to obtain the corresponding depth feature map. In this embodiment, the feature extraction network is VGG-19.

[0097] Step S504: Input the depth feature map into the initial model to obtain the second output result.

[0098] The data training module inputs the deep feature map into the initial model to obtain the corresponding second output result. In this embodiment, the second output result includes the second joint confidence map and the second affinity field. The structure of the initial model is basically the same as that of the source model, and will not be described in detail here.

[0099] Please refer to the following: Figure 6 This is the fourth sub-flowchart of the attitude detection method provided in the first embodiment of this application. Step S310 specifically includes the following steps.

[0100] Step S602: Calculate the first sub-function based on the first joint confidence graph and the second joint confidence graph.

[0101] Step S604: Calculate the second sub-function based on the first part of the affinity field and the second part of the affinity field.

[0102] Step S606: The sum of the first sub-function and the second sub-function is used as the first loss function.

[0103] like Figure 16 As shown, the data training module inputs the camera feature map into the source model, and the first output result includes the first joint confidence map. And the first part of the affinity scene The deep feature map is input into the initial model, and the second output includes the second joint confidence map. Part Two Affinity Field

[0104]

[0105] In this embodiment, T P For each stage in the T stage, a confidence loss function is calculated. C For each stage in the series, an affinity field loss function is calculated. Specifically, the confidence loss...

[0106] loss function Indicates at stage t i The affinity loss function of the time-partial affinity field. and Let P represent the partial affinity field of pixel P in the c-th stage output of the source model and the initial model, respectively. and Let represent the joint confidence maps of pixel P output from the j-th stage in the source model and the initial model, respectively. W is the binary mask of pixel P, with an unlabeled pixel P having a binary mask of 0.

[0107] T PThe sum of the confidence loss functions of all stages in a given stage is the first sub-function, T. C The sum of the affinity field loss functions of all stages in a given stage constitutes the second sub-function. Therefore, the first loss function is expressed as:

[0108]

[0109] Among them, f op This represents the value of the first loss function.

[0110] In the above embodiments, an existing vision-based keypoint detection network, i.e., the source model, is used to detect keypoints in the camera image, and the detection results are used as labels for the initial model training. This allows the knowledge of the existing model, i.e. the source model, to be transferred to the initial model to obtain the target model, thereby solving the problem of the difficulty of manually annotating point clouds on two-dimensional depth maps and significantly improving the annotation efficiency of point cloud information.

[0111] Please refer to the following: Figure 7 This is a first sub-flowchart of the attitude detection method provided in the second embodiment of this application. The difference between the attitude detection method provided in the second embodiment and the attitude detection method provided in the first embodiment is that, after executing step S310, the attitude detection method provided in the second embodiment further includes the following steps.

[0112] Step S702: Input the camera image into the domain adaptation network to obtain the third output result, and calculate the second loss function based on the third output result.

[0113] The data training module also inputs the camera images into the domain adaptation network to obtain a third output. In this embodiment, the domain adaptation network includes a Flatten layer, two fully connected layers, and a softmax layer. The third output represents the classification result of the scene in which the camera image is located.

[0114] The data training module calculates the second loss function based on the third output result.

[0115] The specific process of adapting the camera image input domain to the network to obtain the third output result, and calculating the second loss function based on the third output result, will be described in detail below.

[0116] Step S704: Calculate the target loss function based on the first loss function and the second loss function.

[0117] The data training module calculates the target loss function based on the first loss function and the second loss function. In this embodiment, the data training module calculates the target loss function according to Formula 3. Specifically, Formula 3 is: f = f op -λf aWhere f represents the value of the target loss function, f op f represents the value of the first loss function. a The value of the second loss function is represented by λ, which represents the influence parameter of the domain adaptation network on the target model. The larger λ is, the stronger the domain adaptation network's ability; when λ is 0, it means the domain adaptation network has no effect.

[0118] Step S706: Iteratively train the initial model according to the target loss function to obtain the target model.

[0119] The data training module iteratively trains the initial model according to the target loss function to obtain the target model. In this embodiment, the process of the data training module iteratively training the initial model into the target model is basically the same as the process of iterative training of existing models, and will not be described again here.

[0120] A comparison of the detection errors of the target model obtained with and without a domain adaptation network for 90° pose detection of the target object is shown below. Figure 17 As shown in the figure, the target model trained with the addition of a domain adaptation network exhibits improved prediction accuracy for all joints of the human body. This is because the domain adaptation network can suppress the model's feature extraction part from extracting features belonging to the environment and the individual, allowing the model to eliminate individual differences and environmental interference, extracting more generalized features, thereby increasing the prediction accuracy and stability of the target model.

[0121] Please refer to the following: Figure 8 This is a second sub-flowchart of the attitude detection method provided in the second embodiment of this application. Step S702 specifically includes the following steps.

[0122] Step S802: Input the camera image into the feature extraction network to obtain the camera feature map.

[0123] The data training module inputs camera images into the feature extraction network to obtain corresponding camera feature maps. In this embodiment, the feature extraction network is VGG-19, and the camera feature map includes several feature vectors.

[0124] Step S804: Input the camera feature map into the domain adaptation network to obtain the third output result.

[0125] The data training module inputs the camera feature map into the domain adaptation network to obtain a third output result. This third output result includes several output vectors.

[0126] Step S806: Calculate the second loss function based on the feature vector and the output vector.

[0127] The data training module calculates a second loss function based on the feature vector and the output vector. In this embodiment, the data training module calculates the second loss function according to Formula 4. Specifically, Formula 4 is:

[0128]

[0129] Among them, f a Let J represent the value of the second loss function, J represent the length of the output vector, and D represent the output vector. * Let represent the feature vector, and i represent the i-th element in the vector.

[0130] In the above embodiments, different environments have different impacts on data acquisition. For example, the initial point cloud data and the model trained from camera images acquired in a single environment may have biased judgments on key points when performing pose detection in other environments, resulting in poor performance. Domain adaptation networks can suppress the model's extraction of features that are part of environmental interference. Therefore, adding a domain adaptation network during the training process of the source model on the initial model can make the features extracted by the target model more general, effectively improving the accuracy and stability of the target model.

[0131] Please refer to the following: Figure 18 This is a schematic diagram of the internal structure of the attitude detection module provided in this application embodiment. The attitude detection module 10 includes a memory 11 and a processor 12. The memory 11 is used to store program instructions, and the processor 12 is used to execute the program instructions to implement the above-described attitude detection method.

[0132] In some embodiments, the processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program instructions stored in the memory 11.

[0133] The memory 11 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of a computer device, such as a hard disk. In other embodiments, the memory 11 may be an external storage device of a computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the memory 11 may include both internal and external storage units of the computer device. The memory 11 can be used not only to store application software and various types of data installed on the computer device, such as code implementing attitude detection methods, but also to temporarily store data that has been output or will be output.

[0134] Please refer to the following: Figure 19 This is a schematic diagram of the internal structure of the millimeter-wave radar provided in this embodiment. The millimeter-wave radar 20 includes a main body 21 and an attitude detection module 10, with the attitude detection module 10 disposed on the main body 21. The millimeter-wave radar 20 can be applied to, but is not limited to, motion capture, behavior recognition, and smart elderly care. The specific structure of the attitude detection module 10 is as described in the above embodiment. In this embodiment, the size of the millimeter-wave radar 20 can be as small as the size of a fingernail for easy installation.

[0135] Since the millimeter-wave radar 20 adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.

[0136] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

[0137] The above-listed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A posture detection method characterized by comprising: The posture detection method comprises: receiving millimeter wave point cloud data of a target object; converting the millimeter wave point cloud data into a two-dimensional depth map; inputting the two-dimensional depth map into a target model to obtain key point coordinates of the target object, wherein the target model is obtained by transferring learning from a source model, the source model transfers knowledge to an initial model to obtain the target model; the source model contains knowledge of identifying key points, can label the positions of key points of an input image, and outputs an output result containing key point knowledge of the input image after receiving the input image; the structure of the initial model is basically consistent with that of the source model, a depth map corresponding to the input image is input into the initial model, the initial model outputs a corresponding result, and the key point identification ability of the source model is learned step by step through a first loss function; the first loss function is constructed based on the difference between the output result and the corresponding result; and obtaining the posture of the target object according to the key point coordinates; wherein the posture detection method further comprises: inputting an input image into a domain adaptation network to obtain a third output result, and calculating a second loss function according to the third output result; calculating a target loss function according to the first loss function and the second loss function; and iteratively training the initial model according to the target loss function to obtain the target model; wherein the input image is a camera picture, the camera picture is input into a feature extraction network to obtain a camera feature map, wherein the camera feature map comprises a plurality of feature vectors; the camera feature map is input into the domain adaptation network to obtain the third output result, wherein the third output result comprises a plurality of output vectors; and the second loss function is calculated according to the feature vectors and the output vectors.

2. The attitude detection method according to claim 1, characterized by, Before inputting the two-dimensional depth map into the target model to obtain the key point coordinates of the target object, the posture detection method further comprises: synchronously collecting the same object from the same angle to obtain one-to-one initial point cloud data and the camera picture; standardizing the initial point cloud data to obtain an initial depth map; inputting the camera picture into the source model to obtain a first output result; inputting the initial depth map into an initial model to obtain a second output result; calculating a first loss function according to the first output result and the second output result; and iteratively training the initial model according to the first loss function to obtain the target model.

3. The attitude detection method according to claim 2, characterized by, Inputting the camera picture into the source model to obtain a first output result comprises: inputting the camera picture into a feature extraction network to obtain a camera feature map; and inputting the camera feature map into the source model to obtain the first output result, wherein the first output result comprises a first key point confidence map and a first part affinity field.

4. The attitude detection method according to claim 3, characterized by, Inputting the initial depth map into an initial model to obtain a second output result comprises: inputting the initial depth map into the feature extraction network to obtain a depth feature map; and inputting the depth feature map into the initial model to obtain a second output result, wherein the second output result comprises a second joint confidence map and a second part affinity field.

5. The attitude detection method according to claim 4, characterized by, calculating a first loss function according to the first output result and the second output result comprises: calculating a first sub-function according to the first joint confidence map and the second joint confidence map; calculating a second sub-function according to the first part affinity field and the second part affinity field; and taking a sum of the first sub-function and the second sub-function as the first loss function.

6. The attitude detection method according to claim 1, wherein converting the millimeter wave point cloud data into a two-dimensional depth map comprises: obtaining three-dimensional coordinates of all points in the millimeter wave point cloud data, wherein the three-dimensional coordinates comprise a width value, a height value and a depth value; constructing a two-dimensional coordinate system, wherein the two-dimensional coordinate system comprises a horizontal coordinate and a vertical coordinate; and mapping the width value and the height value into the horizontal coordinate and the vertical coordinate respectively, and converting the corresponding depth value into a pixel value to obtain the two-dimensional depth map.

7. A posture detection module, comprising: the posture detection module comprises: a memory for storing program instructions; and a processor for executing the program instructions to implement the posture detection method according to any one of claims 1 to 5.

8. A millimeter wave radar, characterized by, the millimeter wave radar comprises a main body and the posture detection module according to claim 7, and the posture detection module is arranged on the main body.

Citation Information

Patent Citations

  • Millimeter wave human body intelligent posture detection method, detection device and detection system

    CN115423749A