Method and device for generating three-dimensional key points, robot and storage medium
Patent Information
- Application Number
- CN202211415093.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-11
AI Technical Summary
[0003]现有技术中,实现三维关键点检测方案有很多,通常基于深度学习实现,例如通过图像直接检测三维关键点坐标,或者先对图像进行二维关键点的检测,再利用图像和二维关键点一起计算三维关键点坐标,但现有的三维关键点检测方法在模型精度上存在一定缺陷,导致最终得到的三维关键点结果准确度不高
[0068]本申请提供的目标生成模型可以通过关键点的二维坐标生成三维坐标,从而实现基于二维关键点的三维关键点检测。目标生成模型具体采用有向图卷积层和自注意力层并用的架构,有向图卷积层负责局部特征的捕捉,自注意力层负责全局特征的捕捉,再将局部特征和全局特征有机结合,在保证实时性的情况下提高模型精度,提高三维关键点结果的准确性。
Smart Images

Figure CN115861643B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method, apparatus, robot, and storage medium for generating three-dimensional key points. Background Technology
[0002] The field of keypoint detection includes facial keypoints, human body keypoints, and keypoint detection for specific object categories. Among these, 3D keypoint detection of the human body is a relatively popular, challenging, and widely applied research area.
[0003] In the existing technology, there are many schemes for realizing 3D key point detection, which are usually based on deep learning. For example, the coordinates of 3D key points can be detected directly through the image, or the 2D key points can be detected in the image first, and then the coordinates of 3D key points can be calculated by using the image and the 2D key points together. However, the existing 3D key point detection methods have certain defects in model accuracy, resulting in low accuracy of the final 3D key point results. Summary of the Invention
[0004] This application provides a method, apparatus, robot, and storage medium for generating three-dimensional key points, which can improve the accuracy of generating three-dimensional key points based on two-dimensional key points.
[0005] The first aspect of this application provides a method for generating three-dimensional key points, including:
[0006] Acquire the two-dimensional key point data to be processed;
[0007] The target generation model is invoked, which includes a fully connected input module, a feature extraction module, and a fully connected output module.
[0008] The two-dimensional key point data is input into the target generation model, and the two-dimensional key point data is processed by feature mapping through the fully connected input module to obtain the first feature;
[0009] The first feature is input into the feature extraction module to generate the second feature. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0010] The second feature is processed by keypoint mapping through the fully connected output module to obtain the three-dimensional keypoints output by the target generation model.
[0011] Optionally, the feature extraction module includes at least one sub-module group, the sub-module group including two main sub-modules and one feature overlay sub-module, and the step of inputting the first feature into the feature extraction module to generate the second feature includes:
[0012] The first feature is input into the first main submodule of the submodule group to obtain the intermediate feature;
[0013] The first feature and the intermediate feature are superimposed by the feature superposition submodule, and the superposition result is input to the second main submodule to obtain the second feature output by the submodule group.
[0014] Optionally, the backbone submodule includes a directed graph convolutional layer and a self-attention layer, and the step of inputting the first feature into the first backbone submodule to obtain intermediate features includes:
[0015] The directed graph convolutional layer captures the local features of the first feature, the self-attention layer captures the global features of the first feature, and the intermediate feature is obtained by fusing the local features and the global features.
[0016] Optionally, the directed graph convolutional layer includes a forward graph convolutional layer and a backward graph convolutional layer, and the local features captured by the directed graph convolutional layer for the first feature include:
[0017] The first node feature is extracted through the forward graph convolutional layer, and the first node feature is used to characterize the forward connection relationship in the first feature.
[0018] The second node feature is extracted through the inverse graph convolutional layer, and the second node feature is used to characterize the inverse connection relationship in the first feature.
[0019] By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
[0020] Optionally, the directed graph convolutional layer further includes an autograph convolutional layer. Before fusing the first node features and the second node features to capture local features of the first feature, the method further includes:
[0021] The third node feature is extracted through the self-graph convolutional layer, and the third node feature is used to characterize the connection relationship between itself and itself in the first feature.
[0022] The process of fusing the first node features and the second node features to capture local features of the first feature includes:
[0023] By fusing the features of the first node, the second node, and the third node, local features of the first feature are captured.
[0024] Optionally, the method further includes training the target generation model, wherein training the target generation model includes:
[0025] Obtain a sample dataset, which contains multiple sets of matching two-dimensional and three-dimensional sample data;
[0026] Obtain an initial generative model, which includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0027] Select a set of matching two-dimensional sample data and three-dimensional sample data from the sample dataset, input the two-dimensional sample data into the initial generation model, and obtain the first three-dimensional key points output by the initial generation model;
[0028] A first loss value is calculated based on the first 3D key points and the 3D sample data, and the initial generation model is iteratively trained based on the first loss value to obtain a target generation model. The target generation model is used to generate 3D key points based on 2D key points.
[0029] Optionally, before iteratively training the initial generative model based on the first loss value, the method further includes:
[0030] Obtain the second three-dimensional key points output from the middle segment of the network of the initial generated model;
[0031] Calculate the second loss value based on the second three-dimensional key points and the three-dimensional sample data;
[0032] The iterative training of the initial generative model based on the first loss value includes:
[0033] The initial generative model is iteratively trained based on the first loss value and the second loss value.
[0034] A second aspect of this application provides a three-dimensional key point generation apparatus, comprising:
[0035] The data acquisition unit is used to acquire the two-dimensional key point data to be processed.
[0036] The model invocation unit is used to invoke the target generation model, which includes a fully connected input module, a feature extraction module, and a fully connected output module.
[0037] A data input unit is used to input the two-dimensional key point data into the target generation model;
[0038] The model processing unit is used to perform feature mapping processing on the two-dimensional keypoint data through the fully connected input module to obtain a first feature; input the first feature into the feature extraction module to generate a second feature, the feature extraction module including a directed graph convolutional layer and a self-attention layer; and perform keypoint mapping processing on the second feature through the fully connected output module to obtain the three-dimensional keypoints output by the target generation model.
[0039] Optionally, the feature extraction module includes at least one sub-module group, which includes two main sub-modules and one feature overlay sub-module. The model processing unit is specifically used for:
[0040] The first feature is input into the first main submodule of the submodule group to obtain the intermediate feature;
[0041] The first feature and the intermediate feature are superimposed by the feature superposition submodule, and the superposition result is input to the second main submodule to obtain the second feature output by the submodule group.
[0042] Optionally, the backbone submodule includes directed graph convolutional layers and self-attention layers, and the model processing unit is further used for:
[0043] The directed graph convolutional layer captures the local features of the first feature, the self-attention layer captures the global features of the first feature, and the intermediate feature is obtained by fusing the local features and the global features.
[0044] Optionally, the directed graph convolutional layer includes a forward graph convolutional layer and a backward graph convolutional layer, and the model processing unit is further configured to:
[0045] The first node feature is extracted through the forward graph convolutional layer, and the first node feature is used to characterize the forward connection relationship in the first feature.
[0046] The second node feature is extracted through the inverse graph convolutional layer, and the second node feature is used to characterize the inverse connection relationship in the first feature.
[0047] By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
[0048] Optionally, the directed graph convolutional layer further includes an autograph convolutional layer, and the model processing unit is further used for:
[0049] The third node feature is extracted through the self-graph convolutional layer, and the third node feature is used to characterize the connection relationship between itself and itself in the first feature.
[0050] By fusing the features of the first node, the second node, and the third node, local features of the first feature are captured.
[0051] Optionally, the generation device further includes: a model training unit;
[0052] The model training unit is specifically used for:
[0053] Obtain a sample dataset, which contains multiple sets of matching two-dimensional and three-dimensional sample data;
[0054] Obtain an initial generative model, which includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0055] A set of matching two-dimensional sample data and three-dimensional sample data are randomly selected from the sample dataset. The two-dimensional sample data is input into the initial generation model to obtain the first three-dimensional key points output by the initial generation model.
[0056] A first loss value is calculated based on the first 3D key points and the 3D sample data, and the initial generation model is iteratively trained based on the first loss value to obtain a target generation model. The target generation model is used to generate 3D key points based on 2D key points.
[0057] Optionally, the model training unit is further used for:
[0058] Obtain the second three-dimensional key points output from the middle segment of the network of the initial generated model;
[0059] Calculate the second loss value based on the second three-dimensional key points and the three-dimensional sample data;
[0060] The iterative training of the initial generative model based on the first loss value includes:
[0061] The initial generative model is iteratively trained based on the first loss value and the second loss value.
[0062] A third aspect of this application provides a robot, the robot comprising:
[0063] Processor, memory, input / output units, and bus;
[0064] The processor is connected to the memory, the input / output unit, and the bus;
[0065] The memory stores a program, which the processor calls to execute the first aspect and any optional method for generating three-dimensional key points in the first aspect.
[0066] The fourth aspect of this application provides a computer-readable storage medium storing a program that, when executed on a computer, performs the first aspect and any optional method for generating three-dimensional key points in the first aspect.
[0067] As can be seen from the above technical solutions, this application has the following advantages:
[0068] The target generation model provided in this application can generate three-dimensional coordinates from the two-dimensional coordinates of key points, thereby realizing three-dimensional key point detection based on two-dimensional key points. Specifically, the target generation model adopts an architecture that uses both directed graph convolutional layers and self-attention layers. The directed graph convolutional layers are responsible for capturing local features, while the self-attention layers are responsible for capturing global features. The local and global features are then organically combined to improve model accuracy and the accuracy of the three-dimensional key point results while ensuring real-time performance. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 A schematic diagram of the hardware structure of the robot provided in this application;
[0071] Figure 2 A schematic diagram of the mechanical structure of the robot provided in this application;
[0072] Figure 3 A schematic flowchart of an embodiment of the method for generating three-dimensional key points provided in this application;
[0073] Figure 4 A schematic flowchart of another embodiment of the method for generating three-dimensional key points provided in this application;
[0074] Figure 5 A schematic diagram of the network structure of the target generation model in the method for generating 3D key points provided in this application;
[0075] Figure 6 A schematic diagram of the main sub-modules in the target generation model provided in this application;
[0076] Figure 7 A schematic diagram of the structure of the directed graph convolutional layer in the target generation model provided in this application;
[0077] Figure 8 A schematic diagram of an embodiment of the three-dimensional key point generation device provided in this application;
[0078] Figure 9 This application provides a schematic diagram of the structure of one embodiment of the robot. Detailed Implementation
[0079] This application provides a method, apparatus, robot, and storage medium for generating three-dimensional key points, which can improve the accuracy of generating three-dimensional key points based on two-dimensional key points.
[0080] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. It should be noted that the method for generating three-dimensional key points provided in this application can be applied to terminals as well as servers. For example, a terminal can be a robot, smartphone, computer, tablet computer, smart TV, smartwatch, portable computer terminal, or a desktop computer, etc. For ease of description, the following description will focus on the terminal execution entity.
[0081] In the following description, the use of suffixes such as "module," "component," or "unit" to denote parts is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "component," or "unit" may be used interchangeably.
[0082] Please see Figure 1 , Figure 1 This is a schematic diagram of the hardware structure of a robot 100 according to one embodiment of the present invention. Figure 1 In the illustrated embodiment, robot 100 includes a mechanical unit 101, a communication unit 102, a sensing unit 103, an interface unit 104, a storage unit 105, a control module 110, and a power supply 111. The various components of robot 100 can be connected in any way, including wired or wireless connections. Those skilled in the art will understand that... Figure 1 The specific structure of the robot 100 shown does not constitute a limitation on the robot 100. The robot 100 may include more or fewer parts than shown. Some parts are not essential components of the robot 100 and may be omitted or combined as needed without changing the nature of the invention.
[0083] The following is combined Figure 1 A detailed introduction to each component of Robot 100:
[0084] Mechanical unit 101 is the hardware of robot 100. For example... Figure 1 As shown, the mechanical unit 101 may include a drive board 1011, a motor 1012, and a mechanical structure 1013, such as... Figure 2As shown, the mechanical structure 1013 may include a main body 1014, extendable legs 1015, and feet 1016. In other embodiments, the mechanical structure 1013 may also include an extendable robotic arm (not shown), a rotatable head structure 1017, a rocking tail structure 1018, a cargo-carrying structure 1019, a saddle structure 1020, a camera structure 1021, etc. It should be noted that the various component modules of the mechanical unit 101 can be one or multiple, depending on the specific situation. For example, there may be four legs 1015, and each leg 1015 may be equipped with three motors 1012, resulting in a total of twelve motors 1012.
[0085] The communication unit 102 can be used for receiving and sending signals, and can also communicate with networks and other devices. For example, it can receive instructions from a remote control or other robot 100 to move in a specific direction at a specific speed according to a specific gait, and then transmit these instructions to the control module 110 for processing. The communication unit 102 includes modules such as WiFi, 4G, 5G, Bluetooth, and infrared modules.
[0086] The sensing unit 103 is used to acquire information data about the environment surrounding the robot 100 and to monitor parameter data of various components inside the robot 100, and then sends this data to the control module 110. The sensing unit 103 includes various sensors, such as sensors for acquiring information about the surrounding environment: lidar (for remote object detection, distance determination, and / or velocity determination), millimeter-wave radar (for short-range object detection, distance determination, and / or velocity determination), cameras, infrared cameras, and Global Navigation Satellite System (GNSS). Sensors for monitoring various components inside the robot 100 include: an inertial measurement unit (IMU) (for measuring velocity, acceleration, and angular velocity values), foot sensors (for monitoring the position of the foot's contact point, foot posture, magnitude and direction of the contact force), and temperature sensors (for detecting component temperature). Other sensors that can be configured on the robot 100, such as load sensors, touch sensors, motor angle sensors, and torque sensors, are not detailed here.
[0087] The interface unit 104 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within the robot 100, or it can be used to output to external devices (e.g., data, power, etc.). The interface unit 104 may include a power port, a data port (such as a USB port), a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, etc.
[0088] Storage unit 105 is used to store software programs and various data. Storage unit 105 may mainly include a program storage area and a data storage area. The program storage area may store operating system programs, motion control programs, application programs (such as text editors), etc.; the data storage area may store data generated by the robot 100 during use (such as various sensor data acquired by the sensing unit 103, log file data, etc.). Furthermore, storage unit 105 may include high-speed random access memory, and may also include non-volatile memory, such as disk storage, flash memory, or other volatile solid-state memory.
[0089] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0090] Input unit 107 can be used to receive input numerical or character information. Specifically, input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect user touch operations (such as operations performed by the user using their palm, fingers, or suitable accessories on or near touch panel 1071) and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: touch detection device 1073 and touch controller 1074. Touch detection device 1073 detects the user's touch position and the signal generated by the touch operation, and transmits the signal to touch controller 1074; touch controller 1074 receives touch information from touch detection device 1073, converts it into touch point coordinates, and sends it to control module 110, and can also receive and execute commands from control module 110. In addition to touch panel 1071, input unit 107 may also include other input devices 1072. Specifically, other input devices 1072 may include, but are not limited to, one or more of the following: remote control handles, etc., without any specific limitation here.
[0091] Furthermore, the touch panel 1071 can cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the control module 110 to determine the type of touch event. Subsequently, the control module 110 provides corresponding visual output on the display panel 1061 according to the type of touch event. Although in Figure 1In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components that implement input and output functions respectively. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to implement input and output functions. The specific implementation is not limited here.
[0092] The control module 110 is the control center of the robot 100. It connects all the components of the robot 100 through various interfaces and lines. It controls the robot 100 as a whole by running or executing the software program stored in the storage unit 105 and calling the data stored in the storage unit 105.
[0093] Power supply 111 supplies power to various components. Power supply 111 may include a battery and a power control board. The power control board controls battery charging, discharging, and power consumption management. Figure 1 In the illustrated embodiment, power supply 111 is electrically connected to control module 110. In other embodiments, power supply 111 may also be electrically connected to sensing unit 103 (such as camera, radar, speaker, etc.) and motor 1012. It should be noted that each component may be connected to a different power supply 111, or may be powered by the same power supply 111.
[0094] Based on the above embodiments, specifically, in some embodiments, a terminal device can be used to communicate with the robot 100. When the terminal device communicates with the robot 100, it can send instruction information to the robot 100. The robot 100 can receive the instruction information through the communication unit 102 and, upon receiving the instruction information, can transmit it to the control module 110, so that the control module 110 can process the instruction information to obtain the target speed value. The terminal device includes, but is not limited to, mobile phones, tablets, servers, personal computers, wearable smart devices, and other electrical appliances with image capture capabilities.
[0095] The instruction information can be determined based on preset conditions. In one embodiment, the robot 100 may include a sensing unit 103, which can generate instruction information based on the current environment of the robot 100. The control module 110 can determine whether the current speed value of the robot 100 meets the corresponding preset conditions based on the instruction information. If it does, the robot 100 will maintain its current speed value and current gait; if it does not, the control module 110 will determine a target speed value and a corresponding target gait based on the corresponding preset conditions, thereby controlling the robot 100 to move at the target speed value and the corresponding target gait. Environmental sensors may include temperature sensors, air pressure sensors, vision sensors, and sound sensors. Instruction information may include temperature information, air pressure information, image information, and sound information. The communication method between the environmental sensors and the control module 110 can be wired or wireless. Wireless communication methods include, but are not limited to: wireless networks, mobile communication networks (3G, 4G, 5G, etc.), Bluetooth, and infrared.
[0096] The method for generating 3D key points provided in this application is described below. Please refer to [link / reference]. Figure 3 , Figure 3 An embodiment of the method for generating three-dimensional key points provided in this application includes:
[0097] Step 301: Obtain the two-dimensional key point data to be processed.
[0098] The field of keypoint detection includes facial keypoints, human body keypoints, and keypoint detection for specific object categories (such as hand bones), among which human body keypoint detection is currently the most widely used area. In this embodiment, a method for generating three-dimensional keypoints based on two-dimensional keypoints is designed to achieve three-dimensional keypoint detection.
[0099] The terminal first acquires two-dimensional keypoint data, specifically the coordinates of the two-dimensional keypoints. It should be noted that the method of acquiring these keypoints is not limited here; for example, the terminal can obtain the two-dimensional keypoint data by performing two-dimensional keypoint detection on the image to be detected, or it can directly acquire pre-determined two-dimensional keypoint data. In scenarios where the terminal performs two-dimensional keypoint detection to acquire keypoint data, the terminal can specifically achieve the detection of two-dimensional keypoints through methods such as deep learning or manual labeling.
[0100] Step 302: Call the target generation model, which includes a fully connected input module, a feature extraction module, and a fully connected output module.
[0101] In image processing, keypoints are essentially features, abstract descriptions of a fixed region or spatial physical relationship, describing combinations or contextual relationships within a certain neighborhood. They are not merely point information or representations of a location, but rather represent the combination of context and surrounding neighborhood. In this embodiment, a target generation model is designed to extract features from 2D keypoint data and generate corresponding 3D keypoint results.
[0102] The terminal calls the target generation model to process the two-dimensional keypoint data and obtain the corresponding three-dimensional keypoints. Specifically, the target generation model in this application includes a fully connected input module, a feature extraction module, and a fully connected output module. The fully connected input module and the fully connected output module are multilayer perceptron (MLP) structures, and the feature extraction module includes directed graph convolutional layers and self-attention layers. This target generation model is trained using a sample dataset.
[0103] Step 303: Input the two-dimensional key point data into the target generation model, and perform feature mapping processing on the two-dimensional key point data through the fully connected input module to obtain the first feature.
[0104] The fully connected input module is an MLP structure, which includes at least an input layer, a hidden layer, and an output layer. All layers are fully connected, meaning that any neuron in one layer is connected to all neurons in the next layer. After the terminal inputs 2D keypoint data into the target generation model, the fully connected input module performs feature mapping processing on the 2D keypoint data to obtain the first feature. This feature mapping specifically refers to mapping the 2D keypoint coordinates into a feature matrix.
[0105] In this embodiment and subsequent embodiments, the total number of two-dimensional keypoint coordinates obtained is 17. Then, the input of the fully connected input module is a 17*2 matrix. The fully connected input module is composed of 17 different MLPs. The two-dimensional keypoint data is processed by feature mapping through the fully connected input module to obtain a two-dimensional feature matrix, which is the first feature in this application.
[0106] Step 304: Input the first feature into the feature extraction module to generate the second feature. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0107] The feature extraction module in this application uses a combination of a directed graph convolutional network (GCN) and a self-attention mechanism. Since the human body is naturally a graph structure, existing undirected graph convolutional methods do not separately capture information from the core to the limbs and from the limbs to the core, which is crucial for human pose estimation, especially in scenarios lacking certain keypoints. Experiments demonstrate that the directed GCN outperforms the traditional undirected GCN with three times the parameters.
[0108] The directed graph convolutional layer and the self-attention layer in this embodiment are described below:
[0109] 1. Directed graph convolutional layer:
[0110] Graph Convolutional Networks (GCNs) differ from conventional convolutional networks in that they process graph-structured data, not image-structured data. Furthermore, directed GCNs differ from traditional GCNs in that traditional GCNs do not consider the connection direction between nodes, while directed GCNs categorize node correlations according to direction. This direction includes outward-to-inward, inward-to-outward, and self-defined directions. In this embodiment, the direction of the directed GCN can be defined as the direction from the hip bone point to the edge, and the direction from the edge to the hip bone point, thereby capturing feature information of the human body from the core to the limbs and from the limbs to the core.
[0111] For GCN, the input is typically an n*d matrix, where n represents n vectors, d represents the channel depth of each vector (d dimensions), and there are complex graph connections between the n vectors, represented by an n*n matrix A. Matrix A contains only 0 and 1; 0 indicates no connection between row and column nodes, and 1 indicates a connection. For example, if the value at row 3, column 4 of the n*n matrix is 1, it means that the 3rd and 4th vectors are connected. The GCN calculation formula is AXW, where A is an n*n adjacency matrix, X is the input matrix (n*d), and W is the parameter matrix (d*D), where D is the dimension of the output matrix (which can be defined). Therefore, the input to GCN is an n*d matrix, and the output is an n*D matrix. In directed GCN, the core information is from the outside to the inside. First, the connection matrix A only contains the connection relationships from the outside to the inside. For example, in an n*n connection matrix A, if the value of row 3 and column 4 is 1, it means that the vector from node 3 to node 4 exists. However, if the value of row 4 and column 3 is 0, it means that the vector from node 4 to node 3 does not exist. In this way, the network can know what the connection relationships are between each other, thus capturing the information from the outside to the inside. The same logic applies to the information from the inside to the outside.
[0112] 2. Self-attention layer:
[0113] Attention mechanisms mimic the internal processes of biological observation, aligning internal experience with external sensations to increase the precision of observation in specific areas. Attention mechanisms can quickly extract important features from sparse data. Self-attention mechanisms, an improvement on attention mechanisms, reduce reliance on external information and are better at capturing the internal correlations of data or features. The feature extraction module includes a self-attention layer. Its input is an n*d matrix, which is first passed through three different fully connected layers to obtain three n*d matrices Q, K, and V. Then, the n*n matrix Q*KT is passed through a softmax function and left-multiplied by V to obtain an n*d output matrix.
[0114] The terminal inputs the first feature into the feature extraction module, which captures local features of the first feature through directed graph convolutional layers and global features through self-attention layers. Specifically, in directed GCN, each output vector fuses information from nodes that are forward, backward, or related to that node, thus capturing local features but not global features. In contrast, each output vector of self-attention is calculated by fusing all vectors according to their weights, thus capturing global features. In other words, in directed GCN, each node can only see the features of its neighbors, while self-attention can see the features of all nodes.
[0115] Specifically, the input to the feature extraction module is a 17*64 matrix (first feature), which passes through a directed graph convolutional layer and a self-attention layer to obtain a new 17*64 matrix (second feature). The two matrix features output by the directed graph convolutional layer and the self-attention layer are both 17*64 in shape. The matrix feature finally output by the feature extraction module, i.e. the second feature, contains local feature fusion and global feature encoding, thus making the semantics richer and facilitating the subsequent calculation of 3D keypoint results.
[0116] Step 305: Perform keypoint mapping processing on the second feature through the fully connected output module to obtain the three-dimensional keypoints output by the target generation model.
[0117] The terminal will perform keypoint mapping processing on the two-dimensional key data through a fully connected output module to obtain three-dimensional keypoints. This keypoint mapping processing specifically refers to the process of mapping the feature matrix into three-dimensional keypoints. The fully connected output module is also an MLP structure. The second feature generated in step 304 is a 17*64 matrix feature. The matrix is flattened to obtain a 1*1088 vector, which is then input to a 1088*48 fully connected output module to obtain a 1*48 output vector. After reshaping to 16*3, the three-dimensional coordinates, i.e., the three-dimensional keypoints in this application, can be obtained.
[0118] In this embodiment, the target generation model can generate three-dimensional coordinates from the two-dimensional coordinates of key points, thereby achieving three-dimensional key point detection based on two-dimensional key points. Specifically, the target generation model adopts an architecture that uses both directed graph convolutional layers and self-attention layers. The directed graph convolutional layers are responsible for capturing local features, while the self-attention layers are responsible for capturing global features. The local and global features are then organically combined to improve model accuracy while ensuring real-time performance, thus enhancing the accuracy of the three-dimensional key point results.
[0119] Please see Figure 4 , Figure 4 Another embodiment of the method for generating three-dimensional key points provided in this application includes:
[0120] Step 401: Obtain the two-dimensional key point data to be processed.
[0121] In this embodiment, step 501 is similar to the aforementioned step 301, and will not be described again here.
[0122] Step 402: Call the target generation model. The target generation model includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes at least one sub-module group. The sub-module group includes two backbone sub-modules and one feature stacking sub-module. The backbone sub-module includes a directed graph convolutional layer and a self-attention layer.
[0123] Please see Figure 5 , Figure 5 This is a schematic diagram of the network structure of the target generation model provided in this embodiment. The target generation model includes a fully connected input module 501, a feature extraction module 502, and a fully connected output module 503. The feature extraction module 502 includes at least one sub-module group. Each sub-module group includes two backbone sub-modules 5021 and one feature stacking sub-module 5022. Each backbone sub-module includes a directed graph convolutional layer and a self-attention layer.
[0124] Step 403: Input the two-dimensional key point data into the target generation model, and perform feature mapping processing on the two-dimensional key point data through the fully connected input module to obtain the first feature.
[0125] In this embodiment, step 403 is similar to step 303 in the previous embodiment, and will not be described again here.
[0126] Step 404: Input the first feature into the first main submodule of the submodule group to obtain the intermediate feature.
[0127] Please see Figure 6In this embodiment, each backbone sub-module 5021, i.e., block module, includes a directed graph convolutional layer and a self-attention layer, specifically composed of a directed graph convolutional layer and a self-attention layer connected in series. A detailed description of the directed graph convolutional layer and the self-attention layer is given in step 304, and will not be repeated here. The specific steps for inputting the first feature into the first backbone sub-module of the sub-module group to obtain intermediate features are as follows:
[0128] Step A: Capture the local features of the first feature through a directed graph convolutional layer, capture the global features of the first feature through a self-attention layer, and obtain the intermediate features by fusing the local and global features.
[0129] It should be noted that the order of the directed graph convolutional layer and the self-attention layer in the main submodule 5021 can be reversed. That is, it is possible to run the directed graph convolutional layer first and then the self-attention layer in step A, or to run the self-attention layer first and then the directed graph convolutional layer. No specific restrictions are imposed here.
[0130] For further details, please refer to Figure 7 The directed graph convolutional layer in the backbone submodule 5021 includes a forward GCN, a backward GCN, and its own GCN, specifically composed of the forward GCN, backward GCN, and its own GCN connected in parallel. Step A will be described in detail below based on the specific structure of this directed graph convolutional layer:
[0131] Step A1: Extract the first node features through the forward graph convolutional layer. The first node features are used to represent the forward connection relationships in the first features.
[0132] In this embodiment, the forward direction in the forward graph convolutional layer is defined as the direction from the human hip bone point to the edge. Therefore, the terminal can extract the feature information (first node feature) from the human core to the limbs in the first feature through this forward graph convolutional layer.
[0133] Step A2: Extract the second node features through the inverse graph convolutional layer. The second node features are used to characterize the inverse connection relationship in the first feature.
[0134] In this embodiment, the reverse direction in the reverse graph convolutional layer is defined as the direction from the edge to the hip bone of the human body. Therefore, the terminal can extract the feature information (second node feature) from the limbs to the core of the human body in the first feature through this reverse graph convolutional layer.
[0135] Step A3: Extract the third node feature through the self-graph convolutional layer. The third node feature is used to represent the connection relationship between itself and itself in the first feature.
[0136] The self-graph convolutional layer is an n*n identity matrix that only considers connections between itself. The terminal can extract its own feature information (third node features) from the first feature using this self-graph convolutional layer.
[0137] Step A4: Fuse the features of the first node, the second node, and the third node to capture the local features of the first feature.
[0138] Finally, the terminal fuses the node features extracted from each directed graph convolutional layer to capture the local features of the first feature. It should be noted that the self-GCN in the directed graph convolutional layer can be omitted, that is, step A3 can be omitted. In step A4, only the first node features and the second node features need to be fused to capture the local features of the first feature. Regardless of whether the self-GCN is included, its performance is better than the undirected GCN in the existing technology.
[0139] In short, the only difference between forward GCN and GCN is that the connection matrix A only considers connections spreading from the body across bone points to the edge, while backward GCN does the opposite, only considering connections from the edge to the body across bone points. Self GCN is an n*n identity matrix that only considers connections between itself. Therefore, forward / backward / self GCNs take the same 17*64 matrix as input and output three 17*64 matrices. Summing these matrices yields the 17*64 output matrix of the directed graph convolutional layer, which can then be used as a local feature of the first feature.
[0140] Step 405: Overlay the first feature and intermediate features through the feature overlay submodule, and input the overlay result into the second main submodule to obtain the second feature output by the submodule group.
[0141] Submodule group 502 has a g(x+f(x)) structure, such as Figure 5 As shown, the 17*64 first feature is processed by the first backbone submodule 5021a to obtain a 17*64 output (intermediate feature). The intermediate feature and the input (first feature) of the first backbone submodule 5021a are superimposed by the feature superposition submodule 5022 as a residual to obtain another 17*64 matrix. Then, this matrix is input into the second backbone submodule 5021b to output a 17*64 matrix (no residual is made at this time). This is the structural detail of g(x+f(x)).
[0142] It should be noted that steps 404 to 405 can be repeated multiple times, meaning that multiple sub-module groups 502 can be chained together in the target generation model. It's important to note that what is repeated is the structure of the sub-module group 502; the parameters in different sub-module groups 502 are not shared, each using its own parameters. In practical applications, sub-module groups 502 can be stacked twice, or increased to three, four, etc., depending on the model's real-time performance and accuracy metrics. No specific limitation is made here. Existing resGCN structures reach their performance limit and begin to decline after five stacks due to the common GCN oversmoothing problem. In this embodiment, by stacking sub-module groups 502, g(x+f(x)) forms a special stacked structure. Because the direct backpropagation link from the loss to the input is removed, oversmoothing is effectively avoided, and performance can be further improved even after 20 stacks or more.
[0143] Step 406: Perform keypoint mapping processing on the second feature through the fully connected output module to obtain the three-dimensional keypoints output by the target generation model.
[0144] In this embodiment, step 406 is similar to step 305 in the previous embodiment, and will not be described again here.
[0145] In this embodiment, a directed graph convolutional network (GCN) is used in the target generation model instead of the undirected GCN of existing technologies. Since the human body naturally possesses a graph structure, the target generation model specifically employs an architecture that combines directed graph convolutional layers and self-attention layers. The directed graph convolutional layers capture local features of the human body from the core to the limbs or from the limbs to the core, while the self-attention layers capture global features. This organic combination of local and global features improves model accuracy while ensuring real-time performance, thereby enhancing the accuracy of the 3D keypoint results. Furthermore, this target generation model utilizes a special stacking method with multiple layers of directed graph convolutions to avoid the oversmoothing problem, further improving the accuracy of the 3D keypoint results.
[0146] The training method for the target generation model provided in this application is described in detail below. Please refer to [link / reference]. Figure 6 The training method includes:
[0147] Step 601: Obtain the sample dataset, which contains multiple sets of matching two-dimensional and three-dimensional sample data.
[0148] During the training phase, a dataset needs to be prepared, consisting of multiple sets of matched 2D / 3D coordinate data. Again, using 17 keypoints as an example, the sample dataset should contain multiple sets of sample data, each set including the 2D and 3D coordinates of the 17 keypoints.
[0149] Step 602: Obtain the initial generation model. The initial generation model includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0150] The structure of the initial generation model in this embodiment is as follows: Figure 5 As shown, the data flow in the initial generated model is as follows: Figure 1 , 4 The corresponding embodiments are shown, and will not be repeated here.
[0151] Step 603: Randomly select a set of matching two-dimensional sample data and three-dimensional sample data from the sample dataset, input the two-dimensional sample data into the initial generation model, and obtain the first three-dimensional key points output by the initial generation model.
[0152] Two-dimensional sample data is input into the initial generation model. The two-dimensional sample data is processed by feature mapping through the fully connected input module to obtain the first sample features. The first sample features are input into the feature extraction module to generate the second sample features. The feature extraction module can be stacked multiple times. Finally, the second sample features are processed by key point mapping through the fully connected output module to obtain the first three-dimensional key points output by the initial generation model. The first three-dimensional key points are the refined result of the initial generation model.
[0153] Step 604: Calculate the first loss value based on the first three-dimensional key points and three-dimensional sample data, and iteratively train the initial generation model based on the first loss value to obtain the target generation model. The target generation model is used to generate three-dimensional key point coordinates based on the two-dimensional key point coordinates.
[0154] Two-dimensional sample data is input into the initial generative model to obtain the first three-dimensional key point output by the initial generative model. Then, the first loss value between the first three-dimensional key point and the three-dimensional sample data is calculated. The initial generative model is trained using the first loss value. After that, a new set of two-dimensional sample data and three-dimensional sample data are selected to train the initial generative model again until the training times reach the preset number of times. The trained initial generative model is then determined as the target generative model.
[0155] It should be noted that the first loss value can be calculated using the mean squared difference loss function MSELoss or the cross-entropy loss function. The training method can be either mini-batch gradient descent or stochastic gradient descent, and the specific method is not limited here.
[0156] Furthermore, to accelerate the training process and reduce training difficulty, a loss function can be added in the middle of the network. This means that while training the initial generative model using the refined results described above, the initial generative model is also trained using the coarse results output from the middle of the network. Taking a double-layered model containing two sub-modules as an example, the output of the first sub-module is mapped to 3D coordinates (second 3D keypoints) using an MLP, and the output of the second sub-module is also mapped to 3D coordinates (first 3D keypoints) using an MLP. Loss values are calculated separately using the second and first 3D keypoints and the 3D sample data, respectively, and then the loss values are backpropagated to update the network parameters.
[0157] Please see Figure 6 , Figure 6 One embodiment of the three-dimensional key point generation apparatus provided in this application includes:
[0158] Data acquisition unit 601 is used to acquire two-dimensional key point data to be processed;
[0159] The model invocation unit 602 is used to invoke the target generation model, which includes a fully connected input module, a feature extraction module, and a fully connected output module.
[0160] The data input unit 603 is used to input two-dimensional key point data into the target generation model;
[0161] The model processing unit 604 is used to perform feature mapping processing on the two-dimensional keypoint data through the fully connected input module to obtain the first feature; input the first feature into the feature extraction module to generate the second feature, the feature extraction module including a directed graph convolutional layer and a self-attention layer; and perform keypoint mapping processing on the second feature through the fully connected output module to obtain the three-dimensional keypoints output by the target generation model.
[0162] Optionally, the feature extraction module includes at least one sub-module group, which includes two main sub-modules and one feature overlay sub-module. The model processing unit 604 is specifically used for:
[0163] The first feature is input into the first main submodule of the submodule group to obtain the intermediate feature;
[0164] The first feature and intermediate features are superimposed by the feature superposition submodule, and the superposition result is input into the second main submodule to obtain the second feature output by the submodule group.
[0165] Optionally, the backbone submodule includes directed graph convolutional layers and self-attention layers, and the model processing unit is further used for:
[0166] The first feature is captured by a directed graph convolutional layer, the first feature is captured by a self-attention layer, and the intermediate feature is obtained by fusing the local and global features.
[0167] Optionally, the directed graph convolutional layer includes a forward graph convolutional layer and a backward graph convolutional layer, and the model processing unit 604 is further used for:
[0168] The first node features are extracted by a forward graph convolutional layer. The first node features are used to represent the forward connection relationships in the first features.
[0169] The second node features are extracted by the inverse graph convolutional layer. The second node features are used to characterize the inverse connection relationship in the first feature.
[0170] By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
[0171] Optionally, the directed graph convolutional layer also includes its own graph convolutional layer, and the model processing unit 604 is specifically used for:
[0172] The first node features are extracted by a forward graph convolutional layer. The first node features are used to represent the forward connection relationships in the first features.
[0173] By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
[0174] Optionally, the generation device also includes: a model training unit 605;
[0175] The model training unit 605 is specifically used for:
[0176] Obtain the sample dataset, which contains multiple sets of matching two-dimensional and three-dimensional sample data;
[0177] Obtain the initial generative model, which includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes a directed graph convolutional layer and a self-attention layer.
[0178] Randomly select a set of matching two-dimensional sample data and three-dimensional sample data from the sample dataset, input the two-dimensional sample data into the initial generation model, and obtain the first three-dimensional key points output by the initial generation model;
[0179] The first loss value is calculated based on the first 3D key points and 3D sample data, and the initial generation model is iteratively trained based on the first loss value to obtain the target generation model. The target generation model is used to generate 3D key points based on 2D key points.
[0180] Optionally, the model training unit 605 is also specifically used for:
[0181] Obtain the second 3D keypoints output from the middle segment of the initial generated model network;
[0182] The second loss value is calculated based on the second three-dimensional key points and three-dimensional sample data;
[0183] Iterative training of the initial generative model based on the first loss value includes:
[0184] The initial generative model is iteratively trained based on the first and second loss values.
[0185] In this embodiment, the functions of each unit are the same as described above. Figure 3 or Figure 4 The steps in the method embodiments shown correspond to those in the examples, and will not be repeated here.
[0186] This application also provides a device for generating three-dimensional key points; please refer to [link / reference]. Figure 7 , Figure 7 One embodiment of the robot provided in this application includes:
[0187] Processor 701, memory 702, input / output unit 703, bus 704;
[0188] The processor 701 is connected to the memory 702, the input / output unit 703, and the bus 704;
[0189] The memory 702 stores a program, and the processor 701 calls the program to execute any of the above methods for generating three-dimensional key points.
[0190] This application also relates to a computer-readable storage medium storing a program, characterized in that, when the program is run on a computer, it causes the computer to execute any of the above-described methods for generating three-dimensional key points.
[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0192] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0193] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0194] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0195] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for generating three-dimensional key points, characterized in that, The method includes: The system acquires an image to be detected by a camera or infrared camera, and performs two-dimensional human key point detection on the image to be detected to obtain two-dimensional key point data to be processed. The two-dimensional key point data includes the two-dimensional coordinates of human key points in the image to be detected. The target generation model is invoked, which includes a fully connected input module, a feature extraction module, and a fully connected output module. The two-dimensional key point data is input into the target generation model, and the two-dimensional key point data is processed by feature mapping through the fully connected input module to obtain the first feature; The first feature is input into the feature extraction module to generate the second feature. The feature extraction module includes a directed graph convolutional layer and a self-attention layer. The second feature is processed by keypoint mapping through the fully connected output module to obtain the human body three-dimensional keypoints output by the target generation model. The human body three-dimensional keypoints are used for human pose estimation. The feature extraction module includes at least one sub-module group, which includes two main sub-modules and one feature overlay sub-module. The step of inputting the first feature into the feature extraction module to generate the second feature includes: The first feature is input into the first main submodule of the submodule group to obtain the intermediate feature; The first feature and the intermediate feature are superimposed by the feature superposition submodule, and the superposition result is input to the second backbone submodule to obtain the second feature output by the submodule group. The backbone submodule includes a directed graph convolutional layer and a self-attention layer. The step of inputting the first feature into the first backbone submodule to obtain intermediate features includes: The directed graph convolutional layer captures the local features of the first feature, the self-attention layer captures the global features of the first feature, and the intermediate feature is obtained by fusing the local features and the global features. The directed graph convolutional layer includes a forward graph convolutional layer and a backward graph convolutional layer, and the local features captured by the directed graph convolutional layer for the first feature include: The first node feature is extracted through the forward graph convolutional layer. The first node feature is used to represent the forward connection relationship from the human hip bone point to the edge in the first feature. The second node feature is extracted through the inverse graph convolutional layer. The second node feature is used to represent the inverse connection relationship from the edge to the human hip bone point in the first feature. By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
2. The method according to claim 1, characterized in that, The directed graph convolutional layer further includes an autograph convolutional layer. Before fusing the first node features and the second node features to capture local features of the first feature, the method further includes: The third node feature is extracted through the self-graph convolutional layer, and the third node feature is used to characterize the connection relationship between itself and itself in the first feature. The process of fusing the first node features and the second node features to capture local features of the first feature includes: By fusing the features of the first node, the second node, and the third node, local features of the first feature are captured.
3. The method according to any one of claims 1 to 2, characterized in that, The method further includes training the target generation model, wherein training the target generation model includes: Obtain a sample dataset, which contains multiple sets of matching two-dimensional and three-dimensional sample data; Obtain an initial generative model, which includes a fully connected input module, a feature extraction module, and a fully connected output module. The feature extraction module includes a directed graph convolutional layer and a self-attention layer. A set of matching two-dimensional sample data and three-dimensional sample data are randomly selected from the sample dataset. The two-dimensional sample data is input into the initial generation model to obtain the first three-dimensional key points output by the initial generation model. A first loss value is calculated based on the first 3D key points and the 3D sample data, and the initial generation model is iteratively trained based on the first loss value to obtain a target generation model. The target generation model is used to generate 3D key points based on 2D key points.
4. The method according to claim 3, characterized in that, Before iteratively training the initial generative model based on the first loss value, the method further includes: Obtain the second three-dimensional key points output from the middle segment of the network of the initial generated model; Calculate the second loss value based on the second three-dimensional key points and the three-dimensional sample data; The iterative training of the initial generative model based on the first loss value includes: The initial generative model is iteratively trained based on the first loss value and the second loss value.
5. A device for generating three-dimensional key points, characterized in that, The apparatus is used to perform the method as described in claim 1, the apparatus comprising: The data acquisition unit is used to acquire the image to be detected captured by a camera or an infrared camera, and to perform two-dimensional human key point detection on the image to be detected to obtain two-dimensional key point data to be processed. The two-dimensional key point data includes the two-dimensional coordinates of the human key points in the image to be detected. The model invocation unit is used to invoke the target generation model, which includes a fully connected input module, a feature extraction module, and a fully connected output module. A data input unit is used to input the two-dimensional keypoint data into the target generation model, perform feature mapping processing on the two-dimensional keypoint data through the fully connected input module to obtain a first feature, and input the first feature into the feature extraction module to generate a second feature. The feature extraction module includes a directed graph convolutional layer and a self-attention layer. The model processing unit is used to perform key point mapping processing on the second feature through the fully connected output module to obtain the human body three-dimensional key points output by the target generation model. The human body three-dimensional key points are used for human pose estimation. The feature extraction module includes at least one sub-module group, which includes two main sub-modules and one feature overlay sub-module. The step of inputting the first feature into the feature extraction module to generate the second feature includes: The first feature is input into the first main submodule of the submodule group to obtain the intermediate feature; The first feature and the intermediate feature are superimposed by the feature superposition submodule, and the superposition result is input to the second backbone submodule to obtain the second feature output by the submodule group. The backbone submodule includes a directed graph convolutional layer and a self-attention layer. The step of inputting the first feature into the first backbone submodule to obtain intermediate features includes: The directed graph convolutional layer captures the local features of the first feature, the self-attention layer captures the global features of the first feature, and the intermediate feature is obtained by fusing the local features and the global features. The directed graph convolutional layer includes a forward graph convolutional layer and a backward graph convolutional layer, and the local features captured by the directed graph convolutional layer for the first feature include: The first node feature is extracted through the forward graph convolutional layer. The first node feature is used to represent the forward connection relationship from the human hip bone point to the edge in the first feature. The second node feature is extracted through the inverse graph convolutional layer. The second node feature is used to represent the inverse connection relationship from the edge to the human hip bone point in the first feature. By fusing the features of the first node and the features of the second node, local features of the first feature are captured.
6. A robot, characterized in that, The robot includes: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium having a program stored thereon, the program performing the method as described in any one of claims 1 to 4 when executed on a computer.
Citation Information
Patent Citations
Three-dimensional reconstruction method, apparatus and device, and computer storage medium
CN114219890A