Method for constructing dataset of three-dimensional model, model training method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-19
- Publication Date
- 2026-08-11
Smart Images

Figure CN117173515B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to the fields of computer vision, augmented reality, virtual reality, etc., and particularly to methods for constructing datasets for 3D models, methods for model training, apparatus for constructing datasets, apparatus for model training, electronic devices, storage media, and program products, which can be applied to scenarios such as metaverse, digital humans, and generative artificial intelligence (AIGC). Background Technology
[0002] Computer vision has long been dedicated to understanding human movements and behaviors, finding applications in robotics, healthcare, virtual try-on, AR / VR, and other fields. Significant progress has been made in 2D human pose detection and 3D human pose estimation from single images, thanks to real-world datasets annotated with 2D keypoints and 3D data. Therefore, constructing such datasets remains a crucial problem to solve. Summary of the Invention
[0003] This disclosure provides a method for constructing a dataset for a 3D model, a model training method, a dataset construction apparatus, a model training apparatus, an electronic device, a storage medium, and a program product, which can construct a dataset for model training.
[0004] According to one aspect of this disclosure, a method for constructing a dataset for a three-dimensional model is provided, comprising: establishing a three-dimensional human body model based on an acquired human body image; determining the pose data and shape data of the three-dimensional human body model and adding the three-dimensional human body model to a scene model; determining the contact relationship between the three-dimensional human body model and the scene model based on the mesh data of the three-dimensional human body model and the mesh data of the scene model; and constructing a dataset using the human body image, pose data, shape data and contact relationship.
[0005] According to another aspect of this disclosure, a model training method is provided, comprising: acquiring the dataset mentioned above; and using the dataset to train a model to obtain a human pose estimation model and / or a contact estimation model.
[0006] According to another aspect of this disclosure, a dataset construction apparatus is provided, comprising: a modeling module, an adding module, a determining module, and a construction module. The modeling module is configured to build a three-dimensional human body model based on acquired human body images. The adding module is configured to determine the pose data and shape data of the three-dimensional human body model and add the three-dimensional human body model to a scene model. The determining module is configured to determine the contact relationship between the three-dimensional human body model and the scene model based on the mesh data of the three-dimensional human body model and the mesh data of the scene model. The construction module is configured to construct the dataset using the human body images, pose data, shape data, and contact relationships.
[0007] According to another aspect of this disclosure, a model training apparatus is provided, comprising: a data acquisition module and a model training module. The data acquisition module is configured to acquire the dataset mentioned above. The model training module is configured to train a model using the dataset to obtain a human pose estimation model and / or a contact estimation model.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor and a memory. The memory is communicatively connected to the at least one processor and stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods mentioned above.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods mentioned above.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods mentioned above.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0013] Figure 1 A schematic block diagram of an exemplary system architecture for constructing datasets or training models for 3D models, which can be applied according to the present disclosure, is shown.
[0014] Figure 2 This is a flowchart illustrating a method for constructing a dataset for a three-dimensional model according to a first embodiment of the present disclosure;
[0015] Figure 3 This is a flowchart illustrating a method for constructing a dataset for a three-dimensional model according to a second embodiment of the present disclosure;
[0016] Figure 4 This is a schematic flowchart of a model training method according to a third embodiment of the present disclosure;
[0017] Figure 5 This is a schematic flowchart of a model training method according to the fourth embodiment of the present disclosure;
[0018] Figure 6 This is a schematic diagram of the architecture of the contact estimation model according to the fourth embodiment of the present disclosure;
[0019] Figure 7 This is a schematic block diagram of a dataset construction apparatus according to the fifth embodiment of the present disclosure;
[0020] Figure 8 This is a schematic block diagram of a model training apparatus according to the sixth embodiment of the present disclosure;
[0021] Figure 9 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0023] Figure 1 This is a schematic block diagram of an exemplary system architecture for constructing a dataset or training a model for a 3D model, according to a first embodiment of the present disclosure.
[0024] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0025] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as video applications, live streaming applications, instant messaging tools, email clients, social media platform software, etc.
[0026] The terminal devices 101, 102, and 103 here can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0027] Server 105 can be a server that provides various services, such as a backend server that supports terminal devices 101, 102, and 103. The backend server can analyze and process received data such as human images and feed back the processing results (such as datasets) to the terminal devices.
[0028] It should be noted that the data set construction method or model training method for three-dimensional models provided in this disclosure embodiment can be executed by server 105 or terminal devices 101, 102, 103. Accordingly, the data set construction device or model training device can be set in server 105 or terminal devices 101, 102, 103.
[0029] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0030] Continue to refer to Figure 2 , Figure 2 This is a flowchart illustrating a method for constructing a dataset for a 3D model according to a second embodiment of this disclosure. Figure 2 As shown, the method 200 for constructing a dataset for a 3D model may include the following steps:
[0031] Step 201: Establish a three-dimensional human body model based on the acquired human body images.
[0032] In this embodiment, the executing entity (e.g., a server or terminal device) can acquire human images captured by a camera and perform three-dimensional reconstruction based on the human images to obtain a three-dimensional human model.
[0033] In some optional embodiments of this disclosure, the executing entity can acquire human body images captured from multiple perspectives and perform multi-view processing on the human body images. Figure 3 3D reconstruction yields a 3D human body model. For example, the executing entity captures the target human body using multiple cameras. The number and position of the cameras can be set as needed, with multiple cameras covering as many angles and perspectives of the target human body as possible. Optionally, the camera positions and angles of different cameras can vary significantly, so that the videos captured by different cameras can represent the target human body from different angles and perspectives. After obtaining the human body images, the executing entity can perform multi-view fitting based on the human body images to obtain a 3D human body model. For example, the executing entity can perform 2D keypoint detection on the human body images captured from multiple perspectives, and triangulate the 2D keypoints in the human body images to obtain 3D keypoints, and then perform 3D reconstruction based on the 3D keypoints.
[0034] It's worth noting that acquiring multi-view images of the human body allows for the acquisition of more details from various perspectives, resulting in a more accurate 3D model that reflects the target human body's data. Furthermore, multi-view 3D human reconstruction better captures the contact relationship between the human body and the scene, making the interaction more realistic. This is significant for applications such as augmented reality and virtual reality, enhancing the user's immersion and sense of presence.
[0035] It should be understood that, without departing from the teachings of this disclosure, a synchronous camera or other cameras may be used, and this disclosure does not restrict this choice.
[0036] In some optional embodiments of this disclosure, the process by which the executing entity obtains a three-dimensional human model based on human images from multiple perspectives may include: performing human body recognition on each frame of the human body images captured from each perspective, determining the position information of the identified target human body in each human body image; and, for the target human body, locking the image information corresponding to the target human body in the human body image based on the position information; and, based on the image information, performing multi-view... Figure 3 3D reconstruction techniques are used to obtain a 3D human body model. For example, cameras from different perspectives can capture videos of the target human body within a scene. For the captured multi-view videos, the actor can perform human body recognition using cross-view and cross-temporal methods. Cross-view recognition allows the identification of the same target human body in videos from different perspectives. Cross-temporal recognition allows the tracking of the same target human body within a video sequence. For each identified target human body, a multi-view fitting method is used to reconstruct it.
[0037] It is worth mentioning that the reconstruction of 3D models through multi-view fitting has robustness to the detection of noisy 2D key points, which can improve the quality of the reconstructed 3D human body model.
[0038] It should be understood that, without departing from the teachings of this publication, the executing entity may perform operations such as frame alignment and camera calibration on human images captured from multiple perspectives before performing multi-view fitting, and there are no restrictions here.
[0039] It should be understood that, without departing from the teachings of this disclosure, the implementing entity may also reconstruct the three-dimensional human body model in other ways, such as by taking a depth image of the target human body with a depth camera and then reconstructing the three-dimensional model based on the depth image. This disclosure does not limit the method of three-dimensional reconstruction of the three-dimensional human body model.
[0040] Step 202: Determine the pose and shape data of the 3D human body model, and add the 3D human body model to the scene model.
[0041] In this embodiment, the executing entity can obtain the posture and shape data of a 3D human body model through 3D reconstruction. After completing the 3D reconstruction of the 3D human body model, the executing entity can add the 3D human body model to the scene model based on the human body image in order to determine the contact relationship between the 3D human body model and the scene model.
[0042] In some alternative embodiments of this disclosure, the scene model is obtained through pre-scanning. For example, a three-dimensional coordinate point cloud is generated by measuring the reflection time of points on the scene using an industrial laser scanner. After calibration and registration of the three-dimensional coordinate point cloud data, three-dimensional reconstruction is performed to obtain the scene model.
[0043] It should be understood that scenario models can be constructed in other ways without departing from the teachings of this disclosure, and this disclosure does not impose any restrictions on this.
[0044] Step 203: Determine the contact relationship between the 3D human body model and the scene model based on the mesh data of the 3D human body model and the mesh data of the scene model.
[0045] In this embodiment, after adding the 3D human body model to the scene model, the executing entity can determine the mesh data of the 3D human body model and the mesh data of the scene model respectively, and determine the contact relationship between the 3D human body model and the scene model based on the mesh data of the 3D human body model and the mesh data of the scene model.
[0046] It is worth mentioning that by introducing scene awareness and multi-view fitting for body scene contact estimation, more accurate human pose estimation results can be provided. This will have a positive impact on applications in virtual try-on, human-computer interaction, animation production, and other fields, improving user experience and the naturalness of interaction. Furthermore, by determining contact relationships based on the mesh data of the 3D human body model and the scene model, contact relationships at the mesh accuracy level can be obtained, improving the accuracy of contact relationships.
[0047] In some optional embodiments of this disclosure, the execution entity may convert the scene model into a mesh representation through triangulation, Poisson reconstruction, normal estimation and surface reconstruction, deep learning methods, etc. This disclosure does not limit the process of meshing the scene model.
[0048] In some optional embodiments of this disclosure, the executing entity can input the pose and shape data of the target human body into a parametric human body model (Skinned Multi-Person Linear Model, SMPL) to obtain the mesh data of the three-dimensional human body model. The SMPL model takes the human body pose and shape as input and outputs a three-dimensional human body mesh.
[0049] It should be understood that, without departing from the teachings of this disclosure, the implementing entity may also obtain the mesh data of the three-dimensional human body model through methods such as triangulation, and this disclosure does not restrict the process of meshing the three-dimensional human body model.
[0050] Step 204: Construct a dataset using human images, pose data, shape data, and contact relationships.
[0051] In this embodiment, the executing entity constructs a dataset based on the results of the above steps, focusing on human-scene interaction. This dataset includes the target human's pose data, shape data, and contact relationships indicating whether the target human and the scene are in contact. This data will provide more realistic training and evaluation data for subsequent human pose estimation algorithms.
[0052] In some optional embodiments of this disclosure, contact relationships can be stored as grid-level scene contact tags, meaning the contact relationships reflect the contact status of each grid vertex of the 3D human body model corresponding to the human body image. Recording the contact relationships between the human body and the scene using grid-level scene contact tags can improve the accuracy of the contact relationships.
[0053] According to the first embodiment of this disclosure, in the process of constructing a dataset for model training, the contact relationship between the human body and the scene is confirmed, introducing scene awareness. This provides more realistic and accurate data for research in areas such as human pose estimation in computer vision, further improving the performance and robustness of human pose estimation algorithms, enhancing the effect and experience of computer vision applications, and reducing unrealistic results caused by ignoring the scene and estimating human pose in isolation. By constructing a scene-interaction-based dataset, a richer and more practical data foundation can be provided for fields such as human pose estimation and scene understanding, helping to improve existing datasets, enhance the effectiveness of training models, and promote the development of research and applications in related fields. By fusing human pose and scene information, more accurate scene understanding and scene reconstruction results can be provided. This is very important for research and applications in fields such as 3D reconstruction and computer graphics, helping to improve the accuracy of models and the realism of reconstruction results. Furthermore, determining the contact relationship based on the mesh data of the 3D human model and the mesh data of the scene model can obtain contact relationships at the mesh accuracy level, improving the accuracy of contact relationships.
[0054] See also Figure 3 , Figure 3 This is a flowchart illustrating a method for constructing a dataset for a 3D model according to a second embodiment of this disclosure. Figure 3 As shown, the method 300 for constructing a dataset for a 3D model may include the following steps:
[0055] Step 301: Establish a three-dimensional human body model based on the acquired human body image.
[0056] Step 302: Determine the pose and shape data of the 3D human body model, and add the 3D human body model to the scene model.
[0057] Among them, steps 301 and 302 are related to Figure 2 Steps 201 and 202 in the example are largely the same, and will not be repeated here.
[0058] Step 303: Determine the vertices of the human body mesh based on the mesh data of the 3D human body model.
[0059] In this embodiment, the executing entity can mesh the 3D human body model using the SMPL model to obtain the mesh data of the 3D human body model. The mesh data output by the SMPL model includes the data of the human body mesh vertices; therefore, the executing entity can determine the human body mesh vertices based on the output of the SMPL model.
[0060] Step 304: Determine the scene mesh vertices based on the mesh data of the scene model.
[0061] In this embodiment, during the process of meshing the scene model using methods such as triangulation, the execution entity can determine the vertices of each scene mesh (i.e., scene mesh vertices) and store the data of each scene mesh vertex. During the process of determining contact relationships, the execution entity can read the scene mesh vertex data from the stored mesh data of the acquired model.
[0062] Step 305: Determine the contact relationship between the 3D human body model and the scene model based on the distance between the vertices of the human body mesh and the vertices of the scene mesh.
[0063] In this embodiment, the executing entity can determine whether the human mesh vertices and scene mesh vertices are in contact based on the distance between the human mesh vertices and the scene mesh vertices and a preset distance threshold. For example, if the distance between the human mesh vertices and scene mesh vertices is less than the distance threshold, it can be determined that the human mesh vertices and scene mesh vertices are in contact; if the distance between the human mesh vertices and scene mesh vertices is greater than or equal to the distance threshold, it can be determined that the human mesh vertices and scene mesh vertices are not in contact. Then, the executing entity can determine the contact relationship between the 3D human model and the scene model based on the contact status of the human mesh vertices of the 3D human model and the scene mesh vertices of the scene model.
[0064] In some optional embodiments of this disclosure, the executing entity may determine the scene mesh vertex that is closest to the human body mesh vertex, and determine the contact relationship between the human body mesh vertex and the scene model based on the distance between the human body mesh vertex and the closest scene mesh vertex.
[0065] It should be understood that, without departing from the teachings of this disclosure, the implementing entity can also determine the distances between human mesh vertices and all scene mesh vertices, and determine the contact relationship between the human mesh vertices and the scene model based on the number of distances less than a distance threshold among all determined distances. For example, if the distance between a human mesh vertex and at least M scene mesh vertices is less than the distance threshold, it can be determined that the human mesh vertex is in contact with the scene model; otherwise, it is determined that the human mesh vertex is not in contact with the scene model. M can be set as needed and is not limited here. This disclosure does not limit the conditions for determining contact between human mesh vertices and the scene model.
[0066] In some optional embodiments of this disclosure, different body parts of the 3D human model are configured with distance thresholds for determining whether they are in contact with the scene. The executing entity can determine the distance threshold corresponding to a human mesh vertex based on the body part to which it belongs. For example, body parts include hands and feet; if a mesh vertex is a hand vertex, the distance threshold for the hand is obtained as the distance threshold for that mesh vertex. The executing entity determines whether the distance between a human mesh vertex and any scene mesh vertex is less than the distance threshold. If it is, it can be determined that the human mesh vertex is in contact with the scene model; if it is not, it can be determined that the human mesh vertex is not in contact with the scene model. The executing entity can determine the contact relationship between the 3D human model and the scene model based on the contact status between each human mesh vertex of the target human and the scene model.
[0067] It is worth mentioning that different distance thresholds are configured for different body parts to determine whether they are in contact with the scene model. This fully considers the individual differences of each body part and makes the determined contact relationship more accurate.
[0068] It should be understood that, without departing from the teachings of this disclosure, the judgment logic used by the implementing entity to determine the contact relationship between the 3D human body model and the scene model based on the contact between each human body mesh vertex and the scene model can be set as needed. For example, this judgment logic can indicate: if the number of human body mesh vertices in contact with the scene model among all human body mesh vertices corresponding to a body part is greater than or equal to a preset number, the contact relationship between the 3D human body model and the scene model indicates that the body part is in contact with the scene model; if the number of human body mesh vertices in contact with the scene model among all human body mesh vertices corresponding to a body part is less than a preset number, the contact relationship between the 3D human body model and the scene model indicates that the body part is not in contact with the scene model. The preset number can be set according to the mesh size, the total number of human body mesh vertices corresponding to the body part, etc., and is not limited here. As another example, this judgment logic can indicate: storing the contact between each human body mesh vertex and the scene model as scene contact labels in a vector or other manner to record the contact relationship between the 3D human body model and the scene model. This disclosure does not restrict the judgment logic for determining the contact relationship between the 3D human body model and the scene model based on the contact between each human body mesh vertex of the target human body and the scene model.
[0069] In some optional embodiments of this disclosure, the body parts of the three-dimensional human model mentioned above may include a first body part exposed to the outside and a second body part covered by clothing, wherein the distance threshold corresponding to the second body part is greater than the distance threshold corresponding to the first body part. For example, the first body part exposed to the outside may include hands, head, etc., and the second body part covered by clothing may include feet. For the first body parts such as the head and hands, since they are directly exposed to the outside, the corresponding distance threshold may be, for example, 2.5 cm. For the second body parts such as feet, considering the thickness of the clothing (e.g., shoe soles), the corresponding distance threshold may be, for example, 5 cm.
[0070] It is worth mentioning that classifying body parts based on whether they are covered by other objects, increasing the distance threshold for body parts covered by clothing, and fully considering the influence of other factors on contact relationships can improve the accuracy of determining contact relationships.
[0071] It should be understood that body parts can be classified based on other characteristics without departing from the teachings of this disclosure, and this disclosure does not impose any restrictions on this.
[0072] In some optional embodiments of this disclosure, when determining whether a human mesh vertex is in contact with a scene model, the executing entity considers not only the distance between the human mesh vertex and the scene mesh vertex, but also the angle between the normals of the human mesh vertex and the scene mesh vertex. For example, if the executing entity determines that the distance between the human mesh vertex and the scene mesh vertex is less than a distance threshold, and the angle between the normals of the human mesh vertex and the scene mesh vertex is less than a preset angle, it determines that the human mesh vertex is in contact with the scene model. If it determines that the distance between the human mesh vertex and the scene mesh vertex is greater than or equal to the distance threshold, or that the angle between the normals of the human mesh vertex and the scene mesh vertex is greater than or equal to the preset angle, it determines that the human mesh vertex is not in contact with the scene model. The preset angle can be set as needed; for example, it can be set to 0 degrees, meaning that the contact condition includes the normals of the human mesh vertex and the scene mesh vertex having the same direction. Then, the executing entity determines the contact relationship between the 3D human model and the scene model based on the contact situation between the human mesh vertex and the scene model.
[0073] It is worth mentioning that fully considering the relationship between the normals of the human mesh vertices and the scene mesh vertices to determine whether the human mesh vertices and the scene mesh vertices are in contact can improve the accuracy of the contact relationship determination.
[0074] Step 306: Construct a dataset using human images, pose data, shape data, and contact relationships.
[0075] Step 306 and Figure 2 The example step 204 is largely the same, and will not be repeated here.
[0076] According to the second embodiment of this disclosure, during the construction of the dataset for model training, the contact relationship between the human body and the scene is confirmed, introducing scene awareness. This provides more realistic and accurate data for human pose estimation research in the field of computer vision, further improving the performance and robustness of the human pose estimation algorithm, enhancing the effect and experience of computer vision applications, and reducing unrealistic results caused by ignoring the scene and estimating human pose in isolation. Furthermore, determining the contact relationship based on the mesh data of the 3D human body model and the mesh data of the scene model yields contact relationships at the mesh accuracy level, improving the accuracy of the contact relationships.
[0077] After providing an exemplary description of the dataset construction process, the following provides an exemplary description of how to use the dataset.
[0078] Continue to refer to Figure 4 , Figure 4 This is a schematic flowchart of a model training method according to a third embodiment of this disclosure. Figure 4 As shown, the model training method may include the following steps:
[0079] Step 401: Obtain the dataset for training.
[0080] In this embodiment, the dataset can be constructed using the dataset construction method for 3D models exemplified in the above embodiments. This dataset may include human images, pose data of the target human body within the human images, shape data of the target human body, and the contact relationship between the target human body and the scene.
[0081] It should be understood that the execution subject of this embodiment may be the same as or different from the execution subject of the first embodiment and the second embodiment, and no limitation is made here.
[0082] Step 402: Use the dataset to train the model and obtain the human pose estimation model and / or contact estimation model.
[0083] In some embodiments of this disclosure, the executing entity can use a dataset to train a model to obtain a human pose estimation model. For example, the executing entity can use human images from the dataset as sample data for an initial human pose estimation model, and the pose data of the target human in the human images as sample labels, to train the initial human pose estimation model until it converges. The trained human pose estimation model can then be used to predict human pose.
[0084] It should be understood that, without departing from the teachings of this disclosure, the implementing entity may also use other data in the dataset as part of the training data during the training of the human pose estimation model. For example, the contact relationship between the target human body and the scene may also be used as part of the training data to adjust the estimated human pose. This disclosure does not limit this.
[0085] In some embodiments of this disclosure, the executing entity can use a dataset to train a model to obtain a contact estimation model. For example, the executing entity can use human images from the dataset as sample data for an initial human pose estimation model, and use the contact relationship between the target human body and the scene in the human images as sample labels to train the initial human pose estimation model until the initial contact estimation model converges. The trained contact estimation model can be used to predict the contact relationship between the human body and the scene.
[0086] According to the third embodiment of this disclosure, since the dataset in this embodiment fully considers the interaction between the human body and the scene, the dataset is richer and more accurate, thereby making the estimation results of the trained human pose estimation model and / or contact estimation model more accurate.
[0087] Continue to refer to Figure 5 , Figure 5 This is a schematic flowchart of a model training method according to the fourth embodiment of this disclosure. Figure 5 As shown, the model training method 500 for this contact estimation model may include:
[0088] Step 501: Obtain the dataset for training.
[0089] The dataset includes human images, pose data of the target human in the human images, shape data of the target human, and contact relationship between the target human and the scene.
[0090] Step 502: Determine the position features of the mesh vertices corresponding to the human body image based on the pose data and shape data corresponding to the human body image.
[0091] In this embodiment, the executing entity can input the pose data and shape data corresponding to the human body image into the SMPL model, and determine the mesh vertex position features corresponding to the human body image based on the output of the SMPL model.
[0092] Step 503: Input the human image into the feature extraction neural network model, and concatenate the features output by the feature extraction neural network model with the grid vertex position features to obtain the fused features.
[0093] In this embodiment, the feature extraction neural network model can be, for example, a convolutional neural network, used to extract high-dimensional image features from human images. The features extracted by the feature extraction neural network model can be concatenated with the network vertex position features to obtain the fused features of the human image.
[0094] Step 504: Use the fused features as sample data for the judgment neural network model, and use the contact relationships corresponding to the human body images as sample labels for the judgment neural network model, and train the judgment neural network model.
[0095] In this embodiment, a judgment neural network model is trained to learn to determine whether a human image is in contact with a scene based on the features of a single human image.
[0096] For ease of understanding, the following will use... Figure 6 The schematic diagram of the architecture of the example contact estimation model is provided as an example.
[0097] In this embodiment, considering that the contact area is often occluded, the contact estimation model needs to be able to explore the entire image to obtain evidence. Therefore, this embodiment provides a novel contact estimation model. Figure 6 As shown, the contact estimation model 600 may include a feature extraction neural network model 610 and a decision neural network model 620. The human image 630 is input into the feature extraction neural network model 610 to extract high-dimensional image features X from the human image 630. The feature extraction neural network model 610 may be, for example, a convolutional neural network. The contact estimation model 600 also includes an SMPL model 640. By inputting the pose data and shape data corresponding to the human image 630 into the SMPL model 640, the grid vertex position features corresponding to the target human body within the human image 630 can be obtained. The contact estimation model 600 can concatenate the image features X output by the feature extraction neural network model 610 with the grid vertex position features to obtain fused features q1, q2…q v The fused features are input into the judgment neural network model 620 to obtain the contact probabilities p1, p2...p of each human mesh vertex with the scene. vOptionally, the contact estimation model 600 may further include a binary transformation function 650, which may be, for example, a binary cross-entropy loss function. This binary transformation function 650 can be used to determine the contact vector C based on the contact probability between each human mesh vertex and the scene, and a preset probability threshold. The contact vector can use two values, 0 and 1, to represent no contact and contact, respectively. For example, if the contact probability between a human mesh vertex and the scene is less than the probability threshold, the element corresponding to that human mesh vertex in the contact vector C can be determined to be 0; if the contact probability is greater than or equal to the probability threshold, the element corresponding to that human mesh vertex in the contact vector can be determined to be 1. By training the contact estimation model, the executing agent can use the contact estimation model to obtain a contact vector representing whether each human mesh vertex is in contact with the scene in a single image. The executing agent can then determine the contact relationship between the human body and the scene based on this contact vector. For example, if all elements in the contact vector have a value of 0, it can be determined that the human body is not in contact with the scene. If there is an element with a value of 1 in the contact vector, it can be determined whether the body part corresponding to the human body mesh vertex of that element is in contact with the scene according to the preset judgment logic. The judgment logic can be referred to the illustrative description of the judgment logic mentioned in the second embodiment, and will not be repeated here.
[0098] In some optional embodiments of this disclosure, the determining neural network model may be a multi-layer Transformer model. For example, since the contact area is often occluded, the contact estimation model needs to be able to explore the entire image to obtain evidence. To achieve this, this embodiment introduces a Transformer model to learn non-local relationships, enabling the comprehensive consideration of global information in the image, thereby more accurately predicting human-scene contact.
[0099] According to the fourth embodiment of this disclosure, since the dataset in this embodiment fully considers the interaction between the human body and the scene, its dataset is richer and more accurate, thereby making the estimation results of the trained human pose estimation model and / or contact estimation model more accurate. Furthermore, by learning local relationships in the image through the transformer model, the contact relationship between the human body and the scene can be predicted more accurately.
[0100] See also Figure 7 , Figure 7 This is a schematic block diagram of a dataset construction apparatus according to a fifth embodiment of the present disclosure. Figure 7As shown, the dataset construction apparatus 700 may include a modeling module 710, an adding module 720, a determining module 730, and a construction module 740. The modeling module 710 is configured to create a 3D human model based on acquired human images. The adding module 720 is configured to determine the pose and shape data of the 3D human model and add it to the scene model. The determining module 730 is configured to determine the contact relationship between the 3D human model and the scene model based on the mesh data of the 3D human model and the mesh data of the scene model. The construction module 740 is configured to construct the dataset using the human images, pose data, shape data, and contact relationships.
[0101] In some optional embodiments of this disclosure, the determining module 730 includes: a human body mesh vertex determining submodule, a scene mesh determining submodule, and a relationship determining submodule. The human body mesh vertex determining submodule is configured to determine human body mesh vertices based on the mesh data of the 3D human body model. The scene mesh determining submodule is configured to determine scene mesh vertices based on the mesh data of the scene model. The relationship determining submodule is configured to determine the contact relationship between the 3D human body model and the scene model based on the distance between the human body mesh vertices and the scene mesh vertices.
[0102] In some optional embodiments of this disclosure, different body parts of the 3D human model are configured with distance thresholds for determining whether they are in contact with the scene. The relationship determination submodule includes: a distance threshold determination unit, a first mesh vertex relationship determination unit, and a first contact relationship determination unit. The distance threshold determination unit is configured to determine the distance threshold corresponding to a human mesh vertex based on the body part to which the human mesh vertex belongs. The first mesh vertex relationship determination unit is configured to determine that the human mesh vertex is in contact with the scene model in response to a distance less than the distance threshold between the human mesh vertex and any scene mesh vertex. The first contact relationship determination unit is configured to determine the contact relationship between the 3D human model and the scene model based on the contact situation between the human mesh vertex and the scene model.
[0103] In some optional embodiments of this disclosure, the body parts include a first body part exposed to the outside and a second body part covered by clothing, wherein the distance threshold corresponding to the second body part is greater than the distance threshold corresponding to the first body part.
[0104] In some optional embodiments of this disclosure, the relationship determination submodule includes a second mesh vertex relationship determination unit and a second contact relationship determination unit. The second mesh vertex relationship determination unit is configured to determine that there is contact between the human mesh vertex and the scene mesh vertex in response to the distance between the human mesh vertex and the scene mesh vertex being less than a distance threshold, and the angle between the normals of the human mesh vertex and the scene mesh vertex being less than a preset angle. The second contact relationship determination unit is configured to determine the contact relationship between the 3D human model and the scene model based on the contact situation between the human mesh vertex and the scene model.
[0105] In some optional embodiments of this disclosure, the determination module 730 includes a mesh determination submodule configured to input pose data and shape data into a parameterized human body model to obtain mesh data of the three-dimensional human body model.
[0106] In some optional embodiments of this disclosure, the modeling module 710 includes an acquisition submodule and a reconstruction submodule. The acquisition submodule is configured to acquire human images captured from multiple viewpoints. The reconstruction submodule is configured to perform multi-view reconstruction on the human images. Figure 3 3D reconstruction yields a three-dimensional human body model.
[0107] In some optional embodiments of this disclosure, the reconstruction submodule includes a recognition unit and a fitting unit. The recognition unit is configured to perform human body recognition on each frame of the human body images captured from various viewpoints, determining the positional information of the identified target human body in each human body image. The fitting unit is configured to, for the target human body, lock the image information corresponding to the target human body in the human body image based on the positional information; and, based on the image information, perform multi-view... Figure 3 3D reconstruction technology is used to obtain a three-dimensional human body model.
[0108] It is not difficult to see that this embodiment is a device embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0109] See also Figure 8 , Figure 8 This is a schematic block diagram of a model training apparatus according to the sixth embodiment of this disclosure. Figure 8 As shown, the model training device 800 may include a data acquisition module 810 and a model training module 820. The data acquisition module 810 is configured to acquire the dataset constructed in the above embodiments. The model training module 820 is configured to use the dataset to train the model, thereby obtaining a human pose estimation model and / or a contact estimation model.
[0110] In some optional embodiments of this disclosure, the contact estimation model includes a feature extraction neural network model and a judgment neural network model. The contact estimation model training module includes a position feature determination submodule, a stitching submodule, and a training submodule. The position feature determination submodule is configured to determine the position features of the grid vertices corresponding to the human image based on the pose data and shape data corresponding to the human image. The stitching submodule is configured to input the human image into the feature extraction neural network model and stitch the features output by the feature extraction neural network model with the grid vertex position features to obtain fused features. The training submodule is configured to use the fused features as sample data for the judgment neural network model and the contact relationships corresponding to the human image as sample labels for the judgment neural network model, thereby training the judgment neural network model.
[0111] In some optional embodiments of this disclosure, the neural network model is determined to be a multilayer transformation neural network model.
[0112] It is not difficult to see that this embodiment is a device embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0113] According to a seventh embodiment of this disclosure, an electronic device is provided, comprising: at least one processor and a memory. The memory is communicatively connected to the at least one processor and stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods mentioned in the above embodiments.
[0114] According to an eighth embodiment of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the methods mentioned in the above embodiments.
[0115] According to a ninth embodiment of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods mentioned in the above embodiments.
[0116] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0117] like Figure 9 As shown, device 900 includes a processor 901, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 902 or a computer program loaded from memory 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0118] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; memory 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0119] Processor 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 901 performs the various methods and processes described above, such as methods 200 / 300 / 400 / 500. For example, in some embodiments, methods 200 / 300 / 400 / 500 may be implemented as computer software programs tangibly contained in a machine-readable medium, such as memory 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by processor 901, one or more steps of methods 200 / 300 / 400 / 500 described above may be performed. Alternatively, in other embodiments, processor 901 may be configured to execute method 200 / 300 / 400 / 500 by any other suitable means (e.g., by means of firmware).
[0120] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0121] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0125] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0126] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0127] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for constructing a dataset for a 3D model, comprising: A three-dimensional human body model is created based on the acquired human body images; The pose and shape data of the three-dimensional human body model are determined, and the three-dimensional human body model is added to the scene model. The human body images taken from multiple perspectives are acquired and multi-view three-dimensional reconstruction is performed to establish the three-dimensional human body model. Through the multi-view three-dimensional reconstruction, the pose and shape data of the three-dimensional human body model can be obtained, as well as the contact relationship between the human body and the scene can be captured. Determining the contact relationship between the 3D human body model and the scene model based on the mesh data of the 3D human body model includes: determining human body mesh vertices based on the mesh data of the 3D human body model; determining scene mesh vertices based on the mesh data of the scene model; and determining the contact relationship between the 3D human body model and the scene model based on the distance between the human body mesh vertices and the scene mesh vertices, and a distance threshold corresponding to the body part to which the human body mesh vertex belongs, wherein different body parts of the 3D human body model are respectively configured with distance thresholds for determining whether they are in contact with the scene, and the body parts include a first body part exposed to the outside and a second body part covered by clothing, wherein the distance threshold corresponding to the second body part is greater than the distance threshold corresponding to the first body part; and A dataset is constructed using the human body image, the posture data, the shape data, and the contact relationships.
2. The method according to claim 1, wherein, Determining the contact relationship between the 3D human model and the scene model based on the distance between the vertices of the human body mesh and the vertices of the scene mesh includes: In response to the fact that the distance between the human mesh vertex and the scene mesh vertex is less than a distance threshold, and the angle between the directions of the normal of the human mesh vertex and the normal of the scene mesh vertex is less than a preset angle, it is determined that the human mesh vertex is in contact with the scene model; The contact relationship between the 3D human body model and the scene model is determined based on the contact between the vertices of the human body mesh and the scene model.
3. The method according to claim 1 or 2, further comprising: The posture data and shape data are input into the parameterized human body model to obtain the mesh data of the three-dimensional human body model.
4. The method according to claim 3, wherein, The step of building a three-dimensional human body model based on the acquired human body images includes: Acquire the human body images captured from multiple perspectives; The human body image is reconstructed using multi-view 3D reconstruction to obtain the 3D human body model.
5. The method according to claim 4, wherein, The process of performing multi-view 3D reconstruction on the human body image to obtain the 3D human body model includes: Human body recognition is performed on each frame of the human body images captured from various perspectives to determine the position information of the identified target human body in each of the human body images; and For the target human body, the image information corresponding to the target human body in the human body image is locked according to the location information; based on the image information, the three-dimensional human body model of the target human body is obtained through multi-view three-dimensional reconstruction technology.
6. A model training method, comprising: Obtain the dataset constructed by the method of any one of claims 1 to 5; as well as The dataset is used to train the model to obtain a human pose estimation model and / or a contact estimation model. The contact estimation model includes a feature extraction neural network model and a judgment neural network model. The step of training the model using the dataset to obtain the contact estimation model includes: The position features of the mesh vertices corresponding to the human body image are determined based on the posture data and shape data corresponding to the human body image. The human image is input into the feature extraction neural network model, and the features output by the feature extraction neural network model are concatenated with the grid vertex position features to obtain the fused features; The fused features are used as sample data for the judgment neural network model, and the contact relationships corresponding to the human images are used as sample labels for the judgment neural network model to train the judgment neural network model.
7. The method according to claim 6, wherein, The judgment neural network model is a multi-layer transformation neural network model.
8. An apparatus for constructing a dataset for a three-dimensional model, comprising: The modeling module is configured to create a 3D human body model based on the acquired human body images; The addition module is configured to determine the pose data and shape data of the three-dimensional human body model and add the three-dimensional human body model to the scene model. The module acquires human body images taken from multiple perspectives, performs multi-view three-dimensional reconstruction, and establishes the three-dimensional human body model. Through the multi-view three-dimensional reconstruction, the pose data and shape data of the three-dimensional human body model can be obtained, as well as the contact relationship between the human body and the scene can be captured. A determination module, configured to determine the contact relationship between the 3D human body model and the scene model based on the mesh data of the 3D human body model and the mesh data of the scene model, includes: a human body mesh vertex determination submodule, configured to determine human body mesh vertices based on the mesh data of the 3D human body model; a scene mesh determination submodule, configured to determine scene mesh vertices based on the mesh data of the scene model; and a relationship determination submodule, configured to determine the contact relationship between the 3D human body model and the scene model based on the distance between the human body mesh vertices and the scene mesh vertices, and a distance threshold corresponding to the body part to which the human body mesh vertex belongs, wherein different body parts of the 3D human body model are respectively configured with distance thresholds for determining whether they are in contact with the scene, the body parts include a first body part exposed to the outside and a second body part covered by clothing, the distance threshold corresponding to the second body part is greater than the distance threshold corresponding to the first body part; and The building module is configured to build a dataset using the human body image, the pose data, the shape data, and the contact relationships.
9. The apparatus according to claim 8, wherein, The relationship determination submodule includes: The second mesh vertex relationship determination unit is configured to determine that the human mesh vertex and the scene mesh are in contact in response to the distance between the human mesh vertex and the scene mesh vertex being less than a distance threshold and the angle between the direction of the normal of the human mesh vertex and the normal of the scene mesh vertex being less than a preset angle. The second contact relationship determination unit is configured to determine the contact relationship between the three-dimensional human body model and the scene model based on the contact situation between the human body mesh vertices and the scene model.
10. The apparatus according to claim 8 or 9, wherein, The determining module includes: The mesh determination submodule is configured to input the pose data and the shape data into the parameterized human body model to obtain the mesh data of the three-dimensional human body model.
11. The apparatus according to claim 10, wherein, The modeling module includes: The acquisition submodule is configured to acquire the human body images captured from multiple perspectives; and The reconstruction submodule is configured to perform multi-view 3D reconstruction on the human body image to obtain the 3D human body model.
12. The apparatus according to claim 11, wherein, The reconstruction submodule includes: The recognition unit is configured to perform human body recognition on each frame of the human body images captured from various viewpoints, and determine the position information of the recognized target human body in each of the human body images; and The fitting unit is configured to, for the target human body, lock the image information corresponding to the target human body in the human body image according to the position information; and, based on the image information, obtain the three-dimensional human body model of the target human body through multi-view three-dimensional reconstruction technology.
13. A model training device, comprising: The data acquisition module is configured to acquire the dataset constructed by the method of any one of claims 1 to 5; as well as The model training module is configured to train a model using the dataset to obtain a human pose estimation model and / or a contact estimation model. The contact estimation model includes a feature extraction neural network model and a judgment neural network model, and the contact estimation model training module includes: The position feature determination submodule is configured to determine the position features of the mesh vertices corresponding to the human body image based on the pose data and shape data corresponding to the human body image. The stitching submodule is configured to input the human image into the feature extraction neural network model, and stitch the features output by the feature extraction neural network model with the grid vertex position features to obtain fused features; The training submodule is configured to use the fused features as sample data for the judgment neural network model and the contact relationships corresponding to the human body images as sample labels for the judgment neural network model, and to train the judgment neural network model.
14. The apparatus according to claim 13, wherein, The judgment neural network model is a multi-layer transformation neural network model.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Panoramic three-dimensional scene understanding method based on graph neural network and relation optimization
CN114820932A
Method for capturing human motion in deformable scene based on neural deformation
CN115965765A