Three-dimensional model reconstruction method, device, equipment and computer-readable storage medium

By inputting the human body image into the preset network model, the target coordinates are obtained for three-dimensional model reconstruction, the problem of low accuracy of three-dimensional model reconstruction in the prior art is solved, and higher accuracy is achieved.

CN113327320BActive Publication Date: 2025-05-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110735973.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-05-06
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

The three-dimensional sitting standards obtained by the three-dimensional model reconstruction method in the prior art are not very accurate and cannot accurately represent the positional relationship of key points in the human body.

Method used

By obtaining the model reconstruction instructions sent by the terminal device, the human body image is input to the preset network model, the target coordinates of the human body key points in the three-dimensional space corresponding to the human body image are obtained, and the three-dimensional model is reconstructed based on the target coordinates.

Benefits of technology

The accuracy of three-dimensional model reconstruction is improved, and the three-dimensional coordinates of each key point corresponding to the human body image can be more accurately obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113327320B_ABST
    Figure CN113327320B_ABST
Patent Text Reader

Abstract

The present invention provides a three-dimensional model reconstruction method, device, equipment and computer-readable storage medium, the method comprising: obtaining a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed; according to the model reconstruction instruction, inputting the human body image into a preset network model to obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image; reconstructing the three-dimensional model corresponding to the human body image according to the target coordinates to obtain the target three-dimensional model; and sending the target three-dimensional model to the terminal device. Different from the solution of obtaining the three-dimensional coordinates corresponding to similar two-dimensional coordinates in a preset database in the prior art, the deep features of the human body image can be analyzed through the preset network model, and the three-dimensional coordinates of each key point of the human body corresponding to the human body image can be accurately obtained, thereby improving the accuracy of the three-dimensional model reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a three-dimensional model reconstruction method, device, equipment and computer-readable storage medium. Background Art

[0002] In many scenarios, it is necessary to convert two-dimensional images or scenes into three-dimensional images or scenes. For example, after obtaining character images in two-dimensional space, online games also need to convert task images into three-dimensional space. Alternatively, in many scenarios where emoticon packages are generated, after obtaining a human posture image in two-dimensional space captured by an image acquisition device set on a terminal device, the human posture image can be converted into three-dimensional space to obtain a three-dimensional emoticon package corresponding to the human posture image.

[0003] In order to realize the conversion from a two-dimensional image to a three-dimensional model, a preset three-dimensional database is generally established in the prior art, and the three-dimensional database specifically includes two-dimensional coordinates and three-dimensional coordinates corresponding to the two-dimensional coordinates. After obtaining the human body coordinates in the two-dimensional space, the three-dimensional coordinates corresponding to the two-dimensional coordinates closest to the human body coordinates are obtained in the three-dimensional database according to the human body coordinates. The three-dimensional model is reconstructed according to the three-dimensional coordinates.

[0004] However, when the above method is used to reconstruct a three-dimensional model, the three-dimensional coordinates obtained are all corresponding to the preset two-dimensional coordinates, and the accuracy of obtaining the corresponding three-dimensional coordinates based only on similarity is often not high. Summary of the invention

[0005] The present invention provides a three-dimensional model reconstruction method, device, equipment and computer-readable storage medium, which are used to solve the technical problem that the three-dimensional coordinates obtained by the existing model reconstruction method are not accurate enough.

[0006] A first aspect of the present invention is to provide a three-dimensional model reconstruction method, comprising:

[0007] Acquire a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed;

[0008] According to the model reconstruction instruction, the human body image is input into a preset network model to obtain target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image;

[0009] Reconstructing the three-dimensional model corresponding to the human body image according to the target coordinates to obtain a target three-dimensional model;

[0010] The target three-dimensional model is sent to the terminal device.

[0011] In a possible design, before inputting the human body image into a preset network model according to the model reconstruction instruction, the method further includes:

[0012] Acquire preset data to be trained, wherein the data to be trained includes a plurality of images to be trained, and the images to be trained include complete human body images;

[0013] The preset model to be trained is trained according to the data to be trained to obtain the network model.

[0014] In a possible design, training a preset model to be trained according to the data to be trained includes:

[0015] Obtaining at least one feature information corresponding to each image to be trained;

[0016] Determine a loss function according to the at least one feature information;

[0017] The model to be trained is trained according to the loss function until the model to be trained converges to obtain the network model.

[0018] In a possible design, the at least one feature information includes a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the image to be trained;

[0019] Accordingly, the step of obtaining at least one feature information corresponding to each image to be trained includes:

[0020] The two-dimensional heat map, bone depth feature map and hidden feature map corresponding to the image to be trained are obtained through the preset first sub-model.

[0021] In a possible design, determining the loss function according to the at least one feature information includes:

[0022] Determine a first loss value according to the two-dimensional heat map and the real labeled data corresponding to the training data;

[0023] Determine a second loss value according to the bone depth feature map and a real bone depth feature map generated by preset real data;

[0024] Connecting the two-dimensional heat map, the skeleton depth feature map and the hidden feature map to obtain a target feature map;

[0025] Inputting the target feature map into a second network model to obtain three-dimensional human body key points corresponding to the image to be trained;

[0026] Determine a third loss value according to the three-dimensional human body key points and preset real three-dimensional human body key points;

[0027] The loss function is determined according to the first loss value, the second loss value and the third loss value.

[0028] In a possible design, determining the loss function according to the first loss value, the second loss value, and the third loss value includes:

[0029] The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

[0030] In a possible design, training the model to be trained according to the loss function includes:

[0031] The model to be trained is trained according to the loss function through a gradient descent algorithm.

[0032] In one possible design, the method further includes:

[0033] Acquire a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed;

[0034] Inputting the human body posture image into the network model to obtain the three-dimensional coordinates corresponding to the human body posture image;

[0035] Determine the posture information corresponding to the human posture image according to the three-dimensional coordinates;

[0036] The posture information corresponding to the human posture image is sent to the terminal device.

[0037] A second aspect of the present invention is to provide a three-dimensional model reconstruction device, comprising:

[0038] An acquisition module, used for acquiring a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed;

[0039] A processing module, used for inputting the human body image into a preset network model according to the model reconstruction instruction, and obtaining target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image;

[0040] A reconstruction module, used to reconstruct the three-dimensional model corresponding to the human body image according to the target coordinates to obtain a target three-dimensional model;

[0041] The model sending module is used to send the target three-dimensional model to the terminal device.

[0042] In one possible design, the device further includes:

[0043] An acquisition module, used for acquiring preset data to be trained, wherein the data to be trained includes a plurality of images to be trained, and the images to be trained include complete human body images;

[0044] The training module is used to train the preset model to be trained according to the data to be trained to obtain the network model.

[0045] In one possible design, the training module includes:

[0046] A feature acquisition unit, used to acquire at least one feature information corresponding to each image to be trained;

[0047] A loss function determining unit, configured to determine a loss function according to the at least one feature information;

[0048] A training unit is used to train the model to be trained according to the loss function until the model to be trained converges to obtain the network model.

[0049] In a possible design, the at least one feature information includes a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the image to be trained;

[0050] Accordingly, the feature acquisition unit includes:

[0051] The two-dimensional heat map, the bone depth feature map and the hidden feature map corresponding to the image to be trained are obtained through the preset first sub-model.

[0052] In a possible design, the loss function determination unit is used to:

[0053] Determine a first loss value according to the two-dimensional heat map and the real labeled data corresponding to the training data;

[0054] Determine a second loss value according to the bone depth feature map and a real bone depth feature map generated by preset real data;

[0055] Connecting the two-dimensional heat map, the skeleton depth feature map and the hidden feature map to obtain a target feature map;

[0056] Inputting the target feature map into a second network model to obtain three-dimensional human body key points corresponding to the image to be trained;

[0057] Determine a third loss value according to the three-dimensional human body key points and preset real three-dimensional human body key points;

[0058] The loss function is determined according to the first loss value, the second loss value and the third loss value.

[0059] In a possible design, the loss function determination unit is used to:

[0060] The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

[0061] In one possible design, the training unit is used to:

[0062] The model to be trained is trained according to the loss function through a gradient descent algorithm.

[0063] In one possible design, the device further includes:

[0064] An instruction acquisition module, used to acquire a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed;

[0065] An input module, used to input the human posture image into the network model to obtain the three-dimensional coordinates corresponding to the human posture image;

[0066] A production module, used for determining posture information corresponding to the human body posture image according to the three-dimensional coordinates;

[0067] The sending module is used to send the posture information corresponding to the human posture image to the terminal device.

[0068] A third aspect of the present invention is to provide a three-dimensional model reconstruction device, comprising: a memory, a processor;

[0069] Memory; Memory for storing instructions executable by the processor;

[0070] The processor is configured to execute the three-dimensional model reconstruction method as described in the first aspect.

[0071] A fourth aspect of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the three-dimensional model reconstruction method as described in the first aspect.

[0072] The three-dimensional model reconstruction method, device, equipment and computer-readable storage medium provided by the present invention, after obtaining the model reconstruction instruction sent by the terminal device, input the human body image to be reconstructed into the preset network model according to the model reconstruction instruction, obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image, and reconstruct the three-body image according to the target coordinates to obtain the target three-dimensional model. Different from the solution of obtaining the three-dimensional coordinates corresponding to similar two-dimensional coordinates in a preset database in the prior art, the deep features of the human body image can be analyzed through the preset network model, and the three-dimensional coordinates of each key point of the human body corresponding to the human body image can be accurately obtained, thereby improving the accuracy of the three-dimensional model reconstruction. In addition, the target three-dimensional model can also be sent to the terminal device, so that the user can adjust the three-dimensional model according to actual needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention, and a person skilled in the art can also obtain other drawings based on these drawings.

[0074] Figure 1 A schematic diagram of the system architecture based on which the present invention is based;

[0075] Figure 2 A schematic diagram of a process flow of a three-dimensional model reconstruction method provided in Embodiment 1 of the present invention;

[0076] Figure 3 A schematic diagram of a display interface provided by an embodiment of the present invention;

[0077] Figure 4 Another system architecture diagram provided for an embodiment of the present invention;

[0078] Figure 5 A schematic diagram of a flow chart of a three-dimensional model reconstruction method provided in Embodiment 2 of the present invention;

[0079] Figure 6 A schematic diagram of a flow chart of a three-dimensional model reconstruction method provided in Embodiment 3 of the present invention;

[0080] Figure 7 An application scenario diagram for determining posture information provided by an embodiment of the present invention;

[0081] Figure 8 A schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 4 of the present invention;

[0082] Fig. 9 A schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 5 of the present invention;

[0083] Fig.10 A schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 6 of the present invention;

[0084] Fig.11 This is a schematic diagram of the process of a three-dimensional model reconstruction device provided in Embodiment 7 of the present invention. DETAILED DESCRIPTION

[0085] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0086] In response to the above-mentioned technical problem that the three-dimensional coordinates obtained by the existing model reconstruction method are not accurate enough, the present invention provides a three-dimensional model reconstruction method, device, equipment and computer-readable storage medium.

[0087] It should be noted that the three-dimensional model reconstruction method, device, equipment and computer-readable storage medium provided by the present application can be used in various 3D human body posture conversion scenarios.

[0088] In order to realize the reconstruction of the three-dimensional model, it is first necessary to determine the three-dimensional coordinates in the three-dimensional space. In the prior art, after obtaining the two-dimensional image to be reconstructed, the two-dimensional coordinates corresponding to the two-dimensional image are determined, and the target two-dimensional coordinates with a high similarity to the two-dimensional coordinates are obtained in the three-dimensional database. The three-dimensional model is reconstructed according to the three-dimensional coordinates corresponding to the target two-dimensional coordinates. However, the three-dimensional coordinates obtained in the prior art are generated according to the target two-dimensional coordinates, which cannot accurately represent the positional relationship of the key points of the human body in the two-dimensional image to be reconstructed, resulting in the three-dimensional model constructed according to the three-dimensional coordinates. The accuracy is not high.

[0089] In view of the problems in the prior art, the inventors discovered through research that in order to accurately construct a three-dimensional model corresponding to a two-dimensional image, it is necessary to extract the deep features of the two-dimensional image and determine the three-dimensional coordinates based on the deep features.

[0090] The inventors further discovered that the human body image to be reconstructed can be input into a preset network model according to the model reconstruction instruction sent by the terminal device to obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image, and the three-body image can be reconstructed into a three-dimensional model according to the target coordinates to obtain the target three-dimensional model. Different from the solution in the prior art of obtaining the three-dimensional coordinates corresponding to similar two-dimensional coordinates in a preset database, the preset network model can analyze the deep features of the human body image, accurately obtain the three-dimensional coordinates of each key point of the human body corresponding to the human body image, and improve the accuracy of the three-dimensional model reconstruction.

[0091] Figure 1 Schematic diagram of the system architecture based on the present invention, such as Figure 1 As shown, the system architecture based on the present invention at least includes: a terminal device 1 and a 3D model reconstruction device 2. The 3D model reconstruction device 2 is written in languages ​​such as C / C++, Java, Shell or Python; the terminal device 1 can be, for example, a desktop computer, a tablet computer, etc. The 3D model reconstruction device 2 is connected to the terminal device 1 in communication, so that information can be exchanged with the terminal device 1.

[0092] Figure 2 A schematic diagram of a process flow of a three-dimensional model reconstruction method provided in Embodiment 1 of the present invention is shown in FIG. Figure 2 As shown, the method includes:

[0093] Step 101: Obtain a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed.

[0094] The executor of this embodiment is a 3D model reconstruction device, which is connected to the terminal device for data exchange. It should be noted that the 3D model reconstruction device can be set in the terminal device or can be a device independent of the terminal device.

[0095] In this embodiment, the three-dimensional model reconstruction device can obtain a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction can include a human body image to be reconstructed. The human body image can specifically be a complete human body image or an image of a partial human body area. For example, in the process of game modeling, the human body image can be a complete human body image; in the process of making a user's facial expression package, the human body image can be a facial image of the user. Different human body images can be set according to different needs in actual applications, and the present invention does not limit this.

[0096] Specifically, the user can generate a three-dimensional model reconstruction instruction by triggering a three-dimensional model reconstruction icon set on the display interface of the terminal device. Figure 3A schematic diagram of a display interface provided by an embodiment of the present invention, such as Figure 3 As shown, the user can trigger the 3D model reconstruction icon to generate a corresponding 3D model reconstruction instruction. The user can trigger the 3D model reconstruction icon by single-clicking, double-clicking, long pressing, dragging, etc., and the present invention does not limit this.

[0097] Step 102: according to the model reconstruction instruction, the human body image is input into a preset network model to obtain target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image.

[0098] In this embodiment, after obtaining the three-dimensional model reconstruction instruction sent by the terminal device, the human body image can be input into the preset network model according to the three-dimensional model reconstruction instruction. The network model can specifically include a first sub-model and a second sub-model, the first sub-model is used to perform feature extraction operations on the human body posture image, and the second sub-model is used to determine the three-dimensional coordinates according to the feature information corresponding to the human body posture image. The network model determines the three-dimensional coordinates according to the feature information corresponding to the human body image, so that the three-dimensional coordinates are more closely matched with the human body image, and it can more accurately characterize the features of the human body image.

[0099] Step 103: reconstruct the three-dimensional model corresponding to the human body image according to the target coordinates to obtain a target three-dimensional model.

[0100] In this embodiment, after obtaining the target image corresponding to the human body image, the three-dimensional model corresponding to the human body image can be reconstructed according to the target image to obtain a target three-dimensional model. Since the target three-dimensional model is formed according to the three-dimensional coordinates output by the network model, it can more accurately represent the characteristics of the human body image.

[0101] Step 104: Send the target three-dimensional model to the terminal device.

[0102] In this embodiment, in order to enable the user to edit the target three-dimensional model, after obtaining the target three-dimensional model, the target three-dimensional model can be sent to the terminal device. Accordingly, after obtaining the target three-dimensional model, the terminal device can display the target three-dimensional model on the display interface. The user can edit the target three-dimensional model on the terminal device according to his own needs.

[0103] The three-dimensional model reconstruction method provided in this embodiment, after obtaining the model reconstruction instruction sent by the terminal device, inputs the human body image to be reconstructed into the preset network model according to the model reconstruction instruction, obtains the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image, and reconstructs the three-body image according to the target coordinates to obtain the target three-dimensional model. Different from the solution of obtaining the three-dimensional coordinates corresponding to similar two-dimensional coordinates in a preset database in the prior art, the deep features of the human body image can be analyzed through the preset network model, and the three-dimensional coordinates of each key point of the human body corresponding to the human body image can be accurately obtained, thereby improving the accuracy of the three-dimensional model reconstruction. In addition, the target three-dimensional model can also be sent to the terminal device, so that the user can adjust the three-dimensional model according to actual needs.

[0104] Further, based on the first embodiment, before obtaining the three-dimensional coordinates through the network model, it is first necessary to train the network model. Specifically, before step 102, it also includes:

[0105] Acquire preset data to be trained, wherein the data to be trained includes a plurality of images to be trained, and the images to be trained include complete human body images;

[0106] The preset model to be trained is trained according to the data to be trained to obtain the network model.

[0107] Figure 4 Another system architecture diagram provided for an embodiment of the present invention is as follows: Figure 4 As shown, the system architecture based on the present invention also includes a data server 3, and the 3D model reconstruction device 1 is respectively connected to the terminal device 1 and the data server 3. The data server 3 can be a cloud server, etc., which stores a large amount of data to be trained.

[0108] In this embodiment, in order to realize the training of the model to be trained, first, the 3D model reconstruction device can obtain preset data to be trained from the data server. In order to enable the network model to perform a 3D reconstruction operation on any human body region, the data to be trained includes multiple images to be trained, and each image to be trained includes a complete human body image. Then, the preset model to be trained can be trained with the data to be trained until the model to be trained converges to obtain a preset network model.

[0109] Specifically, the training data set can be randomly divided into a training set and a test set, the training set is used to train the training model, and the test set is used to test the training model. The training is iterated continuously until the training model converges.

[0110] Figure 5A flow chart of a three-dimensional model reconstruction method provided in the second embodiment of the present invention is provided. Based on any of the above embodiments, Figure 5 As shown, the training of the preset model to be trained according to the data to be trained specifically includes:

[0111] Step 201: Obtain at least one feature information corresponding to each image to be trained;

[0112] Step 202: determining a loss function according to the at least one feature information;

[0113] Step 203: Train the model to be trained according to the loss function until the model to be trained converges to obtain the network model.

[0114] In this embodiment, in order to realize the training of the model to be trained, it is also necessary to determine the loss function, and the model to be trained is trained by the loss function. Specifically, it is first necessary to perform feature extraction on each image to be trained to obtain at least one feature information, wherein the at least one feature information includes a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the image to be trained. Accordingly, the network model can specifically include a first sub-model and a second sub-model, and the first sub-model is used to perform feature extraction operations on human posture images, so the two-dimensional heat map, the bone depth feature map, and the hidden feature map corresponding to the image to be trained can be obtained by the preset first sub-model.

[0115] After obtaining at least one feature information corresponding to each image to be trained, a loss function can be determined according to the at least one feature information, and then the model can be trained according to the loss function.

[0116] Further, based on any of the above embodiments, step 202 specifically includes:

[0117] Determine a first loss value according to the two-dimensional heat map and the real labeled data corresponding to the training data;

[0118] Determine a second loss value according to the bone depth feature map and a real bone depth feature map generated by preset real data;

[0119] Connecting the two-dimensional heat map, the skeleton depth feature map and the hidden feature map to obtain a target feature map;

[0120] Inputting the target feature map into a second network model to obtain three-dimensional human body key points corresponding to the image to be trained;

[0121] Determine a third loss value according to the three-dimensional human body key points and preset real three-dimensional human body key points;

[0122] The loss function is determined according to the first loss value, the second loss value and the third loss value.

[0123] In this embodiment, first, L1 loss calculation can be performed on the two-dimensional heat map and the real annotated data corresponding to the training data to obtain a first loss value I1. Further, L1 loss calculation can be performed on the bone depth feature map and the real bone depth feature map generated by the preset real data to obtain a second loss value I2. It should be noted that the bone depth feature map is obtained by interpolating or obtaining the depth of the two key points of each bone in each image to be trained.

[0124] Furthermore, the two-dimensional heat map, the bone depth feature map and the hidden feature map can be connected to obtain a target feature map. The target feature map can then be input into the second network model to obtain the three-dimensional human key points corresponding to the image to be trained. The human key points and the preset real three-dimensional human key points are subjected to L1 loss calculation to obtain a third loss value l3. After obtaining the first loss value, the second loss value and the third loss value, the loss function can be determined according to the first loss value, the second loss value and the third loss value.

[0125] Specifically, determining the loss function according to the first loss value, the second loss value, and the third loss value includes:

[0126] The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

[0127] In this embodiment, each loss value corresponds to a balance coefficient, which can be obtained based on historical experience or set by the user according to actual needs, and the present invention does not limit this. Therefore, the loss function LossL can be determined based on the first loss value, the second loss value, the third loss value, and the balance coefficient corresponding to each loss value. The loss function can be specifically shown in Formula 1:

[0128] Loss L=a1*I1+a2*l2+a3*l3 (1)

[0129] Among them, I1 is the first loss value, a1 is the balance coefficient corresponding to I1, I2 is the second loss value, a2 is the balance coefficient corresponding to I2, I3 is the third loss value, a3 is the balance coefficient corresponding to I3.

[0130] Furthermore, the model to be trained can be trained according to the loss function through a gradient descent algorithm.

[0131] It should be noted that other training methods may also be used to perform training operations on the training model, and the present invention does not limit this.

[0132] The three-dimensional model reconstruction method provided in this embodiment can obtain a preset network model by using a preset data set to be trained to train the model to be trained, thereby providing a basis for subsequent three-dimensional model reconstruction. In addition, since in the model training process, the features of the image to be trained are first extracted, and the loss function is determined based on the feature information, and the model to be trained is trained based on the loss function, the subsequent network model can determine the three-dimensional coordinates according to the feature information of the human body image during the process of determining the three-dimensional coordinates, thereby improving the accuracy and restoration of the three-dimensional coordinates, and thus improving the accuracy of the model.

[0133] Figure 6 The flowchart of the three-dimensional model reconstruction method provided in the third embodiment of the present invention is based on any of the above embodiments. Figure 6 As shown, the method also includes:

[0134] Step 301: obtaining a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed;

[0135] Step 302: input the human posture image into the network model to obtain the three-dimensional coordinates corresponding to the human posture image;

[0136] Step 303, determining posture information corresponding to the human body posture image according to the three-dimensional coordinates;

[0137] Step 304: Send the posture information corresponding to the human posture image to the terminal device.

[0138] In this embodiment, the determination of the posture information of the human body image can also be achieved through the three-dimensional coordinates. Specifically, a human body posture determination instruction sent by a terminal device can be obtained, wherein the human body posture determination instruction includes a human body posture image to be processed. According to the human body posture determination instruction, the human body posture image is input into the network model to obtain the three-dimensional coordinates corresponding to the human body posture image. Then, the posture information corresponding to the human body posture image can be determined according to the three-dimensional coordinates corresponding to the human body posture image. In order to enable the user to clearly determine the recognition result of the human body posture image, the posture information corresponding to the human body posture image can be sent to the terminal device.

[0139] Figure 7 The posture information provided in the embodiment of the present invention determines the application scenario diagram, such as Figure 7As shown, after acquiring the human posture image, the human posture image can first be input into a preset network model. The network model can specifically include a first sub-model and a second sub-model, the first sub-model is used to perform feature extraction operations on the human posture image, and the second sub-model is used to determine the three-dimensional coordinates according to the feature information corresponding to the human posture image. After the human posture image is input into the network model, the three-dimensional coordinates corresponding to the human posture image can be obtained, and then the reconstruction of the three-dimensional model and the determination of the posture information can be realized according to the three-dimensional coordinates. Figure 7 As shown, the posture information corresponding to the human posture image may be half squatting.

[0140] The three-dimensional model reconstruction method provided in this embodiment obtains the human posture determination instruction sent by the terminal device, inputs the human posture image into the network model, obtains the three-dimensional coordinates corresponding to the human posture image, and then can accurately determine the posture information corresponding to the human posture image, thereby improving the accuracy of posture determination.

[0141] Figure 8 This is a schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 4 of the present invention. Figure 8 As shown, the device includes: an acquisition module 41, a processing module 42, a reconstruction module 43 and a model sending module 44, wherein the acquisition module 41 is used to acquire a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed; the processing module 42 is used to input the human body image into a preset network model according to the model reconstruction instruction, and obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image; the reconstruction module 43 is used to reconstruct the three-dimensional model corresponding to the human body image according to the target coordinates to obtain a target three-dimensional model; the model sending module 44 is used to send the target three-dimensional model to the terminal device.

[0142] The three-dimensional model reconstruction device provided in this embodiment, after obtaining the model reconstruction instruction sent by the terminal device, inputs the human body image to be reconstructed into the preset network model according to the model reconstruction instruction, obtains the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image, and reconstructs the three-body image according to the target coordinates to obtain the target three-dimensional model. Different from the solution of obtaining the three-dimensional coordinates corresponding to similar two-dimensional coordinates in a preset database in the prior art, the deep features of the human body image can be analyzed through the preset network model, and the three-dimensional coordinates of each key point of the human body corresponding to the human body image can be accurately obtained, thereby improving the accuracy of the three-dimensional model reconstruction. In addition, the target three-dimensional model can also be sent to the terminal device, so that the user can adjust the three-dimensional model according to actual needs.

[0143] Furthermore, based on the fourth embodiment, the device also includes: an acquisition module and a training module, wherein the acquisition module is used to acquire preset data to be trained, the data to be trained includes multiple images to be trained, and the images to be trained include complete human body images; the training module is used to train a preset model to be trained according to the data to be trained to obtain the network model.

[0144] Fig. 9 This is a schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 5 of the present invention. Based on any of the above embodiments, Fig. 9 As shown, the training module includes: a feature acquisition unit 51, a loss function determination unit 52 and a training unit 53, wherein the feature acquisition unit 51 is used to obtain at least one feature information corresponding to each image to be trained; the loss function determination unit 52 is used to determine the loss function according to the at least one feature information; the training unit 53 is used to train the model to be trained according to the loss function until the model to be trained converges to obtain the network model.

[0145] Further, based on any of the above embodiments, the at least one feature information includes a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the image to be trained;

[0146] Accordingly, the feature acquisition unit includes:

[0147] The two-dimensional heat map, the bone depth feature map and the hidden feature map corresponding to the image to be trained are obtained through the preset first sub-model.

[0148] Further, based on any of the above embodiments, the loss function determination unit is used to:

[0149] Determine a first loss value according to the two-dimensional heat map and the real labeled data corresponding to the training data;

[0150] Determine a second loss value according to the bone depth feature map and a real bone depth feature map generated by preset real data;

[0151] Connecting the two-dimensional heat map, the skeleton depth feature map and the hidden feature map to obtain a target feature map;

[0152] Inputting the target feature map into a second network model to obtain three-dimensional human body key points corresponding to the image to be trained;

[0153] Determine a third loss value according to the three-dimensional human body key points and preset real three-dimensional human body key points;

[0154] The loss function is determined according to the first loss value, the second loss value and the third loss value.

[0155] Further, based on any of the above embodiments, the loss function determination unit is used to:

[0156] The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

[0157] Further, based on any of the above embodiments, the training unit is used to:

[0158] The model to be trained is trained according to the loss function through a gradient descent algorithm.

[0159] Fig.10 This is a schematic diagram of the structure of a three-dimensional model reconstruction device provided in Embodiment 6 of the present invention, based on any of the above embodiments, such as Fig.10 As shown, the device also includes: an instruction acquisition module 61, an input module 62, a production module 63 and a sending module 64, wherein the instruction acquisition module 61 is used to obtain a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed; the input module 62 is used to input the human body posture image into the network model to obtain the three-dimensional coordinates corresponding to the human body posture image; the production module 63 is used to determine the posture information corresponding to the human body posture image according to the three-dimensional coordinates; and the sending module 64 is used to send the posture information corresponding to the human body posture image to the terminal device.

[0160] Fig.11 A schematic diagram of the process flow of a three-dimensional model reconstruction device provided in Embodiment 7 of the present invention, such as Fig.11 As shown, the three-dimensional model reconstruction device includes: a memory 71, a processor 72;

[0161] Memory 71; Memory 71 for storing instructions executable by the processor 72;

[0162] The processor 72 is configured to execute the three-dimensional model reconstruction method as described in any of the above embodiments.

[0163] The memory 71 is used to store programs. Specifically, the program may include program codes, and the program codes include computer operation instructions. The memory 71 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0164] The processor 72 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0165] Optionally, in a specific implementation, if the memory 71 and the processor 72 are implemented independently, the memory 71 and the processor 72 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0166] Optionally, in a specific implementation, if the memory 71 and the processor 72 are integrated on a chip, the memory 71 and the processor 72 can communicate with each other through an internal interface.

[0167] Yet another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the three-dimensional model reconstruction method as described in any of the above embodiments.

[0168] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0169] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional model reconstruction method, characterized in that: include: Acquire a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed; According to the model reconstruction instruction, the human body image is input into a preset network model to obtain target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image; the three-dimensional model corresponding to the human body image is reconstructed according to the target coordinates to obtain a target three-dimensional model; Sending the target three-dimensional model to the terminal device; The step of inputting the human body image into a preset network model to obtain target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image includes: The preset network model includes a first sub-model and a second sub-model; The first sub-model is used to perform feature extraction operations on human posture images; the features output by the first sub-model include: a two-dimensional heat map corresponding to the human image, a bone depth feature map, and a hidden feature map; the bone depth feature map is obtained by interpolating the depths of two key points of each bone in each human image; Connecting the two-dimensional heat map, the bone depth feature map and the hidden feature map to obtain a target feature map; inputting the target feature map into the second sub-model to obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image; Before inputting the human body image into a preset network model according to the model reconstruction instruction, the method further includes: Acquire preset data to be trained, wherein the data to be trained includes a plurality of images to be trained, and the images to be trained include complete human body images; Acquire at least one feature information corresponding to each image to be trained through the first sub-model, wherein the feature information includes a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the human body image; Connecting the two-dimensional heat map, the bone depth feature map and the hidden feature map to obtain a target feature map, inputting the target feature map into the second sub-model to obtain three-dimensional human body key points corresponding to the image to be trained; Determine a first loss value according to the two-dimensional heat map and the real annotated data corresponding to the training data; determine a second loss value according to the bone depth feature map and the real bone depth feature map generated by the preset real data; determine a third loss value according to the three-dimensional human key points and the preset real three-dimensional human key points; Determine a loss function according to the first loss value, the second loss value and the third loss value; The model to be trained is trained according to the loss function until the model to be trained converges, thereby obtaining the network model.

2. The method according to claim 1, characterized in that The determining the loss function according to the first loss value, the second loss value and the third loss value comprises: The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

3. The method according to claim 1 or 2, characterized in that: The training of the model to be trained according to the loss function includes: The model to be trained is trained according to the loss function through a gradient descent algorithm.

4. The method according to claim 1 or 2, characterized in that: The method further comprises: Acquire a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed; Inputting the human body posture image into the network model to obtain the three-dimensional coordinates corresponding to the human body posture image; Determine the posture information corresponding to the human posture image according to the three-dimensional coordinates; The posture information corresponding to the human posture image is sent to the terminal device.

5. A three-dimensional model reconstruction device, characterized in that: include: An acquisition module, used for acquiring a model reconstruction instruction sent by a terminal device, wherein the model reconstruction instruction includes a human body image to be reconstructed; A processing module, used for inputting the human body image into a preset network model according to the model reconstruction instruction, and obtaining target coordinates of key points of the human body in a three-dimensional space corresponding to the human body image; A reconstruction module, used to reconstruct the three-dimensional model corresponding to the human body image according to the target coordinates to obtain a target three-dimensional model; A model sending module, used for sending the target three-dimensional model to the terminal device; The processing module, specifically used for the preset network model, includes a first sub-model and a second sub-model; the first sub-model is used to perform feature extraction operations on human posture images; the features output by the first sub-model include: a two-dimensional heat map, a bone depth feature map, and a hidden feature map corresponding to the human body image; the bone depth feature map is obtained by interpolating the depths of two key points of each bone in each human body image; the two-dimensional heat map, the bone depth feature map, and the hidden feature map are connected to obtain a target feature map; the target feature map is input into the second sub-model to obtain the target coordinates of the key points of the human body in the three-dimensional space corresponding to the human body image; The device is also used for: an acquisition module, used for acquiring preset data to be trained, wherein the data to be trained includes multiple images to be trained, and the images to be trained include complete human images; a feature acquisition module, used for acquiring at least one feature information corresponding to each image to be trained through the first sub-model, wherein the feature information includes a two-dimensional heat map, a bone depth feature map and a hidden feature map corresponding to the human image; a key point recognition module, used for connecting the two-dimensional heat map, the bone depth feature map and the hidden feature map to obtain a target feature map, and inputting the target feature map into the second sub-model to obtain three-dimensional human key points corresponding to the image to be trained; a processing module, used for determining a first loss value according to the two-dimensional heat map and the real annotated data corresponding to the data to be trained; determining a second loss value according to the bone depth feature map and the real bone depth feature map generated by the preset real data; determining a third loss value according to the three-dimensional human key points and the preset real three-dimensional human key points; a loss value determination module, used for determining a loss function according to the first loss value, the second loss value and the third loss value; a training module, used for training the model to be trained according to the loss function until the model to be trained converges to obtain the network model.

6. The device according to claim 5, characterized in that The loss function determination unit is used for: The loss function is determined according to the first loss value, the second loss value, the third loss value and the balance coefficient corresponding to each loss value.

7. The device according to claim 5 or 6, characterized in that The training module is used to: The model to be trained is trained according to the loss function through a gradient descent algorithm.

8. The device according to any one of claims 5 or 6, characterized in that: The device also includes: An instruction acquisition module, used to acquire a human body posture determination instruction sent by a terminal device, wherein the human body posture determination instruction includes a human body posture image to be processed; An input module, used to input the human posture image into the network model to obtain the three-dimensional coordinates corresponding to the human posture image; A production module, used for determining posture information corresponding to the human body posture image according to the three-dimensional coordinates; The sending module is used to send the posture information corresponding to the human posture image to the terminal device.

9. A three-dimensional model reconstruction device, characterized in that: include: Memory, processor; Memory; a memory for storing instructions executable by the processor; The processor is configured to execute the three-dimensional model reconstruction method as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the three-dimensional model reconstruction method according to any one of claims 1 to 4.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the three-dimensional model reconstruction method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Human body key point extraction method and device, readable storage medium and equipment

    CN110532981A

  • Posture detection and video processing method and device, electronic equipment and storage medium

    CN111666917A