Method for constructing three-dimensional grid generation model, three-dimensional grid generation method and device

By acquiring the image and depth features of multiple perspectives, the three-dimensional grid generation model is used to reconstruct the three-dimensional grid of the target object, solving the accuracy problem caused by low voxel resolution and achieving higher accuracy of three-dimensional reconstruction.

CN115147527BActive Publication Date: 2025-06-06ALIBABA INNOVATION PRIVATE LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110351124.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-06-06
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

When the voxel resolution of an object is low, the accuracy of reconstructing the object's three-dimensional grid based on the voxel is low, which cannot meet the actual needs.

Method used

By obtaining target images from multiple different perspectives of the target object, the target voxels and target depth features of the target object are obtained, and a three-dimensional grid generation model is used to generate the target three-dimensional grid of the target object based on these features.

Benefits of technology

The accuracy of the generated three-dimensional grid can be improved, the three-dimensional structure of the object can be reconstructed more accurately, the accuracy problem caused by low voxel resolution is solved, and the incomplete image information caused by occlusion can be handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147527B_ABST
    Figure CN115147527B_ABST
Patent Text Reader

Abstract

The present application provides a method for constructing a three-dimensional mesh generation model, a three-dimensional mesh generation method and a device. Target images of a target object at multiple different perspectives are obtained, target voxels of the target object are obtained according to the target images, target depth features of the target images of each perspective are respectively obtained, and then a target three-dimensional mesh of the target object is obtained according to each target depth feature. The target depth features of the target images of each perspective of the target object can be directly obtained according to the target images of each perspective of the target object. Therefore, the target depth features of the target images of each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the target voxels of the target object are directly or indirectly optimized (such as fine-tuning or refining, etc.) using the target depth features of the target images of each perspective, which can improve the accuracy of the target three-dimensional mesh of the acquired target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for constructing a three-dimensional grid generation model, a device for constructing a three-dimensional grid generation model, a three-dimensional grid generation method, and a three-dimensional grid generation device. Background Art

[0002] Three-dimensional reconstruction is currently widely used in various industries. Its goal is to reconstruct objects in three dimensions through captured images of the objects.

[0003] The main current 3D reconstruction methods include: collecting voxels of an object in an image, and then reconstructing a 3D mesh of the object based on the voxels of the object.

[0004] However, the inventors have found that when the resolution of the voxels of the object is low, the accuracy of reconstructing the three-dimensional grid of the object based on the voxels of the object is low, resulting in that the three-dimensional grid cannot meet actual needs. Summary of the invention

[0005] In order to improve the accuracy of the generated three-dimensional mesh, the present application shows a method for constructing a three-dimensional mesh generation model, a three-dimensional mesh generation method and a device.

[0006] In a first aspect, the present application provides a three-dimensional mesh generation method, the method comprising:

[0007] Acquire target images of a target object from multiple different perspectives;

[0008] Acquire target voxels of the target object according to target images of the target object at multiple different perspectives, and respectively acquire target depth features of the target images of the target object at each perspective;

[0009] A target three-dimensional grid of the target object is acquired according to the target voxels of the target object and the target depth features of the target images of each viewing angle of the target object.

[0010] In a second aspect, the present application shows a method for constructing a three-dimensional mesh generation model, the method comprising:

[0011] Acquire at least one sample data set, the sample data set comprising: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object;

[0012] Construct the network structure of the 3D mesh generation model;

[0013] Using the sample data set to train the network parameters in the three-dimensional mesh generation model until the network parameters converge to obtain the three-dimensional mesh generation model;

[0014] Wherein, the network structure at least includes a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network;

[0015] The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different viewing angles;

[0016] The depth feature extraction network is used to respectively obtain sample depth features of sample images of each viewing angle of the sample object;

[0017] The three-dimensional mesh generation network is used to obtain a predicted three-dimensional mesh of the sample object according to sample voxels of the sample object and sample depth features of sample images of each viewing angle of the sample object.

[0018] In a third aspect, the present application provides a three-dimensional grid generation device, the device comprising:

[0019] A first acquisition module is used to acquire target images of a target object at multiple different viewing angles;

[0020] a second acquisition module, configured to acquire target voxels of the target object according to target images of multiple different perspectives of the target object, and a third acquisition module, configured to respectively acquire target depth features of the target images of each perspective of the target object;

[0021] The fourth acquisition module is used to acquire a target three-dimensional grid of the target object according to the target voxels of the target object and the target depth features of the target images of each viewing angle of the target object.

[0022] In a fourth aspect, the present application provides a device for constructing a three-dimensional mesh generation model, the device comprising:

[0023] A fifth acquisition module, configured to acquire at least one sample data set, wherein the sample data set includes: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object;

[0024] A construction module is used to construct a network structure of a three-dimensional mesh generation model;

[0025] A training module, used to train the network parameters in the three-dimensional mesh generation model using the sample data set until the network parameters converge to obtain the three-dimensional mesh generation model;

[0026] Wherein, the network structure at least includes a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network;

[0027] The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different viewing angles;

[0028] The depth feature extraction network is used to respectively obtain sample depth features of sample images of each viewing angle of the sample object;

[0029] The three-dimensional mesh generation network is used to obtain a predicted three-dimensional mesh of the sample object according to sample voxels of the sample object and sample depth features of sample images of each viewing angle of the sample object.

[0030] In a fifth aspect, the present application shows an electronic device, the electronic device comprising:

[0031] processor;

[0032] a memory for storing processor-executable instructions;

[0033] Wherein, the processor is configured to execute the three-dimensional grid generation method as described in the first aspect.

[0034] In a sixth aspect, the present application illustrates a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the three-dimensional mesh generation method as described in the first aspect.

[0035] In a seventh aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the three-dimensional mesh generation method as described in the first aspect.

[0036] In an eighth aspect, the present application provides an electronic device, the electronic device comprising:

[0037] processor;

[0038] a memory for storing processor-executable instructions;

[0039] Wherein, the processor is configured to execute the method for constructing a three-dimensional mesh generation model as described in the second aspect.

[0040] In a ninth aspect, the present application illustrates a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method for constructing a three-dimensional mesh generation model as described in the second aspect.

[0041] In a tenth aspect, the present application illustrates a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to execute the method for constructing a three-dimensional mesh generation model as described in the second aspect.

[0042] Compared with the prior art, the embodiments of the present application have the following advantages:

[0043] In the present application, target images of the target object at multiple different perspectives are acquired, target voxels of the target object are acquired based on the target images of the target object at multiple different perspectives, and target depth features of the target images of the target object at each perspective are acquired respectively. Then, a target three-dimensional grid of the target object is acquired based on the target voxels of the target object and the target depth features of the target images of the target object at each perspective.

[0044] Through the present application, on the one hand, the target depth features of the target images of the target object at each perspective can be directly acquired based on the target images of the target object at each perspective, and the target depth features of the target images of the target object at each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the target voxels of the target object can be directly or indirectly optimized (for example, fine-tuned or refined) using the target depth features of the target images of the target object at each perspective, thereby improving the accuracy of the acquired target three-dimensional mesh of the target object.

[0045] On the other hand, if the target object in some target images is occluded (at least a part of the target object is occluded, resulting in incomplete image information), in this case, since the present application obtains the target voxels of the target object based on target images of multiple different perspectives of the target object, a large number of image features of the target object can be learned from the target images of multiple different perspectives, thereby completing the image information and solving the problem of incomplete image information caused by occlusion, thereby avoiding the accuracy of the target three-dimensional mesh of the acquired target object being affected by the occlusion problem.

[0046] On the other hand, when obtaining the target three-dimensional mesh of the target object, the three-dimensional mesh to be optimized can be generated according to the target voxels of the target object, and then the three-dimensional mesh to be optimized can be optimized according to the target depth features of the target images of each perspective of the target object to obtain the target three-dimensional mesh.

[0047] It can be seen that the optimization process is an optimization of the three-dimensional grid. The amount of data representing the three-dimensional grid is relatively low and is often lower than the amount of data representing the voxels. Therefore, compared with the method of directly processing the voxels, the method of "first generating the three-dimensional grid to be optimized for the target object based on the target voxels of the target object, and then directly optimizing the three-dimensional grid" in the present application can reduce the amount of computational data involved in the optimization process, thereby improving the optimization efficiency and saving computing resources.

[0048] On the other hand, when optimizing the three-dimensional mesh, a graph convolutional neural network can be used to optimize the three-dimensional mesh. With the help of the powerful image processing capabilities of the graph convolutional neural network, the degree of optimization can be improved, and the accuracy of the target three-dimensional mesh of the target object can be further improved.

[0049] On the other hand, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh to be optimized can be optimized for multiple rounds in sequence according to the target depth features of the target images of each perspective of the target object. Each round of optimization can further optimize the three-dimensional mesh obtained in the previous round of optimization, thereby achieving multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, and further improving the accuracy of the target three-dimensional mesh of the target object.

[0050] On the other hand, since the present application can optimize the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object to improve the accuracy of the three-dimensional mesh, it can support the generation of a lower-accuracy three-dimensional mesh to be optimized when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". Since it supports the generation of a lower-accuracy three-dimensional mesh to be optimized, it can support the use of low-resolution voxels when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". The data volume of the low-resolution voxels is low, thereby reducing the amount of computational data involved in "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object", thereby improving the efficiency of generating the three-dimensional mesh to be optimized for the target object and saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flowchart of the steps of a method for constructing a three-dimensional mesh generation model shown in this application.

[0052] Figure 2 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0053] Figure 3 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0054] Figure 4 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0055] Figure 5 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0056] Figure 6 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0057] Figure 7 It is a schematic diagram of a network structure of a three-dimensional grid generation model shown in this application.

[0058] Figure 8 It is a flowchart of the steps of a three-dimensional grid generation method shown in this application.

[0059] Fig. 9 It is a structural block diagram of a device for constructing a three-dimensional grid generation model shown in this application.

[0060] Fig.10 It is a structural block diagram of a three-dimensional grid generation device shown in this application.

[0061] Fig.11 It is a schematic diagram of the structure of the device shown in this application. DETAILED DESCRIPTION

[0062] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0063] In order to improve the accuracy of the generated 3D mesh, refer to Figure 1 , showing a flow chart of the steps of a method for constructing a three-dimensional mesh generation model in an embodiment of the present invention. The method can be used to construct a three-dimensional mesh generation model, and then the three-dimensional mesh of the object can be generated based on the three-dimensional mesh generation model according to images of the object from multiple different perspectives to improve the accuracy of the generated three-dimensional mesh.

[0064] The method is applied to an electronic device, and the method may specifically include the following steps:

[0065] In step S101, at least one sample data set is acquired, where the sample data set includes: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object.

[0066] The sample objects include three-dimensional objects, such as people, sofas, vehicles, trees, and buildings.

[0067] There may be multiple sample data sets, and different sample data sets may include sample images of different sample objects. In the same sample data set, the perspectives of the sample objects in each sample image are different.

[0068] The viewing angle can be understood as the direction or angle of viewing / collecting the sample object. For example, taking a vehicle as an example, viewing the vehicle from the front of the vehicle is one viewing angle, viewing the vehicle from the left side of the vehicle is one viewing angle, viewing the vehicle from the right side of the vehicle is one viewing angle, viewing the vehicle from the left front side of the vehicle is one viewing angle, viewing the vehicle from the right front side of the vehicle is one viewing angle, viewing the vehicle from the left rear side of the vehicle is one viewing angle, viewing the vehicle from the right rear side of the vehicle is one viewing angle, and viewing the vehicle from the rear of the vehicle is one viewing angle, etc.

[0069] The sample image may include an RGB image, etc. The sample image may also include camera parameters corresponding to the sample object. When acquiring a sample image of the sample object, a camera may be used to acquire the sample image of the sample object. The camera may obtain the relative position relationship between the camera and the sample object as the camera parameters. After the camera acquires the sample image of the sample object, the camera parameters may be added to the sample image.

[0070] The three-dimensional mesh in the present application may include a triangular mesh, a quadrilateral mesh or a pentagonal mesh, etc. The surface of a three-dimensional object can be simulated by the three-dimensional mesh.

[0071] In step S102, a network structure of a three-dimensional mesh generation model is constructed.

[0072] See also Figure 2 In one embodiment of the present application, the network structure of the three-dimensional mesh generation model includes at least a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network.

[0073] The deep feature extraction network may include MVSNet (Multiple View Stereo Net), CascadeMVSNet (Cascade Multiple View Stereo Net) or PointMVSNet (Point Multiple View Stereo Net), etc. Of course, it may also include other networks for extracting deep features, which is not limited in this application.

[0074] The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different perspectives.

[0075] The deep feature extraction network is used to obtain sample deep features of sample images of each viewing angle of the sample object.

[0076] In one embodiment, when acquiring sample depth features of sample images of each viewing angle of a sample object, reference may be made to camera parameters corresponding to the sample object in each sample image to improve the accuracy of the acquired sample depth features of the sample images of each viewing angle of the sample object.

[0077] The three-dimensional mesh generation network is used to obtain a predicted three-dimensional mesh of the sample object according to sample voxels of the sample object and sample depth features of sample images of each viewing angle of the sample object.

[0078] Among them, the input end of the three-dimensional mesh generation model includes the input end of the voxel generation network and the input end of the deep feature extraction network.

[0079] The output of the voxel generation network is connected to the input of the 3D mesh generation network.

[0080] The output of the deep feature extraction network is connected to the input of the 3D mesh generation network.

[0081] The output end of the three-dimensional mesh generation model includes an output end of the three-dimensional mesh generation network.

[0082] In the present application, the network structure of the three-dimensional mesh generation model can be constructed based on demand, and the constructed three-dimensional mesh generation model may be applied to different application scenarios later. Different application scenarios are suitable for different network structures of the three-dimensional mesh generation model.

[0083] In this application, Figure 2 to Figure 7 The network structure of the three-dimensional grid generation model is illustrated by way of example, but is not intended to limit the scope of protection of the present application.

[0084] In step S103, the sample data set is used to train the network parameters in the three-dimensional mesh generation model until the network parameters converge to obtain the three-dimensional mesh generation model.

[0085] In the present application, after the network structure of the three-dimensional mesh generation model is constructed, the network parameters in the three-dimensional mesh generation model can be trained according to at least one sample data set.

[0086] During the training process, sample images of a sample object at multiple different perspectives can be input into the three-dimensional mesh generation model, so that the voxel generation network in the three-dimensional mesh generation model obtains sample voxels of the sample object based on the sample images of the sample object at multiple different perspectives, and then the sample voxels of the sample object are input into the three-dimensional mesh generation network; and, the depth feature extraction network in the three-dimensional mesh generation model is used to respectively obtain sample depth features of the sample images of each perspective of the sample object, and then the sample depth features of the sample images are input into the three-dimensional mesh generation network.

[0087] Afterwards, the 3D mesh generation network in the 3D mesh generation model can obtain a predicted 3D mesh of the sample object according to the sample voxels of the sample object and the sample depth features of the sample images of each viewing angle of the sample object.

[0088] Afterwards, the loss function can be used to adjust the network parameters in the network structure in the 3D mesh generation model based on the predicted 3D mesh of the sample object and the labeled 3D mesh of the sample object until the network parameters in the 3D mesh generation model converge, thereby completing the training and the obtained 3D mesh generation model can be put into use online.

[0089] exist Figure 2 Based on the embodiment shown, in another embodiment of the present application, see Figure 3 , the voxel generation network includes: convolutional neural network, voxel generation layer and voxel fusion layer.

[0090] The convolutional neural network is used to obtain sample convolution features of sample images of each viewing angle of the sample object.

[0091] The voxel generation layer is used to generate perspective voxels corresponding to each perspective of the sample object according to the sample convolution features of the sample image at each perspective of the sample object.

[0092] In one embodiment, when generating the viewing voxels of the sample object corresponding to each viewing angle, the camera parameters corresponding to the sample object in each sample image may be referred to so as to improve the accuracy of the generated viewing voxels of the sample object corresponding to each viewing angle.

[0093] The voxel fusion layer is used to fuse the view voxels of the sample object corresponding to each view to obtain the sample voxels of the sample object.

[0094] Among them, the input end of the voxel generation network includes a convolutional neural network input end.

[0095] The output of the convolutional neural network is connected to the input of the voxel generation layer.

[0096] The output of the voxel generation layer is connected to the input of the voxel fusion layer.

[0097] The output of the voxel generation network includes the output of the voxel fusion layer.

[0098] exist Figure 3 Based on the embodiment shown, in another embodiment of the present application, see Figure 4 ,The 3D mesh generation network includes: a 3D mesh generation layer and a 3D mesh optimization layer.

[0099] The three-dimensional mesh generation layer is used to generate a three-dimensional mesh to be optimized of the sample object according to the sample voxels of the sample object.

[0100] The three-dimensional mesh optimization layer is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images at various viewing angles of the sample object to obtain a predicted three-dimensional mesh.

[0101] The input end of the three-dimensional grid generation layer is connected to the output end of the voxel generation layer.

[0102] The output terminal of the 3D mesh generation layer is connected to the input terminal of the 3D mesh optimization layer.

[0103] The input of the 3D mesh optimization layer is also connected to the output of the deep feature extraction network.

[0104] The output end of the 3D mesh generation network includes the output end of the 3D mesh optimization layer.

[0105] exist Figure 4 Based on the embodiment shown, in another embodiment of the present application, see Figure 5 , the three-dimensional mesh optimization layer includes: graph convolutional neural network and contrast deep feature extraction network.

[0106] The contrast depth feature extraction network is used to obtain sample depth difference information between the depth features of the three-dimensional grid to be optimized and the sample depth features of the sample images of each viewing angle of the sample object.

[0107] The graph convolutional neural network is used to optimize the three-dimensional mesh according to the sample depth difference information to obtain the predicted three-dimensional mesh.

[0108] Among them, the input end of the comparison depth feature extraction network is connected to the output end of the three-dimensional mesh generation network.

[0109] The input end of the contrast deep feature extraction network is also connected to the output end of the deep feature extraction network.

[0110] The output of the contrast deep feature extraction network is connected to the input of the graph convolutional neural network.

[0111] The input of the graph convolutional neural network is also connected to the output of the 3D mesh generation network.

[0112] The output of the 3D mesh optimization layer includes the output of the graph convolutional neural network.

[0113] exist Figure 5 Based on the embodiment shown, in another embodiment of the present application, see Figure 6 , the contrastive depth feature extraction network includes: a neural renderer (Neural Renderer) and a contrastive depth feature extractor (ContrastiveFeature Extractor).

[0114] The neural renderer is used to obtain the deep features of each viewing angle of the 3D mesh to be optimized.

[0115] In one embodiment, when acquiring the depth features of each viewing angle of the 3D mesh to be optimized, the camera parameters corresponding to the sample objects in each sample image may be referred to to improve the accuracy of the acquired depth features of each viewing angle of the 3D mesh to be optimized.

[0116] The contrast depth feature extractor is used to respectively obtain sample depth difference information between the depth features of the three-dimensional grid and the sample depth features of the sample image at the same viewing angle.

[0117] Among them, the input end of the neural renderer is connected to the output end of the 3D mesh generation layer.

[0118] The output of the neural renderer is connected to the input of the contrast deep feature extractor.

[0119] The input of the contrast deep feature extractor is also connected to the output of the deep feature extraction network.

[0120] The output of the contrast deep feature extraction network includes the output of the deep feature extractor.

[0121] exist Figure 4 Based on the embodiment shown, in another embodiment of the present application, see Figure 7 , there are multiple three-dimensional mesh optimization layers, and the multiple three-dimensional mesh optimization layers are cascaded;

[0122] The three-dimensional mesh optimization layer with the first arrangement order is used to optimize the three-dimensional mesh according to the sample depth features of the sample images of each viewing angle of the sample object;

[0123] Among them, in any two adjacent three-dimensional mesh optimization layers arranged in cascade, the three-dimensional mesh optimization layer with a later arrangement order is used to optimize the intermediate three-dimensional mesh output by the three-dimensional mesh optimization layer with a forward arrangement order according to the sample depth features of the sample images of each viewing angle of the sample object.

[0124] The three-dimensional grid optimization layer with the last arrangement order is used to optimize the intermediate three-dimensional grid output with the second last arrangement order (second to last) according to the sample depth features of the sample images of each viewing angle of the sample object to obtain a predicted three-dimensional grid.

[0125] Among any two adjacent three-dimensional grid optimization layers arranged in cascade, the output end of the three-dimensional grid optimization layer arranged in a forward order is connected to the input end of the three-dimensional grid optimization layer arranged in a backward order;

[0126] The output end of the three-dimensional mesh generation layer is connected to the input end of the three-dimensional mesh optimization layer which is the first in the arrangement order;

[0127] The input end of each 3D mesh optimization layer is also connected to the output end of the deep feature extraction network;

[0128] The output end of the three-dimensional mesh generation network includes the output end of the three-dimensional mesh optimization layer which is the last in the arrangement order.

[0129] Among them, through Figure 2 to Figure 7 In the embodiment shown, a plurality of three-dimensional mesh generation models with different network structures can be constructed respectively, and then the three-dimensional mesh generation models with different network structures can be selected according to actual conditions to generate a three-dimensional mesh of the object according to images of the object from multiple different perspectives.

[0130] The three-dimensional mesh generation model constructed by the present application can support obtaining target images of a target object from multiple different perspectives, obtaining target voxels of the target object based on the target images of the target object from multiple different perspectives, and respectively obtaining target depth features of the target images of each perspective of the target object, and then obtaining the target three-dimensional mesh of the target object based on the target voxels of the target object and the target depth features of the target images of each perspective of the target object.

[0131] In this way, on the one hand, the target depth features of the target images of the target object at each perspective can be directly acquired by the three-dimensional mesh generation model constructed by the present application based on the target images of the target object at each perspective, and the target depth features of the target images of the target object at each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the three-dimensional mesh generation model constructed by the present application uses the target depth features of the target images of the target object at each perspective to directly or indirectly optimize the target voxels of the target object (for example, fine-tuning or refining, etc.), thereby improving the accuracy of the target three-dimensional mesh of the acquired target object.

[0132] On the other hand, if the target object in some target images is occluded (at least a part of the target object is occluded, resulting in incomplete image information), in this case, since the three-dimensional mesh generation model constructed by the present application can obtain the target voxels of the target object based on target images of multiple different perspectives of the target object, the three-dimensional mesh generation model constructed by the present application can learn a large number of image features of the target object from target images of multiple different perspectives, thereby completing the image information and solving the problem of incomplete image information caused by occlusion, thereby avoiding the accuracy of the target three-dimensional mesh of the target object acquired due to occlusion problems.

[0133] On the other hand, when obtaining the target three-dimensional mesh of the target object, the three-dimensional mesh generation model constructed by the present application can generate the three-dimensional mesh to be optimized of the target object based on the target voxels of the target object, and then optimize the three-dimensional mesh to be optimized based on the target depth features of the target images of each perspective of the target object to obtain the target three-dimensional mesh.

[0134] It can be seen that the optimization process is an optimization of the three-dimensional grid. The amount of data representing the three-dimensional grid is relatively low and is often lower than the amount of data representing the voxels. Therefore, compared with the method of directly processing the voxels, the method of "the three-dimensional grid generation model constructed in this application first generates the three-dimensional grid to be optimized for the target object based on the target voxels of the target object, and then directly optimizes the three-dimensional grid" in this application can reduce the amount of computational data involved in the optimization process, thereby improving the optimization efficiency and saving computing resources.

[0135] On the other hand, when optimizing the three-dimensional mesh, the graph convolutional neural network in the three-dimensional mesh generation model constructed in the present application can be used to optimize the three-dimensional mesh. With the help of the powerful image processing capabilities of the graph convolutional neural network, the degree of optimization can be improved, and the accuracy of the target three-dimensional mesh of the target object can be further improved.

[0136] On the other hand, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh generation model constructed by the present application can perform multiple rounds of optimization on the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object. Each round of optimization can further optimize the three-dimensional mesh obtained by the previous round of optimization, thereby achieving multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, and further improving the accuracy of the target three-dimensional mesh of the target object.

[0137] On the other hand, since the three-dimensional mesh generation model constructed by the present application can optimize the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object to improve the accuracy of the three-dimensional mesh, it can support the generation of a lower-precision three-dimensional mesh to be optimized when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". Since it supports the generation of a lower-precision three-dimensional mesh to be optimized, it can support the use of low-resolution voxels when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". The amount of data representing the low-resolution voxels is low, thereby reducing the amount of computational data involved in "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object", thereby improving the efficiency of generating the three-dimensional mesh to be optimized for the target object and saving computing resources.

[0138] After the 3D mesh generation model is trained, the trained 3D mesh generation model can be deployed for online application. In this way, if the electronic device needs to generate a target 3D mesh of the target object based on target images of the target object at multiple different perspectives, the electronic device can input the target images of the target object at multiple different perspectives into the 3D mesh generation model trained based on the aforementioned embodiment, so that the 3D mesh generation model generates the 3D mesh of the target object based on the target images of the target object at multiple different perspectives, and outputs the 3D mesh of the target object. In this way, the electronic device can obtain the 3D mesh of the target object.

[0139] Specifically, the specific process of the 3D mesh generation model generating the 3D mesh of the target object according to the target images of the target object at multiple different perspectives can be found in Figure 8 The illustrated embodiment will not be described in detail here.

[0140] For example, refer to Figure 8 , shows a schematic flow chart of a three-dimensional grid generation method of the present application, the method is applied in an electronic device, and the method may include:

[0141] In step S201, target images of a target object at multiple different viewing angles are acquired.

[0142] The target objects include three-dimensional objects, such as people, sofas, vehicles, trees, and buildings.

[0143] The perspectives of the target objects in the respective target images are different.

[0144] The viewing angle can be understood as the direction or angle of viewing / collecting the target object. For example, taking a vehicle as an example, viewing the vehicle from the front of the vehicle is one viewing angle, viewing the vehicle from the left side of the vehicle is one viewing angle, viewing the vehicle from the right side of the vehicle is one viewing angle, viewing the vehicle from the left front side of the vehicle is one viewing angle, viewing the vehicle from the right front side of the vehicle is one viewing angle, viewing the vehicle from the left rear side of the vehicle is one viewing angle, viewing the vehicle from the right rear side of the vehicle is one viewing angle, and viewing the vehicle from the rear of the vehicle is one viewing angle, etc.

[0145] The target image may include an RGB image, etc. The target image may also include camera parameters corresponding to the target object. When acquiring the target image of the target object, a camera may be used to acquire the target image of the target object. The camera may obtain the relative position relationship between the camera and the target object as the camera parameters. After the camera acquires the target image of the target object, the camera parameters may be added to the target image.

[0146] In step S202, target voxels of the target object are acquired according to target images of the target object at multiple different viewing angles, and target depth features of the target images of the target object at each viewing angle are acquired respectively.

[0147] In the present application, target images of a target object at multiple different perspectives can be input into a three-dimensional mesh generation model, so that the three-dimensional mesh generation model uses the depth feature extraction network included therein to perform multi-view depth estimation (Multi View Depth Estimation) on the target images of the target object at multiple different perspectives, thereby obtaining the target depth features of the target images of each perspective of the target object.

[0148] In one embodiment, when acquiring target depth features of target images of each perspective of the target object, camera parameters corresponding to the target object in each target image may be referenced to improve the accuracy of the acquired target depth features of the target images of each perspective of the target object.

[0149] The deep feature extraction network may include MVSNet, CascadeMVSNet or PointMVSNet, etc. Of course, it may also include other networks for extracting deep features, which is not limited in this application.

[0150] The three-dimensional mesh generation model may use the voxel generation network included therein to obtain target voxels of the target object according to target images of the target object at multiple different perspectives.

[0151] Specifically, when obtaining the target voxel of the target object according to the target images of the target object at multiple different perspectives, it can be achieved through the following process, including:

[0152] 11) Obtain target convolution features of target images of target objects at different viewpoints respectively.

[0153] For a target image of a target object from any perspective, the target image can be input into a convolutional neural network in a voxel generation network in a three-dimensional grid generation model, so that the convolutional neural network processes the target image to obtain a target convolution feature of the target image.

[0154] The same is true for the target image of each other perspective of the target object.

[0155] 12) Generate perspective voxels of the target object corresponding to each perspective according to the target convolution features of the target image at each perspective of the target object.

[0156] For the target convolution features of the target image of any perspective of the target object, the target convolution features can be input into the voxel generation layer in the voxel generation network in the three-dimensional grid generation model, so that the voxel generation layer processes the target convolution features to obtain the perspective voxels of the target object corresponding to the perspective. The processing of the target convolution features by the voxel generation layer can refer to the existing processing methods, and the present application is not limited to this.

[0157] The same is true for the target convolution features of the target image of each other view of the target object, so that the target object corresponds to the view voxels of each view.

[0158] In one embodiment, when generating the view voxels of the target object corresponding to each view angle, the camera parameters corresponding to the target object in each target image may be referred to so as to improve the accuracy of the generated view voxels of the target object corresponding to each view angle.

[0159] 13) Fuse the view voxels of the target object corresponding to each view to obtain the target voxel of the target object.

[0160] The present application does not limit the method of fusing multiple viewpoint voxels. In an optional example, assuming that the viewpoint voxels corresponding to each viewpoint of the target object are all three-dimensional matrices (including multiple elements, for example, multiple eigenvectors, etc.), then the elements corresponding to the same position of the target object can be summed in the viewpoint voxels corresponding to each viewpoint of the target object, thereby achieving the fusion of the viewpoint voxels corresponding to each viewpoint of the target object. Alternatively, the weighted average value of the elements corresponding to the same position of the target object is calculated, thereby achieving the fusion of the viewpoint voxels corresponding to each viewpoint of the target object.

[0161] In step S203, a target three-dimensional grid of the target object is acquired according to the target voxels of the target object and the target depth features of the target images of each viewing angle of the target object.

[0162] The three-dimensional mesh in the present application may include a triangular mesh, a quadrilateral mesh or a pentagonal mesh, etc. The surface of a three-dimensional object can be simulated by the three-dimensional mesh.

[0163] This step can be implemented through the following process, including:

[0164] 2031. Generate a three-dimensional mesh to be optimized of the target object according to the target voxels of the target object.

[0165] In the present application, the target voxels of the target object may be cubed to obtain the three-dimensional mesh to be optimized of the target object. For example, according to the target voxel occupancy probability of the target object and a preset binarization voxel threshold, the fused voxels are then converted into the three-dimensional mesh to be optimized of the target object through the cubed operation.

[0166] The present application does not limit the specific method for generating a three-dimensional grid based on voxels, and the details can also refer to the existing methods.

[0167] 2032. Optimize the three-dimensional mesh to be optimized according to the target depth features of the target images at each viewing angle of the target object to obtain a target three-dimensional mesh.

[0168] In one embodiment of the application, step 2032 may be implemented by a process including:

[0169] 11) Obtain target depth difference information between the depth features of the three-dimensional mesh to be optimized and the target depth features of the target images of each perspective of the target object.

[0170] For example, the depth features of each viewing angle of the three-dimensional mesh to be optimized can be obtained. For example, the depth features of each viewing angle of the three-dimensional mesh to be optimized can be obtained according to the neural renderer. In one example, the three-dimensional mesh to be optimized can be input into the neural renderer in the three-dimensional mesh generation network, the three-dimensional mesh optimization layer, and the comparison depth feature extraction network in the three-dimensional mesh generation model, so that the neural renderer processes the three-dimensional mesh to be optimized and obtains the depth features of each viewing angle of the three-dimensional mesh to be optimized, and the "each viewing angle" in the depth features of each viewing angle of the three-dimensional mesh to be optimized corresponds one to one with the "multiple different viewing angles" in the target image of multiple different viewing angles of the target object.

[0171] In one embodiment, when acquiring the depth features of each viewing angle of the 3D mesh to be optimized, the camera parameters corresponding to the target objects in each target image may be referred to to improve the accuracy of the acquired depth features of each viewing angle of the 3D mesh to be optimized.

[0172] Then, the target depth difference information between the depth features of the three-dimensional mesh of the same viewing angle and the target depth features of the target image can be obtained respectively. For example, for any viewing angle, the target depth difference information between the depth features of the three-dimensional mesh of the viewing angle and the depth features of the target image of the viewing angle can be obtained according to the contrast depth feature extractor. In one example, the depth features of the three-dimensional mesh of the viewing angle and the depth features of the target image of the viewing angle can be input into the contrast depth feature extractor in the three-dimensional mesh generation network, the three-dimensional mesh optimization layer, and the contrast depth feature extraction network in the three-dimensional mesh generation model, so that the contrast depth feature extractor processes the depth features of the three-dimensional mesh of the viewing angle and the depth features of the target image of the viewing angle to obtain the target depth difference information between the depth features of the three-dimensional mesh of the viewing angle and the depth features of the target image of the viewing angle. The same is true for each other viewing angle.

[0173] 12) Optimize the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh.

[0174] In the present application, a graph convolutional neural network can be used to optimize the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh.

[0175] For instance, in one example, the target difference depth information can be input into the graph convolutional neural network in the three-dimensional mesh generation network and the three-dimensional mesh optimization layer in the three-dimensional mesh generation model, and the three-dimensional mesh to be optimized can be input into the graph convolutional neural network in the three-dimensional mesh generation network and the three-dimensional mesh optimization layer in the three-dimensional mesh generation model, so that the graph convolutional neural network optimizes the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh, and the accuracy of the target three-dimensional mesh is higher than the accuracy of the three-dimensional mesh to be optimized.

[0176] In another embodiment of the present application, in step S2023, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh to be optimized can be optimized in multiple rounds in sequence according to the target depth features of the target images of each perspective of the target object, so as to achieve multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, thereby further improving the accuracy of the target three-dimensional mesh of the acquired target object.

[0177] Specifically, the three-dimensional grid to be optimized is optimized according to the target depth features of the target images at each viewing angle of the target object to obtain a first intermediate three-dimensional grid;

[0178] The first intermediate three-dimensional grid is optimized according to the target depth features of the target images at various perspectives of the target object to obtain the second intermediate three-dimensional grid. Similarly, the N-1th intermediate three-dimensional grid is optimized according to the target depth features of the target images at various perspectives of the target object to obtain the Nth intermediate three-dimensional grid; N is a positive integer greater than 1; the specific value of N can be determined according to actual conditions, and this application does not limit this.

[0179] The Nth intermediate three-dimensional grid is optimized according to target depth features of target images at various viewing angles of the target object to obtain a target three-dimensional grid.

[0180] In the present application, target images of the target object at multiple different perspectives are acquired, target voxels of the target object are acquired based on the target images of the target object at multiple different perspectives, and target depth features of the target images of the target object at each perspective are acquired respectively. Then, a target three-dimensional grid of the target object is acquired based on the target voxels of the target object and the target depth features of the target images of the target object at each perspective.

[0181] Through the present application, on the one hand, the target depth features of the target images of the target object at each perspective can be directly acquired based on the target images of the target object at each perspective, and the target depth features of the target images of the target object at each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the target voxels of the target object can be directly or indirectly optimized (for example, fine-tuned or refined) using the target depth features of the target images of the target object at each perspective, thereby improving the accuracy of the acquired target three-dimensional mesh of the target object.

[0182] On the other hand, if the target object in some target images is occluded (at least a part of the target object is occluded, resulting in incomplete image information), in this case, since the present application obtains the target voxels of the target object based on target images of multiple different perspectives of the target object, a large number of image features of the target object can be learned from the target images of multiple different perspectives, thereby completing the image information and solving the problem of incomplete image information caused by occlusion, thereby avoiding the accuracy of the target three-dimensional mesh of the acquired target object being affected by the occlusion problem.

[0183] On the other hand, when obtaining the target three-dimensional mesh of the target object, the three-dimensional mesh to be optimized can be generated according to the target voxels of the target object, and then the three-dimensional mesh to be optimized can be optimized according to the target depth features of the target images of each perspective of the target object to obtain the target three-dimensional mesh.

[0184] It can be seen that the optimization process is an optimization of the three-dimensional grid. The amount of data representing the three-dimensional grid is relatively low and is often lower than the amount of data representing the voxels. Therefore, compared with the method of directly processing the voxels, the method of "first generating the three-dimensional grid to be optimized for the target object based on the target voxels of the target object, and then directly optimizing the three-dimensional grid" in the present application can reduce the amount of computational data involved in the optimization process, thereby improving the optimization efficiency and saving computing resources.

[0185] On the other hand, when optimizing the three-dimensional mesh, a graph convolutional neural network can be used to optimize the three-dimensional mesh. With the help of the powerful image processing capabilities of the graph convolutional neural network, the degree of optimization can be improved, and the accuracy of the target three-dimensional mesh of the target object can be further improved.

[0186] On the other hand, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh to be optimized can be optimized for multiple rounds in sequence according to the target depth features of the target images of each perspective of the target object. Each round of optimization can further optimize the three-dimensional mesh obtained in the previous round of optimization, thereby achieving multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, and further improving the accuracy of the target three-dimensional mesh of the target object.

[0187] On the other hand, since the present application can optimize the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object to improve the accuracy of the three-dimensional mesh, it can support the generation of a lower-accuracy three-dimensional mesh to be optimized when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". Since it supports the generation of a lower-accuracy three-dimensional mesh to be optimized, it can support the use of low-resolution voxels when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". The data volume of the low-resolution voxels is low, thereby reducing the amount of computational data involved in "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object", thereby improving the efficiency of generating the three-dimensional mesh to be optimized for the target object and saving computing resources.

[0188] The three-dimensional mesh generated by the present application can be applied to the three-dimensional display scene of the object, for example, it can be applied to the three-dimensional display of the object in the cloud-based exhibition. In one example, there are display objects in the exhibition. When some users cannot go to the exhibition in person, the display objects in the exhibition can be three-dimensionally reconstructed to obtain the three-dimensional mesh of the display objects in the exhibition, and the three-dimensional mesh of the display objects in the exhibition can be uploaded to the cloud, so that the user can download and view the three-dimensional mesh of the display objects in the exhibition from the cloud, so that the user can see the three-dimensional structure of the display objects in the exhibition and understand the display objects in the exhibition.

[0189] The three-dimensional grid generated by this application can also be applied to AR (Augmented Reality) house viewing scenes or VR (Virtual Reality) house viewing scenes. For example, a three-dimensional grid of each object in the house (furniture and appliances, or hard furnishings and soft furnishings, etc.) and a three-dimensional grid of the house can be constructed so that users can see the three-dimensional structure of the house and the three-dimensional structure of the objects in the house, giving users a more realistic house viewing experience.

[0190] The three-dimensional grid generated by this application can also be applied to map scenes. For example, when a user views a map of a certain area, a three-dimensional grid of objects in the area (roads, trees, mountains, rivers, and buildings, etc.) can be generated so that the user can see the three-dimensional structure of the objects in the area on the map, giving the user a more realistic map viewing experience.

[0191] The three-dimensional grid generated by this application can also be applied to 3D printing of objects, digital preservation of objects (such as antiques, etc.), searching for objects in a large number of objects (commodity search or lost person / object search, etc.), human identity recognition, comparison of the shape or volume of at least two objects, and three-dimensional reconstruction of lost objects (such as antiques, etc.).

[0192] In this application, Figure 8 The triggering condition of the process of acquiring the target three-dimensional grid of the target object according to the target images of the target object at multiple different viewing angles may be manually triggered.

[0193] For example, when a user needs to make an electronic device generate a target three-dimensional mesh of a target object, the user can input a generation operation in the electronic device so that the electronic device starts to execute the generation operation according to the generation operation. Figure 2 In the illustrated embodiment, a target three-dimensional grid of the target object is acquired based on target images of the target object at multiple different viewing angles.

[0194] Among them, the generating operation may include a changing operation of changing the posture of the electronic device, an input operation of inputting a specific gesture on the touch screen of the electronic device, an input operation of inputting a specific voice command in the electronic device, an input operation of inputting a specific facial expression on the electronic device, and a click operation of clicking a specific virtual button displayed on the touch screen of the electronic device, etc.

[0195] The method provided in the present application may be a working method of an application in an electronic device (such as a mobile phone, computer, PAD, etc.). Correspondingly, the memory provided by the present invention may be a storage medium of the electronic device. The application calls the hardware such as the camera, memory, processor, etc. of the electronic device to implement the above method.

[0196] Reference Fig. 9 , shows a structural block diagram of an embodiment of a construction device for a three-dimensional mesh generation model of the present application, which may specifically include the following modules:

[0197] A fifth acquisition module 11 is used to acquire at least one sample data set, wherein the sample data set includes: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object;

[0198] A construction module 12 is used to construct a network structure of a three-dimensional mesh generation model;

[0199] A training module 13, used to train the network parameters in the three-dimensional mesh generation model using the sample data set until the network parameters converge to obtain the three-dimensional mesh generation model;

[0200] Wherein, the network structure at least includes a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network;

[0201] The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different viewing angles;

[0202] The depth feature extraction network is used to respectively obtain sample depth features of sample images of each viewing angle of the sample object;

[0203] The three-dimensional mesh generation network is used to obtain a predicted three-dimensional mesh of the sample object according to sample voxels of the sample object and sample depth features of sample images of each viewing angle of the sample object.

[0204] In an optional implementation, the voxel generation network includes: a convolutional neural network, a voxel generation layer, and a voxel fusion layer;

[0205] The convolutional neural network is used to respectively obtain sample convolution features of sample images of each viewing angle of the sample object;

[0206] The voxel generation layer is used to generate perspective voxels corresponding to each perspective of the sample object according to the sample convolution features of the sample image of each perspective of the sample object;

[0207] The voxel fusion layer is used to fuse the view voxels of the sample object corresponding to each view angle to obtain the sample voxels of the sample object.

[0208] In an optional implementation, the three-dimensional grid generation network includes: a three-dimensional grid generation layer and a three-dimensional grid optimization layer;

[0209] The three-dimensional mesh generation layer is used to generate the three-dimensional mesh to be optimized of the sample object according to the sample voxels of the sample object;

[0210] The three-dimensional mesh optimization layer is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images at various viewing angles of the sample object to obtain the predicted three-dimensional mesh.

[0211] In an optional implementation, the three-dimensional mesh optimization layer includes: a graph convolutional neural network and a contrastive deep feature extraction network;

[0212] The contrast depth feature extraction network is used to obtain sample depth difference information between the depth feature of the three-dimensional grid to be optimized and the sample depth features of the sample images of each viewing angle of the sample object;

[0213] The graph convolutional neural network is used to optimize the three-dimensional grid to be optimized according to the sample depth difference information to obtain the predicted three-dimensional grid.

[0214] In an optional implementation, the contrastive depth feature extraction network includes: a neural renderer and a contrastive depth feature extractor;

[0215] The neural renderer is used to obtain the depth features of each viewing angle of the three-dimensional mesh to be optimized;

[0216] The contrast depth feature extractor is used to respectively obtain sample depth difference information between the depth feature of the three-dimensional grid and the sample depth feature of the sample image at the same viewing angle.

[0217] In an optional implementation, there are multiple three-dimensional grid optimization layers, and the multiple three-dimensional grid optimization layers are cascaded;

[0218] The three-dimensional mesh optimization layer with the first arrangement order is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images of each viewing angle of the sample object;

[0219] Among any two adjacent three-dimensional mesh optimization layers arranged in cascade, the three-dimensional mesh optimization layer arranged later is used to optimize the intermediate three-dimensional mesh output by the three-dimensional mesh optimization layer arranged earlier according to the sample depth features of the sample images of each viewing angle of the sample object;

[0220] The three-dimensional grid optimization layer with the last arrangement order is used to optimize the intermediate three-dimensional grid output by the three-dimensional grid optimization layer with the second last arrangement order according to the sample depth features of the sample images of each viewing angle of the sample object to obtain the predicted three-dimensional grid.

[0221] In an optional implementation, the input end of the three-dimensional mesh generation model includes an input end of a voxel generation network and an input end of the deep feature extraction network;

[0222] The output end of the voxel generation network is connected to the input end of the three-dimensional mesh generation network;

[0223] The output end of the deep feature extraction network is connected to the input end of the three-dimensional mesh generation network;

[0224] The output end of the three-dimensional mesh generation model includes the output end of the three-dimensional mesh generation network.

[0225] In an optional implementation, the voxel generation network includes: a convolutional neural network, a voxel generation layer, and a voxel fusion layer;

[0226] The input end of the voxel generation network includes the convolutional neural network input end;

[0227] The output end of the convolutional neural network is connected to the input end of the voxel generation layer;

[0228] The output end of the voxel generation layer is connected to the input end of the voxel fusion layer;

[0229] The output of the voxel generation network includes the output of the voxel fusion layer.

[0230] In an optional implementation, the three-dimensional grid generation network includes: a three-dimensional grid generation layer and a three-dimensional grid optimization layer;

[0231] The input end of the three-dimensional grid generation layer is connected to the output end of the voxel generation layer;

[0232] The output end of the three-dimensional grid generation layer is connected to the input end of the three-dimensional grid optimization layer;

[0233] The input end of the three-dimensional mesh optimization layer is also connected to the output end of the deep feature extraction network;

[0234] The output end of the three-dimensional mesh generation network includes the output end of the three-dimensional mesh optimization layer.

[0235] In an optional implementation, the three-dimensional mesh optimization layer includes: a graph convolutional neural network and a contrastive deep feature extraction network;

[0236] The input end of the comparison depth feature extraction network is connected to the output end of the three-dimensional mesh generation network;

[0237] The input end of the comparison depth feature extraction network is also connected to the output end of the depth feature extraction network;

[0238] The output end of the comparison depth feature extraction network is connected to the input end of the graph convolutional neural network;

[0239] The input end of the graph convolutional neural network is also connected to the output end of the three-dimensional grid generation network;

[0240] The output end of the three-dimensional grid optimization layer includes the output end of the graph convolutional neural network.

[0241] In an optional implementation, the contrastive depth feature extraction network includes: a neural renderer and a contrastive depth feature extractor;

[0242] The input end of the neural renderer is connected to the output end of the three-dimensional mesh generation layer;

[0243] An output terminal of the neural renderer is connected to an input terminal of the contrast depth feature extractor;

[0244] The input end of the comparison depth feature extractor is also connected to the output end of the depth feature extraction network;

[0245] The output end of the comparative deep feature extraction network includes the output end of the deep feature extractor.

[0246] In an optional implementation, there are multiple three-dimensional grid optimization layers, and the multiple three-dimensional grid optimization layers are cascaded;

[0247] Among any two adjacent three-dimensional grid optimization layers arranged in cascade, the output end of the three-dimensional grid optimization layer arranged in a forward order is connected to the input end of the three-dimensional grid optimization layer arranged in a backward order;

[0248] The output end of the three-dimensional grid generation layer is connected to the input end of the three-dimensional grid optimization layer that is ranked first in the arrangement order;

[0249] The input end of each three-dimensional mesh optimization layer is also connected to the output end of the deep feature extraction network respectively;

[0250] The output end of the three-dimensional mesh generation network includes the output end of the three-dimensional mesh optimization layer which is the last in the arrangement order.

[0251] The three-dimensional mesh generation model constructed by the present application can support obtaining target images of a target object from multiple different perspectives, obtaining target voxels of the target object based on the target images of the target object from multiple different perspectives, and respectively obtaining target depth features of the target images of each perspective of the target object, and then obtaining the target three-dimensional mesh of the target object based on the target voxels of the target object and the target depth features of the target images of each perspective of the target object.

[0252] In this way, on the one hand, the target depth features of the target images of the target object at each perspective can be directly acquired by the three-dimensional mesh generation model constructed by the present application based on the target images of the target object at each perspective, and the target depth features of the target images of the target object at each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the three-dimensional mesh generation model constructed by the present application uses the target depth features of the target images of the target object at each perspective to directly or indirectly optimize the target voxels of the target object (for example, fine-tuning or refining, etc.), thereby improving the accuracy of the target three-dimensional mesh of the acquired target object.

[0253] On the other hand, if the target object in some target images is occluded (at least a part of the target object is occluded, resulting in incomplete image information), in this case, since the three-dimensional mesh generation model constructed by the present application can obtain the target voxels of the target object based on target images of multiple different perspectives of the target object, the three-dimensional mesh generation model constructed by the present application can learn a large number of image features of the target object from target images of multiple different perspectives, thereby completing the image information and solving the problem of incomplete image information caused by occlusion, thereby avoiding the accuracy of the target three-dimensional mesh of the target object acquired due to occlusion problems.

[0254] On the other hand, when obtaining the target three-dimensional mesh of the target object, the three-dimensional mesh generation model constructed by the present application can generate the three-dimensional mesh to be optimized of the target object based on the target voxels of the target object, and then optimize the three-dimensional mesh to be optimized based on the target depth features of the target images of each perspective of the target object to obtain the target three-dimensional mesh.

[0255] It can be seen that the optimization process is an optimization of the three-dimensional grid. The amount of data representing the three-dimensional grid is relatively low and is often lower than the amount of data representing the voxels. Therefore, compared with the method of directly processing the voxels, the method of "the three-dimensional grid generation model constructed in this application first generates the three-dimensional grid to be optimized for the target object based on the target voxels of the target object, and then directly optimizes the three-dimensional grid" in this application can reduce the amount of computational data involved in the optimization process, thereby improving the optimization efficiency and saving computing resources.

[0256] On the other hand, when optimizing the three-dimensional mesh, the graph convolutional neural network in the three-dimensional mesh generation model constructed in the present application can be used to optimize the three-dimensional mesh. With the help of the powerful image processing capabilities of the graph convolutional neural network, the degree of optimization can be improved, and the accuracy of the target three-dimensional mesh of the target object can be further improved.

[0257] On the other hand, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh generation model constructed by the present application can perform multiple rounds of optimization on the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object. Each round of optimization can further optimize the three-dimensional mesh obtained by the previous round of optimization, thereby achieving multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, and further improving the accuracy of the target three-dimensional mesh of the target object.

[0258] On the other hand, since the three-dimensional mesh generation model constructed by the present application can optimize the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object to improve the accuracy of the three-dimensional mesh, it can support the generation of a lower-precision three-dimensional mesh to be optimized when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". Since it supports the generation of a lower-precision three-dimensional mesh to be optimized, it can support the use of low-resolution voxels when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". The amount of data representing the low-resolution voxels is low, thereby reducing the amount of computational data involved in "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object", thereby improving the efficiency of generating the three-dimensional mesh to be optimized for the target object and saving computing resources.

[0259] Reference Fig.10 , shows a structural block diagram of an embodiment of a three-dimensional grid generation device of the present application, which may specifically include the following modules:

[0260] A first acquisition module 21 is used to acquire target images of a target object at multiple different viewing angles;

[0261] A second acquisition module 22 is used to acquire target voxels of the target object according to target images of multiple different perspectives of the target object, and a third acquisition module 23 is used to respectively acquire target depth features of the target images of each perspective of the target object;

[0262] The fourth acquisition module 24 is configured to acquire a target three-dimensional grid of the target object according to the target voxels of the target object and the target depth features of the target images of each viewing angle of the target object.

[0263] In an optional implementation, the second acquisition module includes:

[0264] An acquisition submodule, used to respectively acquire target convolution features of target images of each viewing angle of the target object;

[0265] A first generating submodule, configured to generate perspective voxels of the target object corresponding to each perspective according to the target convolution features of the target image of each perspective of the target object;

[0266] The fusion submodule is used to fuse the view voxels of the target object corresponding to each view to obtain the target voxel of the target object.

[0267] In an optional implementation, the fourth acquisition module includes:

[0268] A second generating submodule, configured to generate a to-be-optimized three-dimensional mesh of the target object according to the target voxels of the target object;

[0269] The optimization submodule is used to optimize the three-dimensional mesh to be optimized according to the target depth features of the target images of each viewing angle of the target object to obtain the target three-dimensional mesh.

[0270] In an optional implementation, the optimization submodule includes:

[0271] An acquisition unit, configured to acquire target depth difference information between the depth feature of the three-dimensional mesh to be optimized and the target depth features of the target images of each viewing angle of the target object;

[0272] The first optimization unit is used to optimize the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh.

[0273] In an optional implementation, the acquiring unit includes:

[0274] A first acquisition subunit is used to acquire depth features of each viewing angle of the three-dimensional grid to be optimized;

[0275] The second acquisition subunit is used to respectively acquire target depth difference information between the depth feature of the three-dimensional grid and the target depth feature of the target image at the same viewing angle.

[0276] In an optional implementation, the first acquisition subunit is specifically used to: acquire depth features of each viewing angle of the three-dimensional mesh to be optimized according to a neural renderer.

[0277] In an optional implementation, the second acquisition subunit is specifically used to: for each viewing angle, obtain target depth difference information between the depth features of the three-dimensional grid of the viewing angle and the depth features of the target image of the viewing angle according to a comparison depth feature extractor.

[0278] In an optional implementation, the first optimization unit includes:

[0279] The optimization subunit is used to optimize the three-dimensional grid to be optimized using a graph convolutional neural network according to the target depth difference information to obtain the target three-dimensional grid.

[0280] In an optional implementation, the optimization submodule includes:

[0281] A second optimization unit is used to optimize the three-dimensional grid to be optimized according to the target depth features of the target images of each viewing angle of the target object to obtain a first intermediate three-dimensional grid;

[0282] a third optimization unit, configured to optimize the first intermediate three-dimensional grid according to the target depth features of the target images at each viewing angle of the target object to obtain a second intermediate three-dimensional grid, and so on, optimize the N-1th intermediate three-dimensional grid according to the target depth features of the target images at each viewing angle of the target object to obtain an Nth intermediate three-dimensional grid; N is a positive integer greater than 1;

[0283] The fourth optimization unit is used to optimize the Nth intermediate three-dimensional grid according to the target depth features of the target images of each viewing angle of the target object to obtain the target three-dimensional grid.

[0284] In the present application, target images of the target object at multiple different perspectives are acquired, target voxels of the target object are acquired based on the target images of the target object at multiple different perspectives, and target depth features of the target images of the target object at each perspective are acquired respectively. Then, a target three-dimensional grid of the target object is acquired based on the target voxels of the target object and the target depth features of the target images of the target object at each perspective.

[0285] Through the present application, on the one hand, the target depth features of the target images of the target object at each perspective can be directly acquired based on the target images of the target object at each perspective, and the target depth features of the target images of the target object at each perspective can accurately reflect the depth information of the target object in the target images of each perspective. In this way, the target voxels of the target object can be directly or indirectly optimized (for example, fine-tuned or refined) using the target depth features of the target images of the target object at each perspective, thereby improving the accuracy of the acquired target three-dimensional mesh of the target object.

[0286] On the other hand, if the target object in some target images is occluded (at least a part of the target object is occluded, resulting in incomplete image information), in this case, since the present application obtains the target voxels of the target object based on target images of multiple different perspectives of the target object, a large number of image features of the target object can be learned from the target images of multiple different perspectives, thereby completing the image information and solving the problem of incomplete image information caused by occlusion, thereby avoiding the accuracy of the target three-dimensional mesh of the acquired target object being affected by the occlusion problem.

[0287] On the other hand, when obtaining the target three-dimensional mesh of the target object, the three-dimensional mesh to be optimized can be generated according to the target voxels of the target object, and then the three-dimensional mesh to be optimized can be optimized according to the target depth features of the target images of each perspective of the target object to obtain the target three-dimensional mesh.

[0288] It can be seen that the optimization process is an optimization of the three-dimensional grid. The amount of data representing the three-dimensional grid is relatively low and is often lower than the amount of data representing the voxels. Therefore, compared with the method of directly processing the voxels, the method of "first generating the three-dimensional grid to be optimized for the target object based on the target voxels of the target object, and then directly optimizing the three-dimensional grid" in the present application can reduce the amount of computational data involved in the optimization process, thereby improving the optimization efficiency and saving computing resources.

[0289] On the other hand, when optimizing the three-dimensional mesh, a graph convolutional neural network can be used to optimize the three-dimensional mesh. With the help of the powerful image processing capabilities of the graph convolutional neural network, the degree of optimization can be improved, and the accuracy of the target three-dimensional mesh of the target object can be further improved.

[0290] On the other hand, when optimizing the three-dimensional mesh to be optimized, the three-dimensional mesh to be optimized can be optimized for multiple rounds in sequence according to the target depth features of the target images of each perspective of the target object. Each round of optimization can further optimize the three-dimensional mesh obtained in the previous round of optimization, thereby achieving multi-level optimization of the three-dimensional mesh to be optimized from coarse to fine, and further improving the accuracy of the target three-dimensional mesh of the target object.

[0291] On the other hand, since the present application can optimize the three-dimensional mesh to be optimized according to the target depth features of the target image of each perspective of the target object to improve the accuracy of the three-dimensional mesh, it can support the generation of a lower-accuracy three-dimensional mesh to be optimized when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". Since it supports the generation of a lower-accuracy three-dimensional mesh to be optimized, it can support the use of low-resolution voxels when "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object". The data volume of the low-resolution voxels is low, thereby reducing the amount of computational data involved in "generating the three-dimensional mesh to be optimized for the target object according to the target voxels of the target object", thereby improving the efficiency of generating the three-dimensional mesh to be optimized for the target object and saving computing resources.

[0292] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0293] The present application embodiment provides one or more machine-readable media on which instructions are stored, and when executed by one or more processors, the electronic device executes one or more of the methods described in the above embodiments. In the present application embodiment, the electronic device includes a server, a gateway, a sub-device, etc., and the sub-device is an Internet of Things device or other device.

[0294] The embodiments of the present disclosure may be implemented as an apparatus configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the apparatus may include electronic devices such as servers (clusters), terminal devices such as IoT devices, and the like.

[0295] Fig.11 An exemplary apparatus 1300 that can be used to implement various embodiments described in this application is schematically shown.

[0296] For one embodiment, Fig.11 An exemplary apparatus 1300 is shown having one or more processors 1302, a control module (chip set) 1304 coupled to at least one of the (one or more) processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0297] The processor 1302 may include one or more single-core or multi-core processors, and the processor 1302 may include any combination of general-purpose processors or special-purpose processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 1300 can be used as a server device such as a gateway described in the embodiments of the present application.

[0298] In some embodiments, the device 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 configured to execute the instructions 1314 in combination with the one or more computer-readable media to implement a module to perform the actions described in the present disclosure.

[0299] For one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1302 and / or any suitable device or component in communication with the control module 1304 .

[0300] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0301] The memory 1306 may be used, for example, to load and store data and / or instructions 1314 for the device 1300. For one embodiment, the memory 1306 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 1306 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0302] For one embodiment, control module 1304 may include one or more input / output controllers to provide an interface to NVM / storage device 1308 and input / output device(s) 1310 .

[0303] For example, NVM / storage 1308 may be used to store data and / or instructions 1314. NVM / storage 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0304] NVM / storage device 1308 may include storage resources that are physically part of the device on which apparatus 1300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 1308 may be accessed via input / output device(s) 1310 over a network.

[0305] (One or more) input / output devices 1310 may provide an interface for the apparatus 1300 to communicate with any other appropriate device, and the input / output device 1310 may include a communication component, a phonetic component, a sensor component, etc. The network interface 1312 may provide an interface for the apparatus 1300 to communicate through one or more networks, and the apparatus 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0306] For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers of the control module 1304 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304 to form a system-on-chip (SoC).

[0307] In various embodiments, the device 1300 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0308] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enable the electronic device to execute a method for constructing a three-dimensional mesh generation model as described in one or more of the present application.

[0309] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enable the electronic device to perform one or more three-dimensional mesh generation methods as described in the present application.

[0310] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0311] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0312] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable information processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0313] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable information processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0314] These computer program instructions can also be loaded onto a computer or other programmable information processing terminal device, so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0315] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0316] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0317] The above is a detailed introduction to the construction method of the three-dimensional grid generation model, the three-dimensional grid generation method and the device provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A three-dimensional mesh generation method, It is characterized in that The method comprises: Acquire target images of a target object at multiple different viewing angles; the target object includes a three-dimensional object; Acquire target voxels of the target object according to target images of the target object at multiple different perspectives, and respectively acquire target depth features of the target images of the target object at each perspective; Generating a to-be-optimized three-dimensional mesh of the target object according to the target voxels of the target object; The three-dimensional mesh to be optimized is optimized according to target depth features of target images of each viewing angle of the target object to obtain a target three-dimensional mesh.

2. The method according to claim 1, It is characterized in that The step of acquiring a target voxel of the target object according to a plurality of target images of the target object at different viewing angles comprises: Respectively obtaining target convolution features of target images of each viewing angle of the target object; Generate perspective voxels of the target object corresponding to each perspective respectively according to the target convolution features of the target image of each perspective of the target object; The view voxels of the target object corresponding to each view angle are fused to obtain the target voxels of the target object.

3. The method according to claim 1, It is characterized in that The step of optimizing the three-dimensional mesh to be optimized according to the target depth features of the target images of each viewing angle of the target object to obtain the target three-dimensional mesh comprises: Acquire target depth difference information between the depth feature of the three-dimensional grid to be optimized and the target depth features of the target images of each viewing angle of the target object; The three-dimensional mesh to be optimized is optimized according to the target depth difference information to obtain the target three-dimensional mesh.

4. The method according to claim 3, It is characterized in that The obtaining target depth difference information between the depth feature of the three-dimensional mesh to be optimized and the target depth features of the target images of each viewing angle of the target object includes: Acquire the depth features of each viewing angle of the three-dimensional grid to be optimized; Target depth difference information between the depth features of the three-dimensional grid at the same viewing angle and the target depth features of the target image is obtained respectively.

5. The method according to claim 4, It is characterized in that The obtaining of depth features of each viewing angle of the three-dimensional mesh to be optimized includes: The depth features of each viewing angle of the three-dimensional mesh to be optimized are obtained according to the neural renderer.

6. The method according to claim 4, It is characterized in that The step of respectively acquiring target depth difference information between the depth features of the three-dimensional grid at the same viewing angle and the target depth features of the target image includes: For each viewing angle, target depth difference information between the depth feature of the three-dimensional grid of the viewing angle and the depth feature of the target image of the viewing angle is obtained according to the comparison depth feature extractor.

7. The method according to claim 3, It is characterized in that The step of optimizing the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh includes: The three-dimensional mesh to be optimized is optimized using a graph convolutional neural network according to the target depth difference information to obtain the target three-dimensional mesh.

8. The method according to claim 1, It is characterized in that The step of optimizing the three-dimensional mesh to be optimized according to the target depth features of the target images of each viewing angle of the target object to obtain the target three-dimensional mesh comprises: Optimizing the three-dimensional grid to be optimized according to target depth features of target images at various viewing angles of the target object to obtain a first intermediate three-dimensional grid; The first intermediate three-dimensional grid is optimized according to the target depth features of the target images at each viewing angle of the target object to obtain a second intermediate three-dimensional grid, and the N-1th intermediate three-dimensional grid is optimized according to the target depth features of the target images at each viewing angle of the target object to obtain an Nth intermediate three-dimensional grid; N is a positive integer greater than 1; The Nth intermediate three-dimensional grid is optimized according to the target depth features of the target images at each viewing angle of the target object to obtain the target three-dimensional grid.

9. A method for constructing a three-dimensional mesh generation model. It is characterized in that The method comprises: Acquire at least one sample data set, the sample data set comprising: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object; the sample object comprises a three-dimensional object; Construct the network structure of the 3D mesh generation model; Using the sample data set to train the network parameters in the three-dimensional mesh generation model until the network parameters converge to obtain the three-dimensional mesh generation model; Wherein, the network structure at least includes a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network; The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different viewing angles; The depth feature extraction network is used to respectively obtain sample depth features of sample images of each viewing angle of the sample object; The three-dimensional grid generation network includes: a three-dimensional grid generation layer and a three-dimensional grid optimization layer; The three-dimensional mesh generation layer is used to generate a to-be-optimized three-dimensional mesh of the sample object according to the sample voxels of the sample object; The three-dimensional mesh optimization layer is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images at various viewing angles of the sample object to obtain a predicted three-dimensional mesh.

10. The method according to claim 9, It is characterized in that The voxel generation network includes: a convolutional neural network, a voxel generation layer and a voxel fusion layer; The convolutional neural network is used to respectively obtain sample convolution features of sample images of each viewing angle of the sample object; The voxel generation layer is used to generate perspective voxels corresponding to each perspective of the sample object according to the sample convolution features of the sample image of each perspective of the sample object; The voxel fusion layer is used to fuse the view voxels of the sample object corresponding to each view angle to obtain the sample voxels of the sample object.

11. The method according to claim 9, It is characterized in that The three-dimensional grid optimization layer includes: a graph convolutional neural network and a contrast deep feature extraction network; The contrast depth feature extraction network is used to obtain sample depth difference information between the depth feature of the three-dimensional grid to be optimized and the sample depth features of the sample images of each viewing angle of the sample object; The graph convolutional neural network is used to optimize the three-dimensional grid to be optimized according to the sample depth difference information to obtain the predicted three-dimensional grid.

12. The method according to claim 11, It is characterized in that The contrastive depth feature extraction network includes: a neural renderer and a contrastive depth feature extractor; The neural renderer is used to obtain the depth features of each viewing angle of the three-dimensional mesh to be optimized; The contrast depth feature extractor is used to respectively obtain sample depth difference information between the depth feature of the three-dimensional grid and the sample depth feature of the sample image at the same viewing angle.

13. The method according to claim 9, It is characterized in that There are multiple three-dimensional grid optimization layers, and the multiple three-dimensional grid optimization layers are cascaded; The three-dimensional mesh optimization layer with the first arrangement order is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images of each viewing angle of the sample object; Among any two adjacent three-dimensional mesh optimization layers arranged in cascade, the three-dimensional mesh optimization layer arranged later is used to optimize the intermediate three-dimensional mesh output by the three-dimensional mesh optimization layer arranged earlier according to the sample depth features of the sample images of each viewing angle of the sample object; The three-dimensional grid optimization layer with the last arrangement order is used to optimize the intermediate three-dimensional grid output by the three-dimensional grid optimization layer with the second last arrangement order according to the sample depth features of the sample images of each viewing angle of the sample object to obtain the predicted three-dimensional grid.

14. The method according to claim 9, It is characterized in that The input end of the three-dimensional mesh generation model includes an input end of a voxel generation network and an input end of the deep feature extraction network; The output end of the voxel generation network is connected to the input end of the three-dimensional grid generation network; The output end of the deep feature extraction network is connected to the input end of the three-dimensional mesh generation network; The output end of the three-dimensional mesh generation model includes the output end of the three-dimensional mesh generation network.

15. The method according to claim 14, It is characterized in that The voxel generation network includes: a convolutional neural network, a voxel generation layer and a voxel fusion layer; The input end of the voxel generation network includes the convolutional neural network input end; The output end of the convolutional neural network is connected to the input end of the voxel generation layer; The output end of the voxel generation layer is connected to the input end of the voxel fusion layer; The output of the voxel generation network includes the output of the voxel fusion layer.

16. The method according to claim 14 or 15, It is characterized in that The three-dimensional grid generation network includes: a three-dimensional grid generation layer and a three-dimensional grid optimization layer; The input end of the three-dimensional grid generation layer is connected to the output end of the voxel generation layer; The output end of the three-dimensional grid generation layer is connected to the input end of the three-dimensional grid optimization layer; The input end of the three-dimensional mesh optimization layer is also connected to the output end of the deep feature extraction network; The output end of the three-dimensional mesh generation network includes the output end of the three-dimensional mesh optimization layer.

17. The method according to claim 16, It is characterized in that The three-dimensional grid optimization layer includes: a graph convolutional neural network and a contrast deep feature extraction network; The input end of the comparison depth feature extraction network is connected to the output end of the three-dimensional mesh generation network; The input end of the comparison depth feature extraction network is also connected to the output end of the depth feature extraction network; The output end of the comparison depth feature extraction network is connected to the input end of the graph convolutional neural network; The input end of the graph convolutional neural network is also connected to the output end of the three-dimensional grid generation network; The output end of the three-dimensional grid optimization layer includes the output end of the graph convolutional neural network.

18. The method according to claim 17, It is characterized in that The contrastive depth feature extraction network includes: a neural renderer and a contrastive depth feature extractor; The input end of the neural renderer is connected to the output end of the three-dimensional mesh generation layer; An output terminal of the neural renderer is connected to an input terminal of the contrast depth feature extractor; The input end of the comparison depth feature extractor is also connected to the output end of the depth feature extraction network; The output end of the comparative deep feature extraction network includes the output end of the deep feature extractor.

19. The method according to claim 16, It is characterized in that There are multiple three-dimensional grid optimization layers, and the multiple three-dimensional grid optimization layers are cascaded; Among any two adjacent three-dimensional grid optimization layers arranged in cascade, the output end of the three-dimensional grid optimization layer arranged in a forward order is connected to the input end of the three-dimensional grid optimization layer arranged in a backward order; The output end of the three-dimensional grid generation layer is connected to the input end of the three-dimensional grid optimization layer that is ranked first in the arrangement order; The input end of each three-dimensional mesh optimization layer is also connected to the output end of the deep feature extraction network respectively; The output end of the three-dimensional mesh generation network includes the output end of the three-dimensional mesh optimization layer which is the last in the arrangement order.

20. A three-dimensional grid generation device, It is characterized in that The device comprises: A first acquisition module is used to acquire target images of a target object at multiple different viewing angles; the target object includes a three-dimensional object; a second acquisition module, configured to acquire target voxels of the target object according to target images of multiple different perspectives of the target object, and a third acquisition module, configured to respectively acquire target depth features of the target images of each perspective of the target object; The fourth acquisition module includes: A second generating submodule, configured to generate a to-be-optimized three-dimensional mesh of the target object according to the target voxels of the target object; The optimization submodule is used to optimize the three-dimensional mesh to be optimized according to the target depth features of the target images of each viewing angle of the target object to obtain a target three-dimensional mesh.

21. The device according to claim 20, It is characterized in that The optimization submodule includes: An acquisition unit, configured to acquire target depth difference information between the depth feature of the three-dimensional mesh to be optimized and the target depth features of the target images of each viewing angle of the target object; The first optimization unit is used to optimize the three-dimensional mesh to be optimized according to the target depth difference information to obtain the target three-dimensional mesh.

22. A device for constructing a three-dimensional mesh generation model, It is characterized in that The device comprises: A fifth acquisition module is used to acquire at least one sample data set, wherein the sample data set includes: sample images of a sample object at multiple different viewing angles, and a labeled three-dimensional grid of the sample object; the sample object includes a three-dimensional object; A construction module is used to construct a network structure of a three-dimensional mesh generation model; A training module, used to train the network parameters in the three-dimensional mesh generation model using the sample data set until the network parameters converge to obtain the three-dimensional mesh generation model; Wherein, the network structure at least includes a voxel generation network, a deep feature extraction network and a three-dimensional mesh generation network; The voxel generation network is used to obtain sample voxels of the sample object according to sample images of the sample object at multiple different viewing angles; The depth feature extraction network is used to respectively obtain sample depth features of sample images of each viewing angle of the sample object; The three-dimensional grid generation network includes: a three-dimensional grid generation layer and a three-dimensional grid optimization layer; The three-dimensional mesh generation layer is used to generate a to-be-optimized three-dimensional mesh of the sample object according to the sample voxels of the sample object; The three-dimensional mesh optimization layer is used to optimize the three-dimensional mesh to be optimized according to the sample depth features of the sample images at various viewing angles of the sample object to obtain a predicted three-dimensional mesh.

23. The device according to claim 22, It is characterized in that The three-dimensional grid optimization layer includes: a graph convolutional neural network and a contrast deep feature extraction network; The contrast depth feature extraction network is used to obtain sample depth difference information between the depth feature of the three-dimensional grid to be optimized and the sample depth features of the sample images of each viewing angle of the sample object; The graph convolutional neural network is used to optimize the three-dimensional grid to be optimized according to the sample depth difference information to obtain the predicted three-dimensional grid.

24. An electronic device, It is characterized in that The electronic device comprises: processor; a memory for storing processor-executable instructions; The processor is configured to execute the three-dimensional mesh generation method as described in any one of claims 1-8.

25. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the three-dimensional mesh generation method as described in any one of claims 1-8.

26. An electronic device, It is characterized in that The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to execute the method for constructing a three-dimensional mesh generation model as described in any one of claims 9-19.

27. A non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method for constructing a three-dimensional mesh generation model as described in any one of claims 9 to 19.

Citation Information

Patent Citations

  • Selective surface mesh regeneration for 3-dimensional renderings

    US20160364907A1