Three-dimensional model construction method and device, electronic equipment and readable storage medium
By using semantic segmentation and feature extraction of single-viewpoint images, combined with depth estimation and occupancy probability inference, the problem of high cost and poor versatility of 3D clothing human body model reconstruction equipment is solved, and low-cost, high-precision 3D model construction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER
- Filing Date
- 2023-07-14
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for reconstructing 3D clothed human models require dense camera arrays, resulting in high equipment costs and poor versatility.
A 3D model is constructed by performing semantic segmentation and feature extraction on single-viewpoint images, combined with depth estimation and occupancy probability inference.
It reduces hardware requirements and improves the accuracy and versatility of 3D model construction.
Smart Images

Figure CN116863077B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method and apparatus for constructing three-dimensional models, an electronic device, and a computer-readable storage medium. Background Technology
[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.
[0003] In related technologies, 3D clothing human models are widely used in animation production, game entertainment, e-commerce, virtual reality and other fields. Traditional 3D clothing human model reconstruction methods require dense camera arrays to reconstruct the human body surface through multi-view geometry, which has high equipment costs and is difficult to apply in many scenarios.
[0004] Therefore, the methods for constructing 3D models of target objects, including the human body, in related technologies have high requirements for hardware equipment and poor versatility. Summary of the Invention
[0005] The purpose of this disclosure is to provide a method, apparatus, electronic device, and computer-readable storage medium for constructing three-dimensional models, which can accurately construct three-dimensional models of target objects using single-viewpoint images, with low requirements for hardware devices and good versatility.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0007] This disclosure provides a method for constructing a three-dimensional model, comprising: acquiring a two-dimensional image of a target object; performing semantic segmentation processing on the two-dimensional image of the target object using a semantic segmentation module to determine semantic segmentation features and a semantic segmentation map corresponding to the target object; performing feature extraction processing on the two-dimensional image of the target object using a feature extraction module to determine two-dimensional backbone features corresponding to the two-dimensional image of the target object; performing feature fusion processing on the semantic segmentation map, semantic segmentation features, and two-dimensional backbone features corresponding to the target object to determine fused two-dimensional features of the two-dimensional image of the target object; acquiring multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls within the two-dimensional image of the target object; determining the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object based on the fused two-dimensional features corresponding to the two-dimensional image of the target object; and performing occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point using an occupancy probability inference module to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object according to the spatial occupancy probability corresponding to each three-dimensional point.
[0008] In some embodiments, the method further includes: performing depth estimation processing on the two-dimensional image of the target object through a depth estimation module to determine the depth features and depth map of the two-dimensional image of the target object; wherein, performing feature fusion processing on the semantic segmentation map, the semantic segmentation features, and the two-dimensional backbone features to determine the fused two-dimensional features of the two-dimensional image of the target object includes: performing feature fusion processing on the semantic segmentation map, the semantic segmentation features, the depth map, the depth features, and the two-dimensional backbone features to determine the fused two-dimensional features of the two-dimensional image of the target object.
[0009] In some embodiments, the method further includes: performing depth estimation processing on the two-dimensional image of the target object using a depth estimation module to determine the depth map and depth features corresponding to the two-dimensional image; segmenting the depth map and depth features corresponding to the two-dimensional image of the target object using the semantic segmentation map to determine the depth map and depth features corresponding to the region where the target object is located; wherein, performing feature fusion processing on the semantic segmentation map, the semantic segmentation features, and the two-dimensional backbone features to determine the fused two-dimensional features of the two-dimensional image of the target object includes: fusing the semantic segmentation map, the semantic segmentation features, the depth map and depth features corresponding to the region where the target object is located, and the two-dimensional backbone features to determine the fused two-dimensional features of the two-dimensional image of the target object.
[0010] In some embodiments, the plurality of three-dimensional points include target three-dimensional points; wherein, by performing occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point through an occupancy probability inference module to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, the process includes: determining the distance of the target three-dimensional point from the plane where the two-dimensional image of the target object is located; using the distance of the target three-dimensional point from the plane where the two-dimensional image of the target object is located as the depth of the target three-dimensional point; and determining the spatial occupancy probability of the target three-dimensional point in the three-dimensional model of the target object based on the fused two-dimensional features of the projection point of the target three-dimensional point and the depth of the target three-dimensional point.
[0011] In some embodiments, constructing a three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point includes: extracting three-dimensional points whose spatial occupancy probability is equal to a target value from the plurality of three-dimensional points; generating a three-dimensional surface of the three-dimensional model of the target object based on the three-dimensional points whose spatial occupancy probability is equal to the target value; and constructing the three-dimensional model of the target object based on the three-dimensional surface.
[0012] In some embodiments, acquiring multiple three-dimensional points in a three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls within the two-dimensional image of the target object, includes: acquiring a preset height value; constructing the three-dimensional space based on the two-dimensional image of the target object and the preset height value; and performing uniform sampling or random sampling in the three-dimensional space to obtain the multiple three-dimensional points.
[0013] In some embodiments, the method further includes: acquiring a two-dimensional image of a training object and an actual three-dimensional model of the training object; performing semantic segmentation processing on the two-dimensional image of the training object through the semantic segmentation module to determine the semantic segmentation features and semantic segmentation map corresponding to the training object; performing feature extraction processing on the two-dimensional image of the training object through the feature extraction module to determine the two-dimensional backbone features corresponding to the two-dimensional image of the training object; performing feature fusion on the semantic segmentation map, semantic segmentation features, and two-dimensional backbone features of the training object to determine the fused two-dimensional features of the two-dimensional image of the training object; acquiring three-dimensional training points in the three-dimensional space, wherein the projection points of the three-dimensional training points on the plane where the two-dimensional image of the training object is located fall... In the two-dimensional image of the training object; based on the fused two-dimensional features corresponding to the two-dimensional image of the training object, determine the fused two-dimensional features corresponding to the projection points of the three-dimensional training points in the two-dimensional image of the training object; perform occupancy probability inference processing on the fused two-dimensional features of the three-dimensional training points through the occupancy probability inference module to determine the spatial occupancy probability of the three-dimensional training points in the predicted three-dimensional model of the training object; determine the actual positional relationship between the three-dimensional training points and the actual three-dimensional model of the training object; determine the loss value based on the actual positional relationship between the three-dimensional training points and the actual three-dimensional model and the spatial occupancy probability corresponding to the three-dimensional training points; train the parameter values of the feature extraction module and the occupancy probability inference module through the loss value.
[0014] This disclosure provides a three-dimensional model construction device, including: a two-dimensional image acquisition module, a semantic segmentation module, a two-dimensional backbone feature extraction module, a feature fusion module, a three-dimensional point acquisition module, a fused two-dimensional feature acquisition module, and an occupancy probability inference module.
[0015] The two-dimensional image acquisition module is used to acquire a two-dimensional image of the target object; the semantic segmentation module can be used to perform semantic segmentation processing on the two-dimensional image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object; the two-dimensional backbone feature extraction module can be used to perform feature extraction processing on the two-dimensional image of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object; the feature fusion module can be used to perform feature fusion processing on the semantic segmentation map, semantic segmentation features, and two-dimensional backbone features corresponding to the target object to determine the fused two-dimensional features of the two-dimensional image of the target object; the three-dimensional point acquisition module can... The system is designed to acquire multiple 3D points in 3D space, wherein the projection point of each 3D point onto the plane of the 2D image of the target object falls within the 2D image of the target object. The fused 2D feature acquisition module can be used to determine the fused 2D features corresponding to the projection points of each 3D point in the 2D image of the target object based on the fused 2D features corresponding to the 2D image of the target object. The occupancy probability inference module can be used to perform occupancy probability inference processing on the fused 2D features corresponding to each 3D point through the occupancy probability inference module to determine the spatial occupancy probability of each 3D point in the 3D model of the target object, so as to construct the 3D model of the target object based on the spatial occupancy probability corresponding to each 3D point.
[0016] This disclosure provides an electronic device comprising: a memory and a processor; the memory for storing computer program instructions; and the processor for calling the computer program instructions stored in the memory to implement the three-dimensional model construction method described above.
[0017] This disclosure provides a computer-readable storage medium storing computer program instructions to implement the three-dimensional model construction method as described in any of the preceding embodiments.
[0018] This disclosure provides a computer program product or computer program that includes computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and the processor executes the computer program instructions to implement the aforementioned three-dimensional model construction method.
[0019] The three-dimensional model construction method, apparatus, electronic device, and computer-readable storage medium provided in this disclosure can enhance the recognition ability of the two-dimensional backbone of the two-dimensional image of the target object through semantic segmentation features of the two-dimensional image, thereby improving the recognition ability of the two-dimensional backbone features by combining semantic segmentation features, and thus improving the accuracy of the three-dimensional model construction of the target object.
[0020] This application proposes a semantic segmentation-based depth-aware 3D model reconstruction method. It integrates 2D backbone features with semantic segmentation features before performing 3D point occupancy probability inference, which can improve the representational ability of pixel features during the occupancy probability inference process and improve the accuracy and completeness of 3D model reconstruction.
[0021] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this disclosure. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0023] Figure 1 A schematic diagram of a scenario that can be applied to the three-dimensional model building method or three-dimensional model building apparatus of the present disclosure is shown.
[0024] Figure 2 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment.
[0025] Figure 3 This is a schematic diagram illustrating an occupancy probability reasoning process according to an exemplary embodiment.
[0026] Figure 4 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment.
[0027] Figure 5 This is a schematic diagram illustrating an occupancy probability reasoning process according to an exemplary embodiment.
[0028] Figure 6 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment.
[0029] Figure 7 This is a flowchart illustrating a method for determining space occupancy probability according to an exemplary embodiment.
[0030] Figure 8 This is a flowchart illustrating a deep network model training method according to an exemplary embodiment.
[0031] Figure 9 This is a flowchart illustrating a network model training method according to an exemplary embodiment.
[0032] Figure 10This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment.
[0033] Figure 11 This is a block diagram illustrating a three-dimensional model building apparatus according to an exemplary embodiment.
[0034] Figure 12 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0035] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0036] Those skilled in the art will recognize that embodiments of this disclosure can be a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0037] The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more specific details omitted, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0038] The accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus omitting repeated descriptions of them. Some block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0039] The flowchart shown in the accompanying drawings is merely illustrative and does not necessarily include all content and steps, nor does it require execution in the described order. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0040] In the description of this disclosure, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences; the terms "contains," "includes," and "has" are used to indicate an open-ended meaning of inclusion and refer to the existence of additional elements / components / etc. besides those listed.
[0041] To better understand the above-mentioned objectives, features and advantages of this application, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.
[0042] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0043] This application proposes a method for reconstructing 3D models (such as 3D dressed human models) from a single RGB image. In related technologies, deep learning-based methods for reconstructing dressed human figures, which regress a 3D human representation from an RGB image, can be mainly divided into parametric and non-parametric methods. This application proposes a depth-aware method for reconstructing dressed human figures. This method explicitly estimates the human body depth map to assist in the reconstruction of the human body surface, improving the expressive power of image features. This application also introduces a multi-structure segmentation module for the human body, endowing the 3D surface inference module with more information about the semantic structure of the human body. By explicitly extracting depth features and multi-structure semantic segmentation features from the RGB image, this application improves the representational power of pixel alignment features and enhances the reconstruction effect.
[0044] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0045] Figure 1 A schematic diagram of a scenario that can be applied to the three-dimensional model building method or three-dimensional model building apparatus of the present disclosure is shown.
[0046] Please refer to Figure 1The diagram illustrates an implementation environment provided by an exemplary embodiment of this disclosure.
[0047] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0048] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, desktop computers, wearable devices, virtual reality devices, smart home devices, etc.
[0049] Server 105 can be a server that provides various services, such as a backend management server that supports the devices operated by users using terminal devices 101, 102, and 103. The backend management server can analyze and process received requests and other data, and then feed the processing results back to the terminal devices.
[0050] A server can be a standalone physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This disclosure does not impose any restrictions on this.
[0051] Server 105 may, for example, acquire a two-dimensional image of the target object; server 105 may, for example, perform semantic segmentation processing on the two-dimensional image of the target object through a semantic segmentation module to determine the semantic segmentation features and semantic segmentation map corresponding to the target object; server 105 may, for example, perform feature extraction processing on the two-dimensional image of the target object through a feature extraction module to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object; server 105 may, for example, perform feature fusion processing on the semantic segmentation map, semantic segmentation features, and two-dimensional backbone features corresponding to the target object to determine the fused two-dimensional features of the two-dimensional image of the target object; server 105 may, for example, acquire multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls within the two-dimensional image of the target object; server 105 may, for example, determine the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object based on the fused two-dimensional features corresponding to the two-dimensional image of the target object; server 105 may, for example, perform occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point through an occupancy probability inference module to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point.
[0052] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Server 105 can be a single physical server or a combination of multiple servers. Depending on actual needs, it can have any number of terminal devices, networks, and servers.
[0053] Figure 2 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment. The method provided in this disclosure can be executed by any electronic device with computing power, for example, the method can be executed by the above-described... Figure 1 The execution can be performed by a server or terminal device in the embodiments, or it can be performed by both a server and a terminal device. In the following embodiments, the server is used as the execution subject for illustration, but this disclosure is not limited to this.
[0054] Reference Figure 2 The three-dimensional model construction method provided in this disclosure may include the following steps.
[0055] Step S202: Obtain a two-dimensional image of the target object.
[0056] The target object mentioned above can be any object that needs to be modeled in three dimensions, such as the human body, animals, or buildings. This application does not impose any restrictions on this.
[0057] Two-dimensional images can refer to RGB images acquired through a single viewpoint, etc., and this application does not limit them.
[0058] Step S204: The two-dimensional image of the target object is semantically segmented by the semantic segmentation module to determine the semantic segmentation features and semantic segmentation map corresponding to the target object.
[0059] In some embodiments, the semantic segmentation module can be used to extract the contour of the target object from a two-dimensional image; the semantic segmentation module can be used to segment and extract the limbs of the target object from a two-dimensional image. This application does not limit the semantic segmentation results, and those skilled in the art can set the semantic segmentation purpose according to actual needs.
[0060] In some embodiments, the semantic segmentation module can extract semantic features from a two-dimensional image to obtain semantic segmentation features, and then use a classifier (such as softmax) to process the semantic segmentation features to obtain a semantic segmentation map.
[0061] like Figure 3 As shown, the semantic segmentation module Mseg can perform semantic segmentation processing on the two-dimensional image of the input target object to obtain semantic segmentation features Gseg and semantic segmentation map S.
[0062] In some embodiments, the semantic segmentation module described above can be trained independently. For example, loss can be calculated using the actual segmentation map of the target object and the predicted segmentation map predicted from the two-dimensional image, and then the parameters of the semantic segmentation module can be adjusted separately.
[0063] Step S206: The feature extraction module performs feature extraction processing on the two-dimensional image of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object.
[0064] In some embodiments, a feature extraction module can be used to perform feature extraction processing on the two-dimensional image of the target object to obtain two-dimensional backbone features that can describe the macroscopic overall features of the two-dimensional image.
[0065] Step S208: Perform feature fusion processing on the semantic segmentation map, semantic segmentation features and two-dimensional backbone features corresponding to the target object to determine the fused two-dimensional features of the two-dimensional image of the target object.
[0066] In some embodiments, such as Figure 3 As shown, by performing feature fusion processing on the semantic segmentation map, semantic segmentation features, and two-dimensional backbone features, a fused two-dimensional feature G can be obtained.
[0067] Step S210: Obtain multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls in the two-dimensional image of the target object.
[0068] In some embodiments, a three-dimensional space can be constructed by: obtaining a preset height value; then constructing a three-dimensional space based on a two-dimensional image of the target object and the preset height value; and finally performing uniform or random sampling in the three-dimensional space to obtain multiple three-dimensional points. It is understood that those skilled in the art can obtain the aforementioned three-dimensional space in other ways, and this application does not limit such methods.
[0069] It is understood that any three-dimensional space composed of three-dimensional points in a two-dimensional image of the target object, where the projection points fall, can be the three-dimensional space in this application.
[0070] like Figure 3 As shown, a three-dimensional point X obtained from three-dimensional space can include x and z, where x represents the coordinates of the projection point of the three-dimensional point onto the plane of the two-dimensional image, and z represents the distance (i.e., depth) of the three-dimensional point relative to the plane of the two-dimensional image.
[0071] Step S212: Based on the fused two-dimensional features corresponding to the two-dimensional image of the target object, determine the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object.
[0072] In some embodiments, given the fused two-dimensional features corresponding to each pixel in a two-dimensional image, then the fused two-dimensional features corresponding to the projection points of the three-dimensional points onto the plane of the two-dimensional image (such as...) Figure 3 The pixel alignment feature Gx in the image is also easy to determine, and this application does not impose any restrictions on it.
[0073] Step S214: The occupancy probability reasoning module performs occupancy probability reasoning on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point.
[0074] In some embodiments, probability inference modules (such as occupancy probability inference modules) can be used. Figure 3 The occupancy probability inference module (Mocc303) performs occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object.
[0075] Among them, occupancy probability inference processing can refer to calculating the probability that the three-dimensional point is located in the three-dimensional model.
[0076] In some embodiments, if a 3D point is located in the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 1; if a 3D point is not located in the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 0; if a 3D point is located on the surface of the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 0.5, and so on.
[0077] In some embodiments, constructing a three-dimensional model of a target object based on the spatial occupancy probability corresponding to each three-dimensional point may include the following steps: extracting three-dimensional points with a spatial occupancy probability equal to a target value from multiple three-dimensional points; generating a three-dimensional surface of the three-dimensional model of the target object based on the three-dimensional points with a spatial occupancy probability equal to the target value; and constructing a three-dimensional model of the target object based on the three-dimensional surface.
[0078] The above embodiments can enhance the recognition ability of the two-dimensional backbone of the target object's two-dimensional image by using semantic segmentation features of the two-dimensional image, thereby improving the recognition ability of the two-dimensional backbone features by combining semantic segmentation features, and thus improving the accuracy of constructing the three-dimensional model of the target object.
[0079] Figure 4 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment. The method provided in this disclosure can be executed by any electronic device with computing power, for example, the method can be executed by the above-described... Figure 1 The execution can be performed by a server or terminal device in the embodiments, or it can be performed by both a server and a terminal device. In the following embodiments, the server is used as the execution subject for illustration, but this disclosure is not limited to this.
[0080] Reference Figure 4 The three-dimensional model construction method provided in this disclosure may include the following steps.
[0081] Step S402: Obtain a two-dimensional image of the target object.
[0082] like Figure 5 As shown, a two-dimensional image I of the target object can be obtained.
[0083] Step S404: The semantic segmentation module performs semantic segmentation processing on the two-dimensional image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object.
[0084] like Figure 5 As shown, the semantic segmentation module 501 can perform feature extraction processing on the two-dimensional image I of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object.
[0085] Step S406: The feature extraction module performs feature extraction processing on the two-dimensional image of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object.
[0086] like Figure 5 As shown, the feature extraction module 502 can perform feature extraction processing on the two-dimensional image I of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object.
[0087] Step S408: The depth estimation module performs depth estimation processing on the two-dimensional image of the target object to determine the depth features and depth map of the two-dimensional image of the target object.
[0088] like Figure 5 As shown, the depth estimation module 503 can perform depth estimation processing on the two-dimensional image I of the target object to determine the depth features and depth map of the two-dimensional image of the target object.
[0089] In some embodiments, the depth estimation module can be trained independently. For example, loss can be calculated using the actual depth map of the target object and the predicted depth map predicted from the two-dimensional image, and then the parameters of the depth estimation module can be adjusted separately.
[0090] Step S410 involves performing feature fusion processing on the semantic segmentation map, semantic segmentation features, depth map, depth features, and two-dimensional backbone features to determine the fused two-dimensional features of the target object's two-dimensional image.
[0091] like Figure 5 As shown, the semantic segmentation graph G can be... seg Deep features G depth and two-dimensional backbone features G f Features are obtained through fusion processing, and then fused with depth map D and semantic segmentation map S to determine the fused two-dimensional features G of the two-dimensional image of the target object.
[0092] Step S412: Obtain multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls in the two-dimensional image of the target object.
[0093] In some embodiments, a three-dimensional space can be constructed by: obtaining a preset height value; then constructing a three-dimensional space based on a two-dimensional image of the target object and the preset height value; and finally performing uniform or random sampling in the three-dimensional space to obtain multiple three-dimensional points. It is understood that those skilled in the art can obtain the aforementioned three-dimensional space in other ways, and this application does not limit such methods.
[0094] like Figure 3As shown, obtaining a 3D point X from 3D space can include x and z, where x represents the coordinates of the projection point of the 3D point onto the plane of the 2D image, and z represents the distance of the 3D point relative to the plane of the 2D image.
[0095] Step S414: Based on the fused two-dimensional features corresponding to the two-dimensional image of the target object, determine the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object.
[0096] In some embodiments, given the fused two-dimensional features corresponding to each pixel in a two-dimensional image, then the fused two-dimensional features corresponding to the projection points of the three-dimensional points onto the plane of the two-dimensional image (such as...) Figure 5 Pixel alignment feature G in x It is also easy to determine, and this application does not impose any restrictions on it.
[0097] Step S416: The occupancy probability reasoning module performs occupancy probability reasoning on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point.
[0098] In some embodiments, probability inference modules (such as occupancy probability inference modules) can be used. Figure 5 In step 504, the occupancy probability inference process is performed on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object.
[0099] Among them, the occupancy probability inference processing refers to calculating the probability that the three-dimensional point is located in the three-dimensional model.
[0100] In some embodiments, if a 3D point is located in the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 1; if a 3D point is not located in the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 0; if a 3D point is located on the surface of the 3D model of the target object, then the spatial occupancy probability of the 3D point in the 3D model of the target object can be 0.5, and so on.
[0101] In some embodiments, constructing a three-dimensional model of a target object based on the spatial occupancy probability corresponding to each three-dimensional point may include the following steps: extracting three-dimensional points with a spatial occupancy probability equal to a target value from multiple three-dimensional points; generating a three-dimensional surface of the three-dimensional model of the target object based on the three-dimensional points with a spatial occupancy probability equal to the target value; and constructing a three-dimensional model of the target object based on the three-dimensional surface.
[0102] The above embodiments can enhance the recognition ability of the two-dimensional backbone of the target object's two-dimensional image by using semantic segmentation features of the two-dimensional image, thereby improving the recognition ability of the two-dimensional backbone features by combining semantic segmentation features, and thus improving the accuracy of constructing the three-dimensional model of the target object.
[0103] This embodiment proposes a semantic segmentation depth-aware 3D model reconstruction method. By introducing a depth-aware module into the 3D model reconstruction framework, the representation capability of pixel features is improved, thereby enhancing the accuracy and completeness of 3D model reconstruction.
[0104] Figure 6 This is a flowchart illustrating a three-dimensional model construction method according to an exemplary embodiment. The method provided in this disclosure can be executed by any electronic device with computing power, for example, the method can be executed by the above-described... Figure 1 The execution can be performed by a server or terminal device in the embodiments, or it can be performed by both a server and a terminal device. In the following embodiments, the server is used as the execution subject for illustration, but this disclosure is not limited to this.
[0105] Reference Figure 6 The three-dimensional model construction method provided in this disclosure may include the following steps.
[0106] Step S602: Obtain a two-dimensional image of the target object.
[0107] Step S604: The semantic segmentation module performs semantic segmentation processing on the two-dimensional image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object.
[0108] Step S606: The feature extraction module performs feature extraction processing on the two-dimensional image of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object.
[0109] Step S608: The depth estimation module performs depth estimation processing on the two-dimensional image of the target object to determine the depth map and depth features corresponding to the two-dimensional image.
[0110] Step S610: The depth map and depth features corresponding to the two-dimensional image of the target object are segmented using the semantic segmentation map to determine the depth map and depth features corresponding to the region where the target object is located.
[0111] Step S612: The semantic segmentation map, semantic segmentation features, depth map and depth features corresponding to the region where the target object is located, and two-dimensional backbone features are fused to determine the fused two-dimensional features of the two-dimensional image of the target object.
[0112] Step S614: Obtain multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls in the two-dimensional image of the target object.
[0113] Step S616: Based on the fused two-dimensional features corresponding to the two-dimensional image of the target object, determine the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object.
[0114] Step S618: The occupancy probability reasoning module performs occupancy probability reasoning on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point.
[0115] In the technical solution provided in the above embodiments, when predicting the occupancy probability of three-dimensional points, not only two-dimensional image features are considered, but also the segmentation features corresponding to the outline of the target object, as well as the depth map and depth features corresponding to the area where the target object is located. This improves the representation capability of pixel features and enhances the accuracy and completeness of three-dimensional model reconstruction.
[0116] Figure 7 This is a flowchart illustrating a method for determining space occupancy probability according to an exemplary embodiment.
[0117] In some embodiments, multiple three-dimensional points may include target three-dimensional points. In this embodiment, the determination of spatial occupancy probability will be explained using the target three-dimensional point as an example.
[0118] refer to Figure 7 The above method for determining the probability of space occupancy may include the following steps.
[0119] Step S702: Determine the distance between the target 3D point and the plane containing the 2D image of the target object.
[0120] Step S704: The distance between the target 3D point and the plane containing the 2D image of the target object is taken as the depth of the target 3D point.
[0121] Step S706: Based on the fused two-dimensional features of the projection points of the target three-dimensional points and the depth of the target three-dimensional points, determine the spatial occupancy probability of the target three-dimensional points in the three-dimensional model of the target object.
[0122] The above method can accurately predict the spatial occupancy probability of a target 3D point in the 3D model of the target object.
[0123] Figure 8 This is a flowchart illustrating a deep network model training method according to an exemplary embodiment.
[0124] refer to Figure 8 The above-mentioned deep network model training method may include the following steps.
[0125] Step S802: Obtain the two-dimensional image of the training object and the actual three-dimensional model of the training object.
[0126] Step S804: The semantic segmentation module performs semantic segmentation processing on the two-dimensional image of the training object to determine the semantic segmentation features and semantic segmentation map corresponding to the training object.
[0127] In some embodiments, the semantic segmentation module described above may have been trained separately based on the actual semantic segmentation map.
[0128] Step S806: The feature extraction module performs feature extraction processing on the two-dimensional image of the training object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the training object.
[0129] Step S808: The semantic segmentation map, semantic segmentation features and two-dimensional backbone features of the training object are fused to determine the fused two-dimensional features of the two-dimensional image of the training object.
[0130] Step S810: Obtain three-dimensional training points in three-dimensional space, wherein the projection points of the three-dimensional training points on the plane where the two-dimensional image of the training object is located fall in the two-dimensional image of the training object.
[0131] Step S812: Based on the fused two-dimensional features corresponding to the two-dimensional image of the training object, determine the fused two-dimensional features corresponding to the projection points of the three-dimensional training points in the two-dimensional image of the training object.
[0132] Step S814: The occupancy probability inference module is used to perform occupancy probability inference processing on the fused two-dimensional features of the three-dimensional training points to determine the spatial occupancy probability of the three-dimensional training points in the predicted three-dimensional model of the training object.
[0133] Step S816: Determine the actual positional relationship between the 3D training points and the actual 3D model of the training object.
[0134] Step S818: Determine the loss value based on the actual positional relationship between the 3D training points and the actual 3D model and the spatial occupancy probability corresponding to the 3D training points.
[0135] Step S820: Train the parameter values of the feature extraction module and the occupancy probability inference module using the loss value.
[0136] In this embodiment, all parameters except those of the semantic segmentation module can be adjusted.
[0137] The above embodiments can be used for, for example Figure 3The neural network model shown is trained to more accurately construct a 3D model of the target object.
[0138] It is understood that those skilled in the art can make adjustments based on the above embodiments. Figure 5 The neural network model corresponding to the illustrated embodiment is trained, wherein... Figure 5 The semantic segmentation module and depth estimation module can both be independent and pre-trained.
[0139] In some embodiments, the target object may be a human body (such as a human body to be dressed in an animation or game). This application will use the human body as an example for illustrative purposes.
[0140] In summary, the deep learning-based method for 3D reconstruction of a clothed human body proposed in this embodiment can include three stages: feature extraction, spatial occupancy estimation, and surface reconstruction. Figure 5 The method framework diagram for the feature extraction stage and the spatial occupancy estimation stage is shown. Figure 5 As shown, the feature extraction stage may include a feature extraction module M. f Depth estimation module M depth Human body structure semantic segmentation module M seg For the input image I, the depth estimation module M seg Estimate the depth values of the human body region in the image to obtain the depth map D and intermediate features G. depth Human semantic segmentation module M seg Output human body multi-structure segmentation image S and intermediate features G seg Feature extraction backbone network M f Extracting backbone features G for 3D human reconstruction f The final feature G obtained in the feature extraction stage is composed of (G... f G depth G seg Composed of ,D,S).
[0141] In the space occupancy estimation stage, the occupancy probability inference module M occ Estimating spatial location using feature G Let P(X) be the probability that a location X is located inside the surface of a 3D human body model. The probability value of P(X) is in the range of [0,1], where 1 indicates that the spatial location X is inside the 3D model and 0 indicates that the spatial location X is outside the 3D model.
[0142] In the surface reconstruction stage, isosurfaces with a value of 0.5 are extracted from the uniformly sampled location {X} in space to obtain the 3D reconstructed surface of the clothed human body. The following embodiment will explain how to construct each module:
[0143] Feature extraction module M fDepth estimation module M depth Human body structure semantic segmentation module M seg With a similar convolutional neural network structure, similar to U-net, it features a convolutional encoder and decoder. The network can have five spatial resolution layers. The encoder can contain four convolutional sub-modules with residual connections, each containing four 3×3 convolutional layers, followed by a 2×2 average pooling downsampling layer. The decoder contains four convolutional sub-modules with residual connections and one convolutional layer, followed by a ×2 linear upsampling layer. Residual connections exist between the encoder and decoder between feature maps of the same size.
[0144] Possession Probability Reasoning Module M occ In the spatial occupancy probability reasoning stage, spatial location Coordinates projected onto a two-dimensional image and depth Composition, Probabilistic Reasoning Module M occ Pixel alignment features at position x And depth z, to estimate the probability that position X lies inside the surface of a 3D shape. Probabilistic inference module M occ It consists of fully connected layers.
[0145] The flowchart of the training method involved in this embodiment is attached. Figure 9 As shown.
[0146] The learnable modules in the clothing-based human reconstruction method framework of this application are trained through the following steps:
[0147] S1-1: Training depth estimation module M depth .
[0148] Obtain an RGB human image dataset and its corresponding depth map, and train M in a supervised manner. depth M depth The parameters of the convolutional neural network.
[0149] S1-2: Training the human body segmentation module M seg .
[0150] Obtain the RGB human image dataset and the corresponding human multi-structure semantic segmentation map, and train M in a supervised manner. seg M seg The parameters of the convolutional neural network.
[0151] S1-3: Training Feature Extraction Module M f and the probability of possession reasoning module M occ :
[0152] Fixed module M depand M seg The neural network weights are constructed as shown in the attached figure. Figure 1 The framework shown is trained to obtain the feature extraction module M. f and the probability of possession reasoning module M occ Network parameters.
[0153] The flowchart of the method in the practical application stage involved in this embodiment is attached. Figure 10 As shown.
[0154] S2-1: For the input RGB image of a clothed human body from a single viewpoint, extract image features G using the depth estimation module, human body segmentation module, and feature extraction module.
[0155] For an input single-viewpoint human RGB image i, the module M trained using S1-1 is used. depth Estimate the depth map D and obtain the depth-related features G dept The module M obtained by training using S1-2 seg Estimate the segmentation map S and obtain semantically relevant features G seg The module M obtained by training using S1-3 f Extracting image features G f The final features obtained By (G) f G depth G seg It consists of (D, S).
[0156] S2-2: For a spatial location X(x, z), its pixel alignment feature is Gx. The probability P(X) that the spatial location X is located inside the 3D model is inferred using the occupancy probability inference module Mocc.
[0157] Uniform sampling is performed on the positions in three-dimensional space to obtain a set of position points {X}. For each spatial position... use Obtain the spatial alignment feature at position x Using the module M obtained from steps S1-3 occ Input features Given the depth z, estimate the probability P(X) that the location X is inside the 3D shape.
[0158] S2-3: Using Marching Cube technology, isosurfaces are extracted to obtain the final 3D reconstruction result of the clothed human body.
[0159] For the set {(X, P(X))} obtained in step S2-2, the contour surface with a value of 0.5 is extracted using the Marching Cube technique to obtain the final 3D reconstructed surface of the clothed human body.
[0160] The above embodiments utilize pixel alignment features extracted by convolutional neural networks to estimate the probability that a position in space lies within a shape. In this process, the representational and descriptive capabilities of the features are crucial; however, existing technologies lack sufficient feature extraction capabilities for RGB image inputs, leading to incomplete and inaccurate 3D human reconstruction results. This embodiment introduces additional depth estimation and semantic segmentation modules, providing rich geometrically relevant, depth-related, and semantically relevant information for the pixel alignment features, effectively improving the expressive and descriptive capabilities of the features and enhancing the accuracy of human reconstruction. Using this embodiment, a high-quality 3D model of a clothed human body can be reconstructed from a single-viewpoint RGB image.
[0161] To make the objectives, technical solutions, and advantages of this embodiment clearer, the method for reconstructing a clothed human body based on a single-viewpoint RGB image proposed in this application will be further described in detail below. It should be understood that the specific implementation methods described herein are only for explaining this application and are not intended to limit the scope of protection of this application. The described embodiments are some, but not all, of the embodiments in this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0162] S1-1 can specifically include the following:
[0163] Construct depth estimation module M depth This module has a convolutional encoder and a decoder. The network has five spatial resolution layers. The encoder contains four convolutional sub-modules with residual connections. Each sub-module contains four 3×3 convolutional layers, followed by a 2×2 average pooling downsampling layer. The decoder contains four convolutional sub-modules with residual connections and one convolutional layer, followed by a ×2 linear upsampling layer. There are residual connections between the encoder and decoder between feature maps of the same size. RGB human image datasets and corresponding depth maps are obtained, and module M is trained in a supervised manner. dept Training is performed, and the loss function L is used. depth for:
[0164] L depth (D,D gt ) = H(DD) gt ,α) (1)
[0165]
[0166] Where D is the depth map predicted by the convolutional neural network, D gt Let H(x, α) be the ground truth of the depth map, and let H(x, α) be the Huber loss function, with α taking the value of 0.1.
[0167] S1-2 may include the following:
[0168] Human body segmentation module M seg Network structure and M depth Similarly, we obtain RGB human image datasets and corresponding human multi-structure semantic segmentation maps, and train M in a supervised manner. seg M seg The parameters of the convolutional neural network and the loss function are as follows:
[0169] L seg (S, S) gt )=CE(S,S gt (3)
[0170] Where S is the segmentation map predicted by the convolutional neural network, S gt represents the ground truth for semantic segmentation, and CE is the cross-entropy loss function.
[0171] S1-3 may include the following:
[0172] Fixed module M dep and M seg The neural network weights are constructed as shown in the attached figure. Figure 1 The frame shown, M f Network structure M depth Similarly, M occ The module consists of fully connected layers, and by training it, a feature extraction module M is obtained. f and the probability of possession reasoning module M occ The network parameters are as follows. The input image is I, the estimated spatial location X is P(X), and the loss function is as follows:
[0173] L occ (P(X), O(X)) = |P(X) - O(X)| 2 (4)
[0174] Where P(X) is module M occ The output estimate, O(X), represents the space occupancy of the true value of the 3D model. That is, when X is inside the shape, O(X) is 1, otherwise it is 0.
[0175] Practical application stage.
[0176] S2-1 may include the following:
[0177] For an input single-viewpoint human RGB image I, the module M trained using S1-1 is used. dep Estimate the depth map D and obtain the depth-related features G dept The module M obtained by training using S1-2 segEstimate the segmentation map S and obtain semantically relevant features G seg The module M obtained by training using S1-3 f Extracting image features G f The final features obtained By (G) f G dept G seg It consists of (D, S).
[0178] S2-2 may include the following:
[0179] Uniform sampling is performed on the positions in three-dimensional space to obtain a set of position points {X}. For each spatial position... use Obtain the spatial alignment feature at position x Using the module M obtained from steps S1-3 occ Input features Given the depth z, estimate the probability P(X) that the location X is inside the 3D shape.
[0180] S2-3 may include the following:
[0181] For the set {(X, P(X))} obtained in step S2-2, the contour surface with a value of 0.5 is extracted using Marching Cube technology (a computer graphics algorithm for triangulating isosurfaces represented in grid form) to obtain the final 3D reconstructed surface of the clothed human body.
[0182] The method for reconstructing clothing from a single-viewpoint RGB image proposed in this embodiment has at least the following distinguishing technical features.
[0183] 1. Enhancement of depth-related features and semantic-related features in 3D reconstruction of the human body using implicit clothing.
[0184] Implicit 3D reconstruction methods for dressed humans in related technologies are insufficient in extracting geometric, depth, and semantic information from the human body surface, leading to inaccuracies and incompleteness in human body surface estimation. The technical solution provided in this application introduces an additional convolutional neural network pathway to explicitly extract depth-related and semantic-related features from the input RGB human body image, providing rich feature information for the surface estimation of dressed humans. This application improves the accuracy of 3D reconstruction of dressed humans by enhancing the descriptive and expressive capabilities of the extracted features.
[0185] 2. Training of the multi-path feature extraction network and the human body 3D reconstruction network
[0186] To effectively train the multi-path feature-enhanced clothing human reconstruction framework proposed in this application, this application proposes a training strategy, which first trains the depth estimation module M... depth and human body segmentation module M seg The system was trained independently, and then the depth estimation module and the human segmentation module were integrated into the human 3D reconstruction backbone network. The parameters of the two modules were fixed, and the framework was trained to optimize the feature extraction module M. f and the probability of possession reasoning module M occ Learnable weights. In practical applications, the trained M... depth M seg M f and M occ The constructed framework for reconstructing a clothed human body estimates the probability that the three-dimensional spatial location X is located inside the clothed human body.
[0187] The purpose of this embodiment is to improve the feature representation capability in implicit clothing human reconstruction. By introducing a human depth estimation module and a human multi-structure semantic segmentation module into the method framework, depth and semantic information are explicitly extracted, enhancing the feature representation capability in clothing human surface inference. The method framework of this embodiment includes a feature extraction backbone module, a depth estimation module, and a human semantic segmentation module. It can extract rich structural, depth, and semantic information from RGB images, construct pixel-aligned features, and use these features to infer 3D shape surfaces. The method provided in this application improves the accuracy of reconstructing 3D models of clothing humans from single-viewpoint RGB images.
[0188] This application proposes a semantic segmentation-based depth-aware method for reconstructing dressed human bodies. By introducing a depth-aware module into the implicit dressed human body reconstruction framework, the representation capability of pixel features is improved, thereby enhancing the accuracy and completeness of the 3D dressed human body reconstruction.
[0189] It should be particularly noted that the steps in each embodiment of the above-described 3D model construction method can be interchanged, substituted, added to, or deleted from each other. Therefore, these reasonable permutations and combinations of the 3D model construction method should also fall within the protection scope of this disclosure, and the protection scope of this disclosure should not be limited to the described embodiments.
[0190] Based on the same inventive concept, this disclosure also provides a three-dimensional model construction device, as shown in the following embodiment. Since the principle by which this device solves the problem is similar to that of the method embodiment described above, the implementation of this device embodiment can refer to the implementation of the method embodiment described above, and repeated details will not be elaborated further.
[0191] Figure 11 This is a block diagram illustrating a three-dimensional model building apparatus according to an exemplary embodiment. (Refer to...) Figure 11The three-dimensional model construction device 1100 provided in this embodiment may include: a two-dimensional image acquisition module 1101, a semantic segmentation module 1102, a two-dimensional backbone feature extraction module 1103, a feature fusion module 1104, a three-dimensional point acquisition module 1105, a fused two-dimensional feature acquisition module 1106, and an occupancy probability reasoning module 1107.
[0192] The system includes: a 2D image acquisition module 1101 for acquiring a 2D image of a target object; a semantic segmentation module 1102 for performing semantic segmentation on the 2D image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object; a 2D backbone feature extraction module 1103 for performing feature extraction on the 2D image of the target object to determine the 2D backbone features corresponding to the 2D image of the target object; a feature fusion module 1104 for performing feature fusion on the semantic segmentation map, semantic segmentation features, and 2D backbone features corresponding to the target object to determine the fused 2D features of the 2D image of the target object; and a 3D point acquisition module 1. 105 can be used to acquire multiple 3D points in 3D space, wherein the projection point of each 3D point on the plane of the 2D image of the target object falls in the 2D image of the target object; the fusion 2D feature acquisition module 1106 can be used to determine the fusion 2D features corresponding to the projection points of each 3D point in the 2D image of the target object based on the fusion 2D features corresponding to the 2D image of the target object; the occupancy probability reasoning module 1107 can be used to perform occupancy probability reasoning processing on the fusion 2D features corresponding to each 3D point through the occupancy probability reasoning module to determine the spatial occupancy probability of each 3D point in the 3D model of the target object, so as to construct the 3D model of the target object according to the spatial occupancy probability corresponding to each 3D point.
[0193] It should be noted that the aforementioned two-dimensional image acquisition module 1101, semantic segmentation module 1102, two-dimensional backbone feature extraction module 1103, feature fusion module 1104, three-dimensional point acquisition module 1105, fused two-dimensional feature acquisition module 1106, and occupancy probability inference module 1107 correspond to S202 to S214 in the method embodiment. The examples and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content disclosed in the above method embodiment. It should be noted that these modules, as part of the apparatus, can be executed in a computer system such as a set of computer-executable instructions.
[0194] In some embodiments, the 3D model building apparatus 1101 may further include a depth estimation module.
[0195] The depth estimation module can be used to perform depth estimation processing on the two-dimensional image of the target object to determine the depth features and depth map of the two-dimensional image of the target object.
[0196] The feature fusion module 1104 may include a fused two-dimensional feature determination submodule. This fused two-dimensional feature determination submodule can be used to perform feature fusion processing on the semantic segmentation map, semantic segmentation features, depth map, depth features, and two-dimensional backbone features to determine the fused two-dimensional features of the target object's two-dimensional image.
[0197] In some embodiments, the 3D model building apparatus 1100 may further include a second depth estimation module and a depth feature determination module.
[0198] The second depth estimation module can be used to perform depth estimation processing on the two-dimensional image of the target object through the depth estimation module, so as to determine the depth map and depth features corresponding to the two-dimensional image; the depth feature determination module can be used to segment the depth map and depth features corresponding to the two-dimensional image of the target object through the semantic segmentation map, so as to determine the depth map and depth features corresponding to the region where the target object is located.
[0199] The feature fusion module 1104 may include a deep fusion submodule.
[0200] The deep fusion submodule can be used to fuse semantic segmentation map, semantic segmentation features, depth map and depth features corresponding to the region where the target object is located, and two-dimensional backbone features to determine the fused two-dimensional features of the two-dimensional image of the target object.
[0201] In some embodiments, the plurality of three-dimensional points include target three-dimensional points; wherein, the occupancy probability inference module 1107 may include: a distance determination submodule, a depth determination submodule, and an occupancy probability determination submodule.
[0202] The distance determination submodule can be used to determine the distance between the target 3D point and the plane where the 2D image of the target object is located; the depth determination submodule can be used to take the distance between the target 3D point and the plane where the 2D image of the target object is located as the depth of the target 3D point; the occupancy probability determination submodule can be used to determine the spatial occupancy probability of the target 3D point in the 3D model of the target object based on the fused 2D features of the projection point of the target 3D point and the depth of the target 3D point.
[0203] In some embodiments, the occupancy probability reasoning module 1107 may include: a 3D point extraction submodule, a 3D surface construction submodule, and a 3D model construction submodule.
[0204] Among them, the 3D point extraction submodule can be used to extract 3D points with a spatial occupancy probability equal to the target value from multiple 3D points; the 3D surface construction submodule can be used to generate the 3D surface of the 3D model of the target object based on the 3D points with a spatial occupancy probability equal to the target value; and the 3D model construction submodule can be used to construct the 3D model of the target object based on the 3D surface.
[0205] In some embodiments, the 3D point acquisition module 1105 may include: a preset height acquisition submodule, a 3D space construction submodule, and a 3D point determination submodule.
[0206] The preset height acquisition submodule can be used to acquire a preset height value; the 3D space construction submodule can be used to construct a 3D space based on the 2D image of the target object and the preset height value; and the 3D point determination submodule can be used to perform uniform or random sampling in the 3D space to obtain multiple 3D points.
[0207] In some embodiments, the three-dimensional model construction device 1100 may further include: an actual three-dimensional model determination module, an actual model segmentation module, a training backbone feature determination module, a training fusion module, a training three-dimensional point acquisition module, a two-dimensional alignment feature determination module, an occupancy probability inference module, an actual positional relationship determination module, a loss determination module, and a parameter training module.
[0208] The system includes the following modules: The actual 3D model determination module acquires the 2D image and the actual 3D model of the training object; the actual model segmentation module performs semantic segmentation on the 2D image of the training object using the semantic segmentation module to determine the corresponding semantic segmentation features and semantic segmentation map; the training backbone feature determination module extracts features from the 2D image of the training object using the feature extraction module to determine the corresponding 2D backbone features; the training fusion module fuses the semantic segmentation map, semantic segmentation features, and 2D backbone features of the training object to determine the fused 2D features of the training object's 2D image; and the training 3D point acquisition module acquires 3D training points in 3D space, where the projection of the 3D training points onto the plane of the 2D image of the training object falls within... In the 2D image of the training object; the 2D alignment feature determination module can be used to determine the fused 2D features corresponding to the projection points of the 3D training points in the 2D image of the training object based on the fused 2D features corresponding to the 2D image of the training object; the occupancy probability inference module can be used to perform occupancy probability inference processing on the fused 2D features of the 3D training points through the occupancy probability inference module to determine the spatial occupancy probability of the 3D training points in the predicted 3D model of the training object; the actual positional relationship determination module can be used to determine the actual positional relationship between the 3D training points and the actual 3D model of the training object; the loss determination module can be used to determine the loss value based on the actual positional relationship between the 3D training points and the actual 3D model and the spatial occupancy probability corresponding to the 3D training points; the parameter training module can be used to train the parameter values of the feature extraction module and the occupancy probability inference module through the loss value.
[0209] Since the functions of the device 1100 have been described in detail in their respective method embodiments, they will not be repeated here.
[0210] The modules and / or sub-modules and / or units described in the embodiments of this disclosure can be implemented in software or hardware. The described modules and / or sub-modules and / or units can also be located in a processor. The names of these modules and / or sub-modules and / or units do not, in some cases, constitute a limitation on the module and / or sub-module and / or unit itself.
[0211] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module or program segment containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer program instructions.
[0212] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0213] Figure 12 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. It should be noted that... Figure 12 The illustrated electronic device 1200 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0214] like Figure 12 As shown, the electronic device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage section 1208 into a random access memory (RAM) 1203. The RAM 1203 also stores various programs and data required for the operation of the electronic device 1200. The CPU 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0215] The following components are connected to I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN card, modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to I / O interface 1205 as needed. Removable media 1211, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1210 as needed so that computer programs read from them can be installed into storage section 1208 as needed.
[0216] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing computer program instructions for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1209, and / or installed from removable medium 1211. When the computer program is executed by central processing unit (CPU) 1201, it performs the functions defined above in the system of this disclosure.
[0217] It should be noted that the computer-readable storage medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable computer program instructions. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Computer program instructions contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0218] In another aspect, this disclosure also provides a computer-readable storage medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable storage medium carries one or more programs that, when executed by the device, enable the device to perform the following functions: acquiring a two-dimensional image of a target object; performing semantic segmentation processing on the two-dimensional image of the target object using a semantic segmentation module to determine the semantic segmentation features and semantic segmentation map corresponding to the target object; and performing feature extraction processing on the two-dimensional image of the target object using a feature extraction module to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object.
[0219] The semantic segmentation map, semantic segmentation features, and 2D backbone features corresponding to the target object are fused to determine the fused 2D features of the target object's 2D image. Multiple 3D points in 3D space are acquired, where the projection of each 3D point onto the plane of the target object's 2D image falls within the target object's 2D image. Based on the fused 2D features corresponding to the target object's 2D image, the fused 2D features corresponding to the projection points of each 3D point in the target object's 2D image are determined. An occupancy probability inference module is used to perform occupancy probability inference on the fused 2D features corresponding to each 3D point to determine the spatial occupancy probability of each 3D point within the target object's 3D model, so that the 3D model of the target object can be constructed based on the spatial occupancy probability corresponding to each 3D point.
[0220] According to one aspect of this disclosure, a computer program product or computer program is provided, comprising computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and a processor executes the computer program instructions to implement the methods provided in various optional implementations of the above embodiments.
[0221] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several computer program instructions to cause an electronic device (such as a server or terminal device, etc.) to execute the method according to the embodiments of this disclosure.
[0222] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0223] It should be understood that this disclosure is not limited to the detailed structures, drawing arrangements or implementations shown herein; rather, this disclosure is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.
Claims
1. A method for constructing a three-dimensional model, characterized in that, include: Obtain a two-dimensional image of the target object; The semantic segmentation module performs semantic segmentation processing on the two-dimensional image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object. The feature extraction module performs feature extraction processing on the two-dimensional image of the target object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object. The semantic segmentation map, semantic segmentation features, and two-dimensional backbone features corresponding to the target object are subjected to feature fusion processing to determine the fused two-dimensional features of the two-dimensional image of the target object; Acquire multiple three-dimensional points in three-dimensional space, wherein the projection point of each three-dimensional point on the plane where the two-dimensional image of the target object is located falls in the two-dimensional image of the target object; Based on the fused two-dimensional features corresponding to the two-dimensional image of the target object, determine the fused two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object; The occupancy probability reasoning module performs occupancy probability reasoning on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point. The plurality of three-dimensional points include target three-dimensional points; wherein, by performing occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point through an occupancy probability inference module, the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object is determined, including: Determine the distance between the target 3D point and the plane containing the 2D image of the target object; The distance between the target 3D point and the plane containing the 2D image of the target object is taken as the depth of the target 3D point; Based on the fused two-dimensional features of the projected points of the target three-dimensional point and the depth of the target three-dimensional point, the spatial occupancy probability of the target three-dimensional point in the three-dimensional model of the target object is determined.
2. The method according to claim 1, characterized in that, The method further includes: The depth estimation module performs depth estimation processing on the two-dimensional image of the target object to determine the depth features and depth map of the two-dimensional image of the target object; The process of fusing the semantic segmentation map, the semantic segmentation features, and the two-dimensional backbone features to determine the fused two-dimensional features of the target object's two-dimensional image includes: The semantic segmentation map, the semantic segmentation features, the depth map, the depth features, and the two-dimensional backbone features are fused to determine the fused two-dimensional features of the two-dimensional image of the target object.
3. The method according to claim 1, characterized in that, The method further includes: The depth estimation module performs depth estimation processing on the two-dimensional image of the target object to determine the depth map and depth features corresponding to the two-dimensional image. The semantic segmentation map is used to segment the depth map and depth features corresponding to the two-dimensional image of the target object, so as to determine the depth map and depth features corresponding to the region where the target object is located; The process of fusing the semantic segmentation map, the semantic segmentation features, and the two-dimensional backbone features to determine the fused two-dimensional features of the target object's two-dimensional image includes: The semantic segmentation map, the semantic segmentation features, the depth map and depth features corresponding to the region where the target object is located, and the two-dimensional backbone features are fused to determine the fused two-dimensional features of the two-dimensional image of the target object.
4. The method according to claim 1, characterized in that, Constructing a 3D model of the target object based on the spatial occupancy probability corresponding to each 3D point, including: Extract the three-dimensional points whose spatial occupancy probability is equal to the target value from the plurality of three-dimensional points; Based on the three-dimensional points whose spatial occupancy probability is equal to the target value, generate the three-dimensional surface of the three-dimensional model of the target object; The three-dimensional model of the target object is constructed based on the three-dimensional surface.
5. The method according to claim 1, characterized in that, Acquiring multiple 3D points in 3D space, wherein the projection of each 3D point onto the plane containing the 2D image of the target object falls within the 2D image of the target object, including: Get the preset height value; The three-dimensional space is constructed based on the two-dimensional image of the target object and the preset height value; The plurality of three-dimensional points are obtained by uniform sampling or random sampling in the three-dimensional space.
6. The method according to claim 1, characterized in that, The method further includes: Obtain a two-dimensional image of the training object and the actual three-dimensional model of the training object; The semantic segmentation module performs semantic segmentation processing on the two-dimensional image of the training object to determine the semantic segmentation features and semantic segmentation map corresponding to the training object. The feature extraction module performs feature extraction processing on the two-dimensional image of the training object to determine the two-dimensional backbone features corresponding to the two-dimensional image of the training object. The semantic segmentation map, semantic segmentation features, and two-dimensional backbone features of the training object are fused to determine the fused two-dimensional features of the two-dimensional image of the training object. Obtain three-dimensional training points in the three-dimensional space, wherein the projection points of the three-dimensional training points on the plane where the two-dimensional image of the training object is located fall in the two-dimensional image of the training object; Based on the fused two-dimensional features corresponding to the two-dimensional image of the training object, determine the fused two-dimensional features corresponding to the projection points of the three-dimensional training points in the two-dimensional image of the training object; The occupancy probability inference module performs occupancy probability inference processing on the fused two-dimensional features of the three-dimensional training points to determine the spatial occupancy probability of the three-dimensional training points in the predicted three-dimensional model of the training object. Determine the actual positional relationship between the three-dimensional training points and the actual three-dimensional model of the training object; The loss value is determined based on the actual positional relationship between the 3D training points and the actual 3D model, and the spatial occupancy probability corresponding to the 3D training points. The loss value is used to train the parameter values of the feature extraction module and the occupancy probability inference module.
7. A three-dimensional model construction device, characterized in that, include: The 2D image acquisition module is used to acquire a 2D image of the target object; The semantic segmentation module is used to perform semantic segmentation processing on the two-dimensional image of the target object to determine the semantic segmentation features and semantic segmentation map corresponding to the target object. The two-dimensional backbone feature extraction module is used to perform feature extraction processing on the two-dimensional image of the target object through the feature extraction module, so as to determine the two-dimensional backbone features corresponding to the two-dimensional image of the target object. The feature fusion module is used to perform feature fusion processing on the semantic segmentation map, semantic segmentation features and two-dimensional backbone features corresponding to the target object, so as to determine the fused two-dimensional features of the two-dimensional image of the target object; The 3D point acquisition module is used to acquire multiple 3D points in 3D space, wherein the projection point of each 3D point on the plane where the 2D image of the target object is located falls in the 2D image of the target object. The fusion two-dimensional feature acquisition module is used to determine the fusion two-dimensional features corresponding to the projection points of each three-dimensional point in the two-dimensional image of the target object based on the fusion two-dimensional features corresponding to the two-dimensional image of the target object. The occupancy probability reasoning module is used to perform occupancy probability reasoning on the fused two-dimensional features corresponding to each three-dimensional point to determine the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object, so as to construct the three-dimensional model of the target object based on the spatial occupancy probability corresponding to each three-dimensional point. The plurality of three-dimensional points include target three-dimensional points; wherein, by performing occupancy probability inference processing on the fused two-dimensional features corresponding to each three-dimensional point through an occupancy probability inference module, the spatial occupancy probability of each three-dimensional point in the three-dimensional model of the target object is determined, including: Determine the distance between the target 3D point and the plane containing the 2D image of the target object; The distance between the target 3D point and the plane containing the 2D image of the target object is taken as the depth of the target 3D point; Based on the fused two-dimensional features of the projected points of the target three-dimensional point and the depth of the target three-dimensional point, the spatial occupancy probability of the target three-dimensional point in the three-dimensional model of the target object is determined.
8. An electronic device, characterized in that, include: Memory; as well as A processor coupled to the memory, the processor being used to execute the three-dimensional model construction method as described in any one of claims 1-6 based on computer program instructions stored in the memory.
9. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the three-dimensional model construction method as described in any one of claims 1-6.
Citation Information
Patent Citations
Method and device for constructing three-dimensional semantic map
CN112819893A
High-resolution human body three-dimensional reconstruction method based on multi-view RGBD camera
CN112927348A