Three-dimensional reconstruction method, device and readable storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请提供一种三维重建方法、设备及可读存储介质,用以解决现有技术中符号距离场和颜色信息场的重建速度慢,导致三维重建的速度慢、效率低的问题
[0020] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first or second aspect.
Smart Images

Figure CN115797561B_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and more particularly to a three-dimensional reconstruction method, apparatus, and readable storage medium. Background Technology
[0002] 3D reconstruction is a key component of Extended Reality (XR) applications, and the reconstruction and representation of 3D information about people, objects, and environments is the core technology of 3D reconstruction. Symbolic distance field and color information field are important methods for representing geometric and color information in 3D space, respectively. How to optimize the symbolic distance field and color information field during the 3D reconstruction process is a core aspect of 3D reconstruction and has received continuous research and attention.
[0003] Current 3D reconstruction methods typically employ two multi-layer perceptrons (MLPs) to represent the symbolic distance field and the color information field, respectively. For an object to be reconstructed, a set of observed images of the object is used to optimize the MLPs to reconstruct the object's symbolic distance field and color information field. A 3D model of the object is then built based on the reconstructed symbolic distance field and color information field. However, the inference speed of the MLP is relatively slow, resulting in a slow reconstruction speed for the symbolic distance field and color information field, thus leading to slow 3D reconstruction speed and low efficiency. Summary of the Invention
[0004] This application provides a three-dimensional reconstruction method, device, and readable storage medium to solve the problem of slow reconstruction speed of symbolic distance field and color information field in the prior art, which leads to slow speed and low efficiency of three-dimensional reconstruction.
[0005] Firstly, this application provides a three-dimensional reconstruction method, including:
[0006] Acquire the observation image of the target object to be reconstructed, and the camera parameters corresponding to the observation image;
[0007] Based on the camera parameters, a voxel model corresponding to the 3D scene is constructed. The symbolic distance field and color information field of the 3D scene are initialized based on the voxel model. The symbolic distance field contains the symbolic distance information at the voxel vertices in the voxel model, and the color information field contains the color information at the voxel vertices in the voxel model.
[0008] Based on the camera parameters, the symbolic distance field, and the color information field, the pixels in the observed image are rendered, and the rendered color value of the pixel is determined.
[0009] The loss is calculated based on the rendered color value and the actual color value of the pixel. The symbolic distance field and color information field are then optimized based on the loss to obtain the optimized symbolic distance field and color information field.
[0010] A three-dimensional model of the target object is generated based on the optimized symbolic distance field.
[0011] Secondly, this application provides a three-dimensional reconstruction method, including:
[0012] Acquire observation images of the target product to be reconstructed, as well as the camera parameters corresponding to the observation images;
[0013] A voxel model of the corresponding 3D scene is constructed based on the camera parameters, and the symbolic distance field and color information field of the 3D scene are initialized based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model.
[0014] Based on the camera parameters, the symbolic distance field, and the color information field, the pixels in the observed image are rendered, and the rendered color value of the pixel is determined.
[0015] The loss is calculated based on the rendered color value and the actual color value of the pixel. The symbolic distance field and color information field are then optimized based on the loss to obtain the optimized symbolic distance field and color information field.
[0016] A 3D model of the target product is generated based on the optimized symbolic distance field; the 3D model of the target product is rendered based on the optimized color information field, so as to display the rendered 3D model of the target product on the display page of the target product.
[0017] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0018] The memory stores computer-executed instructions;
[0019] The processor executes computer execution instructions stored in the memory to implement the method described in the first or second aspect.
[0020] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first or second aspect.
[0021] The 3D reconstruction method, device, and readable storage medium provided in this application construct a voxel model of the 3D scene where the target object is located. By storing the symbolic distance values at the voxel vertices in the voxel model, the symbolic distance field of the 3D spatial geometry is explicitly represented. By storing the color information at the voxel vertices in the voxel model, the color information field of the 3D spatial geometry is explicitly represented. The initial symbolic distance field and color information field are optimized using the observed image of the target object to reconstruct the symbolic distance field and color information field of the target object (the 3D scene where it is located). Based on the optimized symbolic distance field, a 3D model of the target object is generated. Compared with the training and inference time required by the multilayer perceptron, this embodiment significantly shortens the time required to optimize the symbolic distance field and color information field, and improves the reconstruction speed of the symbolic distance field and color information field, thereby improving the speed and efficiency of 3D reconstruction. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] Figure 1 A schematic diagram of a three-dimensional reconstructed scene provided for this application;
[0024] Figure 2 A flowchart of a three-dimensional reconstruction method provided as an exemplary embodiment of this application;
[0025] Figure 3 An infographic showing the expansion of the spherical harmonic basis functions of 0 to 2 degrees provided in this application using Cartesian coordinates;
[0026] Figure 4 A flowchart illustrating the optimization of the symbolic distance field and color information field provided for an exemplary embodiment of this application;
[0027] Figure 5 A flowchart illustrating the rendering of an observed image provided in an exemplary embodiment of this application;
[0028] Figure 6 A flowchart of a 3D reconstruction method in an e-commerce scenario provided as an exemplary embodiment of this application;
[0029] Figure 7 A structural diagram of a three-dimensional reconstruction apparatus provided in an exemplary embodiment of this application;
[0030] Figure 8 A structural diagram of a three-dimensional reconstruction apparatus provided in another exemplary embodiment of this application;
[0031] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an example embodiment of this application.
[0032] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0034] First, let me explain the terms used in this application:
[0035] Signed Distance Field (SDF): A method for representing geometric information in three-dimensional space. For a point in three-dimensional space, a scalar (signed distance value) is defined. The absolute value of the scalar is the distance from the point to the nearest surface. The positive / negative sign of the scalar indicates whether the point is outside or inside the object on the surface.
[0036] Radiance Field (RF): A method for representing three-dimensional space, where each point in the three-dimensional space is expressed as its density and color. The density information of each point is a scalar greater than or equal to 0; the color information of each point is an anisotropic vector containing the color of that point in each direction in space, and the color information in each direction can be different.
[0037] Neural Radiance Field (NeRF): A radiation field represented by a neural network, typically using a multilayer perceptron (MLP). The input of an MLP is the three-dimensional coordinates and orientation of a point in the scene, and the output is the density at the input point and the color of the input point in the input orientation.
[0038] Color information field: A method for representing color information in three-dimensional space. For a point in three-dimensional space, the color information of each point is an anisotropic vector, which contains the color of the point in each direction in space, and the color information in each direction can be different.
[0039] A voxel model is a model that uses a large number of regular voxels (such as cubes) to represent a three-dimensional object / three-dimensional space. In this application, it refers to a voxel model that represents three-dimensional space. The three-dimensional space is divided into discrete three-dimensional lattices (i.e., voxels). Only the information at the vertices of the voxels is stored. The information at the non-vertices is obtained through spatial trilinear interpolation. To address the issue of angular continuity, a set of spherical harmonic functions is used as basis functions. Only the coefficients of the basis functions need to be stored.
[0040] Explicit Representation: This refers to directly storing information at each point in a defined field (such as a symbolic distance field or a radiation field). Since positions or angles in three-dimensional space are continuous and cannot be directly stored, this application uses a volume model to explicitly store the symbolic distance field and color information field in three-dimensional space to address the issue of positional continuity.
[0041] Implicit representation: This typically refers to the use of neural networks to implicitly store field information. The coordinates and / or angles of points within the field are input into the neural network for inference, and the inference result is used as the field information at that point. For example, neural radiation fields are a method of implicitly representing radiation fields.
[0042] Spherical harmonics are a set of orthogonal basis functions defined in spherical coordinates. Theoretically, a linear combination of spherical harmonics can approximate any function defined in spherical coordinates within an acceptable error range. Spherical harmonics have wide applications; in computer graphics, they are often used to represent the color information of a point in space in various directions.
[0043] To address the problem of slow 3D reconstruction speed and low efficiency caused by the slow inference speed of multilayer perceptrons (MLPs) used to represent the symbolic distance field and color information field separately, this application provides a 3D reconstruction method. This method constructs a voxel model to explicitly store the symbolic distance field and color information field. Based on camera parameters, the symbolic distance field, and the color information field, pixels in the observed image are rendered, and the rendered color value of each pixel is determined. A loss is calculated based on the rendered color value and the actual color value of the pixel. The symbolic distance field and color information field are then optimized based on the loss, resulting in an optimized symbolic distance field and color information field for the 3D scene. This allows for the reconstruction of the symbolic distance field and color information field of the 3D scene containing the target object. Based on the explicitly stored symbolic distance field and color information field, reconstruction / optimization of the symbolic distance field and color information field can be achieved through simple calculations, further reconstructing the 3D model of the target object. Compared to the inference speed of multilayer perceptrons, this method significantly improves the speed of reconstructing / optimizing the symbolic distance field and color information field, thereby improving the speed and efficiency of 3D reconstruction.
[0044] The subject executing the three-dimensional reconstruction method provided in this application can be an electronic device with network communication, computing and information display functions, including but not limited to terminals, servers, Internet of Things devices, and clusters deployed in the cloud.
[0045] The 3D reconstruction method provided in this application can be widely applied to fields such as e-commerce, extended reality, architecture, home decoration, transportation, and manufacturing. Specifically, it can be applied to the 3D reconstruction of various people, objects, and scenes, including goods in e-commerce, real-world images of intersections / streets in immersive maps, interior scenes in home decoration, parts in industrial manufacturing, and people and scenes in extended reality.
[0046] For example, refer to Figure 1 , Figure 1This application provides a schematic diagram of a 3D reconstruction scene. Taking the 3D reconstruction of goods in the e-commerce field as an example, the user uses an image acquisition device to acquire an observation image of the target product to be reconstructed, and transmits the observation image of the target product and the corresponding camera parameters to an electronic device. The electronic device constructs a voxel model of the corresponding 3D scene based on the camera parameters, and initializes the symbolic distance field and color information field of the 3D scene based on the voxel model. This 3D scene contains the target product, and the symbolic distance field and color information field of the 3D scene together contain the symbolic distance field and color information field of the target product. The electronic device uses the observation image of the target product and the camera parameters, and uses the current symbolic distance field and color information field to render the pixels in the observation image. Based on the rendered color value and the actual color value of the pixel, the symbolic distance field and color information field are optimized. After optimization, an accurate symbolic distance field and color information field of the 3D scene are obtained. The process of optimizing the symbolic distance field is the process of reconstructing the symbolic distance field of the 3D scene containing the target product. Based on the optimized symbolic distance field, the 3D model of the target product can be reconstructed. Furthermore, the 3D model of the target product is rendered based on the optimized color information field to obtain a simulation model (i.e., a color model) of the target product.
[0047] Depending on the needs of the actual application scenario, the electronic device can output a 3D model (usually a white model) of the target product reconstructed based on the optimized symbolic distance field, or it can output a simulation model of the target product rendered based on the optimized color information field.
[0048] For example, electronic devices can provide the reconstructed white model and / or color model of the target product to the server in the e-commerce platform. The server, based on the processing logic of the actual application scenario, displays the white model and / or color model of the target product on the target product's display page when the user triggers the opening of the target product's display page.
[0049] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0050] Figure 2 This is a flowchart illustrating a three-dimensional reconstruction method provided as an exemplary embodiment of this application. The execution subject of the method provided in this application is the electronic device described above for executing the three-dimensional reconstruction method flow. Figure 2 As shown, the specific steps of this method are as follows:
[0051] Step S201: Obtain the observation image of the target object to be reconstructed, and the camera parameters corresponding to the observation image.
[0052] The target object for 3D reconstruction can be a person, an object, a scene, etc.
[0053] For example, 3D virtual character models can be created based on real-life images to serve as intelligent customer service representatives or characters in online games. 3D models of products can be created based on observed images, allowing users to view and understand the product's 3D information. 3D models of intersection scenes can be created based on images of intersections to generate real-world maps or display detailed intersection navigation maps in navigation systems.
[0054] In this embodiment, in order to reconstruct the 3D model of the target object, it is necessary to acquire a set of observation images of the target object. Typically, the target object is placed in a scene, and images of the target object are acquired from multiple different perspectives. This set of observation images corresponds to the same 3D scene and has the same camera coordinate system.
[0055] The camera parameters for the observed images include intrinsic and extrinsic parameters. Intrinsic parameters are the internal parameters of the image acquisition device that acquires the observed images, while extrinsic parameters include the position and orientation of the image acquisition device in the 3D scene. In this embodiment, in order to construct a 3D model of the target object, the acquired observation images of the target object are a set of observation images corresponding to the same 3D scene, that is, under the same camera coordinate system.
[0056] In addition, if a set of observed images corresponds to different camera coordinate systems, one of the camera coordinate systems can be used as the reference for constructing the voxel model. In subsequent calculations, the coordinates of pixels in other different camera coordinate systems are converted to the reference camera coordinate system according to the camera coordinate system corresponding to the observed images, and the symbolic distance field and color information field are reconstructed in the reference coordinate system.
[0057] Step S202: Based on the camera parameters, construct a voxel model corresponding to the 3D scene, and initialize the symbolic distance field and color information field of the 3D scene based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model.
[0058] In this model, the voxel model of the 3D scene divides the 3D scene into multiple 3D grids, each of which is a voxel. Each voxel is a regular geometric shape, such as a cube or cuboid. A voxel has multiple vertices, each of which is a point in 3D space.
[0059] In this embodiment, based on the voxel model in three-dimensional space, the symbolic distance field and color information field are initialized. The resulting symbolic distance field and color information field can be understood as two voxel models with the same voxel division but different information stored at the voxel vertices.
[0060] Specifically, by storing the symbolic distance values at the voxel vertices in the voxel model, the symbolic distance values at any non-vertex location in space can be obtained through spatial trilinear interpolation, resulting in a symbolic distance field that explicitly represents the geometric information of three-dimensional space. The symbolic distance values at the voxel vertices contained in the symbolic distance field can be randomly initialized.
[0061] By storing the color information at the voxel vertices in the voxel model, the color information at any non-vertex location in space can be obtained through spatial trilinear interpolation, resulting in a color information field that explicitly represents the geometric information of three-dimensional space. The color information field containing the color information at the voxel vertices can be randomly initialized.
[0062] In this embodiment, spherical harmonic basis functions can be used to represent the color information of a point in space in various directions. Here, a spherical harmonic function is a set of orthogonal basis functions defined on spherical coordinates. Each spherical harmonic basis function can be indicated by its degree and order. Typically, spherical harmonic basis functions of 0 to 2 degrees use the expansion of Cartesian coordinates as follows: Figure 3 As shown, Figure 3 middle Let denote a spherical harmonic basis function, where the subscript l is the degree of the spherical harmonic basis function and the superscript m is the order of the spherical harmonic basis function, and satisfy . It represents the set of non-negative integers.
[0063] Specifically, nine spherical harmonic basis functions of degrees 0 to 2 are used to represent the color information of the red, green, and blue color channels, respectively. The color information at each voxel vertex in the color information field stores 27 (9 × 3) weight coefficients; that is, the color information at each voxel vertex in the color information field is a vector of weight coefficients for the spherical harmonic basis functions. For a point in the 3D scene, the color value of that point can be calculated based on the weight coefficients of the spherical harmonic basis functions at that point and the direction of the ray corresponding to that point.
[0064] Alternatively, in other alternative embodiments, other basis functions can be used to represent color information. For example, randomly initialized basis functions can be used, or basis functions can be automatically learned using a neural network. The learned basis functions are then used to represent color information.
[0065] After initializing the symbolic distance field and color information field, the symbolic distance field and color information field are iteratively optimized using the observed image of the target object. After optimization, the reconstructed symbolic distance field and color information field can be obtained.
[0066] Step S203: Based on the camera parameters, symbolic distance field, and color information field, render the pixels in the observed image and determine the rendered color value of the pixels.
[0067] In this step, the pixels in the observed image are rendered based on the camera parameters, symbolic distance field, and color information field to obtain the color value of the rendered pixel, which is called the rendered color value.
[0068] Step S204: Calculate the loss based on the rendered color value and the actual color value of the pixel, and optimize the symbolic distance field and color information field based on the loss to obtain the optimized symbolic distance field and color information field.
[0069] In this step, the loss is calculated based on the rendered color value of the pixel and the actual color value of the pixel in the observed image. Based on the loss, the symbolic distance value at the voxel vertex in the symbolic distance field and the color information at the voxel vertex in the color information field are updated through backpropagation.
[0070] Through multiple rounds of iterative optimization, the iteration stops when the convergence condition is met, resulting in the optimized symbolic distance field and color information field. The convergence condition can be set and adjusted according to the needs of the actual application scenario, and is not specifically limited here. For example, the convergence condition may include at least one of the following: the loss decreases significantly and stabilizes, or the loss value is less than or equal to a preset threshold.
[0071] In this embodiment, the symbolic distance field and color information field initialized based on the voxel model are essentially two voxel models with identical voxel partitioning but different stored information at voxel vertices. The process of rendering pixels in the observed image based on the symbolic distance field and color information field is differentiable, and stochastic gradient descent can be used to optimize the symbolic distance field and color information field. By representing the symbolic distance field and color information field of the 3D scene using an explicitly expressed voxel model, and thus leveraging differentiable volume rendering, the symbolic distance field and color information field of the 3D scene can be optimized based on a set of observed images with known camera parameters, significantly improving the optimization speed of the symbolic distance field.
[0072] The optimization process of the symbolic distance field and color information field is also the process of reconstructing the symbolic distance field and color information field of the 3D scene. After optimization, a symbolic distance field that accurately represents the geometric information of the 3D scene and a color information field that accurately represents the color information of the 3D scene can be obtained.
[0073] For example, Figure 4 For a flowchart illustrating the optimization of the symbolic distance field and color information field in this embodiment, please refer to [link / reference]. Figure 4Based on the voxel model initialization, a symbolic distance field and color information field are obtained. Forward result calculation is then performed, that is, based on the camera parameters and the current symbolic distance field and color information field, the observed image is rendered (this rendering process is differentiable), yielding a rendering result that includes the rendered color values of the pixels in the observed image. Based on the rendering result and the ground truth (the actual color values of the pixels in the observed image), the loss function is calculated. Backward gradient propagation is performed using the chain rule, propagating the gradient of the loss function to the symbolic distance field and color information field, updating the information at the voxel vertices in these fields. This optimization process is repeated iteratively multiple times until the convergence condition is met, thus obtaining the final symbolic distance field and color information field.
[0074] Step S205: Generate a 3D model of the target object based on the optimized symbolic distance field.
[0075] After optimization, the optimized symbolic distance field contains the geometric information of the target object in the 3D scene. Based on the optimized symbolic distance field, a 3D model of the target object can be generated, usually a white model.
[0076] Optionally, the electronic device can also display the 3D model of the target object generated based on the optimized symbolic distance field online, or provide it to other devices, systems, or services, so that other devices, systems, or services can display the 3D model of the target object, or continue to modify, create, or overlay other effects based on the 3D model of the target object.
[0077] This step can be implemented using any existing technology that draws a 3D model based on the symbolic distance field, which will not be elaborated here.
[0078] The method in this embodiment constructs a voxel model of the 3D scene where the target object is located. By storing the symbolic distance values at the voxel vertices in the voxel model, the symbolic distance field of the 3D spatial geometry is explicitly represented. By storing the color information at the voxel vertices in the voxel model, the color information field of the 3D spatial geometry is explicitly represented. The initial symbolic distance field and color information field are optimized using the observed image of the target object to reconstruct the symbolic distance field and color information field of the target object (the 3D scene where it is located). Based on the optimized symbolic distance field, a 3D model of the target object is generated. Compared with the training and inference time required by the multilayer perceptron, this embodiment significantly shortens the time required to optimize the symbolic distance field and color information field, and improves the reconstruction speed of the symbolic distance field and color information field, thereby improving the speed and efficiency of 3D reconstruction.
[0079] In one optional embodiment, after generating a 3D model (white model) of the target object based on the optimized symbolic distance field, the 3D model of the target object can be rendered based on the optimized color information field to obtain a colored simulation model of the target object. The electronic device can also display the rendered 3D model of the target object online or provide it to other devices, systems, or services, allowing other devices, systems, or services to display the rendered 3D model of the target object, or to further modify, create, or overlay other effects based on the rendered 3D model of the target object.
[0080] In an alternative embodiment, see Figure 5 In step S203 above, the pixels in the observed image are rendered based on the camera parameters, the symbolic distance field, and the color information field, and the rendered color value of the pixel is determined. Specifically, this can be achieved using the following steps S2031-S2033:
[0081] Step S2031: Based on the camera parameters corresponding to the observed image, determine multiple sampling points on the ray corresponding to the pixel in the observed image.
[0082] Specifically, this step can be implemented in the following way:
[0083] Based on the camera parameters corresponding to the observed image, determine the ray corresponding to the pixel; determine the intersection point between the ray corresponding to the pixel and the voxel model, and determine the nearest intersection point that is closest to the endpoint of the ray, and the farthest intersection point that is farthest from the endpoint of the ray; perform equidistant sampling between the nearest and farthest intersection points on the ray to obtain multiple sampling points.
[0084] Here, the ray corresponding to a pixel is a ray that originates at the camera and passes through the pixel. For an observed image with given camera parameters, let the coordinates of the camera origin be denoted as o. For a pixel in the observed image, let the direction from the camera origin to the pixel be denoted as d. Then, the equation of the ray corresponding to a pixel in the observed image can be expressed as: o + t·d, t > 0. Here, t is a free parameter in the equation.
[0085] Furthermore, since the ray corresponding to the pixel is continuous and extends infinitely, by determining the nearest and farthest intersection points of the ray with the voxel model, the portion between the nearest and farthest intersection points on the ray is the portion in the current 3D scene. By sampling multiple points from this portion, it can be ensured that the sampling points are points in the current 3D scene.
[0086] In this embodiment, multiple sampling points can be obtained by sampling at equal intervals between the nearest and farthest intersection points on the ray, resulting in a uniformly distributed set of sampling points. Optionally, a preset number of sampling points can be obtained based on a pre-set number of sampling points, or multiple sampling points with an interval distance between adjacent sampling points less than or equal to the preset interval distance threshold can be obtained based on a preset interval distance threshold.
[0087] Alternatively, in other embodiments, multiple sampling points can be randomly sampled from the nearest intersection point and the farthest larger point on the ray. The specific sampling method is not specifically limited here.
[0088] Step S2032: Determine the symbol distance value and color information of multiple sampling points based on the symbol distance field and color information field.
[0089] In this step, for each sampling point, multiple neighboring points are determined among the voxel vertices of the voxel model based on the sampling point's 3D coordinates. Trilinear interpolation is then performed based on the symbolic distance values of these neighboring points in the symbolic distance field to determine the symbolic distance value of the sampling point. Furthermore, trilinear interpolation is performed based on the color information of these neighboring points in the color information field to determine the color information of the sampling point.
[0090] For example, for any sampling point, its position in the 3D scene can be determined based on its 3D coordinates, and its 8-neighborhood points can be determined from the voxel vertices of the voxel model. The 8-neighborhood points include the 8 voxel vertices in the voxel model that are closest to the sampling point. The signed distance value at the sampling point is obtained by performing trilinear interpolation based on the signed distance values of the 8 neighboring points. The color information at the sampling point is then obtained by performing trilinear interpolation based on the color information (such as the weight coefficients of the spherical harmonic basis functions) of the 8 neighboring points.
[0091] Let N represent the number of sampling points obtained on the ray corresponding to the pixel, where N is a positive integer. The sampling points are numbered in ascending order of proximity to the ray endpoint, where the i-th sampling point can be represented as: q i =o+t i ·d. Where o represents the coordinates of the camera origin, d represents the direction from the camera origin to the pixel, and i takes the value of a positive integer in the interval [1, N].
[0092] In other embodiments, depending on the specific scenario, more or fewer neighboring points than 8 can be sampled to calculate the symbolic distance value and color information of the sampled points; no specific limitation is made here.
[0093] Step S2033: Render the pixels based on the symbol distance values and color information of multiple sampling points, and determine the rendering color value of the pixels.
[0094] Specifically, in this step, the opacity at each sampling point is calculated based on the symbol distance values of multiple sampling points; the transparency of the line segment from each sampling point to the ray endpoint is calculated based on the opacity at each sampling point; and the rendering color value of the pixel is determined based on the opacity at each sampling point, the transparency of the line segment from each sampling point to the ray endpoint, the color information of each sampling point, and the pixel color rendering equation.
[0095] In this embodiment, the color information at the voxel vertices stored in the color information field is the weighting coefficient of the spherical harmonic basis function. In this step, the color value at the sampling point needs to be calculated based on the color information of the sampling point (i.e., the weighting coefficient of the spherical harmonic basis function) and the ray direction. The pixel is then rendered based on the color value at the sampling point and other information. The i-th sampling point q... i The color value at that location can be represented by c. i Represented by the i-th sampling point q. i The symbolic distance value at that location can be used express.
[0096] In this step, based on the symbol distance values of multiple sampling points and the preset s-density function, the s-density value at each sampling point can be calculated. For any sampling point, the opacity at the sampling point is determined based on the difference between the s-density value at that sampling point and the s-density values at adjacent sampling points. Here, s is a preset hyperparameter.
[0097] For example, for the i-th sampling point q i The following formula (I) can be used to calculate the current sampling point q. i Opacity at:
[0098]
[0099] Where, α i Represents the i-th sampling point q i Opacity at that location. Represents the i-th sampling point q i The symbolic distance value at the location. Represents the (i-1)th sampling point q i-1 The sign distance value at Φ. s (x) represents the s-density function, s is a pre-set hyperparameter that can be set and adjusted according to the needs of the actual application scenario; no specific limitations are made here. max() represents the function to retrieve the maximum value.
[0100] Furthermore, for the i-th sampling point q i Based on the opacity at each sampling point, the q of the i-th sampling point is calculated using the following formula (I).i Transparency of the line segment to the ray endpoint:
[0101]
[0102] Among them, T i Represents the i-th sampling point q i The transparency of the line segment leading to the ray's endpoint. α j Represents the j-th sampling point q j The opacity at that point. Specifically, T1 = 1.
[0103] Furthermore, based on the opacity at each sampling point, the transparency of the line segment from each sampling point to the ray endpoint, and the color value of each sampling point, the following rendering equation (III) is used to determine the rendering color value of the pixel:
[0104]
[0105] in, This represents the rendered color value of a pixel. T i Represents the i-th sampling point q i The transparency of the line segment leading to the ray's endpoint. α i Represents the i-th sampling point q i Opacity at that location. i The i-th sampling point q i The color value at that location.
[0106] In this embodiment, pixel rendering can be achieved based on the symbolic distance field and color information field explicitly expressed using a voxel model, obtaining the rendered color value of each pixel, thereby obtaining the rendering result of the observed image. The above rendering process is computationally fast, and all calculations are differentiable. By comparing the rendering result of the observed image based on the current symbolic distance field and color information field with the ground truth value of the observed image, the loss function value is calculated. Then, the chain rule is used for backpropagation, propagating the gradient of the loss function to the symbolic distance field and color information field, updating the information at the voxel vertices in the symbolic distance field and color information field, and achieving rapid optimization of the symbolic distance field and color information field.
[0107] In an optional embodiment, in step S204 above, the photometric loss can be calculated based on the rendered color value and the actual color value of the pixel. The symbolic distance field and color information field are then optimized based on the photometric loss.
[0108] For example, the photometric loss can be calculated using the following photometric loss function (iv):
[0109]
[0110] Among them, L colorThis represents the photometric loss. M represents the number of pixels rendered in this iteration, and M is a positive integer. This represents the rendered color value of pixel m. C m This represents the actual color value of pixel m, where pixel m represents any pixel rendered in the current iteration. Different values of the subscript m represent different pixels. ‖x‖1 indicates the calculation of the 1-norm of x.
[0111] Optionally, regularization loss and / or smoothing loss of the symbolic distance field can also be calculated. If multiple losses are calculated, a comprehensive loss is determined based on the multiple losses, and the symbolic distance field and color information field are optimized based on the comprehensive loss. By combining multiple losses, the optimization effect of the symbolic distance field and color information field is improved.
[0112] For example, according to the definition of the symbolic distance field, the gradient magnitude at each point is 1. Therefore, based on the symbolic distance field value of the pixel in the current symbolic distance field, the gradient at the voxel vertex in the voxel model is calculated. Based on the gradient at the voxel vertex, the regularization loss is calculated, so that the optimized symbolic distance field is more in line with the characteristics of the symbolic distance field.
[0113] Specifically, the regularization loss can be calculated using the following regularization loss function (V):
[0114]
[0115] Among them, L Eikonal This represents the regularization loss. K represents the number of voxel vertices in the voxel model corresponding to the symbolic distance field, and K is a positive integer. represents the gradient at voxel vertex k in the sign distance field, where voxel vertex k represents any voxel vertex, and different values of k represent different voxel vertexes. ‖x‖2 represents the calculation of the 2-norm of x.
[0116] Furthermore, the symbolic distance field and color information field are optimized based on the photometric loss and regularization loss. For example, the photometric loss and regularization loss can be weighted and summed to determine the comprehensive loss. The symbolic distance field and color information field are then optimized based on the comprehensive loss, making the optimized symbolic distance field more consistent with the characteristics of a symbolic distance field.
[0117] For example, a smoothing loss is calculated based on the difference between the symbolic distance field at a voxel vertex in the current symbolic distance field and the mean of the symbolic distance fields at neighboring points of the same voxel vertex. Based on the photometric loss and the smoothing loss, the symbolic distance field and color information field are optimized to avoid the influence of high-frequency noise during the optimization process, thereby improving the optimization quality and ultimately increasing the accuracy of the optimized symbolic distance field and color information field.
[0118] Specifically, the regularization loss can be calculated using the following smoothing loss function (vi):
[0119]
[0120] Among them, L smooth This represents the smoothing loss. K represents the number of voxel vertices in the voxel model corresponding to the signed distance field. sdf k This represents the symbolic distance value at voxel vertex k in the symbolic distance field. Ω(k) represents the set of 6-neighborhood points of voxel vertex k in the voxel model. sdf h This represents the symbolic distance value at a neighboring point h of voxel vertex k.
[0121] Furthermore, the symbolic distance field and color information field are optimized based on the photometric loss and smoothing loss. For example, the photometric loss and smoothing loss can be weighted and summed to determine the comprehensive loss. The symbolic distance field and color information field are then optimized based on the comprehensive loss, making the optimized symbolic distance field more consistent with the characteristics of a symbolic distance field.
[0122] For example, the photometric loss, regularization loss, and smoothing loss can be calculated simultaneously, and the symbolic distance field and color information field can be optimized based on the photometric loss, regularization loss, and smoothing loss.
[0123] Specifically, the photometric loss, regularization loss, and smoothing loss can be weighted and summed to determine the comprehensive loss. The symbolic distance field and color information field can then be optimized based on the comprehensive loss, so that the optimized symbolic distance field better matches the characteristics of the symbolic distance field.
[0124] For example, based on the photometric loss, regularization loss, and smoothing loss, the overall loss can be calculated using the following formula (VII):
[0125] L = L color +λ·L Eikonal +β·L smooth (seven)
[0126] Where L represents the overall loss. color Indicates luminance loss. L Eikonal L represents the regularization loss. smooth This represents the smoothing loss. λ and β are the pre-set weighting coefficients of the loss function.
[0127] In addition, this example only uses the weight coefficients of 1, λ and β for photometric loss, regularization loss and smoothing loss, respectively. The weight coefficients of each loss can be set according to the actual application scenario and empirical values. No specific limitation is made here.
[0128] In this embodiment, in order to optimize the explicitly represented symbolic distance field, a set of loss functions are defined, including the photometric loss function, the symbolic distance field gradient regularization loss function, and the symbolic distance field smoothing loss function. By minimizing the comprehensive loss, the symbolic distance field is optimized, which can improve the accuracy of the optimized symbolic distance field.
[0129] It should be noted that in other embodiments, other loss functions may be added to the symbolic distance field, or loss functions may be added to the color information field, in order to improve the optimization quality of the symbolic distance field and / or the color information field. Such schemes should also be regarded as variations under the same technical concept of this scheme.
[0130] Based on any of the above method embodiments, the symbolic information field and color information field of the 3D scene containing the target object are optimized. If the target object is not a scene, but a person or object, background information other than the target object can be removed from the observed image of the target object before optimizing the symbolic information field and color information field to segment the image region of the target object. For example, the background can be set to a white background. Optimizing the symbolic information field and color information field based on the observed image after removing background information yields the symbolic information field and color information field of the target object, avoiding the influence of background noise on the symbolic information field and color information field of the target object.
[0131] In addition, if the target object is not a scene, but a person or object, the background information can be left unremoved before optimizing the symbolic information field and color information field. Instead, after the symbolic information field and color information field are optimized, the voxel vertices that do not belong to the target object are segmented from the voxel model, and the information of the voxel vertices that do not belong to the target object is removed from the optimized symbolic information field and color information field to obtain the symbolic information field and color information field of the target object.
[0132] Figure 6 This is a flowchart illustrating a three-dimensional reconstruction method provided as another exemplary embodiment of this application. The three-dimensional reconstruction method provided in this embodiment is applicable to reconstructing three-dimensional models of target products in e-commerce scenarios. Figure 6 As shown, the specific steps of this method are as follows:
[0133] Step S601: Acquire observation images of the target product to be reconstructed, as well as the camera parameters corresponding to the observation images.
[0134] Similar to step S201 above, in order to reconstruct the 3D model of the target product, it is necessary to acquire a set of observation images of the target product. For example, the target product can be placed in a scene, and images of the target product can be acquired from multiple different perspectives to obtain a set of observation images in the same camera coordinate system.
[0135] Optionally, to reconstruct the symbolic distance field and color information field of the target product, background information other than the target product can be removed from the observed image of the target product to segment the image region where the target product is located. For example, the background can be set to a white background. The symbolic information field and color information field are optimized based on the observed image after removing the background information. After optimization, the symbolic information field and color information field of the target product can be obtained, avoiding the influence of background noise on the symbolic information field and color information field of the target product, thus improving the optimization quality.
[0136] Step S602: Construct a voxel model of the corresponding 3D scene based on the camera parameters, and initialize the symbolic distance field and color information field of the 3D scene based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model.
[0137] Step S603: Based on the camera parameters, symbolic distance field, and color information field, render the pixels in the observed image and determine the rendered color value of the pixels.
[0138] Step S604: Calculate the loss based on the rendered color value and the actual color value of the pixel, and optimize the symbolic distance field and color information field based on the loss to obtain the optimized symbolic distance field and color information field.
[0139] In this embodiment, the specific implementation of steps S602-S604 is similar to that of steps S202-S204. Please refer to the relevant content in the above embodiment for details. The specific functions and effects will not be repeated here.
[0140] Step S605: Generate a 3D model of the target product based on the optimized symbolic distance field; render the 3D model of the target product based on the optimized color information field, so as to display the rendered 3D model of the target product on the target product's display page.
[0141] In e-commerce scenarios, the product models displayed to users should be as realistic as possible to facilitate users' selection of the products they need, thereby reducing the costs associated with returns and exchanges due to users' lack of understanding of the products. Therefore, in this embodiment, after generating a 3D model (such as a white model) of the target product based on the optimized symbolic distance field, the 3D model of the target product is rendered based on the optimized color information field to obtain a color simulation model of the target product, which is closer to the real shape and color of the target product.
[0142] After obtaining the rendered 3D model of the target product, in response to the user's trigger operation of opening the target product's display page, the rendered 3D model of the target product is displayed on the target product's display page for the user to view. In addition, for the displayed 3D model of the target product, the user can rotate the 3D model in various directions by dragging, sliding and other operations, so that the user can view the target product (3D model) from various angles.
[0143] In this embodiment, a voxel model of the 3D scene where the target product is located is constructed. By storing the symbolic distance values at the voxel vertices in the voxel model, the symbolic distance field of the 3D spatial geometry is explicitly represented. By storing the color information at the voxel vertices in the voxel model, the color information field of the 3D spatial geometry is explicitly represented. The initial symbolic distance field and color information field are optimized using the observed image of the target product to reconstruct the symbolic distance field and color information field of the target product (in the 3D scene). Based on the optimized symbolic distance field, a 3D model of the target product is generated, and the 3D model of the target product is rendered based on the optimized color information field. This improves the reconstruction speed of the symbolic distance field and color information field of the target product, thereby improving the speed and efficiency of the 3D reconstruction of the target product.
[0144] Figure 7 This is a structural diagram of a three-dimensional reconstruction apparatus provided as an exemplary embodiment of this application. The apparatus provided in this embodiment is used to perform a three-dimensional reconstruction method. Figure 7 As shown, the 3D reconstruction device 70 includes: an image and camera parameter acquisition module 71, a field construction module 72, a field optimization module 73, and a model reconstruction module 74.
[0145] The image and camera parameter acquisition module 71 is used to acquire the observation image of the target object to be reconstructed, as well as the camera parameters corresponding to the observation image.
[0146] The field construction module 72 is used to construct a voxel model corresponding to the 3D scene based on the camera parameters, and initialize the symbolic distance field and color information field of the 3D scene based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model.
[0147] The field optimization module 73 is used to render pixels in the observed image based on camera parameters, symbolic distance field, and color information field, and determine the rendered color value of the pixel; calculate the loss based on the rendered color value and the actual color value of the pixel, and optimize the symbolic distance field and color information field based on the loss to obtain the optimized symbolic distance field and color information field.
[0148] The model reconstruction module 74 is used to generate a three-dimensional model of the target object based on the optimized symbolic distance field.
[0149] In an optional embodiment, when rendering pixels in the observed image based on camera parameters, a symbolic distance field, and a color information field, and determining the rendered color value of each pixel, the field optimization module 73 is further configured to:
[0150] Based on the camera parameters corresponding to the observed image, determine multiple sampling points on the ray corresponding to the pixel in the observed image; based on the symbolic distance field and color information field, determine the symbolic distance value and color information of the multiple sampling points; based on the symbolic distance value and color information of the multiple sampling points, render the pixel and determine the rendering color value of the pixel.
[0151] In an optional embodiment, when determining multiple sampling points on the ray corresponding to a pixel in the observed image based on the camera parameters corresponding to the observed image, the field optimization module 73 is further configured to:
[0152] Based on the camera parameters corresponding to the observed image, determine the ray corresponding to the pixel. The ray corresponding to the pixel is the ray with the camera origin as the endpoint and passing through the pixel. Determine the intersection point between the ray corresponding to the pixel and the voxel model, and determine the nearest intersection point that is closest to the endpoint of the ray, and the farthest intersection point that is farthest from the endpoint of the ray. Perform equidistant sampling between the nearest and farthest intersection points on the ray to obtain multiple sampling points.
[0153] In an optional embodiment, when determining the symbol distance values and color information of multiple sampling points based on the symbol distance field and the color information field, the field optimization module 73 is further configured to:
[0154] For each sampling point, based on the three-dimensional coordinates of the sampling point, multiple neighboring points of the sampling point are determined in the voxel vertices of the voxel model; based on the symbolic distance values of the multiple neighboring points in the symbolic distance field, trilinear interpolation is performed to determine the symbolic distance value of the sampling point; based on the color information of the multiple neighboring points in the color information field, trilinear interpolation is performed to determine the color information of the sampling point.
[0155] In an optional embodiment, when rendering pixels and determining the rendering color value of pixels based on the symbolic distance values and color information of multiple sampling points, the field optimization module 73 is further configured to:
[0156] The opacity at each sampling point is calculated based on the symbolic distance values of multiple sampling points. The transparency of the line segment from each sampling point to the ray endpoint is calculated based on the opacity at each sampling point. The rendered color value of the pixel is determined based on the opacity at each sampling point, the transparency of the line segment from each sampling point to the ray endpoint, the color information of each sampling point, and the pixel color rendering equation.
[0157] In an optional embodiment, when calculating the loss based on the rendered color value and the actual color value of a pixel, the field optimization module 73 is further configured to:
[0158] Calculate the luminance loss based on the rendered color value and the actual color value of each pixel.
[0159] In an optional embodiment, the field optimization module 73 is further configured to calculate at least one of the following losses:
[0160] Based on the symbolic distance values of the pixels in the current symbolic distance field, calculate the gradient at the voxel vertex in the voxel model, and calculate the regularization loss based on the gradient at the voxel vertex; calculate the smoothing loss based on the difference between the symbolic distance field at the voxel vertex in the current symbolic distance field and the mean of the symbolic distance field at the neighboring points of the same voxel vertex.
[0161] In an optional embodiment, when optimizing the symbol distance field and color information field based on the loss value, the field optimization module 73 is further configured to: if multiple losses are calculated, determine a comprehensive loss based on the multiple losses, and optimize the symbol distance field and color information field based on the comprehensive loss.
[0162] In an optional embodiment, when generating a 3D model of the target object based on the optimized symbolic distance field, the model reconstruction module 74 is further configured to:
[0163] If the target object is a scene, a 3D model of the scene is generated based on the optimized symbolic distance field; or, if the target object is an object, a 3D model of the object is generated based on the optimized symbolic distance field.
[0164] In an optional embodiment, after generating a 3D model of the target object based on the optimized symbolic distance field, the model reconstruction module 74 is further configured to:
[0165] The 3D model of the target object is rendered based on the optimized color information field.
[0166] The device provided in this embodiment can be used to execute the three-dimensional reconstruction method based on any of the above embodiments. The specific functions and technical effects that can be achieved will not be described in detail here.
[0167] Figure 8 A structural diagram of a three-dimensional reconstruction apparatus provided as another exemplary embodiment of this application. The apparatus provided in this embodiment is used to perform... Figure 6 The three-dimensional reconstruction method shown. Figure 8 As shown, the 3D reconstruction device 80 includes: a commodity data acquisition module 81, a field construction module 82, a field optimization module 83, and a model reconstruction and rendering module 84.
[0168] Among them, the commodity data acquisition module 81 is used to acquire observation images of the target commodity to be reconstructed, as well as the camera parameters corresponding to the observation images.
[0169] The field construction module 82 is used to construct a voxel model of the corresponding 3D scene based on the camera parameters, and initialize the symbolic distance field and color information field of the 3D scene based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model.
[0170] The field optimization module 83 is used to render pixels in the observed image based on camera parameters, symbolic distance field, and color information field, and determine the rendered color value of the pixel; calculate the loss based on the rendered color value and the actual color value of the pixel, and optimize the symbolic distance field and color information field based on the loss to obtain the optimized symbolic distance field and color information field.
[0171] The model reconstruction and rendering module 84 is used to generate a 3D model of the target product based on the optimized symbolic distance field; and to render the 3D model of the target product based on the optimized color information field, so as to display the rendered 3D model of the target product on the target product's display page.
[0172] In an optional embodiment, before rendering the pixels in the observed image based on camera parameters, symbolic distance field, and color information field, and determining the rendering color value of the pixels, the field optimization module 83 is further configured to: remove background information other than the target product from the observed image of the target product.
[0173] The device provided in this embodiment can be specifically used to perform the above-described... Figure 6 The specific functions and technical effects of the 3D reconstruction method shown will not be elaborated here.
[0174] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an example embodiment of this application. For example... Figure 9 As shown, the electronic device 90 includes a processor 901 and a memory 902 communicatively connected to the processor 901, the memory 902 storing computer-executed instructions.
[0175] The processor executes computer execution instructions stored in the memory to implement the solution provided in any of the above method embodiments. The specific functions and technical effects that can be achieved will not be elaborated here.
[0176] This application also provides a computer-readable storage medium storing computer-executable instructions. When executed by a processor, the computer-executable instructions are used to implement the solution provided in any of the above method embodiments. The specific functions and technical effects to be achieved are not described here.
[0177] This application also provides a computer program product, which includes a computer program stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program to cause the electronic device to perform the solution provided in any of the above method embodiments. The specific functions and technical effects that can be achieved are not described here.
[0178] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0179] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The sequence numbers are merely used to distinguish different operations, and the sequence number itself does not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. "Multiple" means two or more, unless otherwise explicitly specified.
[0180] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0181] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A three-dimensional reconstruction method, characterized in that, include: Acquire the observation image of the target object to be reconstructed, and the camera parameters corresponding to the observation image; Based on the camera parameters, a voxel model corresponding to the 3D scene is constructed. The symbolic distance field and color information field of the 3D scene are initialized based on the voxel model. The symbolic distance field contains the symbolic distance information at the voxel vertices in the voxel model, and the color information field contains the color information at the voxel vertices in the voxel model. Based on the camera parameters, the symbolic distance field, and the color information field, the pixels in the observed image are rendered, and the rendered color value of the pixel is determined. Calculate the photometric loss based on the rendered color value and the actual color value of the pixel. Based on the photometric loss and the regularization loss of the symbolic distance field, the symbolic distance field and the color information field are optimized to obtain the optimized symbolic distance field and color information field; the regularization loss is used to constrain the gradient magnitude at the voxel vertices in the symbolic distance field; A three-dimensional model of the target object is generated based on the optimized symbolic distance field.
2. The method according to claim 1, characterized in that, The step of rendering pixels in the observed image based on the camera parameters, the symbolic distance field, and the color information field, and determining the rendered color value of the pixel, includes: Based on the camera parameters corresponding to the observed image, determine multiple sampling points on the ray corresponding to the pixel in the observed image; Based on the symbol distance field and color information field, determine the symbol distance value and color information of the plurality of sampling points; Based on the symbolic distance values and color information of the multiple sampling points, the pixel is rendered, and the rendering color value of the pixel is determined.
3. The method according to claim 2, characterized in that, The step of determining multiple sampling points on the ray corresponding to the pixel in the observed image based on the camera parameters corresponding to the observed image includes: Based on the camera parameters corresponding to the observed image, the ray corresponding to the pixel is determined. The ray corresponding to the pixel is a ray that has the camera origin as its endpoint and passes through the pixel. Determine the intersection point between the ray corresponding to the pixel and the voxel model, and determine the nearest intersection point that is closest to the endpoint of the ray, and the farthest intersection point that is farthest from the endpoint of the ray; Multiple sampling points are obtained by sampling at equal intervals between the nearest and farthest intersection points on the ray.
4. The method according to claim 2, characterized in that, The step of determining the symbol distance values and color information of the plurality of sampling points based on the symbol distance field and color information field includes: For each sampling point, based on the three-dimensional coordinates of the sampling point, multiple neighborhood points of the sampling point are determined in the voxel vertices of the voxel model; Based on the symbol distance values of the plurality of neighboring points in the symbol distance field, trilinear interpolation is performed to determine the symbol distance value of the sampling point; Based on the color information of the multiple neighboring points in the color information field, trilinear interpolation is performed to determine the color information of the sampling point.
5. The method according to claim 2, characterized in that, The step of rendering the pixel based on the symbolic distance value and color information of the plurality of sampling points, and determining the rendering color value of the pixel, includes: The opacity at each sampling point is calculated based on the symbol distance values of the multiple sampling points. Calculate the transparency of the line segment from each sampling point to the ray endpoint based on the opacity at each sampling point. The rendered color value of the pixel is determined based on the opacity at each sampling point, the transparency of the line segment from each sampling point to the ray endpoint, the color information of each sampling point, and the pixel color rendering equation.
6. The method according to claim 1, characterized in that, Also includes: Based on the symbolic distance field values of the pixels in the current symbolic distance field, calculate the gradient at the voxel vertex in the voxel model, and calculate the regularization loss based on the gradient at the voxel vertex. The smoothing loss is calculated based on the difference between the mean of the symbolic distance field at the voxel vertex in the current symbolic distance field and the mean of the symbolic distance field at the neighboring points of the same voxel vertex.
7. The method according to claim 6, characterized in that, The step of optimizing the symbolic distance field and the color information field based on the photometric loss and the regularization loss of the symbolic distance field includes: The comprehensive loss is determined based on the photometric loss, the regularization loss of the symbolic distance field, and the smoothing loss; The symbol distance field and color information field are optimized based on the comprehensive loss.
8. The method according to claim 1, characterized in that, The step of generating a 3D model of the target object based on the optimized symbolic distance field includes: The target object is a scene, and a 3D model of the scene is generated based on the optimized symbolic distance field. or, If the target object is an object, then a 3D model of the object is generated based on the optimized symbolic distance field.
9. The method according to any one of claims 1-8, characterized in that, After generating the 3D model of the target object based on the optimized symbolic distance field, the process further includes: Based on the optimized color information field, a 3D model of the target object is rendered so as to display the rendered 3D model of the target object on the target object's display page.
10. A three-dimensional reconstruction method, characterized in that, include: Acquire observation images of the target product to be reconstructed, as well as the camera parameters corresponding to the observation images; A voxel model of the corresponding 3D scene is constructed based on the camera parameters, and the symbolic distance field and color information field of the 3D scene are initialized based on the voxel model. The symbolic distance field contains the symbolic distance information of the voxel vertices in the voxel model, and the color information field contains the color information of the voxel vertices in the voxel model. Based on the camera parameters, the symbolic distance field, and the color information field, the pixels in the observed image are rendered, and the rendered color value of the pixel is determined. Calculate the luminance loss based on the rendered color value and the actual color value of the pixel; Based on the photometric loss and the regularization loss of the symbolic distance field, the symbolic distance field and the color information field are optimized to obtain the optimized symbolic distance field and color information field; the regularization loss is used to constrain the gradient magnitude at the voxel vertices in the symbolic distance field; Based on the optimized symbolic distance field, a three-dimensional model of the target product is generated; Based on the optimized color information field, a 3D model of the target product is rendered so that the rendered 3D model of the target product can be displayed on the display page of the target product.
11. The method according to claim 10, characterized in that, Before determining the rendered color value of a pixel by rendering pixels in the observed image based on the camera parameters, the symbolic distance field, and the color information field, the method further includes: Remove background information other than the target product from the observed image of the target product.
12. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-11.