Image reconstruction method, device, electronic device, storage medium and program product

By extracting image feature information to determine the voxel density and feature information of the neural radiation field and generating a reconstructed image, the problems of high storage and computing costs in existing technologies are solved, and the ability to reconstruct and edit three-dimensional scenes from a single image is realized.

CN115018979BActive Publication Date: 2025-09-23SHANGHAI SENSETIME LINGANG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210587057.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-09-23
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

Existing technologies require multi-angle images and camera angle annotation data for 3D scene reconstruction, and can only model a single scene, resulting in high storage and computing costs, and the reconstructed images cannot be edited.

Method used

By acquiring the image to be processed, extracting the characteristic information of the object, determining the voxel density and characteristic information of the neural radiation field, generating a reconstructed image, and using a single image for implicit three-dimensional scene learning, the storage and computing costs are reduced, and image editing is supported.

Benefits of technology

It realizes the reconstruction of three-dimensional scenes based on a single perspective, reduces storage and computing costs, and the reconstructed images are editable, which expands the application scenarios of three-dimensional modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115018979B_ABST
    Figure CN115018979B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image reconstruction method, apparatus, electronic device, storage medium, and program product. The method comprises: acquiring an image to be processed; extracting first characteristic information of an object in the image to be processed; determining, based on the first characteristic information, a first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed; and generating, based on the first voxel density and the second characteristic information, a first reconstructed image corresponding to the image to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to an image reconstruction method, device, electronic device, storage medium, and program product. Background Art

[0002] Understanding the three-dimensional structure of objects in three-dimensional scenes (such as objects, human bodies, animals, etc.), modeling and editing three-dimensional scenes, and synthesizing images from new perspectives are long-term research topics in the field of three-dimensional reconstruction and are of great significance in technical fields such as computer vision. Summary of the Invention

[0003] The present disclosure provides an image reconstruction technology solution.

[0004] According to one aspect of the present disclosure, there is provided an image reconstruction method, comprising:

[0005] Get the image to be processed;

[0006] Extracting first feature information of an object in the image to be processed;

[0007] determining, based on the first characteristic information, a first voxel density and a second characteristic information of a neural radiation field corresponding to the image to be processed;

[0008] A first reconstructed image corresponding to the image to be processed is generated according to the first voxel density and the second feature information.

[0009] By acquiring an image to be processed, extracting first characteristic information of an object in the image to be processed, determining a first voxel density and a second characteristic information of a neural radiation field corresponding to the image to be processed based on the first characteristic information, and generating a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second characteristic information, it is possible to learn implicit three-dimensional scenes based on a single image of a single perspective, reduce the storage cost of the three-dimensional model and the computational cost of three-dimensional modeling, obtain a reconstructed image from a new angle, and edit the objects in the reconstructed image.

[0010] In a possible implementation, determining, based on the first feature information, the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed includes:

[0011] Get the camera's viewing direction information;

[0012] Obtaining location information of the reference point;

[0013] The first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined according to the first feature information, the viewing direction information and the position information of the reference point.

[0014] In this implementation, by obtaining the viewing direction information of the camera, the position information of the reference point is obtained, and based on the first feature information, the viewing direction information and the position information of the reference point, the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed are determined. This can more accurately determine the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed, thereby facilitating more accurate image reconstruction.

[0015] In a possible implementation, obtaining the viewing direction information of the camera includes:

[0016] Get the camera's viewing direction;

[0017] Position encoding is performed based on the viewing direction to obtain viewing direction information of the camera.

[0018] In this implementation, the viewing direction of the camera is obtained and position encoding is performed based on the viewing direction to obtain the viewing direction information of the camera, thereby obtaining higher-dimensional viewing direction information, which is conducive to achieving more accurate image reconstruction.

[0019] In a possible implementation, obtaining the position information of the reference point includes:

[0020] Obtain the three-dimensional coordinates of the reference point;

[0021] Position encoding is performed based on the three-dimensional coordinates of the reference point to obtain position information of the reference point.

[0022] In this implementation, the position information of the reference point is obtained by obtaining the three-dimensional coordinates of the reference point and performing position encoding based on the three-dimensional coordinates of the reference point, thereby obtaining the position information of the reference point of a higher dimension, which is conducive to achieving more accurate image reconstruction.

[0023] In a possible implementation, determining the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information, the viewing direction information, and the position information of the reference point includes:

[0024] determining a second voxel density and third feature information of a neural radiation field corresponding to the object based on the first feature information, the viewing direction information, and the position information of the reference point;

[0025] The first voxel density and the second characteristic information of the neural radiation field corresponding to the image to be processed are determined according to the second voxel density and the third characteristic information.

[0026] In this implementation, the second voxel density and third feature information of the neural radiation field corresponding to the object are determined based on the first feature information, the viewing direction information and the position information of the reference point, and the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined based on the second voxel density and the third feature information, thereby accurately determining the voxel density and feature information of the neural radiation field corresponding to the image to be processed.

[0027] In a possible implementation, generating a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information includes:

[0028] generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information;

[0029] The first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed.

[0030] In this implementation, a first feature map corresponding to the image to be processed is generated based on the first voxel density and the second feature information, and the first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed, thereby accelerating the rendering speed and improving the quality of the reconstructed image.

[0031] In a possible implementation, generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information includes:

[0032] determining, according to the first voxel density, a transparency value of a three-dimensional position in the three-dimensional space corresponding to the image to be processed;

[0033] determining a transmittance of the three-dimensional position according to the transparency value;

[0034] A first feature map corresponding to the image to be processed is generated according to the transparency value, the transmittance and the second feature information.

[0035] In this implementation, the transparency value of the three-dimensional position in the three-dimensional space corresponding to the image to be processed is determined based on the first voxel density, the transmittance of the three-dimensional position is determined based on the transparency value, and the first feature map corresponding to the image to be processed is generated based on the transparency value, the transmittance and the second feature information. In this way, the first feature map can be obtained quickly and accurately, which helps to improve the speed and accuracy of image reconstruction.

[0036] In a possible implementation manner, the first feature information includes first appearance feature information and first transformation feature information.

[0037] In this implementation, by extracting the first appearance feature information and the first transformation feature information of the object in the image to be processed, the information of the object in the image to be processed can be more accurately represented, thereby achieving more accurate three-dimensional reconstruction.

[0038] In a possible implementation manner, the first appearance feature information includes first comprehensive appearance feature information and first shape feature information.

[0039] In this implementation, by extracting the first comprehensive appearance feature information, first shape feature information and first transformation feature information of the object in the image to be processed, the information of the object in the image to be processed can be more accurately represented, thereby further improving the accuracy of three-dimensional reconstruction.

[0040] In a possible implementation, after generating a first reconstructed image corresponding to the image to be processed, the method further includes:

[0041] In response to an editing request for any object in the first reconstructed image, the object is edited to obtain an edited first reconstructed image.

[0042] In this implementation, since feature information of objects in the image to be processed is extracted separately, different objects in the image to be processed are independent and separable, so that objects in the first reconstructed image can be edited in response to an editing request.

[0043] In one possible implementation,

[0044] The extracting the first feature information of the object in the image to be processed includes: extracting the first feature information of the object in the image to be processed by a pre-trained neural network;

[0045] Determining the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed according to the first feature information includes: determining the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed according to the first feature information by the neural network;

[0046] Generating a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information includes: generating a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information through the neural network.

[0047] In this implementation, by acquiring an image to be processed, first characteristic information of an object in the image to be processed is extracted through a pre-trained neural network, and the first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed are determined by the neural network based on the first characteristic information, and a first reconstructed image corresponding to the image to be processed is generated by the neural network based on the first voxel density and the second characteristic information, thereby improving the accuracy and speed of image reconstruction of the image to be processed.

[0048] In a possible implementation, before extracting first feature information of an object in the image to be processed by using a pre-trained neural network, the method further includes:

[0049] Get training images;

[0050] Processing the training image through the neural network to obtain a second reconstructed image corresponding to the training image;

[0051] determining a value of a loss function of the neural network based on the second reconstructed image;

[0052] The neural network is trained according to the value of the loss function.

[0053] In this implementation, the neural network can be trained end-to-end. In this example, a training image is obtained, processed by the neural network to obtain a second reconstructed image corresponding to the training image, and the value of the neural network's loss function is determined based on the second reconstructed image. The neural network is then trained based on the value of the loss function, thereby enabling the neural network to learn the ability to perform image reconstruction.

[0054] In a possible implementation, the loss function includes a first loss function;

[0055] Determining a value of a loss function of the neural network according to the second reconstructed image includes:

[0056] The value of the first loss function is determined according to difference information between the second reconstructed image and the training image.

[0057] In this implementation, the value of the first loss function is determined based on the difference information between the second reconstructed image and the training image, and the neural network is trained based on the value of the first loss function. This allows the degree of difference between the reconstructed image obtained by the neural network and the original image to be reduced through training, thereby improving the accuracy of image reconstruction.

[0058] In one possible implementation, the loss function includes a second loss function;

[0059] Determining a value of a loss function of the neural network according to the second reconstructed image includes:

[0060] Determine the location information of the mask area;

[0061] generating a composite image according to the second reconstructed image, the training image, and the position information of the mask area;

[0062] A value of the second loss function is determined based on the synthesized image.

[0063] In this implementation, by determining the position information of the mask area, a composite image is generated based on the second reconstructed image, the training image and the position information of the mask area, and the value of the second loss function is determined based on the composite image. The neural network is trained based on the value of the second loss function, which can help the neural network learn the ability to reconstruct more accurate image detail information, thereby helping to reconstruct a clearer image.

[0064] In a possible implementation, generating a composite image according to the second reconstructed image, the training image, and the position information of the mask area includes:

[0065] Determining a first image to be synthesized based on the second reconstructed image and the position information of the mask area, wherein a size of the first image to be synthesized is the same as a size of the second reconstructed image, and in the first image to be synthesized, pixel values ​​of pixels outside the mask area are the same as those of the second reconstructed image, and pixel values ​​of pixels within the mask area are empty;

[0066] Determining a second image to be synthesized based on the position information of the training image and the mask area, wherein a size of the second image to be synthesized is the same as a size of the training image, and in the second image to be synthesized, pixel values ​​of pixels within the mask area are the same as those of the training image, and pixel values ​​of pixels outside the mask area are blank;

[0067] The synthesized image is generated according to the first image to be synthesized and the second image to be synthesized.

[0068] In this implementation, the first image to be synthesized is determined based on the second reconstructed image and the position information of the mask area, and the second image to be synthesized is determined based on the training image and the position information of the mask area, wherein the size of the first image to be synthesized is the same as the size of the second reconstructed image, and in the first image to be synthesized, the pixel values ​​of the pixels outside the mask area are the same as those of the second reconstructed image, and the pixel values ​​of the pixels within the mask area are empty, the size of the second image to be synthesized is the same as the size of the training image, and in the second image to be synthesized, the pixel values ​​of the pixels within the mask area are the same as those of the training image, and the pixel values ​​of the pixels outside the mask area are empty, and a synthetic image is generated based on the first image to be synthesized and the second image to be synthesized, and the neural network is trained based on the synthetic image thus generated, which can achieve self-supervised training and help the neural network learn the ability to reconstruct more accurate image detail information.

[0069] In a possible implementation, the second loss function includes a generation loss function and an adversarial loss function;

[0070] Determining a value of the second loss function according to the synthesized image includes:

[0071] Determining a value of a generation loss function based on the synthesized image;

[0072] The value of the adversarial loss function is determined according to the difference information between the synthetic image and the training image.

[0073] In this implementation, by determining the value of the generation loss function based on the synthetic image, determining the value of the adversarial loss function based on the difference information between the synthetic image and the training image, and training the neural network based on the values ​​of the generation loss function and the adversarial loss function, the neural network can learn the ability to reconstruct more accurate image detail information, thereby helping to reconstruct a clearer image.

[0074] According to one aspect of the present disclosure, there is provided an image reconstruction apparatus, comprising:

[0075] A first acquisition module is used to acquire an image to be processed;

[0076] An extraction module, configured to extract first feature information of an object in the image to be processed;

[0077] a first determining module, configured to determine, based on the first characteristic information, a first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed;

[0078] A generating module is used to generate a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information.

[0079] In a possible implementation, the first determining module is configured to:

[0080] Get the camera's viewing direction information;

[0081] Obtaining location information of the reference point;

[0082] The first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined according to the first feature information, the viewing direction information and the position information of the reference point.

[0083] In a possible implementation, the first determining module is configured to:

[0084] Get the camera's viewing direction;

[0085] Position encoding is performed based on the viewing direction to obtain viewing direction information of the camera.

[0086] In a possible implementation, the first determining module is configured to:

[0087] Obtain the three-dimensional coordinates of the reference point;

[0088] Position encoding is performed based on the three-dimensional coordinates of the reference point to obtain position information of the reference point.

[0089] In a possible implementation, the first determining module is configured to:

[0090] determining a second voxel density and third feature information of a neural radiation field corresponding to the object based on the first feature information, the viewing direction information, and the position information of the reference point;

[0091] The first voxel density and the second characteristic information of the neural radiation field corresponding to the image to be processed are determined according to the second voxel density and the third characteristic information.

[0092] In a possible implementation, the generating module is configured to:

[0093] generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information;

[0094] The first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed.

[0095] In a possible implementation, the generating module is configured to:

[0096] determining, according to the first voxel density, a transparency value of a three-dimensional position in the three-dimensional space corresponding to the image to be processed;

[0097] determining a transmittance of the three-dimensional position according to the transparency value;

[0098] A first feature map corresponding to the image to be processed is generated according to the transparency value, the transmittance and the second feature information.

[0099] In a possible implementation manner, the first feature information includes first appearance feature information and first transformation feature information.

[0100] In a possible implementation manner, the first appearance feature information includes first comprehensive appearance feature information and first shape feature information.

[0101] In a possible implementation, the apparatus further includes:

[0102] The editing module is configured to edit any object in the first reconstructed image in response to an editing request for the object, and obtain an edited first reconstructed image.

[0103] In one possible implementation,

[0104] The extraction module is used to: extract first feature information of an object in an image to be processed through a pre-trained neural network;

[0105] The first determining module is configured to determine, through the neural network and based on the first characteristic information, a first voxel density and a second characteristic information of a neural radiation field corresponding to the image to be processed;

[0106] The generating module is used to generate a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information through the neural network.

[0107] In a possible implementation, the apparatus further includes:

[0108] A second acquisition module is used to acquire a training image;

[0109] a processing module, configured to process the training image through the neural network to obtain a second reconstructed image corresponding to the training image;

[0110] a second determining module, configured to determine a value of a loss function of the neural network according to the second reconstructed image;

[0111] A training module is used to train the neural network according to the value of the loss function.

[0112] In a possible implementation, the loss function includes a first loss function;

[0113] The second determining module is used for:

[0114] The value of the first loss function is determined according to difference information between the second reconstructed image and the training image.

[0115] In one possible implementation, the loss function includes a second loss function;

[0116] The second determining module is used for:

[0117] Determine the location information of the mask area;

[0118] generating a composite image according to the second reconstructed image, the training image, and the position information of the mask area;

[0119] A value of the second loss function is determined based on the synthesized image.

[0120] In a possible implementation, the second determining module is configured to:

[0121] Determining a first image to be synthesized based on the second reconstructed image and the position information of the mask area, wherein a size of the first image to be synthesized is the same as a size of the second reconstructed image, and in the first image to be synthesized, pixel values ​​of pixels outside the mask area are the same as those of the second reconstructed image, and pixel values ​​of pixels within the mask area are empty;

[0122] Determining a second image to be synthesized based on the position information of the training image and the mask area, wherein a size of the second image to be synthesized is the same as a size of the training image, and in the second image to be synthesized, pixel values ​​of pixels within the mask area are the same as those of the training image, and pixel values ​​of pixels outside the mask area are blank;

[0123] The synthesized image is generated according to the first image to be synthesized and the second image to be synthesized.

[0124] In a possible implementation, the second loss function includes a generation loss function and an adversarial loss function;

[0125] The second determining module is used for:

[0126] Determining a value of a generation loss function based on the synthesized image;

[0127] The value of the adversarial loss function is determined according to the difference information between the synthetic image and the training image.

[0128] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0129] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.

[0130] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0131] In an embodiment of the present disclosure, by acquiring an image to be processed, first feature information of an object in the image to be processed is extracted, and based on the first feature information, a first voxel density and second feature information of a neural radiation field corresponding to the image to be processed are determined, and based on the first voxel density and the second feature information, a first reconstructed image corresponding to the image to be processed is generated. This enables implicit three-dimensional scene learning based on a single image from a single perspective, reduces the storage cost of the three-dimensional model and the computational cost of three-dimensional modeling, and obtains a reconstructed image from a new angle, and the object in the reconstructed image can be edited.

[0132] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0133] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0134] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0135] Figure 1 A flowchart of an image reconstruction method provided by an embodiment of the present disclosure is shown.

[0136] Figure 2aA schematic diagram showing a mask image in the image reconstruction method provided by an embodiment of the present disclosure.

[0137] Figure 2b A schematic diagram showing a first image to be synthesized in the image reconstruction method provided by an embodiment of the present disclosure.

[0138] Figure 2c A schematic diagram showing a second image to be synthesized in the image reconstruction method provided by an embodiment of the present disclosure.

[0139] Figure 2d A schematic diagram showing a synthesized image in the image reconstruction method provided by an embodiment of the present disclosure.

[0140] Figure 3 A schematic diagram showing a neural network in the image reconstruction method provided by an embodiment of the present disclosure.

[0141] Figure 4a A schematic diagram showing an image to be processed in the image reconstruction method provided by an embodiment of the present disclosure.

[0142] Figures 4b to 4d A schematic diagram showing a reconstructed image of a new perspective obtained by reconstructing an image to be processed in the image reconstruction method provided by an embodiment of the present disclosure.

[0143] Figure 5 A block diagram of an image reconstruction device provided by an embodiment of the present disclosure is shown.

[0144] Figure 6 A block diagram of an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0145] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0146] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0147] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0148] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0149] In related technologies, the establishment of three-dimensional scenes usually requires complex three-dimensional modeling and relies on the input of three-dimensional data such as point cloud data and three-dimensional mesh graphs, so the application scenarios are relatively limited.

[0150] Neural Radiance Field technology combines neural networks with the construction of three-dimensional space. Through an implicit three-dimensional spatial representation, it models a single scene based on multiple two-dimensional input images. However, related technologies rely on images of the same scene from multiple angles and corresponding camera angle annotation data, and can only model a single scene.

[0151] In order to solve technical problems similar to those described above, the embodiments of the present disclosure provide an image reconstruction method, device, electronic device, storage medium and program product, which obtains the image to be processed, extracts the first feature information of the object in the image to be processed, determines the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information, and generates a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information. This enables implicit three-dimensional scene learning based on a single image of a single perspective, reduces the storage cost of the three-dimensional model and the computational cost of three-dimensional modeling, and obtains a reconstructed image at a new angle, and the object in the reconstructed image can be edited.

[0152] The image reconstruction method provided by the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0153] Figure 1 A flow chart of the image reconstruction method provided by an embodiment of the present disclosure is shown. In one possible implementation, the execution subject of the image reconstruction method may be an image reconstruction device. For example, the image reconstruction method may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device, etc. In some possible implementations, the image reconstruction method may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1As shown, the image reconstruction method includes steps S11 to S14.

[0154] In step S11, an image to be processed is obtained.

[0155] In step S12, first feature information of the object in the image to be processed is extracted.

[0156] In step S13, a first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed are determined based on the first characteristic information.

[0157] In step S14, a first reconstructed image corresponding to the image to be processed is generated according to the first voxel density and the second feature information.

[0158] In the embodiments of the present disclosure, the image to be processed may represent an image that needs to be 3D reconstructed. In one possible implementation, the image to be processed may be a 2D image. With this implementation, a 3D scene can be reconstructed based on a single 2D image, thereby expanding the application scenarios of 3D scene modeling. In another possible implementation, the image to be processed may also be a 3D image. In one example, the image to be processed may be denoted as I o1 .

[0159] In an embodiment of the present disclosure, the first feature information of all or part of the objects in the image to be processed can be extracted. In one possible implementation, the first feature information of all objects in the image to be processed can be extracted. In another possible implementation, the first feature information of one or more specified types of objects in the image to be processed can be extracted. The objects can be objects, people, animals, scenes, etc., which are not limited here. In one possible implementation, the number of objects in the image to be processed can be greater than or equal to 2, and the objects in the image to be processed can include: at least one of the objects, people, and animals in the image to be processed, and the scene of the image to be processed. The first feature information of any object in the image to be processed can represent the feature information corresponding to the object. The first feature information can correspond one-to-one to the objects in the image to be processed.

[0160] In one possible implementation, the first feature information includes first appearance feature information and first transformation feature information. The first appearance feature information of any object in the image to be processed may be any information that can represent the appearance features of the object. The first transformation feature information of any object in the image to be processed may be any information that can represent the transformation features of the object. In this implementation, by extracting the first appearance feature information and first transformation feature information of the object in the image to be processed, the information of the object in the image to be processed can be more accurately represented, thereby achieving more accurate three-dimensional reconstruction.

[0161] As an example of this implementation, the first appearance feature information includes first comprehensive appearance feature information and first shape feature information. In this example, the first comprehensive appearance feature information of any object in the image to be processed may be any information that can represent the visual features of the object. The first shape feature information of any object in the image to be processed may be any information that can represent the shape features of the object. In one example, the first comprehensive appearance feature information of the i-th object in the image to be processed can be recorded as The first shape feature information of the i-th object can be recorded as The first appearance feature information of the i-th object can be recorded as Where i∈{1,...,N}, N represents the number of objects in the image to be processed, and N≥1. One or more objects may be scenes in the image to be processed. For example, the scene in the image to be processed may be represented by the first feature information corresponding to one of the N objects.

[0162] In this example, by extracting the first comprehensive appearance feature information, first shape feature information and first transformation feature information of the object in the image to be processed, the information of the object in the image to be processed can be more accurately represented, thereby further improving the accuracy of three-dimensional reconstruction.

[0163] As an example of this implementation, the first transformed feature information of the i-th object can be recorded as (t i ,s i ,r i ). Among them, t i can represent the displacement parameter of the i-th object, and t i It can include displacement parameters in three directions; i Can represent the scaling parameter of the i-th object; r i Can represent the rotation parameters of the i-th object, and r i The Euler angles in three directions may be included. In one example, the first transformation feature information of the i-th object may be transformation feature information in a world coordinate system.

[0164] In one example, the first feature information of the i-th object can be recorded as

[0165] In another possible implementation, the first feature information may be represented by comprehensive visual feature information without distinguishing between appearance feature information and transformation feature information.

[0166] In one possible implementation, the first feature information of the object in the image to be processed can be extracted by an object-aware scene encoder. As an example of this implementation, the object-aware scene encoder can be composed of a Convolutional Neural Network (CNN) backbone and a Transformer (converter) structure head. The Transformer adopts an encoder-decoder architecture. For the image to be processed I o1 , the object-aware scene encoder can predict the first feature information of N objects.

[0167] In the embodiment of the present disclosure, the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed can be converted based on the first feature information. The first voxel density can represent the voxel density of the neural radiation field corresponding to the image to be processed, and the second feature information can represent the feature information of the neural radiation field corresponding to the image to be processed. In an example, the first voxel density of the neural radiation field corresponding to the image to be processed can be recorded as σ, and the second feature information of the neural radiation field corresponding to the image to be processed can be recorded as

[0168] In one possible implementation, determining the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information includes: obtaining the viewing direction information of the camera; obtaining the position information of the reference point; and determining the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information, the viewing direction information and the position information of the reference point.

[0169] In this implementation, the camera may refer to a camera that captures the image to be processed. The camera's viewing direction information may be any information that can represent the camera's viewing direction. For example, the camera's viewing direction may be directly used as the viewing direction information. In another example, the camera's viewing direction may be processed to obtain the viewing direction information.

[0170] A reference point may represent a point with known coordinates. There may be multiple reference points. The position information of any reference point may be any information that can represent the position of the reference point. For example, the coordinates of the reference point may be directly used as the position information of the reference point. In another example, the coordinates of the reference point may be processed to obtain the position information of the reference point.

[0171] In this implementation, by obtaining the viewing direction information of the camera, the position information of the reference point is obtained, and based on the first feature information, the viewing direction information and the position information of the reference point, the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed are determined. This can more accurately determine the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed, thereby facilitating more accurate image reconstruction.

[0172] As an example of this implementation, obtaining the camera's viewing direction information includes: obtaining the camera's viewing direction; and performing position encoding based on the viewing direction to obtain the camera's viewing direction information. In one example, the camera's viewing direction can be predicted using an object-aware scene encoder. In other examples, the camera's viewing direction can also be obtained through methods such as manual annotation.

[0173] In one example, the camera's viewing angle direction d can be obtained, and the camera's viewing angle direction d can be projected and transformed to obtain the transformed viewing angle direction The direction of the transformed view can be Perform position encoding to obtain the camera's viewing direction information Where γ() can be a positional encoding as shown in Equation 1:

[0174] γ(v)=(sin(2 0 πv),cos(2 0 πv),...,sin(2 L πv),cos(2 L πv)) Equation 1.

[0175] In one example, in position encoding of the viewing direction of the camera, L may be equal to 4.

[0176] In this example, the viewing direction of the camera is obtained and position encoding is performed based on the viewing direction to obtain the viewing direction information of the camera, thereby obtaining higher-dimensional viewing direction information, which is conducive to achieving more accurate image reconstruction.

[0177] As an example of this implementation, obtaining the position information of the reference point includes: obtaining the three-dimensional coordinates of the reference point; performing position encoding based on the three-dimensional coordinates of the reference point to obtain the position information of the reference point. In one example, for any reference point, the three-dimensional coordinate x of the reference point in the world coordinate system can be obtained. The three-dimensional coordinate x of the reference point in the world coordinate system can be converted to a coordinate system centered on the reference point to obtain the converted three-dimensional coordinate corresponding to the reference point. The converted three-dimensional coordinates corresponding to the reference point Perform position encoding to obtain the position information of the reference point In one example, the transformed three-dimensional coordinates corresponding to the reference point can be obtained using Formula 1: Perform position encoding. In the position encoding of the three-dimensional coordinates of the reference point, L may be equal to 10.

[0178] In this example, the position information of the reference point is obtained by obtaining the three-dimensional coordinates of the reference point and performing position encoding based on the three-dimensional coordinates of the reference point, thereby obtaining the position information of the reference point of higher dimension, which is conducive to achieving more accurate image reconstruction.

[0179] As an example of this implementation method, the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined based on the first feature information, the viewing direction information and the position information of the reference point, including: determining the second voxel density and third feature information of the neural radiation field corresponding to the object based on the first feature information, the viewing direction information and the position information of the reference point; determining the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the second voxel density and the third feature information.

[0180] In this example, the second voxel density can represent the voxel density of the neural radiation field corresponding to the object, and the third feature information can represent the feature information of the neural radiation field corresponding to the object. For the i-th object among the N objects in the image to be processed, the second voxel density σ of the neural radiation field corresponding to the i-th object can be determined based on the first feature information of the i-th object, the viewing direction information of the camera, and the position information of the reference point. i and the third feature information f i .

[0181] In one example, for the scene of the N objects, a scene decoder can be used to process the first feature information of the scene, the camera's viewing direction information, and the position information of the reference point to obtain the second voxel density and third feature information of the neural radiation field corresponding to the scene. The scene decoder can adopt a multi-layer perceptron (MLP) structure or other decoder structures, which are not limited here.

[0182] In one example, for any object other than the scene in the N objects, an object decoder can be used to process the first feature information of the object, the camera's viewing direction information, and the position information of the reference point to obtain the second voxel density and third feature information of the neural radiation field corresponding to the object. The object decoder can adopt a multi-layer perceptron structure or other decoder structures, which are not limited here.

[0183] In one example, Equation 2 can be used to determine the first voxel density σ of the neural radiation field corresponding to the image to be processed:

[0184]

[0185] In another example, a weighted sum of the second voxel densities of the neural radiation fields corresponding to the objects in the image to be processed may be calculated to obtain the first voxel density of the neural radiation fields corresponding to the image to be processed. In this example, the weights of the second voxel densities of the neural radiation fields corresponding to different objects may be different.

[0186] In one example, Formula 3 can be used to determine the second characteristic information of the neural radiation field corresponding to the image to be processed:

[0187] In this example, the second voxel density and third feature information of the neural radiation field corresponding to the object are determined based on the first feature information, the viewing direction information and the position information of the reference point, and the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined based on the second voxel density and the third feature information, thereby accurately determining the voxel density and feature information of the neural radiation field corresponding to the image to be processed.

[0188] In the embodiment of the present disclosure, a first reconstructed image corresponding to the image to be processed can be rendered based on the first voxel density and the second feature information. The first reconstructed image can represent the reconstructed image corresponding to the image to be processed. In one example, the first reconstructed image can be denoted as I r1 If the image to be processed is an RGB (Red–Green–Blue) image, the first reconstructed image may be an RGB image or a grayscale image. If the image to be processed is a grayscale image, the first reconstructed image may also be a grayscale image. The size of the first reconstructed image may be the same as that of the image to be processed.

[0189] In a possible implementation, generating a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information includes: generating a first feature map corresponding to the image to be processed based on the first voxel density and the second feature information; and upsampling the first feature map to obtain the first reconstructed image corresponding to the image to be processed. In one example, the first voxel density σ and the second feature information can be used to generate a first reconstructed image corresponding to the image to be processed. Use voxel rendering to render the first feature map f corresponding to the image to be processed out The first feature map f out Perform upsampling to obtain the first reconstructed image I corresponding to the image to be processed r1 For example, a neural renderer (NeuralRender) can be used to render the first feature map f out Perform upsampling to obtain the first reconstructed image I corresponding to the image to be processed r1 In this implementation, a first feature map corresponding to the image to be processed is generated based on the first voxel density and the second feature information, and the first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed, thereby accelerating the rendering speed and improving the quality of the reconstructed image.

[0190] As an example of this implementation method, generating a first feature map corresponding to the image to be processed based on the first voxel density and the second feature information includes: determining a transparency value of a three-dimensional position in the three-dimensional space corresponding to the image to be processed based on the first voxel density; determining a transmittance ratio of the three-dimensional position based on the transparency value; and generating a first feature map corresponding to the image to be processed based on the transparency value, the transmittance ratio and the second feature information.

[0191] In one example, the transparency value α of the j-th three-dimensional position on the light beam in the three-dimensional space corresponding to the image to be processed can be determined using Formula 4: j :

[0192]

[0193] Among them, x j represents the three-dimensional coordinates of the j-th three-dimensional position, σ j represents the voxel density of the j-th three-dimensional position, where 1≤j≤N s , N s Indicates the total number of three-dimensional positions on the light beam in the three-dimensional space corresponding to the image to be processed.

[0194] In one example, the transmittance τ of the j-th three-dimensional position in the three-dimensional space corresponding to the image to be processed can be determined using Formula 5: j :

[0195]

[0196] Among them, α k Represents the transparency value of the kth three-dimensional position on the light beam in the three-dimensional space corresponding to the image to be processed.

[0197] In one example, Equation 6 can be used to obtain the first feature map f corresponding to the image to be processed: out :

[0198]

[0199] in, Represents the feature information of the j-th three-dimensional position.

[0200] In this example, the transparency value of the three-dimensional position in the three-dimensional space corresponding to the image to be processed is determined based on the first voxel density, the transmittance of the three-dimensional position is determined based on the transparency value, and the first feature map corresponding to the image to be processed is generated based on the transparency value, the transmittance and the second feature information. In this way, the first feature map can be obtained quickly and accurately, which helps to improve the speed and accuracy of image reconstruction.

[0201] In one possible implementation, the extracting first feature information of the object in the image to be processed includes: extracting the first feature information of the object in the image to be processed through a pre-trained neural network; determining the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information includes: determining the first voxel density and second feature information of the neural radiation field corresponding to the image to be processed based on the first feature information through the neural network; generating the first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information includes: generating the first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information through the neural network.

[0202] In one example, the object-aware scene encoder in the neural network can be used to extract first feature information of N objects in the image to be processed, wherein the N objects include the scene in the image to be processed and objects other than the scene in the image to be processed.

[0203] In one example, for the scene in the N objects, the first feature information of the scene can be processed by the scene decoder in the neural network to obtain the second voxel density and third feature information of the neural radiation field corresponding to the scene.

[0204] In one example, for any object other than the scene among the N objects, the first feature information of the object can be processed by the object decoder in the neural network to obtain the second voxel density and third feature information of the neural radiation field corresponding to the object.

[0205] In one example, a neural renderer in the neural network may generate a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information.

[0206] In this implementation, by acquiring an image to be processed, first characteristic information of an object in the image to be processed is extracted through a pre-trained neural network, and the first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed are determined by the neural network based on the first characteristic information, and a first reconstructed image corresponding to the image to be processed is generated by the neural network based on the first voxel density and the second characteristic information, thereby improving the accuracy and speed of image reconstruction of the image to be processed.

[0207] As an example of this implementation, before extracting the first feature information of the object in the image to be processed by the pre-trained neural network, the method further includes: obtaining a training image; processing the training image by the neural network to obtain a second reconstructed image corresponding to the training image; determining the value of the loss function of the neural network based on the second reconstructed image; and training the neural network based on the value of the loss function. In this example, the training image can be an image without labeled data. That is, in this example, the neural network can perform self-supervised training based on an unlabeled training image set. The training image can be a two-dimensional image. In one example, the training image can be denoted as I o2 In this example, the second reconstructed image may represent a reconstructed image corresponding to the training image. The size of the second reconstructed image may be the same as the size of the training image. In one example, the second reconstructed image may be denoted as I r2 .

[0208] In one example, fourth feature information can be extracted for n objects in a training image, where n is greater than or equal to 2. The fourth feature information can represent feature information of the objects in the training image. The n objects in the training image can include: the scene of the training image, and at least one of an object, a person, or an animal in the training image. The fourth feature information can correspond one-to-one with each object in the training image. The fourth feature information can include second appearance feature information and second change feature information, where the second appearance feature information can include second comprehensive appearance feature information and second shape feature information. For any object in the training image, the third voxel density and fifth feature information of the neural radiation field corresponding to the object can be determined based on the camera's viewing direction information, the position of a reference point, and the fourth feature information of the object. The fourth voxel density and sixth feature information of the neural radiation field corresponding to the training image can be determined based on the third voxel density and fifth feature information of the neural radiation field corresponding to each object in the training image. A second feature map corresponding to the training image can be generated based on the fourth voxel density and sixth feature information. The second feature map can be upsampled to obtain a second reconstructed image corresponding to the training image.

[0209] In this example, the neural network can be trained end-to-end. In this example, a training image is obtained, the training image is processed by the neural network to obtain a second reconstructed image corresponding to the training image, the value of the neural network's loss function is determined based on the second reconstructed image, and the neural network is trained based on the value of the loss function, thereby enabling the neural network to learn the ability to perform image reconstruction.

[0210] In one example, the loss function includes a first loss function; determining the value of the loss function of the neural network based on the second reconstructed image includes: determining the value of the first loss function based on difference information between the second reconstructed image and the training image.

[0211] In one example, the first loss function can be determined using Equation 7: Value:

[0212]

[0213] In this example, the training objective of the neural network may include minimizing the difference information between the second reconstructed image and the training image, that is, minimizing the consistency loss between the second reconstructed image and the training image. In this example, by determining the value of a first loss function based on the difference information between the second reconstructed image and the training image, and training the neural network based on the value of the first loss function, the difference between the reconstructed image obtained by the neural network and the original image can be reduced through training, thereby improving the accuracy of image reconstruction.

[0214] In one example, the loss function includes a second loss function; determining the value of the loss function of the neural network based on the second reconstructed image includes: determining the position information of the mask area; generating a composite image based on the second reconstructed image, the training image and the position information of the mask area; and determining the value of the second loss function based on the composite image.

[0215] The position information of the masked area can be represented by data in the form of a mask image, a two-dimensional matrix, or other data formats, without limitation. The size of the masked area is smaller than the size of the second reconstructed image, and the shape of the masked area can be arbitrary. When the position information of the masked area is represented by a mask image, the size of the mask image can be the same as the size of the second reconstructed image. In the mask image, the pixel value of the masked area can be 255, and the pixel value of the unmasked area can be 0. In this example, the position information of the masked area can be randomly determined.

[0216] In this example, by determining the position information of the mask area, a composite image is generated based on the second reconstructed image, the training image and the position information of the mask area, and the value of the second loss function is determined based on the composite image. The neural network is trained based on the value of the second loss function, which can help the neural network learn the ability to reconstruct more accurate image detail information, thereby helping to reconstruct a clearer image.

[0217] In one example, generating a composite image based on the second reconstructed image, the training image, and the positional information of the mask region includes: determining a first image to be synthesized based on the second reconstructed image and the positional information of the mask region, wherein the size of the first image to be synthesized is the same as the size of the second reconstructed image, and in the first image to be synthesized, the pixel values ​​of pixels outside the mask region are the same as those of the second reconstructed image, and the pixel values ​​of pixels within the mask region are null; determining a second image to be synthesized based on the training image and the positional information of the mask region, wherein the size of the second image to be synthesized is the same as the size of the training image, and in the second image to be synthesized, the pixel values ​​of pixels within the mask region are the same as those of the training image, and the pixel values ​​of pixels outside the mask region are null; and generating the composite image based on the first and second images to be synthesized. The first image to be synthesized may represent an image to be synthesized generated based on the second reconstructed image and the positional information of the mask region, and the second image to be synthesized may represent an image to be synthesized generated based on the training image and the positional information of the mask region.

[0218] Figure 2a A schematic diagram showing a mask image in the image reconstruction method provided by an embodiment of the present disclosure. Figure 2b A schematic diagram showing a first image to be synthesized in the image reconstruction method provided by an embodiment of the present disclosure. Figure 2c A schematic diagram showing a second image to be synthesized in the image reconstruction method provided by an embodiment of the present disclosure. Figure 2d A schematic diagram showing a synthesized image in the image reconstruction method provided by an embodiment of the present disclosure.

[0219] For example, the synthetic image I can be determined using Equation 8. m :

[0220] I m =I o2 *mask+I r2 *(1-mask) Equation 8,

[0221] Among them, mask represents the mask image.

[0222] In the above example, the first and second images to be synthesized can be synthesized by using the second image to be synthesized as the upper layer and the first image to be synthesized as the lower layer. Alternatively, the first and second images to be synthesized can be synthesized by using the second image to be synthesized as the lower layer and the first image to be synthesized as the upper layer.

[0223] In the above example, the first image to be synthesized is determined based on the second reconstructed image and the position information of the mask area, and the second image to be synthesized is determined based on the position information of the training image and the mask area, wherein the size of the first image to be synthesized is the same as the size of the second reconstructed image, and in the first image to be synthesized, the pixel values ​​of the pixels outside the mask area are the same as those of the second reconstructed image, and the pixel values ​​of the pixels within the mask area are empty, the size of the second image to be synthesized is the same as the size of the training image, and in the second image to be synthesized, the pixel values ​​of the pixels within the mask area are the same as those of the training image, and the pixel values ​​of the pixels outside the mask area are empty, and a synthetic image is generated based on the first image to be synthesized and the second image to be synthesized, and the neural network is trained based on the synthetic image thus generated, which can achieve self-supervised training and help the neural network learn the ability to reconstruct more accurate image detail information.

[0224] In another example, generating a composite image based on the second reconstructed image, the training image and the position information of the mask area includes: determining a first image to be synthesized based on the second reconstructed image and the position information of the mask area, wherein the size of the first image to be synthesized is the same as the size of the second reconstructed image, and in the first image to be synthesized, the pixel values ​​of the pixels within the mask area are the same as those of the second reconstructed image, and the pixel values ​​of the pixels outside the mask area are empty; determining a second image to be synthesized based on the training image and the position information of the mask area, wherein the size of the second image to be synthesized is the same as the size of the training image, and in the second image to be synthesized, the pixel values ​​of the pixels outside the mask area are the same as those of the training image, and the pixel values ​​of the pixels within the mask area are empty; generating the composite image based on the first image to be synthesized and the second image to be synthesized.

[0225] In one example, the second loss function includes a generation loss function and an adversarial loss function; determining the value of the second loss function based on the synthetic image includes: determining the value of the generation loss function based on the synthetic image; and determining the value of the adversarial loss function based on the difference information between the synthetic image and the training image.

[0226] In this example, the generation loss function can represent the loss function corresponding to the generator. The training objectives of the generator can include deceiving the discriminator and minimizing the difference between the synthesized image and the training image. In one example, the generation loss function can be determined using Equation 9. Value:

[0227]

[0228] in,

[0229] In this example, the adversarial loss function can represent the loss function corresponding to the discriminator. The training purpose of the discriminator can be to distinguish the synthetic image from the training image. In one example, the adversarial loss function can be determined using Equation 10. Value:

[0230]

[0231] in, is the regularization loss, λ reg for The corresponding weights. In one example, λ reg =10.

[0232] In this example, by determining the value of the generation loss function based on the synthetic image, determining the value of the adversarial loss function based on the difference information between the synthetic image and the training image, and training the neural network based on the values ​​of the generation loss function and the adversarial loss function, the neural network can learn the ability to reconstruct more accurate image detail information, thereby helping to reconstruct a clearer image.

[0233] In one example, the loss function corresponding to the neural network can be determined using Equation 11: Value:

[0234]

[0235] Among them, λ e1 express The corresponding weight, λ g express The corresponding weight, λ d express The corresponding weight.

[0236] In one possible implementation, generating a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information includes: obtaining information of a specified perspective; and generating the first reconstructed image corresponding to the image to be processed based on the first voxel density, the second feature information, and the specified perspective. In this implementation, the specified perspective may be different from the original perspective corresponding to the image to be processed. That is, by employing this implementation, an image with a new perspective can be reconstructed. The reconstructed image with a new perspective can be used to enhance training data for neural networks, enhance the functionality of photography software, and so on.

[0237] In one possible implementation, after generating a first reconstructed image corresponding to the image to be processed, the method further includes: responding to an edit request for any object in the first reconstructed image, editing the object to obtain an edited first reconstructed image. The edits performed may include at least one of moving, enlarging, reducing, and rotating the object. In this implementation, because feature information of the objects in the image to be processed is extracted separately, the different objects in the image to be processed are independent and separable, enabling editing of the objects in the first reconstructed image in response to an edit request.

[0238] In a possible implementation, after generating a first reconstructed image corresponding to the image to be processed, the method further includes: rendering a three-dimensional scene based on the first reconstructed image, for example, a game scene.

[0239] The image reconstruction method provided by the embodiments of the present disclosure can be applied to application scenarios such as computer vision, artificial intelligence, neural radiation field, three-dimensional reconstruction, AR (Augmented Reality), image editing, video editing, and games.

[0240] The image reconstruction method provided by the embodiment of the present disclosure is described below through a specific application scenario.

[0241] In this application scenario, the neural network can be pre-trained. Figure 3 Schematic diagram of a neural network in the image reconstruction method provided by an embodiment of the present disclosure. Figure 3 As shown, the neural network may include an object-aware scene encoder and a neural radiance field-based renderer. The neural radiance field-based renderer may include a scene decoder, an object decoder, a synthesizer, a voxel renderer, and a neural renderer. The object-aware scene encoder, scene decoder, object decoder, and neural renderer may include parameters that need to be trained and updated.

[0242] A training image set for training the neural network can be obtained. The training image set includes a plurality of training images without labeled data. For any training image, fourth feature information of n objects in the training image can be extracted by an object-aware scene encoder. The n objects in the training image may include: the scene of the training image, and objects outside the scene in the training image (e.g., objects, people, animals). The fourth feature information may include second appearance feature information and second change feature information, wherein the second appearance feature information may include second comprehensive appearance feature information and second shape feature information.

[0243] For the scene among the n objects, the fourth feature information of the scene can be processed by the scene decoder to obtain the third voxel density and fifth feature information of the neural radiation field corresponding to the scene. For any object other than the scene among the n objects, the fourth feature information of the object can be processed by the object decoder to obtain the third voxel density and fifth feature information of the neural radiation field corresponding to the object. The synthesizer can determine the fourth voxel density and sixth feature information of the neural radiation field corresponding to the training image based on the third voxel density and fifth feature information of the neural radiation field corresponding to each object in the training image. The voxel renderer can generate a second feature map corresponding to the training image based on the fourth voxel density and sixth feature information. The neural renderer can upsample the second feature map to obtain a second reconstructed image corresponding to the training image.

[0244] According to Formula 11 above, a loss function corresponding to the neural network can be constructed to train the neural network.

[0245] After the neural network training is completed, the two-dimensional image to be processed I o1 Input the neural network and extract the image to be processed I through the object-aware scene encoder o1 The N objects may include the first feature information of the image to be processed I o1 The scene and the image to be processed I o1 Objects outside the scene (such as objects, people, animals). Among them, the first feature information of the i-th object in the N objects can be For the scene in the N objects, the scene decoder can be used to decode the first feature information of the scene and the viewing direction information of the camera. and reference point position information Processing is performed to obtain the second voxel density and third feature information of the neural radiation field corresponding to the scene. For any object other than the scene in the N objects, the object decoder can be used to decode the first feature information of the object and the camera's viewing direction information. and reference point position information The synthesizer can use Equation 2 and Equation 3 to determine the first voxel density σ and the second characteristic information of the neural radiation field corresponding to the image to be processed according to the second voxel density and the third characteristic information of the neural radiation field corresponding to each object in the image to be processed. For the first voxel density σ and the second feature information The voxel renderer can use equations 4 to 6 to render the image to be processed I o1 The corresponding first feature map f outThe neural renderer can process the first feature map f out Perform upsampling to obtain the image to be processed I o1 The corresponding first reconstructed image I r1 . Among them, the first reconstructed image is editable.

[0246] Figure 4a A schematic diagram showing an image to be processed in the image reconstruction method provided by an embodiment of the present disclosure. Figures 4b to 4d A schematic diagram showing a reconstructed image of a new perspective obtained by reconstructing an image to be processed in the image reconstruction method provided by an embodiment of the present disclosure.

[0247] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0248] In addition, the present disclosure also provides an image reconstruction device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any image reconstruction method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method section and will not be repeated here.

[0249] Figure 5 FIG. 1 is a block diagram of an image reconstruction device provided by an embodiment of the present disclosure. Figure 5 As shown, the image reconstruction device includes:

[0250] A first acquisition module 51 is used to acquire an image to be processed;

[0251] An extraction module 52 is configured to extract first feature information of an object in the image to be processed;

[0252] A first determining module 53 is configured to determine a first voxel density and second characteristic information of a neural radiation field corresponding to the image to be processed based on the first characteristic information;

[0253] The generating module 54 is configured to generate a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information.

[0254] In a possible implementation, the first determining module 53 is configured to:

[0255] Get the camera's viewing direction information;

[0256] Obtaining location information of the reference point;

[0257] The first voxel density and second feature information of the neural radiation field corresponding to the image to be processed are determined according to the first feature information, the viewing direction information and the position information of the reference point.

[0258] In a possible implementation, the first determining module 53 is configured to:

[0259] Get the camera's viewing direction;

[0260] Position encoding is performed based on the viewing direction to obtain viewing direction information of the camera.

[0261] In a possible implementation, the first determining module 53 is configured to:

[0262] Obtain the three-dimensional coordinates of the reference point;

[0263] Position encoding is performed based on the three-dimensional coordinates of the reference point to obtain position information of the reference point.

[0264] In a possible implementation, the first determining module 53 is configured to:

[0265] determining a second voxel density and third feature information of a neural radiation field corresponding to the object based on the first feature information, the viewing direction information, and the position information of the reference point;

[0266] The first voxel density and the second characteristic information of the neural radiation field corresponding to the image to be processed are determined according to the second voxel density and the third characteristic information.

[0267] In a possible implementation, the generating module 54 is configured to:

[0268] generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information;

[0269] The first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed.

[0270] In a possible implementation, the generating module 54 is configured to:

[0271] determining, according to the first voxel density, a transparency value of a three-dimensional position in the three-dimensional space corresponding to the image to be processed;

[0272] determining a transmittance of the three-dimensional position according to the transparency value;

[0273] A first feature map corresponding to the image to be processed is generated according to the transparency value, the transmittance and the second feature information.

[0274] In a possible implementation manner, the first feature information includes first appearance feature information and first transformation feature information.

[0275] In a possible implementation manner, the first appearance feature information includes first comprehensive appearance feature information and first shape feature information.

[0276] In a possible implementation, the apparatus further includes:

[0277] The editing module is configured to edit any object in the first reconstructed image in response to an editing request for the object, and obtain an edited first reconstructed image.

[0278] In one possible implementation,

[0279] The extraction module 52 is used to: extract first feature information of an object in the image to be processed through a pre-trained neural network;

[0280] The first determining module 53 is configured to determine, through the neural network and based on the first characteristic information, the first voxel density and the second characteristic information of the neural radiation field corresponding to the image to be processed;

[0281] The generating module 54 is configured to generate a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information through the neural network.

[0282] In a possible implementation, the apparatus further includes:

[0283] A second acquisition module is used to acquire a training image;

[0284] a processing module, configured to process the training image through the neural network to obtain a second reconstructed image corresponding to the training image;

[0285] a second determining module, configured to determine a value of a loss function of the neural network according to the second reconstructed image;

[0286] A training module is used to train the neural network according to the value of the loss function.

[0287] In a possible implementation, the loss function includes a first loss function;

[0288] The second determining module is used for:

[0289] The value of the first loss function is determined according to difference information between the second reconstructed image and the training image.

[0290] In one possible implementation, the loss function includes a second loss function;

[0291] The second determining module is used for:

[0292] Determine the location information of the mask area;

[0293] generating a composite image according to the second reconstructed image, the training image, and the position information of the mask area;

[0294] A value of the second loss function is determined based on the synthesized image.

[0295] In a possible implementation, the second determining module is configured to:

[0296] Determining a first image to be synthesized based on the second reconstructed image and the position information of the mask area, wherein a size of the first image to be synthesized is the same as a size of the second reconstructed image, and in the first image to be synthesized, pixel values ​​of pixels outside the mask area are the same as those of the second reconstructed image, and pixel values ​​of pixels within the mask area are empty;

[0297] Determining a second image to be synthesized based on the position information of the training image and the mask area, wherein a size of the second image to be synthesized is the same as a size of the training image, and in the second image to be synthesized, pixel values ​​of pixels within the mask area are the same as those of the training image, and pixel values ​​of pixels outside the mask area are blank;

[0298] The synthesized image is generated according to the first image to be synthesized and the second image to be synthesized.

[0299] In a possible implementation, the second loss function includes a generation loss function and an adversarial loss function;

[0300] The second determining module is used for:

[0301] Determining a value of a generation loss function based on the synthesized image;

[0302] The value of the adversarial loss function is determined according to the difference information between the synthetic image and the training image.

[0303] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0304] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the above method. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0305] The embodiment of the present disclosure further provides a computer program, comprising a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the above method.

[0306] An embodiment of the present disclosure further provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0307] An embodiment of the present disclosure also provides an electronic device, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0308] The electronic device may be provided as a terminal, a server, or other forms of devices.

[0309] Figure 6 FIG. 1 is a block diagram of an electronic device 1900 provided by an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server. Figure 6 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0310] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (Mac OSX TM ), a multi-user, multi-process computer operating system (Unix TM), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.

[0311] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0312] The present disclosure relates to the field of augmented reality. By acquiring image information of a target object in a real-world environment, the relevant features, states, and attributes of the target object are detected or identified using various vision-related algorithms, thereby achieving an AR effect that combines virtual and real life and matches the specific application. For example, the target object may be a face, limbs, gestures, movements, etc. related to the human body, or an identifier or marker related to an object, or a sandbox, display area, or display items related to a venue or location. Vision-related algorithms may involve visual positioning, SLAM, 3D reconstruction, image registration, background segmentation, key point extraction and tracking of objects, and object pose or depth detection. Specific applications can involve not only interactive scenarios such as guided tours, navigation, explanations, reconstruction, and virtual effect overlay displays related to real scenes or objects, but also special effects processing related to people, such as makeup beautification, body beautification, special effects display, and virtual model display. Detection or identification of the relevant features, states, and attributes of the target object can be achieved using a convolutional neural network. The above-mentioned convolutional neural network is a network model obtained by model training based on a deep learning framework.

[0313] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0314] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0315] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0316] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0317] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0318] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0319] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0320] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0321] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0322] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0323] If the technical solutions of the embodiments of the present disclosure involve personal information, the products applying the technical solutions of the embodiments of the present disclosure have clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solutions of the embodiments of the present disclosure involve sensitive personal information, the products applying the technical solutions of the embodiments of the present disclosure have obtained the individual's separate consent before processing the sensitive personal information, and at the same time meet the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information. The personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0324] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An image reconstruction method, characterized in that: include: Get training images; Processing the training image through a neural network to obtain a second reconstructed image corresponding to the training image; Determining a value of a loss function of the neural network based on the second reconstructed image; wherein the loss function includes a second loss function; determining the value of the loss function of the neural network based on the second reconstructed image includes: determining position information of a mask area; generating a composite image based on the second reconstructed image, the training image, and the position information of the mask area; and determining the value of the second loss function based on the composite image; Training the neural network according to the value of the loss function; Get a single image to be processed; Extracting, by means of the pre-trained neural network, first feature information corresponding one-to-one to N objects in the image to be processed, wherein N is an integer greater than or equal to 1, and the N objects include scenes in the image to be processed and objects outside the scenes in the image to be processed, wherein the objects outside the scenes include at least one type of object, person, or animal; Obtain the camera's viewing direction information and the reference point's position information; Determining, by the neural network, second voxel densities and third feature information of neural radiation fields corresponding one-to-one to the N objects based on the first feature information, the viewing direction information, and the position information of the reference point; Determining, by the neural network, the first voxel density and the second feature information of the neural radiation field corresponding to the image to be processed based on the second voxel density and the third feature information of the neural radiation fields corresponding one to one to the N objects; A first reconstructed image corresponding to the image to be processed is generated by the neural network according to the first voxel density and the second feature information.

2. The method according to claim 1, characterized in that The obtaining of the camera's viewing angle direction information includes: Get the camera's viewing direction; Position encoding is performed based on the viewing direction to obtain viewing direction information of the camera.

3. The method according to claim 1, characterized in that The obtaining of the position information of the reference point includes: Obtain the three-dimensional coordinates of the reference point; Position encoding is performed based on the three-dimensional coordinates of the reference point to obtain position information of the reference point.

4. The method according to any one of claims 1 to 3, characterized in that Generating a first reconstructed image corresponding to the image to be processed according to the first voxel density and the second feature information includes: generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information; The first feature map is upsampled to obtain a first reconstructed image corresponding to the image to be processed.

5. The method according to claim 4, characterized in that Generating a first feature map corresponding to the image to be processed according to the first voxel density and the second feature information includes: determining, according to the first voxel density, a transparency value of a three-dimensional position in the three-dimensional space corresponding to the image to be processed; determining a transmittance of the three-dimensional position according to the transparency value; A first feature map corresponding to the image to be processed is generated according to the transparency value, the transmittance and the second feature information.

6. The method according to any one of claims 1 to 3, characterized in that The first feature information includes first appearance feature information and first transformation feature information.

7. The method according to claim 6, characterized in that The first appearance feature information includes first comprehensive appearance feature information and first shape feature information.

8. The method according to any one of claims 1 to 3, characterized in that After generating the first reconstructed image corresponding to the image to be processed, the method further includes: In response to an editing request for any object in the first reconstructed image, the object is edited to obtain an edited first reconstructed image.

9. The method according to any one of claims 1 to 3, characterized in that The loss function includes a first loss function; Determining a value of a loss function of the neural network according to the second reconstructed image includes: The value of the first loss function is determined according to difference information between the second reconstructed image and the training image.

10. The method according to any one of claims 1 to 3, characterized in that Generating a synthetic image according to the second reconstructed image, the training image, and the position information of the mask area includes: Determining a first image to be synthesized based on the second reconstructed image and the position information of the mask area, wherein a size of the first image to be synthesized is the same as a size of the second reconstructed image, and in the first image to be synthesized, pixel values ​​of pixels outside the mask area are the same as those of the second reconstructed image, and pixel values ​​of pixels within the mask area are empty; Determining a second image to be synthesized based on the position information of the training image and the mask area, wherein a size of the second image to be synthesized is the same as a size of the training image, and in the second image to be synthesized, pixel values ​​of pixels within the mask area are the same as those of the training image, and pixel values ​​of pixels outside the mask area are blank; The synthesized image is generated according to the first image to be synthesized and the second image to be synthesized.

11. The method according to any one of claims 1 to 3, characterized in that The second loss function includes a generation loss function and an adversarial loss function; Determining a value of the second loss function according to the synthesized image includes: Determining a value of a generation loss function based on the synthesized image; The value of the adversarial loss function is determined according to the difference information between the synthetic image and the training image.

12. An image reconstruction device, characterized in that: include: A second acquisition module is used to acquire a training image; a processing module, configured to process the training image through a neural network to obtain a second reconstructed image corresponding to the training image; a second determination module, configured to determine a value of a loss function of the neural network based on the second reconstructed image; wherein the loss function includes a second loss function; and determining the value of the loss function of the neural network based on the second reconstructed image includes: determining position information of a mask region; generating a composite image based on the second reconstructed image, the training image, and the position information of the mask region; and determining the value of the second loss function based on the composite image. A training module, configured to train the neural network according to a value of the loss function; A first acquisition module is used to acquire a single image to be processed; an extraction module, configured to extract, using the pre-trained neural network, first feature information corresponding one-to-one to N objects in the image to be processed, where N is an integer greater than or equal to 1, and the N objects include scenes in the image to be processed and objects outside the scenes in the image to be processed, where the objects outside the scenes include at least one type of object, person, or animal; A first determination module is configured to obtain viewing direction information of a camera and position information of a reference point; determine, through the neural network, a second voxel density and a third feature information of a neural radiation field corresponding one-to-one to the N objects based on the first feature information, the viewing direction information, and the position information of the reference point; and determine, through the neural network, a first voxel density and a second feature information of a neural radiation field corresponding to the image to be processed based on the second voxel density and the third feature information of the neural radiation field corresponding one-to-one to the N objects; A generation module is used to generate a first reconstructed image corresponding to the image to be processed based on the first voxel density and the second feature information through the neural network.

13. An electronic device, characterized in that: include: one or more processors; a memory for storing executable instructions; The one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

15. A computer program product, characterized in that The invention comprises a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is executed in an electronic device, a processor in the electronic device executes the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Image rendering method and device based on neural radiation field, and electronic equipment

    CN113592991A

  • Method and equipment for determining viewpoint path in three-dimensional scene

    CN113628348A