Three-dimensional scene reconstruction method and device, equipment and storage medium

By constructing the neural radiation field of a three-dimensional scene and performing multi-layer rendering, multi-layer spherical shell images are generated, which solves the problem of high NeRF rendering calculations, real-time rendering and wide applicability of three-dimensional scenes on mid- and low-end XR devices.

CN120070719APending Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311605735.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using NeRF to directly render three-dimensional scenes, the calculation amount is too large, making it difficult to support real-time rendering on mid- and low-end XR devices, limiting the applicability of three-dimensional scene reconstruction.

Method used

By constructing the neural radiation field of the three-dimensional scene and determining the light sampling points and color information for each point center, multi-layer rendering technology is used to generate multi-layer spherical shell images to achieve accurate reconstruction of the three-dimensional scene.

Benefits of technology

It reduces the computational overhead of three-dimensional scene reconstruction, ensures real-time rendering of three-dimensional scenes on mid- and low-end XR devices, and improves the comprehensive applicability of three-dimensional scene reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070719A_ABST
    Figure CN120070719A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional scene reconstruction method and device, equipment and a storage medium. The method comprises the following steps: constructing a neural radiation field of a three-dimensional scene according to a multi-view image sequence in the three-dimensional scene; for each given point center, determining a light sampling point under the point center and color information of the light sampling point according to the nerve radiation field; and carrying out multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image reconstructed by the three-dimensional scene under the point center. According to the embodiment of the invention, the method can achieve the accurate reconstruction of the three-dimensional scene in a multi-point center, achieves the local approximation of a virtual scene after the reconstruction of the three-dimensional scene through the multi-layer spherical shell image in the multi-point center, guarantees the real stereoscopic impression of the reconstructed three-dimensional scene, reduces the calculation cost of the reconstruction of the three-dimensional scene, and improves the reconstruction efficiency of the three-dimensional scene. Accurate reconstruction of the three-dimensional scene can also be realized on middle-end and low-end XR equipment, and the comprehensive applicability of reconstruction of the three-dimensional scene is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a three-dimensional scene reconstruction method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of Extended Reality (XR) technology, three-dimensional scene reconstruction technology can provide users with more and more virtual interaction scenarios to enhance the immersive interaction experience of users in three-dimensional scenes.

[0003] Generally, as a method of representing a three-dimensional scene as a radiance field approximated by a neural network, Neural Radiance Field (NeRF) has more advantages in terms of the rendering realism and restoration degree of three-dimensional scenes compared to traditional Multi-View Stereo (MVS) reconstruction methods.

[0004] However, when directly rendering a three-dimensional scene using NeRF, there will be a large amount of rendering calculation, which requires high device processing performance and is difficult to support real-time rendering of three-dimensional scenes on mid- to low-end XR devices, and there are certain limitations for three-dimensional scene reconstruction. Summary of the Invention

[0005] Embodiments of the present application provide a three-dimensional scene reconstruction method, apparatus, device, and storage medium, which can accurately reconstruct a three-dimensional scene under multiple point centers, locally approximate the virtual scene after three-dimensional scene reconstruction through multi-layer spherical shell images under multiple point centers, reduce the computational overhead of three-dimensional scene reconstruction, and ensure the comprehensive applicability of three-dimensional scene reconstruction.

[0006] In a first aspect, embodiments of the present application provide a three-dimensional scene reconstruction method, which includes:

[0007] Construct a neural radiance field of the three-dimensional scene according to a multi-view image sequence within the three-dimensional scene;

[0008] For each given point center, determine the light sampling points and the color information of the light sampling points under the point center according to the neural radiance field;

[0009] Perform multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image of the three-dimensional scene reconstructed under the point center.

[0010] In a second aspect, embodiments of the present application provide a three-dimensional scene reconstruction apparatus, which includes:

[0011] A radiation field construction module for constructing a neural radiation field of the three-dimensional scene according to a multi-view image sequence within the three-dimensional scene;

[0012] A point color determination module for determining, for each given point center, a ray sampling point under the point center and color information of the ray sampling point according to the neural radiation field;

[0013] A three-dimensional scene reconstruction module for performing multi-layer rendering on the color information of the ray sampling points to obtain a multi-layer spherical shell image reconstructed by the three-dimensional scene under the point center.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes:

[0015] A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, and the computer program enables a computer to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions, and the computer program / instructions enable a computer to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0018] Through the technical solution of the present application, first, a multi-view image sequence within a three-dimensional scene is processed to construct a neural radiation field of the three-dimensional scene. Then, for each given point center, the respective ray sampling points and the color information of each ray sampling point under the point center can be determined according to the neural radiation field, and multi-layer rendering is performed on the color information of the respective ray sampling points, so that a multi-layer spherical shell image reconstructed by the three-dimensional scene under the point center can be obtained, thereby realizing accurate reconstruction of the three-dimensional scene under multiple point centers. The virtual scene after the three-dimensional scene reconstruction is locally approximated by the multi-layer spherical shell images under multiple point centers, ensuring the true three-dimensional sense after the three-dimensional scene reconstruction, and reducing the computational overhead of the three-dimensional scene reconstruction. Accurate reconstruction of the three-dimensional scene can also be achieved on mid-range and low-end XR devices, ensuring the comprehensive applicability of the three-dimensional scene reconstruction. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of a three-dimensional scene reconstruction method provided by an embodiment of the present application;

[0021] Figure 2 It is an exemplary schematic diagram of the three-dimensional scene reconstruction process provided by an embodiment of the present application;

[0022] Figure 3 It is a flowchart of a method for rendering multi-layer spherical shell images under each point center provided by an embodiment of the present application;

[0023] Figure 4 It is an exemplary schematic diagram of the principle for setting the depth relationship between adjacent layers provided by an embodiment of the present application;

[0024] Figure 5 It is a flowchart of a method for optimizing the multi-layer spherical shell images reconstructed under any point center of the three-dimensional scene provided by an embodiment of the present application;

[0025] Figure 6 It is an exemplary schematic diagram of the optimization process of the multi-layer spherical shell images reconstructed under any point center of the three-dimensional scene provided by an embodiment of the present application;

[0026] Figure 7 It is a principle block diagram of a three-dimensional scene reconstruction device provided by an embodiment of the present application;

[0027] Figure 8 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] In the embodiments of this application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific way.

[0031] To enhance the diverse interactions of users in a three-dimensional scene, a virtual scene that is exactly the same as the three-dimensional scene can usually be reconstructed to support users to perform various interaction operations within the virtual scene, so as to improve the immersive interaction experience of users in the three-dimensional scene.

[0032] Before introducing the specific technical solutions of this application, the specific application scenarios of the virtual scene reconstructed from the three-dimensional scene of this application can be described accordingly first:

[0033] The virtual scene reconstructed from any three-dimensional scene of this application can be supported to be presented on mobile phones, tablet computers, personal computers, servers, and smart wearable devices, so that users can view the virtual scene reconstructed from the three-dimensional scene. When the virtual scene reconstructed from the three-dimensional scene is presented on an XR device, users can enter the virtual scene reconstructed from the three-dimensional scene by wearing the XR device, so as to perform various interaction operations to achieve diverse interactions of users in the three-dimensional scene.

[0034] XR refers to the combination of the real and the virtual through a computer to create a virtual environment for human-computer interaction. XR is also a general term for various technologies such as Virtual Reality (VR for short), Augmented Reality (AR for short), and Mixed Reality (MR for short). By integrating the visual interaction technologies of the three, it brings an "immersive feeling" of seamless conversion between the virtual world and the real world to the experiencer. XR devices are usually worn on the user's head, so XR devices are also called head-mounted devices.

[0035] VR: A technology for creating and experiencing virtual worlds, which computationally generates a virtual environment. It is a multi-source information (the virtual reality mentioned in this article includes at least visual perception, and may also include auditory perception, tactile perception, motion perception, and even taste perception, olfactory perception, etc.). It realizes the fusion of the virtual environment, an interactive three-dimensional dynamic visual scene and the simulation of entity behaviors, enabling users to immerse themselves in the simulated virtual reality environment and realizing applications in various virtual environments such as maps, games, videos, education, medical treatment, simulation, collaborative training, sales, assisting manufacturing, maintenance, and repair.

[0036] A VR device refers to the terminal that realizes the virtual reality effect, usually in the form of glasses, a Head-Mounted Display (HMD), or contact lenses, for realizing visual perception and other forms of perception. Of course, the form of the virtual reality device is not limited to this and can be further miniaturized or enlarged according to needs.

[0037] AR: An AR scene refers to a simulated scene in which at least one virtual object is superimposed on a physical scene or its representation. For example, an electronic system may have an opaque display and at least one imaging sensor for capturing images or videos of the physical scene, which are representations of the physical scene. The system combines the images or videos with virtual objects and displays the combination on the opaque display. An individual uses the system to indirectly view the physical scene via the images or videos of the physical scene and observes the virtual objects superimposed on the physical scene. When the system uses one or more image sensors to capture images of the physical scene and uses those images to present an AR scene on the opaque display, the displayed images are called video pass-through. Alternatively, an electronic system for displaying an AR scene may have a transparent or translucent display through which an individual can directly view the physical scene. The system can display virtual objects on the transparent or translucent display such that an individual uses the system to observe the virtual objects superimposed on the physical scene. For another example, the system may include a projection system that projects virtual objects into the physical scene. The virtual objects may be projected, for example, onto a physical surface or as a hologram such that an individual uses the system to observe the virtual objects superimposed on the physical scene. Specifically, it is a technology that calculates in real time the camera pose information parameters of a camera in the real world (also known as the three-dimensional world or the real world) during the process of the camera capturing images, and adds virtual elements to the images captured by the camera according to the camera pose information parameters. The virtual elements include, but are not limited to: images, videos, and three-dimensional models. The goal of AR technology is to interact by superimposing the virtual world on the real world on the screen.

[0038] MR: By presenting virtual scene information in a real-world scenario, an information loop for interactive feedback is established among the real world, the virtual world, and the user to enhance the realism of the user experience. For example, integrating sensory inputs created by a computer (such as virtual objects) with sensory inputs from a physical scene or its representation in a simulated scene. In some MR scenes, the sensory inputs created by the computer can adapt to changes in the sensory inputs from the physical scene. Additionally, some electronic systems for presenting an MR scene can monitor orientation and / or position information relative to the physical scene so that virtual objects can interact with real objects (i.e., physical elements from the physical scene or their representations). For example, the system can monitor movement such that a virtual plant appears stationary relative to a physical building.

[0039] Optionally, the XR device described in the embodiments of the present application, also known as a virtual reality device, may include, but is not limited to, the following types:

[0040] 1) Mobile virtual reality device, which supports setting a mobile terminal (such as a smartphone) in various ways (such as a head-mounted display with a dedicated card slot). Through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations for virtual reality functions and outputs data to the mobile virtual reality device. For example, watch virtual reality videos through the APP of the mobile terminal.

[0041] 2) All-in-one virtual reality device, which has a processor for performing relevant calculations for virtual functions, and thus has independent virtual reality input and output functions. It does not need to be connected to a PC or a mobile terminal, and has a high degree of freedom of use.

[0042] 3) Computer-side virtual reality (PCVR) device, which uses the PC side to perform relevant calculations for virtual reality functions and data output. The externally connected computer-side virtual reality device uses the data output by the PC side to achieve the virtual reality effect.

[0043] After introducing the specific application scenarios that the virtual scene reconstructed in the three-dimensional scene reconstruction in this application can support, the following specifically describes a three-dimensional scene reconstruction method provided by an embodiment of this application with reference to the accompanying drawings.

[0044] Currently, when using NeRF to directly render a three-dimensional scene, there will be an excessive amount of rendering calculations, which requires a high processing performance of the device. It is difficult to support the real-time rendering of the three-dimensional scene on mid-range and low-end XR devices, and there are certain limitations for three-dimensional scene reconstruction.

[0045] To solve the above problems, the inventive concept of this application is: First, process the multi-view image sequence in the three-dimensional scene to construct the neural radiance field of the three-dimensional scene. Then, for each given point center, according to the neural radiance field, determine the respective light sampling points and the color information of each light sampling point under this point center, and perform multi-layer rendering on the color information of each light sampling point, then the multi-layer spherical shell image reconstructed under this point center of the three-dimensional scene can be obtained, so as to achieve the accurate reconstruction of the three-dimensional scene under multiple point centers. The virtual scene reconstructed in the three-dimensional scene is locally approximated by the multi-layer spherical shell images under multiple point centers, ensuring the true three-dimensional sense after the three-dimensional scene reconstruction, and reducing the computational cost of the three-dimensional scene reconstruction. The accurate reconstruction of the three-dimensional scene can also be achieved on mid-range and low-end XR devices, ensuring the comprehensive applicability of the three-dimensional scene reconstruction.

[0046] Figure 1The flowchart of a 3D scene reconstruction method provided by an embodiment of this application. This method can be executed by the 3D scene reconstruction device provided by this application, where the 3D scene reconstruction device can be implemented in any software and / or hardware manner. Exemplarily, the 3D scene reconstruction device can be configured in any electronic device such as an XR device, a server, a mobile phone, a tablet computer, a personal computer, a smart wearable device, etc., and this application does not impose any restrictions on the specific type of the electronic device.

[0047] Specifically, as Figure 1 shown, this method may include the following steps:

[0048] S110, construct a neural radiance field of the 3D scene according to the multi-view image sequence within the 3D scene.

[0049] Among them, the 3D scene can be any real environment where the user is located.

[0050] To achieve accurate reconstruction of the 3D scene, first, it is necessary to obtain the 3D shapes, texture information, etc. of various real objects within the 3D scene in all directions. The real objects can be actual items, walls, floors, etc. within the 3D scene. For this purpose, this application can set a camera at any position point within the 3D scene, and through this camera, capture multiple real environment images of the 3D scene from multiple perspectives, thereby forming the multi-view image sequence in this application.

[0051] It can be understood that in order to achieve comprehensive reconstruction of the 3D scene, after the multi-view image sequence in this application is arranged accordingly according to the corresponding shooting perspectives, a panoramic image of the 3D scene can be formed, so as to be able to comprehensively understand the geometric structure and appearance information of the 3D scene subsequently.

[0052] Since there may be weak texture regions in each image within the multi-view image sequence, it is difficult for the traditional MVS reconstruction method to represent the fine and complex geometric structures in the 3D scene, and it is also difficult for the textures of the images in the multi-view image sequence to reproduce the details such as lighting and materials in the 3D scene with high fidelity. As a method that represents the 3D scene as a radiance field approximated by a neural network, NeRF can accurately describe the color information and volume density of each spatial point in the 3D scene in each viewing direction, making it more advantageous than the traditional MVS reconstruction method in terms of the reduction degree of geometric structure and rendering realism.

[0053] Therefore, in this application, when reconstructing any 3D scene, multiple cameras can be set at different positions within the 3D scene or the position of a single camera can be changed, and the orientation of the camera at each position can be set, so as to capture real scene images within the 3D scene from multiple different positions and perspectives, thereby forming the multi-view image sequence in this application.

[0054] Then, the network structure of NeRF in this application is a simple fully-connected network.

[0055] As Figure 2 shown, for a multi-view image sequence, the neural radiance field of the three-dimensional scene can be trained accordingly by processing the geometric structure information and texture features in the real-scene images collected at each view in the multi-view image sequence, so as to construct the neural radiance field of the three-dimensional scene.

[0056] Specifically, for each view image in the multi-view image sequence, by analyzing the camera position and camera orientation corresponding to the view image, corresponding radiation rays can be formed from the camera position towards each pixel point in the imaging plane of the camera, and thus multiple radiation rays under the multi-view image sequence can be obtained. For each radiation ray, a corresponding ray sampling strategy can be adopted to continuously sample multiple ray sampling points on the radiation ray. Then, by processing the texture features in each real-scene image in the multi-view image sequence, the color information and volume density of each ray sampling point on each radiation ray can be predicted, thereby constructing the neural radiance field of the three-dimensional scene.

[0057] Among them, the neural radiance field of the three-dimensional scene can include particle information such as the color, volume density, light intensity, and material of each ray sampling point on each radiation ray in the three-dimensional scene, so as to represent the actual texture of each real object in the three-dimensional scene.

[0058] In some implementable ways, in order to ensure the high-fidelity appearance information of the three-dimensional scene, this application can represent the neural radiance field based on the ray density, and analyze the particle information such as the color, volume density, light intensity, and material of each ray sampling point in the three-dimensional scene, so as to generate the neural radiance field of the three-dimensional scene to represent the high-fidelity appearance information of the three-dimensional scene.

[0059] S120. For each given point center, according to the neural radiance field, determine the ray sampling points and the color information of the ray sampling points under the point center.

[0060] To ensure convenient interaction of the user in the three-dimensional scene, this application can preset multiple default position points in the three-dimensional scene, or support the user to select multiple custom position points in the presented virtual scene after presenting the virtual scene reconstructed from the three-dimensional scene. Then, whether it is a default position point or a custom position point, the position point can be used as the center point, and corresponding multiple user-roaming areas can be delimited respectively according to the preset range size, and then each user-roaming area can be used as multiple points in the three-dimensional scene in this application. Then, the center points of each point are the multiple point centers given in the three-dimensional scene.

[0061] Then, after constructing the neural radiance field of the three-dimensional scene, for each given point center, the present application can perform corresponding processing on the specific position of the point center and each supported camera pose information, so as to determine, according to the neural radiance field of the three-dimensional scene, multiple radiance rays formed by each pixel point in the imaging plane starting from the camera position facing the camera at the point center. For each radiance ray, a corresponding ray sampling strategy can be adopted to continuously sample multiple ray sampling points on the radiance ray, thereby obtaining each ray sampling point under the point center. Moreover, according to the high-fidelity appearance information represented by the neural radiance field of the three-dimensional scene, the color information of each ray sampling point can be determined, and the color information can include the color information (i.e., RGB value) and volume density information of the ray sampling point.

[0062] In some implementable ways, for the accuracy of the color information of the ray sampling points under each point center, the present application can determine the ray sampling points and the color information of the ray sampling points under each point center through the following steps:

[0063] Step 1, for each given point center, determine the radiance rays and the ray sampling points on the radiance rays under the point center according to the neural radiance field.

[0064] For each point center, according to each camera view angle supported by the point center in the neural radiance field, multiple radiance rays can be determined to extend outward from the point center according to the neural radiance field. For each radiance ray, a corresponding ray sampling strategy can be adopted to continuously sample the radiance ray. For example, a ray point is sampled every preset length on the radiance ray, so as to obtain each ray sampling point on each radiance ray as the ray sampling points under the point center.

[0065] Step 2, determine the color information of the ray sampling points according to the neural radiance field.

[0066] Since the neural radiance field can include particle information such as the color, volume density, illumination intensity, and material of each ray sampling point on each radiance ray in the three-dimensional scene, so as to represent the actual texture of each real object in the three-dimensional scene. Therefore, after determining each ray sampling point under each point center, the present application can determine the color information of each ray sampling point under the point center according to the high-fidelity appearance information represented by the neural radiance field of the three-dimensional scene, including the RGB color value and volume density information of the ray sampling point.

[0067] S130, perform multi-layer rendering on the color information of the ray sampling points to obtain a multi-layer spherical shell image reconstructed by the three-dimensional scene under the point center.

[0068] Since in the neural radiance field, through the volume rendering method, the color intensity projected by each ray sampling point on multiple radiance rays onto the imaging plane can be analyzed to render the projection image on the imaging plane. Then, in order to ensure the true three-dimensional sense of the 3D scene reconstruction, the present application can set a corresponding layer at multiple different depths for each point center, and thus multiple layers can be set around each point center.

[0069] Then, for each point center, by comparing the depth of each layer under the point center with the depth of each ray sampling point, the volume rendering method designed in the neural radiance field can be used to determine the color intensity projected by the color information of the partial ray sampling points that can affect each layer among each ray sampling point on the layer, and obtain the rendered image on the layer. In the same way as above, the color information of each ray sampling point under the point center can be projected onto each layer respectively to render each layer under the point center, so as to obtain the rendered image on each layer under the point center, and thus form a multi-sphere image (abbreviated as MSI) reconstructed from the 3D scene under the point center.

[0070] Among them, MSI can be a data format that extends the multi-plane image (abbreviated as MPI) to a 360-degree sphere. At each point center of the 3D scene, by presenting a multi-sphere image in a surrounding manner, the true three-dimensional sense after the 3D scene reconstruction is enhanced.

[0071] It can be understood that the volume rendering formula designed in the neural radiance field in the present application can be:

[0072]

[0073] Among them, n is the serial number of each ray sampling point belonging to the same radiance ray under each point in the order from near to far according to the distance. t n represents the distance between the nth ray sampling point and the point. σ n can be the volume density of the nth ray sampling point, and c n can be the RGB color value of the nth ray sampling point.

[0074] Then, (1 - exp(-σ n (t n+1 - t n ))) can represent the opacity of the nth ray sampling point.

[0075] T(t 1 →t n)It can be the transmittance from the 1st light sampling point to the nth light sampling point when arranging the light sampling points on the same radiation ray under the center of this point position in the order of increasing distance from near to far.

[0076] Among them, δ k represents the distance between adjacent light sampling points, that is, t k+1 -t k .

[0077] Thus, according to the above formula, for each point position center, the color information of some light sampling points among the light sampling points under this point position center that can affect each layer can be projected onto this layer, and the color intensity after projection on this layer can be calculated to obtain the rendered image on each layer under this point position center, thereby composing the multi-layer spherical shell image reconstructed under this point position center of the three-dimensional scene.

[0078] In addition, for any three-dimensional scene, the multi-layer spherical shell images reconstructed under each point position center can support the panoramic roaming function of the user within this three-dimensional scene to display the corresponding roaming images in real time for the user. Therefore, after obtaining the multi-layer spherical shell images reconstructed under each point position center of the three-dimensional scene, this application will also obtain the roaming pose information of the user, and this roaming pose information includes the roaming position and the roaming posture; if the roaming position is within this point position, the corresponding roaming image will be displayed according to the roaming pose information and the multi-layer spherical shell image reconstructed under this point position center.

[0079] That is to say, for the roaming of the user within the three-dimensional scene, this application can obtain the roaming pose information of the user in real time, so as to determine the roaming position and roaming posture of the user within the three-dimensional scene in real time.

[0080] Each point position can be composed of a point position center and a user-roamable area. By judging the roaming position of the user and the boundaries of the roamable areas of each point position, it can be determined that the roaming position is within a certain point position.

[0081] In the case where the roaming position of the user is within a certain point position, this application can first obtain the multi-layer spherical shell image reconstructed under this point position center. Then, by processing the roaming pose information of the user, the roaming viewing angle range of the user within the multi-layer spherical shell image reconstructed under this point position center can be determined, and based on this, the multi-layer local spherical shell images within this roaming viewing angle range of the multi-layer spherical shell image under this point position center are rendered, thereby displaying the corresponding roaming image.

[0082] The technical solution provided by the embodiments of this application first processes the multi-view image sequence in the three-dimensional scene to construct the neural radiance field of the three-dimensional scene. Then, for each given point center, the neural radiance field can be used to determine the light sampling points under this point center and the color information of each light sampling point, and multi-layer rendering is performed on the color information of each light sampling point, so as to obtain the multi-layer spherical shell image reconstructed under this point center of the three-dimensional scene, thereby realizing the accurate reconstruction of the three-dimensional scene under multiple point centers. The virtual scene after the reconstruction of the three-dimensional scene is locally approximated by the multi-layer spherical shell images under multiple point centers, ensuring the true three-dimensional sense after the reconstruction of the three-dimensional scene and reducing the computational overhead of the three-dimensional scene reconstruction. The accurate reconstruction of the three-dimensional scene can also be realized on mid-range and low-end XR devices, ensuring the comprehensive applicability of the three-dimensional scene reconstruction.

[0083] As an alternative implementation solution in this application, since the true object depths are different when viewing the three-dimensional scene from different camera perspectives under each point center of the three-dimensional scene, there are also certain differences in the depths of multiple layers set under each point center, thereby generating multi-layer spherical shell images at different depths. Therefore, in order to ensure the true accuracy after the reconstruction of the three-dimensional scene, this application can describe the specific rendering process of the multi-layer spherical shell images reconstructed under each point center of the three-dimensional scene.

[0084] Figure 3 It is the method flow chart of the rendering process of the multi-layer spherical shell image under each point center provided by the embodiments of this application. As Figure 3 shown, the method can specifically include the following steps:

[0085] S310, for each point center, perform depth rendering on the light sampling points of each radiation ray under this point center to obtain the minimum radiation depth and the maximum radiation depth under this point center.

[0086] For each point center, according to the respective camera perspectives under this point center, multiple radiation rays can be projected from this point center to the imaging plane of each camera, and by using the corresponding light sampling strategy, multiple light sampling points on each radiation ray can be obtained.

[0087] Since each radiation ray under each point center will lose energy due to the occlusion of various particles during the light path propagation. For example, gases will occlude some radiation rays, while opaque solids will occlude all radiation rays, so that each radiation ray under each point center will have different light lengths according to whether it encounters the occlusion of real objects, that is, each radiation ray under each point center will have different radiation depths, and the number of light sampling points on each radiation ray is also limited.

[0088] Therefore, for each point center, after determining each radiation ray under the point center, the present application can, for each radiation ray, determine the distance between each ray sampling point on the radiation ray and the point center as the depth information of each ray sampling point. Then, according to the depth information of the ray sampling points on the same radiation ray, depth rendering can be performed on the ray sampling points on the radiation ray to obtain the radiation depth of the radiation ray.

[0089] Among them, the depth rendering formula for each ray sampling point on each radiation ray can be shown as follows:

[0090]

[0091] Among them, when performing depth rendering on each ray sampling point on a certain radiation ray, discrete summation can be performed with reference to the depth information of any ray sampling point on the radiation ray to obtain the radiation depth of the radiation ray. In the present application, the midpoint of adjacent ray sampling points can also use any ray sampling point t n , and the present application does not limit this.

[0092] Thus, in the same way as above, the radiation depth of each radiation ray under the point center can be determined. Then, by comparing the magnitudes of the radiation depths of the radiation rays under the point center, the minimum radiation depth and the maximum radiation depth under the point center can be determined. Taking the point center as the center, there will be no real objects in the three-dimensional scene in the area less than the minimum radiation depth and the area greater than the maximum radiation depth.

[0093] S320. According to the minimum radiation depth, the maximum radiation depth, and the preset adjacent layer depth relationship, determine the number of layers under the point center and the depth information of each layer.

[0094] For the multi-layer spherical shell images under each point center, it can support the user to move within the user's roamable area represented by the point to view the multi-layer spherical shell images presented in a surrounding manner under the point center.

[0095] It can be understood that the radius of each point should be less than the depth information of the first layer in the multi-layer spherical shell image under the point center, so that the user can comprehensively observe the overall spatial environment of the three-dimensional scene under the point center within the point.

[0096] When the user views the presented multi-layer spherical shell images at each point position, if any adjacent layer is too close, the content of the latter layer in the adjacent layers will have a certain impact on the viewing effect of the content of the previous layer. Therefore, in order to ensure the true accuracy after the three-dimensional scene reconstruction, this application can preset an adjacent layer depth relationship, and by setting the depth interval between adjacent layers within a certain range, cancel the viewing impact of the content of the latter layer on the content of the previous layer.

[0097] As Figure 4 shown, taking the adjacent layers r l and r l+1 centered at a certain point position as an example, r l is the depth information of the previous layer in the adjacent layers, and r l+1 is the depth information of the latter layer in the adjacent layers. The dotted area centered at this point position can be the roamable area represented by this point position, and the initial value of the radius of this point position is the preset expected roam depth r i .

[0098] When observing the same point on the layer r l at different positions within this point position, the observable line of sight will pass through this point on the layer r l and fall on different points on the layer r l+1 . Then, in order to avoid the viewing impact of the content of the latter layer on the content of the previous layer, it is required that the adjacent layer depth relationship should satisfy that the maximum pixel parallax of the observable line of sight of any pixel point of the previous layer in the adjacent layers within this point position when it falls on the latter layer in the adjacent layers is less than or equal to a unit pixel.

[0099] At this time, as Figure 4 shown, for any pixel point on the previous layer r l in the adjacent layers, when observing this pixel point on the layer r l at the two boundary tangent points where the connection line with this pixel point is tangent within the roamable area represented by this point position, the pixel parallax of these two observable lines of sight when they fall on the latter layer r l+1 in the adjacent layers is the largest. Then, it is required that this maximum pixel parallax is less than or equal to a unit pixel.

[0100] From the above content, assuming that the maximum pixel parallax is equal to a unit pixel, we can obtain

[0101] where res is the texture resolution of the preset multi-layer spherical shell image, and d is the arc length corresponding to the maximum pixel parallax on the latter layer r l+1 .

[0102] Assuming then it can be inferred that

[0103] Assume that the included angle formed when the two observable sight lines corresponding to the maximum pixel parallax pass through the pixel point on the previous layer r l is Then it can be speculated that and

[0104] It can be speculated from the above formula that the depth relationship between adjacent layers can be

[0105] Then, according to the formula of the depth relationship between adjacent layers above, multiple layers can be divided between the minimum radiation depth and the maximum radiation depth under the center of each point position, so that the depth of each two adjacent layers can meet the above formula, the depth information of the first layer is greater than or equal to the minimum radiation depth, and the depth information of the last layer is less than or equal to the maximum radiation depth. Thus, the number of layers under the center of this point position and the depth information of each layer can be obtained.

[0106] In some implementable ways, to ensure the accuracy of layer division under the center of each point position, the present application can determine the number of layers under the center of each point position and the depth information of each layer through the following steps: Take the minimum radiation depth as the depth information of the first layer, and take the first layer as the current layer; Execute the multi-layer depth determination step: According to the depth information of the current layer and the preset depth relationship between adjacent layers, determine the depth information of the next layer; Take the next layer as the new current layer, and continue to execute the above multi-layer depth determination step until the depth information of the latest layer is greater than or equal to the maximum radiation depth, so as to obtain the number of layers under the center of the point position and the depth information of each layer.

[0107] That is to say, under the center of each point position, the minimum radiation depth under the center of this point position can be taken as the depth information of the first layer. According to the formula satisfied by the depth relationship between adjacent layers above, the depth information of the second layer can be calculated. Then, take the second layer as the current layer, and continue to calculate the depth information of the third layer according to the formula satisfied by the depth relationship between adjacent layers above, and cycle in turn, continuously calculate the depth information of the next layer until the depth information of a certain latest layer is greater than or equal to the maximum radiation depth under the center of this point position, and then the layer division can be stopped. Thus, according to the layer division result under the center of this point position, the number of layers under the center of this point position and the depth information of each layer can be obtained.

[0108] S330, according to the depth information of each layer and the depth information of the light sampling points, perform multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image reconstructed by the three-dimensional scene under the center of this point position.

[0109] After determining the depth information of each layer at each point center, by comparing the depth information of each layer at the point center with the depth information of each light sampling point at the point center, the partial light sampling points that will affect the rendering result of each layer are determined. Then, the color information of the partial light sampling points that will affect the rendering result of each layer is projected onto the layer to perform color rendering on each layer at the point center, so as to obtain the rendered images on each layer at the point center, and form a multi-layer spherical shell image reconstructed from the three-dimensional scene at the point center.

[0110] In some implementable ways, in order to ensure the accuracy of the multi-layer spherical shell image at each point center, the present application can determine the multi-layer spherical shell image reconstructed from the three-dimensional scene at each point center in the following manner: according to the depth information of each layer and the depth information of the light sampling points, determine the associated light sampling points of each layer; perform volume rendering on the color information of the associated light sampling points of each layer to obtain the multi-layer spherical shell image reconstructed from the three-dimensional scene at the point center.

[0111] That is to say, by comparing the depth information of each layer at the point center with the depth information of each light sampling point at the point, the associated light sampling points of each layer can be determined.

[0112] Among them, since layer rendering can be divided into color rendering and transparency rendering, and the colors of all light sampling points after each layer will have a certain impact on the color rendering result of the layer, the associated light sampling points during the color rendering of each layer can be each light sampling point after the layer.

[0113] And the transparency rendering of each layer is only related to the transparency of each light sampling point between the layer and the next layer, so the associated light sampling points during the transparency rendering of each layer can be each light sampling point between the layer and the next layer.

[0114] Then, perform color rendering and transparency rendering of the light sampling points on each layer respectively. For the associated light sampling points involved in the color rendering of each layer, perform volume rendering on the color information in the color information of the associated light sampling points. Moreover, for the associated light sampling points involved in the transparency rendering of each layer, perform volume rendering on the transparency information in the color information of the associated light sampling points. After completing the color rendering and transparency rendering, the multi-layer spherical shell image reconstructed from the three-dimensional scene at the point center can be obtained.

[0115] Exemplarily, for each layer r m The formula for performing color rendering of the light sampling points can be:

[0116]

[0117] Among them, t 1 = r m ,t N = d far ,d far is the maximum radiation depth, and this represents that the associated sampling points involved in the color rendering of layer r m are all the ray sampling points after this layer r m .

[0118] For each layer r m , the formula for performing the transparency rendering of the ray sampling points can be:

[0119]

[0120] Among them, t 1 = r m ,t N = r m+1 , and this represents that the associated sampling points involved in the transparency rendering of layer r m are all the ray sampling points between this layer r m and the next layer r m+1 .

[0121] The technical solution provided by the embodiments of the present application divides the layers under the center of each point position through the minimum radiation depth, maximum radiation depth, and the preset depth relationship between adjacent layers under the center of each point position, so as to perform multi-layer rendering on the color information of each ray sampling point, obtain a multi-layer spherical shell image reconstructed under the center of each point position of the three-dimensional scene, thereby realizing the accurate reconstruction of the three-dimensional scene under multiple point positions. The virtual scene after the reconstruction of the three-dimensional scene is locally approximated by the multi-layer spherical shell images under multiple point positions, ensuring the true three-dimensional sense after the reconstruction of the three-dimensional scene, and reducing the computational overhead of the three-dimensional scene reconstruction. The accurate reconstruction of the three-dimensional scene can also be realized on mid-range and low-end XR devices, ensuring the comprehensive applicability of the three-dimensional scene reconstruction.

[0122] According to one or more embodiments of the present application, since after obtaining the multi-layer spherical shell images reconstructed under the center of each point position of the three-dimensional scene, it is necessary to present the multi-layer spherical shell images under the center of each point position on a client with relatively low operating performance, so that users can immerse themselves in the virtual scene under the center of each point position. Then, in order to ensure the efficient presentation of the multi-layer spherical shell images reconstructed under the center of each point position of the three-dimensional scene on the client, the present application can preset a number of layers as the preset layer upper limit that the client can support for efficient presentation after the three-dimensional scene reconstruction.

[0123] Then, after determining the number of layers and the depth information of each layer under the center of each point according to the minimum radiation depth, maximum radiation depth, and the preset adjacent layer depth relationship under the center of each point, it is first necessary to determine whether the number of layers is less than or equal to the preset upper limit of the number of layers.

[0124] If the number of layers under the center of a certain point is greater than the preset upper limit of the number of layers, it means that the client cannot ensure the efficient rendering of the multi-layer spherical shell image under the center of this point. Then, for the formula of the preset adjacent layer depth relationship where In this application, the radius of this point in the adjacent layer depth relationship can be reduced to re-determine the number of layers and the depth information of each layer under the center of this point according to the minimum radiation depth, maximum radiation depth, and the preset adjacent layer depth relationship, so that the re-determined number of layers is less than or equal to the preset upper limit of the number of layers.

[0125] That is to say, by continuously reducing the radius of this point in the adjacent layer depth relationship from the initial value represented by the expected roaming depth, the layers are re-divided between the minimum radiation depth and the maximum radiation depth according to the adjacent layer depth relationship again. At this time, in the adjacent layers r l and r l+1 By reducing the radius of this point in the adjacent layer depth relationship, the included angle formed when the two observable sight lines corresponding to the maximum pixel parallax pass through a certain pixel point on the previous layer r l is reduced Thereby increasing the depth interval between adjacent layers, so as to reduce the number of layers divided between the minimum radiation depth and the maximum radiation depth under the center of this point, making it gradually become less than or equal to the preset upper limit of the number of layers, and thus re-determining the number of layers and the depth information of each layer under the center of this point.

[0126] After re-obtaining the number of layers and the depth information of each layer under the center of each point, the color information of the light sampling points can be multi-layer rendered according to the depth information of each layer and the depth information of the light sampling points to obtain the multi-layer spherical shell image reconstructed by the three-dimensional scene under the center of this point.

[0127] However, when presenting the multi-layer spherical shell image reconstructed by the three-dimensional scene obtained at this time under the center of each point to the user, it is only supported to view the multi-layer spherical shell image under the center of this point within the reduced radius range of this point, while viewing the multi-layer spherical shell image under the center of this point within the range from the reduced radius to the expected roaming depth within this point, the multi-layer spherical shell image may be distorted.

[0128] Therefore, to ensure that at any position within each point, the distortion of the multi-layer spherical shell images centered at each point is avoided, after the present application re-divides the layers under the center of a certain point by continuously shrinking the radius of this point in the depth relationship of adjacent layers from the initial value represented by the expected roaming depth and obtaining the multi-layer spherical shell images under the center of this point, the present application also needs to optimize the multi-layer spherical shell images so that they can be presented without distortion within the points set by the expected roaming depth.

[0129] As Figure 5 shown, the process of optimizing the multi-layer spherical shell images reconstructed under the center of any point in the three-dimensional scene can be described as follows:

[0130] S510, determine a multi-view image sample sequence at multiple preset sampling poses within this point according to the neural radiance field.

[0131] If the present application re-divides the layers under the center of a certain point by continuously shrinking the radius of the point in the depth relationship of adjacent layers from the initial value represented by the expected roaming depth. Then, for the center of this point, the present application can re-set multiple sampling poses within the point. Moreover, the present application can analyze the particle information such as color, volume density, illumination intensity, and material of each light ray sampling point on each radiation ray formed under the center of this point in the neural radiance field of the three-dimensional scene, and can regenerate the scene images at each sampling pose, thereby forming a multi-view image sample sequence at multiple sampling poses within this point.

[0132] Among them, the multi-view image sample sequence can represent the real scene images that can be obtained from the three-dimensional scene under each sampling pose within this point.

[0133] S520, perform volume rendering on the intersection color information between the projection rays at each sampling pose and the multi-layer spherical shell images reconstructed under the center of this point, and obtain a multi-view rendering image sequence within the point.

[0134] In the neural radiance field, rays can be projected according to each sampling view within this point. Then, as Figure 6 shown, the projection rays at each sampling pose will intersect with the multi-layer spherical shell images that have been reconstructed under the center of this point, and by means of equirectangular projection (Equirectangular Projection, abbreviated as ERP), determine the coordinate points of the intersections of each projection ray on each spherical shell image, so as to determine the color information of each intersection from the appearance information of the three-dimensional scene.

[0135] Then, for each sampled pose, the following volume rendering formula designed in the neural radiance field can be used to perform volume rendering on the intersection color information between each projected ray at the sampled pose and each spherical shell image layer under the center of this point, and thus the rendering results at each sampled pose within this point can be obtained, forming a multi-view rendering image sequence within this point.

[0136] Exemplarily, the following volume rendering formula designed in the neural radiance field can be:

[0137]

[0138] wherein, n in the multi-layer spherical shell images can be arranged from far to near according to the distance between each layer and the point. c i can be the color information of the intersection point after the ray projected according to a certain sampled pose within this point intersects with each spherical shell image layer i in the multi-layer spherical shell images, and α i is the opacity information corresponding to the intersection point and can be represented by the volume density of this intersection point.

[0139] S530. Optimize the multi-layer spherical shell images reconstructed under the center of this point according to the difference between the multi-view image sample sequence and the multi-view rendering image sequence.

[0140] After obtaining the multi-view rendering image sequence rendered at each sampled pose through the multi-layer spherical shell images, by comparing the difference between the multi-view rendering image sequence and the two images belonging to the same sampled pose in the multi-view image sample sequence, the corresponding loss function is adopted to continuously adjust the color information and opacity information of each pixel point in the multi-layer spherical shell images, so that the two images belonging to the same sampled pose in the multi-view rendering image sequence and the multi-view image sample sequence can be kept consistent, thereby obtaining the optimized multi-layer spherical shell images under the center of this point, supporting the user to accurately view the multi-layer spherical shell images of the three-dimensional scene at any position point within each point without distortion, and ensuring the accuracy of the multi-layer spherical shell images of the three-dimensional scene under the center of each point.

[0141] Figure 7 This is a schematic block diagram of a three-dimensional scene reconstruction device provided by an embodiment of the present application. As Figure 7 shown, the three-dimensional scene reconstruction device 700 may include:

[0142] A radiance field construction module 710, configured to construct a neural radiance field of the three-dimensional scene according to a multi-view image sequence within the three-dimensional scene;

[0143] A point color determination module 720, configured to, for each given point center, determine a ray sampling point under the point center and the color information of the ray sampling point according to the neural radiance field;

[0144] A three-dimensional scene reconstruction module 730 is configured to perform multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image of the three-dimensional scene reconstructed under the center of the point position.

[0145] In some implementable manners, the three-dimensional scene reconstruction module 730 may include:

[0146] A radiation depth determination unit is configured to perform depth rendering on the light sampling points of each radiation ray under the center of the point position to obtain the minimum radiation depth and the maximum radiation depth under the center of the point position;

[0147] A layer division unit is configured to determine the number of layers under the center of the point position and the depth information of each layer according to the minimum radiation depth, the maximum radiation depth, and a preset adjacent layer depth relationship;

[0148] A multi-layer spherical shell image rendering unit is configured to perform multi-layer rendering on the color information of the light sampling points according to the depth information of each layer and the depth information of the light sampling points to obtain a multi-layer spherical shell image of the three-dimensional scene reconstructed under the center of the point position.

[0149] In some implementable manners, the layer division unit may specifically be configured to:

[0150] Use the minimum radiation depth as the depth information of the first layer, and use the first layer as the current layer;

[0151] Execute a multi-layer depth determination step: determine the depth information of the next layer according to the depth information of the current layer and a preset adjacent layer depth relationship;

[0152] Use the next layer as the new current layer, and continue to execute the above multi-layer depth determination step until the depth information of the latest layer is greater than or equal to the maximum radiation depth, so as to obtain the number of layers under the center of the point position and the depth information of each layer.

[0153] In some implementable manners, the multi-layer spherical shell image rendering unit may specifically be configured to:

[0154] Determine the associated light sampling points of each layer according to the depth information of each layer and the depth information of the light sampling points;

[0155] Perform volume rendering on the color information of the associated light sampling points of each layer to obtain a multi-layer spherical shell image of the three-dimensional scene reconstructed under the center of the point position.

[0156] In some implementable ways, the adjacent layer depth relationship satisfies that the maximum pixel parallax when the observable line of sight of any pixel point of the previous layer in the adjacent layers within the point position falls on the subsequent layer in the adjacent layers is less than or equal to one unit pixel.

[0157] In some implementable ways, the initial value of the radius of the point position is a preset expected roaming depth. If the number of layers under the center of the point position is greater than the preset layer upper limit, the three-dimensional scene reconstruction device 700 may further include a layer update module. The layer update module can be used for:

[0158] By reducing the radius of the point position in the adjacent layer depth relationship, to re-determine the number of layers under the center of the point position and the depth information of each layer according to the minimum radiation depth, the maximum radiation depth, and the preset adjacent layer depth relationship, so that the re-determined number of layers is less than or equal to the preset layer upper limit.

[0159] In some implementable ways, the three-dimensional scene reconstruction device 700 may further include a multi-layer spherical shell image optimization module. The multi-layer spherical shell image optimization module can be used for:

[0160] Determine a multi-view image sample sequence at a plurality of preset sampling poses within the point position according to the neural radiance field;

[0161] Perform volume rendering on the intersection color information between the projection rays at each sampling pose and the multi-layer spherical shell image reconstructed under the center of the point position to obtain a multi-view rendering image sequence within the point position;

[0162] Optimize the multi-layer spherical shell image reconstructed under the center of the point position according to the difference between the multi-view image sample sequence and the multi-view rendering image sequence.

[0163] In some implementable ways, the neural radiance field is represented based on light density.

[0164] In some implementable ways, the point position color determination module 720 can be specifically used for:

[0165] For each given point position center, determine the radiation rays under the point position center and the ray sampling points on the radiation rays according to the neural radiance field;

[0166] Determine the color information of the ray sampling points according to the neural radiance field.

[0167] In some implementable ways, the three-dimensional scene reconstruction device 700 may further include an image display module. The image display module can be used for:

[0168] Obtain the roaming pose information of the user, where the roaming pose information includes the roaming position and the roaming attitude;

[0169] If the roaming position is within the point position, display the corresponding roaming rendering image according to the roaming pose information and the multi-layer spherical shell image reconstructed under the center of the point position.

[0170] In the embodiments of the present application, first, a multi-view image sequence in a three-dimensional scene is processed to construct a neural radiance field of the three-dimensional scene. Then, for each given point position center, the respective ray sampling points and the color information of each ray sampling point under the center of the point position can be determined according to the neural radiance field, and multi-layer rendering is performed on the color information of each ray sampling point, so as to obtain the multi-layer spherical shell image reconstructed under the center of the point position in the three-dimensional scene, thereby realizing the accurate reconstruction of the three-dimensional scene under multiple point position centers. The virtual scene after the reconstruction of the three-dimensional scene is locally approximated by the multi-layer spherical shell images under multiple point position centers, ensuring the true three-dimensional sense after the reconstruction of the three-dimensional scene, and reducing the computational overhead of the three-dimensional scene reconstruction. The accurate reconstruction of the three-dimensional scene can also be realized on mid-range and low-end XR devices, ensuring the comprehensive applicability of the three-dimensional scene reconstruction.

[0171] It should be understood that the device embodiments in this application can correspond to the method embodiments in this application, and similar descriptions can refer to the method embodiments in this application. To avoid repetition, they will not be elaborated here.

[0172] Specifically, Figure 7 The shown device 700 can execute any method embodiment provided in this application, and Figure 7 The foregoing and other operations and / or functions of the respective modules in the shown device 700 respectively implement the corresponding processes of the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0173] In the above, the above method embodiments of the present application have been described from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0174] Figure 8 Schematic block diagram of the electronic device provided by an embodiment of the present application.

[0175] As Figure 8 shown, the electronic device 800 may include:

[0176] A memory 810 and a processor 820. The memory 810 is used to store a computer program and transmit the program code to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiment of the present application.

[0177] For example, the processor 820 can be used to execute the above method embodiment according to the instructions in the computer program.

[0178] In some embodiments of the present application, the processor 820 may include, but is not limited to:

[0179] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0180] In some embodiments of the present application, the memory 810 includes, but is not limited to:

[0181] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0182] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 800.

[0183] As Figure 8 shown, the electronic device may further include:

[0184] A transceiver 830, which may be connected to the processor 820 or the memory 810.

[0185] Among them, the processor 820 can control the transceiver 830 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 830 may include a transmitter and a receiver. The transceiver 830 may further include an antenna, and the number of antennas may be one or more.

[0186] It should be understood that the components in the electronic device 800 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0187] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments.

[0188] The embodiments of the present application also provide a computer program product including computer programs / instructions. When the computer programs / instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0189] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0190] The above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A three-dimensional scene reconstruction method, characterized in that, comprising: Constructing a neural radiance field of the three-dimensional scene according to a multi-view image sequence within the three-dimensional scene; For each given point center, determining light sampling points under the point center and color information of the light sampling points according to the neural radiance field; Performing multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image reconstructed for the three-dimensional scene under the point center.

2. The method according to claim 1, characterized in that, The performing multi-layer rendering on the color information of the light sampling points to obtain a multi-layer spherical shell image reconstructed for the three-dimensional scene under the point center includes: Performing depth rendering on the light sampling points of each radiation ray under the point center to obtain the minimum radiation depth and the maximum radiation depth under the point center; Determining the number of layers under the point center and depth information of each layer according to the minimum radiation depth, the maximum radiation depth, and a preset adjacent layer depth relationship; Performing multi-layer rendering on the color information of the light sampling points according to the depth information of each layer and the depth information of the light sampling points to obtain a multi-layer spherical shell image reconstructed for the three-dimensional scene under the point center.

3. The method according to claim 2, characterized in that, The determining the number of layers under the point center and depth information of each layer according to the minimum radiation depth, the maximum radiation depth, and a preset adjacent layer depth relationship includes Taking the minimum radiation depth as the depth information of the first layer, and taking the first layer as the current layer; Performing a multi-layer depth determination step: determining the depth information of the next layer according to the depth information of the current layer and the preset adjacent layer depth relationship; Taking the next layer as the new current layer, and continuing to perform the above multi-layer depth determination step until the depth information of the latest layer is greater than or equal to the maximum radiation depth, so as to obtain the number of layers under the point center and depth information of each layer.

4. The method according to claim 2, characterized in that, The performing multi-layer rendering on the color information of the light sampling points according to the depth information of each layer and the depth information of the light sampling points to obtain a multi-layer spherical shell image reconstructed for the three-dimensional scene under the point center includes: Determining associated light sampling points of each layer according to the depth information of each layer and the depth information of the light sampling points; Performing volume rendering on the color information of the associated light sampling points of each layer to obtain a multi-layer spherical shell image reconstructed for the three-dimensional scene under the point center.

5. The method according to claim 2 or 3, characterized in that, The adjacent layer depth relationship satisfies that for any pixel point of the previous layer in the adjacent layers within the point, the maximum pixel parallax when the observable line of sight of the previous layer falls on the next layer in the adjacent layers is less than or equal to one unit pixel.

6. The method according to claim 5, characterized in that, The initial value of the radius of the point is a preset expected roaming depth. If the number of layers under the center of the point is greater than the preset upper limit of the number of layers, the method further includes: By reducing the radius of the point in the adjacent layer depth relationship, to re-determine the number of layers under the center of the point and the depth information of each layer according to the minimum radiation depth, the maximum radiation depth and the preset adjacent layer depth relationship, so that the re-determined number of layers is less than or equal to the preset upper limit of the number of layers.

7. The method according to claim 1, wherein, the method further includes: Determining a multi-view image sample sequence at a plurality of preset sampling poses within the point according to the neural radiance field; Performing volume rendering on the intersection color information between the projection rays at each sampling pose and the multi-layer spherical shell images reconstructed under the center of the point, to obtain a multi-view rendering image sequence within the point; Optimizing the multi-layer spherical shell images reconstructed under the center of the point according to the difference between the multi-view image sample sequence and the multi-view rendering image sequence.

8. The method according to claim 1, wherein, the neural radiance field is represented based on light density.

9. The method according to claim 1, wherein, the determining, for each given center of a point, the light sampling points under the center of the point and the color information of the light sampling points according to the neural radiance field includes: For each given center of a point, determining the radiation rays under the center of the point and the light sampling points on the radiation rays according to the neural radiance field; Determining the color information of the light sampling points according to the neural radiance field.

10. The method according to claim 1, wherein, the method further includes: Obtaining the roaming pose information of the user, where the roaming pose information includes the roaming position and the roaming attitude; If the roaming position is within the point, displaying the corresponding roaming image according to the roaming pose information and the multi-layer spherical shell images reconstructed under the center of the point.

11. A three-dimensional scene reconstruction device, wherein, it includes: A radiation field construction module, configured to construct a neural radiance field of the three-dimensional scene according to a multi-view image sequence within the three-dimensional scene; A point color determination module, configured to determine, for each given center of a point, the light sampling points under the center of the point and the color information of the light sampling points according to the neural radiance field; A three-dimensional scene reconstruction module, configured to perform multi-layer rendering on the color information of the light sampling points to obtain multi-layer spherical shell images reconstructed under the center of the point of the three-dimensional scene.

12. An electronic device, wherein, it includes: A processor; and A memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the three-dimensional scene reconstruction method according to any one of claims 1-10 by executing the executable instructions.

13. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the three-dimensional scene reconstruction method according to any one of claims 1-10.

14. A computer program product comprising computer programs / instructions, characterized in that, when the computer programs / instructions comprised in the computer program product are run on an electronic device, the electronic device is caused to execute the three-dimensional scene reconstruction method according to any one of claims 1-10.

Citation Information

Cited By

  • Three-dimensional scene simulation system and method based on VR technology

    CN120997460A

  • A three-dimensional scene simulation system and method based on VR technology

    CN120997460B