Virtual scene display method and device, equipment and storage medium

The three-dimensional reconstruction of the real scene through the neural radiation field solves the problem of user roaming in virtual reality, and achieves a higher sense of immersion and three-dimensional roaming experience.

CN120070807APending Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311602928.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the existing virtual reality technology, users lack three-dimensionality roaming in virtual scenes, resulting in poor immersion.

Method used

By using neural radiation fields to reconstruct the real scene three-dimensionally, a virtual scene with a real three-dimensional sense is generated, and a high-fidelity and three-dimensional roaming image is rendered in real time based on the user's roaming position information and perspective information.

Benefits of technology

It improves the three-dimensionality and immersion of users roaming in virtual scenes, allowing users to experience an immersive roaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070807A_ABST
    Figure CN120070807A_ABST
Patent Text Reader

Abstract

The invention provides a virtual scene display method and device, equipment and a storage medium, and the method comprises the steps: obtaining the roaming information of a user in a virtual scene, and the roaming information comprises roaming position information and roaming view angle information; and obtaining a target roaming image according to the roaming position information, the roaming view angle information and target virtual scene data, the target virtual scene data being determined according to the roaming position information and virtual scene data, and the virtual scene data comprising a neural radiation field generated at least partially based on a real scene. The roaming immersion of the user in the virtual scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a method, apparatus, device, and storage medium for displaying a virtual scene. Background Art

[0002] With the continuous development of Extended Reality (XR) technology, XR technology has been widely applied to various scenario experiences, such as roaming in virtual reality scenarios such as houses, tourist attractions, and buildings, enabling the public to roam around the world without leaving home through electronic devices that provide virtual scenes.

[0003] Currently, the virtual scene roaming uses three-degree-of-freedom panoramic images. Therefore, when a user roams in a virtual scene, the user can only view the three-degree-of-freedom panoramic images at fixed roaming points, that is, three-degree-of-freedom roaming is achieved. However, this three-degree-of-freedom roaming makes the user's roaming in the virtual scene lack a sense of three-dimensionality, resulting in a poor immersion in the scene roaming. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, device, and storage medium for displaying a virtual scene, which can improve the immersion of a user roaming in a virtual scene.

[0005] In a first aspect, an embodiment of the present application provides a method for displaying a virtual scene, which is applied to a terminal device. The method includes:

[0006] Obtain roaming information of a user in a virtual scene, where the roaming information includes roaming position information and roaming viewing angle information;

[0007] According to the roaming position information, the roaming viewing angle information, and target virtual scene data, obtain a target roaming image, where the target virtual scene data is determined according to the roaming position information and virtual scene data, and the virtual scene data includes at least part of a neural radiance field generated based on a real scene.

[0008] In a second aspect, an embodiment of the present application provides a device for displaying a virtual scene, which is configured in a terminal device and includes:

[0009] An information acquisition module, configured to obtain roaming information of a user in a virtual scene, where the roaming information includes roaming position information and roaming viewing angle information;

[0010] An image display module, configured to obtain a target roaming image according to the roaming position information, the roaming viewing angle information, and target virtual scene data, where the target virtual scene data is determined according to the roaming position information and virtual scene data, and the virtual scene data includes at least part of a neural radiance field generated based on a real scene.

[0011] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0012] a processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the virtual scene display method described in the first aspect embodiment or its various implementation manners.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, and the computer program causes a computer to execute the virtual scene display method described in the first aspect embodiment or its various implementation manners.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product containing program instructions. When the program instructions run on an electronic device, the electronic device is caused to execute the virtual scene display method described in the first aspect embodiment or its various implementation manners.

[0015] The technical solution provided by the embodiment of the present application obtains the roaming position information and roaming perspective information of the user in the virtual scene, obtains the target virtual scene data according to the roaming position information and roaming perspective information, and further obtains the target roaming image according to the roaming position information, roaming perspective information, and target virtual scene data, so that when the user roams in the virtual scene, it has a stronger three-dimensional sense, thereby improving the immersion of the user in the virtual scene roaming to achieve an immersive roaming experience. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic diagram of an application scenario for an embodiment of the present application;

[0018] Figure 2 It is a schematic flowchart of a virtual scene display method provided by an embodiment of the present application;

[0019] Figure 3 It is a schematic flowchart of a process for obtaining target virtual scene data provided by an embodiment of the present application;

[0020] Figure 4 It is a schematic diagram of a virtual scene obtained by three-dimensional reconstruction of a real scene of a house provided by an embodiment of the present application;

[0021] Figure 5 Schematic diagram of a first - type roaming point provided by an embodiment of the present application, including a central area and a boundary area;

[0022] Figure 6 Flow schematic diagram of another virtual - scene display method provided by an embodiment of the present application;

[0023] Figure 7 Schematic diagram of a user moving from a first - type roaming point C to a second - type roaming point U provided by an embodiment of the present application;

[0024] Figure 8 Schematic block diagram of a virtual - scene display device provided by an embodiment of the present application;

[0025] Figure 9 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. According to the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above - mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.

[0029] In the description of the embodiments of the present application, unless otherwise specified, "a plurality" means two or more, that is, at least two. "At least two" means two or more. "At least one" means one or more.

[0030] To facilitate the understanding of the embodiments of the present application, before describing each embodiment of the present application, some concepts involved in all embodiments of the present application are first explained appropriately as follows:

[0031] 1) Virtual Reality (VR), a technology for creating and experiencing virtual worlds, computes and generates a virtual environment, which is a multi-source information (the virtual reality mentioned in this article includes at least visual perception, and may also include auditory perception, tactile perception, motion perception, and even taste perception, olfactory perception, etc.). It realizes the fusion of the virtual environment, an interactive three-dimensional dynamic visual scene and the simulation of entity behavior, enabling users to immerse themselves in the simulated virtual reality environment and realizing applications in various virtual environments such as maps, games, videos, education, medical treatment, simulation, collaborative training, sales, assisting manufacturing, maintenance and repair.

[0032] 2) Virtual Reality device (VR device), a terminal that realizes the virtual reality effect, usually provided in the form of glasses, head-mounted display (HMD), or contact lenses to achieve visual perception and other forms of perception. Of course, the form in which the virtual reality device is realized is not limited to this, and it can be further miniaturized or enlarged according to actual needs.

[0033] Optionally, the virtual reality devices described in the embodiments of the present application may include but are not limited to the following types:

[0034] 2.1) PC-based Virtual Reality (PCVR) device, which uses the PC to perform relevant calculations and data output for virtual reality functions. The externally connected PCVR device uses the data output by the PC to achieve the virtual reality effect.

[0035] 2.2) Mobile virtual reality device, which supports setting a mobile terminal (such as a smart phone) in various ways (such as a head-mounted display with a dedicated card slot), and through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations for virtual reality functions and outputs data to the mobile virtual reality device. For example, watch virtual reality videos through the APP of the mobile terminal.

[0036] 2.3) All-in-one virtual reality device, which has a processor for performing relevant calculations for virtual functions, thus having independent virtual reality input and output functions, and does not need to be connected to a PC or a mobile terminal, with high freedom of use.

[0037] 3) XR refers to the combination of the real and the virtual through a computer to create a virtual environment for human-computer interaction. XR is also a general term for various technologies such as VR, Augmented Reality (AR for short), and Mixed Reality (MR for short). By integrating the visual interaction technologies of the three, it brings a sense of "immersion" of seamless conversion between the virtual world and the real world to users.

[0038] 4) AR: A technology that, during the process of a camera capturing an image, calculates in real time the camera pose parameters of the camera in the real world (or three-dimensional world, real world), and adds virtual elements to the image captured by the camera according to the camera pose parameters. Virtual elements include, but are not limited to: images, videos, and three-dimensional models. The goal of AR technology is to interact by overlaying the virtual world on the real world on the screen.

[0039] 5) MR: By presenting virtual scene information in a real scene, it builds an interactive feedback information loop among the real world, the virtual world, and the user to enhance the sense of reality of the user experience. For example, integrating the sensory inputs created by a computer (such as virtual objects) with the sensory inputs from a physical set or their representations in a simulated set. In some MR sets, the sensory inputs created by the computer can adapt to changes in the sensory inputs from the physical set. Additionally, some electronic systems for presenting MR sets can monitor the orientation and / or position information relative to the physical set so that virtual objects can interact with real objects (i.e., physical elements from the physical set or their representations). For example, the system can monitor movement so that a virtual plant appears stationary relative to a physical building.

[0040] 6) A virtual scene (also known as a virtual space) is a virtual scene displayed (or provided) when an application runs on an electronic device. The virtual scene can be a simulation environment of the real world, a semi-simulated and semi-fictional virtual scene, or a purely fictional virtual scene. The virtual scene can be any one of a two-dimensional virtual scene, a 2.5D virtual scene, or a three-dimensional virtual scene. The embodiments of the present application do not limit the dimension of the virtual scene. For example, the virtual scene can include the sky, land, sea, etc., and the land can include environmental elements such as deserts and cities.

[0041] 7) A virtual object is an object that interacts in a virtual scene, is controlled by a user or a robot program (such as a robot program based on artificial intelligence), and can be stationary, move, and perform various behaviors in the virtual scene.

[0042] 8) Three degrees of freedom (3Dof) refer to the degrees of freedom with three rotational angles, that is, the ability to rotate on the X, Y, and Z axes, without the ability to move on the X, Y, and Z axes.

[0043] 9) Six degrees of freedom (6Dof) refer to the degrees of freedom with three rotational angles and three position-related degrees of freedom such as up and down, front and back, left and right. That is, in addition to the ability to rotate on the X, Y, and Z axes, it also has the ability to move on the X, Y, and Z axes.

[0044] Considering that when users currently perform panoramic roaming in a virtual scene based on an electronic device, they can only view 3Dof panoramic images at fixed roaming points, making the roaming in the virtual scene lack a sense of three-dimensionality and resulting in a poor immersion in the scene roaming. To solve the above technical problems, the inventive concept of this application is: through the use of a Neural Radiance Field (NeRF for short) to perform three-dimensional reconstruction on a real scene to obtain a virtual scene with a real three-dimensional sense. Then, when the user roams in this virtual scene, corresponding high-fidelity and three-dimensional roaming images are displayed to the user according to the user's roaming information, making the user more three-dimensional when roaming in the virtual scene, thereby improving the immersion of the user in the virtual scene roaming to achieve an immersive roaming experience.

[0045] It should be understood that the technical solution of this application can be applied but is not limited to the following scenarios:

[0046] Such as Figure 1 As shown, the application scenario may include: a terminal device 110 and a server 120. Among them, the terminal device 110 can communicate with the server 120 through a network to achieve data interaction.

[0047] In some alternative embodiments, the above terminal device 110 may be various electronic devices capable of providing a virtual scene and virtual scene roaming functions. Such as VR devices, XR devices, smartphones (such as Android phones, IOS phones, Windows Phone phones, etc.), tablets, laptops, ultra-mobile personal computers (UMPCs), personal digital assistants (PDAs), etc. This application does not make specific restrictions on the type of electronic device. It should be understood that the terminal device 110 in this application may also be referred to as a user equipment (UE), a terminal, or a user device, etc., and no restrictions are made here.

[0048] When the above terminal device 110 is a virtual scene product such as a VR device or an XR device, it is preferably a head-mounted display (HMD) under a VR device or an XR device.

[0049] In some alternative embodiments, the above server 120 can be various types of servers, such as traditional servers and cloud servers, etc. Among them, the traditional server can be, but is not limited to: a file server and a database server.

[0050] In this application scenario, the terminal device 110 can obtain the three-dimensional model data (i.e., virtual scene data) corresponding to the virtual scene from the server 120 or the storage module of the terminal device 110 itself based on the user's trigger operation for any virtual scene, enabling the user to enter the virtual scene, and rendering the corresponding roaming image according to the user's roaming information in the virtual scene, that is, rendering the corresponding target virtual scene image based on the user's real-time pose in the virtual scene, so that the user can obtain a more three-dimensional and immersive roaming experience in the virtual scene.

[0051] It should be noted that Figure 1 the terminal device 110 and the server 120 shown in Figure 1 are only illustrative. The quantity and type of the terminal device 110 and the server 120 can be adjusted according to actual usage needs, and are not limited to

[0052] shown.

[0053] Figure 2 It is a schematic flowchart of a method for displaying a virtual scene provided by an embodiment of the present application. The method for displaying a virtual scene provided by the present application can be executed by a display device for a virtual scene. The display device for the virtual scene can be composed of hardware and / or software and can be integrated into the terminal device. As Figure 2 shown, the method can include the following steps:

[0054] S101, obtain the user's roaming information in the virtual scene, where the roaming information includes roaming position information and roaming perspective information.

[0055] In the present application, the virtual scene can be understood as a three-dimensional model obtained by performing three-dimensional geometric space reconstruction on any real scene. Among them, the real scene can be a house, a tourist attraction or other buildings, etc.

[0056] The above-mentioned three-dimensional geometric space reconstruction of the real scene can be performed using a multi-view stereo (MVS) reconstruction method, or can be performed using NeRF, etc. The present application does not impose any restrictions on the three-dimensional reconstruction method of the real scene.

[0057] When considering three-dimensional reconstruction of a real scene using the MVS reconstruction method, there may be weak texture regions in each image of the multi-view image sequence, making it difficult for the MVS reconstruction method to represent the subtle and complex geometric structures in the real scene. Moreover, it is also difficult to reproduce the details of illumination and materials in the real scene with high fidelity in the texture of each image in the multi-view image sequence. NeRF, on the other hand, is a method that represents a real scene as a neural radiance field approximated by a neural network, which can accurately describe the color information and volume density (i.e., volumetric density) of each spatial point in the real scene in each viewing direction, making it more advantageous than the MVS reconstruction method in terms of the realism of rendering and the reduction degree of complex geometric structures. Therefore, this application preferably uses NeRF to reconstruct the real scene to obtain a three-dimensional solid model with a high-fidelity appearance and a more refined geometric structure.

[0058] In some alternative embodiments, when using NeRF to reconstruct a real scene, this application may first set a collection device with a shooting function, such as a camera or a camcorder, at at least one position in the real scene, so as to perform omnidirectional image collection of the real scene from multiple perspectives through the collection device at each position, and obtain multiple frames of real environment images in the real scene. Then, the multiple frames of real environment graphics are used as the multi-view image sequence for three-dimensional reconstruction in this application.

[0059] Furthermore, NeRF performs three-dimensional reconstruction of the real scene based on the multi-view image sequence to obtain a corresponding virtual scene. It should be understood that using NeRF to perform three-dimensional reconstruction of the real scene is to construct the neural radiance field of the real scene, so as to achieve high-precision and high-fidelity modeling of the real scene through the neural radiance field.

[0060] As an alternative implementation, NeRF obtains the neural radiance field of the real scene based on the multi-view image sequence in the real scene. The specific implementation process is as follows: By analyzing the positions and orientations of the collection devices corresponding to each perspective image in the multi-view image sequence, corresponding radiance rays are formed from the position of the collection device towards each pixel point in the imaging plane of the collection device, so as to obtain multiple radiance rays under the multi-view image sequence. Furthermore, for each radiance ray, a ray sampling strategy is adopted, and multiple ray sampling points are continuously sampled on the radiance ray. Then, by processing the texture features in each real scene image in the multi-view image sequence, the color information and volumetric density of each ray sampling point on each radiance ray are predicted, thereby obtaining the corresponding neural radiance field of the real scene.

[0061] It should be understood that after the above multi-view image sequence is sorted according to the shooting perspectives, a panoramic image of the real scene can be formed, so as to comprehensively understand the geometric structure and appearance information of the real scene.

[0062] Considering that although NeRF has more advantages in terms of the restoration degree of geometric structure and rendering realism, since the real scene is not a strict surface structure, it is difficult to directly estimate the geometric structure by extracting the zero-value surface. Therefore, in order to accurately describe the geometric structure and real rendering texture of the real scene, this application can adopt two different methods according to the differences between the geometric structure and the image texture to represent two different neural radiance fields that can independently predict the geometric information and appearance information of the real scene. Then, using the multi-view image sequence in the real scene, train the two neural radiance fields represented by these two different methods initially constructed to obtain the geometric neural radiance field and the appearance neural radiance field in this application respectively.

[0063] Among them, the geometric neural radiance field can use the collected multi-view images as training samples to train the initially constructed neural radiance field, and continuously analyze the scene geometric structure inside it during the training process, so that after the training is completed, the geometric neural radiance field can be used to accurately estimate the geometric information of the real scene. The appearance neural radiance field uses the collected multi-view images as training samples to train the initially constructed neural radiance field, and continuously analyze the scene appearance information inside it during the training process, so that after the training is completed, the appearance neural radiance field can be used to accurately estimate the high-fidelity appearance information of the real scene.

[0064] That is to say, the geometric neural radiance field can be expressed as the relative distance between each ray sampling point on each radiance ray in the real scene and the nearest real object in the real scene, so as to represent the geometric structure of each real object in the real scene.

[0065] The above-mentioned appearance neural radiance field can include particle information such as the color, volume density, light intensity, material, etc. of each ray sampling point on each radiance ray in the real scene, so as to represent the actual texture of each real object in the real scene.

[0066] In some optional implementation manners, in order to ensure a highly restored geometric structure of the real scene, this application can represent the geometric neural radiance field based on the Signed Distance Fields (SDF for short). Among them, SDF means to establish a spatial field, and the value of each voxel in this spatial field can represent the distance between this voxel and the nearest geometric image in the spatial field (that is, the geometric surface graph of each real object in the real scene). If the voxel is inside the real object, the distance is set to negative. If the voxel is outside the real object, the distance is set to positive. And if the voxel is on the boundary of the geometric surface graph of the real object, the distance is set to zero.

[0067] Therefore, when training a geometric neural radiance field based on SDF representation, corresponding geometric feature processing can be performed on a multi-view image sequence to analyze the distance between each ray sampling point in the real scene and the nearest geometric surface graph of each real object, so as to extract each ray sampling point with a distance of zero, and then the geometric neural radiance field of the real scene can be obtained.

[0068] To ensure the high-fidelity appearance information of the real scene, this application can represent the appearance neural radiance field based on ray density, so as to analyze particle information such as the color, volume density, illumination intensity, and material of each ray sampling point in the real scene, and then the appearance neural radiance field of the real scene can be generated to represent the high-fidelity appearance information of the real scene.

[0069] As an optional implementation, this application can obtain the geometric neural radiance field of the real scene through the following steps:

[0070] Step 1: Input the multi-view image sequence in the real scene into the trained geometric prior model to obtain the geometric prior information of the real scene.

[0071] To ensure the accuracy of the geometric structure reconstruction of the real scene, this application can additionally train a geometric prior model to analyze information such as the spatial structure and depth of the real scene through image feature processing, and use this as an additional constraint condition during the training of the geometric neural radiance field to ensure the accuracy of the geometric neural radiance field of the real scene.

[0072] Therefore, for the multi-view image sequence in the real scene, this application can first input the multi-view image sequence into the trained geometric prior model respectively, so as to perform corresponding feature processing on each real environment image in the multi-view image sequence through the geometric prior model in advance to predict the geometric prior information of the real scene, such as normal vectors, depth information, etc., thereby providing an additional constraint for the training of the geometric neural radiance field to ensure the construction accuracy of the geometric neural radiance field of the real scene.

[0073] Step 2: Obtain the geometric neural radiance field of the real scene according to the multi-view image sequence and the geometric prior information.

[0074] After obtaining the geometric prior information of the real scene, comprehensive analysis can be performed on the multi-view image sequence and the geometric prior information, so as to perform corresponding processing on the geometric structure features in each real environment image in the multi-view image sequence under the additional constraint of the geometric prior information, thereby accurately constructing the geometric neural radiance field of the real scene.

[0075] Thus, by structuring each geometric information represented in the geometric neural radiance field of the real scene, a three-dimensional mesh model of the real scene (denoted as Mesh) can be generated, facilitating subsequent user roaming in the three-dimensional mesh model.

[0076] In some alternative embodiments, the three-dimensional mesh model of the real scene in this application can also be generated based on the MVS reconstruction method or other methods, and this application does not impose any restrictions on this.

[0077] It should be noted that the three-dimensional reconstruction of the real scene to obtain a virtual scene in this application can be implemented on a terminal device, or alternatively, it can also be implemented on a server side communicatively connected to the terminal device, and this application does not impose any restrictions on this. Considering that the performance of various devices or components of the terminal device is lower than that of the server side, and various processing processes need to be carried out on the terminal device, while the reconstruction of the virtual scene requires a large amount of computing resources. Therefore, this application preferably configures the operation of three-dimensional reconstruction of the real scene on the server side, so as to reduce the resource occupancy of the terminal device, relieve the computing pressure of the terminal device, and provide conditions for improving the performance of the terminal device.

[0078] When specifically executing step S101, the user can select any virtual scene from multiple virtual scenes provided by the terminal device, so that the terminal device wakes up the virtual scene in the sleep state and controls the user to enter the virtual scene. Alternatively, the terminal device can obtain the three-dimensional model data (virtual scene data) of the virtual scene from the server side, load the virtual scene based on the virtual scene data, and control the user to enter the loaded virtual scene. Furthermore, the user can perform roaming operations within the virtual scene.

[0079] Considering that multiple users can enter a virtual scene simultaneously, each user can enter the virtual scene in the form of a virtual object when entering, so that a user can see the walking actions or other roaming interaction operations of other users within the virtual scene. The virtual object corresponding to each user above can be automatically assigned by the terminal device according to the default rules, or it can also be a personalized virtual object created by the user based on the object creation function provided by the terminal device, etc., and this application does not impose any restrictions on this.

[0080] In some alternative embodiments, when the user enters the virtual scene, optionally, the user first enters the birth roaming point, and then the user can move from the birth roaming point to other roaming points, so that the user can start roaming operations within the virtual scene from the birth roaming point, making the roaming images corresponding to the virtual scene more orderly and providing a better experience.

[0081] The above-mentioned birth roaming point can be a roaming point located at the entrance of the virtual scene, or a roaming point at the center position, or a roaming point at any other position, etc. The present application does not impose any restrictions on this.

[0082] It should be understood that the above-mentioned roaming point can be a specific position point or a roaming area. When the roaming point is a roaming area, the roaming area can be a cylindrical space, a spherical space, etc. Among them, when the roaming area is a cylindrical space, the cylindrical space is centered on a position point, with a first distance as the radius to determine two parallel upper and lower bases, and a second distance as the height to construct. When the roaming area is a spherical space, the spherical space is centered on a position point and obtained with a third distance as the radius. The above-mentioned first distance, second distance, and third distance are all adjustable parameters, which can be specifically set flexibly according to the scene roaming requirements.

[0083] In some alternative embodiments, when the user roams in the virtual scene, a movement operation will be triggered. For example, when the user moves from the birth roaming point to other roaming points, the movement direction can be controlled by a joystick to move towards the target roaming point, or the movement can also be achieved by clicking on the target roaming point, or the movement can also be achieved when the user gazes at the target roaming point through eye tracking for a preset duration, etc. Among them, controlling the movement with a joystick is applicable to operations under a touch screen, keyboard, or gamepad. Clicking to control the movement is applicable to operations under finger touch click or mouse click. The above-mentioned target roaming point can be understood as the point where the user wants to move.

[0084] When the user roams in the virtual scene, the terminal device can obtain the user's roaming information in real time, that is, the roaming position information and the roaming perspective information. Furthermore, based on the obtained roaming position information and roaming perspective information, the corresponding roaming image is displayed to the user, so as to display the corresponding virtual scene image based on the user's location, thereby providing a more immersive and three-dimensional roaming experience for the user. Among them, the roaming perspective information includes the field of view angle and the line of sight direction.

[0085] In the present application, the user's roaming pose information can be determined based on the inertial measurement unit data obtained by the inertial measurement unit (Inertial Measurement Unit, abbreviated as IMU) in the terminal device, or can be obtained through calculation of the environmental images collected by the image acquisition device of the terminal device. For the specific determination of the roaming pose information, reference can be made to the existing solutions, and no more elaboration is provided here.

[0086] S102, obtain the target roaming image according to the roaming position information, the roaming perspective information, and the target virtual scene data.

[0087] Among them, the target virtual scene data is determined according to the roaming position information and the virtual scene data, and the virtual scene data includes at least part of the neural radiance field generated based on the real scene.

[0088] Considering that there are some important positions or areas in the real scene that are often observed by users, such as aisles, doorways, or living rooms, etc., and some unimportant positions or areas that are not often observed by users, such as corners, turns, etc. Therefore, when reconstructing the real scene using the neural radiance field in this application, it is optional to perform high-precision and high-fidelity reconstruction on the important positions or areas in the real scene that are often observed by users, while using traditional 3D reconstruction methods for the unimportant positions or areas that are not often observed by users, such as the MVS reconstruction method or the binocular stereo vision reconstruction method, etc. This can reduce the reconstruction calculation cost for the real scene. Of course, in order to obtain a high-fidelity and high-precision virtual scene, this application can also use the neural radiance field to model each position or area in the real scene, and this application does not make any restrictions on this.

[0089] The above virtual scene data can be understood as the 3D model data for the 3D reconstruction of any real scene.

[0090] The above neural radiance field is to represent the real scene as a radiance field approximated by a neural network.

[0091] In order to obtain a roaming image corresponding to the user's roaming information, this application first obtains the target virtual scene data from the virtual scene data corresponding to the virtual scene according to the roaming position information. Then, according to the roaming position information and the roaming viewing angle information, it determines the image corresponding to the target virtual scene data, so as to render the image and display the obtained target roaming image.

[0092] That is to say, based on the user's roaming position information and roaming viewing angle information in the virtual scene, this application can display a virtual scene image corresponding to the roaming position information and roaming viewing angle information to the user, so that the user can obtain a three-dimensional and immersive roaming experience when roaming in the virtual scene.

[0093] The technical solution provided by the embodiments of this application obtains the user's roaming position information and roaming viewing angle information in the virtual scene, obtains the target virtual scene data according to the roaming position information and roaming viewing angle information, and then obtains the target roaming image according to the roaming position information, roaming viewing angle information, and target virtual scene data, so that when the user roams in the virtual scene, it is more three-dimensional, thereby improving the immersion of the user in the virtual scene roaming to achieve an immersive roaming experience.

[0094] In another alternative implementation scenario, considering that there are roaming points set in the virtual scene, when the present application obtains the target virtual scene based on the roaming position information and roaming perspective information of the user in the virtual scene, it can first determine which roaming point the user is currently located at according to the roaming position information of the user, and then obtain the target virtual scene data based on this roaming point, so as to realize obtaining the target virtual scene data based on the roaming point where the user's roaming position information is located. The following combines Figure 3 , and specifically describes the obtaining of the target virtual scene data.

[0095] As Figure 3 shown, the method includes the following steps:

[0096] S201, obtain the roaming information of the user in the virtual scene, where the roaming information includes the roaming position information and the roaming perspective information.

[0097] S202, determine whether the roaming position information is located at a first type of roaming point. If so, execute step S203; otherwise, execute step S204.

[0098] For the convenience of the user's roaming interaction operation in the virtual scene, the present application can set multiple first type of roaming points and multiple second type of roaming points in the virtual scene, so that the user can perform roaming operations in the virtual scene based on these roaming points.

[0099] Among them, the first type of roaming points can be multiple default points preset in the real scene. The second type of roaming points can be multiple custom points selected by the user in the virtual scene after the virtual scene reconstructed from the real scene is displayed, and the position information corresponding to each custom point is obtained. It should be noted that the user who selects multiple custom points here specifically refers to the content producer who performs three-dimensional reconstruction on the real scene.

[0100] It should be understood that both the above-mentioned first type of roaming points and the second type of roaming points can support the user's 6-degree-of-freedom roaming operation at this point. That is, when the user roams at any roaming point, a 6-degree-of-freedom roaming method of changing the roaming position and changing the roaming perspective can be realized. It should be understood that there can also be only the first type of roaming points in the virtual scene and no second type of roaming points. That is, the user can only switch between the first type of roaming points in the virtual scene.

[0101] In addition, the above-mentioned first type of roaming points can also support the personalized setting requirements of content producers. Specifically, content producers can mark key points in the virtual scene and use the marked key points as the first type of roaming points. In this way, when the default points preset in the virtual scene do not meet the scene production requirements, content producers can personalize the default points in the virtual scene to meet different virtual scene production requirements.

[0102] The above-mentioned second type of roaming points specifically refers to the points other than the first type of roaming points. That is, the second type of roaming points is any point other than the first type of roaming points. And this arbitrary point can be understood as any position point.

[0103] Exemplarily, assuming the real scene is a certain house, then the virtual scene of this house can be as Figure 4 shown. Figure 4 In the virtual scene shown, there are multiple first type of roaming points 310 and multiple second type of roaming points 320.

[0104] Therefore, after obtaining the roaming position information and roaming perspective information of the user in the virtual scene, the present application can determine which roaming point the user is currently located in based on the roaming position information. Furthermore, target virtual scene data is obtained according to the roaming point where the user is currently located.

[0105] Considering that each first type of roaming point in the virtual scene corresponds to a roamable area. And this roamable area can be a circular area that can be used for roaming determined with a certain position point as the center point and a preset observation distance as the radius. Among them, the preset observation distance is an adjustable parameter and is flexibly adjusted according to roaming requirements.

[0106] Therefore, when the present application determines which roaming point the user is currently located in based on the roaming position information, it can first determine the boundary position points of the roamable areas corresponding to each first type of roaming point, and compare the user's roaming position information with the boundary position points of the roamable areas corresponding to each first type of roaming point. If the user's roaming position information is within the position interval formed by the boundary position points of the roamable areas corresponding to any first type of roaming point, it is determined that the user is currently located in this first type of roaming point. If the user's roaming position information is not within the position area formed by the boundary position points of the roamable areas corresponding to all first type of roaming points, it is determined that the user is currently located in any second type of roaming point.

[0107] S203, if the roaming position information is located in a first type of roaming point, the obtained target virtual scene data includes the spherical shell data corresponding to the current first type of roaming point, and this spherical shell data is generated according to the neural radiance field.

[0108] Among them, the current first - type roaming point can be understood as when the roaming position information of the user is within the roamable area of a certain first - type roaming point, this first - type roaming point is the current first - type roaming point. That is to say, the first - type roaming point where the user is currently located.

[0109] Considering that the first - type roaming points can be preset default points, and these default points can support the user to perform 6 - degree - of - freedom roaming with position changes and perspective changes at these points. Also, because there is spherical shell data corresponding to each first - type roaming point in the virtual scene, the target virtual scene data obtained according to the roaming position information includes the spherical shell data corresponding to the current first - type roaming point.

[0110] In this application, the spherical shell data is multi - layer spherical shell data (Multi - Sphere Image, abbreviated as MSI) generated based on the neural radiance field for the first - type roaming points. That is, the current first - type roaming point is input into the neural radiance field to generate multi - layer spherical shell data corresponding to the center point of the current first - type roaming point through the neural radiance field.

[0111] Among them, the multi - layer spherical shell data MSI is a data format that extends multi - layer plane images (Multi - Plane Image, abbreviated as MPI) to a 360 - degree spherical surface. When performing three - dimensional reconstruction of the real scene in this application, by reconstructing a multi - layer spherical shell image at each first - type roaming point in the real scene, it is possible to present a multi - layer spherical shell image in a surrounding manner at each first - type roaming point, thereby enhancing the real three - dimensional sense after the real scene reconstruction.

[0112] In some alternative embodiments, considering that the first - type roaming points in the virtual scene can include: a central area and a boundary area. Optionally, the central area can be a circular area determined with the center point of the roamable area corresponding to the first - type roaming point as the center and a certain distance value less than the radius of the roamable area as the radius. The above - mentioned distance value is an adjustable parameter as long as it is less than the radius of the roamable area. Correspondingly, the boundary area can be an annular area that is concentric with the central area obtained by subtracting the central area from the roamable area corresponding to the first - type roaming point.

[0113] Exemplarily, such as Figure 5As shown in the figure, assume that the center point of the roamable area corresponding to the first type of roaming point is (X, Y), and the radius of the roamable area is 2 meters. Then, for the central area of the first type of roaming point, a circular area Q1 can be determined with the center point (X, Y) as the center and a distance value of 1.5 meters, which is less than the radius of the roamable area, as the radius. Correspondingly, for the boundary area of the first type of roaming point, the circular area Q1 can be subtracted from the roamable area corresponding to the first type of roaming point to obtain an annular area Q2 that is concentric with the circular area Q. The concentric point of the circular area Q1 and the annular area Q2 is the center point (X, Y) of the first type of roaming point.

[0114] It should be understood that the sizes of the central area and the boundary area of the first type of roaming point can be flexibly adjusted according to actual roaming requirements, and the present application does not make any settings in this regard.

[0115] Therefore, when the roaming position information is located at the first type of roaming point, the steps for obtaining the target virtual scene data may include the following:

[0116] Step S11: Detect whether the roaming position information is within the central area of the current first type of roaming point. If the roaming position information is within the central area of the current first type of roaming point, then execute step S12. If the roaming position information is within the boundary area of the current first type of roaming point, then execute step S13.

[0117] In some alternative embodiments, first, determine the boundary position points of the central area of the current first type of roaming point and the boundary position points of the boundary area of the current first type of roaming point. Then, compare the roaming position information with the boundary position points of the central area and the boundary position points of the boundary area respectively. If the roaming position information is within the position interval formed by the boundary position points of the central area, it is determined that the user is within the central area of the current first type of roaming point. If the roaming position information is within the position interval formed by the boundary position points of the boundary area, it is determined that the user is within the boundary area of the current first type of roaming point.

[0118] Step S12: If it is detected that the roaming position information is within the central area of the current first type of roaming point, obtain the spherical shell data corresponding to the current first type of roaming point as the target virtual scene data.

[0119] Step S13: If it is detected that the roaming position information is within the boundary area of the current first type of roaming point, obtain the spherical shell data corresponding to the current first type of roaming point and the grid data corresponding to the roaming position information, and use the spherical shell data and the grid data as the target virtual scene data.

[0120] The above grid data is grid data generated based on a three-dimensional grid model of a real scene. That is, the roaming position information is input into the three-dimensional grid model to output grid data corresponding to the roaming position information through the three-dimensional grid model. Among them, the grid data can be represented as Mesh data.

[0121] Since there are multiple position points in the roamable area corresponding to the first type of roaming point in the virtual scene of this application, and each position point corresponds to a Mesh data. Therefore, when it is determined that the roaming position information of the user is within the boundary area of any first type of roaming point, this application will obtain the spherical shell data corresponding to the first type of roaming point where the user is currently located, and the grid data of the position point corresponding to the user in the roamable area of this first type of roaming point.

[0122] That is to say, the target virtual scene data obtained by this application according to the roaming position information may include, in addition to the spherical shell data, the grid data corresponding to the roaming position information.

[0123] S204, if the roaming position information is located at the second type of roaming point, the obtained target virtual scene data includes the grid data corresponding to the roaming position information.

[0124] Considering that the second type of roaming point is other points except the default point, and each second type of roaming point does not have corresponding virtual scene data, but each position point in the roamable area corresponding to each second type of roaming point has corresponding virtual scene data. And, the virtual scene data corresponding to each position point in the roamable area corresponding to each second type of roaming point is specifically grid data. Therefore, when it is determined that the roaming position information of the user is located at a certain second type of roaming point, the obtained target virtual scene data is the grid data corresponding to the roaming position information.

[0125] S205, obtain a target roaming image according to the roaming position information, the roaming viewing angle information, and the target virtual scene data.

[0126] After obtaining the target virtual scene data according to the roaming position information, this application can determine the image corresponding to the target virtual scene data according to the roaming position information and the roaming viewing angle information, so as to render the image and display the rendered target roaming image.

[0127] In some alternative embodiments, considering that the roaming position information may be located at a certain first type of roaming point or a certain second type of roaming point, so this application obtains a target roaming image according to the roaming position information, the roaming viewing angle information, and the target virtual scene data, which may include the following situations:

[0128] In the first case, when the target virtual scene data includes the spherical shell data corresponding to the current first type of roaming point, the panoramic image corresponding to the spherical shell data is determined according to the roaming position information and the roaming viewing angle information, and the target roaming image is obtained by rendering the panoramic image.

[0129] In the second case, when the target virtual scene data includes the spherical shell data corresponding to the current first type of roaming point and the grid data corresponding to the roaming position information, the panoramic image corresponding to the spherical shell data and the target image corresponding to the grid data are determined according to the roaming position information and the roaming viewing angle information. Then, the panoramic image corresponding to the spherical shell data and the target image corresponding to the grid data are mixed and rendered to obtain the target roaming image.

[0130] Among them, the mixed rendering of the panoramic image corresponding to the spherical shell data and the target image corresponding to the grid data can be weighted mixed rendering or other fusion renderings, etc., and this application does not impose any restrictions on this.

[0131] In the third case, when the target virtual scene data includes the grid data corresponding to the roaming position information, the target image corresponding to the grid data is determined according to the roaming position information and the roaming viewing angle information. Then, the target image is rendered to obtain the target roaming image.

[0132] In this application, the rendering of the images determined in the above three cases is specifically volume rendering.

[0133] Moreover, when the panoramic image determined in the above several cases corresponds to the multi-layer spherical shell data, that is, the panoramic image is a multi-layer spherical shell image, the volume rendering of this multi-layer spherical shell image in this application can be implemented by the following formula:

[0134]

[0135] Among them, c r is the RGB color value finally presented on the target roaming image for the pixel point where the user's observation light (multiple sampled light rays emitted from the user's eyes) intersects the i-th layer of the spherical shell layer. ∑ is the summation symbol, n is the n-layer spherical shell layer corresponding to the center point of the user's roaming position information, 1 ≤ n ≤ N, N is an integer greater than 2, and n is sorted in the order from far to near according to the distance between the spherical shell layer and the user's eyes. α i is the opacity of the pixel point where the user's observation light intersects the i-th layer of the spherical shell layer. c i is the RGB color value of the pixel point where the user's observation light intersects the i-th layer of the spherical shell layer. j is the multi-layer spherical shell layer located between the i-th layer of the spherical shell layer and the roaming position information. ∏ is the product symbol. α j is the opacity of the pixel point where the user's observation light intersects the j-th layer of the spherical shell plane.

[0136] In some alternative embodiments, considering that a communication connection is established between the terminal device and the server, the terminal device can send the roaming location information and roaming perspective information of the user to the server. Furthermore, the server determines the target virtual scene data based on the roaming location information, and obtains the target roaming image based on the roaming location information, roaming perspective information, and target virtual scene data. After obtaining the target roaming image, the server can send the target roaming image to the terminal device, so that the terminal device can directly display the target roaming image, thereby avoiding the image rendering operation of the terminal device and improving the display speed of the roaming image.

[0137] When the target roaming image obtained by the server based on the roaming location information, roaming perspective information, and target virtual scene data sent by the terminal device is an optional multi-layer spherical shell image, the server can perform volume rendering on the multi-layer spherical shell image based on the volume rendering formula designed in the neural radiance field generated based on the real scene. The specific formula is as follows:

[0138]

[0139] Where n is the serial number of each ray sampling point on the same ray under the first type of roaming point position arranged in the order from near to far according to the distance. n t represents the distance between the nth ray sampling point and the first type of roaming point position. n σ can be the volume density of the nth ray sampling point. n c can be the RGB color value of the nth ray sampling point.

[0140] Then, (1 - exp(-σ n (t n+1 - t n ))) can represent the opacity of the nth ray sampling point.

[0141] T(t 1 →t n ) can be the transmittance from the 1st ray sampling point to the nth ray sampling point when the ray sampling points on the same ray under the first type of roaming point position are arranged in the order from near to far according to the distance.

[0142] Where δ k represents the distance between adjacent ray sampling points, that is, t k+1 - t k .

[0143] By following the above formula, for each roaming point of the first type, the color information of the partial light sampling points among the light sampling points at this point that can affect each layer can be projected onto this layer, and the color intensity after projection on this layer can be calculated to obtain the rendered image on each layer under this roaming point of the first type, thereby obtaining the multi-layer spherical shell image reconstructed under this roaming point of the first type.

[0144] The technical solution provided by the embodiments of the present application obtains the roaming position information and roaming perspective information of the user in the virtual scene, obtains the target virtual scene data according to the roaming position information and roaming perspective information, and further obtains the target roaming image according to the roaming position information, roaming perspective information, and target virtual scene data, making the user more three-dimensional when roaming in the virtual scene, thereby improving the immersion of the user's roaming in the virtual scene to achieve an immersive roaming experience.

[0145] In an optional implementation scenario of the present application, considering that the first type of roaming points set in the virtual scene includes: a central area and a boundary area. Then, when the roaming position information of the user is located in the boundary area of any first type of roaming point, by performing spatial compression on the roaming position information of the user, it can be ensured that the user can move freely within the current first type of roaming point without exceeding the roamable area corresponding to this first type of roaming point, thereby ensuring that the quality of the displayed roaming image is always in the best state, and there will be no problems such as distortion or deformation of the displayed roaming image due to the roaming position information of the user being in the boundary area of the current first type of roaming point or about to move out of the boundary area. The following combines Figure 6 , and further illustrates the display method of the virtual scene provided by the embodiments of the present application.

[0146] As Figure 6 shown, the method may include the following steps:

[0147] S301, obtain the roaming information of the user in the virtual scene, where the roaming information includes roaming position information and roaming perspective information.

[0148] S302, in response to the roaming position information being located in the boundary area of the first type of roaming point, perform spatial compression on the roaming position information according to the non-linear compression method, and obtain the spherical shell data corresponding to the current first type of roaming point and the grid data corresponding to the roaming position information, so as to use the spherical shell data and the grid data as the target virtual scene data.

[0149] S303, obtain the target roaming image according to the roaming position information, roaming perspective information, and target virtual scene data.

[0150] Considering the characteristics of the spherical shell image, when approaching or exceeding the roamable area, problems such as distortion or deformation may occur in the rendering result because the conditions of local approximation are no longer satisfied. Therefore, when it is determined that the roam position information of the user is located in the boundary area of the current first type of roam point, by performing non-linear compression on the roam position information of the user, it is realized that the closer the user's moving position is to the boundary position of the roamable area, the more intense the spatial compression is, so that the actual moving distance of the user is smaller, enabling the user to move freely and smoothly within the roamable area of the first type of roam point without exceeding the roamable area, and thus high-fidelity roam images that do not suffer from distortion or deformation can always be displayed to the user.

[0151] In some alternative embodiments, the non-linear compression of the roam position information of the user according to the non-linear compression method can be achieved through the following formula:

[0152]

[0153] where contract(x) is the position information obtained after spatial compression of the roam position information of the user, x is the roam position information of the user, || || is the modulus symbol, a is the spatial compression parameter, and this parameter is an adjustable parameter, r i is the radius of the roamable area of the current first type of roam point.

[0154] Exemplarily, assume that the user is initially located within the central area of the first type of roam point B. When the user moves from the central area of the first type of roam point B to the boundary area of the first type of roam point B, the above pose compression formula is used to perform spatial compression on the roam position information of the user in the boundary area of the first type of roam point B, so that the user does not move out of the roamable area of the first type of roam point B, thereby ensuring that the user can always see high-fidelity roam images at the first type of roam point.

[0155] In some alternative embodiments, the present application further includes: in response to the selection operation of any new roam point, controlling the user to jump from the current roam point to the new roam point, and switching the roam image corresponding to the current roam point to the roam image corresponding to the new roam point; wherein, the current roam point is the first type of roam point or the second type of roam point, and the new roam point is another first type of roam point or another second type of roam point other than the current roam point.

[0156] Exemplarily, as Figure 7 shown, assume that the user is currently located in the roamable area of the first type of roam point C. Then the user can click on the second type of roam point U in the virtual scene, so that the user jumps from the first type of roam point C to the second type of roam point U, and at the same time switches the roam image corresponding to the first type of roam point C to the roam image corresponding to the second type of roam point U.

[0157] That is, the user can select any roaming point except the current roaming point within the virtual scene for point jump. Thus, when the user jumps from the current roaming point to a new roaming point, the roaming image corresponding to the new roaming point where the user is located after the jump can be viewed.

[0158] Among them, the way for the user to select any roaming point except the current roaming point can be achieved by means such as a joystick, clicking, or eye movement tracking. For specific details, reference can be made to the part of the user's movement operation in the virtual scene in the foregoing embodiments, and no further elaboration will be made here.

[0159] The technical solution provided by the embodiments of the present application obtains the roaming position information and roaming perspective information of the user in the virtual scene, obtains the target virtual scene data according to the roaming position information and roaming perspective information, and then obtains the target roaming image according to the roaming position information, roaming perspective information, and target virtual scene data, making the user more stereoscopic when roaming in the virtual scene, thereby improving the immersion of the user in virtual scene roaming to achieve an immersive roaming experience. In addition, when it is detected that the roaming position after the user's movement is located in the boundary area of the default roaming point, spatial compression is performed on the roaming position after the user's movement so that the user will not move out of the roamable area of the first type of roaming point, thereby ensuring that the user can always see a high-fidelity roaming image at this roaming point.

[0160] Next, with reference to the attached Figure 8 , a display device for a virtual scene proposed in the embodiments of the present application will be described. Figure 8 It is a schematic block diagram of a display device for a virtual scene provided by the embodiments of the present application.

[0161] As Figure 8 shown, the display device 400 for the virtual scene includes: an information acquisition module 410 and an image display module 420.

[0162] Among them, the information acquisition module 410 is used to acquire the roaming information of the user in the virtual scene, and the roaming information includes roaming position information and roaming perspective information;

[0163] The image display module 420 is used to obtain a target roaming image according to the roaming position information, the roaming perspective information, and the target virtual scene data. The target virtual scene data is determined according to the roaming position information and the virtual scene data, and the virtual scene data includes at least part of the neural radiance field generated based on the real scene.

[0164] One or more alternative implementations of the embodiments of the present application. The virtual scene includes at least one first type of roaming point, the roaming position information is located at the first type of roaming point, the target virtual scene data includes spherical shell data corresponding to the current first type of roaming point, and the spherical shell data is generated according to the neural radiance field.

[0165] One or more alternative implementations of the embodiments of the present application. The spherical shell data is multi-layer spherical shell data.

[0166] One or more alternative implementations of the embodiments of the present application. The target virtual scene data further includes grid data corresponding to the roaming position information.

[0167] One or more alternative implementations of the embodiments of the present application. The first type of roaming point includes: a central area and a boundary area. When the roaming position information is located in the central area of the first type of roaming point, the target virtual scene data includes spherical shell data corresponding to the current first type of roaming point; when the roaming position information is located in the boundary area of the first type of roaming point, the target virtual scene data includes spherical shell data corresponding to the current first type of roaming point and grid data corresponding to the roaming position information.

[0168] One or more alternative implementations of the embodiments of the present application. The apparatus 400 further includes:

[0169] A compression module, configured to perform spatial compression on the roaming position information in a non-linear compression manner in response to the roaming position information being located in the boundary area of the first type of roaming point.

[0170] One or more alternative implementations of the embodiments of the present application. The virtual scene further includes at least one second type of roaming point, the roaming position information is located at the second type of roaming point, and the target virtual scene data includes grid data corresponding to the roaming position information.

[0171] One or more alternative implementations of the embodiments of the present application. The image display module 420 is specifically configured to:

[0172] In response to the target virtual scene data including spherical shell data corresponding to the current first type of roaming point and grid data corresponding to the roaming position information, perform hybrid rendering on the spherical shell data corresponding to the current first type of roaming point and the grid data corresponding to the roaming position information to obtain the target roaming image.

[0173] One or more alternative implementations of the embodiments of the present application. The apparatus 400 further includes:

[0174] An image switching module, configured to respond to a selection operation of any new roaming point, control the user to jump from the current roaming point to the new roaming point, and switch the roaming image corresponding to the current roaming point to the roaming image corresponding to the new roaming point;

[0175] Wherein, the current roaming point is a first type of roaming point or a second type of roaming point, and the new roaming point is another first type of roaming point or another second type of roaming point other than the current roaming point.

[0176] It should be understood that the device embodiments and the foregoing method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 8 The illustrated device 400 can execute Figure 2 The corresponding method embodiments, and the foregoing and other operations and / or functions of each module in the device 400 are respectively for implementing Figure 2 The corresponding processes in each method in, and for the sake of brevity, will not be elaborated here.

[0177] The device 400 of the embodiments of the present application has been described above from the perspective of functional modules in conjunction with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the first aspect of the embodiments of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in software form. Combining the steps of the method in the first aspect disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the method embodiments in the first aspect above.

[0178] Figure 9 It is a schematic block diagram of an electronic device provided by the embodiments of the present application. As Figure 9 shown, the electronic device 500 may include:

[0179] A memory 510 and a processor 520. The memory 510 is used to store a computer program and transmit the program code to the processor 520. In other words, the processor 520 can call and run the computer program from the memory 510 to implement the method for displaying a virtual scene in the embodiments of the present application.

[0180] For example, the processor 520 can be used to execute the display method embodiments of the above virtual scenario according to the instructions in the computer program.

[0181] In some embodiments of the present application, the processor 520 may include but is not limited to:

[0182] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0183] In some embodiments of the present application, the memory 510 includes but is not limited to:

[0184] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0185] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 510 and executed by the processor 520 to implement the method for displaying a virtual scene provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.

[0186] As Figure 9 shown, the electronic device 500 may further include:

[0187] a transceiver 530, which may be connected to the processor 520 or the memory 510.

[0188] Among them, the processor 520 may control the transceiver 530 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 530 may include a transmitter and a receiver. The transceiver 530 may further include an antenna, and the number of antennas may be one or more.

[0189] It should be understood that each component in the electronic device is connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0190] The present application further provides a computer-readable storage medium for storing a computer program, and the computer program enables a computer to execute the method for displaying a virtual scene described in the above method embodiments.

[0191] Embodiments of the present application further provide a computer program product including program instructions. When the program instructions run on an electronic device, the electronic device is enabled to execute the method for displaying a virtual scene described in the above method embodiments.

[0192] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0193] Those of ordinary skill in the art will realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0194] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in an electrical, mechanical, or other form.

[0195] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, each functional module can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0196] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0197] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for displaying a virtual scene, characterized in that, applied to a terminal device, the method includes: Obtain the roaming information of the user in the virtual scene, where the roaming information includes roaming position information and roaming perspective information; According to the roaming position information, the roaming perspective information, and the target virtual scene data, obtain a target roaming image, where the target virtual scene data is determined according to the roaming position information and the virtual scene data, and the virtual scene data includes at least part of the neural radiance field generated based on the real scene.

2. The method according to claim 1, characterized in that, The virtual scene includes at least one first - type roaming point, the roaming position information is located at the first - type roaming point, the target virtual scene data includes the spherical shell data corresponding to the current first - type roaming point, and the spherical shell data is generated according to the neural radiance field.

3. The method according to claim 2, characterized in that, The spherical shell data is multi - layer spherical shell data.

4. The method according to claim 2, characterized in that, The target virtual scene data further includes the grid data corresponding to the roaming position information.

5. The method according to claim 4, characterized in that, The first - type roaming points include: a central area and a boundary area. When the roaming position information is located in the central area of the first - type roaming point, the target virtual scene data includes the spherical shell data corresponding to the current first - type roaming point; when the roaming position information is located in the boundary area of the first - type roaming point, the target virtual scene data includes the spherical shell data corresponding to the current first - type roaming point and the grid data corresponding to the roaming position information.

6. The method according to claim 5, characterized in that, The method further includes: In response to the roaming position information being located in the boundary area of the first - type roaming point, perform spatial compression on the roaming position information in a non - linear compression manner.

7. The method according to claim 2, characterized in that, The virtual scene further includes at least one second - type roaming point, the roaming position information is located at the second - type roaming point, and the target virtual scene data includes the grid data corresponding to the roaming position information.

8. The method according to claim 5, characterized in that, The obtaining of the target roaming image according to the roaming position information, the roaming perspective information, and the target virtual scene data includes: In response to the target virtual scene data including the spherical shell data corresponding to the current first - type roaming point and the grid data corresponding to the roaming position information, perform hybrid rendering on the spherical shell data corresponding to the current first - type roaming point and the grid data corresponding to the roaming position information to obtain the target roaming image.

9. The method according to any one of claims 1 - 8, characterized in that, The method further includes: In response to a selection operation of any new roaming point, control the user to jump from the current roaming point to the new roaming point, and switch the roaming image corresponding to the current roaming point to the roaming image corresponding to the new roaming point; Wherein, the current roaming point is a first type of roaming point or a second type of roaming point, and the new roaming point is another first type of roaming point or another second type of roaming point other than the current roaming point.

10. A display device for a virtual scene, characterized in that, configured in a terminal device, comprising: an information acquisition module, configured to acquire roaming information of a user in a virtual scene, where the roaming information includes roaming position information and roaming viewing angle information; an image display module, configured to obtain a target roaming image according to the roaming position information, the roaming viewing angle information, and target virtual scene data, where the target virtual scene data is determined according to the roaming position information and virtual scene data, and the virtual scene data includes at least a part of a neural radiance field generated based on a real scene.

11. An electronic device, characterized in that, comprising: a processor and a memory, where the memory is configured to store a computer program, and the processor is configured to call and run the computer program stored in the memory to execute the virtual scene display method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, configured to store a computer program, where the computer program causes a computer to execute the virtual scene display method according to any one of claims 1-9.