Scene rendering method and device, computer equipment and storage medium

By creating and training geometric representation data and texture representation networks for three-dimensional scenes, the low efficiency problem caused by the large amount of processing in NeRF technology is solved, and efficient and high-quality scene rendering is achieved.

CN120833430APending Publication Date: 2025-10-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410472929.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-18
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

NeRF technology has a large processing volume, resulting in low processing efficiency and high time consumption, making it difficult to meet the needs of efficient scene rendering.

Method used

By creating geometric representation data and a texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints, training description parameters and color values ​​of sample Gaussian points, and using the trained geometric representation data and texture representation network to render the scene, the volume rendering process is reduced.

Benefits of technology

The processing efficiency of scene rendering is improved, the quality of rendered scene images is enhanced, and the requirements of processing efficiency and image quality are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833430A_ABST
    Figure CN120833430A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a scene rendering method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps of creating geometric representation data and a texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints; determining a description parameter and a color value of a sample Gaussian point through the geometric representation data and the texture representation network, and determining a second scene image corresponding to the sample viewpoint based on the description parameter, the color value and a position parameter of the sample viewpoint; training geometric representation data and a texture representation network based on an error between a first scene image and a second scene image corresponding to the sample viewpoint; and performing scene rendering on the three-dimensional scene through the trained geometric representation data and texture representation network. The processing efficiency is improved, and the quality of the scene image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computer, in particular to a scene rendering method and device, computer device and storage medium. BACKGROUND

[0002] Scene rendering technology is a technology of converting a three-dimensional scene into a two-dimensional scene image. In the fields of animation production, film special effects or game development, scene rendering technology plays an extremely important role.

[0003] NeRF (Nerual Radiance Fields) is a technology of using a neural network to implicitly represent a three-dimensional scene. NeRF is trained based on scene images corresponding to at least two viewpoints of a three-dimensional scene, thereby implicitly storing the structure of the three-dimensional scene in the NeRF. In the NeRF, a mapping relationship between the positions of three-dimensional points in the three-dimensional scene and the positions of the viewpoints, and the color values and density values of the three-dimensional points is established. Therefore, given the position of any new viewpoint, based on the positions of each three-dimensional point in the three-dimensional scene and the position of the new viewpoint, the color values and density values of each three-dimensional point can be determined by the NeRF, thereby rendering a scene image corresponding to the new viewpoint. However, the processing amount of the NeRF technology is very large, resulting in very high time consumption and very low processing efficiency. SUMMARY

[0004] Embodiments of the present application provide a scene rendering method, device, computer device and storage medium, which improve the processing efficiency and the quality of the scene image. The technical solution is as follows:

[0005] In one aspect, a scene rendering method is provided, the method comprising:

[0006] creating a geometric representation data and a texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints, the first scene images being images of the three-dimensional scene at the first viewpoints, the geometric representation data comprising a plurality of Gaussian points and description parameters of each Gaussian point, the description parameters at least comprising three-dimensional coordinates, the texture representation network being used to determine a color value of any Gaussian point based on the three-dimensional coordinates of the Gaussian point;

[0007] determining description parameters and a color value of a sample Gaussian point through the geometric representation data and the texture representation network, and determining a second scene image corresponding to a sample viewpoint based on the description parameters, the color value and position parameters of the sample viewpoint, the sample Gaussian point comprising at least one Gaussian point in the geometric representation data, the sample viewpoint comprising at least one first viewpoint;

[0008] train the geometry representation data and the texture representation network based on an error between the first scene image corresponding to the sample viewpoint and the second scene image;

[0009] render the three-dimensional scene based on the trained geometry representation data and the texture representation network.

[0010] In another aspect, a scene rendering apparatus is provided, the apparatus comprising:

[0011] a creating module configured to create geometry representation data and a texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints, the first scene images being images of the three-dimensional scene at the first viewpoints, the geometry representation data comprising a plurality of Gaussian points and description parameters of each of the Gaussian points, the description parameters comprising at least three-dimensional coordinates, the texture representation network being configured to determine a color value of any of the Gaussian points based on three-dimensional coordinates of the any of the Gaussian points;

[0012] a training module configured to determine description parameters and a color value of a sample Gaussian point based on the geometry representation data and the texture representation network, the sample Gaussian point comprising at least one of the Gaussian points, and determine a second scene image corresponding to a sample viewpoint based on the description parameters, the color value, and a position parameter of the sample viewpoint, the sample viewpoint comprising at least one of the first viewpoints;

[0013] the training module is further configured to train the geometry representation data and the texture representation network based on an error between the first scene image corresponding to the sample viewpoint and the second scene image;

[0014] a scene rendering module configured to render the three-dimensional scene based on the trained geometry representation data and the texture representation network.

[0015] In a possible implementation, the creating module comprises:

[0016] a geometry creating unit configured to create first geometry representation data of the three-dimensional scene based on at least two of the first scene images, the first geometry representation data comprising a plurality of feature points in the three-dimensional scene and three-dimensional coordinates of each of the feature points in the three-dimensional scene;

[0017] a converting unit configured to convert each of the feature points in the first geometry representation data into a Gaussian point to obtain second geometry representation data, the second geometry representation data comprising a plurality of the Gaussian points and description parameters of each of the Gaussian points;

[0018] a network creating unit, configured to determine a preset network as the texture representation network, and initialize network parameters in the texture representation network.

[0019] In a possible implementation, the sample Gaussian points include each Gaussian point in the geometry representation data, and the training module includes:

[0020] a parameter obtaining unit, configured to obtain a description parameter of each Gaussian point from the geometry representation data;

[0021] a color value determining unit, configured to respectively determine a color value of each Gaussian point based on a three-dimensional coordinate of each Gaussian point by using the texture representation network;

[0022] a projecting unit, configured to project each Gaussian point into an image plane corresponding to the sample viewpoint based on the description parameter and the color value of each Gaussian point and a position parameter of the sample viewpoint, to obtain the second scene image corresponding to the sample viewpoint.

[0023] In a possible implementation, the description parameter further includes transparency.

[0024] The projecting unit is configured to, for each pixel point in the image plane, determine an associated Gaussian point based on a position parameter of the pixel point and the position parameter of the sample viewpoint, the associated Gaussian point being a Gaussian point projected into a position of the pixel point according to an observation direction of the sample viewpoint; mix color values of each associated Gaussian point based on transparency of each associated Gaussian point, to obtain a color value of the pixel point; and determine the second scene image based on the color value of each pixel point in the image plane.

[0025] In a possible implementation, the texture representation network is a tensor radiance field, and the training module includes:

[0026] a parameter obtaining unit, configured to obtain a description parameter of the sample Gaussian point from the geometry representation data;

[0027] a spatial decomposition unit, configured to perform spatial decomposition on a three-dimensional coordinate of the sample Gaussian point, to obtain a component of the three-dimensional coordinate, and determine a texture feature value of the sample Gaussian point based on the component of the three-dimensional coordinate;

[0028] a mapping unit, configured to map the texture feature value of the sample Gaussian point and a position parameter of the sample viewpoint based on network parameters in the texture representation network, to obtain a color value of the sample Gaussian point.

[0029] In a possible implementation, the training module is configured to train the description parameters of at least one Gaussian point in the geometry representation data and the network parameters in the texture representation network based on the error.

[0030] In a possible implementation, the scene rendering module is configured to determine the description parameters and color values of each Gaussian point by using the trained geometry representation data and the trained texture representation network, and determine a scene image corresponding to the second viewpoint based on the description parameters, the color values, and a position parameter of the second viewpoint.

[0031] In another aspect, a computer device is provided, which includes a processor and a memory, and the memory stores at least one computer program, which is loaded and executed by the processor to implement the operations performed by the scene rendering method according to the above aspects.

[0032] In another aspect, a computer readable storage medium is provided, which stores at least one computer program, which is loaded and executed by a processor to implement the operations performed by the scene rendering method according to the above aspects.

[0033] In another aspect, a computer program product is provided, which includes a computer program, which is loaded and executed by a processor to implement the operations performed by the scene rendering method according to the above aspects.

[0034] The scheme provided by the embodiments of the present application trains geometry representation data and a texture representation network based on a first scene corresponding to at least two first viewpoints, improves the accuracy of the geometry representation data and the texture representation network, so that the trained geometry representation data and the texture representation network are used together to describe a three-dimensional scene, and when performing scene rendering, the volume rendering process in the NeRF technology does not need to be performed for Gaussian points, the processing amount is reduced, the time consumption is saved, the processing efficiency is improved, and the texture representation network is used to describe the color of the three-dimensional scene, the trained texture representation network can ensure the texture continuity of the three-dimensional scene, and thus the quality of the rendered scene image is improved. Therefore, the embodiments of the present application combine the use of the geometry representation data and the texture representation network, and meet the demand for processing efficiency and the quality of the scene image. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0036] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0037] Figure 2 is a flowchart of a scene rendering method provided by an embodiment of the present application;

[0038] Figure 3 is a schematic diagram of a processing flow provided by an embodiment of the present application;

[0039] Figure 4 is a flowchart of another scene rendering method provided by an embodiment of the present application;

[0040] Figure 5 is a schematic diagram of a processing flow of making a 3D photo provided by an embodiment of the present application;

[0041] Figure 6 is a schematic diagram of a processing flow of making a movie resource containing "bullet time" provided by an embodiment of the present application;

[0042] Figure 7 is a structural schematic diagram of a scene rendering apparatus provided by an embodiment of the present application;

[0043] Figure 8 is a structural schematic diagram of another scene rendering apparatus provided by an embodiment of the present application;

[0044] Figure 9 is a structural schematic diagram of a terminal provided by an embodiment of the present application;

[0045] Figure 10 is a structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0046] To make the objectives, technical solutions and advantages of embodiments of the present application clearer, the following will further describe the embodiments of the present application with reference to the accompanying drawings.

[0047] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, a first scene image can be referred to as a second scene image, and similarly, a second scene image can be referred to as a first scene image.

[0048] At least two refers to two or more than two, for example, at least two scene images can be two scene images, three scene images, or any integer greater than or equal to two scene images. Each refers to each of the at least two, for example, each scene image refers to each of the at least two scene images, and if the at least two scene images are three scene images, each scene image refers to each of the three scene images.

[0049] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are fully authorized by the user or relevant aspects, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0050] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0051] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely used in downstream tasks of various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, etc.

[0052] Pre-Training Model (PTM), also known as cornerstone model, large model, refers to a deep neural network (DNN) with large parameters. It is trained on a large amount of unlabeled data. PTM extracts common features from data using the function approximation capability of large parameter DNN. Through fine tuning, parameter-efficient fine-tuning (PEFT), prompt tuning and other technologies, PTM is suitable for downstream tasks. Therefore, pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language model, visual model, speech model and multi-modal model according to the data modality processed, among which multi-modal model refers to a model that establishes feature representation of two or more data modalities. Pre-training model is an important tool for outputting artificial intelligence generated content, and can also be used as a general interface connecting multiple specific task models.

[0053] Computer Vision (CV) is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further to do image processing, so that the computer processing becomes an image more suitable for human observation or transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, and tries to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Large model technology brings important changes to the development of computer vision technology. Swin-Transformer (a kind of visual converter), ViT (Vision Transformer, a kind of image classification model), V-MoE (Vision-Mixture of Experts, vision mixture of experts network), MAE (Masked AutoEncoders, masked autoencoders) and other pre-training models in the field of vision can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition, optical character recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3 Dimension, three-dimensional) technology, virtual reality, augmented reality and other technologies.

[0054] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content, conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0055] The solutions provided in the embodiments of this application involve artificial intelligence technology, which is specifically illustrated by the following embodiments:

[0056] First, the terms involved in the embodiments of this application are introduced as follows:

[0057] 1. MLP (Multi-Layer Perception), which includes multiple fully connected layers, is a fully connected multi-layer neural network.

[0058] 2. NeRF (Nerual Radiance Fields) is a scene reconstruction and new viewpoint rendering technology. It is trained on scene images of the same scene from different viewpoints. It can implicitly model the structure of the 3D scene in the MLP network. Therefore, given any new viewpoint, it can be sampled through inverse ray tracing technology and rendered into the scene image corresponding to the new viewpoint.

[0059] 3DGS (3D Guassians splatting) is a scene reconstruction and new viewpoint rendering technology. It is trained on scene images corresponding to the same scene but different viewpoints. It can model the structure of the 3D scene using a certain number of 3D Gaussian distributions in 3D space. Therefore, given any new viewpoint, it can render the image of the new viewpoint through forward rasterization projection.

[0060] The method provided by the embodiments of the present application is used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the terminal is a smart phone, a computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, and the like, but is not limited thereto. Optionally, the server is a standalone physical server, or is a server cluster or a distributed system composed of multiple physical servers, or is a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and big data and artificial intelligence platform. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0061] In a possible implementation manner, the computer program related to the embodiments of the present application can be deployed to execute on one computer device, or to execute on multiple computer devices located in one place, or to execute on multiple computer devices distributed in multiple places and interconnected through a communication network, and the multiple computer devices distributed in multiple places and interconnected through the communication network can constitute a blockchain system.

[0062] In a possible implementation manner, the computer device in the embodiments of the present application is a node in the blockchain system, which can store the geometric representation data and texture representation network data of the three-dimensional scene in the blockchain, and then the node or a node corresponding to other devices in the blockchain can query the data stored in the blockchain by accessing the blockchain.

[0063] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application, referring to Figure 1 The implementation environment includes a terminal 101 and a server 102, and the terminal 101 and the server 102 are connected through a wired network or a wireless network.

[0064] The terminal 101 is installed and runs a client 111, which can be a game client, a video sharing client, and the like. When the terminal 101 runs the client 111, a user interface of the client 111 is displayed on a screen of the terminal 101. The terminal 101 is a terminal used by a user 121.

[0065] Optionally, the terminal 101 can be one of a plurality of terminals, and the terminal 101 includes a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a smart household appliance, a vehicle terminal, a flying device, a VR (Virtual Reality) device, an AR (Augmented Reality) device, and the like, but is not limited thereto.

[0066] Those skilled in the art can know that the number of the terminals can be more or less. For example, the terminals can be only one, or the terminals can be 6 or 8 or more. The number of the terminals and the type of the devices are not limited in the embodiments of the present application.

[0067] Figure 1 Only one terminal is shown, but in different embodiments, there are a plurality of other terminals 103 that can access the server 102. Optionally, there is also one or more terminals 103 that are the corresponding terminals of the developers, and a development and editing platform of the client is installed on the terminal 103. The developers can edit and update the client on the terminal 103, and transmit the updated client installation package to the server 102 through a wired or wireless network. The terminal 101 can download the client installation package from the server 102 to implement the update of the client.

[0068] The terminal 101 and the other terminals 103 are connected to the server 102 through a wired network or a wireless network.

[0069] The server 102 includes at least one of a server, a plurality of servers, a cloud computing platform, and a virtualization center. The server 102 is used to provide a background service for the client. Optionally, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or the server 102 and the terminal 101 adopt a distributed computing architecture to perform collaborative computing.

[0070] In the embodiment of the present application, the terminal 101 acquires first scene images corresponding to at least two first viewpoints of a three-dimensional scene, and sends the first scene images to the server 102. The server 102 trains geometric representation data and texture representation network of the three-dimensional scene based on the acquired first scene images by using the method provided in the embodiment of the present application. Then, the terminal 101 selects a second viewpoint from the three-dimensional scene, where the number of the second viewpoints is one or more, and the one or more second viewpoints can be the same as any one or more first viewpoints or different from each first viewpoint. The terminal 101 sends the second viewpoint to the server 102. The server 102 can perform scene rendering on the three-dimensional scene by using the trained geometric representation data and texture representation network, to obtain a scene image corresponding to the second viewpoint, and sends the scene image to the terminal 101. The terminal 101 can display the scene image corresponding to the second viewpoint, so that a user of the terminal 101 can observe the three-dimensional scene from the position of the second viewpoint.

[0071] It should be noted that the above implementation environment is only an example, and the interaction mode of the terminal 101 or the server 102 can also be other modes, for example, the server 102 sends the trained geometric representation data and texture representation network to the terminal 101 for storage, and then the terminal 101 performs scene rendering on the three-dimensional scene by using the trained geometric representation data and texture representation network. Alternatively, the method provided in the embodiment of the present application can also be executed by the terminal 101 or the server 102 alone, or by other computer devices, which is not limited in the embodiment of the present application.

[0072] Figure 2 is a flowchart of a scene rendering method provided in the embodiment of the present application. The embodiment of the present application is executed by a computer device, which is a device such as the terminal 101 or the server 102 as shown in Figure 1 . Referring to Figure 2 , the method comprises the following steps.

[0073] 201. The computer device creates geometric representation data and texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints.

[0074] The three-dimensional scene can be a scene of any object, for example, the object can include buildings, houses, furniture, vehicles, characters, plants or animals, etc. In order to reconstruct the three-dimensional scene, the computer device needs to acquire first scene images corresponding to at least two first viewpoints.

[0075] The viewpoint refers to a point for observing a three-dimensional scene, and the position of the viewpoint can refer to a position of a camera device for shooting the three-dimensional scene. The first viewpoint is an arbitrary viewpoint in a three-dimensional space, and the first scene image is an image of the three-dimensional scene at the first viewpoint, that is, a two-dimensional image obtained by observing the three-dimensional scene at the position of the first viewpoint. One first scene image corresponds to one first viewpoint, and the first scene image is used to describe the three-dimensional scene observed at the position of the first viewpoint. At least two first scene images corresponding to at least two first viewpoints describe the three-dimensional scene observed from different positions.

[0076] In a possible implementation, a camera, a mobile phone, or the like is used to shoot a photo of the three-dimensional scene at the position of at least two viewpoints, so that the first scene images corresponding to the at least two viewpoints are obtained. Alternatively, a camera, a mobile phone, or the like is used to shoot a video of the three-dimensional scene, and the camera device is moved during shooting, so that the camera device shoots the three-dimensional scene from different positions. The video obtained by shooting includes the first scene images corresponding to at least two viewpoints. The first scene images corresponding to the at least two viewpoints are extracted from the video.

[0077] It should be noted that, in the embodiments of the present application, the scene image of the three-dimensional scene obtained is referred to as a first scene image, and the viewpoint corresponding to the first scene image is referred to as a first viewpoint, but this does not constitute a limitation on the scene image and the viewpoint. The first viewpoint is an arbitrary viewpoint in a three-dimensional space, and the number of the first viewpoints is an arbitrary integer greater than 1.

[0078] In the embodiments of the present application, the geometric structure of the three-dimensional scene is described by using geometric representation data, and the color of the three-dimensional scene is described by using a texture representation network. First, the computer device creates the geometric representation data and the texture representation network. At this time, the geometric representation data and the texture representation network are not accurate enough and do not match the three-dimensional scene. Then, based on the first scene images corresponding to the at least two first viewpoints of the three-dimensional scene, the geometric representation data and the texture representation network can be trained, so that the geometric representation data and the texture representation network that match the three-dimensional scene are obtained.

[0079] The geometric representation data includes a plurality of Gaussian points and description parameters of each Gaussian point. The geometric representation data is point cloud data composed of a plurality of Gaussian points with description parameters. Depending on the structural complexity of the three-dimensional scene, the number of Gaussian points included in the geometric representation data is different. Each Gaussian point refers to a point in the three-dimensional scene. Unlike ordinary feature points, a Gaussian point is a three-dimensional Gaussian distribution, which is an ellipsoid in appearance. The Gaussian point can also be referred to as a three-dimensional Gaussian distribution, a normal distribution, a Gaussian distribution function, and the like. Embodiments of the present application do not limit this. The description parameters of the Gaussian point at least include three-dimensional coordinates, which represent the position of the Gaussian point in the three-dimensional scene and can be the three-dimensional coordinates of the center point of the ellipsoid. The three-dimensional coordinates represent the mean of the three-dimensional Gaussian distribution. The three-dimensional coordinates can be in the form of a vector, and the size of the vector is 3x1. For example, the three-dimensional coordinates are [x, y, z]. The three-dimensional space includes an X-axis, a Y-axis, and a Z-axis. x represents the coordinate of the Gaussian point on the X-axis, y represents the coordinate of the Gaussian point on the Y-axis, and z represents the coordinate of the Gaussian point on the Z-axis.

[0080] In addition, the description parameters of the Gaussian point can further include at least one of the following: covariance and transparency. The covariance is a positive semi-definite matrix with a size of 3x3. The covariance of the Gaussian point represents the covariance of the three-dimensional Gaussian distribution. The covariance is in the form of a matrix. The elements on the diagonal line of the covariance matrix represent the variance of the Gaussian point in the X-axis, Y-axis, and Z-axis directions. The elements on the non-diagonal line represent the correlation between different random variables subject to the Gaussian distribution. The transparency represents the transparency of the color of the Gaussian point and can be represented in the form of a percentage or a decimal in the interval [0, 1]. In another embodiment, the transparency can be replaced by opacity, which also represents the transparency of the color of the Gaussian point.

[0081] The texture representation network is used to determine the color value of any Gaussian point based on the three-dimensional coordinates of the Gaussian point. The texture representation network can be a convolutional neural network, an MLP network, or other types of networks, etc. The texture representation network includes network parameters. Based on the network parameters and the three-dimensional coordinates of the Gaussian point, the color value of the Gaussian point can be determined. For example, the texture representation network includes at least one network layer. The network parameters of the texture representation network include weights and bias terms in each network layer, or weights between any two adjacent network layers, etc.

[0082] When creating the geometric representation data and the texture representation network, the computer device initializes the description parameters of each Gaussian point in the geometric representation data and the network parameters in the texture representation network.

[0083] 202、the computer device determines the description parameters and color values of the sample Gaussian points based on the geometric representation data and the texture representation network, and determines a second scene image corresponding to the sample viewpoint based on the description parameters, the color values and the position parameters of the sample viewpoint.

[0084] After the geometric representation data and the texture representation network are created, the computer device iteratively trains the geometric representation data and the texture representation network one or more times, so as to obtain the geometric representation data and the texture representation network capable of accurately describing the three-dimensional scene. The steps 202 and 203 of the embodiments of the present application illustrate the processing process of one training round.

[0085] In one training round, the computer device determines the sample Gaussian points and the sample viewpoint, the sample Gaussian points include at least one Gaussian point in the geometric representation data, and the sample viewpoint includes at least one first viewpoint. The description parameters and color values of the sample Gaussian points can be determined based on the geometric representation data and the texture representation network, and then the second scene image corresponding to the sample viewpoint can be determined based on the description parameters, the color values of the sample Gaussian points and the position parameters of the sample viewpoint.

[0086] Exemplarily, the color values of the Gaussian points include RGB (Red Green Blue) values, or the color values include RGB values, material properties, shadow parameters and any other parameters affecting the color of the pixel points. The color values of the Gaussian points can be set according to requirements, which are not limited in the embodiments of the present application.

[0087] Exemplarily, each Gaussian point in the geometric representation data is taken as the sample Gaussian point in each training, and any first viewpoint is taken as the sample viewpoint in each training. Alternatively, in one training round, the sample Gaussian points can include part of the Gaussian points in the geometric representation data, so as to reduce the number of Gaussian points and reduce the processing amount.

[0088] In a possible implementation, different Gaussian points can be taken as the sample Gaussian points in different training rounds, so that all the Gaussian points can be covered after multiple training rounds. The processing process is the same for different sample Gaussian points, which will not be described herein again.

[0089] 203、the computer device trains the geometric representation data and the texture representation network based on the error between the first scene image and the second scene image corresponding to the sample viewpoint.

[0090] The first scene image is a real scene image obtained by shooting a three-dimensional scene, and the second scene image is a scene image predicted by the current geometry representation data and the texture representation network. Since the current geometry representation data and the texture representation network have not been trained completely, the accuracy is insufficient, which causes an error between the second scene image and the first scene image. Then, the geometry representation data and the texture representation network are trained based on the error, so as to improve the accuracy of the geometry representation data and the texture representation network.

[0091] After the geometry representation data and the texture representation network are trained once or multiple times, the trained geometry representation data and the texture representation network can accurately describe the three-dimensional scene, and the accuracy meets the requirement. Subsequently, the three-dimensional scene can be scene rendered by using the trained geometry representation data and the texture representation network.

[0092] 204. The computer device scene renders the three-dimensional scene by using the trained geometry representation data and the texture representation network.

[0093] The processing flow of the embodiment of the present application is shown in Figure 3 The processing flow includes three stages. The first stage is an initialization stage, in which the geometry representation data and the texture representation network are initialized. The second stage is a forward process, in which the scene is rendered by using the geometry representation data and the texture representation network based on the position parameters of the given sample viewpoint and the description parameters of the sample Gaussian point, to obtain the second scene image corresponding to the sample viewpoint. The third stage is a reverse process, in which the geometry representation data and the texture representation network are trained based on the error between the second scene image obtained by rendering and the real first scene image, so as to obtain the trained geometry representation data and the texture representation network. After the training is completed, the scene can be rendered by using the trained geometry representation data and the texture representation network.

[0094] The method provided by the embodiment of the present application trains the geometry representation data and the texture representation network based on the first scene corresponding to at least two first viewpoints, improves the accuracy of the geometry representation data and the texture representation network, so that the trained geometry representation data and the texture representation network are used together to describe the three-dimensional scene. In scene rendering, the volume rendering process in the NeRF technology does not need to be performed for the Gaussian point, the processing amount is reduced, the time consumption is saved, the processing efficiency is improved, and the trained texture representation network can ensure the texture continuity of the three-dimensional scene, so as to improve the quality of the rendered scene image. Therefore, the geometry representation data and the texture representation network are used together in the embodiment of the present application, which meets the requirements of processing efficiency and scene image quality.

[0095] On the basis of the above embodiment, Figure 4is a flowchart of another scene rendering method provided by the embodiment of the present application. The embodiment of the present application is executed by a computer device, which is a device such as terminal 101 or server 102 shown in Figure 1

[0096] Referring to Figure 4 , the method comprises:

[0097] 401. The computer device creates geometric representation data and texture representation network of the three-dimensional scene based on the first scene images corresponding to the at least two first viewpoints.

[0098] The process of step 401 is the same as that of step 201 described above, and repeated content is not described here.

[0099] In a possible implementation, step 401 comprises:

[0100] 4011. The computer device creates first geometric representation data of the three-dimensional scene based on the at least two first scene images, the first geometric representation data comprising a plurality of feature points in the three-dimensional scene and three-dimensional coordinates of each feature point in the three-dimensional scene.

[0101] The first scene images corresponding to the at least two first viewpoints show the three-dimensional scene under different viewpoints, and therefore the first geometric representation data of the three-dimensional scene can be reconstructed based on the at least two first scene images, the first geometric representation data being used to describe the geometric structure of the three-dimensional scene.

[0102] Exemplarily, the feature points are extracted from the at least two first scene images, and two-dimensional coordinates of each feature point in the first scene image to which the feature point belongs are determined, wherein the feature points can be points satisfying extraction conditions, such as points located at the edges of lines, points with color values satisfying preset color value conditions, or points with gradients satisfying preset gradient conditions, etc. The feature points in different first scene images are matched, thereby determining the same feature points in different first scene images, and three-dimensional coordinates of each feature point in the three-dimensional scene are determined based on the two-dimensional coordinates of the mutually matched feature points in the first scene images in the at least two first scene images, thereby obtaining the first geometric representation data.

[0103] In a possible implementation, the SfM (Structure from Motion) method is used to estimate the first geometric representation data from the at least two first scene images. Exemplarily, the computer device directly calls the COLMAP (a tool for three-dimensional reconstruction) library to execute the step of estimating the first geometric representation data from the at least two first scene images by using the SfM method. In addition, the computer device can also determine the first geometric representation data by using other algorithms, which are not limited by the embodiment of the present application.

[0104] ​4012、convert each feature point in the first geometric representation data into a Gaussian point to obtain second geometric representation data, the second geometric representation data comprising a plurality of Gaussian points and a description parameter of each Gaussian point.

[0105] The difference between the feature point and the Gaussian point is that the feature point is a point, and the Gaussian point is a Gaussian distribution. The computer device creates a description parameter for each Gaussian point, and the description parameter comprises a three-dimensional coordinate, and can further comprise a covariance, a transparency or an opacity, etc. The description parameter of the Gaussian point can be randomly determined by the computer device, and the description parameter is trained in a subsequent training process to improve the accuracy of the description parameter.

[0106] 4013、determine a preset network as the texture representation network, and initialize network parameters in the texture representation network.

[0107] The computer device can set a preset network, such as a convolutional neural network or an MLP network, etc. In the initialization stage, the preset network is determined as the texture representation network, and the network parameters in the texture representation network are initialized.

[0108] In a possible implementation manner, the computer device initializes the network parameters in the texture representation network by using an xavier initialization algorithm. In addition, the computer device can also initialize the network parameters in the texture representation network by using other algorithms, which are not limited in the embodiments of the present application.

[0109] The embodiments of the present application realize the initialization of the geometric representation data and the texture representation data by performing steps 4011-4013, provide a good foundation for the subsequent training process, and further ensure that the subsequent training process is stably performed, avoiding problems such as gradient disappearance or gradient explosion.

[0110] 402、the computer device obtains the description parameter of each Gaussian point from the geometric representation data, and determines the color value of each Gaussian point by the texture representation network based on the three-dimensional coordinate of each Gaussian point.

[0111] After the geometric representation data and the texture representation network are created, the computer device iteratively trains the geometric representation data and the texture representation network one or more times, and then obtains the geometric representation data and the texture representation network which can accurately describe the three-dimensional scene. The steps 402 and 403 of the embodiments of the present application illustrate the processing process of one training round.

[0112] In one training round, the computer device determines the description parameters and color values of the sample Gaussian points through the geometry representation data and the texture representation network. The description parameters of the Gaussian points include three-dimensional coordinates, and the input of the texture representation network is the three-dimensional coordinates of the Gaussian points, and the output is the color values of the Gaussian points. Therefore, after obtaining the three-dimensional coordinates of the sample Gaussian points from the geometry representation data, the color values of the sample Gaussian points can be determined through the texture representation network.

[0113] In some embodiments, the texture representation network is a tensor radiation field, and the texture representation network introduces a spatial decomposition technique. A texture field in a three-dimensional space can be represented by a 4D (4 Dimension) tensor, and the entire three-dimensional space is divided into a 3D grid. Assuming that the resolutions of the coordinate axes (X, Y, Z) along three perpendicular directions are (I, J, K) respectively, and the texture dimension is represented as P, the 4D texture field can be represented as g c ∈R I×J×KXP With the aid of the spatial decomposition technique, the 4D tensor can be decomposed into a combination of low-dimensional tensors. Among them, I, J, K can be set according to the complexity of the three-dimensional scene, for example, I, J, K are all 128.

[0114] Therefore, determining the description parameters and color values of the sample Gaussian points through the geometry representation data and the texture representation network includes: obtaining the description parameters of the sample Gaussian points from the geometry representation data, spatially decomposing the three-dimensional coordinates of the sample Gaussian points to obtain components of the three-dimensional coordinates, and determining the texture feature values of the sample Gaussian points based on the components of the three-dimensional coordinates; and mapping the texture feature values of the sample Gaussian points and the position parameters of the sample viewpoint based on the network parameters in the texture representation network to obtain the color values of the sample Gaussian points.

[0115] Among them, the components of the sample Gaussian points can be determined by using the spatial decomposition technique, and the computer device can determine the basis corresponding to each component. Each component is weighted and summed according to the basis corresponding to each component to obtain the texture feature value. Then, the color value of the sample Gaussian point can be obtained by mapping through the network parameters in the texture representation network. Among them, the above mapping can be a linear mapping method or a nonlinear mapping method.

[0116] In one possible implementation, the three-dimensional space is decomposed into three two-dimensional planes and three one-dimensional vectors by using a VM (vector-matrix) decomposition technique, which can be represented as:

[0117]

[0118] Among them, g c represents the texture feature value, r represents the serial number of the tensor, and the value range of r is [1, R c ], R crepresents an hyper-parameter, which can be set to 8-96, for example, the hyper-parameter is 16, which will affect the number of bases of the texture feature, represents a projection vector of the Gaussian point along the X-axis in the three-dimensional space, represents a projection matrix of the Gaussian point along the YZ plane in the three-dimensional space, represents a projection vector of the Gaussian point along the Y-axis in the three-dimensional space, represents a projection matrix of the Gaussian point along the XZ plane in the three-dimensional space, represents a projection vector of the Gaussian point along the Z-axis in the three-dimensional space, represents a projection matrix of the Gaussian point along the XY plane in the three-dimensional space, b 3r-2 , b 3r-1 and b 3r represents a base of the texture feature, the value of which can be set by the computer device, represents an outer product symbol. After the spatial texture field is established, for the three-dimensional coordinates of each Gaussian point in the three-dimensional space, the projection vector of the Gaussian point along the X-axis, the Y-axis and the Z-axis, and the projection matrix of the Gaussian point along the YZ plane, the XZ plane and the XY plane can be calculated by projection, so as to calculate the texture feature value of the Gaussian point.

[0119] In the above VM decomposition technology, the three-dimensional space is decomposed into projection vectors and projection matrices. In another possible implementation manner, the three-dimensional space can be entirely decomposed into projection vectors, and no longer be decomposed into projection matrices. The processing process is the same as the above VM decomposition technology, and will not be described here.

[0120] In the embodiments of the present application, the sample Gaussian point includes each Gaussian point, that is, the color value of each Gaussian point is considered in the process of determining the second scene image. Therefore, the color value of each Gaussian point is obtained in step 402.

[0121] 403. The computer device projects each Gaussian point into the image plane corresponding to the sample viewpoint based on the description parameter and the color value of each Gaussian point and the position parameter of the sample viewpoint, to obtain the second scene image corresponding to the sample viewpoint.

[0122] The position parameter of the sample viewpoint is used to describe the position of the sample viewpoint in the three-dimensional space. The position parameter can include at least one of a pitch angle, a yaw angle, and a roll angle. In one possible implementation, the position parameter includes the three angles: the pitch angle, the yaw angle, and the roll angle. The position parameter can be represented as a rotation matrix of the sample viewpoint relative to a world coordinate system, and the rotation matrix is a 3x3 orthogonal matrix.

[0123] The position parameter can also include a distance between the sample viewpoint and the origin of the three-dimensional space, or the position parameter can not include the distance, and the distance can be set to a fixed value by the computer device. After the pitch angle, the yaw angle, and the distance between the viewpoint and the origin of the three-dimensional space are determined, the position of the sample viewpoint in the three-dimensional space can be determined.

[0124] The image plane corresponding to the sample viewpoint is an image plane used to render a scene image corresponding to the sample viewpoint, and the observation direction of the sample viewpoint is the direction of the main axis of the camera device of the sample viewpoint, which can be determined by the rotation matrix of the sample viewpoint. The image plane corresponding to the sample viewpoint is an imaging plane perpendicular to the main axis of the camera device. The size of the image plane can be a preset fixed size. Each Gaussian point can be projected into the image plane, wherein each Gaussian point can be projected into one or more pixel points in the image plane, and each pixel point can correspond to one or more Gaussian points. After the description parameters, color values, and the image plane of each Gaussian point are determined, each Gaussian point with the description parameters such as color values and transparency can be projected into the image plane, so that the color values of each pixel point in the image plane are determined, and the color values of each pixel point in the image plane constitute the second scene image.

[0125] In some possible implementations, the description parameters of the Gaussian points further include transparency, and step 403 includes: for each pixel point in the image plane, determining an associated Gaussian point based on the position parameter of the pixel point and the position parameter of the sample viewpoint, the associated Gaussian point being a Gaussian point projected to the position of the pixel point according to the observation direction of the sample viewpoint, mixing the color values of each associated Gaussian point based on the transparency of each associated Gaussian point to obtain the color value of the pixel point, and determining the second scene image based on the color values of each pixel point in the image plane.

[0126] The position parameter of the pixel point can be a three-dimensional coordinate of the pixel point in a three-dimensional space, or a two-dimensional coordinate of the pixel point in an image plane, and the like. Based on the position parameter of the pixel point, the three-dimensional coordinate of the Gaussian point, the position parameter of the sample viewpoint, and the internal parameter of the shooting device of the sample viewpoint, it can be determined which pixel point in the image plane is projected by each Gaussian point. The internal parameter of the shooting device can be obtained from a shooting device that shoots the scene image corresponding to the sample viewpoint. Moreover, considering that different Gaussian points in the three-dimensional space can be projected to the same pixel point in the image plane, that is, one pixel point can correspond to multiple associated Gaussian points, in principle, when the sample viewpoint observes the three-dimensional scene, the color value of the pixel point should be determined by the color values and the transparency of the multiple associated Gaussian points. Therefore, the computer device mixes the color values of each associated Gaussian point based on the transparency of each associated Gaussian point to obtain the color value of the pixel point.

[0127] In a possible implementation, considering that for the multiple associated Gaussian points corresponding to the same pixel point, the distances between different associated Gaussian points and the pixel point are different, the influence degrees on the color value of the pixel point are also different. Therefore, in order to more realistically present the observed three-dimensional scene in the scene image, the computer device can determine the depth value of each associated Gaussian point based on the distance between each associated Gaussian point and the image plane (that is, the distance between each associated Gaussian point and the imaging camera), mix the color values of each associated Gaussian point based on the depth value and the transparency of each associated Gaussian point to obtain the color value of the pixel point. For example, the associated Gaussian points are sorted according to the depth values of the associated Gaussian points, and the transparency of each associated Gaussian point is taken as the weight of the color value of each associated Gaussian point, so that the color values of each associated Gaussian point are weighted and mixed according to the order from near to far in combination with the weight of each associated Gaussian point. Alternatively, based on the depth value of each associated Gaussian point, the associated Gaussian points greater than a preset depth value are removed, so that the influence of the color values of the associated Gaussian points far away from the sample viewpoint is ignored.

[0128] In a possible implementation, a rendering method in the 3DGS technology is used to project the Gaussian points to the image plane through affine transformation, and then the color values of the Gaussian points are weighted and fused according to the depth values and the transparency of the Gaussian points, so as to obtain the second scene image.

[0129] 404、The computer device trains the geometric representation data and the texture representation network based on the error between the first scene image and the second scene image corresponding to the sample viewpoint.

[0130] The error can include a difference between color values of pixels at the same position in the first scene image and the second scene image, or translation information or rotation information between the first scene image and the second scene image, and embodiments of the present application do not limit the specific content of the error.

[0131] In some embodiments, step 404 includes training the description parameters of at least one Gaussian point in the geometric representation data and the network parameters in the texture representation network based on the error. The training target is that, for the same sample viewpoint and sample Gaussian point, the error between the second scene image and the first scene image determined by the trained geometric representation data and the texture representation network is smaller than the error before training, so as to reduce the error between the scene image rendered by the geometric representation data and the texture representation network and the real scene image. In this way, after one or more training, error minimization is achieved, and accurate geometric representation data and texture representation network are obtained.

[0132] The description parameters of the Gaussian point include three-dimensional coordinates, covariance, transparency, or opacity, etc. In order to create dense Gaussian points that can accurately express a three-dimensional scene, each of the above description parameters can be optimized in the training process, and the number of Gaussian points can also be dynamically adjusted in an adaptive manner in the training process. The computer device can determine the conditions that the Gaussian points satisfy and the adjustment manner of the Gaussian points under the condition.

[0133] In a possible implementation manner, the computer device uses a gradient descent algorithm for training, and therefore calculates the gradient of the error, adjusts the description parameters of at least one Gaussian point and the network parameters of the texture representation network based on the gradient. When the gradient of the error corresponding to the Gaussian point is greater than a preset gradient, it indicates that the error caused by the Gaussian point is large, and the Gaussian point is filtered. Alternatively, the variance of the Gaussian point is determined, and when the variance of the Gaussian point is greater than a first preset variance, a segmentation operation is performed on the Gaussian point, and when the variance of the Gaussian point is less than a second preset variance, a cloning operation is performed on the Gaussian point. The second preset variance is less than the first preset variance, and the values of the preset gradient, the first preset variance, and the second preset variance can be determined by the computer device.

[0134] The segmentation operation on the Gaussian point means that the three-dimensional coordinates and the transparency or opacity of the Gaussian point are kept unchanged, the original covariance of the Gaussian point is updated to a quotient of the original covariance and a preset transformation parameter, and the value of the preset transformation parameter is greater than 1, so as to reduce the covariance of the Gaussian point. The value of the preset transformation parameter is determined by the computer device, and can be a numerical value such as 1.6 or 2. The cloning of the Gaussian point means that a Gaussian point with the same description parameters as the original Gaussian point is added.

[0135] 405、The computer device determines the description parameters and color values of each Gaussian point through the trained geometry representation data and the trained texture representation network, and determines the scene image corresponding to the second viewpoint based on the description parameters, the color values and the position parameters of the second viewpoint.

[0136] Step 405 is the same as step 202 and steps 402-403 described above, and the difference is that the geometry representation data and the texture representation network used have been trained, and in addition, the internal parameter corresponding to the second viewpoint needs to be set by the user or the computer device, which is used to represent the internal parameter of the shooting device for shooting the three-dimensional scene at the second viewpoint. As for other same parts, the embodiments of the present application will not be described again.

[0137] In the embodiments of the present application, the second viewpoint can be the same as the first viewpoint, or different from the first viewpoint, and the number of the second viewpoints can be one or more. When the computer device obtains the scene image corresponding to the second viewpoint, it can display the scene image corresponding to the second viewpoint for the user to watch, or send the scene image corresponding to the second viewpoint to the device providing the scene image corresponding to the first viewpoint, or perform other processing operations on the scene image corresponding to the second viewpoint.

[0138] In one possible implementation, the computer device combines the first scene images corresponding to at least two first viewpoints and the scene image corresponding to the second viewpoint to obtain a scene image sequence of the three-dimensional scene, and displays the scene image sequence. During the display process, the user can perform a rotation operation on the three-dimensional scene, and the computer device determines the viewpoint corresponding to the rotation operation based on the rotation operation performed by the user, so as to extract the scene image corresponding to the viewpoint from the scene image sequence and display the scene image, thereby realizing the effect that the user observes the three-dimensional scene from different viewpoints, and the viewpoint is determined by the operation of the user, thereby improving the flexibility and meeting the personalized needs of the user.

[0139] In another possible implementation, instead of considering at least two first viewpoints, a plurality of second viewpoints are selected from the three-dimensional scene, the scene images corresponding to the plurality of second viewpoints are combined to obtain a scene image sequence of the three-dimensional scene, and the scene image sequence is displayed. During the display process, the user can perform a rotation operation on the three-dimensional scene, and the computer device determines the viewpoint corresponding to the rotation operation based on the rotation operation performed by the user, so as to extract the scene image corresponding to the viewpoint from the scene image sequence and display the scene image, thereby realizing the effect that the user observes the three-dimensional scene from different viewpoints, and the viewpoint is determined by the operation of the user, thereby improving the flexibility and meeting the personalized needs of the user.

[0140] The method provided in the embodiments of the present application trains the geometry representation data and the texture representation network based on the first scene corresponding to at least two first viewpoints, improves the accuracy of the geometry representation data and the texture representation network, so that the trained geometry representation data and texture representation network are used together to describe a three-dimensional scene, and when scene rendering is performed, the volume rendering process in the NeRF technology does not need to be performed for the Gaussian points, the processing amount is reduced, time consumption is saved, the processing efficiency is improved, and the quality of the rendered scene image is improved. Moreover, the texture continuity of the three-dimensional scene can be ensured by using the trained texture representation network to describe the color of the three-dimensional scene, so that the quality of the rendered scene image is improved, and the number of network layers in the texture representation network does not need to be limited, that is, the number of network layers in the texture representation network can be appropriately reduced, so that the processing amount of the texture representation network is reduced. Therefore, the embodiments of the present application combine the use of the geometry representation data and the texture representation network, meet the demand for processing efficiency and the quality of the scene image. Moreover, due to the reduction of the processing amount, the computing resources and storage resources required by the method provided in the embodiments of the present application are also low, which is suitable for being applied to a computer device with limited resources, and the application range is expanded.

[0141] Moreover, in the embodiments of the present application, the scene rendering is performed in the forward projection mode, the second scene image corresponding to the sample viewpoint is obtained, the volume rendering process in the NeRF technology is avoided, and the processing amount in the rendering process is greatly reduced, so that the processing speed is improved.

[0142] Moreover, considering that one pixel point in the image plane can correspond to multiple associated Gaussian points, in principle, when the three-dimensional scene is observed from the sample viewpoint, the color value of the pixel point should be determined by the color values and the transparency of the multiple associated Gaussian points. Therefore, the computer device mixes the color values of each associated Gaussian point based on the transparency of each associated Gaussian point to obtain the color value of the pixel point, so that the color in the rendered scene image conforms to the actual situation, and the accuracy is improved.

[0143] Moreover, the embodiments of the present application provide a scene reconstruction method combining the 3D Gaussian and spatial decomposition technology, the geometry structure of the three-dimensional scene is fitted by a certain number of Gaussian points, and the color value of each Gaussian point is determined by using the low-dimensional space representation after spatial decomposition, so that the texture of the three-dimensional scene is reconstructed, and fast and high-quality scene rendering is realized.

[0144] The related NeRF technology adopts an MLP network, and in order to ensure accuracy, the MLP network is composed of multiple fully connected layers, and the structure is relatively complex. Whether it is the process of training the MLP network or the process of rendering the scene through the trained MLP network, the processing amount is very large. Moreover, volume rendering needs to be performed in the NeRF technology, and in the volume rendering process, each pixel point needs to sample a light ray and calculate the color value and density value of each sampling point, which also leads to a doubling of the processing amount. Therefore, the NeRF technology is very time-consuming and has very low processing efficiency, and is not suitable for application in application scenarios with high time requirements, so the application scenarios are very limited.

[0145] The related 3DGS technology adopts a forward projection method in the rendering stage, avoids the volume rendering process in the NeRF technology, greatly reduces the processing amount in the rendering process, and greatly improves the processing speed. However, since the 3DGS technology uses discrete Gaussian points to represent a three-dimensional scene, the texture continuity of the three-dimensional scene cannot be guaranteed when the scene is reconstructed, which leads to the fact that the rendered scene image is prone to mutation and defects when the viewing angle is rotated, and cannot meet the demand for high-quality rendered images.

[0146] In the embodiments of the present application, the geometric representation data is used to describe the structure of the three-dimensional scene, and when the scene is rendered through the geometric representation data, the volume rendering process in the NeRF technology does not need to be performed, and only the Gaussian points need to be projected onto the image plane corresponding to the viewpoint, which improves the processing speed. Moreover, since the geometric representation data includes discrete Gaussian points, the texture continuity of the three-dimensional scene cannot be guaranteed when the scene is rendered through the geometric representation data, which leads to the fact that the rendered scene image is prone to mutation and defects when the viewing angle is rotated, and cannot meet the demand for high-quality rendered images. The color of the three-dimensional scene is described by using the texture representation network, and the texture continuity of the three-dimensional scene can be guaranteed after the texture representation network is trained, so that the rendered scene image is prevented from appearing mutation and defects when the viewing angle is rotated, and the quality of the scene image is improved. The geometric representation data and the texture representation network are combined in the embodiments of the present application, which meets the demand for processing efficiency and the quality of the scene image, is suitable for being applied to a scene with high time requirements or a computer device with limited resources, expands the application range, and improves the flexibility.

[0147] On the basis of the above-mentioned various embodiments, the application scenarios of the embodiments of the present application will be described as follows.

[0148] Taking the application scenario of making a 3D photo as an example, the 3D photo is a simulation of the principle of human eyes looking at the world, and uses the visual difference between two human eyes and the optical refraction principle to make people directly see a three-dimensional image in an image plane. The 3D photo can reflect the two-dimensional relationship of the objects up and down and left and right in the photo, and can make the eyes visually see the three-dimensional relationship of the objects up and down, left and right, and front and back.

[0149] Figure 5 is a schematic diagram of a processing flow for making a 3D photo provided by an embodiment of the present application, referring to Figure 5 The processing flow includes:

[0150] 501. The terminal takes pictures of the target object from the positions of at least two first viewpoints to obtain first scene images corresponding to the at least two first viewpoints.

[0151] The target object is a three-dimensional object.

[0152] 502. The terminal sends the first scene images corresponding to the at least two first viewpoints to the server.

[0153] 503. The server receives the first scene images corresponding to the at least two first viewpoints, and creates geometric representation data and texture representation network of the target object based on the first scene images corresponding to the at least two first viewpoints.

[0154] 504. The server determines the description parameters and color values of sample Gaussian points through the geometric representation data and the texture representation network, and determines a second scene image corresponding to a sample viewpoint based on the description parameters, the color values and the position parameters of the sample viewpoint.

[0155] 505. The server trains the geometric representation data and the texture representation network based on the error between the first scene image and the second scene image corresponding to the sample viewpoint.

[0156] 506. The server determines at least two second viewpoints.

[0157] 507. The server determines the description parameters and color values of each Gaussian point through the trained geometric representation data and the trained texture representation network, and determines a scene image corresponding to each second viewpoint based on the description parameters, the color values and the position parameters of each second viewpoint.

[0158] 508. The server makes a 3D photo of the target object based on the scene images corresponding to the at least two second viewpoints.

[0159] 509. The server sends the 3D photo to the terminal.

[0160] 510. The terminal receives and displays the 3D photo of the target object. When viewing the 3D photo, the user of the terminal can feel the spatial structural relationship of the target object through the eyes, achieving the effect of viewing a three-dimensional stereoscopic image of the target object.

[0161] Taking the application scenario of making a movie resource containing "bullet time" as an example, "bullet time" is mainly used to simulate variable speed special effects, such as slow motion and time freeze effects. The realization of "bullet time" usually relies on multiple camera devices shooting at the same time, and then creating a time freeze or slow motion visual effect through computer synthesis.

[0162] Figure 6 is a schematic diagram of a processing flow for making a movie resource containing "bullet time" provided by an embodiment of the present application, referring to Figure 6 The processing flow includes:

[0163] 601. During the process of the person performing the movie making an action, the person is photographed using camera devices at positions of at least two first viewpoints, obtaining first scene images corresponding to the at least two first viewpoints.

[0164] 602. The computer device obtains the first scene images corresponding to the at least two first viewpoints.

[0165] 603. The computer device creates geometric representation data and texture representation network of the person based on the first scene images corresponding to the at least two first viewpoints.

[0166] 604. The computer device determines the description parameters and color values of the sample Gaussian points through the geometric representation data and the texture representation network, and determines the second scene images corresponding to the sample viewpoints based on the description parameters, the color values and the position parameters of the sample viewpoints.

[0167] 605. The computer device trains the geometric representation data and the texture representation network based on the error between the first scene images and the second scene images corresponding to the sample viewpoints.

[0168] 606. The computer device determines at least two second viewpoints.

[0169] 607. The computer device determines the description parameters and color values of each Gaussian point through the trained geometric representation data and the trained texture representation network, and determines the scene images corresponding to each second viewpoint based on the description parameters, the color values and the position parameters of each second viewpoint.

[0170] 608. The computer device adds the first scene images corresponding to each first viewpoint and the second scene images corresponding to each second viewpoint in the movie resource being made, and stores the movie resource in the computer device.

[0171] Since the scene image corresponding to each viewpoint of the character is added in the movie resource, when the movie resource is played, a time pause effect and a rotating effect of the character can be presented, so that the user can experience the feeling of slowly watching the character from different viewpoints when watching the movie resource.

[0172] It should be noted that the application embodiment only takes the above application scenario as an example, and the application embodiment can also be applied to a panoramic roaming scenario. Panoramic roaming refers to a technology of using panoramic technology to display a three-dimensional scene in all directions in a brand-new perspective and an intuitive feeling of being on the scene. A user can control the direction of observing the panorama by touch or a mouse and a keyboard, and can move left, right, near and far, so as to obtain the feeling of observing the three-dimensional scene in a real environment. The scene rendering method provided by the application embodiment can also be applied to other scenarios, and the application embodiment does not limit the application scenario.

[0173] Figure 7 FIG. 1 is a structural schematic diagram of a scene rendering device provided by the application embodiment. Referring to FIG. 1, Figure 7 The device includes:

[0174] The creating module 701 is configured to create geometric representation data and texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints. The first scene image is an image of the three-dimensional scene at the first viewpoint. The geometric representation data includes a plurality of Gaussian points and description parameters of each Gaussian point. The description parameters at least include three-dimensional coordinates. The texture representation network is used to determine a color value of any Gaussian point based on the three-dimensional coordinates of the Gaussian point.

[0175] The training module 702 is configured to determine the description parameters and the color value of a sample Gaussian point through the geometric representation data and the texture representation network, and determine a second scene image corresponding to a sample viewpoint based on the description parameters, the color value and position parameters of the sample viewpoint. The sample Gaussian point includes at least one Gaussian point in the geometric representation data. The sample viewpoint includes at least one first viewpoint.

[0176] The training module 702 is further configured to train the geometric representation data and the texture representation network based on an error between the first scene image and the second scene image corresponding to the sample viewpoint.

[0177] The scene rendering module 703 is configured to perform scene rendering on the three-dimensional scene through the trained geometric representation data and the texture representation network.

[0178] In a possible implementation manner, referring to FIG. 1, Figure 8 The creating module 701 includes:

[0179] The geometry creating unit 711 is configured to create first geometry representation data of the three-dimensional scene based on the at least two first scene images, the first geometry representation data comprising a plurality of feature points in the three-dimensional scene and three-dimensional coordinates of each feature point in the three-dimensional scene.

[0180] The conversion unit 721 is configured to convert each feature point in the first geometry representation data into a Gaussian point to obtain second geometry representation data, the second geometry representation data comprising a plurality of Gaussian points and description parameters of each Gaussian point.

[0181] The network creating unit 731 is configured to determine a preset network as the texture representation network and initialize network parameters in the texture representation network.

[0182] In a possible implementation, the sample Gaussian point comprises each Gaussian point in the geometry representation data, refer to Figure 8 The training module 702 comprises:

[0183] The parameter obtaining unit 712 is configured to obtain the description parameters of each Gaussian point from the geometry representation data.

[0184] The color value determining unit 722 is configured to respectively determine the color value of each Gaussian point by the texture representation network based on the three-dimensional coordinates of each Gaussian point.

[0185] The projection unit 732 is configured to project each Gaussian point into an image plane corresponding to the sample viewpoint based on the description parameters and the color value of each Gaussian point and the position parameters of the sample viewpoint to obtain a second scene image corresponding to the sample viewpoint.

[0186] In a possible implementation, the description parameters further comprise transparency.

[0187] The projection unit 732 is configured to, for each pixel point in the image plane, determine an associated Gaussian point based on the position parameters of the pixel point and the position parameters of the sample viewpoint, the associated Gaussian point being a Gaussian point projected to a position where the pixel point is located in the observation direction of the sample viewpoint; mix the color value of each associated Gaussian point based on the transparency of each associated Gaussian point to obtain the color value of the pixel point; and determine the second scene image based on the color value of each pixel point in the image plane.

[0188] In a possible implementation, the texture representation network is a tensor radiance field, refer to Figure 8 The training module 702 comprises:

[0189] The parameter obtaining unit 712 is configured to obtain the description parameters of the sample Gaussian point from the geometry representation data.

[0190] The spatial decomposition unit 742 is configured to perform spatial decomposition on the three-dimensional coordinates of the sample Gaussian point to obtain components of the three-dimensional coordinates, and determine a texture feature value of the sample Gaussian point based on the components of the three-dimensional coordinates.

[0191] The mapping unit 752 is configured to map the texture feature value of the sample Gaussian point and the position parameter of the sample viewpoint based on network parameters in the texture representation network to obtain a color value of the sample Gaussian point.

[0192] In a possible implementation, the training module 702 is configured to train, based on the error, the description parameter of at least one Gaussian point in the geometry representation data and the network parameters in the texture representation network.

[0193] In a possible implementation, the scene rendering module 703 is configured to determine the description parameter and the color value of each Gaussian point by using the trained geometry representation data and the trained texture representation network, and determine a scene image corresponding to the second viewpoint based on the description parameter, the color value, and the position parameter of the second viewpoint.

[0194] The scheme provided by the embodiments of the present application trains the geometry representation data and the texture representation network based on the first scene corresponding to at least two first viewpoints, improves the accuracy of the geometry representation data and the texture representation network, so that the trained geometry representation data and the trained texture representation network are used together to describe the three-dimensional scene, the volume rendering process in the NeRF technology does not need to be performed for the Gaussian point when the scene is rendered, the processing amount is reduced, the time consumption is saved, the processing efficiency is improved, and the texture continuity of the three-dimensional scene is ensured by using the texture representation network to describe the color of the three-dimensional scene, so that the quality of the rendered scene image is improved. Therefore, the embodiments of the present application combine the geometry representation data and the texture representation network, and meet the demand for processing efficiency and the quality of the scene image.

[0195] It should be noted that the scene rendering apparatus provided in the above embodiments is only used as an example for the division of the above functional modules, and in actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the above described functions. In addition, the scene rendering apparatus and the scene rendering method provided in the above embodiments belong to the same concept, and the specific implementation process is described in detail in the method embodiments, which will not be repeated here.

[0196] The embodiments of the present application also provide a computer device, which includes a processor and a memory, and the memory stores at least one computer program, which is loaded and executed by the processor to implement the operations performed in the scene rendering method of the above embodiments.

[0197] Optionally, the computer device is provided as a terminal. Figure 9 A structure diagram of a terminal 900 provided by an example embodiment of the present application is shown.

[0198] The terminal 900 includes a processor 901 and a memory 902.

[0199] The processor 901 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 901 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), a PLA (Programmable Logic Array). The processor 901 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 901 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 901 can further include an AI (Artificial Intelligence) processor for processing machine learning related computing operations.

[0200] The memory 902 can include one or more computer-readable storage media, which can be non-transitory. The memory 902 can also include a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one computer program for being executed by the processor 901 to implement a scene rendering method provided by a method embodiment of the present application.

[0201] In some embodiments, the terminal 900 can further optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 903 through a bus, a signal line, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, and a power supply 907.

[0202] The peripheral interface 903 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902 and the peripheral interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902 and the peripheral interface 903 can be implemented on a separate chip or circuit board, and the present embodiments are not limited in this regard.

[0203] The radio frequency circuit 904 is configured to receive and send RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals to electromagnetic signals for transmission, or converts electromagnetic signals received to electrical signals. Optionally, the radio frequency circuit 904 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 904 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to, a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 can also include NFC (Near Field Communication) related circuitry, and the present application is not limited in this regard.

[0204] The display screen 905 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 is further configured to capture touch signals on or above the surface of the display screen 905. The touch signals can be input to the processor 901 as control signals for processing. In this case, the display screen 905 can be further configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 905 can be one, disposed on the front panel of the terminal 900; in other embodiments, the display screen 905 can be at least two, respectively disposed on different surfaces of the terminal 900 or in a folding design; in other embodiments, the display screen 905 can be a flexible display screen, disposed on a curved surface or a folding surface of the terminal 900. Even, the display screen 905 can be disposed in an irregular shape other than a rectangle, i.e., a special-shaped screen. The display screen 905 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.

[0205] The camera assembly 906 is configured to capture images or videos. Optionally, the camera assembly 906 includes a front camera and a rear camera. The front camera is disposed on the front panel of the terminal 900, and the rear camera is disposed on the back of the terminal 900. In some embodiments, the rear camera is at least two, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to realize the background blur function by fusing the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function by fusing the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 906 can further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. The dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0206] The power supply 907 is configured to supply power to each component in the terminal 900. The power supply 907 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 907 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0207] Those skilled in the art can understand that the structure shown in the above description is not a limitation on the terminal 900, and the terminal 900 can include more or fewer components than those shown in the figure, or combine certain components, or use a different arrangement of components. Figure 9 Those skilled in the art can understand that the structure shown in the above description is not a limitation on the terminal 900, and the terminal 900 can include more or fewer components than those shown in the figure, or combine certain components, or use a different arrangement of components.

[0208] Optionally, the computer device is provided as a server. Figure 10 is a structural schematic diagram of a server provided by an embodiment of the present application. The server 1000 can be quite different in configuration or accuracy, and can include one or more processors (Central Processing Units, CPUs) 1001 and one or more memories 1002. The memory 1002 stores at least one computer program, which is loaded and executed by the processor 1001 to implement the method provided by each method embodiment described above. Of course, the server can also have a wired or wireless network interface, a keyboard, an input and output interface, and other components for implementing device functions, and will not be described here.

[0209] The embodiment of the present application further provides a computer readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to implement the operations performed by the scene rendering method of the above embodiment.

[0210] The embodiment of the present application further provides a computer program product, which includes a computer program. The computer program is loaded and executed by a processor to implement the operations performed by the scene rendering method of the above embodiment.

[0211] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by a program instructing relevant hardware to complete, and the program can be stored in a computer readable storage medium, such as a read-only memory, a magnetic disk or an optical disk.

[0212] The above is only an optional embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of scene rendering, characterized by, The method comprises: creating geometric representation data and texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints, the first scene images being images of the three-dimensional scene at the first viewpoints, the geometric representation data comprising a plurality of Gaussian points and description parameters of each Gaussian point, the description parameters at least comprising three-dimensional coordinates, and the texture representation network being used to determine color values of any Gaussian point based on three-dimensional coordinates of the Gaussian point; determining description parameters and color values of sample Gaussian points comprising at least one Gaussian point in the geometric representation data by using the geometric representation data and the texture representation network, and determining a second scene image corresponding to a sample viewpoint based on the description parameters, the color values and position parameters of the sample viewpoint; training the geometric representation data and the texture representation network based on errors between the first scene images corresponding to the sample viewpoint and the second scene image; performing scene rendering on the three-dimensional scene by using the trained geometric representation data and the texture representation network.

2. The method of claim 1, wherein, The method comprises: creating first geometric representation data of the three-dimensional scene based on at least two first scene images, the first geometric representation data comprising a plurality of feature points in the three-dimensional scene and three-dimensional coordinates of each feature point in the three-dimensional scene; converting each feature point in the first geometric representation data into a Gaussian point to obtain second geometric representation data, the second geometric representation data comprising a plurality of Gaussian points and description parameters of each Gaussian point; determining a preset network as the texture representation network and initializing network parameters in the texture representation network.

3. The method of claim 1, wherein, The sample Gaussian points comprise each Gaussian point in the geometric representation data, and the method comprises: obtaining description parameters of each Gaussian point from the geometric representation data; determining color values of each Gaussian point by using the texture representation network based on three-dimensional coordinates of each Gaussian point; projecting each Gaussian point into an image plane corresponding to the sample viewpoint based on the description parameters and color values of each Gaussian point and position parameters of the sample viewpoint to obtain the second scene image corresponding to the sample viewpoint.

4. The method of claim 3, wherein, The description parameters further comprise transparency. The method comprises: projecting each Gaussian point into an image plane corresponding to the sample viewpoint based on the description parameters and color values of each Gaussian point and position parameters of the sample viewpoint to obtain the second scene image corresponding to the sample viewpoint. For each pixel point in the image plane, based on the position parameter of the pixel point and the position parameter of the sample viewpoint, a relevant Gaussian point is determined, the relevant Gaussian point being a Gaussian point projected to the position of the pixel point in the observation direction of the sample viewpoint; based on the transparency of each relevant Gaussian point, the color value of each relevant Gaussian point is mixed to obtain the color value of the pixel point; Based on the color value of each pixel point in the image plane, the second scene image is determined.

5. The method of claim 1, wherein, The texture representation network is a tensor radiation field, and the determination of the description parameter and the color value of the sample Gaussian point through the geometry representation data and the texture representation network comprises: The description parameter of the sample Gaussian point is obtained from the geometry representation data; The three-dimensional coordinates of the sample Gaussian point are spatially decomposed to obtain components of the three-dimensional coordinates, and the texture feature value of the sample Gaussian point is determined based on the components of the three-dimensional coordinates; The texture feature value of the sample Gaussian point and the position parameter of the sample viewpoint are mapped based on the network parameter in the texture representation network to obtain the color value of the sample Gaussian point.

6. The method of claim 1, wherein, The training of the geometry representation data and the texture representation network based on the error between the first scene image corresponding to the sample viewpoint and the second scene image comprises: Based on the error, the description parameter of at least one Gaussian point in the geometry representation data and the network parameter in the texture representation network are trained.

7. The method according to any one of claims 1 to 6, characterized in that, The scene rendering of the three-dimensional scene through the trained geometry representation data and the trained texture representation network comprises: The description parameter and the color value of each Gaussian point are determined through the trained geometry representation data and the trained texture representation network, and the scene image corresponding to the second viewpoint is determined based on the description parameter, the color value and the position parameter of the second viewpoint.

8. A scene rendering apparatus, characterized by comprising: The device comprises: A creation module is configured to create geometry representation data and a texture representation network of a three-dimensional scene based on first scene images corresponding to at least two first viewpoints, the first scene image being an image of the three-dimensional scene at the first viewpoint, the geometry representation data comprising a plurality of Gaussian points and a description parameter of each Gaussian point, the description parameter at least comprising three-dimensional coordinates, and the texture representation network being configured to determine a color value of any Gaussian point based on the three-dimensional coordinates of the Gaussian point; A training module is configured to determine a description parameter and a color value of a sample Gaussian point through the geometry representation data and the texture representation network, and to determine a second scene image corresponding to a sample viewpoint based on the description parameter, the color value and the position parameter of the sample viewpoint, the sample Gaussian point comprising at least one Gaussian point in the geometry representation data, and the sample viewpoint comprising at least one first viewpoint; The training module is further configured to train the geometry representation data and the texture representation network based on an error between the first scene image corresponding to the sample viewpoint and the second scene image. a scene rendering module, configured to perform scene rendering on the three-dimensional scene by using the trained geometry representation data and the texture representation network.

9. A computer device, comprising: The computer device comprises a processor and a memory, and the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the scene rendering method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the scene rendering method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by the scene rendering method according to any one of claims 1 to 7.