Cloud rendering method and device based on neural network compression rendering information

By using a high-performance rendering engine and neural network to extract models in the cloud, extracting and compressing drawing information, and transmitting it to mobile devices for reconstruction, it solves the problem of low latency and high frame rate rendering on mobile devices, and achieves high-quality image rendering.

CN115423925BActive Publication Date: 2025-06-06LIGHT CLOUD(HANGZHOU)TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210101473.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-06-06
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

It is difficult for the prior art to achieve high-quality image rendering with low latency and high frame rates on mobile devices, especially in three-dimensional scenes. Traditional methods require sacrificing image quality and resolution to ensure a smooth gaming experience.

Method used

Using a neural network-based cloud rendering method, the cloud uses a high-performance rendering engine to render three-dimensional scene data, extract and quantify the drawing information and transmit it to the client. The client uses the reconstruction model to reconstruct the transmitted drawing information and local rendering results.

Benefits of technology

It realizes high-quality image rendering with low latency and high frame rate on mobile devices, reducing the demand for network bandwidth and ensuring image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423925B_ABST
    Figure CN115423925B_ABST
Patent Text Reader

Abstract

The present invention discloses a cloud rendering method and device based on neural network compression of drawing information, comprising: the cloud side uses an extraction model constructed based on a neural network to extract information from the first input data, obtains drawing information and transmits it to the client; wherein the first input data includes a rendering of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene illumination information and a rendering of a historical frame; the client side uses a reconstruction model constructed based on a neural network to reconstruct the second input data to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene illumination information and a reconstruction of a historical frame. The method and device obtain real-time high-quality rendering images while maintaining low latency and high frame rate on the client by transmitting drawing information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of real-time rendering, and in particular relates to a cloud rendering method and device based on neural network compression of drawing information. Background Art

[0002] With the popularity of mobile devices and the improvement of their computing power, the demand for real-time rendering of realistic images on mobile devices has gradually increased. But at the same time, the continuous development of graphics rendering technology has also put higher and higher requirements on the computing power of devices. The growth of computing performance of mobile devices has not kept pace with people's pursuit of higher quality images. Under the existing hardware and graphics rendering technology, practitioners have to balance the picture quality, picture resolution and latency. For products such as mobile games, ensuring low latency and stable frame rate is the first thing to ensure, so many products have to sacrifice picture quality and resolution on mobile devices to ensure a smoother gaming experience for players.

[0003] In order to provide mobile device users with a better experience, some people have thought of utilizing the powerful computing power of the cloud. For example, patent document CN1856819A discloses a system and method for network transmission of graphic data of distributed applications. After the drawing operation is completed on a server with strong computing power, the drawing result is compressed and transmitted to the mobile device. This method solves the problem of low image quality and low resolution in real-time rendering of mobile devices, but the patent does not solve the problem of how to obtain low-latency and high-frame-rate images.

[0004] Patent document CN101971625A discloses a system and method for compressing streaming interactive video. In order to address the importance of low latency, it proposes a streaming transmission (video streaming) method based on improved traditional video encoding. However, this method requires a large amount of network bandwidth to achieve real-time transmission of high-frame rate and high-resolution images. The actual application of three-dimensional scenes cannot provide a large amount of network bandwidth, so it is not applicable. Summary of the invention

[0005] In view of the above, the purpose of the present invention is to provide a cloud rendering method and device based on neural network compression of drawing information, which can obtain real-time high-quality rendering images while maintaining low latency and high frame rate on the client by transmitting drawing information.

[0006] To achieve the above-mentioned object of the invention, an embodiment provides a cloud rendering method based on neural network compression drawing information, comprising the following steps:

[0007] The cloud uses an extraction model built based on a neural network to extract information from the first input data, obtains drawing information and transmits it to the client; wherein the first input data includes a rendering of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and a rendering of a historical frame;

[0008] The client reconstructs the second input data using a reconstruction model built based on a neural network to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information, and a reconstructed image of a historical frame.

[0009] In one embodiment, the cloud uses a rendering engine to render the three-dimensional scene data through a rendering pipeline to obtain a rendering image of the current frame, a rendering image of the historical frame, three-dimensional scene geometry information, and three-dimensional scene motion information; the client processes the three-dimensional scene data through the rendering pipeline to obtain three-dimensional scene geometry information and three-dimensional scene motion information.

[0010] In one embodiment, the three-dimensional scene geometry information is obtained from the rendering pipeline of the three-dimensional scene data, including position map, depth map, normal map, and material map; the three-dimensional scene motion information is obtained from the rendering pipeline of the three-dimensional scene data, including the three-dimensional scene motion information between two adjacent frames, or the three-dimensional scene motion information between the current frame and each historical frame; the lighting information includes light source parameters, wherein the light source parameters include at least one of the light source type, light source shape, light source position, light direction, light intensity, and ambient light map; the lighting information also includes the encoding vector of the light source parameters after encoding. In one embodiment, the cloud quantizes the drawing information, and the quantized drawing information is output to the client after entropy encoding.

[0011] In one embodiment, the extraction model and the reconstruction model need to be parameter optimized before application. During the parameter optimization, the difference between the reconstructed image corresponding to the three-dimensional scene data and the high-quality rendered image rendered by the rendering pipeline using the three-dimensional scene data is minimized, and the transmission bandwidth of the drawing information is minimized as an optimization goal to optimize the extraction model parameters and the reconstruction model parameters.

[0012] In one embodiment, when optimizing the image prediction model parameters, the loss function Loss corresponding to the optimization target is:

[0013]

[0014] Among them, i is the image index, loss iis the difference between the prediction result and the label corresponding to the i-th sample data, B represents the average / peak bandwidth required to transmit the residuals of this series of k images, λ is the weight parameter, which is a real number greater than 0. The larger the value, the smaller the bandwidth required to transmit the image residuals.

[0015] To achieve the above-mentioned purpose of the invention, the embodiment further provides a cloud rendering device based on neural network compression drawing information, including a cloud and a client for realizing data transmission therewith;

[0016] The cloud is deployed with an extraction model based on a neural network, which is used to extract information from the first input data, obtain drawing information and transmit it to the client, wherein the first input data includes a rendering of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and a rendering of a historical frame;

[0017] The client is deployed with a reconstruction model built based on a neural network, which is used to reconstruct the second input data to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and a reconstruction image of a historical frame.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] The cloud uses a high-performance rendering engine to quickly render the three-dimensional scene data to obtain a rendering image, and uses an extraction model built based on a neural network to extract the drawing information in the rendering image; the drawing information is quantized and compressed and then transmitted to the client. Since the drawing information data volume is small, the demand for network bandwidth is greatly reduced, thereby ensuring low latency and high frame rate for data transmission; the client uses a reconstruction model built based on a neural network to combine the drawing information and the three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information obtained by local rendering, and the reconstruction results of historical frames to reconstruct the reconstructed image. Since the reconstructed image uses the drawing information from the rendering image, the image quality of the reconstructed image is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 is a flow chart of a cloud rendering method based on neural network compression drawing information provided by an embodiment;

[0022] Figure 2It is a structural schematic diagram of a cloud rendering device based on neural network compression drawing information provided by an embodiment. DETAILED DESCRIPTION

[0023] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0024] After research, it is found that current graphics programs have complex drawing pipelines, and many intermediate results of the drawing pipeline can also be transmitted through the network. The bandwidth required to transmit these intermediate information will be much lower than the rendering obtained by directly transmitting the rendering. For the drawing pipeline, not all processes consume computing resources in the same way, so some simple rendering processes can be placed on the client. Specifically, the geometric information such as the position, normal, depth, material, etc. of objects in the three-dimensional scene, and the motion information of objects in the picture between different frames represented by motion vectors can be obtained in the rendering process of the local client. These geometric information and motion information have a great impact on the resolution of the final rendering, but the performance requirements for generating these geometric information and motion information are not high. Relatively higher performance drawing information can be calculated in the cloud and compressed and transmitted to the local. After the above research, the embodiment provides a cloud rendering method and device based on neural network compression of drawing information, which transmits drawing information while maintaining low latency and high frame rate on the client to obtain real-time high-quality renderings.

[0025] Figure 1 is a flow chart of a cloud rendering method based on neural network compression drawing information provided by an embodiment. Figure 1 As shown, the cloud rendering method based on neural network compression drawing information provided by the embodiment includes the following steps:

[0026] Step 1: The cloud uses a high-performance rendering pipeline to render the 3D scene data to obtain the rendering image of the current frame, 3D scene geometry information, 3D scene motion information, and process the lighting information of the current 3D scene.

[0027] The cloud has powerful computing power, so it can use a complete realistic rendering pipeline to render three-dimensional scene data in real time to obtain a rendered image. At the same time, it can also output three-dimensional scene geometric information such as image spatial position, depth, normal vector, material, etc. as an intermediate product of the rendering pipeline. Among them, the spatial position, depth, normal vector, and material are represented by position map, depth map, normal map, and material map.

[0028] In an embodiment, the three-dimensional scene motion information is used to link the current frame with the historical frames. Generally, the motion vectors of the objects in the current frame and the historical frames are used as the three-dimensional scene motion information. During the rendering process, the three-dimensional scene motion information between two adjacent frames represented by the motion vector or the three-dimensional scene motion information between the current frame and each historical frame can also be exported as an intermediate product of the rendering pipeline.

[0029] In one possible implementation, the illumination information may include light source parameters, which include at least one of the light source type, light source shape, light source position, light direction, light intensity, and ambient light map. When it is a point light source, the light source parameters include light source position and light intensity. When it is a parallel light source, the light source parameters include light direction and light intensity. When it is a surface light source, the light source parameters include light source position, light source shape, light direction, and light intensity distribution. When applied, the light source parameters are directly input into the shading information extraction model for extracting shading information. In another possible implementation, the three-dimensional scene illumination information may also include a coded vector of the light source parameters after encoding. The light source parameters are mapped to a latent space by encoding the light source parameters to obtain a coded vector of the light source parameters in the latent space, and then the coded vector is input into the shading information extraction model for extracting shading information.

[0030] Step 2: The cloud uses an extraction model built based on a neural network to calculate the input rendering of the current frame, 3D scene geometry information, 3D scene motion information, 3D scene lighting information, and rendering of historical frames to extract drawing information and transmit it to the client.

[0031] The extraction model built based on neural networks has the function of information extraction and compression for input data, mainly extracting drawing information from renderings. In order to reduce the coupling of drawing information with other information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and renderings of historical frames are also input as auxiliary data. After these input data are calculated by the extraction model, the drawing information of the current frame is obtained.

[0032] In an embodiment, in order to reduce transmission bandwidth and improve transmission efficiency, the cloud side quantizes the drawing information. Preferably, the cloud side uses a uniform scalar quantizer to quantize the drawing information to reduce the amount of data transmitted during network transmission. The quantized drawing information is further compressed using entropy coding and then transmitted to the client.

[0033] Since the rendering information output by the extraction model is a vector y composed of floating-point numbers, in order to reduce the bandwidth required to transmit this vector to the local, a uniform scalar quantizer is used to map the floating-point values ​​of the rendering information to integer values ​​with fewer bits. Calculate and quantify drawing information After entropy coding compression, the redundancy in the drawing information representation is further reduced.

[0034] Step 3: The client uses a low-performance rendering pipeline to render the 3D scene data to obtain 3D scene geometry information and 3D scene motion information.

[0035] The client has a low-performance rendering pipeline, which only needs to calculate the position, depth, normal vector, material and other 3D scene geometric information in the high-resolution final image space of the same 3D scene as the cloud. The 3D scene geometric information is also represented by position map, depth map, normal map and material map. The client reuses the reconstruction map of the historical frame using motion vectors, and the rendering process ends here, eliminating the calculation link of the drawing information that consumes the most, saving the client's computing overhead.

[0036] Step 4: Use the device to calculate the input drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information, and reconstruction images of historical frames using a reconstruction model built based on a neural network to obtain a reconstructed image.

[0037] In the embodiment, the reconstructed image of the historical frame is a reconstructed image of the client at the historical time, and may be a reconstruction result at any historical moment, or may be a superposition result of reconstruction images at multiple historical times.

[0038] The mobile terminal uses a decoder to decode the drawing information transmitted from the cloud, restores the drawing information of the current frame through the decoding and dequantization process of entropy coding, and then uses the reconstruction model to reconstruct the drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and reconstruction images of historical frames to obtain a reconstructed image.

[0039] In the embodiment, the extraction model and reconstruction model constructed based on the neural network need to be optimized before being applied. In the embodiment, the optimization goals include (1) reducing the transmission bandwidth of the drawing information, and (2) reducing the deviation of the high-quality rendering and reconstruction of the three-dimensional scene data. In order to achieve these two optimization goals, the average bandwidth after compression of a section of the drawing information flow can be measured, denoted as B (bandwidth), in Kbps. The L1, L2, MSE, RMSE, SSIM, PSNR and other functions can be used to measure the client's reconstruction image R i With high-quality renderings i The difference between the two functions is collectively called loss i For a video with k frames, the loss function corresponding to the optimization objective is:

[0040] The loss function Loss corresponding to the optimization objective is:

[0041]

[0042] λ is a parameter greater than 0, which is used to control the weight of the two optimization objectives. The smaller λ is, the better the reconstruction image of the reconstruction model is. The larger λ is, the smaller the bandwidth required to transmit the drawing information is, and the greater the compression of the drawing information by the extraction model is. By adjusting the size of λ, a trade-off can be made between the quality of the reconstruction image and the bandwidth required to transmit the drawing information.

[0043] Figure 2 Schematic diagram of the structure of a cloud rendering device based on neural network compression rendering information provided by an embodiment. Figure 2 As shown, the cloud rendering device provided by the embodiment includes: a cloud and a client for realizing data transmission therewith; the cloud is deployed with an extraction model constructed based on a neural network, which is used to extract information from first input data, obtain drawing information and transmit it to the client, wherein the first input data includes a rendering image of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and a rendering image of a historical frame; the client is deployed with a reconstruction model constructed based on a neural network, which is used to reconstruct the second input data to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information and a reconstructed image of a historical frame.

[0044] It should be noted that the cloud rendering device provided in the embodiment and the cloud rendering method provided in the above embodiment belong to an inventive concept. In the cloud rendering device, the method of obtaining input data, the method of constructing the extraction model and the reconstruction model, and the image rendering method of the client and the cloud are the same as those of the cloud rendering method provided in the above embodiment, and will not be repeated here.

[0045] The cloud rendering method and device based on neural network compression of drawing information provided in the above-mentioned embodiments use a high-performance rendering engine on the cloud to quickly render the three-dimensional scene data to obtain a rendering image, and use an extraction model built based on a neural network to extract the drawing information in the rendering image; the drawing information is quantized and compressed and then transmitted to the client. Since the drawing information data volume is small, the demand for network bandwidth is greatly reduced, while ensuring low latency and high frame rate for data transmission; the client uses a reconstruction model built based on a neural network to combine the drawing information and the three-dimensional scene geometry information obtained by local rendering, the three-dimensional scene motion information, and the reconstruction results of historical frames to reconstruct a reconstructed image. Since the reconstructed image uses the drawing information from the rendering image, the image quality of the reconstructed image is guaranteed.

[0046] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A 3D scene cloud rendering method based on neural network compression drawing information, It is characterized in that The following steps are involved: The cloud uses an extraction model built based on a neural network to extract information from the first input data, and the obtained drawing information is compressed and transmitted to the client; wherein the first input data includes a rendering of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene lighting information, and a rendering of a historical frame; The client reconstructs the second input data using a reconstruction model constructed based on a neural network to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene illumination information, and a reconstructed image of a historical frame; The extraction model and the reconstruction model need to be parameter optimized before application. During the parameter optimization, the difference between the reconstructed image corresponding to the three-dimensional scene data and the high-quality rendering image rendered by the rendering pipeline using the three-dimensional scene data is minimized, and the transmission bandwidth of the drawing information is minimized as an optimization goal to optimize the extraction model parameters and the reconstruction model parameters.

2. The cloud rendering method based on neural network compression rendering information according to claim 1, It is characterized in that The cloud uses a rendering engine to render the 3D scene data through a rendering pipeline to obtain a rendering of the current frame, a rendering of the historical frame, 3D scene geometry information, and 3D scene motion information; The client processes the 3D scene data through the rendering pipeline to obtain 3D scene geometry information and 3D scene motion information.

3. The cloud rendering method based on neural network compression rendering information according to claim 1 or 2, It is characterized in that The three-dimensional scene geometric information is obtained from a rendering pipeline of the three-dimensional scene data, including a position map, a depth map, a normal map, and a material map; The 3D scene motion information is obtained from a rendering pipeline of the 3D scene data, including the 3D scene motion information between two adjacent frames, or the 3D scene motion information between the current frame and each historical frame; The three-dimensional scene illumination information includes light source parameters, wherein the light source parameters include at least one of light source type, light source shape, light source position, light direction, light intensity, and ambient light map; The three-dimensional scene illumination information also includes a coding vector obtained by encoding light source parameters.

4. The cloud rendering method based on neural network compression drawing information according to claim 1, It is characterized in that The cloud quantizes the drawing information, and the quantized drawing information is output to the client after entropy coding.

5. The cloud rendering method based on neural network compression drawing information according to claim 1, It is characterized in that When optimizing the parameters of the image prediction model, the loss function corresponding to the optimization target is for: ; Where i is the image index, is the difference between the prediction result and the label corresponding to the i-th sample data, represents the average / peak bandwidth required to transmit the residuals of this series of k images, is a weight parameter, which is a real number greater than 0. The larger the value, the smaller the bandwidth required to transmit the image residual.

6. A cloud rendering device based on neural network compression rendering information, It is characterized in that Includes the cloud and the client that implements data transmission with it; The cloud is deployed with an extraction model based on a neural network, which is used to extract information from the first input data, and the obtained drawing information is compressed and transmitted to the client, wherein the first input data includes a rendering of the current frame, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene illumination information, and a rendering of a historical frame; The client is deployed with a reconstruction model built based on a neural network, which is used to reconstruct the second input data to obtain a reconstructed image, wherein the second input data includes drawing information, three-dimensional scene geometry information, three-dimensional scene motion information, three-dimensional scene illumination information and a reconstruction image of a historical frame; The extraction model and the reconstruction model need to be parameter optimized before application. During the parameter optimization, the difference between the reconstructed image corresponding to the three-dimensional scene data and the high-quality rendering image rendered by the rendering pipeline using the three-dimensional scene data is minimized, and the transmission bandwidth of the drawing information is minimized as an optimization goal to optimize the extraction model parameters and the reconstruction model parameters.

7. The cloud rendering device based on neural network compression rendering information as claimed in claim 6, It is characterized in that The cloud is used to render the three-dimensional scene data through a rendering pipeline to obtain a rendering image of a current frame, a rendering image of a historical frame, three-dimensional scene geometry information, and three-dimensional scene motion information; The client is used to process the three-dimensional scene data through a rendering pipeline to obtain three-dimensional scene geometry information and three-dimensional scene motion information.

8. The cloud rendering device based on neural network compression rendering information as claimed in claim 6, It is characterized in that The three-dimensional scene geometric information is obtained from a rendering pipeline of the three-dimensional scene data, including a position map, a depth map, a normal map, and a material map; The 3D scene motion information is obtained from a rendering pipeline of the 3D scene data, including the 3D scene motion information between two adjacent frames, or the 3D scene motion information between the current frame and each historical frame; The three-dimensional scene illumination information includes light source parameters, wherein the light source parameters include at least one of light source type, light source shape, light source position, light direction, light intensity, and ambient light map; The three-dimensional scene illumination information also includes a coding vector obtained by encoding light source parameters.

9. The cloud rendering device based on neural network compression rendering information as claimed in claim 6, It is characterized in that The cloud quantizes the drawing information, and the quantized drawing information is entropy-coded and then output to the client.

Citation Information

Patent Citations

  • System and method for network transmission of graphical data through a distributed application

    CN1856819A

  • System and method for compressing streaming interactive video

    CN101971625A

  • Three-dimensional reconstruction method, device and system

    CN113362450A