Image rendering method and apparatus, and device
Through the two-stage rendering method, the first rendering engine is used to generate the initial rendering image, and the picture quality is improved through the neural network rendering model of the second rendering engine, solving the problem of large amount of real-time rendering and low efficiency of image rendering, and achieving efficient and low-computing resources rendering effect.
Patent Information
- Application Number
- PCT/CN2024/117279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2024-09-05
- Publication Date
- 2025-06-05
AI Technical Summary
Real-time image rendering is difficult to meet in scenarios with high computational volume, consumes a lot of computing resources and affects rendering efficiency, especially in scenarios with high latency requirements.
A two-stage rendering method is adopted: first, the first rendering engine is used to render the basic image based on the picture data to be rendered to obtain the initial rendering image; then the second rendering engine, including the neural network rendering model, perform picture quality processing on the initial rendering image to generate a high-quality target rendering image.
While satisfying the user's visual experience, it significantly reduces the amount of computing in the image rendering process, saves computing resources, and accelerates the rendering process and improves rendering efficiency.
Smart Images

Figure CN2024117279_05062025_PF_FP_ABST
Abstract
Description
Image rendering method, device and equipment
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 1, 2023, with application number 202311657628.4 and application name “A rendering method and device”, the entire contents of which are incorporated by reference into this application; this application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on January 25, 2024, with application number 202410109086.5 and application name “A image rendering method, device and equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of image processing technology, and in particular to an image rendering method, apparatus, and device. Background Art
[0004] Real-time image rendering technology is an important research topic in 3D computer graphics. It can output 3D scene data and environmental data into rendered images that simulate the real world or virtual scenes through calculation. It has been widely used in scenarios such as games and 3D applications. Graphics processing units (GPUs) can usually be used for real-time image rendering.
[0005] Since real-time image rendering requires a lot of computation, it usually consumes a lot of computing resources, affects rendering efficiency, and cannot meet scenarios with high latency requirements. Therefore, in the field of real-time image rendering, how to reduce the computing resources required for image rendering and improve rendering efficiency is an urgent problem to be solved.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an image rendering method, apparatus, and device, which can reduce the computing resources required for image rendering and improve rendering efficiency.
[0008] In a first aspect, an embodiment of the present application provides an image rendering method, which can be executed by a computing device, or a chip, chip system or circuit in a computing device. The computing device can be a cloud server or a terminal device. The method may include: the computing device obtains an initial rendered image, and the initial rendered image is obtained by using a first rendering engine to perform basic image rendering based on the image data to be rendered. After obtaining the initial rendered image, the computing device uses a second rendering engine to process the picture quality of the initial rendered image to obtain a target rendered image, and the picture quality parameter value of the target rendered image is higher than the picture quality parameter value of the initial rendered image. The second rendering engine includes a neural network rendering model.
[0009] Since the amount of computation performed in the process of rendering a rendered image with higher picture quality using traditional image rendering methods is very large, an embodiment of the present application provides an image rendering method. After obtaining a low-quality initial rendered image obtained by basic image rendering using a first rendering engine, a second rendering engine is used to process the picture quality of the low-quality initial rendered image through a neural network rendering model to obtain a high-quality target rendered image. While satisfying the user's visual experience, compared with traditional image rendering methods, the amount of computation in the image rendering process can be greatly reduced, saving computing resources; at the same time, it can also accelerate the image rendering process and improve rendering efficiency.
[0010] In one possible implementation, the picture quality parameters may include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
[0011] In one possible implementation, assuming that the aforementioned initial rendered image is obtained based on the to-be-rendered image data at a first moment, after obtaining the initial rendered image, the computing device may input the initial rendered image and a rendered image obtained at a moment before the first moment into a second rendering engine to obtain a target rendered image.
[0012] In the above implementation method, the rendered image obtained at the previous moment is the high-quality target rendering image at the previous moment. The high-quality target rendering image at the previous moment can provide more image information. Combining the high-quality target rendering image at the previous moment to render the image at the current moment can improve the temporal stability of the presented image.
[0013] In one possible implementation, the above-described image rendering method can be performed by a first device. After obtaining a target rendered image, the first device transmits the target rendered image to a second device for display. The first device and the second device can be different computing devices. For example, the first device can be a cloud server, and the second device can be a terminal device. Performing image rendering tasks in the cloud can reduce the consumption of computing resources on the terminal device and improve rendering efficiency.
[0014] In one possible implementation, a neural network rendering model is trained based on multiple sets of training samples. Each set of training samples includes multiple training images. The multiple training images are rendered based on training data of the same scene at the same time. The multiple training images have different image quality parameter values.
[0015] In the above implementation method, image rendering is performed based on training data of the same time and the same scene to obtain multiple training images with different picture quality parameter values. Model training is performed using multiple training images with different picture quality parameter values to obtain a neural network rendering model for image rendering.
[0016] In one possible implementation, each of the multiple sets of training samples further includes training reference information generated based on the training data. Exemplarily, the training reference information includes part or all of the following: a depth map, a normal map, camera position parameters, and a motion vector map.
[0017] In the above implementation, the training reference information is used to assist in model training, which can improve the image rendering effect of the obtained neural network rendering model.
[0018] In a possible implementation, multiple groups of training samples are obtained based on training data at multiple consecutive moments.
[0019] In one possible implementation, before acquiring the initial rendered image, the computing device may present a model selection interface for displaying at least one trained neural network rendering model. Upon receiving a model selection operation, the computing device may load the neural network rendering model specified in the model selection operation from the at least one trained neural network rendering model into the second rendering engine.
[0020] In the above implementation, multiple neural network rendering models can be saved in the computing device, and different neural network rendering models can output target rendered images with different picture qualities. In actual applications, users can choose appropriate neural network rendering models according to actual needs, thereby flexibly meeting the various needs of users.
[0021] In a second aspect, an embodiment of the present application provides an image rendering device, which may include:
[0022] A first rendering engine is used to perform basic image rendering based on the image data to be rendered to obtain an initial rendered image;
[0023] The second rendering engine is used to process the picture quality of the initial rendered image to obtain a target rendered image; the picture quality parameter value of the target rendered image is higher than the picture quality parameter value of the initial rendered image, and the second rendering engine includes a neural network rendering model.
[0024] In one possible implementation, the picture quality parameters include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
[0025] In a possible implementation, the second rendering engine may be configured to process the initial rendered image and a rendered image obtained at a moment before the first moment to obtain a target rendered image.
[0026] In a possible implementation, the image rendering apparatus may be provided in the first device, and the image rendering apparatus may further include an image transmission module, which is configured to transmit the target rendered image to the second device for display after obtaining the target rendered image.
[0027] In one possible implementation, a neural network rendering model is trained based on multiple groups of training samples; each group of training samples in the multiple groups of training samples includes multiple training images; the multiple training images are obtained by image rendering based on training data at the same time; and the picture quality parameter values of the multiple training images are different.
[0028] In a possible implementation, each group of training samples in the multiple groups of training samples further includes training reference information generated based on the training data.
[0029] In a possible implementation, the training reference information includes part or all of the following: a depth map, a normal map, camera position parameters, and a motion vector map.
[0030] In a possible implementation, multiple groups of training samples are obtained based on training data at multiple consecutive moments.
[0031] In one possible implementation, the image rendering device may further include a human-computer interaction module, which may be used to: present a model selection interface; the model selection interface is used to display at least one trained neural network rendering model; when a model selection operation is received, the neural network rendering model specified by the model selection operation in at least one trained neural network rendering model is loaded into the second rendering engine.
[0032] In a third aspect, an embodiment of the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes any one of the image rendering methods provided in the first aspect.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to enable a computer to execute any one of the image rendering methods provided in the first aspect.
[0034] In a fifth aspect, an embodiment of the present application provides a computer program product comprising computer-executable instructions, wherein the computer-executable instructions are used to enable a computer to execute any one of the image rendering methods provided in the first aspect.
[0035] The technical effects that can be achieved in any of the second to fifth aspects mentioned above can refer to the description of the beneficial effects in the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of an application scenario of an image rendering method provided by an embodiment of the present application;
[0037] FIG2 is a schematic diagram of the internal structure of a computing device provided in an embodiment of the present application;
[0038] FIG3 is an interaction diagram between various modules in a model training process provided by an embodiment of the present application;
[0039] FIG4 is a schematic diagram of a training image generation process provided in an embodiment of the present application;
[0040] FIG5 is a schematic diagram of training a rendering model to be trained provided by an embodiment of the present application;
[0041] FIG6 is a flowchart of an image rendering method provided by an embodiment of the present application;
[0042] FIG7 is a schematic diagram of the internal structure of another computing device provided in an embodiment of the present application;
[0043] FIG8 is an interaction diagram between various modules in an image rendering process provided by an embodiment of the present application;
[0044] FIG9 is a schematic diagram of an image rendering process provided by an embodiment of the present application;
[0045] FIG10 is a structural block diagram of an image rendering device provided in an embodiment of the present application;
[0046] FIG11 is a structural block diagram of another image rendering device provided in an embodiment of the present application;
[0047] FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0048] FIG13 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0050] Before introducing the specific solutions provided by the embodiments of the present application, some of the terms in the present application are explained to facilitate understanding by those skilled in the art, and the terms in the present application are not limited.
[0051] (1) Three-dimensional (3D) applications: 3D applications are programs that can use the computing resources of the device to render 3D images. The computing resources can refer to the GPU or other processors.
[0052] In the embodiments of the present application, "multiple" refers to two or more. In view of this, in the embodiments of the present application, "multiple" can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two or more. For example, including at least one means including one, two or more, and does not limit which ones are included. For example, including at least one of A, B and C, then the included ones may be A, B, C, A and B, A and C, B and C, or A, B and C. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0053] Unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects.
[0054] The image rendering method provided in the embodiments of the present application can be applied to the image rendering system shown in Figure 1. As shown in Figure 1, the image rendering system may include a cloud server 100 and a terminal device 200 in communication with the cloud server 100. Figure 1 only illustrates one terminal device. In some embodiments, the cloud server 100 may be connected to multiple terminal devices, and the terminal devices and the cloud server 100 may be connected via a communication network.
[0055] The cloud server 100 can be a physical device such as a physical server, or a virtual instance such as a virtual machine or container. The terminal device 200 can be a mobile terminal device such as a mobile phone, tablet computer, laptop computer, smart wearable device, smart car device, or a personal computer, virtual machine, container, multimedia player, smart home appliance, artificial intelligence device, e-reader, or Internet of Things device.
[0056] A 3D application is deployed on the cloud server 100. For example, the 3D application may be a game. A client for the 3D application is installed on the terminal device 200. The user can log in to the 3D application through the client, and the display screen of the 3D application can be obtained by performing image rendering. When the user inputs an operation for the 3D application through the client, the terminal device 200 can send a rendering request to the cloud server 100 based on the operation input by the user. The rendering request is used to instruct the cloud server 100 to generate the display screen of the 3D application required by the client. The cloud server 100 can perform the corresponding image rendering operation based on the rendering request, and transmit the rendered display screen to the terminal device 200, so that the terminal device 200 displays the display screen.
[0057] Taking into account the large amount of computation required for real-time image rendering, in order to reduce the computing resources required for image rendering and improve rendering efficiency, an embodiment of the present application provides an image rendering method, which can be executed by the cloud server 100 in Figure 1. The cloud server can use a first rendering engine to perform basic image rendering based on the image data to be rendered to obtain an initial rendered image, and the image quality of the initial rendered image is relatively low. The cloud server can use a second rendering engine to process the image quality of the initial rendered image to obtain a target rendered image, and transmit the target rendered image to a terminal device for display. The second rendering engine includes a neural network rendering model, and the image quality parameter value of the target rendered image is higher than the image quality parameter value of the initial rendered image, that is, the image quality of the target rendered image is higher. By processing the image quality of the initial rendered image with the second rendering engine, not only can a high-quality target rendered image be obtained, but also rendering efficiency can be improved, the amount of computation can be reduced, and computing resources can be saved.
[0058] For example, as shown in Figure 1, after receiving a rendering request from a client via a terminal device, the 3D application on the cloud server can use a first rendering engine to render the image based on the image data to be rendered, obtaining an initial rendered image. To reduce computational complexity, the image quality of the initial rendered image can be lower. The 3D application can then input the initial rendered image into a neural network rendering model in a second rendering engine. The neural network rendering model then processes the image quality of the initial rendered image to obtain a high-quality target rendered image. The target rendered image is then transmitted to the terminal device for display via an image encoding output module on the cloud server.
[0059] It should be noted that the image rendering system shown in FIG1 is only an exemplary illustration of the application scenario of the present application. The image rendering method provided in the embodiment of the present application is not limited to the application scenario shown in FIG1 , and can also be applied to other application scenarios. For example, the image rendering method provided in the embodiment of the present application can be executed by any computing device; or, in actual applications, the image rendering method provided in the embodiment of the present application can be executed by a rendering device, and the rendering device can be implemented by software or by hardware.
[0060] As an example of a software functional unit, a rendering device may include code running on a compute instance. A compute instance may be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Furthermore, the computing device may be one or more. For example, the rendering device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone (AZ) or in different AZs, with each AZ comprising a data center or multiple geographically close data centers. Typically, a region may include multiple AZs. Similarly, the multiple hosts / virtual machines / containers running the code can be distributed in the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is located within a region. Cross-region communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0061] As an example of a hardware functional unit, a rendering device may include at least one computing device, such as a server. Alternatively, the rendering device may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented as a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The multiple computing devices included in the rendering device may be distributed in the same region or in different regions. The multiple computing devices included in the rendering device may be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the rendering device may be distributed in the same VPC or in multiple VPCs. The multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0062] The image rendering method provided in the embodiment of the present application uses a trained neural network rendering model to process the picture quality of the initial rendered image to obtain a target rendered image. For ease of understanding, the following first introduces the training process of the neural network rendering model. For the training process of the neural network rendering model, the embodiment of the present application also provides a model training method. The model training method provided in the embodiment of the present application can be executed by any one or more computing devices. The computing device that executes the model training method and the computing device that executes the image rendering method can be the same computing device or a different computing device. The computing device can obtain multiple groups of training samples, each of the multiple groups of training samples can include multiple training images, and the multiple training images in a group of training samples are obtained by image rendering based on the training data at the same time, and the values of the picture quality parameters of the multiple training images are different. The computing device can use multiple groups of training samples to train the rendering model to be trained to obtain a neural network rendering model.
[0063] In some embodiments, as shown in FIG2 , the computing device 300 is provided with a first human-computer interaction module, a first rendering control module, one or more rendering modules, a data storage module, a model storage module, and a training module. In actual use, the model storage module can be set in the computing device, or can be set outside the computing device, or can be set in other computing devices. In one embodiment, the computing device 300 shown in FIG2 can be a computing device cluster, including multiple computing devices, and different functional modules can be set on different computing devices. During the model training process, the interaction process between the various functional modules, as shown in FIG3 , can include the following steps:
[0064] S301: A first human-computer interaction module generates a rendering strategy according to a value of an image quality parameter input by a user.
[0065] The first human-computer interaction module can receive the value of the picture quality parameter input by the user through the input component of the computing device, and generate a rendering strategy according to the value of the picture quality parameter. The picture quality parameters may include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters. Exemplarily, the computing device can display a parameter setting interface, and the parameter setting interface may include multiple dialog boxes, such as a dialog box corresponding to image resolution, a dialog box corresponding to anti-aliasing parameters, a dialog box corresponding to shadow effect parameters, a dialog box corresponding to ray tracing effect parameters, and the like. The parameter value input by the user can be received through the dialog box. For example, the dialog box corresponding to the image resolution can receive the parameter value of the image resolution input by the user, and so on.
[0066] The first human-computer interaction module may generate a rendering strategy based on the value of the image quality parameter input by the user. The rendering strategy is used to indicate the image quality of the training images generated by the rendering module. In some embodiments, the first human-computer interaction module may receive multiple sets of image quality parameter values input by the user. The generated rendering strategy may be used to instruct different rendering modules to generate training images of different image qualities, thereby obtaining multiple training images, where the training images refer to the rendered images used to train the model.
[0067] S302: The first human-computer interaction module transmits a rendering strategy to the first rendering control module.
[0068] S303: The first rendering control module controls the rendering module a to generate a training image 1a.
[0069] S304: The first rendering control module controls the rendering module b to generate a training image 1b.
[0070] In some embodiments, the first human-computer interaction module can launch a cloud rendering instance to perform rendering. The cloud rendering instance can launch a 3D application, which can run the first rendering control module. In other words, the first rendering control module can be called by the 3D application to control the image rendering process. For example, as shown in FIG4 , the first rendering control module can be a rendering plug-in of the 3D application. The rendering plug-in can control rendering modules a and b to generate training images according to a rendering strategy based on training data at the same time or the same scene, and generate training images with different image effects for the same frame of training data. In other words, the first rendering control module controls rendering module a to render the image based on the training data at time t1 according to a first set of values of the image quality parameters in the rendering strategy, generating training image 1a. The first rendering control module controls rendering module b to render the image based on the training data at time t1 according to a second set of values of the image quality parameters in the rendering strategy, generating training image 1b. The first set of values and the second set of values are two sets of values of the image quality parameters input by the user, and the first set of values is lower than the second set of values. Therefore, the resulting training image 1a is a low-quality rendered image, and the training image 1b is a high-quality rendered image. This embodiment uses the generation of two training images as an example. In other embodiments, multiple training images can be generated based on training data from the same frame at the same time. The image quality parameters of the multiple training images differ, or in other words, the image quality of the multiple training images differs. The multiple training images constitute a set of training samples. In an optional embodiment, the first rendering control module can control multiple rendering modules to generate multiple training images. In another optional embodiment, the first rendering control module can control a single rendering module to generate multiple training images.
[0071] Step S303 and step S304 are executed in a loop. The first rendering control module can control the rendering module to generate multiple groups of training samples based on the training data at multiple moments. Each group of training samples can include multiple training images.
[0072] In some embodiments, the first rendering control module may further control the rendering module to generate training reference information based on the training data, that is, each of the multiple sets of training samples may include, in addition to multiple training images, training reference information. Exemplarily, the training reference information may include some or all of the reference images such as the depth map, normal map, camera position parameters, and motion vector map. As shown in FIG4 , the training reference information may include the depth map and normal map corresponding to the low-quality rendered image 1a and the depth map and normal map corresponding to the high-quality rendered image 1b.
[0073] S305 : The first rendering control module saves the training image 1 a and the training image 1 b to the data storage module.
[0074] After obtaining the multiple sets of training samples, the first rendering control module may save the multiple sets of training samples to the data storage module.
[0075] S306: The first human-computer interaction module sends a training start instruction to the training module.
[0076] In some embodiments, after generating a sufficient number of training samples, the first human-computer interaction module may send a start training instruction to the training module based on a start training operation input by the user to trigger model training. In other embodiments, the first rendering control module may count the number of generated training samples, and after the number of training samples reaches a set number, may send a notification to the first human-computer interaction module. Upon receiving the notification, the first human-computer interaction module may send a start training instruction to the training module to trigger model training. The set number may be a value set by the user.
[0077] S307: The data storage module transmits the training image to the training module.
[0078] The training module starts training and obtains training images from the data storage module.
[0079] S308: The training module uses the training image to train the rendering model to be trained to obtain a trained neural network rendering model.
[0080] The training module receives the start training instruction sent by the first human-computer interaction module, obtains multiple groups of training samples from the data storage module, and uses the multiple training images contained in each group of training samples to train the rendering model to be trained to obtain a trained neural network rendering model. The rendering model to be trained can be any deep learning network model or AI model. Exemplarily, the rendering model to be trained can be a network structure of a generative adversarial nets (GAN), and the rendering model to be trained can include two neural networks, one of which is a generator network (generator network) and the other is a discriminator network (discriminator network); the generator network can use low-quality training images in multiple groups of training samples to continuously generate new images, and the discriminator network can judge whether the image generated by the generator network meets the requirements based on the high-quality training images in the multiple groups of training samples.
[0081] For example, in an optional embodiment, the training module may input low-quality training images from a set of training samples into the rendering model to be trained to obtain an image to be verified output by the rendering model to be trained; alternatively, the training reference information from a set of training samples may be input into the rendering model to be trained together with the low-quality training images, and the training reference information and the low-quality training images may be processed by the generator network of the rendering model to be trained to generate an image to be verified. The image to be verified is then compared with a high-quality training image by a discriminator network, and a loss value is determined based on the difference between the image to be verified and the high-quality training image, wherein the high-quality training image and the low-quality training image belong to the same set of training samples. The network parameters of the rendering model to be trained are adjusted based on the loss value, and the steps of selecting a set of training samples and inputting the low-quality training images into the rendering model to be trained are repeated for iterative training, and the loss value is recalculated until the change in the calculated loss value is within a set threshold, or the number of iterations reaches a set number. In this case, the rendering model to be trained is considered to have converged, resulting in a trained neural network rendering model.
[0082] In another optional embodiment, the training module can input the high-quality training image at the previous moment and the low-quality training image at the next moment of the training samples from at least two consecutive moments into the rendering model to be trained, and the generator network of the rendering model to be trained processes the high-quality training image at the previous moment and the low-quality training image at the next moment to generate an image to be verified. The discriminator network then compares the image to be verified with the high-quality training image at the next moment, and determines the loss value based on the difference between the image to be verified and the high-quality training image at the next moment. As shown in Figure 5, the high-quality training image (N-1)-b at moment N-1 and the low-quality training image Na at moment N can be input into the rendering model to be trained, and the generator network in the rendering model to be trained outputs the image to be verified at moment N. The loss value is determined based on the difference between the image to be verified at moment N and the high-quality training image Nb at moment N. The network parameters of the rendering model to be trained are adjusted according to the loss value, and iterative training is performed. The loss value is calculated again until the change in the calculated loss value is within the set amplitude threshold, or the number of iterations reaches the set number. In this case, the rendering model to be trained can be considered to have converged, and a trained neural network rendering model is obtained.
[0083] S309: The training module saves the neural network rendering model to the model storage module.
[0084] After obtaining the trained neural network rendering model, the training module can save the trained neural network rendering model to the model storage module.
[0085] In some optional embodiments, the training module can train multiple neural network rendering models separately, and save multiple versions of the neural network rendering models in the model storage module. The picture quality of the rendered images generated by different versions of the neural network rendering models can be different.
[0086] In some embodiments, after obtaining the trained neural network rendering model, the computing device may use the trained neural network rendering model to perform the image rendering process. In other embodiments, after obtaining the trained neural network rendering model, the first computing device may transmit the trained neural network rendering model to the second computing device, and the second computing device may use the trained neural network rendering model to perform the image rendering process. Exemplarily, the first computing device may be a cloud server, and the second computing device may be a terminal device. That is to say, the image rendering method provided in the embodiments of the present application can be executed by any computing device, which may be a computing device that performs neural network rendering model training, or other computing devices; the computing device may be a cloud server or a terminal device.
[0087] As shown in FIG6 , the execution process of the image rendering method provided in the embodiment of the present application may include the following steps:
[0088] S601: Acquire an initial rendering image.
[0089] The initial rendered image is obtained by performing basic image rendering based on the image data to be rendered using a first rendering engine, and the image quality of the initial rendered image is low. The first rendering engine can be set in a computing device that executes the image rendering method, or it can be set in another computing device. For example, the first rendering engine can be set in computing device A, and the first rendering engine is used to perform basic image rendering based on the image data to be rendered to obtain an initial rendered image. Computing device B can be connected to computing device A via a wired network or a wireless network, obtain the initial rendered image from computing device A, and then further process the initial rendered image using the second rendering engine.
[0090] S602: Process the picture quality of the initial rendered image using a second rendering engine to obtain a target rendered image.
[0091] The second rendering engine includes a neural network rendering model, which can be used to process the picture quality of the initial rendered image to obtain a target rendered image. The value of the picture quality parameter of the target rendered image is higher than the value of the picture quality parameter of the initial rendered image. The picture quality parameters may include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
[0092] In an optional embodiment, as shown in FIG7 , a computing device 700 that performs the image rendering process may be provided with multiple functional modules such as a second human-computer interaction module, a second rendering control module, one or more rendering modules, a model storage module, and an image encoding output module. Exemplarily, the computing device 700 may be the above-mentioned cloud server or other network-side device. In one embodiment, the computing device 700 shown in FIG7 may be a computing device cluster, including multiple computing devices, and different functional modules may be provided on different computing devices. During the image rendering process, the interaction process between the various functional modules, as shown in FIG8 , may include the following steps:
[0093] S801: The second human-computer interaction module generates a rendering strategy.
[0094] Exemplarily, when a plurality of rendering models are stored in the model storage module, the second human-computer interaction module may present a model selection interface, which is used to display at least one rendering model that can be selected by the user. The second human-computer interaction module may receive a model selection operation input by the user through the input component of the computing device, and use the rendering model specified by the model selection operation in at least one rendering model as a neural network rendering model to generate a rendering strategy including the neural network rendering model selected by the user. In some optional embodiments, the second human-computer interaction module may also receive a picture quality parameter value input by the user through the input component of the computing device, and generate a rendering strategy including the picture quality parameter value and the neural network rendering model selected by the user, wherein the picture quality parameter value is used to indicate the picture quality of the initial rendered image generated by the rendering module.
[0095] S802: The second human-computer interaction module transmits a rendering strategy to the second rendering control module.
[0096] S803: The second rendering control module controls the rendering module c to configure the value of the picture quality parameter.
[0097] S804: The model storage module transmits the neural network rendering model to the rendering module c.
[0098] The rendering module c includes a first rendering engine and a second rendering engine. The rendering module c obtains the neural network rendering model from the model storage module and loads the neural network rendering model into the second rendering engine.
[0099] In some embodiments, the second human-computer interaction module can start a cloud rendering instance to perform rendering. The cloud rendering instance can start a 3D application, and the 3D application can run the second rendering control module, or in other words, the second rendering control module can be called by the 3D application to control the image rendering process. For example, the 3D application can be a game application, and the second rendering control module can be called by the game application to control the rendering of each frame of the game. Exemplarily, after the second rendering control module receives the rendering strategy transmitted by the second human-computer interaction module, it obtains the picture quality parameter value and the neural network rendering model indicated by the rendering strategy, controls the rendering module c to configure the picture quality parameter value for the first rendering engine, and controls the rendering module c to obtain the user-specified neural network rendering model from the model storage module and load it into the second rendering engine.
[0100] S805: The rendering module c generates an initial rendering image.
[0101] The second rendering control module can control the first rendering engine in the rendering module C to generate a low-quality initial rendering image according to the value of the image quality parameter carried in the rendering strategy. The first rendering engine in the rendering module C can generate a low-quality initial rendering image based on the image data to be rendered at the current moment. The value of the image quality parameter is used to indicate how the first rendering engine currently needs to render, such as what rendering resolution, anti-aliasing level, lighting and shadow effects to use, so that the initial rendering image output by the rendering output matches the training input of the selected neural network rendering model as much as possible to achieve the optimal subsequent image rendering effect.
[0102] S806, the rendering module c processes the picture quality of the initial rendered image through the neural network rendering model to obtain a target rendered image.
[0103] In some embodiments, the second rendering engine in the rendering module c can call a trained neural network rendering model, input the initial rendering image into the neural network rendering model, and obtain a target rendering image output by the neural network rendering model, where the value of the picture quality parameter of the target rendering image is higher than the value of the picture quality parameter of the initial rendering image.
[0104] In other embodiments, the second rendering engine in the rendering module c can call a trained neural network rendering model, input the initial rendering image at the current moment and the target rendering image obtained at the previous moment into the neural network rendering model, and obtain the target rendering image at the current moment output by the neural network rendering model. For example, as shown in Figure 9, the high-quality target rendering image at moment N-1 (i.e., the previous moment) and the low-quality initial rendering image at moment N (i.e., the current moment) can be input into the neural network rendering model, and the high-quality target rendering image at moment N can be output by the neural network rendering model. At the next moment, the high-quality target rendering image at moment N is continued to be input into the neural network rendering model as the high-quality target rendering image at the previous moment. The high-quality target rendering image at the previous moment can also be referred to as the rendered image at the previous moment. The high-quality target rendering image at the previous moment can provide more image information. Combining the high-quality target rendering image at the previous moment to render the image at the current moment can improve the temporal stability of the presented image and avoid image flickering or jumps.
[0105] S807: The rendering module c transmits the target rendered image to the image encoding output module.
[0106] After obtaining the target rendered image, the rendering module c transmits the target rendered image to the image encoding output module. The target rendered image can be media encoded by the image encoding output module and transmitted to the client of the terminal device through media transmission; in other words, the target rendered image is transmitted to the client of the terminal device through the streaming capability of the image encoding output module.
[0107] In actual use, the above steps S805, S806 and S807 can be executed in a loop to generate multiple target rendering images at multiple consecutive moments within a period of time, and transmitted to the client of the terminal device in the form of a video stream, so that the terminal device can display the video screen of the game.
[0108] In the above embodiment, both the first rendering engine and the second rendering engine are provided in the rendering module c. The rendering module c can call the neural network rendering model through the second rendering engine, input the initial rendered image output by the first rendering engine into the neural network rendering model, and obtain the target rendered image. In other embodiments, the first rendering engine can be provided in the rendering module c, and the second rendering engine can be provided in the image coding output module. After the first rendering engine generates the initial rendered image, the rendering module c can transmit the initial rendered image to the image coding output module. The image coding output module can call the neural network rendering model through the second rendering engine, input the initial rendered image into the neural network rendering model, obtain the target rendered image, and then transmit the target rendered image to the client of the terminal device.
[0109] The above embodiments are described using the computing device being a cloud server as an example. In other embodiments, the computing device may also be a terminal device, and the computing device may not include an image encoding output module but include a display module. After obtaining the target rendered image, the rendering module c may transmit the target rendered image to the display module for display.
[0110] The computing device used to perform image rendering and the computing device used for model training mentioned above can be the same computing device or different computing devices. When the computing device used to perform image rendering and the computing device used for model training are the same computing device, that is, when the computing device 300 shown in Figure 2 and the computing device 700 shown in Figure 7 are the same computing device, the above-mentioned first human-computer interaction module and the second human-computer interaction module can be the same human-computer interaction module; the above-mentioned first rendering control module and the second rendering control module can be the same rendering control module; the above-mentioned rendering module a, rendering module b and rendering module c can be integrated into one rendering module.
[0111] The image rendering method provided in the embodiment of the present application uses a neural network rendering model to process a low-quality initial rendered image to obtain a high-quality target rendered image. While satisfying the user's visual experience, compared with traditional image rendering methods, it can significantly reduce the amount of computation required in the image rendering process, saving computing resources; it can also accelerate the image rendering process and improve rendering efficiency. On the premise of obtaining a target rendered image of the same picture quality, using the same computing resources, if a traditional image rendering method can render 30 frames per second, the image rendering method provided in the embodiment of the present application can render 60 frames per second, doubling the rendering efficiency and meeting the needs of gaming scenarios or other practical application scenarios with high latency requirements.
[0112] In conjunction with the above method embodiments, embodiments of the present application also provide an image rendering device. In some embodiments, as shown in FIG10 , the image rendering device 1000 may include a first rendering engine 1010 and a second rendering engine 1020. The image rendering device 1000 can be used to implement the functions of the above method embodiments, thereby achieving the beneficial effects of the above method embodiments.
[0113] Among them, the first rendering engine 1010 is used to perform basic image rendering based on the picture data to be rendered to obtain an initial rendered image; the second rendering engine 1020 is used to process the picture quality of the initial rendered image to obtain a target rendered image; the picture quality parameter value of the target rendered image is higher than the picture quality parameter value of the initial rendered image, and the second rendering engine includes a neural network rendering model.
[0114] In an optional embodiment, the picture quality parameters include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
[0115] In an optional embodiment, as shown in FIG11 , the image rendering device 1000 may further include a human-computer interaction module 1030, which may be used to: present a model selection interface; the model selection interface is used to display at least one trained neural network rendering model; and when a model selection operation is received, the neural network rendering model specified by the model selection operation in at least one trained neural network rendering model is loaded into the second rendering engine.
[0116] In an optional embodiment, the image rendering apparatus may be provided in the first device. As shown in FIG11 , the image rendering apparatus may further include an image transmission module 1040 , which is configured to transmit the target rendered image to the second device for display after obtaining the target rendered image.
[0117] The first rendering engine 1010, the second rendering engine 1020, the human-computer interaction module 1030, and the image transmission module 1040 can all be implemented via software or hardware. For example, the implementation of the second rendering engine 1020 will be described below using the second rendering engine 1020 as an example. Similarly, the implementation of the first rendering engine 1010, the human-computer interaction module 1030, and the image transmission module 1040 can refer to the implementation of the second rendering engine 1020.
[0118] As an example of a software functional unit, the second rendering engine 1020 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the second rendering engine 1020 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same AZ or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0119] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0120] As an example of a hardware functional unit, the second rendering engine 1020 may include at least one computing device, such as a server. Alternatively, the second rendering engine 1020 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0121] The multiple computing devices included in the second rendering engine 1020 can be distributed in the same region or in different regions. The multiple computing devices included in the A module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the A module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0122] It should be noted that in other embodiments, the second rendering engine 1020 can be used to execute any step in the image rendering method, the first rendering engine 1010 can be used to execute any step in the image rendering method, the human-computer interaction module 1030 can be used to execute any step in the image rendering method, and the image transmission module 1040 can be used to execute any step in the image rendering method. The steps that the first rendering engine 1010, the second rendering engine 1020, the human-computer interaction module 1030, and the image transmission module 1040 are responsible for implementing can be specified as needed. The full functionality of the image rendering device is achieved by having the first rendering engine 1010, the second rendering engine 1020, the human-computer interaction module 1030, and the image transmission module 1040 respectively implement different steps in the image rendering method. In other embodiments, the image rendering device may also include more or fewer functional modules, and this application is not limited to this.
[0123] The present application also provides a computing device 1200. As shown in FIG12 , computing device 1200 can be used to implement the functions of the image rendering apparatus described in the above embodiments and includes a bus 1201, a processor 1202, a memory 1203, and a communication interface 1204. Processor 1202, memory 1203, and communication interface 1204 communicate with each other via bus 1201. Computing device 1200 can be a cloud server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 1200.
[0124] Bus 1201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG12 shows only one line, but this does not imply a single bus or type of bus. Bus 1201 may include a path for transmitting information between various components of computing device 1200 (e.g., memory 1203, processor 1202, and communication interface 1204).
[0125] The processor 1202 may include any one or more processors such as a CPU, a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0126] The memory 1203 may include a volatile memory, such as a random access memory (RAM). The processor 1202 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0127] Memory 1203 stores executable program code. Processor 1202 executes the executable program code to implement the functions of the first rendering engine 1010, the second rendering engine 1020, the human-computer interaction module 1030, and the image transmission module 1040, thereby implementing the image rendering method. In other words, memory 1203 stores instructions for executing the image rendering method.
[0128] The communication interface 1204 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1200 and other devices or a communication network.
[0129] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. At least one computing device in the computing device cluster cooperates to implement the functions of the computing device in the above embodiments. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0130] As shown in Figure 13, the computing device cluster includes at least one computing device 1200. The memory 1203 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the image rendering method.
[0131] In some possible implementations, the memory 1203 of one or more computing devices 1200 in the computing device cluster may also store some instructions for executing the image rendering method. In other words, the combination of one or more computing devices 1200 can jointly execute the instructions for executing the image rendering method.
[0132] It should be noted that the memory 1203 in different computing devices 1200 in the computing device cluster can store different instructions, each used to perform a portion of the functions of the computing device. In other words, the instructions stored in the memory 1203 in different computing devices 1200 can implement the functions of one or more units in the first rendering engine 1010, the second rendering engine 1020, the human-computer interaction module 1030, and the image transmission module 1040.
[0133] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network.
[0134] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the above-described image rendering method.
[0135] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to execute an image rendering method, or instructs a computing device to execute an image rendering method.
[0136] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0137] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.
[0138] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0140] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. An image rendering method, characterized in that: include: Get the initial rendered image; The initial rendered image is obtained by performing basic image rendering based on the picture data to be rendered using the first rendering engine; The picture quality of the initial rendered image is processed by a second rendering engine to obtain a target rendered image; the picture quality parameter value of the target rendered image is higher than the picture quality parameter value of the initial rendered image; and the second rendering engine includes a neural network rendering model.
2. The method according to claim 1, characterized in that The picture quality parameters include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
3. The method according to claim 1 or 2, characterized in that: The initial rendered image is obtained based on the to-be-rendered picture data at the first moment; and the picture quality of the initial rendered image is processed by the second rendering engine to obtain the target rendered image, including: The initial rendered image and a rendered image obtained at a moment before the first moment are input into the second rendering engine to obtain the target rendered image.
4. The method according to any one of claims 1 to 3, characterized in that The method is performed by a first device; after obtaining the target rendered image, the method further includes: The target rendered image is transmitted to a second device for display.
5. The method according to claim 4, characterized in that The first device is a cloud server, and the second device is a terminal device.
6. The method according to any one of claims 1 to 5, characterized in that: The neural network rendering model is obtained by training based on multiple groups of training samples; each group of training samples in the multiple groups of training samples includes multiple training images; the multiple training images are obtained by image rendering based on training data at the same time; the picture quality parameter values of the multiple training images are different.
7. The method according to claim 6, characterized in that Each group of training samples in the multiple groups of training samples also includes training reference information generated based on the training data.
8. The method according to claim 7, characterized in that The training reference information includes part or all of the following: a depth map, a normal map, a camera position parameter, and a motion vector map.
9. The method according to any one of claims 6 to 8, characterized in that: The multiple groups of training samples are obtained based on training data at multiple consecutive moments.
10. The method according to any one of claims 1 to 9, characterized in that: Before obtaining the initial rendered image, the method further includes: Presenting a model selection interface; the model selection interface is used to display at least one trained neural network rendering model; When a model selection operation is received, a neural network rendering model specified by the model selection operation among the at least one trained neural network rendering model is loaded into the second rendering engine.
11. An image rendering device, characterized in that: The device comprises: A first rendering engine, used for performing basic image rendering based on the picture data to be rendered to obtain an initial rendered image; The second rendering engine is used to process the picture quality of the initial rendering image to obtain a target rendering image; the target rendering The picture quality parameter value of the image is higher than the picture quality parameter value of the initial rendered image, and the second rendering engine includes a neural network rendering model.
12. The device according to claim 11, characterized in that The picture quality parameters include some or all of the following: image resolution, anti-aliasing parameters, shadow effect parameters, and ray tracing effect parameters.
13. The device according to claim 11 or 12, characterized in that The initial rendering image is obtained based on the to-be-rendered picture data at the first moment; the second rendering engine is specifically used for: The initial rendered image and a rendered image obtained at a moment before the first moment are processed to obtain the target rendered image.
14. The device according to any one of claims 11 to 13, characterized in that The neural network rendering model is obtained by training based on multiple groups of training samples; each group of training samples in the multiple groups of training samples includes multiple training images; the multiple training images are obtained by image rendering based on training data at the same time; the picture quality parameter values of the multiple training images are different.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 10.
16. A computer-readable storage medium, characterized in that: The storage medium stores a computer program or an instruction. When the computer program or the instruction is executed by the communication device, the method according to any one of claims 1 to 10 is implemented.
17. A computer program product, characterized in that When the computer program product is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method, device and system for improving rendering efficiency based on deep learning
CN112419467A
Video rendering processing method and device, equipment and storage medium
CN115035230A
Rendering data acquisition method and electronic equipment
CN115546019A
Image processing method, electronic device and computer program product
CN115690525A
Image generation method and device, electronic equipment and computer readable medium
CN116664738A