A method and apparatus for accelerating rendering
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2026-08-11
AI Technical Summary
但蒙特卡洛求解同样依赖于有效的样本,有效的样本越多,渲染结果质量越高,但是时间也会越长
Smart Images

Figure CN116137047B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of graphics and image processing, and in particular to a method and apparatus for accelerating rendering. Background Technology
[0002] As the performance of terminal display devices improves, the industry's requirements for rendering resolution are increasing. Although the computing power of graphics processing units (GPUs) has continued to grow rapidly in recent years, the efficiency of obtaining high-resolution, high-quality rendering results is still insufficient.
[0003] Currently, to accelerate the rendering process, the Monte Carlo method is mainly used to solve the rendering equations in computer graphics (CG) content rendering. However, the Monte Carlo solution also depends on valid samples; the more valid samples, the higher the quality of the rendering result, but the longer it will take. For example, a single frame in some high-quality movies requires 1000 hours of rendering by a single central processing unit (CPU).
[0004] Therefore, improving the efficiency of obtaining high-resolution, high-quality rendering results has become an urgent problem to be solved in the industry. Summary of the Invention
[0005] This application provides a method and apparatus for accelerating rendering, so as to improve the efficiency of obtaining high-resolution, high-quality rendering results.
[0006] To achieve the above objectives, this application adopts the following technical solution:
[0007] In a first aspect, a method for accelerating rendering is provided, which may include: acquiring a first rendering result and a second rendering result of an image frame at a first resolution; wherein the first rendering result is the final rendering result and the second rendering result is an intermediate rendering result; the number of samples per pixel (SPP) of the first rendering result is greater than or equal to a first preset value; acquiring a first feature of the first rendering result and the second rendering result at a second resolution, wherein the first feature is a feature obtained by fusing the first rendering result and the second rendering result; acquiring a third rendering result of the image frame at a third resolution, wherein the third resolution is greater than the first resolution and the third rendering result is the final rendering result; acquiring a second feature of the third rendering result at a second resolution; acquiring a fourth rendering result of the image frame at a fourth resolution, wherein the fourth rendering result fuses the first feature and the second feature, and the pixels of the fourth rendering result converge to a second preset value.
[0008] The accelerated rendering method provided in this application, due to the high SPP rendering result at the first resolution, can ensure that the rendering output contains richer details and is closer to theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering result; the rendering result at the third resolution can supplement the information missing in the rendering result at the low resolution; therefore, the feature of the fusion of the rendering result at the first resolution and the third rendering result is finally used as the final rendering result. Under the premise of fast rendering speed, it ensures that the rendering output contains richer details and is closer to theoretical convergence, while also ensuring lower noise in the rendering result, thus achieving efficient acquisition of high-quality rendering results.
[0009] Furthermore, the above-mentioned solution in this application scales the rendering results at multiple scales to a second resolution, which can adaptively adjust the fusion size of inputs at different scales according to the computing power of the current device, significantly reducing computing power requirements and improving the applicability of the algorithm. The adaptive fusion network can better fuse inputs at different scales using deep learning models, achieving higher quality output.
[0010] In one possible implementation, the accelerated rendering method provided in this application may further include: obtaining a fifth rendering result of an image frame at a fifth resolution, wherein the fifth rendering result is an intermediate rendering result, and the SPP of the fifth rendering result is less than or equal to a third preset value; the fifth resolution is greater than the fourth resolution; obtaining a third feature of the fifth rendering result at a second resolution; correspondingly, in this implementation, obtaining a fourth rendering result of an image frame at a fourth resolution includes: performing feature fusion on the first feature, the second feature, and the third feature, followed by adaptive fusion to obtain the fourth rendering result. The above-mentioned solution of this application, by introducing multiple rendering targets and rendering inputs of different scales, can effectively supplement the high-frequency information loss existing in low-resolution upsampling to high-resolution, enabling the output result to retain more texture details. Simultaneously, the low-resolution high SPP input not only improves the noise reduction effect but also enhances the convergence of the final result, obtaining a higher quality high-resolution output; using a higher-resolution intermediate rendering result can effectively supplement high-quality geometric texture information, making the final result's detail richness closer to the native rendering result. By referencing higher resolution (fifth resolution) rendering results, high-frequency information such as geometric textures in the rendering results can be better supplemented, thereby improving the quality of the rendering results; moreover, the fifth rendering result is an intermediate rendering result, which is highly efficient to obtain, thus improving the efficiency of obtaining high-quality rendering results.
[0011] In another possible implementation, the third resolution is greater than or equal to the target resolution of the image frame. By using the third resolution branch, the high-frequency information loss that occurs when upsampling from low resolution to high resolution is effectively compensated for, allowing the output to retain more texture details.
[0012] In another possible implementation, the fourth resolution is the target resolution of the image frame, which allows for efficient acquisition of high-quality rendering results at the target resolution.
[0013] In another possible implementation, obtaining the fourth rendering result of the image frame at the fourth resolution includes: mapping the first and second features of the historical frames of the image frame to the image frame according to the motion vector, performing feature fusion with the first and second features of the image frame, and then performing adaptive fusion to obtain the fourth rendering result. The historical frames include one or more image frames preceding the image frame. The multi-frame fusion in the above-described scheme of this application can use the feature information of historical frames to correct errors in the sampling of the current image frame, while effectively supplementing the information of the current image frame. Multi-frame fusion can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0014] In another possible implementation, the accelerated rendering method provided in this application may further include: obtaining rendering parameters, including rendering quality and / or rendering duration; and determining a second resolution based on the rendering parameters. This method can adaptively adjust the fusion size of inputs at different scales according to the computing power of the current device, significantly reducing computing power requirements and expanding the applicability of the algorithm. The adaptive fusion network can better fuse inputs at different scales using deep learning models, achieving higher quality output.
[0015] In another possible implementation, the intermediate rendering results described in this application include: a normal map, or a depth map.
[0016] In another possible implementation, feature fusion is performed through a trained neural network, or adaptive fusion.
[0017] In another possible implementation, the solution provided in this application can be deployed in the cloud and provided to users as a cloud service.
[0018] In another possible implementation, the accelerated rendering method provided in this application may further include: obtaining rendering parameters input by the user, which may include one or more of a first resolution to a fifth resolution, or an upsampling factor, or a rendering duration, or a rendering quality.
[0019] Secondly, another method for accelerating rendering is provided, which may include: obtaining a sixth rendering result of an image frame at a sixth resolution, wherein the sixth rendering result is the final rendering result; wherein the SPP of the sixth rendering result is greater than or equal to a fourth preset value; obtaining sub-pixel information of the image frame, wherein the sub-pixel information is used to describe the characteristics of sub-pixels in the image frame, and the size of the sub-pixels is smaller than the size of the pixels; upsampling the sixth rendering result to a seventh resolution based on the sub-pixel information to obtain a seventh rendering result; wherein the seventh resolution is greater than the sixth resolution; and obtaining an eighth rendering result of the image frame based on the seventh rendering result, wherein the pixels of the eighth rendering result converge to a fifth preset value.
[0020] The accelerated rendering method provided in this application, due to the high SPP rendering results at sixth resolution, can ensure that the rendered output contains richer details and more closely approximates theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering results. By introducing sub-pixel information as upsampled pixel values, it better conforms to the rendering data distribution, increases the effective information of the input, and ultimately obtains high-resolution results with richer details that are more in line with the native rendered image. Therefore, the final rendering result, while maintaining high rendering speed, ensures that the rendered output contains richer details, more closely approximates theoretical convergence, and conforms to the rendering data distribution, thus achieving efficient acquisition of high-quality rendering results.
[0021] In one possible implementation, obtaining the eighth rendering result of the image frame based on the seventh rendering result includes: mapping the seventh rendering result of the historical frames of the image frame to the image frame according to motion vectors, performing feature fusion with the seventh rendering result of the image frame, and then performing adaptive fusion to obtain the eighth rendering result; wherein, the historical frames include one or more image frames preceding the image frame. The multi-frame fusion in the above-described scheme of this application can utilize the feature information of historical frames to correct errors in the sampling of the current image frame, while effectively supplementing the information of the current image frame. Multi-frame fusion can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0022] In one possible implementation, the accelerated rendering method provided in this application may further include: obtaining a ninth rendering result of an image frame at a sixth resolution, wherein the ninth rendering result is an intermediate rendering result; upsampling the ninth rendering result to a seventh resolution to obtain a tenth rendering result; and correspondingly, obtaining an eighth rendering result of the image frame based on the seventh rendering result, including: performing feature fusion on the seventh and tenth rendering results, and then extracting a fourth feature as the eighth rendering result. By introducing the intermediate rendering result at the sixth resolution, missing information is supplemented and the quality of the rendering result is improved without significantly increasing the rendering time.
[0023] In another possible implementation, the accelerated rendering method provided in this application may further include: obtaining a ninth rendering result of an image frame at a sixth resolution, wherein the ninth rendering result is an intermediate rendering result; upsampling the ninth rendering result to a seventh resolution to obtain a tenth rendering result; correspondingly, obtaining an eighth rendering result of the image frame based on the seventh rendering result, including: performing feature fusion on the seventh and tenth rendering results, extracting features to obtain a fourth feature; mapping the fourth feature of the historical frames of the image frame to the image frame according to motion vectors, and performing feature fusion with the fourth feature of the image frame, and then performing adaptive fusion to obtain the eighth rendering result. The historical frames include one or more image frames preceding the current image frame. The multi-frame fusion in the above-described scheme of this application can use the feature information of historical frames to correct errors in the sampling of the current image frame, while effectively supplementing the information of the current image frame. Multi-frame fusion can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0024] In another possible implementation, the intermediate rendering results described in this application include: a normal map, or a depth map.
[0025] In another possible implementation, the seventh resolution is the target resolution of the image frame.
[0026] In another possible implementation, feature fusion is performed through a trained neural network, or adaptive fusion.
[0027] In another possible implementation, the solution provided in this application can be deployed in the cloud and provided to users as a cloud service.
[0028] In another possible implementation, the accelerated rendering method provided in this application may further include: obtaining rendering parameters input by the user, which may include one or more of the sixth to tenth resolutions, or an upsampling factor, or a rendering duration, or a rendering quality.
[0029] Thirdly, an apparatus for accelerating rendering is provided, which may include: a first acquisition unit, a second acquisition unit, a third acquisition unit, and a fourth acquisition unit. Wherein:
[0030] The first acquisition unit is used to acquire a first rendering result and a second rendering result of an image frame at a first resolution; wherein the first rendering result is the final rendering result and the second rendering result is an intermediate rendering result; the SPP of the first rendering result is greater than or equal to a first preset value.
[0031] The second acquisition unit is used to acquire the first rendering result and the first feature of the second rendering result at the second resolution. The first feature is the feature after the first rendering result and the second rendering result are fused.
[0032] The first acquisition unit is further configured to acquire the third rendering result of the image frame at the third resolution. The third resolution is greater than the first resolution; the third rendering result is the final rendering result.
[0033] The third acquisition unit is used to acquire the second feature of the third rendering result at the second resolution.
[0034] The fourth acquisition unit is used to acquire the fourth rendering result of the image frame at the fourth resolution. The fourth rendering result integrates the first feature and the second feature, and the pixels of the fourth rendering result converge to the second preset value.
[0035] The accelerated rendering apparatus provided in this application, due to the high SPP rendering result at the first resolution, can ensure that the rendering output contains richer details and is closer to theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering result; the rendering result at the third resolution can supplement the information missing in the rendering result at the low resolution; therefore, the feature of the fusion of the rendering result at the first resolution and the third rendering result is finally used as the final rendering result, which, while maintaining high rendering speed, ensures that the rendering output contains richer details and is closer to theoretical convergence, and also ensures lower noise in the rendering result, thus achieving efficient acquisition of high-quality rendering results.
[0036] In one possible implementation, the first acquisition unit is further configured to acquire a fifth rendering result of the image frame at a fifth resolution, wherein the fifth rendering result is an intermediate rendering result, and the SPP of the fifth rendering result is less than or equal to a third preset value; the fifth resolution is greater than the fourth resolution. The accelerated rendering apparatus may further include a fifth acquisition unit configured to acquire a third feature of the fifth rendering result at a second resolution. Correspondingly, the fourth acquisition unit may specifically be configured to: perform feature fusion on the first feature, second feature, and third feature, and then perform adaptive fusion to obtain the fourth rendering result. The above-described solution of this application, by introducing multiple rendering targets and rendering inputs of different scales, can effectively supplement the high-frequency information loss existing in low-resolution upsampling to high-resolution, enabling the output result to retain more texture details. Simultaneously, the low-resolution, high-SPP input not only improves the noise reduction effect but also enhances the convergence of the final result, obtaining a higher-quality high-resolution output; using a higher-resolution intermediate rendering result can effectively supplement high-quality geometric texture information, making the final result's detail richness closer to the native rendering result. By referencing higher resolution (fifth resolution) rendering results, high-frequency information such as geometric textures in the rendering results can be better supplemented, thereby improving the quality of the rendering results; moreover, the fifth rendering result is an intermediate rendering result, which is highly efficient to obtain, thus improving the efficiency of obtaining high-quality rendering results.
[0037] In another possible implementation, the third resolution is greater than or equal to the target resolution of the image frame. By using the third resolution branch, the high-frequency information loss that occurs when upsampling from low resolution to high resolution is effectively compensated for, allowing the output to retain more texture details.
[0038] In another possible implementation, the fourth resolution is the target resolution of the image frame, which allows for efficient acquisition of high-quality rendering results at the target resolution.
[0039] In another possible implementation, the first and second features of historical frames of the image frame are mapped to the image frame according to motion vectors, and then fused with the first and second features of the image frame. Adaptive fusion is then performed to obtain the fourth rendering result. The historical frames include one or more image frames preceding the current image frame. The multi-frame fusion scheme in this application can utilize the feature information of historical frames to correct errors in the sampling of the current image frame, while effectively supplementing the information of the current image frame. Multi-frame fusion ensures the temporal stability of the rendering result and reduces problems such as temporal aliasing.
[0040] In another possible implementation, the accelerated rendering device may further include a sixth acquisition unit and a determination unit. The sixth acquisition unit acquires rendering parameters, including rendering quality and / or rendering duration. The determination unit determines a second resolution based on the rendering parameters acquired by the sixth acquisition unit. This adaptive fusion network can adjust the fusion size of inputs at different scales according to the computing power of the current device, significantly reducing computing power requirements and expanding the applicability of the algorithm. The adaptive fusion network can better fuse inputs at different scales using deep learning models, achieving higher quality output.
[0041] In another possible implementation, the intermediate rendering results described in this application include: a normal map, or a depth map.
[0042] In another possible implementation, feature fusion is performed through a trained neural network, or adaptive fusion.
[0043] In another possible implementation, the device for accelerating rendering can be deployed in the cloud and offered to users as a cloud service.
[0044] It should be noted that the accelerated rendering apparatus provided in the third aspect, used to execute the method described in the first aspect or any possible implementation thereof, can achieve the same effect as the scheme described in the first aspect, and its specific implementation will not be elaborated one by one.
[0045] Fourthly, another apparatus for accelerating rendering is provided, which may include: a first acquisition unit, a second acquisition unit, a magnification unit, and a third acquisition unit. Wherein:
[0046] The first acquisition unit is used to acquire the sixth rendering result of the image frame at the sixth resolution. The sixth rendering result is the final rendering result. The SPP of the sixth rendering result is greater than or equal to the fourth preset value.
[0047] The second acquisition unit is used to acquire sub-pixel information of the image frame. The sub-pixel information is used to describe the characteristics of sub-pixels in the image frame, and the size of the sub-pixel is smaller than the size of the pixel.
[0048] The magnification unit is used to upsample the sixth rendering result to a seventh resolution based on the sub-pixel information acquired by the second acquisition unit, thereby obtaining a seventh rendering result. The seventh resolution is greater than the sixth resolution.
[0049] The third acquisition unit is used to acquire the eighth rendering result of the image frame based on the seventh rendering result obtained by the magnification unit, and the pixels of the eighth rendering result converge to the fifth preset value.
[0050] The accelerated rendering apparatus provided in this application, due to its high SPP rendering results at sixth resolution, ensures that the rendered output contains richer details and more closely approximates theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering results. By introducing sub-pixel information as upsampled pixel values, it better conforms to the rendering data distribution, increases the effective information of the input, and ultimately obtains high-resolution results with richer details that are more consistent with the native rendered image. Therefore, the final rendering result, while maintaining high rendering speed, ensures that the rendered output contains richer details, more closely approximates theoretical convergence, and conforms to the rendering data distribution, thus achieving efficient acquisition of high-quality rendering results.
[0051] In one possible implementation, the third acquisition unit can specifically be used to: map the seventh rendering result of the historical frames of the image frame to the image frame according to the motion vector, perform feature fusion with the seventh rendering result of the image frame, and then perform adaptive fusion to obtain the eighth rendering result. Here, the historical frames include one or more image frames preceding the image frame. The multi-frame fusion in the above-described scheme of this application can utilize the feature information of historical frames to correct errors in the sampling of the current image frame, while effectively supplementing the information of the current image frame. Multi-frame fusion can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0052] In one possible implementation, the first acquisition unit can also be used to: acquire the ninth rendering result of the image frame at the sixth resolution, where the ninth rendering result is an intermediate rendering result. The magnification unit can also be used to: upsample the ninth rendering result to the seventh resolution to obtain the tenth rendering result. The third acquisition unit can specifically be used to: perform feature fusion on the seventh and tenth rendering results, and then extract features to obtain a fourth feature, which serves as the eighth rendering result. By introducing the intermediate rendering result at the sixth resolution, missing information is supplemented, and the quality of the rendering result is improved without significantly increasing the rendering time.
[0053] In one possible implementation, the first acquisition unit can also be used to: acquire the ninth rendering result of the image frame at the sixth resolution, where the ninth rendering result is an intermediate rendering result. The magnification unit can also be used to: upsample the ninth rendering result to the seventh resolution to obtain the tenth rendering result. The third acquisition unit can specifically be used to: perform feature fusion on the seventh and tenth rendering results, extract features to obtain the fourth feature, map the fourth feature of the historical frames of the image frame to the image frame according to the motion vector, and perform feature fusion with the fourth feature of the image frame, and then perform adaptive fusion to obtain the eighth rendering result; wherein, the historical frames include one or more image frames before the image frame. The multi-frame fusion in the above scheme of this application can use the feature information of historical frames to correct certain errors in the sampling of the current image frame, and at the same time effectively supplement the information of the current image frame. Multi-frame fusion can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0054] In another possible implementation, the intermediate rendering results described in this application include: a normal map, or a depth map.
[0055] In another possible implementation, the seventh resolution is the target resolution of the image frame.
[0056] In another possible implementation, feature fusion is performed through a trained neural network, or adaptive fusion.
[0057] In another possible implementation, the rendering group can be deployed in the cloud and offered to users as a cloud service.
[0058] It should be noted that the accelerated rendering apparatus provided in the fourth aspect, used to execute the method described in the second aspect or any possible implementation thereof, can achieve the same effect as the scheme described in the second aspect, and its specific implementation will not be elaborated one by one.
[0059] Fifthly, this application provides a rendering device that can implement the functions described in the method examples of the first or second aspect or any possible implementation thereof. These functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions. This image registration device can exist in the form of a chip.
[0060] In one possible implementation, the rendering device includes a processor and a transceiver. The processor is configured to support the rendering device in performing the corresponding functions described in the above methods. The transceiver supports communication between the rendering device and other devices. The rendering device may also include a memory coupled to the processor, which stores necessary program instructions and data for the rendering device.
[0061] In a sixth aspect, a computer-readable storage medium is provided, including instructions that, when executed on a computer, cause the computer to perform the accelerated rendering method provided in the first aspect or any possible implementation thereof.
[0062] In a seventh aspect, a computer program product containing instructions is provided that, when run on a computer, causes the computer to perform the accelerated rendering method provided in the first aspect or any possible implementation thereof.
[0063] Eighthly, this application provides a chip system including a processor and potentially a memory for implementing the corresponding functions described in the above methods. The chip system may be composed of chips or may include chips and other discrete devices.
[0064] It should be noted that any of the possible implementations of any of the above aspects can be combined, provided that the solutions do not contradict each other. Attached Figure Description
[0065] Figure 1 This application provides a schematic diagram of the architecture of a rendering system.
[0066] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0067] Figure 3 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;
[0068] Figure 4 A schematic diagram of the overall process of a method for accelerating rendering provided in an embodiment of this application;
[0069] Figure 5 A flowchart illustrating a method for accelerating rendering provided in an embodiment of this application;
[0070] Figure 6 A schematic diagram illustrating the principle of upsampling provided in an embodiment of this application;
[0071] Figure 7 A schematic diagram illustrating the principle of downsampling provided in an embodiment of this application;
[0072] Figure 8 A schematic diagram illustrating a scenario of fusing current image frames and historical frames, provided as an embodiment of this application;
[0073] Figure 9 This is a schematic diagram of a rendering result transformation scenario provided in an embodiment of this application;
[0074] Figure 10 A flowchart illustrating another method for accelerating rendering provided in an embodiment of this application;
[0075] Figure 11 A flowchart illustrating another method for accelerating rendering provided in an embodiment of this application;
[0076] Figure 12 A schematic flowchart illustrating another method for accelerating rendering provided in an embodiment of this application;
[0077] Figure 13 A schematic flowchart illustrating another method for accelerating rendering provided in an embodiment of this application;
[0078] Figure 14 A flowchart illustrating the accelerated rendering method combining historical frames provided in an embodiment of this application;
[0079] Figure 15 A schematic diagram of the structure of an apparatus for accelerating rendering provided in an embodiment of this application;
[0080] Figure 16 A schematic diagram of another apparatus for accelerating rendering provided in an embodiment of this application;
[0081] Figure 17 A schematic diagram of the structure of another apparatus for accelerating rendering provided in an embodiment of this application;
[0082] Figure 18 A schematic diagram of the structure of another device for accelerating rendering provided in an embodiment of this application;
[0083] Figure 19 This is a schematic diagram of another device for accelerating rendering provided in an embodiment of this application. Detailed Implementation
[0084] The terms "comprising" and "having," and any variations thereof, used in the description of the embodiments of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0085] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0086] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. "And / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone.
[0087] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first resolution," "second resolution," etc., are used to distinguish different resolution values, not to describe the size of the resolution.
[0088] To facilitate understanding, the relevant terms and concepts involved in the embodiments of this application will be introduced below.
[0089] Image rendering refers to the process by which a computer converts shapes stored in memory into their corresponding representations drawn on the screen. Scenes and entities are represented in three dimensions, which is closer to the real world and easier to manipulate and transform. However, most graphics display devices are two-dimensional rasterized displays and pixelated printers. Therefore, image rendering is the process of converting three-dimensional light energy transfer processing into a two-dimensional image. The representation of a three-dimensional scene from N-dimensional raster and pixelated representation is image rendering—or rasterization. For example, the image rendering process may include: converting three-dimensional (3D) coordinates to two-dimensional (2D) coordinates, and then converting the 2D coordinates into actual colored pixels.
[0090] Resolution: Determines the level of detail in an image. Generally, the higher the image resolution, the more pixels it contains, and the clearer the image.
[0091] Target resolution: This refers to the desired final resolution during image rendering. For example, the target resolution can be determined by the display capabilities of the display device, or by the purpose of image rendering.
[0092] Subpixel: A unit smaller than a pixel. A pixel is the basic unit of an image, often referred to as the physical resolution of the image. Subpixels are a further refinement of pixels at the algorithmic level.
[0093] Upsampling, also known as image upsampling or image interpolation, primarily aims to enlarge the original image so that it can be displayed on higher resolution display devices. For example, upsampling can employ interpolation methods, which insert new elements between pixels using a suitable interpolation algorithm based on the existing image pixels.
[0094] Subsampling, also known as image downsampling, primarily aims to make an image fit the size of the display area or generate a thumbnail of the corresponding image. For example, for an image I of size M*N, downsampling it by a factor of s yields a (M / s)*(N / s) resolution image. If considering a matrix image, it transforms the image within an s*s window of the original image into a single pixel, where the value of this pixel is the average of all pixels within the window.
[0095] Whether it is upsampling or downsampling, there are many sampling methods, such as nearest neighbor interpolation, bilinear interpolation, mean interpolation, median interpolation, etc. This application does not limit these methods.
[0096] Depth: Used to represent the distance of each pixel from the observer. In some examples, the relationship between the distance of an object from the user and the depth value can be: the closer the object is to the user, the smaller the depth value of the object's pixels; the farther the object is from the user, the larger the depth value of the object's pixels. Alternatively, the closer the object is to the user, the larger the depth value of the object's pixels; the farther the object is from the user, the smaller the depth value of the object's pixels. The relationship between depth and the distance of an object from the observer is not limited in the embodiments of this application.
[0097] The Render Equation is an integral equation. In computer graphics, the goal of realistic rendering is to solve this equation. It is the theoretical foundation of all global illumination methods (ray tracing, path tracing, radiosity, etc.), and its expression is as follows:
[0098]
[0099] Monte Carlo method (or Monte Carlo integral): This is a general term for a class of algorithms that solve problems by using random sampling. The problem sought is the probability of a random event or the expected value of a random variable. The Monte Carlo method is used to estimate integrals, and it plays a very important role in graphics rendering. Its expression is as follows:
[0100]
[0101] Image features: Information used to represent the characteristics of an image. Image features possess properties such as repeatability, distinguishability, centrality, and efficiency, while also being able to cope with the effects of changes in image brightness, scale, rotation, and affine transformations. For example, in computer vision, corners are often used as image features. Of course, this application does not limit the specific content and form of image features. Image features can be extracted using models.
[0102] Feature fusion refers to merging the features of multiple images to obtain the features of a single image, which is then used as the features of the fused image. Feature fusion can involve calculating mathematical values for the features or other methods, which are not limited in the embodiments of this application. The mathematical values can be superposition or other operations.
[0103] Adaptive fusion refers to the adaptive fusion of different features from multiple images according to rules to obtain the features of a single image. Different features can refer to features that are partially present in the images involved in the fusion but not in others. For example, features can be assigned a full red color, and adaptive fusion can be performed according to the weights of different features. Of course, in practical applications, the specific adaptive fusion scheme can be configured according to actual needs, and this application does not limit this.
[0104] The final rendering result refers to the two-dimensional image obtained after the rendering process is completed. The final rendering result can be directly presented to the user.
[0105] Intermediate rendering results: These refer to intermediate images during the rendering process, which can be images in a single dimension. For example, intermediate rendering results can be normal maps, depth maps, or other types of images. Intermediate rendering results are also known as gbuffer information. gbuffer information can include albedo, normal, position, etc.
[0106] Since the embodiments of this application involve the application of neural networks, for ease of understanding, the relevant terms and concepts such as neural networks involved in the embodiments of this application will be introduced below.
[0107] (1) Neural network (NN)
[0108] A neural network is a machine learning model, a machine learning technique that simulates the neural network of the human brain to achieve artificial intelligence-like capabilities. The input and output of a neural network can be configured according to actual needs, and the network can be trained using sample data to minimize the error between its output and the actual output corresponding to the sample data. A neural network can be composed of neural units, which can refer to... The arithmetic unit that takes an intercept of 1 as input can output the following:
[0109] (Equation 1)
[0110] Where s = 1, 2, ..., n, n is a natural number greater than 1. for The weight, 'f' represents the bias of the neural unit. 'f' represents the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together; that is, the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from the local receptive field, which can be a region composed of several neural units.
[0111] (2) Deep Neural Networks
[0112] Deep neural networks (DNNs), also known as multilayer neural networks, can be understood as neural networks with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three layers based on their position: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs appear complex, the operation of each layer is actually quite simple, resembling a linear relationship as follows: ,in, It is the input vector. It is the output vector. It is an offset vector. It is the weight matrix (also called coefficients). It's an activation function. Each layer simply applies the input vector... The output vector is obtained through such a simple operation. Because DNNs have many layers, the coefficients... and offset vector The number of these parameters is therefore quite large. These parameters are defined in DNNs as follows: [as coefficients] For example: Suppose in a three-layer DNN, the linear coefficient from the fourth neuron in the second layer to the second neuron in the third layer is defined as... The superscript 3 represents the coefficient. The index corresponds to the level number, specifically the output index 2 for the third level and the input index 4 for the second level. In summary: Level L... The coefficients from the k-th neuron in layer 1 to the j-th neuron in layer L are defined as follows: It's important to note that the input layer does not have... Parameters. In deep neural networks, more hidden layers allow the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can accomplish more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers in the trained deep neural network (a weight matrix formed by vectors W from many layers).
[0113] (3) Convolutional Neural Network
[0114] A CNN (Convolutional Neural Network) is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as using a trainable filter to convolve with an input image or a convolutional feature map. A convolutional layer is a layer of neurons in a CNN that performs convolution processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to some of its neighboring neurons. A convolutional layer typically contains several feature maps, each composed of rectangularly arranged neural units. Neural units within the same feature map share weights, which are the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The underlying principle is that the statistical information of one part of an image is the same as that of other parts. This means that image information learned in one part can also be used in another part. Therefore, the same learned image information can be used for all locations in the image. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0115] Convolutional kernels can be initialized as matrices of random size, and during the training of a convolutional neural network, they can learn appropriate weights. Furthermore, sharing weights directly reduces the number of connections between layers in the convolutional neural network, while also lowering the risk of overfitting.
[0116] (4) Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, the layers are fully connected from the input layer to the hidden layer and then to the output layer, but the nodes within each layer are unconnected. While this type of ordinary neural network has solved many difficult problems, it is still powerless against many others. For example, to predict the next word in a sentence, you generally need to use the preceding words because the words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is related to the previous output. Specifically, the network remembers previous information and applies it to the calculation of the current output. That is, the nodes within the hidden layer are no longer unconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time step. Theoretically, RNNs can process sequential data of any length. Training RNNs is the same as training traditional CNNs or DNNs. It also uses the backpropagation algorithm, but with one key difference: if an RNN is expanded, its parameters, such as W, are shared; however, this is not the case with traditional neural networks as illustrated above. Furthermore, in the gradient descent algorithm, the output at each step depends not only on the network at the current step but also on the states of the network in the previous several steps. This learning algorithm is called Backpropagation Through Time (BPTT).
[0117] Since we already have convolutional neural networks (CNNs), why do we need recurrent neural networks (RNNs)? The reason is simple. CNNs rely on the fundamental assumption that elements are independent of each other, and that input and output are also independent—like a cat and a dog. However, in the real world, many elements are interconnected. For example, stock prices fluctuate over time. Or, imagine someone saying, "I love traveling, and my favorite place is Yunnan. I definitely want to go there someday." Humans know the answer to this question is "Yunnan." Humans can infer from context, but how can machines do the same? This is where RNNs come in. RNNs aim to give machines the ability to remember, just like humans. Therefore, the output of an RNN depends on both the current input information and historical memory information.
[0118] (6) Loss function
[0119] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value. Based on the difference, we update the weight vector of each layer (usually pre-configuring parameters before the initial update). For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network predicts the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss.
[0120] As mentioned earlier, with the improvement of terminal display device performance, the industry's requirements for rendering result resolution are increasing. Increased rendering resolution inevitably leads to longer rendering times. High-resolution rendering results are also known as high-quality rendering results.
[0121] Currently, in order to save computing power and improve rendering efficiency to obtain high-quality rendering results, the industry generally adopts two different strategies for image rendering: to obtain high-quality rendering results while minimizing rendering time, thereby improving rendering efficiency.
[0122] Strategy 1: High-resolution low SPP rendering combined with high-quality Monte Carlo denoising method. This type of method has problems with insufficient convergence and blurry results.
[0123] Strategy 2: High SPP rendering at low resolution combined with high-quality supersampling methods. This type of method will also result in blurring issues.
[0124] The following is a brief explanation of the current mainstream image acceleration rendering technologies in the industry.
[0125] One of the image acceleration rendering technologies is NVIDIA ®The newly released Optix denoise technology is a tool for resolving noise in low-sample-count rendered images using neural network methods. By employing GPU-accelerated deep learning algorithms, it significantly reduces the time required to render high-quality images. The rendered images produced by this approach are visually noise-free. The algorithm enables real-time interaction within the rendering engine, providing ultra-fast interactive feedback and enhancing the user experience. While this algorithm offers fast denoising and is geared towards real-time applications, its denoising effect is generally limited. Although it improves rendering efficiency, the rendering quality is not particularly high.
[0126] Another image acceleration rendering technique is Intel® Open Image Denoise, a high-performance, high-quality open-source library for ray-traced image denoising. This technology includes a series of deep learning-based denoising filters, enabling high-quality, high-efficiency Monte Carlo denoising and significantly reducing the number of SPPs required during rendering. The algorithm can achieve rendering denoising with varying SPP numbers, ranging from 1 SPP to near-convergence. This algorithm outperforms Optic Denoise in denoising quality. However, it focuses on denoising in real-time scenes and takes longer than Optic Denoise; while achieving high-quality rendering, its rendering efficiency is lower.
[0127] Another image acceleration rendering technique combines contrast adaptive sharpening (CAS) with upsampling to achieve high-resolution rendering results from low-resolution, high-quality results. The CAS algorithm adaptively adjusts the sharpening level based on contrast, enhancing internal object details while preserving most high-contrast edges. Since upsampling inevitably results in the loss of high-frequency information, causing blurring of the output image, combining CAS with upsampling enhances the high-frequency information in the upsampled result, achieving a high-quality upsampling effect. Building upon CAS, FidelityFX Super Resolution (FSR) employs advanced optimization and upgrade technologies to improve frame rates without requiring users to upgrade their graphics cards, delivering a high-quality, high-resolution gaming experience. In FSR's "Performance" mode, some games can achieve up to 2.5 times the performance at 4K resolution, achieving ultra-high-quality edges and details, providing a gaming experience close to the original resolution. To achieve higher performance on a 4K screen, the CAS or FSR algorithms mentioned above can be used. Running them at 1800p or 80% resolution produces results close to the native image with almost no loss of quality. However, for higher resolutions, the loss of quality becomes significantly greater. The CAS and FSR algorithms are still limited to an upsampling factor of less than 2x, which is a significant shortcoming compared to other super-resolution algorithms, affecting rendering quality.
[0128] Another image rendering technique is the mature upsampling scheme in the CG field. This scheme is based on temporal anti-aliasing (TAA) upsampling. The core of this algorithm is to increase the output resolution by adjusting the position of the sampling points and the magnification factor while keeping the input resolution constant, thereby improving the rendering quality. Because it can guarantee a relatively fast upsampling speed and a certain level of effect, it is currently used in commercial engines. However, this algorithm has the problem of requiring fine-tuning of algorithm parameters manually, which may lead to blurry upsampling results, ghosting, temporal delays, jitter, and other issues, resulting in lower rendering quality.
[0129] Based on this, embodiments of this application provide a method for accelerating rendering. This method uses multi-scale rendering resolution as input and fuses feature information from different scales to generate high-resolution, high-quality rendering results with good convergence, thereby achieving efficient and high-quality image rendering. Alternatively, it uses low-resolution rendering results combined with sub-pixel information upsampling to accelerate image rendering and achieve efficient high-resolution image rendering.
[0130] The accelerated rendering method provided in this application embodiment can be applied to... Figure 1 The rendering system shown. For example... Figure 1 As shown, the rendering system includes a rendering engine 101, a rendering acceleration device 102, and a display device 103.
[0131] The rendering engine 101 is used to obtain the final rendering result and / or intermediate rendering result according to the resolution, and the rendering acceleration device 102 is used to execute the scheme provided in this application, and efficiently obtain the high-quality final rendering result according to the output of the rendering engine 101, and display the rendering result by the display device 103.
[0132] For example, the accelerated rendering device 102 can be deployed in the cloud. The display device 103 requests the accelerated rendering device 102 in the cloud to execute the solution provided in this application to efficiently obtain high-quality rendering results, which are then displayed by the display device 103. Of course, the accelerated rendering device 102 can be deployed in other locations on the network, and this embodiment of the application does not limit this.
[0133] It should be noted that the rendering engine 101 and the rendering acceleration device 102 can be centrally deployed as a single device or separately deployed. This application embodiment does not impose specific limitations on the architecture deployment of the rendering system.
[0134] For example, the accelerated rendering method provided in this application can be used in the following scenarios:
[0135] Application Scenario 1: Real-time supersampling in games. In a range of high-quality game applications, including AAA games, the accelerated rendering method provided in this application can be used to first reduce the resolution of the rendered image using supersampling technology. Then, the supersampling algorithm can quickly upsample the image while maintaining image quality. The upsampling time accounts for a relatively small percentage of the rendering time, which can ensure image quality without affecting the rendering speed of the game screen, thereby improving the rendering frame rate and significantly enhancing game quality and the player's gaming experience.
[0136] Application Scenario 2: High-quality CG content production (e.g., movies, animations, promotional videos). Existing solutions for generating high-quality CG content involve setting very high SPP values to achieve near-convergence and extremely low noise levels in the final rendering result. However, this process requires long rendering times and extremely high computing power. The solution provided in this application enables efficient and rapid generation of high-quality content, significantly reducing computing power requirements, lowering development costs, and shortening the development cycle.
[0137] Application Scenario 3: Interior Design. In the field of interior design, designers need to present high-quality renderings of their designs in real time based on user needs. This allows for timely adjustments and modifications to the plans, accelerating the design process and increasing user participation and the quality of the final design. After completing the overall design, designers may also need to convert the design plans into corresponding high-quality rendered videos for users to view. The solution provided in this application can accelerate the generation of high-quality demonstration videos, shorten the design cycle, and ultimately improve the user experience.
[0138] It should be noted that the above application scenarios are merely illustrative examples and are not intended to limit the application scenarios of the proposed solution.
[0139] For example, the display device 103 described above can be an electronic device. Specifically, the electronic device can be a large-screen display device, mobile phone, laptop computer, tablet computer, in-vehicle device, wearable device (such as smartwatch), ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), artificial intelligence device, and other terminal devices with display functions. This application embodiment does not limit the specific type of electronic device.
[0140] In this application, the structure of the electronic device can be as follows: Figure 2 As shown. Figure 2 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0141] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0142] The processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. For example, in this application, the processor 110 may acquire a first playback speed, optionally related to user playback settings; acquire first information, including image information and / or audio information of the video; and obtain a second playback speed based on the first playback speed and the first information.
[0143] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0144] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0145] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0146] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.
[0147] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0148] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0149] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0150] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0151] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0152] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0153] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0154] The display screen 194 of the electronic device 100 can display a series of graphical user interfaces (GUIs), which serve as the main screen of the electronic device 100. Generally, the size of the display screen 194 of the electronic device 100 is fixed, and only a limited number of controls can be displayed on the display screen 194. A control is a GUI element, a software component contained in an application, that controls all the data processed by the application and the interactive operations related to that data. Users can interact with controls through direct manipulation, thereby reading or editing information related to the application. Generally, controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0155] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0156] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0157] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0158] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0159] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0160] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0161] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0162] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. For example, in this embodiment, processor 110 can execute the video playback method provided in this application by executing the instructions stored in internal memory 121 to obtain the playback speed of the video played by electronic device 100. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phone book, etc.). In addition, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor.
[0163] Electronic device 100 can implement audio functions through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor. Examples include voice playback in video, music playback, and recording.
[0164] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0165] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0166] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0167] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0168] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface, a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface, or another.
[0169] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0170] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios. The gyroscope sensor 180B can also determine whether the electronic device 100 is in a moving state.
[0171] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0172] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0173] The accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in various directions (generally three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device, applied to applications such as screen orientation switching and pedometers. It can also be used to determine whether electronic device 100 is in motion.
[0174] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0175] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.
[0176] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.
[0177] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0178] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0179] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0180] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.
[0181] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0182] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0183] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0184] In addition, an operating system runs on top of these components. Examples include Apple's iOS, Google's Android, and Microsoft's Windows. Applications can be installed and run on this operating system.
[0185] The operating system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, cloud architecture, or other architectures. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.
[0186] Figure 3 This is a software structure block diagram of the electronic device 100 according to an embodiment of this application.
[0187] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0188] The application layer can include a series of application packages. For example... Figure 3 As shown, the application package can include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS. For example, when taking a photo, the camera application can access the camera interface management service provided by the application framework layer.
[0189] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 3As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0190] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0191] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0192] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0193] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0194] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0195] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0196] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0197] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0198] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0199] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0200] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0201] The media library supports playback and recording of various commonly used audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as: Moving Picture Experts Group (MPEG) 4, H.264, MP3, Advanced Audio Coding (AAC), Adaptive Multi Rate (AMR), Joint Photographic Experts Group (JPEG), and Portable Network Graphic Format (PNG).
[0202] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0203] A two-dimensional (2D) graphics engine is a graphics engine for 2D drawing.
[0204] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0205] It should be noted that although the embodiments in this application are described using the Android® system as an example, the basic principles are equally applicable to systems based on iOS. ® or Windows ® Electronic devices with operating systems, etc.
[0206] On the one hand, embodiments of this application provide a method for accelerating rendering, which can be executed by an apparatus for accelerating rendering (an apparatus that performs the method for accelerating rendering provided in this application) to efficiently obtain high-quality rendered images.
[0207] The device for accelerating rendering can first obtain rendering parameters and then execute the following method for accelerating rendering according to the obtained rendering parameters.
[0208] In one possible implementation, the device for accelerating rendering can obtain rendering parameters input by the user through an interactive interface. These rendering parameters may include at least one of the following: any resolution value, rendering quality, or rendering time in the embodiments described below.
[0209] In another possible implementation, the above rendering parameters can also be configured fixed parameters, which do not require user input. The device that accelerates rendering can obtain the rendering parameters by reading the fixed parameters.
[0210] This application does not limit the method of obtaining rendering parameters in its embodiments.
[0211] The overall flow of the accelerated rendering method provided in this application embodiment is as follows: Figure 4 As shown, the user first sets the corresponding rendering parameters in the rendering engine's interactive interface, which may include, but are not limited to, one or more of the following: upsampling factor, number of SPPs, different resolutions, output type, etc. After the user completes the settings, the rendering engine renders according to the user's settings and outputs the final rendering results at one or more resolutions, as well as the required intermediate rendering results. These outputs are then used as input information to the accelerated rendering device provided in this application. The accelerated rendering device provided in this application executes the scheme provided in this application to complete model inference in a very short time and obtain the final high-quality rendering result at high resolution. The output results of the accelerated rendering device are then processed through a series of post-processing algorithms and finally displayed on the monitor or stored for later use.
[0212] The flow of the accelerated rendering method provided in this application embodiment can be as follows: Figure 5 As shown, it may include:
[0213] S501, the device for accelerating rendering acquires a first rendering result and a second rendering result of an image frame at a first resolution.
[0214] The first resolution can be a preset value or a user-input value lower than the target rendering resolution. The target rendering resolution refers to the final resolution of the rendered image.
[0215] Optionally, the first resolution can be a fixed value, or the first resolution can be a resolution calculated based on the expected rendering quality or the expected rendering duration. This application does not limit the specific value of the first resolution or the method of obtaining it.
[0216] For example, a relationship between the desired rendering quality or desired rendering duration and the first resolution is pre-configured. The device for accelerating rendering in S501 can determine the first resolution based on this relationship. The specific content of this relationship is not described in the embodiments of this application, but can be configured according to actual needs.
[0217] The first rendering result is the final rendering result. The second rendering result is an intermediate rendering result.
[0218] Specifically, the SPP of the first rendering result is greater than or equal to the first preset value.
[0219] For example, the SPP of the first rendering result can be 1024 SPP. A higher SPP ensures that the rendered output contains richer details and is closer to the theoretical convergence value.
[0220] The first preset value can be configured according to actual needs, and this application embodiment does not limit it.
[0221] For example, the first preset value can be the lowest SPP threshold value for high-quality images provided by the industry or enterprise. The first rendering result is a high-quality rendering result.
[0222] Furthermore, the selection of the SPP value when obtaining the first rendering result can be implemented according to actual needs, as long as it is greater than or equal to the first preset value. This application embodiment does not specifically limit the SPP value when obtaining the first rendering result.
[0223] In one possible implementation, the device for accelerating rendering in S501 can use the intermediate rendering result as the second rendering result during the process of obtaining the first rendering result.
[0224] In another possible implementation, the device for accelerating rendering in S501 can render in parallel to obtain the first rendering result and the second rendering result.
[0225] S502, the device for accelerating rendering acquires a first rendering result and a second rendering result, and a first feature at a second resolution.
[0226] The second resolution is used for scaling, which scales information from different scales to the same size. The second resolution can be a fixed parameter, a content from the rendering parameters obtained by the aforementioned accelerated rendering device, or a resolution determined based on the content of the rendering parameters obtained by the aforementioned accelerated rendering device.
[0227] In one possible implementation, the choice of scaling size depends on limitations on the computational load and time of the algorithm; therefore, the second resolution can be determined based on rendering parameters. For example, the accelerated rendering method provided in this application embodiment may include: the accelerated rendering apparatus acquiring rendering parameters, including rendering quality and / or rendering duration; and the accelerated rendering apparatus determining a second resolution based on the acquired rendering parameters.
[0228] For example, when the algorithm needs to run with lower computational cost and faster computation time, a lower second resolution can be determined to downsample the input of medium and high resolution branches to the second resolution, thereby reducing computational cost and speeding up computation time. When computational resources are sufficient, a higher second resolution can be determined to upsample the input of low and medium resolution branches to the second resolution, thereby obtaining more accurate output results.
[0229] In another possible implementation, the second resolution can be the target resolution of the image frame.
[0230] The first feature is the feature resulting from the fusion of the first rendering result and the second rendering result.
[0231] For example, the device for accelerating rendering in S502 can extract features from the first rendering result and the second rendering result respectively, then perform feature fusion, and then scale the fused features to the second resolution to obtain the first feature.
[0232] The device for accelerating rendering can extract features from the rendering result using a feature extraction model. This application does not limit the specific implementation of feature extraction. For example, features can be extracted using a residual convolution model or an attention model.
[0233] Optionally, the feature extraction model is used to extract feature information of the rendering results under the corresponding path. In order to reduce the number of parameters, the feature extraction models under different paths can share weights.
[0234] Furthermore, in embodiments of this application, the device for accelerating rendering can perform feature fusion using a trained neural network.
[0235] The type of neural network used for feature fusion and the specific results are not specifically limited in this application embodiment, and can be configured according to actual needs.
[0236] For example, the upsampling used in the embodiments of this application may include, but is not limited to, sub-pixel upsampling, transposed convolution, bilinear upsampling, or others. The sub-pixel upsampling module uses the opposite operation of the deshuffle module to transform channel-dimensional information into spatial dimension.
[0237] like Figure 6 The diagram illustrates the principle of sub-pixel upsampling. Figure 6 As shown, a 2 An image of size C×H×W is upsampled to an image of size C×aH×aW, where a represents the upsampling factor, C represents the feature map data, H represents the height of the feature map, and W represents the width of the feature map.
[0238] For example, the downsampling methods used in the embodiments of this application include, but are not limited to, deshuffle, bilinear, bicubic, and maxpooling. Among them, deshuffle downsampling transforms spatially distributed feature information to the channel level in a lossless manner.
[0239] like Figure 7 As shown, this illustrates the principle of deshuffle downsampling. Figure 7 As shown, the image C×aH×aW is downsampled to an image a2C×H×W.
[0240] S503, the device for accelerating rendering acquires the third rendering result of the image frame at the third resolution.
[0241] The third resolution is greater than the first resolution. The third resolution can be greater than or equal to the target resolution of the image frame being rendered. The third rendering result is the final rendering result.
[0242] Specifically, the specific value of the third resolution can be configured according to actual needs. However, this application embodiment does not limit the specific value of the third resolution.
[0243] For example, a higher third resolution configuration allows for the addition of more high-frequency information during rendering, improving the quality of the rendering results, but it also increases rendering complexity and rendering time. Conversely, a lower third resolution configuration allows for the addition of only some high-frequency information during rendering, but it reduces rendering complexity and shortens rendering time.
[0244] S504, the device for accelerating rendering acquires a third rendering result with a second feature at a second resolution.
[0245] For example, the device for accelerating rendering in S504 can extract features from the third rendering result and scale it to a second resolution to obtain a second feature.
[0246] S505, the device for accelerating rendering acquires a fourth rendering result of an image frame at a fourth resolution, the fourth rendering result incorporating the first feature and the second feature.
[0247] In this process, the pixels of the fourth rendering result converge to a second preset value. The second preset value is greater than or equal to the image convergence threshold. In practical applications, the specific value of the image convergence threshold can be configured according to actual needs, and this embodiment does not limit this.
[0248] In one possible implementation, the fourth resolution is the target resolution of the image frame.
[0249] In another possible implementation, the fourth resolution can be smaller than the target resolution of the image frame. The device for accelerating rendering can obtain the rendering result of the image frame at the target resolution by repeatedly executing the process S501 to S505, incrementing the value of the fourth resolution each time the process S501 to S505 is executed, until the fourth resolution is increased to the resolution of the image frame.
[0250] Specifically, in S505, the accelerated rendering device can perform feature fusion between the first feature and the second feature, and use the result of feature fusion as the fourth rendering result.
[0251] When the fourth resolution is the target resolution of the image frame, the fourth rendering result is the high-quality rendering result of that image frame, and the device that accelerates the rendering can output it.
[0252] It should be noted that since the rendering results at the first resolution and the third resolution are both scaled to the second resolution, the S505 needs to scale the merged features to the fourth resolution to obtain the fourth rendering result.
[0253] In one possible implementation, the accelerated rendering device in S505 can combine the features of historical frames of the current image frame during the rendering process to improve the stability and accuracy of the rendering result. Specifically, the accelerated rendering device in S505 obtains the fourth rendering result of the image frame at the fourth resolution by: mapping the first and second features of the historical frames of the current image frame to the image frame according to the motion vector, performing feature fusion with the first and second features of the current image frame, and then performing adaptive fusion to obtain the fourth rendering result.
[0254] Here, historical frames include one or more image frames preceding the image frame. The first and second features of a historical frame can refer to the features resulting from the fusion (historical accumulation) of the first and second features of each historical frame.
[0255] For example, the first and second features of a historical frame can be the final rendering result of the previous image frame of the current image frame.
[0256] For example, a historical frame Xt-1 can contain the accumulated results of multiple different SPPs, which are sequentially mapped to the current image frame through motion vector information.
[0257] Furthermore, in embodiments of this application, the device for accelerating rendering can perform adaptive feature fusion through a trained neural network.
[0258] For example, during the training phase of a neural network, the denoised rendering result at high resolution and high SPP (e.g., 720p, 4096 SPP) can be used as the ground truth.
[0259] The type of neural network used for adaptive feature fusion and the specific results are not specifically limited in this application embodiment, and can be configured according to actual needs.
[0260] For example, a reconstruction network can be used for adaptive feature fusion.
[0261] For example, the device for accelerating rendering can learn the fusion weights corresponding to the current image frame and the historical frames by reconstructing the network, and adaptively fuse feature information to correct errors in the features of the historical frames and the current image frame, thereby improving the fusion effect of the historical frames and the current frame.
[0262] For example, the reconstruction network can adopt an encoder-decoder network (such as the UNet encoder-decoder structure). The network mainly consists of convolutional layers, pooling layers, activation functions, and upsampling modules. The encoder part downsamples the features through pooling layers, expanding the network's receptive field while extracting and fusing features. The decoder part upsamples the low-resolution features through the upsampling module. The upsampled features are then fused with the corresponding scale encoder features through skip connections for subsequent feature extraction.
[0263] For example, the reconstruction network can employ a scale-invariant deep feature extraction network (such as DenseNet).
[0264] Optionally, to improve performance, the reconstructed network can employ residual modules. To reduce computation, the reconstructed network can use a lower network depth (e.g., the number of convolutional layers in the network) and a lower network width (e.g., the number of convolutional kernels in each convolutional layer), while employing depthwise convolution to further reduce network parameters and computation.
[0265] In this process, the feature information of historical frames is back-mapped to the current image frame through motion vector information. In order to maintain size consistency, the motion vector can be obtained by bilinear upsampling.
[0266] Figure 8 This illustrates a scenario where the current image frame is fused with historical frames.
[0267] The accelerated rendering method provided in this application, due to the high SPP rendering result at the first resolution, can ensure that the rendering output contains richer details and is closer to theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering result; the rendering result at the third resolution can supplement the information missing in the rendering result at the low resolution; therefore, the feature of the fusion of the rendering result at the first resolution and the third rendering result is finally used as the final rendering result. Under the premise of fast rendering speed, it ensures that the rendering output contains richer details and is closer to theoretical convergence, while also ensuring lower noise in the rendering result, thus achieving efficient acquisition of high-quality rendering results.
[0268] Furthermore, Figure 9 This illustrates a transformation scenario for the rendering result. The rendering result used in this embodiment can be transformed as follows: Figure 9 The color space conversion (RGB to YUV) shown in (a) and tone mapping, the fourth rendering result obtained by the device for accelerating rendering, can be processed respectively as follows: Figure 9 (b) In the process of inverse tone mapping and color space conversion (YUV to RGB), the final output result in RGB format is obtained.
[0269] Further, optionally, to better supplement high-frequency information such as geometric textures in the rendering results, such as Figure 10 As shown, the accelerated rendering method provided in this application may further include S506 and S507:
[0270] S506, The device for accelerating rendering acquires the fifth rendering result of the image frame at the fifth resolution, wherein the SPP of the fifth rendering result is less than or equal to the third preset value.
[0271] The fifth resolution can be higher than the fourth resolution. The fifth rendering result can be an intermediate rendering result.
[0272] For example, the SPP of the fifth rendering result can be 4SPP. Lower SPP at high resolutions can achieve faster rendering speeds.
[0273] The third preset value can be configured according to actual needs, and this application embodiment does not limit it.
[0274] For example, the third preset value can be the highest SPP threshold value provided by the industry or enterprise for low-quality images. The fifth rendering result is a low-quality rendering result.
[0275] Furthermore, the selection of the SPP value when obtaining the fifth rendering result can be implemented according to actual needs, as long as it is less than or equal to the third preset value. This application embodiment does not specifically limit the SPP value when obtaining the fifth rendering result.
[0276] S507, the device for accelerating rendering acquires the third feature of the fifth rendering result at the second resolution.
[0277] For example, the device for accelerating rendering in S507 can extract features from the fifth rendering result and scale it to the second resolution to obtain the third feature.
[0278] In one possible implementation, when the accelerated rendering method provided in this application includes S506 and S507, obtaining the fourth rendering result of the image frame at the fourth resolution in S505 can be specifically implemented as follows: after feature fusion of the first feature, the second feature and the third feature, adaptive fusion is performed to obtain the fourth rendering result.
[0279] exist Figure 10 In the illustrated method for accelerating rendering, the rendering result at the fifth resolution only contains gbuffer information rendered at high resolution. Compared to obtaining the final rendering result through shaders, gbuffer information can be calculated more quickly, and information such as normal and albedo can provide richer high-frequency information such as geometric textures.
[0280] In one possible implementation, the accelerated rendering device in S505 can combine the features of historical frames of the current image frame during the rendering process to improve the stability and accuracy of the rendering result. Specifically, the accelerated rendering device in S505 obtains the fourth rendering result of the image frame at the fourth resolution by: mapping the first, second, and third features of the historical frames of the current image frame to the image frame according to the motion vector, performing feature fusion with the first, second, and third features of the current image frame, and then performing adaptive fusion to obtain the fourth rendering result.
[0281] Among them, the first feature, second feature and third feature of the historical frame can refer to the features after the first feature, second feature and third feature of each historical frame are fused (historically accumulated).
[0282] It should be noted that the adaptive fusion process of the current image frame and historical frames has been described in detail in S505, and will not be repeated here.
[0283] Further, optionally, to better supplement high-frequency information such as geometric textures in the rendering results, such as Figure 11 As shown, the accelerated rendering method provided in this application may further include S508:
[0284] S508, the device for accelerating rendering acquires the intermediate rendering result of the image frame at the third resolution and acquires its fifth feature at the second resolution.
[0285] In one possible implementation, when the accelerated rendering method provided in this application includes S508, obtaining the fourth rendering result of the image frame at the fourth resolution in S505 can be specifically implemented as follows: after feature fusion of the first feature, the second feature and the fifth feature, adaptive fusion is performed to obtain the fourth rendering result.
[0286] In one possible implementation, the accelerated rendering device in S505 can combine the features of historical frames of the current image frame during the rendering process to improve the stability and accuracy of the rendering result. Specifically, the accelerated rendering device in S505 obtains the fourth rendering result of the image frame at the fourth resolution by: mapping the first, second, and fifth features of the historical frames of the current image frame to the image frame according to the motion vector, performing feature fusion with the first, second, and fifth features of the current image frame, and then performing adaptive fusion to obtain the fourth rendering result.
[0287] Among them, the first feature, second feature and fifth feature of the historical frame can refer to the features after the first feature, second feature and fifth feature of each historical frame are fused (historically accumulated).
[0288] It should be noted that the adaptive fusion process of the current image frame and historical frames has been described in detail in S505, and will not be repeated here.
[0289] The above-mentioned solution in this application can effectively supplement the high-frequency information loss that exists when upsampling from low resolution to high resolution by introducing multiple rendering targets and rendering inputs of different scales. This allows the output result to retain more texture details. At the same time, the low-resolution high SPP input can not only improve the noise reduction effect, but also improve the convergence of the final result, resulting in a higher quality high-resolution output. Using higher resolution rendering intermediate results can effectively supplement high-quality geometric texture information, making the final result more closely resemble the native rendering result in terms of detail richness.
[0290] Furthermore, the multi-frame fusion in the above-mentioned scheme of this application can use the feature information of historical frames to correct the errors in the sampling of the current image frame, and at the same time effectively supplement the information of the current image frame. The fusion of multiple frames can ensure the temporal stability of the rendering result and reduce problems such as temporal aliasing.
[0291] Furthermore, the above-mentioned solution in this application uses a neural network for adaptive fusion, which can adaptively fuse different feature information. Through this network, errors in the features of historical frames and current frames can be corrected, solving problems such as ghosting, jitter, and frame dragging caused by traditional fixed or heuristic weight settings, and improving the fusion accuracy of historical frames and current frames.
[0292] Furthermore, the above-mentioned solution in this application scales the rendering results at multiple scales to a second resolution, which can adaptively adjust the fusion size of inputs at different scales according to the computing power of the current device, significantly reducing computing power requirements and improving the applicability of the algorithm. The adaptive fusion network can better fuse inputs at different scales using deep learning models, achieving higher quality output.
[0293] Furthermore, the rendering results in the scheme provided in this application (except for the fourth rendering result) are all undenoised rendering results containing Monte Carlo noise. Using noisy input can avoid the problem of high-frequency information loss and input blurring caused by denoising algorithms. Combining super-resolution and denoising tasks can integrate more layers of input information and avoid the impact of flaws from a single task on subsequent tasks.
[0294] As can be seen from the above description, the solution provided in this application can utilize the different levels of information provided by low-resolution high spp (LRHS) and high-resolution low spp (HRLS) inputs to generate high-quality rendering results after high-resolution noise reduction.
[0295] In terms of results, the above-mentioned scheme in this application, using only medium (4SPP) and low (1024SPP) model inference at a 2×2 upsampling ratio, achieves results comparable to native 1024SPP rendering at medium resolution, with some areas even showing significantly better detail, resulting in a rendering speed improvement of more than 3 times. At a 4×4 upsampling ratio, it can achieve a rendering speed improvement of up to 20 times while maintaining high-quality output results.
[0296] In terms of results, the above-mentioned scheme of this application can achieve richer output details by using a larger gbuffer input. The results obtained by using high (fifth resolution), medium (third resolution), and low resolution (first resolution) input model inference can be similar to the 4096 SPP rendering results, and present better noise levels and detail information in areas such as complex mirrors and some geometric contours.
[0297] In terms of rendering efficiency, the above-mentioned solution in this application can achieve a speed improvement of 3-20 times. While saving computing power, it can further improve the gap between the output result and the theoretical convergence value, improve the quality of the final rendered image, and is more suitable for the production of offline high-quality CG content.
[0298] Compared to commonly used denoising algorithms in rendering engines such as Optex Denoise and Open Image Denoise, the above-mentioned solution in this application can provide higher quality denoising effects and improved output results. Compared to deep learning denoising algorithms, the above-mentioned solution in this application can provide richer textures, geometry, and other high-frequency details, achieving significantly superior results. Compared to traditional upsampling algorithms such as CAS upscale and FSR, the above-mentioned solution in this application can provide a higher upsampling ratio and results that are closer to realistic rendering output. Compared to the DLSS algorithm, the above-mentioned solution in this application can achieve high-quality and high-efficiency rendering in offline scenes, while the former focuses on fast rendering in real-time scenes.
[0299] On the other hand, embodiments of this application provide another method for accelerating rendering, such as... Figure 12 As shown, the rendering provided in this application embodiment may include:
[0300] S1201, The device for accelerating rendering acquires the sixth rendering result of the image frame at the sixth resolution, and the sixth rendering result is the final rendering result.
[0301] The sixth resolution can be a preset value or a user-input value lower than the target rendering resolution. The target rendering resolution refers to the final resolution of the rendered image.
[0302] Optionally, the sixth resolution can be a fixed value, or it can be a resolution calculated based on the expected rendering quality or the expected rendering duration. This application does not limit the specific value of the sixth resolution or the method of obtaining it.
[0303] For example, a relationship between the desired rendering quality or desired rendering duration and the sixth resolution is pre-configured. The device for accelerating rendering in S1201 can determine the sixth resolution based on this relationship. The specific content of this relationship is not described in the embodiments of this application, but can be configured according to actual needs.
[0304] Among them, the number of intra-pixel samples (SPP) of the sixth rendering result is greater than or equal to the fourth preset value.
[0305] For example, the SPP of the sixth rendering result can be 1024 SPP. A higher SPP ensures that the rendered output contains richer details and is closer to the theoretical convergence value.
[0306] The fourth preset value can be configured according to actual needs, and this application embodiment does not limit it.
[0307] For example, the fourth preset value can be the lowest SPP threshold value for high-quality images provided by the industry or enterprise. The sixth rendering result is a high-quality rendering result.
[0308] Furthermore, the selection of the SPP value when obtaining the sixth rendering result can be implemented according to actual needs, as long as it is greater than or equal to the fourth preset value. This application embodiment does not specifically limit the SPP value when obtaining the sixth rendering result.
[0309] S1202, The device for accelerating rendering acquires subpixel information of image frames.
[0310] The subpixel information is used to describe the characteristics of subpixels in an image frame, and the size of a subpixel is smaller than the size of a pixel.
[0311] For example, subpixel information includes jitter information, which can be obtained through camera shake.
[0312] For example, jitter information can be generated using low-difference sequences such as Halton.
[0313] S1203, the accelerated rendering device upsamples the sixth rendering result to the seventh resolution based on the subpixel information to obtain the seventh rendering result.
[0314] The seventh resolution is greater than the sixth resolution.
[0315] In one possible implementation, the seventh resolution can be the target resolution of the image frame.
[0316] In one possible implementation, the seventh resolution can be smaller than the target resolution of the image frame.
[0317] It should be noted that the specific value of the seventh resolution can be configured according to actual needs or by the user, and this application embodiment does not limit this.
[0318] For example, subpixel information can specifically describe one or more of the following information about a subpixel: its position in the image, color, brightness, or other relevant information. The device for accelerating rendering can upsample the sixth rendering result to a seventh resolution based on the subpixel information. This upsampling can be camera jitter upsampling, that is, mapping the subpixel to the seventh resolution according to its specific position to achieve image magnification. The specific mapping process is not limited in this embodiment.
[0319] Optionally, blank pixels in the Jitter upscale upsampling results can be filled using interpolation methods such as bilinear interpolation.
[0320] S1204. The accelerated rendering device obtains the eighth rendering result of the image frame based on the seventh rendering result, and the pixels of the eighth rendering result converge to the fifth preset value.
[0321] The fifth preset value is greater than or equal to the image convergence threshold. In practical applications, the specific value of the image convergence threshold can be configured according to actual needs, and this embodiment does not limit this.
[0322] Optionally, the device for accelerating rendering in S1204 can obtain the eighth rendering result of the image frame based on the seventh rendering result, which can be achieved through any of the following three schemes, but not limited to:
[0323] Option 1: The device for accelerating rendering will directly use the seventh rendering result as the eighth rendering result.
[0324] In another possible implementation, the accelerated rendering device in S1024 obtains the eighth rendering result of the image frame based on the seventh rendering result. Specifically, the accelerated rendering device maps the seventh rendering result of the historical frames of the image frame to the current image frame according to the motion vector, performs feature fusion with the seventh rendering result of the current image frame, and then performs adaptive fusion to obtain the eighth rendering result. The historical frames include one or more image frames preceding the current image frame.
[0325] It should be noted that the fusion process of the current image frame and the historical frame can be referred to the relevant content described in S505 above, and will not be repeated here.
[0326] Option 2: The device for accelerating rendering acquires the ninth rendering result of the image frame at the sixth resolution. The ninth rendering result is an intermediate rendering result. The ninth rendering result is upsampled to the seventh resolution to obtain the tenth rendering result. After feature fusion of the seventh rendering result and the tenth rendering result, feature extraction is performed to obtain the fourth feature, which is used as the eighth rendering result.
[0327] Option 3: The accelerated rendering device acquires the ninth rendering result of the image frame at the sixth resolution. The ninth rendering result is an intermediate rendering result. The ninth rendering result is upsampled to the seventh resolution to obtain the tenth rendering result. After feature fusion of the seventh and tenth rendering results, feature extraction is performed to obtain the fourth feature. The fourth feature of the historical frame of the image frame is mapped to the image frame according to the motion vector, and after feature fusion with the fourth feature of the image frame, adaptive fusion is performed to obtain the eighth rendering result.
[0328] In this process, the feature information of historical frames is back-mapped to the current image frame through motion vector information. In order to maintain size consistency, the motion vector can be obtained by bilinear upsampling.
[0329] For example, in a static scene, the motion vector information is 0, the rendering results of different SPPs corresponding to the current image frame Xt and the historical frame Xt-1, and the corresponding gbuffer information. The historical frame Xt-1 can contain the accumulated results of multiple different SPPs.
[0330] In options 2 and 3, such as Figure 13 As shown, the method for accelerating rendering provided in this application may also include S1205 and S1206.
[0331] S1205, The device for accelerating rendering acquires the ninth rendering result of the image frame at the sixth resolution.
[0332] In one possible implementation, the device for accelerating rendering can use the intermediate rendering result as the ninth rendering result while obtaining the sixth rendering result.
[0333] In another possible implementation, the device for accelerating rendering can render in parallel to obtain the sixth and ninth rendering results.
[0334] S1206, The accelerated rendering device upsamples the ninth rendering result to the seventh resolution to obtain the tenth rendering result.
[0335] For example, the ninth rendering result may include albedo, normal, position, etc. In S1206, the seventh resolution can be sampled using linear upsampling methods such as bilinear or bicubic to obtain the tenth rendering result.
[0336] Figure 14 The flowchart illustrates a method for accelerating rendering by incorporating historical frames. The process of obtaining the fourth feature of a historical frame is the same as the process of obtaining the fourth feature of the current image frame.
[0337] In one possible implementation, the seventh resolution is less than the target resolution of the image frame. The device that accelerates rendering can execute the process from S1201 to S1204 multiple times in a loop, incrementing the value of the seventh resolution each time the process from S1201 to S1204 is executed, until the seventh resolution is increased to the resolution of the image frame, thereby obtaining the rendering result of the image frame at the target resolution.
[0338] Furthermore, the rendering results used in the embodiments of this application can be as follows: Figure 9 The color space conversion (RGB to YUV) shown in (a) and tone mapping, the eighth rendering result obtained by the accelerated rendering device, can be processed respectively through, as shown in... Figure 9 (b) In the process of inverse tone mapping and color space conversion (YUV to RGB), the final output result in RGB format is obtained.
[0339] The accelerated rendering method provided in this application, due to the high SPP rendering results at sixth resolution, can ensure that the rendered output contains richer details and more closely approximates theoretical convergence while maintaining high rendering speed, and also ensures lower noise in the rendering results. By introducing sub-pixel information as upsampled pixel values, it better conforms to the rendering data distribution, increases the effective information of the input, and ultimately obtains high-resolution results with richer details that are more in line with the native rendered image. Therefore, the final rendering result, while maintaining high rendering speed, ensures that the rendered output contains richer details, more closely approximates theoretical convergence, and conforms to the rendering data distribution, thus achieving efficient acquisition of high-quality rendering results.
[0340] Compared to general super-resolution algorithms, the above-mentioned scheme in this application can achieve faster inference speed; compared to the traditional TAAU algorithm, it can reduce artifacts and blurring problems, and generate higher quality frame fusion results; compared to CAS upscale and FSR algorithms, it can provide a higher upsampling ratio.
[0341] It should be noted that the execution order of the various steps in the accelerated rendering method provided in this application embodiment can be configured according to actual needs. The accompanying drawings only provide one possible implementation order and are not a limitation on the execution order of each step.
[0342] The above primarily describes the solutions provided in the embodiments of this application from the perspective of the device's working principle. It is understood that the aforementioned accelerated rendering device, in order to achieve the above functions, includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0343] This application embodiment can divide the accelerated rendering apparatus provided in this application into functional modules based on the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. The module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0344] When dividing each function into modules according to its corresponding function. Figure 15 A schematic diagram of a possible structure of the accelerated rendering apparatus involved in the above embodiments is shown. The accelerated rendering apparatus 1500 can be a functional module or a chip. For example... Figure 15 As shown, the accelerated rendering device 1500 may include: a first acquisition unit 1501, a second acquisition unit 1502, a third acquisition unit 1503, and a fourth acquisition unit 1504. The first acquisition unit 1501 is used to perform... Figure 5 or Figure 10 The processes S501 and S503 are described in the text; or, the first acquisition unit 1501 can also be used to execute... Figure 10 S506; the second acquisition unit 1502 is used to execute Figure 5 or Figure 10 The process S502; the third acquisition unit 1503 is used to execute Figure 5 or Figure 10 The process S504; the fourth acquisition unit 1504 is used to execute Figure 5 or Figure 10 The process S505 is described above. All relevant content regarding each step in the above method embodiment can be found in the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0345] Furthermore, such as Figure 16 As shown, the accelerated rendering apparatus 1500 may further include a fifth acquisition unit 1505, used for performing... Figure 10 The process S507.
[0346] When using integrated units, Figure 17 A possible structural diagram of the accelerated rendering apparatus involved in the above embodiments is shown. The accelerated rendering apparatus 1700 may include: a processing module 1701 and a communication module 1702. The processing module 1701 is used to control and manage the operation of the accelerated rendering apparatus 1700, and the communication module 1702 is used to communicate with other devices. For example, the processing module 1701 is used to execute... Figure 5 or Figure 10 Any one of the processes S501 to S505, or execution Figure 10 The process is S506 or S507. The accelerated rendering apparatus 1700 may also include a storage module 1703 for storing the program code and data of the accelerated rendering apparatus 1700.
[0347] The processing module 1701 can be a processor or controller. For example, it can be a CPU, general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processing module 1701 can also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc. The communication module 1702 can be a communication port, or a transceiver, transceiver circuit, or communication interface, etc. Alternatively, the aforementioned communication interface can enable communication with other devices through the aforementioned transceiver components. The aforementioned transceiver components can be implemented by antennas and / or radio frequency devices.
[0348] As mentioned above, the accelerated rendering apparatus 1500 or accelerated rendering apparatus 1700 provided in the embodiments of this application can be used to implement the corresponding functions in the methods implemented in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the embodiments of this application.
[0349] When dividing each function into modules according to its corresponding function. Figure 18 A schematic diagram of a possible structure of the accelerated rendering apparatus involved in the above embodiments is shown. The accelerated rendering apparatus 1800 can be a functional module or a chip. For example... Figure 18 As shown, the accelerated rendering device 1800 may include: a first acquisition unit 1801, a second acquisition unit 1802, a magnification unit 1803, and a third acquisition unit 1804. The first acquisition unit 1801 is used to perform... Figure 12 The process S1201; the second acquisition unit 1802 is used to execute Figure 12 The process S1202; the amplification unit 1803 is used to execute Figure 12 The process S1203; the third acquisition unit 1804 is used to execute Figure 12 The process S1204 is described above. All relevant content regarding each step in the above method embodiment can be found in the functional descriptions of the corresponding functional modules, and will not be repeated here.
[0350] When using integrated units, Figure 19 A possible structural schematic diagram of the accelerated rendering apparatus involved in the above embodiments is shown. The accelerated rendering apparatus 1900 may include: a processing module 1901 and a communication module 1902. The processing module 1901 is used to control and manage the operation of the accelerated rendering apparatus 1900, and the communication module 1902 is used to communicate with other devices. For example, the processing module 1901 is used to execute... Figure 12The process includes any one of steps S1201 to S1204. The accelerated rendering apparatus 1900 may also include a storage module 1903 for storing the program code and data of the accelerated rendering apparatus 1900.
[0351] The processing module 1901 can be a processor or controller. For example, it can be a CPU, general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processing module 1901 can also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc. The communication module 1902 can be a communication port, or a transceiver, transceiver circuit, or communication interface, etc. Alternatively, the aforementioned communication interface can enable communication with other devices through the aforementioned transceiver components. The aforementioned transceiver components can be implemented by antennas and / or radio frequency devices.
[0352] As mentioned above, the accelerated rendering apparatus 1800 or accelerated rendering apparatus 1900 provided in the embodiments of this application can be used to implement the corresponding functions in the methods implemented in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the embodiments of this application.
[0353] As another form of this embodiment, a non-transitory computer-readable storage medium is provided, including program code that, when executed, performs the method described in the above method embodiment.
[0354] As another form of this embodiment, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to execute the method described in the above method embodiment.
[0355] This application provides another chip system, which includes a processor for implementing the technical methods of the embodiments of the present invention. In one possible design, the chip system further includes a memory for storing program instructions and / or data necessary for the embodiments of the present invention. In another possible design, the chip system further includes a memory for the processor to call application code stored in the memory. This chip system may be composed of one or more chips, or may include chips and other discrete devices; this application does not specifically limit this.
[0356] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0357] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0358] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0359] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0360] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0361] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for accelerating rendering, characterized in that, The method includes: Obtain the sixth rendering result of the image frame at the sixth resolution, and the sixth rendering result is the final rendering result; wherein, the number of intra-pixel samples (SPP) of the sixth rendering result is greater than or equal to the fourth preset value; Obtain subpixel information of the image frame, wherein the subpixel information is used to describe the characteristics of subpixels in the image frame, and the size of the subpixel is smaller than the size of the pixel; The subpixel information is used as the upsampled pixel value, and the sixth rendering result is upsampled to the seventh resolution to obtain the seventh rendering result; wherein, the seventh resolution is greater than the sixth resolution; Based on the seventh rendering result, the eighth rendering result of the image frame is obtained, and the pixels of the eighth rendering result converge to the fifth preset value.
2. The method according to claim 1, characterized in that, The step of obtaining the eighth rendering result of the image frame based on the seventh rendering result includes: The seventh rendering result of the historical frame of the image frame is mapped to the image frame according to the motion vector, and after feature fusion with the seventh rendering result of the image frame, adaptive fusion is performed to obtain the eighth rendering result; wherein, the historical frame includes one or more image frames before the image frame.
3. The method according to claim 1, characterized in that, The method further includes: obtaining the ninth rendering result of the image frame at the sixth resolution, wherein the ninth rendering result is an intermediate rendering result; upsampling the ninth rendering result to the seventh resolution to obtain the tenth rendering result; The step of obtaining the eighth rendering result of the image frame based on the seventh rendering result includes: After feature fusion of the seventh rendering result and the tenth rendering result, feature extraction is performed to obtain the fourth feature; The fourth feature of the historical frame of the image frame is mapped to the image frame according to the motion vector, and after feature fusion with the fourth feature of the image frame, adaptive fusion is performed to obtain the eighth rendering result; wherein, the historical frame includes one or more image frames before the image frame.
4. The method according to claim 3, characterized in that, The intermediate rendering results include: normal maps, or depth maps.
5. The method according to any one of claims 1-4, characterized in that, The seventh resolution is the target resolution of the image frame.
6. The method according to any one of claims 1-4, characterized in that, Feature fusion can be performed using a trained neural network, or adaptive fusion.
7. An apparatus for accelerating rendering, characterized in that, The device includes: The first acquisition unit is used to acquire the sixth rendering result of the image frame at the sixth resolution, wherein the sixth rendering result is the final rendering result; wherein the number of intra-pixel samples (SPP) of the sixth rendering result is greater than or equal to a fourth preset value. The second acquisition unit is used to acquire sub-pixel information of the image frame, wherein the sub-pixel information is used to describe the characteristics of sub-pixels in the image frame, and the size of the sub-pixel is smaller than the size of the pixel; The magnification unit is used to use the sub-pixel information acquired by the second acquisition unit as an upsampled pixel value to upsample the sixth rendering result to a seventh resolution, thereby obtaining a seventh rendering result; wherein the seventh resolution is greater than the sixth resolution; The third acquisition unit is used to acquire the eighth rendering result of the image frame based on the seventh rendering result obtained by the magnification unit, wherein the pixels of the eighth rendering result converge to a fifth preset value.
8. The apparatus according to claim 7, characterized in that, The third acquisition unit is specifically used for: The seventh rendering result of the historical frame of the image frame is mapped to the image frame according to the motion vector, and after feature fusion with the seventh rendering result of the image frame, adaptive fusion is performed to obtain the eighth rendering result; wherein, the historical frame includes one or more image frames before the image frame.
9. The apparatus according to claim 7, characterized in that, The first acquisition unit is further configured to: acquire the ninth rendering result of the image frame at the sixth resolution, wherein the ninth rendering result is an intermediate rendering result; The magnification unit is also used to: upsample the ninth rendering result to the seventh resolution to obtain the tenth rendering result; The third acquisition unit is specifically used to: perform feature fusion on the seventh rendering result and the tenth rendering result, and then extract features to obtain the fourth feature; The fourth feature of the historical frame of the image frame is mapped to the image frame according to the motion vector, and after feature fusion with the fourth feature of the image frame, adaptive fusion is performed to obtain the eighth rendering result; wherein, the historical frame includes one or more image frames before the image frame.
10. The apparatus according to claim 9, characterized in that, The intermediate rendering results include: normal maps, or depth maps.
11. The apparatus according to any one of claims 7-10, characterized in that, The seventh resolution is the target resolution of the image frame.
12. The apparatus according to any one of claims 7-10, characterized in that, Feature fusion can be performed using a trained neural network, or adaptive fusion.
13. A rendering device, characterized in that, The rendering device includes: a processor and a memory; The memory is connected to the processor; The memory is used to store computer instructions, and when the processor executes the computer instructions, the rendering device performs the method as described in any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, Includes instructions that, when run on a computer, cause the computer to perform the method of any one of claims 1 to 6.
15. A computer program product containing instructions, characterized in that, When it is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Rendering method and device
CN107330966A
Dynamic enhancement / reduction of graphical image data resolution
US5982373A