Distributed rendering processing method, system, device, equipment and storage medium
Through the distributed rendering processing method, the rendering task is split and multiple rendering nodes and noise reduction nodes are used for parallel processing, which solves the calculation accuracy and memory space limitations of the GPU rendering solution in the prior art, and realizes efficient rendering and noise reduction processing, which improves system throughput and reduces costs.
Patent Information
- Application Number
- CN202411686184.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-11-22
AI Technical Summary
In the prior art, GPU-based rendering schemes have limited computing accuracy, limited video memory space, and high hardware costs, resulting in low overall system throughput.
The distributed rendering processing method is adopted to realize efficient rendering and noise reduction processing by splitting the target rendering task into multiple sub-rendering tasks and using multiple rendering nodes and noise reduction nodes for parallel processing.
It improves the system throughput, reduces rendering costs, can adapt to any rendering workflow, and deploys in high-performance computing clusters to achieve efficient image rendering and noise reduction processing.
Smart Images

Figure CN119180743B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a distributed rendering processing method, apparatus, system, device, and storage medium. Background Art
[0002] In the field of computer graphics, in order to improve the rendering solution quality of offline ray tracing, a combined device of a renderer (mainly for rendering) + a denoiser (mainly for denoising) can be adopted. Here, both the renderer and the denoiser can be implemented using a Central Processing Unit (CPU) or a Graphics Processing Unit (GPU); the technology solution based on the CPU is mature and suitable for processing complex scenes and can complete accurate global illumination calculations; while the technology solution based on the GPU has developed relatively late. Thanks to the fact that the GPU has more parallel processing units, the GPU solves problems faster than the CPU. However, limited by the capabilities of the processing units, the calculation accuracy is limited, and the video random access memory (VRAM) space of the GPU is limited. For example, the space of the VRAM is limited and it is difficult to handle complex scenes with large amounts of data; in addition, the hardware cost of the GPU is much higher than that of the CPU. Summary of the Invention
[0003] The present disclosure provides a distributed rendering processing method, system, apparatus, device, and storage medium to solve or alleviate one or more technical problems in the prior art.
[0004] In a first aspect, the present disclosure provides a distributed rendering processing method, including:
[0005] Determine N sub-rendering tasks obtained after splitting a target rendering task; where N is an integer greater than or equal to 2;
[0006] Invoke M1 rendering nodes to render the N sub-rendering tasks in parallel using the M1 rendering nodes to obtain N sub-rendering results; M1 is an integer less than or equal to N;
[0007] Invoke M2 denoising nodes to denoise the N sub-rendering results in parallel using the M2 denoising nodes to obtain N denoising processing results; M2 is an integer less than or equal to M1;
[0008] Obtain a target rendering result for the target rendering task based on the N denoising processing results.
[0009] In a second aspect, the present disclosure provides a distributed rendering processing system, including:
[0010] A scheduling node, configured to determine N sub-rendering tasks obtained after splitting a target rendering task; where N is an integer greater than or equal to 2; and configured to call M1 rendering nodes and M2 noise reduction nodes; M1 is an integer less than or equal to N; M2 is an integer less than or equal to M1;
[0011] M1 rendering nodes, configured to render the N sub-rendering tasks in parallel to obtain N sub-rendering results;
[0012] M2 noise reduction nodes, configured to perform noise reduction on the N sub-rendering results in parallel to obtain N noise reduction processing results;
[0013] Wherein, the rendering nodes among the M1 rendering nodes are further configured to obtain a target rendering result for the target rendering task based on the N noise reduction processing results.
[0014] In a third aspect, the present disclosure provides a distributed rendering processing apparatus, including:
[0015] A splitting module, configured to determine N sub-rendering tasks obtained after splitting a target rendering task; where N is an integer greater than or equal to 2;
[0016] A processing module, configured to call M1 rendering nodes to render the N sub-rendering tasks in parallel by using the M1 rendering nodes to obtain N sub-rendering results; M1 is an integer less than or equal to N; call M2 noise reduction nodes to perform noise reduction on the N sub-rendering results in parallel by using the M2 noise reduction nodes to obtain N noise reduction processing results; M2 is an integer less than or equal to M1; and obtain a target rendering result for the target rendering task based on the N noise reduction processing results.
[0017] In a fourth aspect, there is provided an electronic device, including:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0021] In a fifth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0022] In a sixth aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements the method according to any one of the embodiments of the present disclosure.
[0023] The beneficial effects of the technical solutions provided by the present disclosure at least include:
[0024] The solution of the present disclosure provides a new distributed heterogeneous rendering-denoising architecture, which can be deployed in a high-performance computing cluster. Thus, image rendering and denoising processing can be efficiently implemented, and any rendering workflow can be adapted. At the same time, the throughput of the entire system is effectively improved.
[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0026] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided by the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0027] Figure 1 is a schematic structural diagram of a distributed rendering processing system according to an embodiment of the present application;
[0028] Figure 2 is a schematic flowchart of a distributed rendering processing method according to an embodiment of the present application Figure 1 ;
[0029] Figure 3 is a flowchart illustration of a distributed rendering processing method in a specific example according to an embodiment of the present application Figure 1 ;
[0030] Figure 4 is a flowchart illustration of a distributed rendering processing method in a specific example according to an embodiment of the present application Figure 2 ;
[0031] Figure 5 is a schematic diagram of the segmentation effect of a target image according to an embodiment of the present application;
[0032] Figure 6 is a comparison diagram of the denoising processing effects of sub-images according to an embodiment of the present application;
[0033] Figure 7 is a schematic diagram of the effect of the rendered image after result merging according to an embodiment of the present application;
[0034] Figure 8 It is a schematic structural diagram of a distributed rendering device according to an embodiment of the present application;
[0035] Figure 9 It is a block diagram of an electronic device for implementing the distributed rendering processing method of the embodiments of the present disclosure. Detailed implementation manners
[0036] The present disclosure will be further described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0037] As used herein, the term "and / or" merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C. The terms "first" and "second" as used herein refer to and distinguish multiple similar technical terms, and do not mean to limit the order or limit to only two. For example, the first feature and the second feature refer to two categories / two features, and the first feature may be one or more, and the second feature may also be one or more.
[0038] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0039] In the field of computer graphics, in order to improve the rendering solution quality of offline ray tracing, a combined device of a renderer (mainly for rendering) + a denoiser (mainly for denoising) can be adopted. Here, both the renderer and the denoiser can be implemented using a CPU or a GPU; the technology solution based on the CPU is mature and suitable for processing complex scenes and can complete accurate global illumination calculations; while the technology solution based on the GPU has developed relatively late. Thanks to the fact that the GPU has more parallel processing units, the GPU solves problems faster than the CPU. However, limited by the capabilities of the processing units, the calculation accuracy is limited. In addition, the video memory space of the GPU is limited. For example, the space of the Video Random Access Memory (VRAM) is limited and it is difficult to handle complex scenes with large amounts of data; in addition, the hardware cost of the GPU is much higher than that of the CPU.
[0040] In order to balance rendering quality and the speed of rendering delivery, in one example, a heterogeneous combination device can also be adopted, that is, a CPU renderer (i.e., a renderer implemented using a CPU) + a GPU denoiser (i.e., a denoiser implemented using a GPU). Here, the CPU renderer can make full use of the computer memory to complete the precise calculation of complex scenes; while the GPU denoiser can load a pre-trained deep learning model, take the result of the CPU renderer as the algorithm input, and quickly solve it. For example, in one example, the workflow can be specifically as follows: Input scene data to the CPU renderer, calculate and solve to obtain a set of rendered channel components, for example, including RGB channels, diffuse channels, refraction channels, reflection channels, and specular channels, etc.; further, take the obtained channel components as the input of the deep learning model to infer and obtain the denoised image after synthesizing the channels.
[0041] If the above heterogeneous combination device is installed and deployed on a single machine, the following problems will occur:
[0042] In the scenario of completing batch rendering tasks, all computers need to be installed with high-performance CPUs and at least one GPU, and the overall machine hardware cost is high; moreover, limited by the serial structure of the workflow, the utilization rate of the GPU will also be reduced. For example, during the calculation and solution stage of the CPU renderer, the GPU is in an idle state, and at this time, the overall system throughput is low.
[0043] Based on this, to solve the above problems, the present disclosure provides a distributed rendering processing method, which can efficiently implement image rendering and denoising processing through a heterogeneous acceleration architecture, thereby effectively improving the throughput of the entire system.
[0044] The present disclosure provides a distributed rendering processing system, as Figure 1 shown, the system specifically includes:
[0045] A scheduling node 101, configured to determine N sub-rendering tasks obtained after splitting a target rendering task; where N is an integer greater than or equal to 2; and configured to call M1 rendering nodes and M2 denoising nodes; M1 is an integer less than or equal to N; M2 is an integer less than or equal to M1;
[0046] M1 rendering nodes 102, configured to render the N sub-rendering tasks in parallel to obtain N sub-rendering results;
[0047] M2 denoising nodes 103, configured to perform denoising on the N sub-rendering results in parallel to obtain N denoising processing results;
[0048] Among them, the rendering node in the M1 rendering nodes is further configured to obtain a target rendering result for the target rendering task based on the N noise reduction processing results.
[0049] For example, in one example, the scheduling node can be mainly configured to receive a rendering request from outside the system (for requesting to execute a target rendering task), perform splitting of the target rendering task, schedule rendering nodes or noise reduction nodes, and return results to the outside, etc.
[0050] Further, the rendering node can be mainly configured to receive the sub-rendering tasks assigned by the scheduling node, and after completing the rendering, initiate a noise reduction request to the scheduling node; further, it can also be configured to, through the coordination of the scheduling node, merge the final results to obtain a target rendering result and return it to the scheduling node.
[0051] Further, the noise reduction node is mainly configured to receive the noise reduction tasks assigned by the scheduling node, and after completing the noise reduction processing, return the noise reduction processing results to the rendering node.
[0052] That is to say, the solution of the present disclosure provides a new distributed heterogeneous rendering-noise reduction architecture, which can be deployed in a high-performance computing cluster. For example, there are a certain number of GPU servers and a certain number of CPU servers in the cluster. Thus, it can efficiently adapt to the image rendering workflow, and at the same time, it can also efficiently adapt to the video (multiple-frame image) rendering workflow.
[0053] Specifically, Figure 2 is a schematic flowchart of a distributed rendering processing method according to an embodiment of the present application Figure 1 . This method is optionally applied to an electronic device, such as an electronic device such as a personal computer, a server, a server cluster, etc.
[0054] Further, this method at least includes at least part of the following content. As Figure 2 shown, it includes:
[0055] Step S201: Determine N sub-rendering tasks obtained after splitting the target rendering task through a scheduling node.
[0056] Here, N is an integer greater than or equal to 2.
[0057] Step S202: Call M1 rendering nodes through a scheduling node to render the N sub-rendering tasks in parallel by using the M1 rendering nodes, and obtain N sub-rendering results.
[0058] Here, M1 is an integer less than or equal to N.
[0059] Further, in one example, the number of rendering nodes to be called is related to the value of N. For example, M1 is equal to N.
[0060] Further, in another example, the value of N can also be determined based on the number of currently available rendering nodes. At this time, N can be equal to M1. At this time, a sub-rendering task can correspond to one rendering node and one noise reduction node. In this way, a foundation is laid for quickly completing the rendering task.
[0061] It can be understood that in practical applications, the value of N can also be determined based on the total task volume of the target rendering task, or N is a preset empirical value, etc. The solution of the present disclosure does not make specific limitations on this.
[0062] Step S203: Through the scheduling node, call M2 noise reduction nodes to use the M2 noise reduction nodes to perform noise reduction on the N sub-rendering results in parallel, and obtain N noise reduction processing results.
[0063] Here, M2 is an integer less than or equal to M1.
[0064] Further, in one example, the number of noise reduction nodes to be called is related to at least one of the following: the value of N, the number of rendering nodes. For example, in one example, M2 is equal to M1. Or, in another example, the values of N, M1, and M2 are all equal.
[0065] Or, further, in another example, the number of noise reduction nodes to be called may also be related to the number of currently available noise reduction nodes in the rendering system, or the number of noise reduction nodes that can currently execute noise reduction tasks.
[0066] Step S204: Obtain the target rendering result for the target rendering task based on the N noise reduction processing results. For example, in one example, the scheduling node schedules a rendering node among the M1 rendering nodes to perform result merging, that is, obtain the N noise reduction processing results, and obtain the target rendering result for the target rendering task based on the N noise reduction processing results.
[0067] That is to say, the solution of the present disclosure provides a distributed rendering system. This rendering system can call multiple rendering nodes and multiple noise reduction nodes. In this way, image rendering and noise reduction processing can be efficiently implemented through a heterogeneous acceleration architecture, thereby effectively improving the throughput of the entire system, and further reducing the rendering cost.
[0068] Here, in a specific example, the multiple processing nodes included in the rendering system. Correspondingly, the M1 rendering nodes and the M2 noise reduction nodes are at least part of the multiple processing nodes included in the rendering system.
[0069] Further, in a specific example, the multiple processing nodes at least include multiple CPU nodes and multiple GPU nodes. Further, the M1 rendering nodes are at least some of the multiple CPU nodes and multiple GPU nodes; that is, in this example, some of the rendering nodes are GPU nodes and some are CPU nodes.
[0070] Further, in another example, the M2 noise reduction nodes are at least some of the multiple GPUs. That is, in this example, the noise reduction nodes are GPU nodes.
[0071] Here, it should be noted that the above-mentioned CPU nodes may specifically be CPU servers, and correspondingly, the GPU nodes may specifically be GPU servers.
[0072] Further, in a specific example, different task splitting (or segmentation) methods can be adopted for different rendering tasks, so as to facilitate subsequent fast image rendering and noise reduction processing; specifically, the above step S201 (that is, determining the N sub-rendering tasks obtained after splitting the target rendering task) may specifically include:
[0073] Step S201-1: When the target rendering task is an image rendering task, based on the image features of the target image targeted by the image rendering task, the target image is split into N sub-images to obtain N sub-rendering tasks. For example, each sub-image corresponds to a sub-rendering task.
[0074] Or,
[0075] Step S201-2: When the target rendering task is a video rendering task, based on the number of frames of the video frame sequence targeted by the video rendering task, the video frame sequence is split into N sub-sequences to obtain N sub-rendering tasks. For example, each sub-sequence corresponds to a sub-rendering task.
[0076] It should be noted that in the solution of the present disclosure, in order to accelerate the sub-workflow of noise reduction processing, noise reduction processing can be directly performed on the sub-space rendering results (that is, the above-mentioned sub-rendering results). Further, directly performing a noise reduction algorithm on the segmented sub-space may cause the obtained results to be prone to boundary noise, affecting the final stitching effect. Based on this, in a specific example, before using the M1 rendering nodes to parallelly render the N sub-rendering tasks, the rendering nodes can perform boundary compensation processing on the sub-rendering tasks that they need to process themselves, so as to effectively avoid boundary noise.
[0077] Further, in a specific example, the above-mentioned use of the rendering nodes to perform boundary compensation processing on the sub-rendering tasks that they need to process themselves may specifically include:
[0078] In the case where the target rendering task is an image rendering task, at this time, a rendering node can be utilized, and based on the image features of the sub-image targeted by the sub-rendering task that needs to be processed itself, height compensation and / or width compensation can be performed on the boundary region of the sub-image; here, the sub-image is a partial region in the target image targeted by the image rendering task. For example, after segmenting the target image, multiple sub-images are obtained.
[0079] Or,
[0080] In the case where the target rendering task is a video rendering task, at this time, a rendering node can be utilized, and based on the number of frames of the subsequence targeted by the sub-rendering task that needs to be processed itself, frame sequence compensation can be performed on the subsequence, where the subsequence is a subsequence in the video frame sequence targeted by the video rendering task. For example, after segmenting the video frame sequence, multiple subsequences are obtained.
[0081] Here, in a specific example, for an image rendering task, the height compensation and / or width compensation for the boundary region of the sub-image described above may specifically include:
[0082] Performing perspective transformation on the sub-image based on the target perspective; where the target perspective is related to the acquisition perspective of the target image. For example, the target perspective is the camera perspective. In this way, the problem of reduced rendering effect caused by perspective deviation can be effectively avoided. Further, height compensation and / or width compensation are performed on the boundary region of the sub-image after perspective transformation.
[0083] Here, in another specific example, for a video rendering task, the frame sequence compensation for the subsequence described above may specifically include: extending a preset number of video frames before and / or after the subsequence. Here, it should be noted that the preset number can be determined based on actual needs, and the present disclosure scheme does not limit this.
[0084] Further, in order to further improve the rendering effect, in an example, after obtaining N noise reduction processing results, the boundary compensation part in the noise reduction processing results can also be removed.
[0085] Further, obtaining the target rendering result for the target rendering task based on the N noise reduction processing results described above may specifically include: performing a merging process on the N noise reduction processing results after removing the boundary compensation part to obtain the target rendering result for the target rendering task.
[0086] The following further elaborates on the present disclosure scheme in combination with specific examples. Specifically:
[0087] This example provides a new distributed heterogeneous rendering-denoising system, which can be deployed in a high-performance computing cluster. For example, there are a certain number of GPU servers and a certain number of CPU servers in the cluster; this system can adapt to both image rendering workflows and video (multiple-frame image) rendering workflows.
[0088] Furthermore, in this example, the input of the workflow is a rendering request, and its output is the target rendering result.
[0089] Here, the rendering request encapsulates a data entity of the "target rendering task". For example, it includes the complete rendering scene data and the constraint information of the rendering result.
[0090] For example, in one example, the rendering scene data includes unstructured data such as 3D mesh models (Meshes) and texture maps in a 3D scene, as well as the transformation matrix of the model in the scene coordinate system; furthermore, for video rendering tasks, the rendering scene data can be further segmented, for example, segmented into a frame sequence. Furthermore, the rendering request can also include camera data in the rendering scene; furthermore, for video rendering tasks, the camera data can also be independently stored in each frame.
[0091] In addition, in another example, the rendering request can also contain constraint information of the rendering result, such as resolution (height and width pixel values), the encoding method of the result, etc.; furthermore, for video rendering, it can further include information such as frame rate and bit rate.
[0092] It should be noted that the above content related to the rendering request is only for illustrative purposes. In actual applications, it can also be set according to actual needs, and the present disclosure solution does not limit this.
[0093] Furthermore, the workflow of this example is deployed in a distributed computing cluster. Furthermore, according to functional roles, the nodes in the distributed computing cluster can be divided into: 1) Core nodes, also known as business function nodes. For example, it can include scheduling nodes, rendering nodes, and denoising nodes; 2) Auxiliary nodes, also known as infrastructure nodes. For example, it can include queue nodes and storage nodes.
[0094] Specifically, the functions of each node are briefly described as follows:
[0095] The scheduling node can be mainly used to receive rendering requests from outside the system, and perform segmentation of the target rendering task, schedule rendering nodes or denoising nodes, and return results to the outside.
[0096] A rendering node can be mainly used to receive sub-rendering tasks assigned by a scheduling node, and after completing the rendering, initiate a denoising request to the scheduling node; further, it can also be used to merge the final rendering results through the coordination of the scheduling node. For example, it can merge the denoising results obtained after each denoising node performs denoising processing to obtain a target rendering task, and then send it back to the scheduling node for the scheduling node to feedback the target rendering task to an external system.
[0097] A denoising node is mainly used to receive denoising tasks assigned by a scheduling node, and after completing the denoising process, send back the denoising result to the rendering node.
[0098] Here, in one example, the distributed computing cluster includes at least one scheduling node, N rendering nodes, and N denoising nodes. Further, in one example, the rendering nodes and the scheduling nodes are in one-to-one correspondence. For example, for the i-th denoising node, the scheduling node can schedule the i-th denoising node to perform denoising on the sub-rendering result (also called the i-th sub-rendering result) obtained by the i-th rendering node; further, after the i-th denoising node performs denoising on the sub-rendering result of the i-th rendering node and obtains a denoising result (which can be called the i-th denoising result), the obtained i-th denoising result can be sent back to the i-th rendering node.
[0099] A queue node is mainly used to perform the duties of a message queue and assist the scheduling node in implementing the rendering and denoising task queues;
[0100] A storage node mainly forms a storage cluster to provide large-scale storage of unstructured data such as rendering models and textures. Further, it can also support data storage for the queue node upward.
[0101] Further, the workflow of this example can include multiple processing units, which can be deployed in the above core nodes (business function nodes). For example, as Figure 3 shown, it can specifically include a rendering request splitting unit, a splitting fault tolerance compensation unit, a rendering processing unit, a denoising processing unit, and a result merging unit.
[0102] Further, the following combines Figure 3 and Figure 4 to elaborate on the specific processing flows of each unit in detail:
[0103] The rendering request splitting unit: It is deployed on the scheduling node. This rendering request splitting unit can split a target rendering task into several sub-tasks (corresponding to the above sub-rendering tasks); further, it can also attach splitting metadata to each sub-task.
[0104] Further, in one example, the rendering request splitting unit may adopt different splitting strategies according to the type of the rendering task (image rendering or video rendering):
[0105] For an image rendering task, the rendering area may be split in a rectangular splitting manner according to the resolution (for example, the resolution of the target image to be rendered). For example, a set of splitting values is preset according to the image resolution; the rendering request splitting unit splits the target image targeted by the target rendering task into sub-images, that is, sub-tasks. At the same time, the sub-tasks may be numbered in the upper left - lower right order to ensure that the target image can be stitched according to the numbers. Further, in one example, each sub-task may be scheduled to a rendering node for rendering. Based on this, a total of rendering nodes are scheduled. At this time, the solution area of each rendering node is of the target image. Further, the target total rendering result can finally be obtained based on the numbers.
[0106] For a video rendering task, the video frame sequence to be rendered may be split into several sub-sequences according to the number of frames (for example, the number of frames of the video frame sequence to be rendered). For example, a set of splitting values is preset according to the video resolution; the rendering request splitting unit splits the video frame sequence targeted by the target rendering task into N sub-sequences, that is, N sub-tasks. Further, in one example, each sub-task may be scheduled to a rendering node for parallel rendering. Based on this, a total of N rendering nodes are scheduled. At this time, the number of rendered frames of each rendering node is of the length of the original frame sequence (the sequence length of the video frame sequence to be rendered).
[0107] It should be noted that if integer division splitting cannot be achieved, the remainder may be appended to the last sub-task. For example, for a video rendering task with an original frame sequence length of L, the sub-sequence length of the first N - 1 sub-tasks is L DIV N’, and the sub-sequence length of the Nth sub-task is L DIV N + L MOD N; the video frame sequence can finally be stitched according to the frame numbers.
[0108] Split Fault Tolerance Compensation Unit: Deployed on the rendering nodes. After being processed by the aforementioned unit, the subtasks scheduled to each rendering node result in a subspace of the target solution (image rendering task: sub-pixel interval; video rendering task: sub-sequence interval). The computational workload for the renderer to solve has been reduced, which is the parallel acceleration of the rendering sub-workflow. To further accelerate the denoising sub-workflow, denoising will be directly performed on the subspace rendering results. However, when performing the denoising algorithm on the segmented subspace, the resulting image is prone to boundary noise, affecting the final stitched effect (image: visible segmentation boundary; video: visible flickering frames). This unit executes the following fault tolerance compensation strategies according to the type of rendering task:
[0109] For the regions obtained by splitting the image rendering task, region padding is performed. For example, for each sub-image, based on the perspective of the center point of the target image, the offset and frustum of the scene camera of the sub-image can be modified.
[0110] For example, in one example, the transformation matrix can be:
[0111]
[0112] Here, W refers to the width of the target image to be rendered; H refers to the height of the target image to be rendered.
[0113] For example, in one example, the pixel matrix of the sub-image can be multiplied by the above transformation matrix to obtain the sub-image after perspective transformation. Further, region compensation is performed on the sub-image after perspective transformation.
[0114] Further, let the resolution of the target image be pixels. Then the segmentation interval of each subtask (i.e., the segmentation interval of each sub-image) is pixels.
[0115] Further, denote the height compensation and width compensation of the subtask as and , respectively. Then, according to the resolution of the subtask (i.e., the resolution of the sub-image), the actual rendering width of the subtask (i.e., the actual rendering width of the sub-image) is pixels. Correspondingly, the actual rendering height of the subtask (i.e., the actual rendering height of the sub-image) is pixels.
[0116] Here, The value range of . Further, in one example, it can take the value of . For example, in one example, it takes the value of 2.
[0117] Correspondingly, The value range is . Further, in one example, the value may be For example, in one example, the value is 2.
[0118] For the subsequences obtained by the video rendering task, frame padding is performed. For example, for each subsequence, n frames are extended forward and backward respectively as the actual rendering subsequence of the current rendering node.
[0119] It should be noted that, for the first and last subsequences, only one direction of expansion is required. For example, for the first subsequence, backward expansion is performed, and for the last subsequence, forward expansion is performed.
[0120] Here, it should be noted that, in the merging stage of the final results, the above compensated boundaries may be clipped and discarded.
[0121] Rendering processing unit: deployed on the rendering node, mainly used to call the renderer to complete the rendering task. Depending on the workload, the rendering node can be a high-performance CPU server or a GPU server; it should be noted that the GPU server as a rendering node can also be used as the processing logic for executing the subsequent noise reduction processing unit. For example, in a scene, the currently idle rendering node (the rendering node implemented by the GPU server) can be used as a noise reduction node.
[0122] Denoising processing unit: deployed on the denoising node. Each rendering node can solve a set of rendering channel components for the subtask, for example, in one example, it can include RGB channels, diffuse channels, refraction channels, reflection channels, and specular channels. The components of these channels can be stored in a preset format, such as the OpenEXR 2.0 format, which supports multi-part data storage. For example, the metadata of each channel component can be stored separately in the file header.
[0123] Furthermore, after the aforementioned rendering processing unit completes the solution of the subtask, it encapsulates the relevant data and initiates a denoising request to the scheduling node. Accordingly, the scheduling node selects an idle denoising node to forward the denoising request. Furthermore, after the denoising processing unit responds to the denoising request, it pulls the rendering channel component from the corresponding rendering node through point-to-point (P2P) transmission, and calls the denoiser to complete the denoising. Furthermore, the denoising processing result is transmitted back to the corresponding rendering node.
[0124] For example, if Figure 5As shown, according to the resolution of the target image to be rendered (for example, ), divide the target image according to to obtain 3 sub-images (i.e., sub-image 1, sub-image 2, and sub-image 3). The original pixels of each sub-image are . At this time, the fault tolerance threshold for regional boundary compensation can be set to , then the actual pixels of each sub-image are . The specific comparison effect after noise reduction is as shown in Figure 6 .
[0125] Result merging unit: Deployed on the rendering node. After all the rendering and noise reduction sub-workflows of a task are completed, this result merging unit starts to merge the final result. Its steps include:
[0126] Discard the fault tolerance compensation. For example, crop the redundant interval and eliminate the redundant frames;
[0127] Result merging encoding. For example, according to the result constraints in the rendering request, merge the encoded images or merge the encoded frame sequences. The final product is a complete rendered image or video (corresponding to the target rendering result described above).
[0128] For example, continue to take Figure 5 as an example, perform a merging process on each sub-image after noise reduction processing to obtain the complete rendered image shown in Figure 7 .
[0129] The solution of the present disclosure can be applied to complex rendering tasks with large scenes, multiple objects, and a large number of texture maps. Moreover, rendering results such as ultra-high-definition pictures (8K and above) and ultra-clear videos (4K and above) can be obtained.
[0130] The solution of the present disclosure makes full use of the heterogeneous acceleration architecture. Through the collaborative work of the CPU server and the GPU server, an efficient image rendering and noise reduction processing workflow is realized. Moreover, the solution of the present disclosure can handle both single-image rendering and video rendering.
[0131] Further, the solution of the present disclosure can also reduce the pre-rendering sampling time and compensate for the rendering image quality through the heterogeneous acceleration noise reduction algorithm, which is of great significance for improving the rendering efficiency and the overall visual performance:
[0132] First, the overall system cost is low; the solution of the present disclosure introduces a distributed batch processing architecture, and the overall system cost can be controlled to the lowest by adjusting the ratio of the CPU server and the GPU server.
[0133] Second, strong adaptability; the solution of the present disclosure can adapt to video rendering workflows and image rendering workflows. Meanwhile, it can adaptively select the sharding method. Further, the solution of the present disclosure decouples from the renderer and the noise reducer. In this way, it can be adjusted according to the actual scenario requirements, further enhancing the adaptability. For example, the renderer of the solution of the present disclosure can completely isolate or co-locate the rendering nodes and the noise reduction nodes according to different implementation routes (CPU / GPU).
[0134] Third, high execution efficiency; the solution of the present disclosure can fully utilize the maximum throughput of the heterogeneous acceleration system.
[0135] The solution of the present disclosure also provides a distributed rendering processing device, as Figure 8 shown, including:
[0136] A splitting module 801, configured to determine N sub-rendering tasks obtained after splitting a target rendering task; where N is an integer greater than or equal to 2;
[0137] A processing module 802, configured to call M1 rendering nodes to render the N sub-rendering tasks in parallel by using the M1 rendering nodes to obtain N sub-rendering results; M1 is an integer less than or equal to N; call M2 noise reduction nodes to perform noise reduction on the N sub-rendering results in parallel by using the M2 noise reduction nodes to obtain N noise reduction processing results; M2 is an integer less than or equal to M1; and obtain a target rendering result for the target rendering task based on the N noise reduction processing results.
[0138] In a specific example of the solution of the present disclosure, the M1 rendering nodes and the M2 noise reduction nodes are at least part of the multiple processing nodes included in the rendering system.
[0139] In a specific example of the solution of the present disclosure, the multiple processing nodes at least include multiple CPU nodes and multiple GPU nodes; the M1 rendering nodes are at least part of the multiple CPU nodes and the multiple GPU nodes;
[0140] and / or,
[0141] the M2 noise reduction nodes are at least part of the multiple GPUs.
[0142] In a specific example of the solution of the present disclosure, the splitting module 801 is specifically configured to:
[0143] In the case where the target rendering task is an image rendering task, split the target image into N sub-images based on the image features of the target image for the image rendering task to obtain N sub-rendering tasks;
[0144] Or,
[0145] In the case where the target rendering task is a video rendering task, the video frame sequence is split into N sub-sequences based on the number of frames of the video frame sequence targeted by the video rendering task, so as to obtain N sub-rendering tasks.
[0146] In a specific example of the present disclosure solution, the processing module 802 is further configured to:
[0147] Before rendering the N sub-rendering tasks in parallel using M1 rendering nodes, use the rendering nodes to perform boundary compensation processing on the sub-rendering tasks that need to be processed by themselves.
[0148] In a specific example of the present disclosure solution, the processing module 802 is specifically configured to:
[0149] In the case where the target rendering task is an image rendering task, use the rendering nodes, and based on the image features of the sub-image targeted by the sub-rendering task that needs to be processed by itself, perform height compensation and / or width compensation on the boundary region of the sub-image; wherein, the sub-image is a partial region in the target image targeted by the image rendering task;
[0150] Or,
[0151] In the case where the target rendering task is a video rendering task, use the rendering nodes, and based on the number of frames of the sub-sequence targeted by the sub-rendering task that needs to be processed by itself, perform frame sequence compensation on the sub-sequence, wherein the sub-sequence is a sub-sequence in the video frame sequence targeted by the video rendering task.
[0152] In a specific example of the present disclosure solution, the processing module 802 is specifically configured to:
[0153] Perform perspective conversion on the sub-image based on the target perspective; wherein, the target perspective is related to the acquisition perspective of the target image;
[0154] Perform height compensation and / or width compensation on the boundary region of the sub-image after perspective conversion.
[0155] In a specific example of the present disclosure solution, the processing module 802 is specifically configured to:
[0156] Expand a preset number of video frames before and / or after the sub-sequence.
[0157] In a specific example of the present disclosure solution, the processing module 802 is further configured to: after obtaining N noise reduction processing results, remove the boundary compensation part in the noise reduction processing results.
[0158] In a specific example of the present disclosure solution, the processing module 802 is specifically configured to:
[0159] Combine the N denoising processing results after removing the boundary compensation part to obtain a target rendering result for the target rendering task.
[0160] For the specific functions and example descriptions of the modules of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0161] Figure 9 It is a structural block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 9 shown, the electronic device includes: a memory 910 and a processor 920. The memory 910 stores a computer program that can run on the processor 920. The number of the memory 910 and the processor 920 can be one or more. The memory 910 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes the method provided in the above method embodiments. The electronic device may further include: a communication interface 930, configured to communicate with external devices and perform data interaction and transmission.
[0162] If the memory 910, the processor 920, and the communication interface 930 are implemented independently, the memory 910, the processor 920, and the communication interface 930 can be interconnected through a bus and complete communication with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0163] Optionally, in specific implementation, if the memory 910, the processor 920, and the communication interface 930 are integrated on a chip, the memory 910, the processor 920, and the communication interface 930 can complete communication with each other through an internal interface.
[0164] It should be understood that the above-mentioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0165] Further, optionally, the above-mentioned memory can include a read-only memory and a random access memory, and can also include a non-volatile random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can include a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can include a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Date SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).
[0166] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., Digital Versatile Disc (DVD)), or a semiconductor medium (e.g., Solid State Disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the present disclosure can be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0167] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.
[0168] In the description of the embodiments of the present disclosure, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0169] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. "And / or" herein is merely a description of the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0170] In the description of the embodiments of the present disclosure, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, "a plurality of" means two or more.
[0171] The above are only exemplary embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A distributed rendering processing method, comprising: Determine N sub-rendering tasks obtained by splitting the target rendering task; where N is an integer greater than or equal to 2; Calling M1 rendering nodes to render the N sub-rendering tasks in parallel using the M1 rendering nodes to obtain N sub-rendering results; M1 is an integer less than or equal to N; Calling M2 denoising nodes to perform denoising on the N sub-rendering results in parallel using the M2 denoising nodes to obtain N denoising processing results; the denoising node is a GPU node; M2 is an integer less than or equal to M1; Obtaining a target rendering result for the target rendering task based on the N denoising processing results; Before using M1 rendering nodes to render the N sub-rendering tasks in parallel, the method further includes: When the target rendering task is an image rendering task, the rendering node is used to perform height compensation and / or width compensation on the boundary area of the sub-image based on the image features of the sub-image targeted by the sub-rendering task to be processed by itself; wherein the sub-image is a partial area of the target image targeted by the image rendering task; the resolution of the target image is pixels; the resolution of the sub-image is Pixels, X and Y represent the preset segmentation values; The performing height compensation and / or width compensation on the boundary area of the sub-image includes: Performing a perspective conversion on the sub-image based on a target perspective; wherein the target perspective is related to a capture perspective of the target image; The height of the boundary area of the sub-image after the perspective conversion is compensated so that the actual rendering height of the sub-image after the height compensation is Pixels; The value range is ; is the height compensation value; and / or, The width of the sub-image after the perspective conversion is compensated so that the actual rendering width of the sub-image after the width compensation is Pixels; The value range is ; is the width compensation value.
2. The method according to claim 1, wherein: The M1 rendering nodes and the M2 denoising nodes are at least part of a plurality of processing nodes included in the rendering system.
3. The method according to claim 2, wherein: The multiple processing nodes include at least multiple CPU nodes and multiple GPU nodes; the M1 rendering nodes are at least part of the multiple CPU nodes and the multiple GPU nodes.
4. The method according to claim 1, wherein: The determining of N sub-rendering tasks obtained by splitting the target rendering task includes: When the target rendering task is an image rendering task, based on image features of the target image targeted by the image rendering task, split the target image into N sub-images to obtain N sub-rendering tasks; or, When the target rendering task is a video rendering task, the video frame sequence is split into N sub-sequences based on the number of frames of the video frame sequence targeted by the video rendering task to obtain N sub-rendering tasks.
5. The method according to any one of claims 1 to 4, wherein: Before using M1 rendering nodes to render the N sub-rendering tasks in parallel, the method further includes: When the target rendering task is a video rendering task, the rendering node is used to perform frame sequence compensation on the subsequence based on the number of frames of the subsequence targeted by the sub-rendering task that needs to be processed by itself, wherein the subsequence is a subsequence in the video frame sequence targeted by the video rendering task.
6. The method according to claim 5, wherein: The performing frame sequence compensation on the subsequence includes: Extend a preset number of video frames before and / or after the subsequence.
7. The method according to any one of claims 1 to 4, wherein: After obtaining N noise reduction processing results, the method further includes: Remove the boundary compensation part in the noise reduction result.
8. The method according to claim 7, wherein: The obtaining a target rendering result for the target rendering task based on the N denoising processing results includes: The N denoising results after removing the boundary compensation part are merged to obtain a target rendering result for the target rendering task.
9. A distributed rendering processing system, comprising: A scheduling node, used to determine N sub-rendering tasks obtained after splitting the target rendering task; wherein N is an integer greater than or equal to 2; and used to call M1 rendering nodes and M2 denoising nodes; the denoising node is a GPU node; M1 is an integer less than or equal to N; M2 is an integer less than or equal to M1; M1 rendering nodes, used to render the N sub-rendering tasks in parallel to obtain N sub-rendering results; M2 denoising nodes, used to perform denoising on the N sub-rendering results in parallel to obtain N denoising processing results; The rendering nodes in the M1 rendering nodes are further used to obtain a target rendering result for the target rendering task based on the N noise reduction processing results; The rendering nodes in the M1 rendering nodes are further used to perform height compensation and / or width compensation on the boundary area of the sub-image based on the image features of the sub-image targeted by the sub-rendering task to be processed by the rendering nodes themselves when the target rendering task is an image rendering task; wherein the sub-image is a partial area in the target image targeted by the image rendering task; and the resolution of the target image is pixels; the resolution of the sub-image is Pixels, X and Y represent the preset segmentation values; The rendering nodes in the M1 rendering nodes are specifically used to perform perspective conversion on the sub-image based on the target perspective, wherein the target perspective is related to the acquisition perspective of the target image; and to perform height compensation on the boundary area of the sub-image after the perspective conversion, so that the actual rendering height of the sub-image after the height compensation is Pixels; The value range is ; is a height compensation value; and / or, performing width compensation on the boundary area of the sub-image after the perspective conversion, so that the actual rendering width of the sub-image after the width compensation is Pixels; The value range is ; is the width compensation value.
10. A distributed rendering processing device, comprising: A splitting module, used to determine N sub-rendering tasks obtained by splitting the target rendering task; wherein N is an integer greater than or equal to 2; A processing module, configured to call M1 rendering nodes to render the N sub-rendering tasks in parallel using the M1 rendering nodes to obtain N sub-rendering results; M1 is an integer less than or equal to N; call M2 denoising nodes to denoise the N sub-rendering results in parallel using the M2 denoising nodes to obtain N denoising processing results; the denoising node is a GPU node; M2 is an integer less than or equal to M1; obtain a target rendering result for the target rendering task based on the N denoising processing results; Wherein, the processing module is also used for: When the target rendering task is an image rendering task, the rendering node is used to perform height compensation and / or width compensation on the boundary area of the sub-image based on the image features of the sub-image targeted by the sub-rendering task to be processed by itself; wherein the sub-image is a partial area of the target image targeted by the image rendering task; the resolution of the target image is pixels; the resolution of the sub-image is Pixels, X and Y represent the preset segmentation values; The processing module is specifically used to convert the sub-image according to the target viewing angle, wherein the target viewing angle is related to the acquisition viewing angle of the target image; and to perform height compensation on the boundary area of the sub-image after the viewing angle conversion, so that the actual rendering height of the sub-image after the height compensation is Pixels; The value range is ; is a height compensation value; and / or, performing width compensation on the boundary area of the sub-image after the perspective conversion, so that the actual rendering width of the sub-image after the width compensation is Pixels; The value range is ; is the width compensation value.
11. The device according to claim 10, wherein: The processing module is further used for: When the target rendering task is a video rendering task, the rendering node is used to perform frame sequence compensation on the subsequence based on the number of frames of the subsequence targeted by the sub-rendering task that needs to be processed by itself, wherein the subsequence is a subsequence in the video frame sequence targeted by the video rendering task.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image display method, electronic equipment and storage medium
CN116243880A
Three-dimensional scene model rendering method and device and distributed rendering server
CN117115326A
Medical image noise reduction method and device, computer equipment and storage medium
CN117333396A