Data rendering method of distributed rendering system and related equipment
Through the distributed rendering system, edge server clusters and cloud servers are used to jointly handle rendering tasks, solving the problem of inefficiency of centralized cloud servers and achieving efficient and real-time data rendering.
Patent Information
- Application Number
- CN202510142488.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-27
AI Technical Summary
Centralized cloud servers are less efficient when dealing with the rendering needs of a large number of concurrent user terminals, resulting in data rendering waiting and unable to meet the needs of efficient real-time rendering.
A distributed rendering system is adopted, and edge server clusters and cloud servers are used to handle rendering tasks together. By analyzing the rendering requirements instructions, the rendering tasks are divided into multiple subtasks, and under the optimization of the resource allocation model, it is distributed to multiple rendering nodes for parallel processing, and the final merge result is sent to the terminal user.
It improves the efficiency and reliability of data rendering, reduces the rendering waiting time for end users, and makes full use of the parallel computing power of edge server clusters.
Smart Images

Figure CN120216163A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data rendering method and related devices for a distributed rendering system. Background Art
[0002] Under the application framework of the metaverse, each participating user can produce and edit 3D digital content, and real-time rendering will serve as an important technology to provide functional support and task guarantee for metaverse applications. With the popularization of applications related to Extended Reality (XR) technology, the demand for large-scale image rendering and real-time interaction will become increasingly important and urgent. Therefore, it is particularly important to develop a suitable rendering method to improve the data rendering rate.
[0003] In related technologies, in order to further improve the data rendering rate, the rendering process of the user terminal is usually placed on a cloud server with high computing and processing capabilities for data rendering, and after the data rendering is completed, the data rendering result is transmitted back to the user terminal for display. However, as the number of user terminals performing rendering tasks further increases, and the types of rendering tasks between each user terminal may also be different, the centralized cloud server cannot efficiently meet the rendering needs of a large number of concurrent user terminals due to the limited number of open rendering instances. Therefore, data rendering waiting often occurs, resulting in low efficiency of data rendering using only a centralized cloud server. Summary of the Invention
[0004] Embodiments of this application provide a data rendering method and related devices for a distributed rendering system, which can improve the data rendering efficiency.
[0005] To achieve the above object, a first aspect of the embodiments of this application proposes a data rendering method for a distributed rendering system. The distributed rendering system includes an edge server cluster and a cloud server. The edge server cluster includes multiple rendering nodes. The method is applied to the edge server cluster and includes:
[0006] Respond to a rendering requirement instruction corresponding to a rendering task issued by a terminal user, and parse the rendering requirement instruction to obtain rendering requirement information;
[0007] Based on the rendering requirement information and the number of rendering nodes, divide the rendering task into multiple rendering subtasks, and send the rendering subtasks to the cloud server. The rendering subtasks include subtask data and one of the rendering nodes as a working rendering node;
[0008] Input all the sub-task data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing to obtain the optimized working parameters and optimized image pixels for each working rendering node;
[0009] Receive the sub-rendering data sent by the cloud server for each rendering sub-task, and perform data rendering on the corresponding working rendering node based on the sub-rendering data, the optimized working parameters, and the optimized image pixels to obtain the rendering sub-results corresponding to the rendering sub-tasks;
[0010] Merge all the rendering sub-results to obtain the target rendering result of the rendering task, and send the target rendering result to the end user.
[0011] In some embodiments, the rendering requirement information includes a plurality of view frustums, the number of view frustum pixels for each view frustum, and the total number of rendering pixels. The dividing of the rendering task based on the rendering requirement information and the number of rendering nodes to obtain a plurality of rendering sub-tasks includes:
[0012] Divide all the view frustums equally to obtain a plurality of first view frustum subsets, and generate a pixel rendering ratio factor based on the reciprocal of the number of rendering nodes;
[0013] When the number of the first subsets is less than the number of rendering nodes, based on the sum of the number of view frustum pixels of all the view frustums in the first view frustum subsets, obtain the first divided pixel number, and based on the ratio of the first divided pixel number to the total number of rendering pixels, obtain the first divided pixel ratio;
[0014] Based on the first divided pixel ratio, the pixel rendering ratio factor, the plurality of first view frustum subsets, and the number of rendering nodes, obtain a plurality of the rendering sub-tasks.
[0015] In some embodiments, the obtaining of a plurality of the rendering sub-tasks based on the first divided pixel ratio, the pixel rendering ratio factor, the plurality of first view frustum subsets, and the number of rendering nodes includes:
[0016] When the first divided pixel ratio is less than the pixel rendering ratio factor, generate the plurality of rendering sub-tasks based on the plurality of first view frustum subsets;
[0017] When the first divided pixel ratio is not less than the pixel rendering ratio factor, divide all the view frustums in each first view frustum subset equally to obtain a plurality of second view frustum subsets, and based on the ratio of the sum of the number of view frustum pixels of all the view frustums in the second view frustum subsets to the total number of rendering pixels, obtain the second divided pixel ratio;
[0018] When the second divided pixel ratio is less than the pixel rendering ratio factor, generate the multiple rendering subtasks based on the multiple second frustum subsets.
[0019] In some embodiments, the subtask data includes a rendering number, and merging all the rendering sub-results to obtain the target rendering result of the rendering task includes:
[0020] Receive the rendering completion instructions sent by all the working rendering nodes;
[0021] Based on the rendering number, sequentially splice each of the rendering sub-results to obtain the target rendering result.
[0022] In some embodiments, the generating step of the resource allocation model includes:
[0023] Generate a state space based on the node hardware parameters of the multiple rendering nodes and multiple subtask data parameters;
[0024] Generate an action space based on the working parameters of the multiple rendering nodes and image pixel parameters;
[0025] Obtain the rendering time function corresponding to the execution of the rendering task, and generate a reward model based on the rendering time function;
[0026] Obtain the initial action network parameters of the action network and the initial evaluation network parameters of the evaluation network;
[0027] Generate an initial resource allocation model based on the state space, the action space, the reward model, the initial action network parameters, and the initial evaluation network parameters, perform multiple rounds of training on the initial resource allocation model, and obtain the resource allocation model based on the trained initial resource allocation model.
[0028] In some embodiments, the obtaining the rendering time function corresponding to the execution of the rendering task includes:
[0029] Obtain the communication delay function during the execution of the rendering task, where the communication delay function includes the first communication time between the end user and the edge server cluster and the second communication time for dividing the rendering task;
[0030] Obtain the maximum rendering time function of the multiple rendering nodes, and obtain the rendering time function based on the maximum rendering time function and the communication delay function.
[0031] In some embodiments, the training step of the initial resource allocation model includes:
[0032] Obtain the rendering state parameters of the current iteration, input the rendering state parameters into the action network to obtain resource action parameters, and obtain updated rendering state parameters based on the resource action parameters;
[0033] Input the rendering state parameters and the updated rendering state parameters into the evaluation network to obtain the state value corresponding to the rendering state parameters and the updated state value corresponding to the updated rendering state parameters;
[0034] Update the initial action network parameters and the initial evaluation network parameters based on the mean square error, the state value, the updated state value, and the resource action parameters.
[0035] To achieve the above object, a second aspect of the embodiments of the present application proposes a data rendering device for a distributed rendering system. The distributed rendering system includes an edge server cluster and a cloud server. The edge server cluster includes multiple rendering nodes. The method is applied to the edge server cluster. The device includes:
[0036] A response parsing module, configured to parse the rendering requirement instruction corresponding to the rendering task issued by the end user to obtain rendering requirement information;
[0037] A task division module, configured to divide the rendering task based on the rendering requirement information and the number of rendering nodes to obtain multiple rendering subtasks, and send the rendering subtasks to the cloud server. The rendering subtasks include subtask data and one of the rendering nodes as the working rendering node;
[0038] A resource optimization module, configured to input all the subtask data and the node hardware parameters corresponding to the working rendering node into a resource allocation model for data processing to obtain optimized working parameters and optimized image pixels for each working rendering node;
[0039] A rendering processing module, configured to receive the sub-rendering data sent by the cloud server for each rendering subtask, and perform data rendering on the corresponding working rendering node based on the sub-rendering data, the optimized working parameters, and the optimized image pixels to obtain a rendering sub-result corresponding to the rendering subtask;
[0040] A combination module, configured to merge all the rendering sub-results to obtain a target rendering result of the rendering task, and send the target rendering result to the end user.
[0041] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the data rendering method of the distributed rendering system as described in the first aspect.
[0042] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the data rendering method of the distributed rendering system as described in the first aspect above.
[0043] The data rendering method and related devices proposed by the embodiments of the present application. The distributed rendering system includes an edge server cluster and a cloud server. The edge server cluster includes multiple rendering nodes. The method is applied to the edge server cluster and includes: First, in response to a rendering requirement instruction corresponding to a rendering task issued by an end user, parse the rendering requirement instruction to obtain rendering requirement information; Second, based on the rendering requirement information and the number of rendering nodes, divide the rendering task into multiple rendering subtasks, and send the rendering subtasks to the cloud server. The rendering subtasks include subtask data and a rendering node serving as a working rendering node; Then, input all the subtask data and the node hardware parameters corresponding to the working rendering node into a resource allocation model for data processing to obtain the optimized working parameters and optimized image pixels of each working rendering node; Next, receive the sub-rendered data sent by the cloud server for each rendering subtask, and perform data rendering on the corresponding working rendering node based on the sub-rendered data, optimized working parameters, and optimized image pixels to obtain the rendering sub-results corresponding to the rendering subtasks; Finally, merge all the rendering sub-results to obtain the target rendering result of the rendering task, and send the target rendering result to the end user. The embodiments of the present application use multiple edge server clusters near the end user as processing rendering devices to reduce the transmission delay of rendering data; In addition, use a larger number of edge server clusters compared to the cloud server to execute rendering tasks for end users, thereby reducing the possibility of end users waiting for data rendering; And, appropriately divide the rendering task into subtasks based on the number of rendering nodes in the edge server cluster, so that the rendering task can be processed simultaneously among multiple rendering nodes to improve the processing efficiency of data rendering; And, after dividing multiple rendering subtasks, use a resource allocation model pre-constructed with the goal of improving data rendering efficiency to plan appropriate optimized working parameters and optimized image pixels for each rendering node, thereby greatly improving the efficiency and reliability of data rendering in the distributed rendering system.
[0044] Other features and advantages of the present application will be described in the subsequent specification, and will, in part, be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification, the claims, and the drawings. Description of the Drawings
[0045] Figure 1 is a schematic structural diagram of a distributed rendering system provided by an embodiment of the present application.
[0046] Figure 2 is a flowchart of a data rendering method for a distributed rendering system provided by another embodiment of the present application.
[0047] Figure 3 is Figure 2 a flowchart of step 202 in
[0048] Figure 4 is Figure 3 a flowchart of step 303 in
[0049] Figure 5 is a flowchart block diagram of a load balancing task partitioning strategy based on image pixel data provided by another embodiment of the present application.
[0050] Figure 6 is a flowchart of generating a resource allocation model provided by another embodiment of the present application.
[0051] Figure 7 is Figure 6 a flowchart of step 603 in
[0052] Figure 8 is a schematic structural diagram of a decision neural network provided by another embodiment of the present application.
[0053] Figure 9 is a working system diagram of an initial resource allocation model provided by an embodiment of the present application.
[0054] Figure 10 is a specific concept schematic diagram when implementing a traditional A3C algorithm and the A3C algorithm provided by the present application in an embodiment of the present application.
[0055] Figure 11 is Figure 2 a flowchart of step 205 in
[0056] Figure 12 is a schematic structural diagram of a data rendering device for a distributed rendering system provided by an embodiment of the present application.
[0057] Figure 13It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0058] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0061] Under the application framework of the metaverse, each participating user can produce and edit 3D digital content, and real-time rendering will serve as an important technology to provide functional support and task guarantee for metaverse applications. Therefore, the demand for high-fidelity large-scale image rendering and real-time interaction based on XR applications will become increasingly important and urgent. Traditional computer image rendering uses a single machine to perform distributed 3D drawing using multi-pipeline resources between the CPU and GPU, and this method has been widely used in a variety of mature commercial software. Due to the increasing complexity of image models and the transmission and communication delays caused by large-scale data transmission, the distributed network architecture combined with edge computing technology, using the physical structure of the network and resource scheduling algorithms to achieve parallel and rapid iterative updates of large-scale data and tasks, is an important research content and direction for improving network efficiency. To further improve the image rendering efficiency, enhance the transmission efficiency of large-scale high-fidelity audio and video, and enhance the interaction performance in virtual reality applications, the research on distributed accelerated rendering technology has received extensive attention in the industrial and academic fields in recent years.
[0062] The frame rate of the screen in the virtual world usually requires no less than 30fps. Such a high refresh rate requires greater computing resources in the face of real-time transmission scenarios of high-quality images. In the application background of the metaverse, with the improvement of the model scale and image quality, the 3D model calculation and rendering tasks will face huge challenges. Although high-performance GPUs can relieve the pressure of rendering tasks, it is still very difficult to complete real-time rendering tasks by implementing the rendering of the entire large-scale and highly realistic scene on a single computing resource. Therefore, distributed rendering is carried out using a distributed parallel computing network. By splitting the entire rendering task and assigning each subtask to different computing nodes, the rendering efficiency can be effectively improved and the quality of the network-rendered images can be enhanced.
[0063] In related technologies, in order to further improve the data rendering rate, the rendering process of the user terminal is usually placed on a cloud server with high computing processing capabilities for data rendering. After the data rendering is completed, the data rendering result is then transmitted back to the user terminal for display. However, as the number of user terminals performing rendering tasks further increases, and the types of rendering tasks between each user terminal may also be different, due to the limited number of open rendering instances of the centralized cloud server, it cannot efficiently meet the rendering requirements of a large number of concurrent user terminals. Therefore, data rendering waiting often occurs, resulting in low efficiency of data rendering using only a centralized cloud server. In addition, real-time rendering of high-quality 3D digital scenes requires large professional graphics cards and other hardware devices, which consume a large amount of cloud computing resources and bring great rendering pressure to the rendering server. A large number of 3D models need to store a large amount of data on the cloud server, and it is difficult to store model files on a single machine.
[0064] To improve the data rendering efficiency, the embodiments of the present application use multiple edge server clusters near the end user as processing rendering devices to reduce the transmission delay of rendering data; in addition, use a larger number of edge server clusters compared to cloud servers to perform rendering tasks for end users, thereby reducing the possibility of data rendering waiting for end users; and, appropriately divide subtasks for the rendering task based on the number of rendering nodes of the edge server cluster, so that the rendering task can be processed collaboratively among multiple rendering nodes simultaneously to improve the processing efficiency of data rendering; and, after dividing multiple rendering subtasks, use a resource allocation model constructed in advance with the goal of improving data rendering efficiency to plan appropriate optimization working parameters and optimized image pixels for each rendering node, thereby greatly improving the efficiency and reliability of data rendering in the distributed rendering system.
[0065] To better describe the data rendering method and related devices of the distributed rendering system provided by the present application, the distributed rendering system of the data rendering method applied to the distributed rendering system is first described below. Refer toFigure 1 , is a schematic structural diagram of a distributed rendering system provided by an embodiment of the present application. As Figure 1 shown in, the distributed rendering system includes an edge server cluster and a cloud server. Among them, the edge server cluster includes edge servers, a front-end middleware cluster, and a rendering cluster.
[0066] Among them, the edge server is used to communicate with the terminal device, receive the rendering requirement instruction corresponding to the real-time rendering task of the terminal device, and then parse the rendering requirement instruction and transmit it to the front-end middleware cluster.
[0067] After receiving the parsed rendering requirement instruction, the front-end middleware cluster splits the rendering task to obtain multiple rendering task agents (i.e., rendering subtasks), and then sends these rendering subtasks to the corresponding rendering nodes in the rendering cluster. The front-end middleware cluster is also used to combine the rendering sub-results corresponding to all rendering subtasks to obtain the rendering result corresponding to the rendering task, and send the rendering result to the terminal device for display through the edge server.
[0068] The rendering cluster includes multiple rendering nodes that can be used to execute different rendering tasks. Each rendering node includes engine adaptation, a rendering engine, a cache, and rendering acceleration. After receiving the corresponding rendering subtask, these rendering nodes obtain the corresponding rendering data from the cloud server and perform the rendering of the corresponding task data. When the rendering is completed, the rendering sub-results are sent back to the front-end middleware cluster.
[0069] The cloud server is used to provide the original data to multiple rendering nodes in the rendering cluster. The rendering nodes themselves will also generate data to improve the rendering efficiency.
[0070] Based on the above-mentioned distributed rendering system, the data rendering method and related devices provided by the embodiments of the present application will be further described below. The data rendering method provided by the embodiments of the present application can be applied to the edge server cluster in the distributed rendering system.
[0071] The data rendering method of the distributed rendering system in the embodiments of the present application will be specifically described below. Refer to Figure 2 , which is an optional flowchart of the data rendering method of the distributed rendering system provided by the embodiments of the present application. Figure 2 The method in may include but is not limited to steps 201 to 205. At the same time, it can be understood that the order of steps 201 to 205 in this embodiment is not specifically limited, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added. Figure 2 The order of steps 201 to 205 in is not specifically limited, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.
[0072] Step 201: In response to a rendering requirement instruction corresponding to a rendering task issued by an end user, parse the rendering requirement instruction to obtain rendering requirement information.
[0073] The following provides a detailed description of Step 201.
[0074] In some embodiments, after an end user generates a rendering task based on display screen requirements or image rendering requirements, the rendering requirement instruction corresponding to the rendering task is sent to the edge server cluster for data rendering.
[0075] When the edge server cluster receives the rendering requirement instruction, it first parses the rendering requirement instruction at the edge server or the front-end middleware cluster to obtain rendering requirement information that clarifies the user's needs. These rendering requirement information include the user's location, orientation, area of interest, multiple viewing frustums that need to be rendered, the number of pixels in each viewing frustum, and the total number of rendering pixels, etc.
[0076] Step 202: Based on the rendering requirement information and the number of rendering nodes, divide the rendering task into multiple rendering subtasks, and send the rendering subtasks to the cloud server.
[0077] The following provides a detailed description of Step 202.
[0078] In some embodiments, after obtaining the rendering requirement information at the front-end middleware cluster of the edge server cluster, the rendering task of the end user will be further divided according to the rendering requirement information and the number of rendering nodes corresponding to all rendering nodes in the rendering cluster to obtain multiple appropriate rendering subtasks and determine the corresponding rendering nodes (i.e., working rendering nodes), so as to facilitate subsequent sending the subtask data corresponding to these rendering subtasks to the corresponding working rendering nodes for parallel processing of multiple rendering subtasks simultaneously, and sending them to the cloud server so that the cloud server can send the corresponding rendering data to each working rendering node, thereby greatly improving the efficiency of data rendering.
[0079] Among them, the working rendering node refers to each rendering node that processes the corresponding rendering subtask.
[0080] The following will further describe how to appropriately divide the rendering task.
[0081] Refer to Figure 3 , divide the rendering task into multiple rendering subtasks based on the rendering requirement information and the number of rendering nodes, including the following steps 301 to 303.
[0082] Step 301: Divide all viewing frustums evenly to obtain multiple first viewing frustum subsets, and generate a pixel rendering ratio factor based on the reciprocal of the number of rendering nodes.
[0083] Step 302: When the number of the first subsets is less than the number of rendering nodes, based on the sum of the cone pixel numbers of all the cones in the first frustum subsets, obtain the first divided pixel number, and based on the ratio of the first divided pixel number to the total number of rendering pixels, obtain the first divided pixel ratio.
[0084] Step 303: Based on the first divided pixel ratio, the pixel rendering ratio factor, multiple first frustum subsets, and the number of rendering nodes, obtain multiple rendering subtasks.
[0085] The following gives a detailed description of Steps 301 to 303.
[0086] In some embodiments, after obtaining the rendering requirement information and the number of rendering nodes N, the preposed middleware cluster evenly divides all the frustums of the rendering task into two parts to obtain two first frustum subsets of the first division; then compare the number of the first subsets and the number of rendering nodes N.
[0087] In addition, the preposed middleware cluster also generates a pixel rendering ratio factor based on the reciprocal of the number of rendering nodes N, that is, 1 / N.
[0088] When the number of the first subsets (i.e., 2) is greater than or equal to the number of rendering nodes N, at this time, the number of available rendering nodes in the rendering cluster is relatively small, that is, the rendering parallel ability of the edge server cluster is limited. At this time, directly generate corresponding rendering subtasks for the rendering pixels required by each divided first frustum subset, and allocate available rendering nodes to these rendering subtasks one by one as the corresponding working rendering nodes.
[0089] When the number of the first subsets (i.e., 2) is less than the number of rendering nodes N, at this time, the number of available rendering nodes in the rendering cluster is relatively large, and it can be further determined whether it is necessary to further divide the frustums to split out more rendering subtasks, so as to utilize the parallel processing ability of more rendering nodes to improve the data rendering rate.
[0090] Based on this situation, further, based on the sum of the cone pixel numbers of all the cones in each first frustum subset, obtain the corresponding first divided pixel number Nl in each first frustum subset, and based on the ratio of the first divided pixel number Nl to the total number of rendering pixels Nm, obtain the first divided pixel ratio Nl / Nm.
[0091] Next, further based on the first divided pixel ratio Nl / Nm, the pixel rendering ratio factor 1 / N, multiple first frustum subsets, and the number of rendering nodes N, determine whether it is necessary to further divide the first frustum subsets after the initial division, that is, determine whether it is necessary to further divide the rendering task, which is specifically described as follows.
[0092] Reference Figure 4 , based on the first division pixel ratio, the pixel rendering ratio factor, multiple first frustum subsets, and the number of rendering nodes, multiple rendering subtasks are obtained, including the following steps 401 to 403.
[0093] Step 401: When the first division pixel ratio is less than the pixel rendering ratio factor, multiple rendering subtasks are generated based on multiple first frustum subsets.
[0094] Step 402: When the first division pixel ratio is not less than the pixel rendering ratio factor, all frustums in each first frustum subset are evenly divided to obtain multiple second frustum subsets. Based on the ratio of the sum of the frustum pixel numbers of all frustums in the second frustum subset to the total number of rendering pixels, the second division pixel ratio is obtained.
[0095] Step 403: When the second division pixel ratio is less than the pixel rendering ratio factor, multiple rendering subtasks are generated based on multiple second frustum subsets.
[0096] The following is a detailed description of steps 401 to 403.
[0097] In some embodiments, for the case where the number of first subsets is less than the number of rendering nodes N, the corresponding first division pixel ratio Nl / Nm within each first frustum subset is further compared with the pixel rendering ratio factor 1 / N to determine whether further frustum subdivision is required.
[0098] When the first division pixel ratio Nl / Nm at this time is less than the pixel rendering ratio factor 1 / N, the proportion of the rectangular area of the rendering image that needs to be pixel-rendered in the total proportion within the evenly divided first frustum subset at this time is already relatively small. At this time, there is no need to further divide the frustum, and corresponding rendering subtasks can be directly generated based on the rendering pixels required for each divided first frustum subset, and available rendering nodes are assigned to these rendering subtasks one by one as the corresponding working rendering nodes.
[0099] When the first division pixel ratio Nl / Nm at this time is not less than the pixel rendering ratio factor 1 / N, the proportion of the rectangular area of the rendering image that needs to be pixel-rendered in the total proportion within the evenly divided first frustum subset at this time is still relatively large. At this time, further frustum division is required. Next, all frustums in each first frustum subset are evenly divided again to obtain multiple second frustum subsets; then, it is determined again whether the number of second frustums in the second frustum subset exceeds the number of rendering nodes N. When the number of second frustums exceeds the number of rendering nodes N, multiple rendering subtasks are directly generated based on the divided multiple second frustum subsets.
[0100] When the number of second frustums does not exceed the number N of rendering nodes, again, according to the ratio of the sum Nl2 of the frustum pixel numbers of all frustums in the second frustum subset to the total number of rendering pixels, the second division pixel ratio Nl2 / Nm is obtained. Then, the second division pixel ratio Nl2 / Nm is compared with the pixel rendering ratio factor 1 / N to determine again whether further division is required.
[0101] Refer to Figure 5 , which is a flowchart of a load balancing task division strategy based on image pixel data provided by an embodiment of the present application. As Figure 5 shown, for the real-time rendering task requirements of end users, the embodiment of the present application adopts a rendering task division and screen division scheme based on a load balancing strategy, specifically including: Real-time distributed rendering has high requirements for real-time performance and low requirements for precise division of load balancing. Therefore, the adopted task division strategy is based on coarse-grained load balancing among rendering nodes, uses the number of pixels to estimate the computing load, and the task division algorithm process of the task scheduling module is as Figure 5 shown. If there are N distributed rendering nodes in the rendering cluster, the total number of rectangular areas of the rendering image to be divided is less than N, and each distributed rendering node renders the corresponding divided rectangular area. Before rendering each frame of the image, it is necessary to re-divide the image data space to ensure that the computing load of each rectangular area is roughly the same, that is, the number of pixels in each rectangular area is of roughly the same order of magnitude.
[0102] Through the above steps 301 to 303, and steps 401 to 403, using the number of rendering nodes corresponding to the available rendering nodes in the rendering cluster as the division matching factor, the frustums of the rendering tasks are gradually and equally divided until multiple rendering subtasks that can make full use of the parallel computing power of the rendering cluster are obtained, thereby effectively improving the data rendering efficiency of the distributed system.
[0103] In this embodiment, an adaptive task division strategy is adopted to perform pixel division on each frame of the image in the rendering task. Further, the simplest embodiment of this strategy is to divide the image into lines, and perform pixel division according to the line sequence scheduler. The rendering node number corresponds to the line number of the divided rendering subtask. When the number of divided lines exceeds the number of rendering nodes, the next rendering subtask of the first rendering node is to render the line numbered (1 + N).
[0104] Step 203: Input all subtask data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing to obtain the optimized working parameters and optimized image pixels of each working rendering node.
[0105] The following provides a detailed description of step 203.
[0106] In some embodiments, in order to further improve the data rendering efficiency of the distributed rendering system, after the rendering tasks are partitioned, resource management and allocation are also performed. The operations of resource management and allocation include the working frequency, bandwidth, memory size, time, etc. of the CPU of each working rendering node when executing the rendering subtasks.
[0107] To improve the network rendering computing performance and resource utilization rate, the embodiment of the present application is based on a deep reinforcement learning scheme of the asynchronous advantage actor-critic algorithm (A3C). Utilizing the characteristic that the A3C algorithm network structure contains multiple agents that match multiple rendering nodes, a resource allocation model corresponding to the rendering task processing of the distributed rendering system is pre-constructed to optimize the resource management of multiple working rendering nodes in the rendering cluster for processing rendering subtasks, so as to make full use of the multi-core structure of the CPUs of multiple working rendering nodes for parallel acceleration calculation, thereby improving the resource optimization efficiency.
[0108] It can be understood that the embodiment of the present application can also adopt other reinforcement learning algorithms to construct the resource allocation model, such as the A2C algorithm, DQN algorithm, DDPG algorithm, PPO algorithm, and so on.
[0109] Next, how to construct the resource allocation model will be further described.
[0110] Refer to Figure 6 , the generation steps of the resource allocation model include the following steps 601 to step 605.
[0111] Step 601: Generate a state space based on the node hardware parameters of multiple rendering nodes and multiple subtask data parameters.
[0112] Step 602: Generate an action space based on the working parameters of multiple rendering nodes and the image pixel parameters.
[0113] Step 603: Obtain the rendering time function corresponding to the execution of the rendering task, and generate a reward model based on the rendering time function.
[0114] Next, steps 601 to 603 will be described in detail.
[0115] In some embodiments, first, a state space of the A3C algorithm framework is generated based on the node hardware parameters (including the working frequency, bandwidth, etc. of the CPU) of each rendering node in the rendering cluster, multiple subtask data parameters (including the number of image bits required for the divided rendering subtasks to be processed), and the number of available rendering nodes.
[0116] Then, an action space of the A3C algorithm framework is generated based on adjustable working parameters (including frequency and bandwidth) of each rendering node in the rendering cluster and image pixel parameters adopted when processing rendering subtasks.
[0117] In addition, in order to further fully improve the data rendering efficiency of real-time rendering tasks, a reward model for training the A3C algorithm framework will be generated based on the required rendering time function corresponding to the execution of the rendering task. How to determine this rendering time function will be further described below.
[0118] Referring to Figure 7 , obtaining the rendering time function corresponding to the execution of the rendering task includes the following steps 701 to 702.
[0119] Step 701: Obtain the communication delay function during the execution of the rendering task.
[0120] Step 702: Obtain the maximum rendering time function of multiple rendering nodes, and obtain the rendering time function based on the maximum rendering time function and the communication delay function.
[0121] In some embodiments, the time used for data communication between the end user and the edge server during the execution of the rendering task in the distributed rendering system is used as the first communication time t1, the time used for data communication between the edge server and the front-end middleware cluster is used as the intermediate communication time t2, and the time used for the front-end middleware cluster to perform rendering task segmentation, send rendering subtasks, and splice and synthesize rendering sub-results is used as the second communication time t3.
[0122] Then, based on the sum of the first communication time t1, the intermediate communication time t2, and the second communication time t3, the communication delay function T = t1 + t2 + t3 during the execution of the rendering task in the distributed rendering system is obtained.
[0123] Next, when multiple rendering nodes in the rendering cluster process data of rendering subtasks in the refrigerator, the maximum rendering time function R is determined by the maximum value of the rendering time-consuming among N parallel rendering nodes, and its mathematical formula is expressed as the following formula.
[0124]
[0125] Where r n represents the time taken for the nth rendering node to complete the rendering task assigned by the front-end middleware.
[0126] Then, based on the sum of the maximum rendering time function and the communication delay function, the rendering time function that can be used to evaluate the performance index of the distributed rendering system is obtained.
[0127] Through the above steps 701 to 702, the rendering time function obtained by using the communication delay function for data transmission, task division, and sub-result synthesis in the distributed rendering system and the maximum rendering time function corresponding to the parallel processing of multiple rendering nodes is used as a performance metric for evaluating the distributed rendering system. This facilitates subsequent model training of the A3C algorithm using the reward model generated from this rendering time function, and can effectively improve the reliability of the obtained resource allocation model in optimizing the rendering efficiency.
[0128] Step 604: Obtain the initial action network parameters of the action network and the initial evaluation network parameters of the evaluation network.
[0129] Step 605: Generate an initial resource allocation model based on the state space, action space, reward model, initial action network parameters, and initial evaluation network parameters, perform multiple rounds of training on the initial resource allocation model, and obtain a resource allocation model based on the trained initial resource allocation model.
[0130] The following provides a detailed description of steps 604 to 605.
[0131] Refer to Figure 8 , which is a schematic structural diagram of a decision neural network provided by an embodiment of the present application. As shown in Figure 8 , in the A3C algorithm framework, two fully connected neural networks, namely the action (Actor) network and the evaluation (Critic) network, are used for decision-making. Therefore, it is also necessary to generate the initial action network parameters of the action network and the initial evaluation network parameters of the evaluation network.
[0132] Then, generate an initial resource allocation model based on the state space, action space, reward model, initial action network parameters, and initial evaluation network parameters that match the distributed rendering system. Then, perform multiple rounds of training on the initial resource allocation model, and obtain the actually used resource allocation model based on the action (Actor) network in the trained initial resource allocation model.
[0133] The following will further describe how to train this initial resource allocation model.
[0134] Refer to Figure 9 , the training steps of the initial resource allocation model include the following steps 801 to 803.
[0135] Step 901: Obtain the rendering state parameters of the current iteration, input the rendering state parameters into the action network to obtain resource action parameters, and obtain updated rendering state parameters based on the resource action parameters.
[0136] Step 902: Input the rendering state parameters and the updated rendering state parameters into the evaluation network to obtain the state value corresponding to the rendering state parameters and the updated state value corresponding to the updated rendering state parameters.
[0137] Step 903: Update the initial action network parameters and the initial evaluation network parameters based on the mean square error, the state value, the updated state value, and the resource action parameters.
[0138] The following provides a detailed description of Steps 901 to 903.
[0139] In some embodiments, the specific logical process of the training process of the initial resource allocation model based on the A3C algorithm includes the following steps.
[0140] The algorithm inputs include: the number of iteration rounds T, the state feature dimension n, where there are p state values (t1, t2,..., t p p), that is, p rendering resource coefficients, the action value A, (where there are p action values (Δt1, Δt2,..., Δt p p)), that is, the change amounts of 4 rendering resource coefficients (the working frequency f, bandwidth B, number of image bits b, and number of rendering nodes N of each rendering node), the learning rates α and β of the two networks, the decay factor γ, the exploration rate ∈, the evaluation network structure, and the action network structure.
[0141] The algorithm outputs include: minimizing the overall rendering delay, the optimal action network parameters θ', and the optimal evaluation network parameters θ v '.
[0142] Initialization: Randomly initialize the resource coefficients (state, S), the overall rendering delay and its regularization term (reward, R), and the change amount of the resource coefficients (action, A).
[0143] Iteration: For i in [1, T], the iteration is described in detail as follows.
[0144] (1) In each iteration, input the rendering state parameters S of the current iteration into the action network to obtain the resource action parameters A output by the action network, and obtain new resource coefficients (i.e., updated rendering state parameters S') based on the resource action parameters A, and feedback the overall rendering delay and its regularization term (T + R + λS); λ is the regularization parameter.
[0145] (2) The evaluation network uses the rendering state parameters (S) and the updated rendering state parameters (S') as inputs, and outputs the state value V(S) corresponding to the rendering state parameters and the updated state value evaluation value V(S') corresponding to the updated rendering state parameters, where V is the output value of the evaluation network.
[0146] (3) Calculate the time difference error: δ = R + YV(S’) - V(S); Y is the decay factor.
[0147] (4) Use the mean square error method to calculate the loss of the evaluation network and update the weight parameters θv’ of the Critic network.
[0148] Among them, the loss function of the evaluation network is: loss = ∑(T + R + YV(S’) - V(S))2.
[0149] The update formula for the initial evaluation network weight parameters is: is the derivative symbol.
[0150] (5) Calculate the loss of the action network and update the weight parameters θ’ of the action network.
[0151] Among them, the loss function of the action network is: loss = logπθ’(S t , A)R.
[0152] The initial action network parameter update formula is:
[0153] (6) Record the minimum rendering delay and the corresponding optimal state up to this iteration.
[0154] Return: the minimum overall rendering delay, the optimal state, the Critic network parameter θ v ’, the Actor network weight parameter θ’.
[0155] The above is the AC algorithm provided by the embodiments of this application. It should be noted that for the update of the action network, the loss function and the network parameter update formula are changed on the basis of the original A3C algorithm.
[0156] It can be understood that the original loss function of the action network is: loss = logπθ’(S t , A).
[0157] The original action network weight parameter update formula is:
[0158] The changes in these two formulas are to multiply the overall rendering delay (T + R) in the loss function or the network parameter iteration formula. These two changes bind the direct correlation between the action and the rendering delay and accelerate the training convergence speed of the initial resource allocation network.
[0159] Based on the above training steps, the detailed training schematic process of the initial resource allocation model will be further described in detail below.
[0160] The embodiments of this application optimize the design for minimizing the overall latency of a distributed rendering system composed of cloud, edge, and terminal. For the first time, a machine learning method is used to minimize the total latency of the distributed rendering network. The generation steps of the constructed resource allocation model include the following descriptions.
[0161] (1) Set the value range of the resource coefficient according to the material parameters of the distributed rendering system: the working frequency f and bandwidth B of each rendering node's CPU, the number of image bits b of the rendering subtask after task division, the number of rendering nodes N, and the overall rendering latency.
[0162] (2) Take the rendering resource coefficient as p state values (t1, t2,..., t p ) of the A3C algorithm, take the change amount of the rendering resource coefficient as p action values (Δt1, Δt2,..., Δt p ) of the A3C algorithm, and take the overall rendering latency and its regularization term as the reward function (T + R + λS).
[0163] (3) Initialize the global neural network parameters and local neural network parameters in the set A3C algorithm model respectively. Among them, the parameters of the Actor network and the Critic network in the global neural network are represented by θ and θ v respectively; the parameters of the Actor network and the Critic network in the local neural network are represented by θ' and θ v ' respectively, accumulate the gradients dθ ← 0 and dθ v ← 0, and finally obtain the initialized A3C algorithm model (i.e., the initial resource allocation model).
[0164] (4) The local neural network interacts with the environment. Input the current state set s t = {t1, t2,..., t p} into the local neural network. According to the current policy π(ax|s t , θ'), obtain the action a t . To better explore, a Gaussian exploration strategy is adopted at each moment. Therefore, the output of the Actor network consists of two parts. One is the average value μ(s t ; θ') of the action, and the other output is the standard deviation σ(s t ; θ') of the action. Then, the change amount a t of the resource coefficient is randomly sampled from the normal distribution N(μ(s t ; θ'), σ(s t ; θ')). Apply the current action a t to the distributed rendering minimum latency model to obtain the current reward r t and the state set s of the next momentt+1 ; Finally, based on the above relevant data s t , a t , r t and s t+1 , a training dataset (s t , a t , r t , s t+1 ) is obtained for the training process of the neural network.
[0165] (5) The evaluation network uses the resource coefficient (S) and the updated resource coefficient (S’) as inputs, outputs the evaluation values V(S) and V(S’) of the network for these two states, and calculates the TD error: δ = T + R + λS + γV(S’) - V(S).
[0166] (6) Based on the original Actor-Critic algorithm, the direct influence of the reward function (T + R + λS) on it is considered, and the loss function and the network parameter update formula are changed. The formula is:
[0167] where loss is the loss function of the Actor network, πθ’ is the probability corresponding to the Gaussian probability sampling action in the Actor network output action set, S t is the state value at the t-th iteration (i.e., t1, t2,..., t p ), A is the Actor network output action set, (T + R + λS) is the reward function, λ is the regularization term parameter, θ is the Actor network fully connected parameter, δ is the TD error calculated by the Actor network and the Critic network, and α is the learning rate parameter.
[0168] (7) Repeat (4 - 6) until the convergence condition of the Actor-Critic model is reached ( or t = t max , that is, when the state value is at the boundary of the optimization limit or the number of iterations reaches the set maximum number of times t = t max , the convergence can be stopped).
[0169] (8) Use the gradient descent method to update the parameter θ of the state value function of the global network v , and the update formula is as follows.
[0170]
[0171] (9) Use the gradient ascent method to update the policy π parameter of the global network. To overcome the situation where the algorithm converges to a local optimum, the reward function is added to the objective function and directly bound to the policy π to improve the convergence accuracy. The gradient of the entire objective function includes the regularization term related to the policy parameter. The formula used is as follows.
[0172]
[0173] (10) Based on the gradients at each moment calculated above, further sum them up to calculate the cumulative gradients dθ and dθ v , based on the calculated dθ and dθ v update the global neural network parameters θ and θ v complete the update, evaluate the network weight update: θ v = θ v - βdθ v , update the action network weight: θ = θ - αdθ, where α and β are update coefficients.
[0174] (11) The local neural network synchronizes the parameters of the global neural network, that is, θ’ = θ and θ v ’ = θ v , and set the cumulative gradient to 0, that is, dθ ← 0 and dθ v ← 0, and then the local neural network (worker) interacts with the subsequent environment.
[0175] (12) Repeat steps (1) to (11) in this way until the training requirements are met. Finally, obtain the trained A3C algorithm model, and use the action network at this time as the resource allocation model.
[0176] Refer to Figure 10 , which is a specific conceptual schematic diagram when implementing a traditional A3C algorithm provided by an embodiment of the present application and the A3C algorithm provided by the present application. As Figure 10 shown, the experience replay method in the present application adopts the strategy of double combined experience of cumulative experience and single experience. Since traditional A3C algorithms are all committed to solving the global maximum reward, the experience replay mechanism of traditional A3C algorithms all adopts cumulative experience replay. While pursuing the global maximum reward in the present application, it is also necessary to focus on minimizing the overall rendering latency (T+R) each time. Therefore, the double combined experience of cumulative experience and single experience is adopted. The number of Workers trained in the present application is 12, and the global iteration times are set to 1000 times.
[0177] Through the above steps 601 to 605, and steps 801 to step 803, using the adaptive learning characteristics of the multi-agent reinforcement learning algorithm, combined with the detailed parameters of the distributed rendering system, and using the rendering time function required to execute the rendering task as the learning target, a resource allocation model that can effectively improve the rendering efficiency of processing rendering tasks is trained, and thus the data rendering efficiency and reliability in practical applications are effectively improved.
[0178] In some embodiments, when actually processing the rendering tasks of end users, after the front-end middleware divides the tasks into multiple rendering subtasks, it inputs all the subtask data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing, obtains the optimized working parameters (including the optimized CPU working frequency and bandwidth) and optimized image pixels for each working rendering node, and then sends these optimized working parameters and optimized image pixels to the corresponding working rendering nodes.
[0179] Step 204: Receive the sub-rendering data sent by the cloud server for each rendering subtask, and based on the sub-rendering data, optimized working parameters, and optimized image pixels, perform data rendering on the corresponding working rendering node to obtain the rendering sub-result corresponding to the rendering subtask.
[0180] The following provides a detailed description of step 204.
[0181] In some embodiments, after each working rendering node receives the subtask data, optimized working parameters (including the optimized CPU working frequency and bandwidth), and optimized image pixels corresponding to the divided rendering subtasks, and receives the sub-rendering data sent by the cloud server for each rendering subtask, these working rendering nodes will perform parallel data rendering based on the sub-rendering data, optimized working parameters, and optimized image pixels. After the data rendering is completed and the rendering sub-result corresponding to the rendering subtask is obtained, the rendering sub-result will be sent back to the front-end middleware.
[0182] Step 205: Merge all the rendering sub-results to obtain the target rendering result of the rendering task, and send the target rendering result to the end user.
[0183] The following provides a detailed description of step 205.
[0184] In some embodiments, after the front-end middleware receives the rendering sub-results sent back by all the working rendering nodes, it will merge all the rendering sub-results to obtain the target rendering result of the rendering task, and send the target rendering result to the end user for display through the edge server.
[0185] The following will further describe how to merge these rendering sub-results.
[0186] Refer to Figure 11 , merging all the rendering sub-results to obtain the target rendering result of the rendering task includes the following steps 1101 to 1102.
[0187] Step 1101: Receive the rendering completion instructions sent by all the working rendering nodes.
[0188] Step 1102: Based on the rendering numbers, splice each rendering sub-result in sequence to obtain the target rendering result.
[0189] The following provides a detailed description of steps 1101 to 1102.
[0190] As mentioned in the above description, the subtask data obtained from each divided rendering subtask includes rendering numbers sorted in order. Therefore, after the front-end middleware receives the rendering completion instructions sent by all working rendering nodes, it will sequentially splice these rendering subtasks based on the rendering numbers carried in the returned rendering subtask results, so as to obtain the target rendering result, avoiding the situation of incorrect splicing results, thereby improving the reliability of data rendering in the distributed rendering system.
[0191] The data rendering method and related devices of the distributed rendering system proposed in the embodiments of this application. The distributed rendering system includes an edge server cluster and a cloud server. The edge server cluster includes multiple rendering nodes. The method is applied to the edge server cluster and includes: First, in response to a rendering requirement instruction corresponding to a rendering task issued by an end user, parse the rendering requirement instruction to obtain rendering requirement information; Second, evenly divide all view frustums to obtain multiple first view frustum subsets, and generate a pixel rendering ratio factor based on the reciprocal of the number of rendering nodes. When the number of first subsets is less than the number of rendering nodes, based on the sum of the view frustum pixel numbers of all view frustums in the first view frustum subset, obtain a first divided pixel number, and based on the ratio of the first divided pixel number to the total number of rendering pixels, obtain a first divided pixel ratio. When the first divided pixel ratio is less than the pixel rendering ratio factor, generate multiple rendering subtasks based on the multiple first view frustum subsets. When the first divided pixel ratio is not less than the pixel rendering ratio factor, evenly divide all view frustums in each first view frustum subset to obtain multiple second view frustum subsets, and based on the ratio of the sum of the view frustum pixel numbers of all view frustums in the second view frustum subset to the total number of rendering pixels, obtain a second divided pixel ratio. When the second divided pixel ratio is less than the pixel rendering ratio factor, generate multiple rendering subtasks based on the multiple second view frustum subsets, and send the rendering subtasks to the cloud server. The rendering subtasks include subtask data and a rendering node serving as a working rendering node; Then, generate a state space based on the node hardware parameters of multiple rendering nodes and multiple subtask data parameters, generate an action space based on the working parameters of multiple rendering nodes and image pixel parameters, obtain a communication delay function during the execution of the rendering task. The communication delay function includes a first communication time between the end user and the edge server cluster and a second communication time for dividing the rendering task, obtain the maximum rendering time function of multiple rendering nodes, and based on the maximum rendering time function and the communication delay function, obtain a rendering time function, and generate a reward model based on the rendering time function, obtain the initial action network parameters of the action network and the initial evaluation network parameters of the evaluation network, generate an initial resource allocation model based on the state space, action space, reward model, initial action network parameters, and initial evaluation network parameters, and perform multiple rounds of training on the initial resource allocation model, and based on the trained initial resource allocation model, obtain a resource allocation model, input all subtask data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing, and obtain the optimized working parameters and optimized image pixels of each working rendering node; Next, receive the sub-rendered data sent by the cloud server for each rendering subtask, and perform data rendering on the corresponding working rendering node based on the sub-rendered data, optimized working parameters, and optimized image pixels to obtain the rendering sub-results corresponding to the rendering subtasks;Finally, receive the rendering completion instructions sent by all working rendering nodes, sequentially splice each rendering sub-result based on the rendered sub-result numbers to obtain the target rendering result, and send the target rendering result to the end user.
[0192] In the embodiment of the present application, multiple edge server clusters near the end user are used as processing rendering devices to reduce the transmission delay of rendering data; in addition, a larger number of edge server clusters compared to cloud servers are used to perform rendering tasks for end users, thereby reducing the possibility of end users waiting for data rendering; and, the rendering tasks are appropriately divided into subtasks based on the number of rendering nodes in the edge server cluster, so that the rendering tasks can be processed simultaneously and collaboratively among multiple rendering nodes to improve the processing efficiency of data rendering; and, after dividing multiple rendering subtasks, a resource allocation model pre-constructed with the goal of improving data rendering efficiency is used to plan appropriate optimization working parameters and optimized image pixels for each rendering node; in addition, the number of rendering nodes corresponding to the available rendering nodes in the rendering cluster is used as a division matching factor, and the frustum of the rendering task is gradually divided equally until multiple rendering subtasks that can make full use of the parallel computing power of the rendering cluster are obtained, thereby effectively improving the data rendering efficiency of the distributed system; and, the rendering time function obtained by using the communication delay function for data transmission, task division, and sub-result synthesis in the distributed rendering system, and the maximum rendering time function corresponding to the parallel processing of multiple rendering nodes is used as a performance indicator for evaluating the distributed rendering system, so as to facilitate subsequent model training of the A3C algorithm using the reward model generated by this rendering time function, and can effectively improve the reliability of the obtained resource allocation model in optimizing rendering efficiency; and, using the adaptive learning characteristics of the multi-agent reinforcement learning algorithm, combined with the detailed parameters of the distributed rendering system, and using the rendering time function required to execute the rendering task as the learning goal, a resource allocation model that can effectively improve the rendering efficiency of processing rendering tasks is trained, thereby greatly improving the efficiency and reliability of data rendering in the distributed rendering system.
[0193] The embodiment of the present application also provides a data rendering device for a distributed rendering system, which can implement the data rendering method of the above-mentioned distributed rendering system. Refer to Figure 12 and this device 1200 includes:
[0194] A response parsing module 1210, configured to parse the rendering requirement instruction in response to the rendering requirement instruction corresponding to the rendering task sent by the end user to obtain rendering requirement information;
[0195] The task division module 1220 is configured to divide the rendering task based on the rendering requirement information and the number of rendering nodes, obtain multiple rendering subtasks, and send the rendering subtasks to the cloud server. The rendering subtasks include subtask data and a rendering node serving as a working rendering node.
[0196] The resource optimization module 1230 is configured to input all the subtask data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing, and obtain the optimized working parameters and optimized image pixels for each working rendering node.
[0197] The rendering processing module 1240 is configured to receive the sub-rendering data sent by the cloud server for each rendering subtask, and perform data rendering on the corresponding working rendering node based on the sub-rendering data, the optimized working parameters, and the optimized image pixels, to obtain the rendering sub-results corresponding to the rendering subtasks.
[0198] The combination module 1250 is configured to merge all the rendering sub-results to obtain the target rendering result of the rendering task, and send the target rendering result to the end user.
[0199] In some embodiments, the task division module 1220 is further configured to:
[0200] Divide all the view frustums evenly to obtain multiple first view frustum subsets, and generate a pixel rendering ratio factor based on the reciprocal of the number of rendering nodes.
[0201] When the number of the first subsets is less than the number of rendering nodes, based on the sum of the view frustum pixel numbers of all the view frustums in the first view frustum subsets, obtain the first division pixel number, and based on the ratio of the first division pixel number to the total number of rendering pixels, obtain the first division pixel ratio.
[0202] Based on the first division pixel ratio, the pixel rendering ratio factor, the multiple first view frustum subsets, and the number of rendering nodes, obtain multiple rendering subtasks.
[0203] In some embodiments, the task division module 1220 is further configured to:
[0204] When the first division pixel ratio is less than the pixel rendering ratio factor, generate multiple rendering subtasks based on the multiple first view frustum subsets.
[0205] When the first division pixel ratio is not less than the pixel rendering ratio factor, divide all the view frustums in each first view frustum subset evenly to obtain multiple second view frustum subsets, and based on the ratio of the sum of the view frustum pixel numbers of all the view frustums in the second view frustum subsets to the total number of rendering pixels, obtain the second division pixel ratio.
[0206] When the second division pixel ratio is less than the pixel rendering ratio factor, generate multiple rendering subtasks based on the multiple second view frustum subsets.
[0207] In some embodiments, the combination module 1250 is further configured to:
[0208] Receive the rendering completion instructions sent by all working rendering nodes;
[0209] Based on the rendering numbers, splice each rendering sub-result in sequence to obtain the target rendering result.
[0210] In some embodiments, the rendering processing module 1240 is further configured to:
[0211] Generate a state space based on the node hardware parameters of multiple rendering nodes and multiple subtask data parameters;
[0212] Generate an action space based on the working parameters of multiple rendering nodes and image pixel parameters;
[0213] Obtain the rendering time function corresponding to the execution of the rendering task, and generate a reward model based on the rendering time function;
[0214] Obtain the initial action network parameters of the action network and the initial evaluation network parameters of the evaluation network;
[0215] Generate an initial resource allocation model based on the state space, action space, reward model, initial action network parameters, and initial evaluation network parameters, perform multiple rounds of training on the initial resource allocation model, and obtain a resource allocation model based on the trained initial resource allocation model.
[0216] In some embodiments, the rendering processing module 1240 is further configured to:
[0217] Obtain the communication delay function during the execution of the rendering task, where the communication delay function includes the first communication time between the end user and the edge server cluster and the second communication time for dividing the rendering task;
[0218] Obtain the maximum rendering time function of multiple rendering nodes, and obtain the rendering time function based on the maximum rendering time function and the communication delay function.
[0219] In some embodiments, the rendering processing module 1240 is further configured to:
[0220] Obtain the rendering state parameters of the current iteration, input the rendering state parameters into the action network to obtain resource action parameters, and obtain updated rendering state parameters based on the resource action parameters;
[0221] Input the rendering state parameters and the updated rendering state parameters into the evaluation network to obtain the state value corresponding to the rendering state parameters and the updated state value corresponding to the updated rendering state parameters;
[0222] Update the initial action network parameters and the initial evaluation network parameters based on the mean squared error, state value, updated state value, and resource action parameters.
[0223] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, the specific implementation manners of the data rendering device of the distributed rendering system are basically the same as those of the data rendering method of the above distributed rendering system, and will not be elaborated here.
[0224] In the embodiments of the present application, the data rendering device of the distributed rendering system uses multiple edge server clusters near the end user as processing and rendering devices to reduce the transmission delay of rendering data; in addition, uses a larger number of edge server clusters compared to cloud servers to perform rendering tasks for end users, thereby reducing the possibility of end users waiting for data rendering; and appropriately divides sub-tasks for the rendering task based on the number of rendering nodes in the edge server cluster, so that the rendering task can be processed simultaneously by multiple rendering nodes to improve the processing efficiency of data rendering; and, after dividing multiple rendering sub-tasks, uses a resource allocation model pre-constructed with the goal of improving data rendering efficiency to plan appropriate optimization working parameters and optimized image pixels for each rendering node; in addition, uses the number of rendering nodes corresponding to the available rendering nodes in the rendering cluster as a division matching factor, and gradually evenly divides the frustum of the rendering task until multiple rendering sub-tasks that can make full use of the parallel computing power of the rendering cluster are obtained, thereby effectively improving the data rendering efficiency of the distributed system; and, uses the communication delay function for data transmission, task division, and sub-result synthesis in the distributed rendering system, and the rendering time function obtained from the maximum rendering time function corresponding to the parallel processing of multiple rendering nodes as a performance indicator for evaluating the distributed rendering system, so as to facilitate subsequent model training of the A3C algorithm using the reward model generated by this rendering time function, and can effectively improve the reliability of the obtained resource allocation model in optimizing rendering efficiency; and, uses the adaptive learning characteristics of the multi-agent reinforcement learning algorithm, combines the detailed parameters of the distributed rendering system, and uses the rendering time function required to execute the rendering task as the learning goal, thereby training a resource allocation model that can effectively improve the rendering efficiency of processing rendering tasks, and further greatly improving the efficiency and reliability of data rendering in the distributed rendering system.
[0225] The embodiments of the present application also provide an electronic device, including:
[0226] At least one memory;
[0227] At least one processor;
[0228] At least one program;
[0229] The program is stored in a memory, and a processor executes the at least one program to implement the data rendering method of the distributed rendering system described above in the present application. The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA for short), an in-vehicle computer, etc.
[0230] Please refer to Figure 13 , Figure 13 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0231] A processor 1301, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0232] A memory 1302, which can be implemented in forms such as a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1302 and are called by the processor 1301 to execute the data rendering method of the distributed rendering system in the embodiments of the present application;
[0233] An input / output interface 1303, which is used to implement information input and output;
[0234] A communication interface 1304, which is used to implement communication and interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0235] A bus 1305, which transmits information between various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304);
[0236] Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 are communicatively connected to each other inside the device through the bus 1305.
[0237] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the data rendering method of the above-mentioned distributed rendering system.
[0238] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0239] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0240] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0241] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0242] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0243] In the description of the present application and the above-mentioned accompanying drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0244] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0245] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0246] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0247] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0248] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0249] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A data rendering method of a distributed rendering system, characterized in that: The distributed rendering system includes an edge server cluster and a cloud server, the edge server cluster includes a plurality of rendering nodes, the method is applied to the edge server cluster, and the method includes: In response to a rendering requirement instruction corresponding to a rendering task issued by a terminal user, parsing the rendering requirement instruction to obtain rendering requirement information; Dividing the rendering task based on the rendering requirement information and the number of rendering nodes to obtain a plurality of rendering subtasks, and sending the rendering subtasks to the cloud server, wherein the rendering subtasks include subtask data and a rendering node as a working rendering node; Inputting all the subtask data and the node hardware parameters corresponding to the working rendering nodes into the resource allocation model for data processing to obtain optimized working parameters and optimized image pixels of each working rendering node; Receiving sub-rendering data sent by the cloud server for each of the rendering sub-tasks, and performing data rendering on the corresponding working rendering node based on the sub-rendering data, the optimized working parameters and the optimized image pixels, to obtain a rendering sub-result corresponding to the rendering sub-task; All the rendering sub-results are merged to obtain a target rendering result of the rendering task, and the target rendering result is sent to the end user.
2. The data rendering method of the distributed rendering system according to claim 1, characterized in that: The rendering requirement information includes a plurality of view cones, the number of view cone pixels of each view cone, and the total number of rendering pixels. The rendering task is divided based on the rendering requirement information and the number of rendering nodes to obtain a plurality of rendering subtasks, including: Dividing all the view cones equally to obtain a plurality of first view cone subsets, and generating a pixel rendering ratio factor based on the inverse of the number of rendering nodes; When the number of the first subsets is less than the number of the rendering nodes, obtaining a first number of divided pixels based on the sum of the numbers of the cone pixels of all the cones in the first cone subset, and obtaining a first divided pixel ratio based on the ratio of the first number of divided pixels to the total number of rendering pixels; A plurality of rendering subtasks are obtained based on the first divided pixel ratio, the pixel rendering ratio factor, the plurality of first view cone subsets, and the number of rendering nodes.
3. The data rendering method of the distributed rendering system according to claim 2, characterized in that: The obtaining of the plurality of rendering subtasks based on the first divided pixel ratio, the pixel rendering ratio factor, the plurality of first view cone subsets and the number of rendering nodes comprises: When the first divided pixel ratio is less than the pixel rendering ratio factor, generating the plurality of rendering subtasks based on the plurality of first view cone subsets; When the first divided pixel ratio is not less than the pixel rendering ratio factor, all the cones in each of the first cone subsets are equally divided to obtain a plurality of second cone subsets, and a second divided pixel ratio is obtained based on a ratio of the sum of the number of cone pixels of all the cones in the second cone subset to the total number of rendered pixels; When the second divided pixel ratio is less than the pixel rendering ratio factor, the plurality of rendering subtasks are generated based on the plurality of second view cone subsets.
4. The data rendering method of the distributed rendering system according to claim 3, characterized in that: The subtask data includes a rendering number, and the merging of all the rendering sub-results to obtain a target rendering result of the rendering task includes: Receiving rendering completion instructions issued by all the working rendering nodes; Based on the rendering number, each of the rendering sub-results is sequentially spliced to obtain the target rendering result.
5. The data rendering method of the distributed rendering system according to claim 1, characterized in that: The step of generating the resource allocation model comprises: Generate a state space based on node hardware parameters of a plurality of the rendering nodes and a plurality of subtask data parameters; Generate an action space based on working parameters of a plurality of rendering nodes and image pixel parameters; Obtaining a rendering time function corresponding to executing a rendering task, and generating a reward model based on the rendering time function; Obtaining initial action network parameters of the action network and initial evaluation network parameters of the evaluation network; An initial resource allocation model is generated based on the state space, the action space, the reward model, the initial action network parameters and the initial evaluation network parameters, and the initial resource allocation model is trained for multiple rounds, and the resource allocation model is obtained based on the trained initial resource allocation model.
6. The data rendering method of the distributed rendering system according to claim 5, characterized in that: The step of obtaining a rendering time function corresponding to executing a rendering task includes: Acquire a communication delay function in a rendering task execution process, wherein the communication delay function includes a first communication time between the terminal user and the edge server cluster and a second communication time of the rendering task; The maximum rendering time function of the plurality of rendering nodes is obtained, and the rendering time function is obtained based on the maximum rendering time function and the communication delay function.
7. The data rendering method of the distributed rendering system according to claim 5, characterized in that: The training step of the initial resource allocation model includes: Obtaining rendering state parameters of the current iteration, and inputting the rendering state parameters into the action network to obtain resource action parameters, and obtaining updated rendering state parameters based on the resource action parameters; Inputting the rendering state parameter and the updated rendering state parameter into the evaluation network to obtain a state value corresponding to the rendering state parameter and an updated state value corresponding to the updated rendering state parameter; Based on the mean square error, the state value, the updated state value and the resource action parameter, the initial action network parameter and the initial evaluation network parameter are updated.
8. A data rendering device of a distributed rendering system, characterized in that: The distributed rendering system includes an edge server cluster and a cloud server, the edge server cluster includes a plurality of rendering nodes, the method is applied to the edge server cluster, and the device includes: A response parsing module, used for responding to a rendering requirement instruction corresponding to a rendering task issued by a terminal user, parsing the rendering requirement instruction to obtain rendering requirement information; A task division module, used for dividing the rendering task based on the rendering requirement information and the number of rendering nodes to obtain a plurality of rendering subtasks, and sending the rendering subtasks to the cloud server, wherein the rendering subtasks include subtask data and a rendering node as a working rendering node; A resource optimization module, used for inputting all the subtask data and the node hardware parameters corresponding to the working rendering nodes into a resource allocation model for data processing, so as to obtain optimized working parameters and optimized image pixels of each working rendering node; A rendering processing module, configured to receive sub-rendering data issued by the cloud server for each of the rendering sub-tasks, and perform data rendering at the corresponding working rendering node based on the sub-rendering data, the optimized working parameters and the optimized image pixels, to obtain a rendering sub-result corresponding to the rendering sub-task; A combining module is used to combine all the rendering sub-results to obtain a target rendering result of the rendering task, and send the target rendering result to the terminal user.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the data rendering method of the distributed rendering system according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data rendering method of the distributed rendering system according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Distributed cloud rendering task dynamic allocation method based on reinforcement learning
CN122285295A