A data-oriented particle system task design and scheduling method and device

Through data-oriented particle system task design and scheduling methods, the limitations of GPU parallel processing in the existing technology are solved, and parallel processing and computing efficiency of cross-particle systems are improved.

CN119292780BActive Publication Date: 2025-06-13ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411384342.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-06-13
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

The prior art has limitations in GPU parallel processing during particle system computing, and cannot fully utilize the task correlation between different particle systems, resulting in inefficient utilization of computing resources and reduced cache consistency.

Method used

Through data-oriented particle system task design and scheduling methods, a particle system prototype containing data areas and task areas is created, the data areas are merged and directed acyclic task dependency graph is generated, and mapped to the GPU computing task manager to achieve highly optimized parallel scheduling execution.

Benefits of technology

Parallel processing across particle systems is realized, memory access mode and cache utilization are optimized, and the computing efficiency of particle systems is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119292780B_ABST
    Figure CN119292780B_ABST
Patent Text Reader

Abstract

The present invention discloses a data-oriented particle system task design and scheduling method and device, including: creating a particle system prototype for each particle system, which includes a data area and a task area. The data area declares the attributes represented by the particles as components, and the task area declares all the tasks executed by the particle system. Moreover, the tasks in the task area have an execution sequence relationship, and each task operates on at least one component; after performing dependency parsing on the tasks in different task stages in each task area and generating a directed acyclic task dependency graph, it is added to multiple global task blocks divided by task stages. The task dependency graphs within each global task block are merged into a unique directed acyclic dependency graph; after merging the unique directed acyclic dependency graphs of all global task blocks and mapping them to the GPU computing task manager, highly optimized parallel scheduling execution is achieved on the GPU, thereby optimizing the memory access pattern, improving the cache utilization rate, and further enhancing the computing efficiency of the particle system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of real-time rendering, and particularly relates to a data-oriented particle system task design and scheduling method and device. Background Art

[0002] In the field of real-time rendering, the particle system is an important part of the scene, and often generates a lot of overhead due to the update of a large amount of particle data. Currently, the idea of using GPU (Graphics Processing Unit) parallelization is used to optimize data calculation.

[0003] However, the current GPU parallelization processing for the particle system calculation process has limitations, mainly reflected in its limitation within the scope of a single particle system class. The disadvantage of this method is that when creating multiple particle system effect classes, such as two different particle systems, the Unreal engine will perform independent parallelization processing on them. This approach ignores the possible task correlations between different particle systems, thus failing to fully utilize the potential optimization space.

[0004] This limitation may not only lead to inefficient use of computing resources, but also may not fully utilize cache coherence and reduce the storage efficiency of data. In complex scenes, the collaborative optimization between multiple particle systems will become a key factor in improving rendering efficiency and computing performance. Therefore, exploring cross-particle system parallelization strategies and how to effectively utilize the task correlations between systems will be an important direction in the future.

[0005] In the system practice of computing optimization, Unity's Data-Oriented Technology Stack (DOTS for short) is a paradigm worthy of attention. DOTS adopts the design pattern of the Entity Component System (ECS), providing an innovative solution for data storage and task description. The core advantage of this architecture lies in its data-oriented design concept. By decomposing game objects into entities, components, and systems, where entities are responsible for marking data, components are responsible for storing data, and systems are responsible for performing operations on data, separating data and behavior, DOTS can organize and manage data more effectively, thus achieving more efficient memory access and cache utilization.

[0006] Existing Unity DOTS implements a data-oriented task scheduling mechanism on the CPU side. The scheduler analyzes the dependencies between tasks, creates a task graph, and then efficiently allocates these tasks on the available CPU cores. This method is merely a data storage and task scheduling strategy on the CPU, designed using the multi-core features of the CPU, but lacks consideration for the parallelization of data-oriented systems on modern GPUs and does not perform specific optimizations for the data update tasks of the particle system.

[0007] CUDA Graph is an advanced feature introduced by NVIDIA to optimize complex computing workflows on the GPU, allowing developers to organize a series of CUDA operations into a graph structure to separate definition and execution. Summary of the Invention

[0008] In view of the above, the object of the present invention is to provide a data-oriented particle system task design and scheduling method and device, which can achieve parallel processing of multiple particle systems by the GPU, optimize the memory access pattern, improve the cache utilization rate, and thus enhance the computing efficiency of the particle system.

[0009] To achieve the above object of the invention, a data-oriented particle system task design and scheduling method provided by an embodiment includes the following steps:

[0010] Create a particle system prototype for each particle system, which includes a data area and a task area. In the data area, declare the attributes represented by the particles in the particle system as components, merge the data areas of all particle system prototypes, and globally store the components according to the array structure. In the task area, declare all the tasks executed by the particle system in different task stages, and there is an execution order for the tasks in the task area, and each task operates on at least one component;

[0011] After performing dependency parsing on the tasks in different task stages in the task area of each particle system prototype and generating a directed acyclic task dependency graph, add it to multiple global task blocks divided by task stages. The task dependency graph within each global task block aims to ensure the original dependency relationship and merge as many task nodes as possible, and is merged into a unique directed acyclic dependency graph corresponding to each global task block;

[0012] After merging the unique directed acyclic dependency graphs corresponding to all global task blocks, map them to the GPU computing task manager to achieve highly optimized parallel scheduling and execution on the GPU.

[0013] In the present invention, tasks in the task area have a sequential relationship. If they read from or modify the same component, the later task depends on the earlier task. Merging the tasks and their dependency relationships in the unique directed acyclic dependency graph corresponding to all global task blocks can be directly mapped onto a CUDA Graph, enabling highly optimized parallel execution on the GPU. This integration not only reduces the communication overhead between the CPU and the GPU but also allows the GPU to manage task execution more intelligently, such as overlapping computation with memory transfer.

[0014] Preferably, the operations performed by each task on at least one component include a read operation or a modification operation. When a modification operation is performed, the component memory in the data area is modified, and the corresponding partial memory of the component in the global storage is synchronously modified.

[0015] Preferably, the global task blocks divided according to the task phases include generation task blocks corresponding to the generation task phase, update task blocks corresponding to the update task phase, and rendering task blocks corresponding to the rendering task phase.

[0016] Preferably, each task is authorized to operate on at least one component, and each task can only operate on the at least one authorized component.

[0017] Preferably, when performing dependency resolution for tasks in different task phases in the task area and generating a directed acyclic task dependency graph, each node in the graph represents a task, and the edges between the nodes represent the dependency relationships between tasks. Furthermore, the graph can represent the task execution order of the particle system prototype within the task blocks. Define the operation set corresponding to each task to be composed of authorized components. When the second task is executed after the first task and the intersection of the authorized components in the two operation sets is not empty, it is considered that the second task depends on the first task, and a directed edge pointing from the node corresponding to the first task to the node corresponding to the second task is constructed.

[0018] Preferably, aiming to ensure the original dependency relationship and merge as many task nodes as possible, the task dependency graphs within each global task block are merged into a unique directed acyclic dependency graph corresponding to each global task block, including:

[0019] Aiming to ensure the original dependency relationship and merge as many task nodes as possible, identify the same task nodes in the task dependency graphs that need to be merged, and represent the same task nodes as the same task node in the unique directed acyclic dependency graph, while retaining the original in-edge and out-edge relationships of the task nodes. When a cycle appears in the merged unique directed acyclic dependency graph, revoke the previous operation of merging task nodes until the cycle disappears.

[0020] Preferably, when undoing the operation of merging task nodes before, select to split the task nodes. The strategy for selecting the split task nodes is as follows: only consider the edges of the task nodes within the connection loop and the new merged task nodes within the loop, and select to split the task nodes that have only one type of edge as the incoming edge and another type of edge as the outgoing edge; when there are no such task nodes, select the task nodes with an in-degree of 0 for a certain type of edge.

[0021] To achieve the above-mentioned invention objective, the embodiment also provides a data-oriented particle system task design and scheduling device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned data-oriented particle system task design and scheduling method.

[0022] To achieve the above-mentioned invention objective, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned data-oriented particle system task design and scheduling method.

[0023] To achieve the above-mentioned invention objective, the embodiment also provides a computer product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned data-oriented particle system task design and scheduling method.

[0024] Compared with the prior art, the beneficial effects of the present invention at least include:

[0025] Through the data-oriented particle system task design, separate data and behavior, which crosses the boundaries of traditional particle system categories, can identify and merge the same behaviors in different systems, resolve dependency relationships, and implement GPU parallelization of data update tasks based on CUDA Graph. Merge data and tasks from a cross-particle system perspective, thereby reducing the number of task scheduling times of the particle system and optimizing the scheduling logic, further optimizing the memory access mode of the GPU program, improving the cache utilization rate, and ultimately improving the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0027] Figure 1 is a flowchart of the data-oriented particle system task design and scheduling method provided by the embodiment;

[0028] Figure 2It is the flowchart of data update in the data-oriented particle system task design and scheduling method provided by the embodiment;

[0029] Figure 3 It is the corresponding relationship diagram of the task area, task block, and particle system prototype provided by the embodiment;

[0030] Figure 4 It is the schematic diagram of task merging across particle systems provided by the embodiment. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0032] See Figure 1 , a data-oriented particle system task design and scheduling method provided by the embodiment includes the following steps:

[0033] S1. Create a particle system prototype including a data area and a task area for each particle system. Declare the attributes represented by the particles in the particle system as components in the data area, merge the data areas of all particle system prototypes, and globally store the components according to the array structure.

[0034] In the embodiment, when the user creates a new type of particle system, an independent particle system prototype is created. The particle system prototype includes a data area describing all the attributes of the particles it manages and a task area describing how to modify these attributes and implement the particle system effect.

[0035] The data area stores all the particles of the particle system. The particles are composed of a series of different types of attributes, including attributes such as position, velocity, and color. A particle attribute of a particle system class is declared as a component within a particle system prototype.

[0036] From the user's perspective, editing the data area only adds or removes components to a certain particle system prototype. From the perspective of the global data manager, it is to add or remove components to the global data area or add new memory to the existing components. This is because for all data areas, the global data manager will merge the same components in the data area and store them together. If the identification names and types of two components are the same, they will be regarded as the same component. After merging, the memory size is the sum of the sizes of the original two components.

[0037] In the global data manager, these components are organized into a structure that facilitates GPU access. Each component is stored as an array on the GPU and then stored in a data layout conforming to the Struct of Array (SoA). When processing specific attributes (such as only updating the positions of all particles), SoA allows for contiguous memory access, which is very friendly to the cache mechanisms of modern CPUs and GPUs. In addition, Single Instruction Multiple Data (SIMD) operations can be more easily applied to the SoA layout because data of the same type is stored contiguously. Moreover, when only partial attributes need to be accessed, SoA can reduce unnecessary data transfers and improve memory bandwidth utilization, which is particularly beneficial for achieving efficient coalesced memory access.

[0038] S2, the task area declares all the tasks that the particle system executes in different task phases, and the task constraints in the task area have an execution order relationship. Each task operates on at least one component.

[0039] In an embodiment, the task area includes all the tasks that the particle system needs to execute in different task phases of a frame. Usually, it includes tasks in three phases: particle generation, particle update, and particle rendering. For example, in the particle generation task phase, an initial velocity can be assigned to the particles; in the particle update task phase, the color of each particle can be modified according to its life, and if the particle life exceeds the lifespan, it will be destroyed; in the particle rendering task phase, the particle orientation can be aligned with the viewing direction, etc.

[0040] Each task is an independent functional unit that is specifically responsible for reading or modifying one or more specific particle attributes (components). The task defines the behavior of the particles in different phases. The task constraints in the task area have a clear execution order. When multiple tasks operate on the same component, a dependency relationship will be formed, and the tasks at the back depend on the tasks at the front. It should be noted that the component modified by a task is the component memory of its prototype, corresponding to a part of the memory of this component in the global data manager. The specific data update process is as Figure 2 shown, read the data from the global data area, execute the CUDA Graph to update the data, and submit the updated data to the rendering end of the particle system for rendering.

[0041] Each task clearly defines at least one component that it is authorized to access and operate on, and a task can only operate on at least one authorized component. The execution of the task needs to ensure data consistency and correctness. This data-based task design method provides a clear structure and execution process, which is beneficial to the modularity, scalability, and performance optimization of the system. It is particularly suitable for complex particle systems and can better manage the behaviors and attributes of a large number of particles.

[0042] In S3, after performing dependency parsing on the tasks in different task phases in the task area of each particle system prototype and generating a directed acyclic task dependency graph, it is added to multiple global task blocks divided by task phases.

[0043] In the embodiment, the task area of each particle system prototype is divided into three phases. Then, for all particle system prototypes, they can be divided into three global task blocks according to the phases of the particle systems: the generation task block corresponding to the generation task phase, the update task block corresponding to the update task phase, and the rendering task block corresponding to the rendering task phase. This division reflects the basic workflow of all particle systems within one frame, from creating their respective new particles, to updating all their particles, and then to displaying their particles.

[0044] The corresponding relationship among the task area, task blocks, and particle system prototypes is as Figure 3 shown. Each global task block contains multiple particle system prototypes. Each particle system prototype corresponds to a series of tasks within a single global task block, and these tasks correspond to the tasks defined in the corresponding phase of the task area of this particle system prototype.

[0045] For each particle system prototype within each task block, the dependency relationship among its tasks can be represented by a directed acyclic task dependency graph (DAG). Specifically, when performing dependency parsing on the tasks in different task phases in the task area and generating the task dependency graph, each node in the graph represents a task, and the connecting edges between nodes represent the dependency relationship among tasks. To specifically illustrate the generation method of the task dependency graph, it is defined that the operation set corresponding to each task consists of authorized components. Define the operation set S2 corresponding to the second task S2 = {component C1, component C2,...}, indicating the set of components that the second task S2 will operate on, and the operation set S1 corresponding to the task S1 = {component C1, component C3,...}, indicating the set of components that the task S1 will operate on. If in the original task area, when the second task S2 is executed after the first task S1, and the intersection of the authorized components in the two operation sets is not empty, that is, the operation set S1 ∩ operation set S2 is not equal to the empty set, then it is considered that the second task S2 depends on the first task S1, and a directed connecting edge is constructed from the node corresponding to the first task S1 to the node corresponding to the second task S2.

[0046] For example, in the update task area of particle system 1, the original task order is defined as: {S2 -> S3 -> S4 -> S1}, the operation set S1 = {C1}, the operation set S2 = {C2, C3}, the operation set S3 = {C1, C2, C4, C5}, and the operation set S4 = {C3, C4}. As Figure 4As shown, to generate a task dependency graph, first generate the shown table by dividing it by each component. Each column in the table represents a task (i.e., an operation set) for the operation component, and the order of the tasks in the column is a subset of the original task order. Then, each order requirement in the table is the dependency relationship defined above, and thus corresponds to a directed edge in the generated task dependency graph. After generating directed edges for all the dependencies in the table, a directed acyclic task dependency graph of the particle system prototype in the task area can be obtained.

[0047] In addition, the task node S3 points to the task node S1, indicating that task S1 depends on task S3. Since the dependency relationship maintains the order of the tasks and the original task order is clear, the generated directed Figure 1 must be acyclic. This graph structure clearly shows the execution order of the tasks.

[0048] This structure based on task blocks, particle system prototypes, and the corresponding directed acyclic dependency graph provides a clear, flexible, and efficient execution framework for a single particle system. It can not only accurately express complex particle behaviors but also provides a basis for the subsequent merging, optimization, and parallelization of tasks.

[0049] For S4, the task dependency graphs within each global task block are merged into a unique directed acyclic dependency graph corresponding to each global task block with the goal of maintaining the original dependency relationship and merging as many task nodes as possible.

[0050] In the embodiment, for each global task block, the task dependency graphs of different particle system prototypes are merged to find common task structures. During the merging, the same task nodes in the two graphs are identified. Here, the same means performing the same logical operations on the same components, and the same task nodes are represented as the same task node in the new graph.

[0051] The merged new task node is different from the original task nodes in that the components processed by the original two task nodes are merged. Specifically: the executed commands are the same, so there is no need to change the commands; because the components are the same but the particle system prototypes they belong to are different, the storage areas operated by the two tasks are two non - overlapping areas of the same global memory. Therefore, the processed data sets are merged, and the same commands are used to operate on a larger area of the same memory. Thus, without changing the computational amount, the GPU call instructions can be reduced and the memory access pattern can be optimized.

[0052] For all nodes of the unique directed acyclic dependency graph, the original incoming and outgoing edge relationships need to be preserved to satisfy all dependencies of at least two original task dependency graphs simultaneously. However, this may lead to the generation of new dependency paths. If a cycle appears in the merged unique directed acyclic dependency graph, it indicates the introduction of circular dependencies, meaning that the previously merged task nodes need to be restored to two task nodes. Therefore, try to sequentially undo the previous merge node operations until the cycle in the new graph disappears.

[0053] The specific logic of the merge and undo operations is as Figure 4 shown. For two particle systems within the global task block, first, respectively, according to the component content of the task operations, parse the dependency relationships of their respective tasks. After the parsing is completed, generate two directed acyclic task dependency graphs G 1 and G 2 . Then, for these two task dependency graphs G 1 and G 2 , identify the same task nodes, merge them into a new directed graph (the intermediate result in Figure 4 ), and mark the incoming edges in the new graph to indicate which graph they come from. Detect a cycle {S1->S2->S3->S1}, and split the nodes on the cycle. Since S1, S2, and S3 are all merged nodes, there are three splitting schemes, forming graphs G 12-1 , G 12-1' , G 12-1'' . And there are still cycles in graphs G 12-1 and G 12-1'' , which need to be further split. After there are no cycles, merge the edges with the same start and end points to obtain the final splitting result.

[0054] While ensuring the original dependency relationships, as many task nodes as possible should be merged. Eventually, all directed acyclic task dependency graphs should be merged into the unique directed acyclic dependency graph of the global task block. That is to say, when there is a cycle after the merge, the number of nodes split to resolve the circular dependencies should be as small as possible.

[0055] In the embodiment, first split all cycles that contain only two task nodes, that is, the cycle where node A points to B and node B points to A. Any one of the nodes can be selected for splitting. Then, for other cycles, the strategy for selecting task nodes to split is: only consider the edges connecting the task nodes within the cycle and the new task nodes merged within the cycle, and select to split the task node that has only one type of edge as the incoming edge and the other type of edge as the outgoing edge. In Figure 4 , select to split task node S2 because when only considering the edges connecting S1, S2, and S3, all the incoming edges of S2 come from G 2 and all the outgoing edges come from G 1This strategy may not be optimal, but it can ensure that the current loop is solved. If there is no such task node, select a task node with a certain edge in-degree of 0. Because the original graph is acyclic, such a node must exist.

[0056] The process of generating a unique directed acyclic dependency graph helps to identify and eliminate unnecessary dependencies and simplify the system structure. By merging task nodes, memory usage can be optimized and data transfer can be reduced. The process of merging dependency graphs can reveal the common dependencies between different prototypes, which helps to optimize the overall task scheduling strategy and reduce task switching overhead.

[0057] S5, merges the unique directed acyclic dependency graphs corresponding to all global task blocks and maps them to the GPU computing task manager to achieve highly optimized parallel scheduling execution on the GPU.

[0058] In the embodiment, the directed acyclic dependency graphs of the three global task blocks can be merged into a final directed acyclic graph according to the execution stage, that is, the directed acyclic graph of the generation task block, the update task block, and the rendering task block are executed in sequence. The merged graph still maintains the correct execution order and dependency relationship, reflecting the workflow of all particle systems. The tasks represented by the directed acyclic graph and their dependencies are expressed using a GPU computing task manager (such as CUDA Graph), and each task node is a CUDA kernel. CUDA Graph allows the entire computing Figure 1 Submitting tasks to the GPU at a time greatly reduces the CPU scheduling overhead. In addition, by pre-defining the task dependency graph, the GPU can schedule and execute tasks more intelligently. This method reduces the GPU's idle time and improves overall utilization.

[0059] The embodiment further provides a data-oriented particle system task design and scheduling device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the device is used to implement the data-oriented particle system task design and scheduling method, which specifically includes the following steps:

[0060] S1, create a particle system prototype containing a data area and a task area for each particle system. The data area declares the properties represented by the particles in the particle system as components, merges the data areas of all particle system prototypes, and stores the components globally according to the array structure;

[0061] S2, the task area declares all tasks performed by the particle system in different task stages, and the task constraints in the task area have an execution order relationship, and each task operates on at least one component;

[0062] S3. After performing dependency parsing on the tasks in different task phases in the task area of each particle system prototype and generating a directed acyclic task dependency graph, add it to multiple global task blocks divided by task phases;

[0063] S4. The task dependency graphs within each global task block are merged into a unique directed acyclic dependency graph corresponding to each global task block with the goal of ensuring the original dependency relationship and merging as many task nodes as possible;

[0064] S5. After merging the unique directed acyclic dependency graphs corresponding to all global task blocks, map them to the GPU computing task manager to achieve highly optimized parallel scheduling and execution on the GPU.

[0065] The computing device provided by the embodiment, at the hardware level, in addition to including a processor and a memory, also includes other hardware required for other services such as an internal bus, a network interface, and a memory. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the data-oriented particle system task design and scheduling method described in the above S1 - S5. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.

[0066] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above data-oriented particle system task design and scheduling method, specifically including the following steps:

[0067] S1. Create a particle system prototype for each particle system, which includes a data area and a task area. Declare the attributes represented by the particles in the particle system as components in the data area, merge the data areas of all particle system prototypes, and globally store the components according to an array structure;

[0068] S2. Declare all the tasks executed by the particle system in different task phases in the task area, and there is an execution sequence relationship among the task constraints in the task area. Each task operates on at least one component;

[0069] S3. After performing dependency parsing on the tasks in different task phases in the task area of each particle system prototype and generating a directed acyclic task dependency graph, add it to multiple global task blocks divided by task phases;

[0070] S4. The task dependency graphs within each global task block are merged into a unique directed acyclic dependency graph corresponding to each global task block with the goal of ensuring the original dependency relationship and merging as many task nodes as possible;

[0071] S5. After merging the unique directed acyclic dependency graphs corresponding to all global task blocks and mapping them to the GPU computing task manager, highly optimized parallel scheduling execution is achieved on the GPU.

[0072] In an embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.

[0073] The embodiment also provides a computer product that includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned data-oriented particle system task design and scheduling method, including the following steps:

[0074] S1. Create a particle system prototype for each particle system that includes a data area and a task area. Declare the attributes represented by the particles in the particle system as components in the data area, merge the data areas of all particle system prototypes, and globally store the components according to an array structure.

[0075] S2. Declare all the tasks executed by the particle system in different task phases in the task area, and there is an execution order relationship among the task constraints in the task area. Each task operates on at least one component.

[0076] S3. After performing dependency parsing on the tasks in different task phases in the task area of each particle system prototype and generating a directed acyclic task dependency graph, add it to multiple global task blocks divided according to task phases.

[0077] S4. With the goal of ensuring the original dependency relationship and merging as many task nodes as possible, merge the task dependency graphs within each global task block into a unique directed acyclic dependency graph corresponding to each global task block.

[0078] S5. After merging the unique directed acyclic dependency graphs corresponding to all global task blocks and mapping them to the GPU computing task manager, highly optimized parallel scheduling execution is achieved on the GPU.

[0079] The above embodiment manages the global particle system using a data-oriented concept, separating data and behavior. It crosses the boundaries of traditional particle system categories, can identify and merge the same behaviors in different systems, resolve dependency relationships, and implement GPU parallelization of data update tasks based on CUDA Graph. This innovative design can not only optimize the memory access pattern but also improve the cache utilization rate, thereby significantly enhancing the computing efficiency of the entire system.

[0080] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data-oriented particle system task design and scheduling method, characterized in that: The following steps are involved: For each particle system, create a particle system prototype containing a data area and a task area, where the data area declares the properties represented by particles in the particle system as components, merges the data areas of all particle system prototypes, and stores components globally according to the array structure. The task area declares all tasks performed by the particle system in different task stages, and the task constraints in the task area have an execution order, and each task operates on at least one component; After performing dependency parsing on the tasks of different task stages in the task area of ​​each particle system prototype and generating a directed acyclic task dependency graph, the tasks are added to multiple global task blocks divided by task stages. The task dependency graph in each global task block aims to ensure the original dependency relationship and merge as many task nodes as possible, and merge into a unique directed acyclic dependency graph corresponding to each global task block; The unique directed acyclic dependency graphs corresponding to all global task blocks are merged and mapped to the GPU computing task manager to achieve highly optimized parallel scheduling execution on the GPU.

2. The data-oriented particle system task design and scheduling method according to claim 1, characterized in that: The operation performed by each task on at least one component includes a read operation or a modification operation. When performing a modification operation, the component memory in the data area is modified, and part of the memory of the corresponding component in the global storage is modified synchronously.

3. The data-oriented particle system task design and scheduling method according to claim 1, characterized in that: The global task blocks divided according to the task stages include a generation task block corresponding to the generation task stage, an update task block corresponding to the update task stage, and a rendering task block corresponding to the rendering task stage.

4. The data-oriented particle system task design and scheduling method according to claim 1, characterized in that: Each task is authorized to operate at least one component, and each task can only operate at least one authorized component.

5. The data-oriented particle system task design and scheduling method according to claim 1, characterized in that: When performing dependency parsing for tasks at different task stages in the task area and generating a directed acyclic task dependency graph, each node in the graph represents a task, and the edges between the nodes represent the dependency relationship between the tasks. The operation set corresponding to each task is defined to be composed of authorization components. When the second task is executed after the first task and the intersection of the authorization components in the two operation sets is not empty, the second task is considered to be dependent on the first task, and a directed edge is constructed from the node corresponding to the first task to the node corresponding to the second task.

6. The data-oriented particle system task design and scheduling method according to claim 1, characterized in that: The task dependency graph in each global task block is merged into a unique directed acyclic dependency graph corresponding to each global task block, with the goal of ensuring the original dependency relationship and merging as many task nodes as possible, including: With the goal of ensuring the original dependency and merging as many task nodes as possible, identify the same task nodes in the task dependency graph that need to be merged, and represent the same task nodes as the same task nodes in the unique directed acyclic dependency graph, while retaining the original in-edge and out-edge relationships of the task nodes. When a cycle appears in the unique directed acyclic dependency graph after the merger, undo the previous operation of merging task nodes until the cycle disappears.

7. The data-oriented particle system task design and scheduling method according to claim 6, characterized in that: When undoing the previous operation of merging task nodes, choose to split the task nodes. The strategy for splitting task nodes is: only consider the edges connecting the task nodes in the ring and the new task nodes merged in the ring, and choose to split the task nodes that have only one type of edge as the input edge and another type of edge as the output edge; when there is no such task node, choose a task node with a certain edge input degree of 0.

8. A data-oriented particle system task design and scheduling device, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the data-oriented particle system task design and scheduling method described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the data-oriented particle system task design and scheduling method described in any one of claims 1-7 is implemented.

10. A computer product comprising a computer program, characterized in that When the computer program is executed by a processor, the data-oriented particle system task design and scheduling method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Pipeline parallelization method for coarse-grained streaming application

    CN103377035A

  • Fractal image generation and rendering method based on game engine and CPU parallel processing

    CN105787865A