Program branch compiling optimization system oriented to SIMT architecture
By collecting dynamic thread distribution information in the SIMT architecture, calculating the thread activity probability at the branch, and performing reconvergence optimization, the thread differentiation problem caused by conditional branches is solved, and the execution efficiency of GPGPU is improved.
Patent Information
- Application Number
- CN202510223956.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-20
AI Technical Summary
In the SIMT architecture, conditional branches cause thread differentiation, disrupt execution mode, causing execution units to wait or idle, significantly reducing GPGPU execution efficiency.
By collecting thread distribution information during dynamic running of the program, the thread activity probability at different branches is calculated, and re-aggregation optimization is performed based on this probability to generate the optimized Split and Join instructions.
It realizes the advance aggregation of threads, makes full use of the originally idle SIMT computing unit, and improves the code execution efficiency.
Smart Images

Figure CN120179256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of program branch compilation optimization, and specifically relates to a program branch compilation optimization system for SIMT architecture. Background Art
[0002] With the rapid development of the fields of artificial intelligence and high-performance computing, general-purpose graphics processing units (GPGPUs) based on the single-instruction multiple-thread (SIMT) architecture have received unprecedented attention. The SIMT architecture combines the advantages of the traditional single-instruction multiple-data (SIMD) architecture and provides higher flexibility and finer-grained control on this basis, enabling programmers to write thread-level parallel code, thereby making full use of the parallel processing capabilities of the GPU.
[0003] In the SIMT architecture, each thread is allowed to have independent behaviors and paths. However, at this time, conditional branches in the instruction stream may become one of the key factors affecting the performance of the GPGPU. Different threads may follow different execution paths based on conditional branches, thus disturbing the execution mode of SIMT, resulting in a large number of execution units waiting or being idle, significantly reducing the execution efficiency of the GPGPU.
[0004] The traditional processing method for conditional branches in the SIMT architecture is the single-instruction multiple-thread stack (SIMT Stack), which will use a way of serializing the execution of different threads to handle conditional branches. When encountering a warp divergence, the SIMT stack will save the non-jumped branches to the stack, use an active mask to select the threads to be executed in the jumped branches, and when an execution of a jumped branch is completed, load the non-jumped branches from the stack and switch the active mask to execute. At the same time, the warp divergence processing based on the SIMT stack is a software-hardware cooperation, and the compiler needs to generate Split instructions and Join instructions during the compilation process. The Split instruction is located before the branch instruction to save the active mask and the branch re-convergence point, and the Join instruction is used as an identifier of the branch re-convergence point.
[0005] However, the traditional processing method is an analysis and optimization method starting from the static structure of the program, which only ensures the correctness of program execution but ignores the dynamic execution characteristics of the program. In nested branches, there may be a huge difference in the number of active threads in different branches, and the possible early convergence may not be utilized, resulting in the idle and waste of computing resources. Summary of the Invention
[0006] In view of the requirements and deficiencies in the current technological development, the present invention provides a program branch compilation optimization system for the SIMT architecture. By collecting the thread distribution information during the dynamic operation of the program, it can calculate the thread active probability at different branches and perform re-convergence optimization based on this probability to generate optimized Split and Join instruction pairs.
[0007] For the program branch compilation optimization system for the SIMT architecture of the present invention, the technical solution adopted to solve the above technical problems is as follows:
[0008] A program branch compilation optimization system for the SIMT architecture, its implementation involves the following structure:
[0009] The compiler backend is responsible for taking the intermediate representation of the code as input, performing hardware-related optimizations and code generation work, and converting the intermediate representation into binary code;
[0010] The video memory is responsible for storing the binary code converted by the compiler backend;
[0011] The SIMT execution unit is responsible for reading the binary code from the video memory and executing the corresponding instructions;
[0012] The SIMT stack, when the SIMT execution unit executes the Split instruction, is responsible for saving the addresses and active masks of the un-jumped branches. After the SIMT execution unit executes the Join instruction, it is responsible for loading the previously saved unexecuted branches and executing them, or setting the active mask for branch re-convergence;
[0013] The branch performance profiling module is responsible for recording the distribution of active threads on different branch paths after the SIMT execution unit encounters a warp divergence;
[0014] The branch optimization module is responsible for obtaining the recorded information of the branch performance profiling module, analyzing and optimizing the input intermediate representation of the code, and replacing the input of the compiler backend with the optimized intermediate representation of the code.
[0015] Optionally, each record of the involved branch performance profiling module contains three parts of encoding, which are respectively:
[0016] The instruction address of the branch jump, which is used to uniquely identify a branch jump operation;
[0017] The number of active threads executing this branch jump, which is used to reflect the actual number of threads participating in the execution during this branch jump;
[0018] The total number of threads reaching this branch jump, which is used to represent how many threads in total have executed to this branch jump point.
[0019] Further optionally, each time the involved SIMT execution unit executes different instructions at a conditional branch, it queries the active mask of the execution. If the active mask is not all 1s, it means that thread divergence has occurred, and corresponding processing needs to be performed on the records of the branch performance profiling module, that is, creating or updating the corresponding records;
[0020] Use the instruction address of this branch jump to search in all the saved records:
[0021] a) If there is no corresponding record, create a new record for the instruction address of this branch jump and save the number of bits encoded as 1 in the active mask as the number of active threads for this instruction address;
[0022] b) If a record for the instruction address of the corresponding branch jump is found, add the number of active threads in the record to the number of bits encoded as 1 in the active mask.
[0023] Further optionally, the involved branch optimization module obtains the record information of the branch performance profiling module, analyzes and optimizes the input intermediate representation of the code. The workflow to achieve this operation is as follows:
[0024] S1. Traverse the input intermediate representation of the code, find the block where the unprocessed conditional jump instruction is located, and use it as the current processing node;
[0025] S2. Insert a Split instruction before the conditional jump instruction. The first parameter of the inserted Split instruction is the judgment condition of this thread, and the second parameter is the immediate post-dominator node of this thread;
[0026] S3. Insert a re-convergence code block before the immediate post-dominator node of the block where the conditional jump instruction is located;
[0027] S4. Save the predecessor nodes of the basic block corresponding to the current processing node to the processing queue;
[0028] S5. Take out the predecessor nodes of the basic block corresponding to the current processing node from the processing queue, and calculate the thread activity rate reaching the current processing node;
[0029] S6. Judge whether the activity rate is greater than a pre-set threshold. If it is, jump to step S7; otherwise, jump to step S8;
[0030] S7. Insert a re-convergence code block before the current processing node, and then jump to step S4;
[0031] S8. Judge whether the processing queue is empty. If it is, jump to step S9; otherwise, jump to step S5;
[0032] S9. Determine whether all the intermediate code representations have been traversed. If so, end the process; otherwise, jump to step S1.
[0033] Preferably, execute step S3. Insert a re-convergence code block before the immediate post-dominator node of the current processing node, which is the block where the conditional jump instruction is located. The specific steps for inserting the re-convergence code block are as follows:
[0034] S3.1. Create a new basic block as the re-convergence code block. The last instruction in the basic block is an unconditional jump instruction, and the jump address of the unconditional jump instruction is the immediate post-dominator node being processed. The second-to-last instruction in the basic block is set as a join instruction.
[0035] S3.2. Traverse the predecessor nodes of the immediate post-dominator node being processed, and modify the jump instruction address of the last unconditional jump instruction in the predecessor nodes to the re-convergence code block.
[0036] S3.3. Traverse the predecessor nodes of the re-convergence code block, find the variables that are modified in more than two predecessor nodes, and for each found variable, add a PHI instruction at the head of the re-convergence code block to determine the specific value of the variable according to the source of the variable during actual runtime.
[0037] Preferably, execute step S7. Insert a re-convergence code block before the current processing node. The specific steps for inserting the re-convergence code block are as follows:
[0038] S7.1. Create a new basic block as the re-convergence code block. The last instruction in the basic block is an unconditional jump instruction, and the jump address of the unconditional jump instruction is the current processing node. The second-to-last instruction in the basic block is set as a join instruction.
[0039] S7.2. Traverse the predecessor nodes of the current processing node, and modify the jump instruction address of the last unconditional jump instruction in the predecessor nodes to the re-convergence code block.
[0040] S7.3. Traverse the predecessor nodes of the re-convergence code block, find the variables that are modified in more than two predecessor nodes, and for each found variable, add a PHI instruction at the head of the re-convergence code block to determine the specific value of the variable according to the source of the variable during actual runtime.
[0041] Preferably, execute step S5. When calculating the thread activity rate, iteratively access the successor nodes from the block where the current processing conditional jump instruction is located until the immediate post-dominator node is traversed. For each node visited on the path, query the number of active threads and the total number of threads reaching the specified jump address from the branch performance profiling module according to the instruction address at the starting position of the node.
[0042] Further preferably, step S5 is executed. The thread activity rate of the node is calculated by dividing the sum of the active threads reaching the node by the maximum value of the total number of threads queried this time. The value range of the calculated activity rate is between 0 and 1.
[0043] A program branch compilation optimization system for SIMT architecture according to the present invention has the following beneficial effects compared with the prior art:
[0044] 1. By collecting the thread distribution information during the dynamic operation of the program, the present invention can calculate the thread activity probability at different branches, and perform re-convergence optimization based on this probability to generate optimized Split and Join instruction pairs; the compiled and optimized code can be pre-converged according to the active thread probability in different branches, and the converged threads can make full use of the originally idle SIMT arithmetic units, ultimately improving the execution efficiency of the code.
[0045] 2. The present invention adopts a software and hardware collaborative method, and uses the branch performance profiling module to collect accurate thread distribution, thus effectively utilizing the runtime information of the program and improving the accuracy of the calculated branch activity rate.
[0046] 3. Based on the general SIMT architecture, the present invention adds a branch performance profiling module to the existing architecture, which is compatible with other classic designs such as the SIMT stack and execution unit of the architecture, and the additional hardware overhead is small. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the system architecture schematic diagram of Embodiment 1 of the present invention;
[0048] Figure 2 is the working flowchart of the branch optimization module in Embodiment 1 of the present invention;
[0049] Figure 3 is the specific example diagram of compilation optimization in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the technical solutions, technical problems solved and technical effects of the present invention more clearly understood, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.
[0051] Embodiment 1:
[0052] Refer to the attached Figure 1 , this embodiment proposes a program branch compilation optimization system for SIMT architecture, and its implementation involves the following structure:
[0053] The compiler backend is responsible for taking the intermediate representation of the code as input, performing hardware-related optimizations and code generation work, and converting the intermediate representation into binary code;
[0054] The video memory is responsible for storing the binary code converted by the back end of the compiler.
[0055] The SIMT execution unit is responsible for reading the binary code from the video memory and executing the corresponding instructions.
[0056] The SIMT stack is responsible for saving the addresses of the un-jumped branches and the active masks when the SIMT execution unit executes the Split instruction. After the SIMT execution unit executes the Join instruction, it is responsible for loading the previously saved unexecuted branches and executing them, or setting the active mask for branch re-convergence.
[0057] The branch performance profiling module is responsible for recording the distribution of active threads on different branch paths after the SIMT execution unit encounters warp divergence.
[0058] The branch optimization module is responsible for obtaining the recorded information of the branch performance profiling module, analyzing and optimizing the input intermediate representation of the code, and replacing the input of the back end of the compiler with the optimized intermediate representation of the code.
[0059] In this embodiment, each record of the branch performance profiling module contains three parts of encoding, namely:
[0060] The instruction address of the branch jump, which is used to uniquely identify a branch jump operation.
[0061] The number of active threads executing this branch jump, which is used to reflect the number of threads actually participating in the execution during this branch jump.
[0062] The total number of threads reaching this branch jump, which is used to indicate how many threads in total have executed to this branch jump point.
[0063] Each time the involved SIMT execution unit executes different instructions of the conditional branch, it queries the executed active mask. If the active mask is not all 1, it means that warp divergence has occurred, and corresponding processing needs to be performed on the records of the branch performance profiling module, that is, creating or updating the corresponding records.
[0064] Use the instruction address of this branch jump to search in all the saved records:
[0065] a) If there is no corresponding record, a record of the instruction address of this branch jump is newly created, and the number of bits encoded as 1 in the active mask is saved as the number of active threads of this instruction address.
[0066] b) If a record of the instruction address of the corresponding branch jump is found, the number of active threads in the record is accumulated with the number of bits encoded as 1 in the active mask.
[0067] In this embodiment, the involved branch optimization module obtains the record information of the branch performance profiling module, analyzes and optimizes the input intermediate representation of the code, and refers to the appendix Figure 2 , and the workflow for implementing this operation is as follows:
[0068] S1. Traverse the input intermediate representation of the code to find the block where the unprocessed conditional jump instruction is located, and use it as the current processing node.
[0069] S2. Insert a Split instruction before the conditional jump instruction. The first parameter of the inserted Split instruction is the judgment condition of the thread, and the second parameter is the immediate post-dominator node of the thread.
[0070] S3. Insert a re-convergence code block before the immediate post-dominator node of the block where the conditional jump instruction is located.
[0071] The specific steps for inserting the re-convergence code block are as follows:
[0072] S3.1. Create a new basic block as the re-convergence code block. The last instruction in the basic block is an unconditional jump instruction, and the jump address of the unconditional jump instruction is the immediate post-dominator node being processed. The second-to-last instruction in the basic block is set as a join instruction;
[0073] S3.2. Traverse the predecessor nodes of the immediate post-dominator node being processed, and modify the jump instruction address of the last unconditional jump instruction in the predecessor nodes to the re-convergence code block;
[0074] S3.3. Traverse the predecessor nodes of the re-convergence code block to find the variables that are modified in more than two predecessor nodes. For each found variable, add a PHI instruction at the head of the re-convergence code block to determine the specific value of the variable according to the source of the variable during actual runtime.
[0075] S4. Save the predecessor nodes of the basic block corresponding to the current processing node to the processing queue.
[0076] S5. Take out the predecessor nodes of the basic block corresponding to the current processing node from the processing queue, and calculate the thread activity rate reaching the current processing node.
[0077] When calculating the thread activity rate, iterate and access the successor nodes from the block where the currently processed conditional jump instruction is located until the immediate post-dominator node is traversed. For each node accessed on the path, query the number of active threads and the total number of threads reaching the specified jump address from the branch performance profiling module according to the instruction address at the starting position of the node.
[0078] Use the sum of the number of active threads reaching the node divided by the maximum value of the total number of threads queried this time as the thread activity rate of the node. The calculated activity rate ranges from 0 to 1.
[0079] S6. Determine whether the activity rate is greater than a pre-set threshold. If so, jump to step S7; otherwise, jump to step S8.
[0080] S7. Insert a re-convergence code block before the current processing node, and then jump to step S4.
[0081] The specific steps for inserting the re-convergence code block are as follows:
[0082] S7.1. Create a new basic block as the re-convergence code block. The last instruction in the basic block is an unconditional jump instruction, and the jump address of the unconditional jump instruction is the current processing node. The second-to-last instruction in the basic block is set as a join instruction.
[0083] S7.2. Traverse the predecessor nodes of the current processing node, and modify the jump instruction address of the last unconditional jump instruction in the predecessor nodes to the re-convergence code block.
[0084] S7.3. Traverse the predecessor nodes of the re-convergence code block, find the variables that are modified in more than two predecessor nodes. For each found variable, add a PHI instruction at the head of the re-convergence code block to determine the specific value of the variable according to the source of the variable during actual runtime.
[0085] S8. Determine whether the processing queue is empty. If so, jump to step S9; otherwise, jump to step S5.
[0086] S9. Determine whether all the code intermediate representations have been traversed. If so, end the process; otherwise, jump to step S1.
[0087] Based on the program branch compilation optimization system of this embodiment, refer to the appendix Figure 3 , use basic block 1, basic block 2, basic block 3, basic block 4, basic block 5, and basic block 6 to represent different basic blocks (Basic Block) in the code intermediate representation. The inside of a basic block is a group of sequentially executed instructions, and the last instruction of the basic block is a jump instruction. The arrows in the figure represent the jump directions pointed to by the jump instructions in the basic blocks, and the numbers beside the arrows represent the thread activity rate for a specific jump branch. Re-convergence code block 1 and re-convergence code block 2 respectively represent the re-convergence code blocks inserted in S7 and S3.
[0088] Observe the appendix Figure 3In the example diagram, there are two arrows starting from basic block 1, indicating that the jump instruction at the end of basic block 1 is a conditional jump instruction. The immediate post-dominator node of basic block 1 in the control flow graph is basic block 6. Therefore, in step S3, a re-convergence code block 2 is created and inserted before basic block 6. The purpose of this operation is to reasonably adjust the program execution flow and lay a foundation for more efficient execution in the future. Then, the concept of a threshold is introduced. In this optimization, the threshold is set to 95%. When further analyzing the predecessor nodes of the basic block, it is found that the predecessor nodes of the basic block (i.e., the re-convergence code block 2) are basic block 4 and basic block 5. Here, the thread activity rate of basic block 4 is 98%. This data is crucial for subsequent optimization decisions. Since the thread activity rate of basic block 4 exceeds the set threshold of 95%, according to the optimization strategy, step S7 is entered to insert the re-convergence code block 1.
[0089] Before optimization, it can be found that there are obvious efficiency problems in the program execution. 25% of the code from basic block 2 and 73% of the code from basic block 3 will be executed in the SIMT (Single Instruction Multiple Threads) execution unit respectively. This execution method results in a large number of arithmetic units in the SIMT execution unit being idle, greatly wasting hardware resources.
[0090] However, after the above series of optimization operations, the situation has been significantly improved. After optimization, a total of 98% (assuming that the thread activity rates of basic block 2 and basic block 3 can be combined and calculated after optimization, and the sum of the two is 98%) of the threads from basic block 2 and basic block 3 will run together in the SIMT execution unit. In this way, the arithmetic units in the SIMT execution unit can be more fully utilized, and the originally idle arithmetic units are effectively activated, thus significantly improving the utilization rate of the SIMT execution unit and ultimately enhancing the execution efficiency of the entire program.
[0091] In summary, by using a program branch compilation optimization system for SIMT architecture of the present invention, by collecting the thread distribution information during the dynamic operation of the program, calculating the thread activity probability at different branches, and performing re-convergence optimization based on this probability, the converged threads can fully utilize the originally idle SIMT arithmetic units, ultimately improving the execution efficiency of the code.
[0092] The above uses specific individual examples to elaborate in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of this technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the patent protection scope of the present invention.
Claims
1. A program branch compilation optimization system for SIMT architecture, characterized in that: Its implementation involves the following structure: The compiler backend is responsible for taking the intermediate representation of the code as input, performing hardware-related optimization and code generation, and converting the intermediate representation into binary code; Video memory, responsible for storing binary code converted by the compiler backend; SIMT execution unit, responsible for reading binary code from video memory and executing corresponding instructions; The SIMT stack is responsible for saving the addresses and active masks of unjumped branches when the SIMT execution unit executes the Split instruction. After the SIMT execution unit executes the Join instruction, it is responsible for loading the previously saved unexecuted branches and executing them, or setting the active mask for branch reconvergence. The branch performance analysis module is responsible for recording the distribution of active threads on different branch paths after the SIMT execution unit encounters thread bundle differentiation; The branch optimization module is responsible for obtaining the recorded information of the branch performance analysis module, analyzing and optimizing the input code intermediate representation, and replacing the input of the compiler backend with the optimized code intermediate representation.
2. A program branch compilation optimization system for SIMT architecture according to claim 1, characterized in that: Each record of the branch performance analysis module contains three parts of code, namely: The instruction address of the branch jump is used to uniquely identify a branch jump operation; The number of active threads executing the branch jump is used to reflect the number of threads actually participating in the execution when the branch jumps; The total number of threads that reach the branch jump is used to indicate how many threads have executed to this branch jump point.
3. A program branch compilation optimization system for SIMT architecture according to claim 2, characterized in that: Each time the SIMT execution unit executes different instructions of the conditional branch, it queries the active mask of the execution. If the active mask is not all 1, it means that thread differentiation has occurred, and the records of the branch performance analysis module need to be processed accordingly, that is, the corresponding records need to be created or updated; Use the branch jump instruction address to search in all saved records: a) If no corresponding record exists, a new record of the instruction address of the branch jump is created, and the number of bits encoded as 1 in the active mask is saved as the number of active threads of the instruction address; b) If a record of the instruction address corresponding to the branch jump is found, the number of active threads in the record and the number of bits encoded as 1 in the active mask are accumulated.
4. A program branch compilation optimization system for SIMT architecture according to claim 3, characterized in that: The branch optimization module obtains the recorded information of the branch performance analysis module, analyzes and optimizes the input code intermediate representation, and the workflow for implementing this operation is as follows: S1, traverse the intermediate representation of the input code, find the block where the unprocessed conditional jump instruction is located, and use it as the current processing node; S2. Insert a Split instruction before the conditional jump instruction, wherein the first parameter of the Split instruction is the judgment condition of the thread, and the second parameter is the immediately subsequent controlling node of the thread; S3, inserting a reconvergent code block before the immediately following dominant node of the block where the conditional jump instruction is located; S4, save the predecessor node of the basic block corresponding to the current processing node to the processing queue; S5. Take out the predecessor node of the basic block corresponding to the current processing node from the processing queue, and calculate the activity rate of the thread reaching the current processing node; S6, determine whether the activity rate is greater than a preset threshold, if yes, jump to step S7, otherwise jump to step S8; S7, inserting a reconvergence code block before the current processing node, and then jumping to step S4; S8, determine whether the processing queue is empty, if so, jump to step S9, otherwise jump to step S5; S9. Determine whether all code intermediate representations have been traversed. If so, end the process; otherwise, jump to step S1.
5. A program branch compilation optimization system for SIMT architecture according to claim 4, characterized in that: Execute step S3 to insert a reconvergent code block before the block where the conditional jump instruction is located, that is, before the immediately subsequent controlling node of the current processing node. The specific steps of inserting the reconvergent code block are as follows: S3.
1. Create a new basic block as a reconvergent code block. The last instruction in the basic block is an unconditional jump instruction. The jump address of the unconditional jump instruction is the immediately following dominant node being processed. The second to last instruction in the basic block is set as a join instruction. S3.2, traverse the predecessor node of the immediately subsequent dominating node being processed, and modify the jump instruction address of the last unconditional jump instruction in the predecessor node to a reconvergent code block; S3.
3. Traverse the predecessor nodes of the reconvergent code block to find variables that are modified in more than two predecessor nodes. For each variable found, add a PHI instruction to the header of the reconvergent code block to determine the specific value of the variable based on the source of the variable at actual runtime.
6. A program branch compilation optimization system for SIMT architecture according to claim 5, characterized in that: Execute step S7 to insert a re-convergence code block before the current processing node. The specific steps of inserting the re-convergence code block are as follows: S7.
1. Create a new basic block as a reconvergent code block. The last instruction in the basic block is an unconditional jump instruction. The jump address of the unconditional jump instruction is the current processing node. The second to last instruction in the basic block is set as a join instruction. S7.2, traverse the predecessor nodes of the current processing node, and modify the jump instruction address of the last unconditional jump instruction in the predecessor node to a reconvergent code block; S7.
3. Traverse the predecessor nodes of the reconvergent code block to find variables that are modified in more than two predecessor nodes. For each variable found, add a PHI instruction to the header of the reconvergent code block to determine the specific value of the variable based on the source of the variable at actual runtime.
7. A program branch compilation optimization system for SIMT architecture according to claim 4, characterized in that: Execute step S5, when calculating the thread activity rate, iteratively access the successor node from the block where the conditional jump instruction currently being processed is located until the immediately subsequent dominating node is traversed. For each node visited on the path, the branch performance analysis module is used to query the number of active threads and the total number of threads that have reached the specified jump address based on the instruction address at the starting position of the node.
8. A program branch compilation optimization system for SIMT architecture according to claim 7, characterized in that: Step S5 is executed to calculate the thread activity rate of the node by dividing the sum of the number of active threads reaching the node by the maximum total number of threads queried this time. The calculated activity rate has a value range between 0 and 1.
Citation Information
Cited By
SIMT architecture branch processing system and method based on structured nodes
CN121411827A
A structured node-based simt architecture branch processing system and method
CN121411827B