Heterogeneous parallel computing method and device

By distributing operators to heterogeneous engine groups in heterogeneous computing systems and adopting an innovative on-chip synchronization mechanism, the contradiction between heterogeneous computing power parallelism and synchronization event overhead is solved, and high-performance heterogeneous parallel computing is achieved.

CN120066744AActive Publication Date: 2025-05-30VASTAI TECH (SHANGHAI) INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510541779.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the prior art, while improving the parallelism of heterogeneous computing power, it is difficult to effectively reduce the overhead of synchronization events, resulting in performance bottlenecks.

Method used

By distributing operators to the corresponding heterogeneous engine group and adopting an innovative on-chip synchronization mechanism, including in-group synchronization and inter-group synchronization, the pulse signal is completed by using counters to compare command words and operators to complete pulse signals, and fast event synchronization is achieved.

Benefits of technology

It significantly improves the performance of heterogeneous computing systems, reduces synchronization waiting time, and improves the utilization rate of hardware resources and computing throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066744A_ABST
    Figure CN120066744A_ABST
Patent Text Reader

Abstract

The invention provides a heterogeneous parallel computing method and device. The method comprises the following steps: distributing each operator to a corresponding heterogeneous engine group according to different computing power requirements of the operators; numbering the operators distributed to each heterogeneous engine group; a data dependency relationship among operators in each heterogeneous engine group is analyzed, a counter comparison command word is inserted into an operator queue, and the data dependency relationship is recorded by the number of each operator in each heterogeneous engine; performing intra-group synchronization on each sub-engine in the same heterogeneous engine group; and performing inter-group synchronization on each heterogeneous engine group. According to the technical scheme provided by the invention, the overhead of the synchronization event can be reduced while the computing power of the GPU and the AI accelerator is fully utilized, so that the utilization rate of hardware resources is remarkably improved, and the overall performance of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of parallel computing technology, and in particular, to a method and device for heterogeneous parallel computing. Background Art

[0002] In a heterogeneous computing system, the collaborative computing between a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU), an Artificial Intelligence (AI) computing accelerator is a key technology for achieving efficient parallel processing. The CPU is usually responsible for controlling the flow and executing serial tasks, while dispatching compute-intensive operators to the GPU or AI accelerator for computing, and obtaining the computation results by creating synchronization events. In this process, the on-chip general controller is responsible for docking the tasks dispatched by the CPU, scheduling hardware resources, and responding to event synchronization operations. The latency overhead of event synchronization control is the main factor affecting system performance, and its latency range may vary from dozens of microseconds to thousands of microseconds.

[0003] Specifically, the latency mainly comes from the following aspects: 1. Command submission overhead: The communication latency generated when the CPU dispatches tasks to the GPU or AI accelerator; 2. Hardware scheduling latency: The task scheduling time inside the GPU or AI accelerator; 3. Event signal transmission latency: The synchronization overhead between the internal computing engines of the GPU or AI accelerator; 4. Latency for the CPU to obtain the event status: including the time for querying or interrupt response.

[0004] In recent years, with the diversified development of AI models, in order to adapt to the needs of different computing tasks, GPUs and AI accelerators generally adopt heterogeneous computing power architectures, such as integrating multiple computing resources such as tensor computing engines, vector computing engines, vector computing engines, and communication engines. These resources are usually uniformly scheduled by the on-chip general controller to improve computing efficiency.

[0005] In the programming models of GPUs and AI accelerators, a task queue (Stream) is the basic unit of task scheduling. The operators within the same Stream are strictly executed in order, while the operators of different Streams can be asynchronously and parallelly computed. For example, in a simple AI model containing three operators (A, B, C), if there is no data dependency between operator A and operator B, while operator C depends on the output results of operator A and operator B, the traditional scheduling method has the following problems: 1. Single-Stream Scheduling: If all operators are executed in the same stream, the general controller of the GPU or AI accelerator needs to schedule operators A, B, and C sequentially in order, resulting in a total latency that is the sum of the execution times of the three operators and the time of one synchronization event (the startup and execution time of operator A + the startup and execution time of operator B + the startup and execution time of operator C + the time of one CPU synchronization event). This mode cannot fully utilize the parallelism of the hardware computing power. Especially when the GPU or AI accelerator has the ability to execute operators A and B in parallel, its performance bottleneck is more significant.

[0006] 2. Multi-Stream Scheduling: Allocate operators A and B to different streams to improve parallelism, and solve the data dependency problem by increasing event synchronization on the CPU side. Then the total latency can be optimized to the execution time of the slower one of operators A and B plus the execution time of operator C and the synchronization overhead [max (the startup and execution time of operator A + the time of one CPU synchronization event, the startup and execution time of operator B + the time of one CPU synchronization event) + the startup and execution time of operator C + the time of one CPU synchronization event]. However, although this method increases the parallelism of the on-chip computing power resources, it increases the CPU synchronization event time. Especially, with the increase in the complexity of the AI model, frequent event synchronization will lead to additional latency accumulation, which may instead become a key factor restricting performance.

[0007] Therefore, in the prior art, how to improve the parallelism of heterogeneous computing power while reducing the overhead of synchronization events has become a technical problem to be solved urgently. Summary of the Invention

[0008] In view of this, the present invention provides a method and device for heterogeneous parallel computing to solve the above-mentioned technical problems in the prior art.

[0009] According to one aspect of the present invention, a method for heterogeneous parallel computing is provided, wherein the method includes: Distribute each operator to the corresponding heterogeneous engine group according to the different computing power requirements of the operators, and each heterogeneous engine group includes multiple sub-engines for executing operators; Number the operators distributed to each heterogeneous engine group; Analyze the data dependency relationship between the operators in each heterogeneous engine group, and insert a counter comparison command word into the operator queue; Perform intra-group synchronization on the sub-engines in the same heterogeneous engine group, and the intra-group synchronization includes: Each sub-engine generates its own corresponding group completion flag of the sub-engine after completing the operator calculation; Receive all the corresponding group completion flags of the sub-engines to achieve intra-group synchronization; Perform inter-group synchronization for each heterogeneous engine group, and the inter-group synchronization includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing the operator calculation; Receive the operator completion pulse signals of each heterogeneous engine group and use a counter to count. The counter value of the counter is the serial number of the operators that have been processed by each heterogeneous engine group; Each heterogeneous engine group obtains the configuration threshold of the counter according to the data dependency relationship; Compare the counter value with the configuration threshold. When the counter value is greater than or equal to the configuration threshold, an enabling signal for the heterogeneous engine group is issued to achieve inter-group synchronization.

[0010] According to another aspect of the present invention, there is provided a device for heterogeneous parallel computing, wherein the device includes a compiler and heterogeneous engine groups, The compiler includes: An operator attribution distribution module, configured to distribute each operator to the corresponding heterogeneous engine group according to the different computing power requirements of the operators. Each heterogeneous engine group includes a plurality of sub-engines for executing operators; An operator sorting module, configured to number the operators distributed to each heterogeneous engine group; An inter-operator dependency analysis module, configured to analyze the data dependency relationship between each operator in each heterogeneous engine group and insert a counter comparison command word into the operator queue; The heterogeneous engine group includes: An intra-group hardware synchronization sub-module, configured to perform intra-group synchronization for each sub-engine in the same heterogeneous engine group. The intra-group synchronization includes: Each sub-engine generates its own sub-engine corresponding group completion flag after completing the operator calculation; Receive all sub-engine corresponding group completion flags to achieve intra-group synchronization; An inter-group hardware synchronization sub-module, configured to perform inter-group synchronization for each heterogeneous engine group. The inter-group synchronization includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing the operator calculation; Receive the operator completion pulse signals of each heterogeneous engine group and use a counter to count. The counter value of the counter is the serial number of the operators that have been processed by each heterogeneous engine group; Each heterogeneous engine group obtains the configuration threshold of the counter according to the data dependency relationship; Compare the counter value with the configuration threshold. When the counter value is greater than or equal to the configuration threshold, an enabling signal for the heterogeneous engine group is issued to achieve inter-group synchronization.

[0011] As can be seen from the above technical solutions, the technical solutions provided by the present invention can significantly improve the performance of heterogeneous computing systems through an innovative on-chip synchronization mechanism, and it has at least the following advantages: 1. Through the on-chip synchronization mechanism, operators without data dependencies can be identified and their parallel computing can be enhanced on the GPU or AI accelerator, thereby significantly improving the utilization rate of hardware resources and accelerating the task execution efficiency; 2. By adopting an optimized on-chip synchronization circuit design, the traditional event synchronization overhead is reduced from dozens of microseconds to dozens of nanoseconds, greatly reducing the synchronization waiting time and improving the overall computing throughput; 3. By introducing multiple sets of synchronization mechanisms, the system can support multiple processes to simultaneously call on-chip hardware computing resources, achieve more efficient resource reuse, and improve the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, but do not constitute a limitation to the technical solutions of the present invention.

[0013] Figure 1 Shows an AI model structure including three operators; Figure 2 Shows a method of executing three operators using the same task queue; Figure 3 Shows a method of executing three operators using multiple task queues; Figure 4 Shows a schematic diagram of the hardware resources used in an exemplary embodiment of the present invention; Figure 5 Shows a schematic diagram of compiler-assisted operator dispatch in the method provided by an exemplary embodiment of the present invention; Figure 6 Shows a schematic diagram of single-engine intra-group synchronization in the method provided by an exemplary embodiment of the present invention; Figure 7 Shows a schematic diagram of multi-engine inter-group synchronization in the method provided by an exemplary embodiment of the present invention; Figure 8 Shows a schematic diagram of the software resources used in an exemplary embodiment of the present invention; Figure 9 Shows a block diagram of the structure of the device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is not intended as any limitation on the present invention and its application or use. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present invention thorough and complete and to fully convey the scope of the present invention to those skilled in the art.

[0015] Unless explicitly stated, if the number of elements is not specifically defined, the element can be one or more. The term "plural / a number of" means two or more, the term "based on" should be interpreted as "at least partially based on", the term "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.

[0016] Please refer to Figures 1 to 3 , which shows the synchronous operator calculation method adopted in the background art.

[0017] Figure 1 shows an AI model structure including three operators, where there is no data dependency between operator A and operator B, while the execution of operator C depends on the outputs of operator A and operator B.

[0018] Figure 2 shows a method of executing three operators using the same task queue (Stream). The CPU creates operators A, B, and C of the AI model in the same task queue and dispatches them to the GPU or AI accelerator for execution. The operators within the same task queue are strictly executed in order, and the scheduling process of the on-chip general controller is as follows: 1. Execute operator A: The CPU dispatches operator A to the GPU or AI accelerator, starts the calculation and waits for completion, and the time consumed is denoted as T A (including the start-up delay and execution time of operator A).

[0019] 2. Execute operator B: After operator A is completed, the GPU or AI accelerator starts to execute operator B, and the time consumed is denoted as T B (including the start-up delay and execution time of operator B).

[0020] 3. Execute operator C: Since operator C depends on the outputs of operator A and operator B, it needs to wait for the completion of the former two before starting to execute, and the time consumed is denoted as T C (including the start-up delay and execution time of operator C).

[0021] 4. CPU synchronization event: The CPU obtains the final result through query or interruption, and the synchronization overhead is denoted as T sync .

[0022] It can be seen from this that the total time consumed T 1=T A + T B + T C + T sync In this scheduling and execution method, although the synchronization overhead is low, the parallelism of computing power cannot be fully utilized, resulting in low computing power utilization.

[0023] Figure 3 Shows a method of using multiple task queues to execute three operators. The CPU dispatches operator A and operator B to different Streams (Stream0 and Stream1) respectively to fully utilize the parallel computing power. However, operator C still depends on the outputs of both and needs to wait for the synchronization to complete before execution. The scheduling process of the on-chip general controller is as follows: 1. Execute operators A and B in parallel: The CPU dispatches operator A to Stream0, taking time T A plus the synchronization overhead T syncA for a sum.

[0024] The CPU dispatches operator B to Stream1, taking time T B plus the synchronization overhead T syncB for a sum.

[0025] 2. Synchronous waiting: The CPU needs to wait for both operator A and operator B to complete, so the time taken is max(T A +T syncA , T B +T syncB ).

[0026] 3. Execute operator C separately: After both operator A and operator B are completed, the CPU dispatches operator C to Stream2, taking time T C plus the synchronization overhead T syncC for a sum.

[0027] It can be seen that the total time T 2 = max(T A +T syncA , T B +T syncB ) + T C + T sync . This scheduling and execution method will result in additional CPU synchronization events (T syncA , T syncB , T sync ). In complex AI models, frequent synchronization may lead to cumulative delays and affect system performance.

[0028] In view of this, the present invention provides a heterogeneous parallel computing method and apparatus, which can further reduce the synchronization delay while improving the system parallelism and optimize the overall performance of the system.

[0029] Please refer to Figure 4 , which shows a schematic diagram of the hardware resources adopted in the exemplary embodiment of the present invention.

[0030] Figure 4 In , three groups of heterogeneous engine groups A[m], B[n], and C[k] are taken as examples for illustration, but it does not limit the present invention. In actual implementation, the number of heterogeneous engine groups is not limited to three groups.

[0031] In this exemplary hardware resource architecture, the on-chip general controller is removed, and according to the computing power resource granularity, a task parser and an intra-group hardware synchronization sub-module are respectively equipped for each engine group to maximize the parallelism of the on-chip homogeneous computing power. At the same time, one or more inter-group hardware synchronization sub-modules are also provided on the chip to maximize the parallelism of the single heterogeneous computing power.

[0032] Please refer to Figure 5 , which shows a schematic diagram of operator dispatch in the method provided by the exemplary embodiment of the present invention.

[0033] The present invention provides a heterogeneous parallel computing method. Specifically, the method provided by the present invention includes: Distribute each operator to the corresponding heterogeneous engine group according to the different computing power requirements of the operator. Each heterogeneous engine group includes multiple sub-engines for executing the operator; Number the operators distributed to each heterogeneous engine group; Analyze the data dependence relationship between the operators in each heterogeneous engine group, and insert a counter comparison command word into the operator queue, where the data dependence relationship is recorded by the numbers of the operators in each heterogeneous engine; Perform intra-group synchronization for the sub-engines in the same heterogeneous engine group; Perform inter-group synchronization for each heterogeneous engine group.

[0034] Please refer to Figure 6 together, which shows a schematic diagram of intra-group synchronization in a single engine group in the method provided by the exemplary embodiment of the present invention. Among them, the intra-group synchronization includes: Each sub-engine generates its own sub-engine corresponding group completion flag after completing the operator calculation; Receive all the sub-engine corresponding group completion flags to achieve intra-group synchronization.

[0035] It should be noted that Figure 6This is only exemplary, which illustrates a set of intra-group synchronization resources. However, in actual implementation, multiple sets of intra-group synchronization resources can be introduced and the completion status of different operators can be distinguished by using the operator number (ID), so as to fully utilize the computing resources and achieve parallelism of multiple homogeneous operators in the same task queue.

[0036] Please refer to Figure 7 together, which shows a schematic diagram of inter-group synchronization of multiple engines provided by an exemplary embodiment of the present invention. Among them, the inter-group synchronization includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing the operator calculation; Receive the operator completion pulse signals of each heterogeneous engine group and use a counter to count. The counter value of the counter is the serial number of the operators that have been processed by each heterogeneous engine group; It should be noted that Figure 7 This is only exemplary, which illustrates a set of inter-group synchronization resources. However, in actual implementation, multiple sets of inter-group synchronization resources can be introduced and the execution status of operators in different task queues can be distinguished by using the task queue number (stream ID). At this time, the counter threshold comparison part needs to select the correct counter value according to the task queue number for comparison.

[0037] Each heterogeneous engine group obtains the configured threshold of the counter according to the data dependency relationship; Compare the counter value with the configured threshold. When the counter value is greater than or equal to the configured threshold, an enabling signal for the heterogeneous engine group is issued to achieve fast event synchronization at the cycle level.

[0038] Using the method provided by the technical solution of the present invention to execute the three operators A, B, and C in the technical background, the total time consumption is shortened to max(T A , T B ) + T C +T sync , which fully utilizes the computing power resources on the GPU and AI accelerator chips and does not introduce CPU event synchronization time. When the computing scale is more complex and larger, its performance benefits will be more prominent.

[0039] In a preferred embodiment, the inter-group synchronization further includes shielding the sub-engines that do not participate in the execution in each heterogeneous engine group to further save the system execution time.

[0040] In a preferred embodiment, the inter-group synchronization further includes broadcasting the corresponding operator completion pulse signal to all other heterogeneous engine groups when all the sub-engines participating in the calculation in a certain heterogeneous engine group have completed the calculation.

[0041] In addition, those skilled in the art know that software can also be used for acceleration, that is, a controller and an operator status counter are provided for each heterogeneous engine group, and then multiple atomic counters are managed in the memory space visible to the heterogeneous engines in combination with software and controller firmware (firmware) to achieve this.

[0042] In the solution of using software for acceleration, each operation of the counter is on the order of dozens of cycles, but the total overhead increases linearly with the number of counter operations, which will make the synchronization delay and overhead 100 times larger than that of the hardware acceleration scheme, and the level from thousands of cycles to tens of thousands of cycles varies depending on the busyness of the memory space.

[0043] Please refer to Figure 8 , which shows a schematic diagram of the software resources used in the exemplary embodiments of the present invention.

[0044] Still taking three heterogeneous engine groups A[m], B[n], and C[k] as an example for illustration, in this exemplary resource architecture, the heterogeneous engine groups A[m], B[n], and C[k] are each equipped with a controller CTRL, an operator status counter Status, and multiple groups of atomic counter resources, and they use controller firmware and a software compiler to achieve operator synchronization and data dependency management.

[0045] Next, taking the heterogeneous engine group A[m] as an example, the intra-group synchronization of the heterogeneous engine group is exemplarily described. The intra-group synchronization operations of other heterogeneous engine groups are the same as this, so they will not be elaborated.

[0046] When the compiler generates an operator task list, it needs to insert an atomic operation on the shared memory after each operator. This operation can be an atomic counter or a similar method to indicate that all sub-engines participating in the calculation of the current operator have completed the corresponding calculation. If all sub-engines A[0], A[1],..., A[m - 1] in the heterogeneous engine group A[m] participate in the current operator calculation, the threshold of the atomic counter is m. If only x sub-engines participate in the calculation, the threshold of the atomic counter is x. When multiple operators are in parallel, multiple groups of atomic counters need to be provided to ensure the correct operation of the system function.

[0047] In the case where x sub-engines participate in the calculation, when all these x sub-engines complete the calculation tasks, the sub-engine task parser performs an atomic operation on the shared memory, and the intra-group synchronization overhead time is denoted as T_atomic×x.

[0048] The controller CTRL_A of the heterogeneous engine group A[m] determines the completion status of the current operator by monitoring the atomic counter corresponding to the operator.

[0049] Next, the inter-group synchronization of each heterogeneous engine group is described.

[0050] When the controller CTRL of each heterogeneous engine group detects that an operator is completed, it updates the completed operator status counter of each. It should be noted that the operator status counters of each heterogeneous engine group should be visible to the controllers of other heterogeneous engine groups.

[0051] The compiler expresses the inter-group data dependence relationship of the operators to the controllers of each heterogeneous engine group in the form of event management. Then, the controllers of each heterogeneous engine group view the operator status counters of each to resolve the data dependence relationship. The cost of each view is on the order of dozens of cycles, and the total cost will increase linearly with the number of views.

[0052] Please refer to Figure 9 , which shows the structural block diagram of the device provided by the exemplary embodiment of the present invention.

[0053] The present invention also provides a device for heterogeneous parallel computing. Specifically, the device provided by the present invention includes a compiler and heterogeneous engine groups. Among them, the compiler includes an operator attribution distribution module, an operator sorting module, and an inter-operator dependence relationship analysis module.

[0054] Specifically, the operator attribution distribution module is configured to distribute each operator to the corresponding heterogeneous engine group according to the different computing power requirements of the operators. Each heterogeneous engine group includes multiple sub-engines for executing operators; the operator sorting module is configured to number the operators distributed to each heterogeneous engine group; the inter-operator dependence relationship analysis module is configured to analyze the data dependence relationship between the operators in each heterogeneous engine group and insert a counter comparison command word into the operator queue.

[0055] The heterogeneous engine group includes an intra-group hardware synchronization sub-module and an inter-group hardware synchronization sub-module. Each heterogeneous engine group may also include a task parser.

[0056] Specifically, the intra-group hardware synchronization sub-module is configured to perform intra-group synchronization on the sub-engines in the same heterogeneous engine group; the inter-group hardware synchronization sub-module is configured to perform inter-group synchronization on each heterogeneous engine group.

[0057] Among them, the intra-group synchronization performed by the intra-group hardware synchronization sub-module includes: Each sub-engine generates its own sub-engine corresponding group completion flag after completing the operator calculation; Receiving all the sub-engine corresponding group completion flags to achieve intra-group synchronization; Among them, the intra-group synchronization performed by the inter-group hardware synchronization sub-module includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing the operator calculation; Receive the operator completion pulse signals of each heterogeneous engine group and use a counter to count. The counter value of the counter is the sequence number of the operators that have been processed and completed by each heterogeneous engine group; Each heterogeneous engine group obtains the configuration threshold of the counter according to the data dependency relationship; Compare the counter value with the configuration threshold. When the counter value is greater than or equal to the configuration threshold, an enabling signal for the heterogeneous engine group is issued to achieve inter-group synchronization.

[0058] The inter-group hardware synchronization sub-module is also configured to mask the sub-engines that do not participate in the execution in each heterogeneous engine group.

[0059] The inter-group hardware synchronization sub-module is also configured to broadcast its corresponding operator completion pulse signal to all other heterogeneous engine groups when all the sub-engines participating in the calculation in a certain heterogeneous engine group have completed the calculation.

[0060] Specifically, the compiler can implement the following functions: Operator attribution dispatch: Allocate operators to the corresponding heterogeneous engine groups according to the computing power requirements (such as tensor / vector calculation); Operator sorting and analysis of dependencies between operators: Generate a sequence number for each operator as the basis for comparing dependencies in the inter-group hardware synchronization module; Dynamic computing power allocation: Apply for the corresponding computing power resources according to the operator computing power requirements. In the case of homogeneous computing power scenarios, allocate the operators to the same heterogeneous engine group, and the intra-group synchronization module manages the execution order; in the case of heterogeneous computing power scenarios, allocate the operators to different heterogeneous engine groups and rely on the inter-group synchronization circuit for coordination. Specifically, still taking the model structure in the background technology Figure 1 as an example, when operators A and B use homogeneous computing power, the compiler can distribute the homogeneous computing power within a single engine group to operators A and B according to the computing power requirements, and use the intra-group hardware synchronization sub-module to record the completion status of single or multiple homogeneous operators; when operators A and B use heterogeneous computing power, the compiler distributes operators A and B to the corresponding heterogeneous engine groups, and each heterogeneous engine group uses the inter-group hardware synchronization sub-module to record the completion status of the operators.

[0061] It should be understood that Figure 9 the device shown in

[0062] Although specific functions have been discussed above with reference to specific modules, it should be noted that the functions of each module in the technical solution of the present invention can also be implemented by dividing them into multiple modules, and / or at least some functions of multiple modules can be combined into a single module for implementation. The manner in which a specific module in the technical solution of the present invention performs an action includes that the specific module itself performs the action, or the specific module calls or otherwise accesses an action-performing module (or performs the action in combination with the specific module). Therefore, the specific module that performs the action may include the specific module itself that performs the action and / or another module that performs the action and is called or otherwise accessed by the specific module.

[0063] The technical solution described in the present invention is not limited to the specific examples of the described technical solution. The explanations and descriptions of the present invention in the foregoing and the accompanying drawings are not restrictive. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can also be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, the scope of protection required by the present invention is defined by the claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are covered by the protection scope of the present invention.

Claims

1. A heterogeneous parallel computing method, characterized in that: The method comprises: According to the different computing power requirements of operators, each operator is distributed to the corresponding heterogeneous engine group. Each heterogeneous engine group includes multiple sub-engines for executing operators. Number the operators distributed to each heterogeneous engine group; Parse the data dependencies between operators in each heterogeneous engine group and insert counter comparison command words into the operator queue; Performing intra-group synchronization on each sub-engine in the same heterogeneous engine group, wherein the intra-group synchronization includes: Each sub-engine generates its own sub-engine corresponding group completion flag after completing the operator calculation; Receive the corresponding group completion flags of all sub-engines to achieve synchronization within the group; Inter-group synchronization is performed on each heterogeneous engine group, and the inter-group synchronization includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing operator calculation; Receiving operator completion pulse signals of each heterogeneous engine group and counting them using a counter, wherein the counter value of the counter is the sequence number of the operator that has been processed by each heterogeneous engine group; Each heterogeneous engine group obtains a configuration threshold of a counter according to the data dependency relationship; The counter value is compared with the configured threshold. When the counter value is greater than or equal to the configured threshold, a heterogeneous engine group enable signal is issued to achieve inter-group synchronization.

2. The method according to claim 1, characterized in that: The method also includes removing the on-chip master controller.

3. The method according to claim 1, characterized in that: The inter-group synchronization also includes shielding sub-engines that do not participate in execution in each heterogeneous engine group.

4. The method according to claim 1, characterized in that: The inter-group synchronization also includes broadcasting corresponding operator completion pulse signals to all other heterogeneous engine groups when all sub-engines participating in the calculation in a certain heterogeneous engine group have completed the calculation.

5. The method according to claim 1, characterized in that The data dependency is recorded by the numbers of the operators in each heterogeneous engine.

6. A heterogeneous parallel computing device, characterized in that: The device includes a compiler and a heterogeneous engine group. The compiler comprises: The operator attribution distribution module is configured to distribute each operator to a corresponding heterogeneous engine group according to the different computing power requirements of the operator, and each heterogeneous engine group includes multiple sub-engines for executing the operator; An operator sorting module, configured to number the operators distributed to each heterogeneous engine group; An inter-operator dependency analysis module is configured to parse the data dependency between operators in each heterogeneous engine group and insert a counter comparison command word into the operator queue; The heterogeneous engine group includes: The intra-group hardware synchronization submodule is configured to perform intra-group synchronization on each sub-engine in the same heterogeneous engine group. The intra-group synchronization includes: After completing the operator calculation, each sub-engine generates its own sub-engine corresponding group completion flag; Receive the corresponding group completion flags of all sub-engines to achieve synchronization within the group; The inter-group hardware synchronization submodule is configured to perform inter-group synchronization on each heterogeneous engine group, and the inter-group synchronization includes: Each heterogeneous engine group generates its own operator completion pulse signal after completing operator calculation; Receiving operator completion pulse signals of each heterogeneous engine group and counting them using a counter, wherein the counter value of the counter is the sequence number of the operator that has been processed by each heterogeneous engine group; Each heterogeneous engine group obtains a configuration threshold of a counter according to the data dependency relationship; The counter value is compared with the configured threshold. When the counter value is greater than or equal to the configured threshold, a heterogeneous engine group enable signal is issued to achieve inter-group synchronization.

7. The device according to claim 6, characterized in that The device does not include an on-chip overall controller.

8. The device according to claim 6, characterized in that Each heterogeneous engine group also includes a task parser.

9. The device according to claim 6, characterized in that The inter-group hardware synchronization submodule is further configured to shield sub-engines that do not participate in execution in each heterogeneous engine group.

10. The device according to claim 6, characterized in that The inter-group hardware synchronization submodule is further configured to broadcast corresponding operator completion pulse signals to all other heterogeneous engine groups when all sub-engines participating in calculation in a certain heterogeneous engine group have completed calculation.

11. The device according to claim 6, characterized in that The data dependency is recorded by the numbers of the operators in each heterogeneous engine.

Citation Information

Patent Citations

  • FPGA accelerator and chip for federated learning and privacy calculation

    CN114416182A

  • AI computing power enabled subsystem

    CN119089956A

  • Task scheduling method and related system

    CN119862000A

  • Graph based heterogeneous parallel processing system

    US20190122415A1

  • Ai processor acceleration method and system, and chip

    WO2025055752A1