Task synchronization method and apparatus, electronic device, and storage medium

By establishing a local mapping table between a unique identifier and a barrier identifier between the source chip and the target chip, cross-chip task synchronization is achieved at the hardware level, which solves the problem of low task synchronization efficiency in the existing technology and improves cross-chip parallel computing capabilities.

CN120429131BActive Publication Date: 2025-11-21SUZHOU YIZHU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510933265.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-21
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing task synchronization solutions have low task synchronization efficiency and cross-chip parallel computing capabilities when multiple computing chips work together, mainly due to the reduced efficiency caused by frequent switching between user space and kernel space.

Method used

By receiving task packets carrying unique identifiers from both the source chip and the target chip, randomly assigning barrier identifiers and establishing local mapping tables, and sending a synchronization ready signal through the inter-chip communication path after the target chip completes its local operations, the source chip queries the mapping table based on the unique identifier to trigger a synchronization signal, thereby achieving cross-chip task synchronization and avoiding frequent software-level switching.

Benefits of technology

It improves the synchronization efficiency and parallel computing capabilities of cross-chip tasks, simplifies the synchronization process, and enhances the task pipeline processing capabilities of computing chips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429131B_ABST
    Figure CN120429131B_ABST
Patent Text Reader

Abstract

The present disclosure provides a task synchronization method and device, electronic equipment and storage medium. The method comprises: a source chip and a target chip respectively receiving a task package carrying a unique identifier; the source chip randomly assigns a first barrier identifier to the task package and establishes a local mapping table of the unique identifier and the first barrier identifier; the target chip randomly assigns a second barrier identifier to the same task package and establishes a local mapping table of the unique identifier and the second barrier identifier, and the first barrier identifier and the second barrier identifier are independent of each other; after completing the local operation, the target chip sends a synchronization ready signal containing the unique identifier to the source chip through an inter-chip communication channel; the source chip queries the local mapping table according to the unique identifier, obtains the corresponding first barrier identifier and triggers a synchronization signal; when the source chip counts the number of synchronization signals of the first barrier identifier to meet a preset condition, a cross-chip task is executed, thereby improving the task synchronization efficiency and parallel computing capability of the cross-chip task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of computer technology, and specifically relates to a task synchronization method, apparatus, electronic device and storage medium. Background Technology

[0002] In parallel computing tasks, especially when multiple computing chips work together, task synchronization is a crucial step in ensuring data consistency and task coordination. Existing task synchronization schemes mainly rely on software-level implementations, specifically including: 1. After completing its task, the data receiving side notifies the sending side of its readiness status via a software interface or protocol, informing it that the task is complete and the next operation can proceed. 2. Upon detecting this readiness status information, the data sending side distributes the task from the sending side. While this approach achieves basic task synchronization, it requires frequent switching between user space and kernel space, which reduces the efficiency of cross-chip task synchronization and cross-chip parallel computing capabilities. Summary of the Invention

[0003] In view of the above problems, this disclosure provides a task synchronization method, apparatus, electronic device and storage medium, which aim to improve the task synchronization efficiency and parallel computing capability of cross-chip tasks.

[0004] According to a first aspect of the present disclosure, a task synchronization method is provided, comprising:

[0005] The source chip and the target chip each receive a task packet carrying a unique identifier;

[0006] The source chip randomly assigns a first barrier identifier to the task package and establishes a local mapping table between the unique identifier and the first barrier identifier;

[0007] The target chip randomly assigns a second barrier identifier to the same task package and establishes a local mapping table between the unique identifier and the second barrier identifier. The first barrier identifier and the second barrier identifier are independent of each other.

[0008] After the target chip completes its local operation, it sends a synchronization ready signal containing a unique identifier to the source chip through the inter-chip communication path.

[0009] The source chip queries the local mapping table based on the unique identifier, obtains the corresponding first barrier identifier, and triggers a synchronization signal.

[0010] When the source chip counts the number of synchronization signals identified by the first barrier and meets the preset conditions, it executes a cross-chip task.

[0011] Optionally, the first barrier identifier and the second barrier identifier are only valid within the chip that assigns the identifier, and the synchronization ready signal transmitted across chips does not include the barrier identifier.

[0012] Optionally, when both chips receive the same task packet:

[0013] The source chip parses the task packet to obtain the target chip identifier, and the target chip parses the same task packet to obtain the source chip identifier.

[0014] After the target chip completes its local operation, it sends a synchronization ready signal to the chip corresponding to the source chip identifier.

[0015] After the source chip receives the synchronization ready signal and completes synchronization, it sends the target data to the chip corresponding to the target chip identifier.

[0016] Optionally, when more than two chips receive the same task packet and form a closed-loop data stream:

[0017] Each chip parsing task packet obtains the upstream chip identifier, and the chip identifier points to the following: the upstream chip identifier of the first chip points to the last chip, the upstream chip identifier of the (k+1)th chip points to the kth chip, and k is a positive integer;

[0018] The synchronization ready signal is sent in the opposite direction to the data flow: the (k+1)th chip sends a synchronization ready signal to the kth chip, and after the kth chip completes synchronization, it sends the target data to the (k+1)th chip, where k is the sequence number in the chip loop.

[0019] Optionally, the preset conditions include:

[0020] The number of synchronization signals needs to reach a threshold related to the target chip.

[0021] Optionally, the trigger synchronization signal includes:

[0022] The unique identifier in the received synchronization ready signal is stored in an idle identifier storage unit, and the valid bit of the corresponding unit is activated in the status register.

[0023] Retrieve the unique identifier from the local mapping table and match it with the identifier storage unit associated with the valid bit of the status register;

[0024] When a match is successful, a synchronization signal is triggered and a clear command is sent to the status control interface to reset the corresponding valid bit in the status register and delete the corresponding matching entry from the local mapping table.

[0025] If no match is found, the current matching process is suspended, and the scan is restarted after a preset time without blocking the processing of subsequent task packages.

[0026] Optionally, the inter-chip communication path uses Ethernet or PCIe protocol to transmit synchronization ready signals.

[0027] Optionally, the task synchronization method further includes:

[0028] When the task package performs local synchronization at the source chip, it directly triggers the synchronization signal corresponding to its first barrier identifier.

[0029] According to a second aspect of the present disclosure, a task synchronization apparatus is provided, comprising:

[0030] The local mapping management module is used to randomly assign a barrier identifier that is only valid on this chip to the received task packet with a unique identifier, and maintain a mapping table from unique identifier to barrier identifier;

[0031] The synchronization signal processing module is used to respond to the unique identifier in the external synchronization ready signal, query the local mapping table to obtain the barrier identifier, and then trigger the synchronization signal.

[0032] The collaborative execution module is used to initiate cross-chip tasks when the number of synchronization signals of a specified barrier identifier meets a preset condition.

[0033] Optionally, the synchronization signal processing module includes:

[0034] Multiple identifier storage units are used to store the unique identifier in the externally received synchronization ready signal;

[0035] The status register has a status bit corresponding to an identifier storage unit, which is used to independently indicate the validity of the unique identifier in the corresponding identifier storage unit;

[0036] The status control interface is used to respond to clear commands that specify status bits.

[0037] The asynchronous query unit is used to obtain a unique identifier from the local mapping table and match it with the identifier storage unit associated with the valid bit of the status register. When a match is successful, a synchronization signal is triggered and the corresponding valid bit of the status register is reset through the status control interface. When no match is found, the current matching process is suspended and the scan is restarted after a preset time without blocking the processing of subsequent task packages.

[0038] The local mapping management module is also used to delete the corresponding matching entry from the local mapping table when a match is successful.

[0039] According to a third aspect of the present disclosure, an electronic device is provided, including a processor, a memory, and a program stored in the memory and executable on the processor. The processor includes a task synchronization device as described in any of the preceding claims, and the program, when executed by the processor, implements the steps of the method.

[0040] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program or instructions which, when executed by a processor, implement the steps of the method.

[0041] The embodiments disclosed herein bring the following beneficial effects:

[0042] The task synchronization method provided in this embodiment involves a source chip and a target chip receiving task packets carrying unique identifiers. The source chip randomly assigns a first barrier identifier to the task packet and establishes a local mapping table between the unique identifier and the first barrier identifier. The target chip randomly assigns a second barrier identifier to the same task packet and establishes a local mapping table between the unique identifier and the second barrier identifier. The first barrier identifier and the second barrier identifier are independent of each other. After completing its local operation, the target chip sends a synchronization ready signal containing the unique identifier to the source chip through an inter-chip communication path. The source chip queries its local mapping table based on the unique identifier, obtains the corresponding first barrier identifier, and triggers the synchronization signal. When the source chip counts the number of synchronization signals with the first barrier identifier to meet a preset condition, it executes the cross-chip task. In this way, by utilizing the local mapping table between the barrier identifier and the unique identifier of the task packet at the hardware level, the synchronization ready signal obtained from the outside is converted into the corresponding first barrier identifier and the synchronization signal is triggered, thus realizing the synchronization of cross-chip tasks. This avoids frequent switching between user space and kernel space, improving the task synchronization efficiency and cross-chip parallel computing capability of cross-chip tasks.

[0043] Other features and advantages of embodiments of this disclosure will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of this disclosure. The objects and other advantages of embodiments of this disclosure are realized and obtained through the structures particularly pointed out in the description and the drawings.

[0044] To make the above-described objects, features and advantages of the embodiments of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0045] The above and other objects, features and advantages of the present disclosure will become clearer from the following description of embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0046] Figure 1 This is an application framework diagram of a task synchronization method provided according to an embodiment of the present disclosure;

[0047] Figure 2 This is a flowchart illustrating a task synchronization method according to an embodiment of the present disclosure;

[0048] Figure 3 This is a schematic diagram of the structure of a task synchronization device according to an embodiment of the present disclosure;

[0049] Figure 4 This is a schematic diagram of the process for triggering a synchronization signal according to an embodiment of the present disclosure;

[0050] Figure 5 This is a schematic diagram illustrating the data flow between two chips when they receive the same task packet according to an embodiment of the present disclosure.

[0051] Figure 6 This is a schematic diagram of the data flow between chips when more than two chips receive the same task packet according to an embodiment of the present disclosure;

[0052] Figure 7 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure. Detailed Implementation

[0053] Various embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. In the various drawings, the same elements are indicated by the same or similar reference numerals. For clarity, the various portions in the drawings are not drawn to scale.

[0054] The following terms are used in this article:

[0055] Task synchronization refers to the process of multiple processes running on multiple computing chips cooperating to complete a task. This requires each process to wait for another process to reach its corresponding barrier object before continuing execution. A barrier object acts as a rendezvous point in the program; when a process needs to wait for other processes, it can run to the barrier object. Once all processes have reached the barrier object, it is removed, thus synchronizing the processes. A barrier identifier is used to identify the barrier object.

[0056] Aggregate communication, also known as group communication or aggregate communication, is a collective communication behavior in which multiple processes running on multiple computing chips in a computing cluster participate in communication to form a process group and execute computing tasks. Currently, the Collective Communication Library (CCL) is commonly used to support aggregate communication between multiple computing chips. For example, a computing task may be divided into multiple subtasks, each performing aggregate operations on different computing chips. Synchronization and data migration operations may be performed between the computing chips hosting the subtasks when necessary, thus completing the entire computing task. Aggregate operations are used to implement data arithmetic operations, such as finding minimum and maximum values, summation, logical AND operations, and other user-defined computational algorithms. Aggregate operations include reduce, all-reduce, and reduce-scatter.

[0057] Figure 1 This is an application framework diagram of a task synchronization method provided according to an embodiment of the present disclosure. Figure 1 As shown, the application framework 100 of the task synchronization method in this embodiment includes a host 110 and multiple computing chips 120 (e.g., computing chip 0 to computing chip 3).

[0058] The computing chip 120 is a processor suitable for executing deep learning algorithms, such as a graphics processing unit (GPU), tensor processing unit (TPU), neural network processing unit (NPU), or deep learning processing unit (DPU). Multiple computing chips 120 can be of the same type or different types. The computing chip 120 includes a command processor 121, a data transmission module 122, and multiple computing modules 123 (e.g., computing modules 0 to 7). The computing modules 123 are responsible for executing computational tasks of related subtasks across chip tasks. The computing modules 123 can be ordinary computing modules or in-memory computing modules (CIM). In-memory computing modules reduce data transmission latency and improve computational efficiency by integrating computational logic into storage units. It should be noted that the number of computing chips 120 and the number of computing modules 123 configured within each computing chip 120 can be other quantities, and this disclosure does not limit this.

[0059] In some embodiments, multiple computing chips 120 share hardware and software resources and are uniformly scheduled by a host 110 to collaboratively complete one or more cross-chip tasks. For example, training and inference of a neural network. It is understood that the host 120 can divide a cross-chip task into multiple subtasks and assign these subtasks to different computing chips 120 for execution. Although the subtasks are logically independent, they may have data dependencies. For example, the input of subtask B may be the output of subtask A. Therefore, the task synchronization method of this disclosure embodiment can be used to achieve synchronized operation of cross-chip tasks among the computing chips 120.

[0060] Figure 2 This is a flowchart illustrating a task synchronization method according to an embodiment of the present disclosure. Figure 2 As shown, the task synchronization method of this disclosure includes:

[0061] In step S210, the source chip and the target chip respectively receive a task packet carrying a unique identifier.

[0062] In some embodiments, host 110 sends the same task package (e.g., a CCL task package) to both the source chip and the target chip that need to perform a cross-chip task. The source chip and the target chip each receive a task package carrying a unique identifier (UID). It should be noted that the task package is generated by host 110 according to the requirements of the cross-chip task, indicating the synchronization relationship and data migration path between the various computing chips participating in the task. The unique identifier carried in the task package is used to uniquely identify the task package. In some embodiments, the task package sent by host 110 also includes a source chip identifier and a target chip identifier. The source chip identifier is used to identify the source chip that performs the cross-chip task together with the target chip. The target chip identifier is used to identify the target chip that performs the cross-chip task together with the source chip. The source chip and the target chip can be respectively... Figure 1 One or more computing chips 120.

[0063] In step S220, the source chip randomly assigns a first barrier identifier to the task package and establishes a local mapping table between the unique identifier and the first barrier identifier.

[0064] Refer again Figure 1 ,like Figure 1 As shown, the command processor 121 includes a task synchronization device 300. Figure 3 This is a schematic diagram of a task synchronization device provided according to an embodiment of the present disclosure. Figure 3As shown, the task synchronization device 300 includes a local mapping management module 310, a synchronization signal processing module 320, and a cooperative execution module 330. In some embodiments, in the computing chip 120 (e.g., source chip and target chip), the local mapping management module 310 is used to randomly assign a barrier identifier valid only on the local chip to the received task packet with a unique identifier, and maintain a mapping table from unique identifier to barrier identifier. The synchronization signal processing module 320 is used to trigger a synchronization signal after querying the local mapping table to obtain the barrier identifier in response to the unique identifier in the external synchronization ready signal. The cooperative execution module 330 is used to start a cross-chip task when the number of synchronization signals for a specified barrier identifier meets a preset condition. In some embodiments, the synchronization signal processing module 320 includes multiple identifier storage units 321 (e.g., storage units UID REG0 to UID REG63), a status register (UID STS) 322, a status control interface (UID CLR) 323, and an asynchronous query unit 324. The multiple identifier storage units 321 are used to store the unique identifier in the externally input synchronization ready signal. Each status bit in the status register 322 corresponds to an identifier storage unit 321, and each status bit is used to independently indicate the validity of the unique identifier in the corresponding identifier storage unit 321. The status control interface 323 is used to respond to a clear command specifying a status bit. The asynchronous query unit 324 is used to obtain a unique identifier from the local mapping table and perform a matching search with the identifier storage unit 321 associated with the valid bit of the status register 322. When a match is successful, a synchronization signal is triggered and the corresponding valid bit of the status register 322 is reset through the status control interface 323. When no match is found, the current matching process is suspended, and the scan is re-engaged after a preset time without blocking subsequent task packet processing. The local mapping management module 310 is also used to delete the corresponding matching entry from the local mapping table when a match is successful.

[0065] The following is combined with Figures 1 to 3 The task synchronization method of the present disclosure will be further described.

[0066] In some embodiments, the local mapping management module 310 in the source chip manages a large number of barrier identifiers (BO_ID), provides a barrier identifier configuration interface register, and can also perform resource management on the barrier identifiers in their respective computing chips, performing initialization / arrival / wait operations. In some embodiments, the local mapping management module 310 in the source chip randomly assigns a first barrier identifier (BO_ID) to the task package and establishes a local mapping table between the unique identifier and the first barrier identifier.

[0067] In step S230, the target chip randomly assigns a second barrier identifier to the same task package, and the first barrier identifier and the second barrier identifier are independent of each other.

[0068] In some embodiments, the local mapping management module 310 in the target chip manages a large number of barrier identifiers (BO_ID), provides a barrier identifier configuration interface register, and can also perform resource management on the barrier identifiers in their respective computing chips, performing initialization / arrival / wait operations. In some embodiments, the local mapping management module 310 in the target chip randomly assigns a second barrier identifier to the same task package. It should be noted that the first barrier identifier and the second barrier identifier are independent of each other. The first barrier identifier and the second barrier identifier are only valid within the chip that assigned the identifier, that is, the first barrier identifier is only valid within the source chip, and the second barrier identifier is only valid within the target chip.

[0069] In step S240, after the target chip completes its local operation, it sends a synchronization ready signal containing a unique identifier to the source chip through the inter-chip communication path.

[0070] In some embodiments, after the computing module 123 in the target chip completes the computing task of the corresponding subtask of the cross-chip task (i.e., completes the local operation), it uses the data transmission module 122 in the target chip to send a synchronization ready signal containing a unique identifier from the target chip to the source chip through the inter-chip communication path. In some embodiments, the inter-chip communication path uses Ethernet or PCIe protocol to transmit the synchronization ready signal. It should be noted that the synchronization ready signal transmitted across chips does not contain a barrier identifier. The second barrier identifier randomly assigned by the local mapping management module 310 in the target chip for the same task package will not be transmitted between chips and is only valid within the target chip. In some embodiments, the data transmission module 122 may include a first transmission module (RDMA) and / or a second transmission module (SDMA). Specifically, after the computing module 123 in the target chip completes the local operation, it can use the first transmission module in the target chip to send a synchronization ready signal containing a unique identifier from the target chip to the source chip through the inter-chip communication path.

[0071] Understandably, by limiting the barrier identifier to be valid only within the chip that assigned it, each chip can independently manage synchronization resources without having to share the global barrier identifier space. This reduces the data load of cross-chip communication (only the unique identifier is transmitted, not the barrier identifier), avoids the risk of resource conflicts caused by the exposure of the barrier identifier, and supports different chips to dynamically allocate barrier identifiers according to their own resource pools, thereby improving system scalability and security.

[0072] In step S250, the source chip queries the local mapping table based on the unique identifier, obtains the corresponding first barrier identifier, and triggers a synchronization signal.

[0073] In some embodiments, the command processor 121 in the source chip receives a synchronization ready signal containing a unique identifier sent by the target chip. It performs protocol parsing on the synchronization ready signal to obtain the unique identifier, and then queries a local mapping table based on the unique identifier to obtain the corresponding first barrier identifier and trigger the synchronization signal. It is understood that by utilizing a local mapping table between barrier identifiers and unique identifiers of task packets at the hardware level, the synchronization ready signal obtained from the outside is converted into the corresponding first barrier identifier and the synchronization signal is triggered, achieving cross-chip task synchronization. Task synchronization operations no longer rely on frequent switching at the software level, but instead, the command processor 121 directly manages synchronization resources, improving the efficiency of cross-chip task synchronization and enhancing the cross-chip parallel computing capabilities of the computing chips. By establishing a local mapping table between barrier identifiers and unique identifiers, the synchronization process of upper-layer applications is simplified, making task synchronization operations more flexible and efficient, and facilitating pipelined processing.

[0074] Figure 4 This is a schematic diagram illustrating the process of triggering a synchronization signal according to an embodiment of the present disclosure. Figure 4 As shown, the method for triggering a synchronization signal in this embodiment of the present disclosure includes:

[0075] In step S410, the unique identifier in the received synchronization ready signal is stored in an idle identifier storage unit, and the valid bit of the corresponding unit is activated in the status register.

[0076] In some embodiments, the command processor 121 in the source chip stores the unique identifier in the received synchronization ready signal into an idle identifier storage unit 321, and activates the valid bit of the corresponding identifier storage unit 321 in the status register 322 to independently indicate the validity of the unique identifier in the corresponding identifier storage unit 321.

[0077] In step S420, a unique identifier is obtained from the local mapping table and matched with the identifier storage unit associated with the valid bit of the status register.

[0078] In some embodiments, the asynchronous query unit 324 in the source chip obtains a unique identifier from the local mapping table in the source chip and performs a matching retrieval with the unique identifier stored in the identifier storage unit 321 associated with the valid bit in the status register 322.

[0079] In step S430, when a match is successful, a synchronization signal is triggered and a clear command is sent to the status control interface to reset the corresponding valid bit in the status register and delete the corresponding matching entry from the local mapping table.

[0080] In some embodiments, when the asynchronous query unit 324 in the source chip successfully matches, it triggers a synchronization signal and sends a clear command to the status control interface 323. The status control interface 323 responds to the specified status bit in the clear command and resets the corresponding valid bit in the status register. In some embodiments, the status control interface 323 responds to the clear command and clears the corresponding identifier storage unit 321. In some embodiments, the local mapping management module 310 also deletes the corresponding matching entry from the local mapping table in the source chip when a match is successful.

[0081] Understandably, each status bit in the status register 322 corresponds to an identifier storage unit 321, used to independently indicate the validity of the unique identifier within the corresponding identifier storage unit. This design ensures physical isolation in the status management of each synchronization ready signal, preventing global synchronization blockage due to a single signal anomaly. Simultaneously, the binding mechanism between status bits and identifier storage units enables fine-grained hardware-level control. The status control interface 323 responds to the specified status bit in the clear instruction, precisely resetting the valid bits of the synchronized status register 322 and clearing the identifier storage unit 321. This ensures the dynamic release of status register and identifier storage unit resources, deletes the corresponding local mapping table entry, and maintains consistency between the local mapping table and the information stored in the status register 322 and identifier storage unit 321.

[0082] In step S440, if no match is found, the current matching process is suspended, and the scan is restarted after a preset time without blocking the processing of subsequent task packages.

[0083] In some embodiments, the asynchronous query unit 324 in the source chip suspends the current matching process when no match is found, waits for a preset time, and then re-scans without blocking subsequent task processing. It is understood that the asynchronous query unit 324 performs a unique identifier matching retrieval, triggers a synchronization signal upon successful matching, and immediately clears the corresponding status bit through the status control interface 323. This ensures the real-time responsiveness of task synchronization operations and avoids redundant data accumulation in the mapping table through the automatic deletion mechanism of matching entries. When no match is found, a non-blocking retry strategy is adopted, which not only prevents the risk of deadlock in the task synchronization process but also achieves elastic reclamation of resource usage through periodic scanning at preset time intervals. This significantly improves the throughput of synchronization signals in a multi-chip environment while maintaining the continuous operation of the task pipeline.

[0084] It should be noted that, prior to step S260, the task synchronization method of this embodiment further includes: directly triggering the synchronization signal corresponding to its first barrier identifier when the task package is locally synchronized in the source chip. In some embodiments, when the computing module 123 in the source chip finishes executing the computing task of the corresponding subtask of the cross-chip task (that is, completing the local synchronization of the source chip), the command processor 121 in the source chip directly triggers the synchronization signal corresponding to its first barrier identifier.

[0085] In step S260, when the source chip counts the number of synchronization signals of the first barrier identifier to meet the preset condition, a cross-chip task is executed.

[0086] In some embodiments, the collaborative execution module 330 in the source chip initiates a cross-chip task when the number of synchronization signals for a specified barrier identifier meets a preset condition. The preset condition includes that the number of synchronization signals must reach a threshold related to the target chip. In some embodiments, the local mapping management module 310 in the source chip also performs an initialization operation on the number of synchronization signals corresponding to the first barrier identifier, so as to initialize the number of synchronization signals corresponding to the first barrier identifier to a threshold related to the target chip. It should be noted that the threshold related to the target chip can be understood as the number of target chips performing task synchronization operations with the same source chip. It is understood that by dynamically associating the number of synchronization signals corresponding to the first barrier identifier with the number of target chips performing task synchronization operations with the same source chip, it is possible to adapt to the needs of cross-chip parallel tasks of different scales. When multiple target chips participate in task synchronization operations, the preset condition is adjusted to suit dynamically networked distributed computing environments.

[0087] Figure 5 This is a schematic diagram illustrating the data flow between two chips when they receive the same task packet according to an embodiment of this disclosure. In some embodiments, such as... Figure 5 As shown, the source chip and the target chip can be respectively... Figure 1A computing chip 120 is used. The source chip and the target chip receive the same task packet. Task synchronization operations are performed between the source and target chips. The source chip parses the task packet to obtain the target chip identifier and a unique identifier; the target chip parses the same task packet to obtain the source chip identifier and a unique identifier. After completing its local operation, the target chip sends a synchronization ready signal to the chip corresponding to the source chip identifier. After receiving the synchronization ready signal and completing synchronization, the source chip sends the target data to the chip corresponding to the target chip identifier. It can be understood that the number of synchronization signals corresponding to the first barrier identifier can be initialized to 2. After completing local synchronization in the source chip, the command processor 121 in the source chip directly triggers the synchronization signal corresponding to its first barrier identifier, and the collaborative execution module 330 in the source chip decrements the number of synchronization signals corresponding to the first barrier identifier by 1. After the target chip completes its local operation, the synchronization signal processing module 320 in the source chip responds to the unique identifier in the external synchronization ready signal, obtains the corresponding first barrier identifier, and triggers the synchronization signal. The collaborative execution module 330 in the source chip decrements the number of synchronization signals corresponding to the first barrier identifier by 1. When the collaborative execution module 330 in the source chip counts that the number of synchronization signals corresponding to the current first barrier identifier is 0, it starts a cross-chip task and sends the target data from the source chip to the target chip through the data transmission module 122 (e.g., the second transmission module).

[0088] Figure 6 This is a schematic diagram of the data flow structure between chips when more than two chips receive the same task packet according to an embodiment of the present disclosure. Figure 6 This illustrates the data flow between chips when more than two chips receive the same task packet and form a closed-loop data stream. For example... Figure 6 As shown, n chips receive the same task packet and form a closed-loop data stream, where n is a positive integer greater than 2. The cascaded source and target chips perform task synchronization operations. Each chip parses the task packet to obtain the upstream chip identifier, and the chip identifiers point to the following: the upstream chip identifier of the first chip points to the last chip, the upstream chip identifier of the (k+1)th chip points to the kth chip, that is, the upstream chip identifier of each chip points to the source chip identifier that is the data sender. The synchronization ready signal is sent in the opposite direction to the data stream: the (k+1)th chip sends a synchronization ready signal to the kth chip, and after the kth chip completes synchronization, it sends the target data to the (k+1)th chip, where k is the sequence number in the chip loop, k is a positive integer, and k is less than n. It can be understood that for a closed-loop data stream, the upstream chip is... Figure 5 The source chip shown, the next hop chip of this upstream chip is... Figure 5 The target chip is shown. The specific process of performing cross-chip task synchronization operations between an upstream chip and its next-hop chip according to the task synchronization method in this embodiment of the disclosure is as follows: Figure 5The specific process for performing cross-chip task synchronization between the source chip and the target chip is the same, so it will not be described again here.

[0089] Each computing chip can receive multiple task packets. This computing chip can act as a source chip for some task packets or as a target chip for others. The computing chip assigns barrier identifiers to each task packet and establishes a local mapping table between the barrier identifier of the task packet used as the source chip and the chip identifier. For example, computing chip 1 receives CCL0 and CCL1 task packets, computing chip 2 receives CCL0 task packet, and computing chip 3 receives CCL1 task packet. For CCL0 task packet, computing chip 1 is the source chip and computing chip 2 is the target chip. Therefore, computing chip 1 establishes a local mapping table between the unique identifier of CCL0 task packet and its local barrier identifier; computing chip 2 does not need to establish such a table. Similarly, for CCL1 task packet, computing chip 1 is the target chip and computing chip 3 is the source chip. Therefore, computing chip 3 establishes a local mapping table between the unique identifier of CCL1 task packet and its local barrier identifier; computing chip 1 does not need to establish such a table. This disclosure also provides an electronic device, such as... Figure 7 As shown, it includes a memory 720, a processor 710, and a program stored in the memory 720 and executable on the processor 710. The processor includes the task synchronization device described above. When the program is executed by the processor, it implements the steps of the method described above and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0090] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, this disclosure also provides a storage medium storing a computer program or instructions that, when executed by a processor, can implement the various processes of the embodiments of the above methods.

[0091] Since the instructions stored in the storage medium can execute the steps of the method provided in the embodiments of this disclosure, the beneficial effects achievable by the method provided in the embodiments of this disclosure can be realized, as detailed in the preceding embodiments, and will not be repeated here. Specific implementations of the above operations can be found in the preceding embodiments, and will not be repeated here.

[0092] In summary, according to the embodiments of this disclosure, the source chip and the target chip respectively receive task packets carrying unique identifiers. The source chip randomly assigns a first barrier identifier to the task packet and establishes a local mapping table between the unique identifier and the first barrier identifier. The target chip randomly assigns a second barrier identifier to the same task packet and establishes a local mapping table between the unique identifier and the second barrier identifier. The first barrier identifier and the second barrier identifier are independent of each other. After the target chip completes its local operation, it sends a synchronization ready signal containing a unique identifier to the source chip through the inter-chip communication path. The source chip queries the local mapping table according to the unique identifier, obtains the corresponding first barrier identifier, and triggers the synchronization signal. When the source chip counts the number of synchronization signals with the first barrier identifier to meet a preset condition, it executes the cross-chip task. In this way, by using the local mapping table between the barrier identifier and the unique identifier of the task packet at the hardware level, the synchronization ready signal obtained from the outside is converted into the corresponding first barrier identifier and the synchronization signal is triggered, realizing the synchronization of cross-chip tasks. This avoids frequent switching between user space and kernel space, improves the task synchronization efficiency of cross-chip tasks, and enhances the cross-chip parallel computing capability.

[0093] Finally, it should be noted that the above embodiments are merely examples for clearly illustrating this disclosure and are not intended to limit the implementation. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of this disclosure.

Claims

1. A task synchronization method, comprising: Multiple computing chips receive task packets carrying unique identifiers, wherein one computing chip receives at least one task packet, and the same computing chip serves as either a source chip or a target chip in different task packets. When the computing chip is used as the source chip for the task package, a first barrier identifier is randomly assigned to the task package that is only valid within it, and a local mapping table between the unique identifier and the first barrier identifier is established. When the computing chip is the target chip of the task package, a second barrier identifier is randomly assigned to the task package that is only valid within it. The first barrier identifier and the second barrier identifier are independent of each other and are not visible to each other. After the target chip completes its local operation, it sends a synchronization ready signal containing a unique identifier to the source chip through the inter-chip communication path. The synchronization ready signal does not contain a barrier identifier. The source chip receives synchronization ready signals from multiple task packages, and queries the local mapping table based on the unique identifier to obtain the first barrier identifier corresponding to different task packages and trigger the corresponding synchronization signal. When the source chip counts the number of synchronization signals identified by the first barrier, which meets the threshold related to the number of target chips, a cross-chip task is executed.

2. The task synchronization method according to claim 1, wherein, When two chips receive the same task packet: The source chip parses the task packet to obtain the target chip identifier, and the target chip parses the same task packet to obtain the source chip identifier. After the target chip completes its local operation, it sends a synchronization ready signal to the chip corresponding to the source chip identifier. After the source chip receives the synchronization ready signal and completes synchronization, it sends the target data to the chip corresponding to the target chip identifier.

3. The task synchronization method according to claim 1, wherein, When more than two chips receive the same task packet and form a closed-loop data stream: Each chip parsing task packet obtains the upstream chip identifier, and the chip identifier points to the following: the upstream chip identifier of the first chip points to the last chip, the upstream chip identifier of the (k+1)th chip points to the kth chip, and k is a positive integer; The synchronization ready signal is sent in the opposite direction to the data flow: the (k+1)th chip sends a synchronization ready signal to the kth chip, and after the kth chip completes synchronization, it sends the target data to the (k+1)th chip, where k is the sequence number in the chip loop.

4. The task synchronization method according to claim 1, wherein, The trigger synchronization signal includes: The unique identifier in the received synchronization ready signal is stored in an idle identifier storage unit, and the valid bit of the corresponding unit is activated in the status register. Retrieve the unique identifier from the local mapping table and match it with the identifier storage unit associated with the valid bit of the status register; When a match is successful, a synchronization signal is triggered and a clear command is sent to the status control interface to reset the corresponding valid bit in the status register and delete the corresponding matching entry from the local mapping table. If no match is found, the current matching process is suspended, and the scan is restarted after a preset time without blocking the processing of subsequent task packages.

5. The task synchronization method according to claim 1, wherein, The inter-chip communication path uses Ethernet or PCIe protocol to transmit synchronization ready signals.

6. The task synchronization method according to claim 1, wherein, Also includes: When the task package performs local synchronization at the source chip, it directly triggers the synchronization signal corresponding to its first barrier identifier.

7. A task synchronization device, comprising: The local mapping management module is used to randomly assign a barrier identifier that is only valid within this chip to the received task packet with a unique identifier, and maintain a mapping table from unique identifier to barrier identifier. The synchronization signal processing module is used to respond to the unique identifier in the external synchronization ready signal, query the local mapping table to obtain the barrier identifier, and then trigger the synchronization signal. The synchronization ready signal does not contain the barrier identifier. The collaborative execution module is used to initiate cross-chip tasks when the number of synchronization signals of a specified barrier identifier meets a threshold related to the number of target chips.

8. The task synchronization device according to claim 7, wherein, The synchronization signal processing module includes: Multiple identifier storage units are used to store the unique identifier in the externally received synchronization ready signal; The status register has a status bit corresponding to an identifier storage unit, which is used to independently indicate the validity of the unique identifier in the corresponding identifier storage unit; The status control interface is used to respond to clear commands that specify status bits. The asynchronous query unit is used to obtain a unique identifier from the local mapping table and match it with the identifier storage unit associated with the valid bit of the status register. When a match is successful, a synchronization signal is triggered and the corresponding valid bit of the status register is reset through the status control interface. When no match is found, the current matching process is suspended and the scan is restarted after a preset time without blocking the processing of subsequent task packages. The local mapping management module is also used to delete the corresponding matching entry from the local mapping table when a match is successful.

9. An electronic device comprising a processor, a memory, and a program stored in the memory and executable on the processor, the processor including a task synchronization device as claimed in any one of claims 7-8, wherein the program, when executed by the processor, implements the steps of the method as claimed in any one of claims 1 to 6.

10. A computer-readable storage medium storing a computer program or instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Task processing method, chip, multi-chip module, electronic equipment and storage medium

    CN116339944A