Circuit, method and system for inter-chip communication

By introducing a task scheduling mechanism into the inter-chip communication circuit, the computing unit allows the computing unit to suspend the current task and perform new tasks under specific events, solving the problem of low inter-chip communication efficiency and improving the computing efficiency of the multi-processor system.

CN114691312BActive Publication Date: 2025-09-02CAMBRICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011624931.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-09-02
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

In multiprocessor systems, the inter-chip communication efficiency in the prior art is low, resulting in waste of computing resources and increased communication delay, and the ideal linear acceleration cannot be achieved.

Method used

By introducing a task scheduling method of scheduling units and computing units into the inter-chip communication circuit, the computing unit is allowed to suspend the current task and perform new tasks when a specific event occurs, reducing communication delay and resource waste.

Benefits of technology

It improves the efficiency of inter-chip communication, reduces the waste of computing resources, and improves the overall processing capability of multiprocessor systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691312B_ABST
    Figure CN114691312B_ABST
Patent Text Reader

Abstract

The present disclosure provides a circuit, method, and system for inter-chip communication. The method can be implemented in a computing device, wherein the computing device can be included in a combined processing device, which can also include a universal interconnect interface and other processing devices. The computing device interacts with the other processing devices to jointly complete user-specified computing operations. The combined processing device can also include a storage device, which is connected to the computing device and the other processing devices respectively and is used to store data from the computing device and the other processing devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to the field of inter-chip communication of multi-processors. Background Art

[0002] In neural network training, if a single machine takes T time to train a neural network of size X, then when N identical machines train the same network, ideally, the training time should be T / N, which is also known as the ideal linear speedup ratio. However, ideal linear speedup is unrealistic because it introduces communication overhead. While the computational component can be linearly accelerated, the communication component (such as the AllReduce algorithm) is inherent and cannot be eliminated.

[0003] Therefore, in order to improve the computing power and operating efficiency of the chip, it is necessary to improve the efficiency of inter-chip communication.

[0004] Furthermore, in the prior art, the current task will release the right to use the computing unit only after the execution of the kernel of the current task is completed, resulting in a problem of wasting computing resources of the computing unit. Summary of the Invention

[0005] According to a first aspect of the present disclosure, a method for scheduling tasks in an inter-chip communication circuit is provided, wherein the inter-chip communication circuit includes a first scheduling unit and a first operation unit, and the method includes: receiving first task description information from the first scheduling unit through the first operation unit, and executing a first task according to the first task description information; suspending execution of the first task at the first operation unit in response to the generation of a first specific event; and executing a second task at the first operation unit in response to suspending execution of the first task.

[0006] According to a second aspect of the present disclosure, a method for performing task scheduling in an inter-chip communication circuit is provided, wherein the inter-chip communication circuit includes a second scheduling unit, a second operation unit, and a second storage unit. The method includes: receiving third task description information from the second scheduling unit through the second operation unit; extracting to-be-processed data from the second storage unit through the second operation unit, and executing a third task on the to-be-processed data according to the third task description information; suspending execution of the third task at the second operation unit in response to the generation of a second specific event; and executing a fourth task at the second operation unit in response to the suspension of execution of the third task.

[0007] According to a third aspect of the present disclosure, a circuit for inter-chip communication is provided, comprising a first scheduling unit and a first operation unit, wherein the first operation unit is configured to: receive first task description information from the first scheduling unit, and execute a first task according to the first task description information; suspend the execution of the first task in response to the generation of a first specific event; and execute a second task in response to the suspension of the execution of the first task.

[0008] According to a fourth aspect of the present disclosure, a circuit for inter-chip communication is provided, comprising a second scheduling unit, a second operation unit, and a second storage unit, wherein the second operation unit is configured to receive third task description information from the second scheduling unit; extract data to be processed from the second storage unit, and execute a third task on the data to be processed according to the third task description information; suspend execution of the third task in response to the generation of a second specific event; and execute a fourth task in response to the suspension of execution of the third task.

[0009] According to a fifth aspect of the present disclosure, a chip is provided, comprising the circuit as described above.

[0010] According to a sixth aspect of the present disclosure, a system for inter-chip communication is provided, including a first chip and a second chip.

[0011] According to a seventh aspect of the present disclosure, an electronic device is provided, comprising the chip or system as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0013] Figure 1 A schematic diagram of an inter-chip communication system according to an embodiment of the present disclosure is shown;

[0014] Figure 2 A schematic diagram showing a circuit for inter-chip communication according to another embodiment of the present disclosure is shown;

[0015] Figure 3 A schematic diagram showing a circuit for inter-chip communication according to another embodiment of the present disclosure is shown;

[0016] Figure 4 A schematic diagram showing a circuit for inter-chip communication according to another embodiment of the present disclosure is shown;

[0017] Figure 5A schematic diagram of a system for performing inter-chip communication according to another embodiment of the present disclosure is shown;

[0018] Figure 6 A method for performing inter-chip communication according to one embodiment of the present disclosure is shown;

[0019] Figure 7 A combined processing device is shown;

[0020] Figure 8 An exemplary board is provided;

[0021] Figure 9a and Figure 9b A method for performing inter-chip communication in an inter-chip communication circuit according to one embodiment of the present disclosure is shown;

[0022] Figure 10a and Figure 10b A method for performing inter-chip communication in an inter-chip communication circuit according to another embodiment of the present disclosure is shown;

[0023] Figure 11 An application scenario of hibernating (suspending) and waking up a task in execution in the present disclosure is shown. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0025] It should be understood that the terms "first," "second," "third," and "fourth," etc. in the claims, specification, and drawings of the present disclosure are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0026] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0027] The above is a detailed introduction to the embodiments of the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, changes or modifications made by those skilled in the art based on the ideas of the present disclosure, on the specific implementation methods and application scope of the present disclosure, all fall within the scope of protection of the present disclosure. In summary, the contents of this specification should not be understood as limiting the present disclosure.

[0028] Figure 1 A schematic diagram of an inter-chip communication system according to an embodiment of the present disclosure is shown.

[0029] like Figure 1 As shown, the system includes chip 1 and chip 2, wherein chip 1 includes a first scheduling unit JS (JobScheduler) 1, a first computing unit TC1, a sending unit TX, a first memory management subunit SMMU (Memory Management Unit) 11, a second memory management subunit SMMU12 and a first storage unit LLC / HBM1; chip 2 includes a second scheduling unit JS2, a second computing unit TC2, a receiving unit RX, a third memory management subunit SMMU21, a fourth memory management subunit SMMU22 and a second storage unit LLC / HBM2.

[0030] The arithmetic units TC1 and TC2 may be various types of processing cores, such as an IPU (Image Processing Unit) and the like.

[0031] like Figure 1 As shown, the first scheduling unit JS1 receives task description information (e.g., task descriptor) from the host 1. The task description information includes the task ID, task category, data size, data address, parameter size, processing core (e.g., first computing unit) configuration information, processing core address information, and task splitting information, etc. It should be understood that when receiving information from the host for the first time, it also includes receiving data to be processed from the host. During operation, data can be transmitted between chips, and the first scheduling unit JS1 only receives task description information, without having to receive data to be processed every time.

[0032] The first scheduling unit JS1 loads the received task description information into the first computing unit TC1. After receiving the loaded task descriptor, the first computing unit TC1 feeds back a response message to the first scheduling unit JS1 to indicate successful reception.

[0033] The first computing unit TC1 can split the task into multiple subtasks (jobs) according to the task description information, and schedule and distribute the task to at least one processor core of the first computing unit according to the granularity of the split subtasks, so that at least one processing core of the first computing unit can process the task in parallel.

[0034] After processing received data, the first processing unit TC1 can also store the processed data in the first storage unit LLC (Last Level Cache) / HBM1 (High Bandwidth Memory) via the communication bus. The first memory management subunit SMMU11 is responsible for implementing memory allocation, address translation, and other functions in the first storage unit LLC / HBM1.

[0035] Next, the sending unit Tx obtains the processed data from the first storage unit LLC / HBM1 through the second memory management unit SMMU12, and, triggered by the first scheduling unit JS1, transmits the processed data to the chip 2 based on the first inter-chip communication description information 1 received from the host 1. The inter-chip communication description information 1 is used to describe the communication task between chips.

[0036] The receiving unit Rx in chip 2 receives processed data from the sending unit of chip 1 based on the second inter-chip communication description information 2 received from the host, and stores the received processed data in the second storage unit LLC / HBM2 of chip 2 via the communication bus through the third memory management subunit SMMU21.

[0037] After receiving the processed data, the receiving unit notifies the second scheduling unit JS2. The second scheduling unit JS2 receives the task description template associated with the second inter-chip communication description information sent by the host 1 from the host 2, and determines the communication task according to the second inter-chip communication description information and the task description template.

[0038] The second processing unit TC2 obtains the processed data stored in the second storage unit LLC / HBM2 through the fourth memory management unit SMMU22 and processes the data.

[0039] exist Figure 1In the illustrated solution, the transmitting unit Tx plays a master role relative to the first computing unit TC1 , and is responsible for extracting data from the first storage unit LLC / HBM1 and sending the data off-chip, without being controlled by the first computing unit TC1 .

[0040] exist Figure 1 In the solution shown, data needs to be cached and read inside the chip before being sent and processed. The time required to read the cache will easily lead to an extension of the communication time, and thus a decrease in the processing capability of the multi-chip system.

[0041] In addition, if a fault or time delay occurs when reading data from the first storage unit LLC / HBM1 in chip 1, it will affect the sending unit Tx from sending data from chip 1 to chip 2, which will cause chip 2 to wait for a long time.

[0042] Figure 2 FIG. 1 shows a schematic diagram of a circuit for inter-chip communication according to another embodiment of the present disclosure. Figure 2 As shown, the circuit includes: a first scheduling unit 211, a first operation unit 212 and a sending unit 213, wherein the first scheduling unit 211 is configured to receive first task description information; the first operation unit 212 is configured to receive the first task description information from the first scheduling unit 211, and process the first data according to the first task description information to obtain first processed data; the first operation unit 212 is further configured to transmit the first processed data to the sending unit 213; the sending unit 213 is configured to send the first processed data to the outside of the chip.

[0043] like Figure 1 The system shown differs in that, in Figure 2 In the system shown, the first scheduling unit 211 is only responsible for transmitting the first task description information to the first computing unit 212 , without scheduling and controlling the sending unit 213 .

[0044] The first computing unit 212 receives the first task description information and processes the first data according to the first task description information. The processed first data can be directly sent to the sending unit 213 without the sending unit 213 obtaining the processed data from the storage unit.

[0045] In this embodiment, the sending unit 213 does not act as the main controller, but rather transmits data under the control of the first computing unit 212. In another embodiment, the sending unit 213 transmits the first processed data in response to receiving the first processed data under the control of the first computing unit 212. In this embodiment, the sending unit has a relatively simple function, thereby simplifying the function and / or structure of the sending unit 213.

[0046] Furthermore, in Figure 1 In the embodiment shown, the host directly sends the inter-chip communication description information to the sending unit 213, and the sending unit 213 interacts with the transceiver units on other chips based on the inter-chip communication description information received from the host. Figure 2 In the illustrated embodiment, the communication of the sending unit 213 is directly controlled by the first operation unit 212. In other words, the first operation unit 212 may include inter-chip communication description information, thereby facilitating control of the communication between the sending unit 213 and the outside.

[0047] The first data mentioned above may come from the host, or may come from data generated after being processed by other chips.

[0048] Figure 3 A schematic diagram of a circuit for inter-chip communication according to another embodiment of the present disclosure is shown.

[0049] like Figure 3 As shown, the circuit of the present disclosure further includes a first storage unit 214 , and the first operation unit 212 is further configured to transmit the first processed data to the first storage unit 214 so as to cache the first processed data.

[0050] In the above technical solution of the present disclosure, in addition to sending the first processed data to the sending unit 213, the first operation unit 212 also sends the first processed data to the first storage unit 214 for storage, so as to facilitate further use of the first processed data.

[0051] According to one embodiment of the present disclosure, the first storage unit may include a first memory management subunit 2141 and a first cache subunit 2142. The first memory management subunit 2141 is configured to manage the storage of the first processed data in the first cache subunit 2142. The first memory management subunit 2141 is responsible for implementing functions such as memory allocation and address translation in the storage unit. The first cache subunit 2142 may be an on-chip cache that is responsible for caching data processed by the first operation unit 212.

[0052] It can be seen that, with Figure 1 Compared to the embodiment shown in Figure 3 In the illustrated embodiment, the sending unit 213 does not read data from the first cache sub-unit 2142 , but directly sends the first processed data received from the first operation unit 212 to other chips.

[0053] This approach reduces or eliminates the time overhead generated by the sending unit 213 reading data from the first cache sub-unit 2142, thereby improving communication efficiency.

[0054] In addition, the sending unit 213 does not need to obtain data from the first cache subunit 2142, so there is no need to call a new memory management subunit. Figure 1 There are significant differences in the storage and reading of data using two memory management subunits as shown in FIG.

[0055] Figure 4 A schematic diagram of a circuit for inter-chip communication according to another embodiment of the present disclosure is shown.

[0056] like Figure 4 As shown, the circuit may include: a second scheduling unit 421 , a second operation unit 422 , a receiving unit 423 and a second storage unit 424 .

[0057] The receiving unit 423 can receive data from outside the circuit and send the data to the second storage unit 424 for buffering. At the same time, the receiving unit 423 can also notify the second scheduling unit 421 of the event of receiving the data, so that the second scheduling unit 421 knows that the data has entered the circuit (chip).

[0058] The second scheduling unit 421 may receive the second task description information and instruct the second computing unit 422 to process the data received by the receiving unit 423. The second task description information may be received from the host and loaded into the second computing unit 422.

[0059] The second task description information may be independent of the first task description information described above, or may be associated with the first task description information. For example, the first task description information and the second task description information may be two different parts of one task description information.

[0060] The second operating unit 422 can be configured to obtain data received by the receiving unit 423 and stored in the second storage unit 424 from the second storage unit 424, and receive the second task description information from the second scheduling unit 421; after receiving the second task description information and data, the first processed data can be processed according to the second task description information to obtain second processed data.

[0061] and Figure 1 The embodiment shown differs in that Figure 1 In the embodiment shown in FIG. 1 , the host needs to send inter-chip communication description information to the receiving unit Rx to control the receiving unit Rx to receive and forward data; and in FIG. Figure 4 In the illustrated embodiment, the receiving unit 423 does not need to receive inter-chip communication description information from the host, but only notifies the second scheduling unit 421 of the received data.

[0062] Furthermore, in Figure 1 In the embodiment shown, the receiving unit Rx is controlled by the host as a slave; Figure 4 In the illustrated embodiment, the receiving unit 423 acts as a master to perform operations such as data reception, notification, and storage.

[0063] According to one embodiment of the present disclosure, the second storage unit 424 may include a second memory management subunit 4241 , a third memory management subunit 4242 , and a second cache subunit 4243 .

[0064] exist Figure 4 In the embodiment, the second memory management sub-unit 4241 can manage the storage of data from the receiving unit 423 to the second cache sub-unit 4243.

[0065] The third memory management sub-unit 4242 can manage the transmission of data from the second cache sub-unit 4243 to the second operation unit 422 .

[0066] Figure 2-Figure 4 The circuit shown may be implemented in a chip or in other devices.

[0067] Figure 5 A schematic diagram of a system for performing inter-chip communication according to another embodiment of the present disclosure is shown.

[0068] like Figure 5 As shown, the inter-chip communication system of the present disclosure may include a first chip 510 and a second chip 520, wherein the first chip 510 may include a first scheduling unit 511, a first operation unit 512 and a sending unit 513; the second chip 520 may include a second scheduling unit 521, a second operation unit 522, a receiving unit 523 and a second storage unit 524.

[0069] exist Figure 5 In the system shown, the first scheduling unit 511 receives first task description information, for example, the first task description information may be received from a first host.

[0070] The first computing unit 512 may receive the first task description information from the first scheduling unit 511 , and process the first data according to the first task description information to obtain first processed data.

[0071] exist Figure 5 In the example, the first data may be received by the first scheduling unit 511 from the first host, or may be received directly or indirectly from other chips. The present disclosure does not limit the source of the first data. For example, when the system is first started or initialized, the first scheduling unit 511 may receive the first data from the first host together with the first task description information. When the first data enters the first task description information, Figure 5 After the system shown, processing can be performed between the individual chips.

[0072] After processing the first data according to the first task description information, the first operation unit 512 may generate first processed data, and then directly send the generated first processed data to the sending unit 513 so as to send the first processed data to the second chip 520 .

[0073] Optionally, the first chip 510 may further include a first storage unit 514 , wherein the first operation unit 512 is further configured to transmit the first processed data to the first storage unit 514 so as to cache the first processed data.

[0074] Before, simultaneously with, or after the first processing data is sent from the first computing unit 512 to the sending unit 513, the first processing data may be cached in the first storage unit 514 via the communication bus for subsequent use. If the first processing data is needed in the future, the corresponding data can be read from the first storage unit 514 without having to be received from the first host.

[0075] The first storage unit 514 may include a first memory management subunit 5141 and a first cache subunit 5142 . The first storage unit 510 may manage the storage of the first processed data in the first cache subunit 5142 through the first memory management subunit 5141 .

[0076] After receiving the first processed data, the sending unit 513 may send the first processed data to the second chip 520. Figure 1 The embodiment shown differs in that Figure 1 In the illustrated embodiment, the sending unit Tx plays a master role, but in this embodiment, the sending unit 513 is controlled by the first operation unit 512 and does not play a master role.

[0077] The receiving unit 523 in the second chip 520 may receive the first processed data from the first chip 510 (specifically, the sending unit 513 in the first chip 510 ).

[0078] After receiving the first processing data, the first processing data is sent to the second storage unit 524, and a message of receiving the first processing data is notified to the second scheduling unit 521.

[0079] The second scheduling unit 521 receives the second task description information. The second task description information and the first task description information can be independent or related to each other, for example, they can be different sub-tasks in the same overall task.

[0080] Next, the second scheduling unit 521 may instruct the second computing unit 522 to perform further processing on the first processed data based on the second task description information. It should be understood that the phrase "instructing the second computing unit 522 to process the first processed data" herein means that the second scheduling unit 521 sends an instruction to the second processing unit 522 to start processing, rather than necessarily sending the first processed data itself to the second processing unit 522.

[0081] The second operation unit 522 may receive the second task description information from the second scheduling unit 521 , obtain the first processed data from the second storage unit 524 , and process the first processed data according to the second task description information to obtain second processed data.

[0082] The second storage unit 524 may include a second memory management subunit 5241, a third memory management subunit 5242, and a second cache subunit 5243. The second memory management subunit 5241 may manage the storage of the first processed data from the receiving unit 523 to the second cache subunit 5243, while the third memory management subunit 5242 may manage the transmission of data from the second cache subunit 5243 to the second computing unit 522.

[0083] exist Figure 5 In the system shown, the receiving unit 523 does not need to receive the communication task description information from the second host, but can directly receive the first processing data from the first chip 510 .

[0084] It should be understood that, for the sake of clarity, in the above text, when the first chip 510 only serves as a data sending role, it includes a sending unit 513, and when the second chip 520 only serves as a data receiving role, it includes a receiving unit 523. However, in actual applications and products, the sending unit 513 and the receiving unit 523 are usually combined into a transceiver unit, which is responsible for both receiving and sending. Therefore, although the sending unit 513 and the receiving unit 523 are shown as two different entities in this article, they are essentially the same entity in actual applications and products.

[0085] Furthermore, the first arithmetic unit 512 and the second arithmetic unit 522 can be the same arithmetic unit, the first scheduling unit 511 and the second scheduling unit 521 can be the same scheduling unit, and the first storage unit 514 and the second storage unit 524 can be the same storage unit. They differ only in how the chips they are in function as data transmitters and receivers. For example, the first storage unit 514 can actually have the same internal structure as the second storage unit 524. In other words, the first chip 510 and the second chip 520 have the same structure and are structurally identical during mass production.

[0086] Figure 6 A method for inter-chip communication according to an embodiment of the present disclosure is shown, including: in operation S610, receiving first task description information through a first scheduling unit; in operation S620, processing first data according to the first task description information through a first computing unit to obtain first processed data; in operation S630, transmitting the first processed data to a sending unit through the first computing unit; and in operation S640, sending the first processed data to outside the chip through the sending unit.

[0087] Figure 9a A method for scheduling tasks in an inter-chip communication circuit according to another embodiment of the present disclosure is shown. Figure 2-Figure 3 To describe in detail Figure 9a-9b The method described.

[0088] like Figure 9a As shown, the inter-chip communication circuit may include a first scheduling unit 211 and a first operation unit 212 (as shown in FIG. Figure 2 and Figure 3As shown), the method includes: in operation S910, receiving first task description information from the first scheduling unit 211 through the first operation unit 212, and executing the first task according to the first task description information; in operation S920, at the first operation unit 212, suspending the execution of the first task in response to the generation of a first specific event; in operation S930, executing the second task in response to suspending the execution of the first task through the first operation unit 212.

[0089] Typically, the first computing unit 212 may receive a task description from the first scheduling unit 211 and execute the task described in the first task description, such as performing communication, computing, task loading, and the like.

[0090] When the first computing unit 212 is executing the first task, a specific event may cause the task execution to be interrupted. In this case, the first computing unit 212 suspends the interrupted task and records the point where the task was interrupted locally in the first computing unit 212. In addition, the point where the task was interrupted is also recorded in the first scheduling unit 212, so that both the first computing unit 212 and the first scheduling unit 212 can know the location where the task was interrupted.

[0091] In the prior art, if a task is interrupted, the first computing unit 212 may stop processing and wait for the task to resume. For example, the first computing unit 212 will only relinquish its right to use a task kernel after the task kernel has completed execution. If the task kernel has not yet completed execution, the first computing unit 212 remains in a waiting state. This obviously wastes computing resources of the first computing unit 212.

[0092] In the technical solution disclosed in the present invention, the first computing unit 212 does not need to stop working after suspending the previous task, but can start executing a new task.

[0093] The new task may be pre-stored in the first computing unit 212 , so that the first computing unit 212 obtains the new task locally and executes it after suspending the previous task. The new task may also be scheduled by the first scheduling unit 211 .

[0094] It should be understood that the scheduling of a new task by the first scheduling unit 211 to the first computing unit 212 is not necessarily dependent on whether the first computing unit 212 has suspended the previous task. For example, before the first computing unit 212 suspends the previous task, the first scheduling unit 211 may first send a new task description to the first computing unit 212. Once the first computing unit 212 suspends the previous task, it may immediately begin executing the newly scheduled task.

[0095] In another embodiment, the first operation unit 212 may notify the first scheduling unit 211 within a time period before being suspended that it will suspend the current task within a specific time period; thus, after receiving the notification, the first scheduling unit 211 may send a new task description information to the first operation unit 212 within the time period, and once the first operation unit 212 suspends the previous task, it may immediately start executing the newly scheduled task.

[0096] In another embodiment, Figure 9b In this embodiment, operation S930 may include: at operation S931, in response to the first operation unit 212 suspending execution of the first task, sending second task description information to the first operation unit 212 at the first scheduling unit 211; and at operation S933, in response to receiving the second task description information, executing the second task at the first operation unit 212.

[0097] In this embodiment, the first scheduling unit 211 can monitor whether the first computing unit 212 is about to suspend a task. Once the first scheduling unit 211 detects that the first computing unit 212 has suspended a task, in order to avoid the first computing unit 212 wasting computing resources due to stopping work, new task information can be scheduled to the first computing unit 212 so that the first computing unit 212 can start executing another task after suspending a task, thereby improving computing efficiency.

[0098] It is important to understand that Figure 9b The dotted line in the figure indicates that the notification of "suspend" may or may not exist, that is, the scheduling of a new task by the first scheduling unit does not necessarily depend on whether the first operation unit 211 has suspended the task. In addition, in the above operations, the order is not necessarily as shown by the numbers, but may vary according to actual conditions. For example, if the first scheduling unit 211 needs to respond to the first operation unit 212 suspending the execution of the first task before sending the second task description information to the first operation unit 212, then operation S931 is after operation S920, but if the first scheduling unit 211 does not need to rely on whether the first operation unit 212 is suspended to send the second task description information to the first operation unit 212, then operation S931 may be before, at the same time or after operation S920.

[0099] Further Figure 2 and Figure 3As shown, the inter-chip communication circuit may further include a sending unit 213, at which processed data is received from the first computing unit and the processed data is sent out of the chip, wherein suspending the execution of the first task in response to the generation of the first specific event includes: suspending the execution of the first task in response to the sending unit being blocked from sending the processed data.

[0100] The sending unit 213 is responsible for sending data and messages from one chip to another. While sending data, the sending unit 213 may experience data back pressure, which can cause the first computing unit 212 to stop processing the data. Data back pressure can occur in a variety of situations, including congestion in the channel leading to the downstream chip, preventing data or messages from being sent normally; insufficient storage capacity in the downstream chip, preventing it from receiving new data or messages; or insufficient processing power in the downstream chip, preventing it from further processing received data or messages. In the prior art, once data back pressure occurs, the first computing unit 212 temporarily stops processing and waits for the back pressure to end. Once the back pressure ends, the first computing unit resumes processing the current task. This process can easily waste processing power. In the technical solution of the present disclosure, when data back pressure occurs at the sending unit 213, the first scheduling unit 211 instructs the first computing unit 212 to suspend the current task, records the location of the suspension, and instructs the first computing unit 212 to execute a new task. This significantly improves the efficiency of the first computing unit 212.

[0101] Further Figure 3 As shown, the inter-chip communication circuit may further include a first storage unit 214, wherein the first storage unit 214 is configured to receive processed data from the first operation unit 212 to cache the processed data, wherein suspending the execution of the first task in response to the generation of the first specific event includes: suspending the execution of the first task in response to the first storage unit 214 failing to cache the processed data.

[0102] As above combined Figure 3 As described, the first processed data of the first operation unit 212 is not only sent to the sending unit 213 for transmission to other chips, but also sent to the first storage unit 214 for storage for further use.

[0103] The first storage unit 214 may be unable to further store data for various reasons, such as when the data size of a task is too large, causing the first storage unit 214 to be unable to accommodate the large amount of data. In this case, the first computing unit 212 may pause operations to wait for the data in the first storage unit 214 to be transferred to another location, and then resume operations on the same task when the first storage unit 214 becomes available. Unnecessary idle time in the first computing unit 212 is obviously undesirable.

[0104] According to one embodiment of the present disclosure, Figure 3 As shown, the first storage unit 214 may include a first memory management subunit 2141 and a first cache subunit 2142, and the first memory management subunit 2141 may be configured to manage the storage of the processed data on the first cache subunit 2142; wherein the execution of the first task is suspended in response to at least one of the first memory management subunit 2141 and the first cache subunit 2142 failing to cache the processed data.

[0105] The failure to cache the processed data of the first operation unit 212 may also be caused by a failure of one or both of the first memory management subunit 2141 and the first cache subunit 2142 .

[0106] The above describes that other resources outside the first computing unit 212 cannot transmit or store processed data well, causing the first computing unit 212 to suspend the currently executed task, but the suspension of the currently executed task by the first computing unit 212 is not entirely caused by reasons outside the first computing unit 212 itself.

[0107] According to one embodiment of the present disclosure, suspending execution of the first task in response to generation of the first specific event includes: suspending execution of the first task in response to a suspend instruction being included in the first task.

[0108] According to the above embodiment, in some cases, the task executed by the first operation unit 212 itself (for example, a task kernel) may contain an instruction that actively instructs the first operation unit 212 to suspend, so that when the task executes the instruction, the first operation unit 212 can stop working and suspend the current task according to the instruction. In the prior art, as described above, a new task is executed only after all the kernels are executed. However, in the embodiment of the present disclosure, when the first scheduling unit 211 detects that the first operation unit 212 stops operating or is suspended, it re-schedules a new task for the first operation unit 212 to fully utilize the computing power of the first operation unit 212.

[0109] To facilitate the resumption of suspended tasks, according to one embodiment of the present disclosure, a task execution list may be established at the first computing unit and the first scheduling unit, the task execution list including at least the location where the first task is suspended.

[0110] Whenever a task is suspended, a breakpoint for task execution will be generated in the task. Whenever a task is suspended, the location where the task is suspended can be stored in the first computing unit 212 and / or the first scheduling unit 211. For example, an index can be used to point to the information required when the suspended task is to be executed, including but not limited to the id of the task, the address where the task is to be executed, the data required when continuing to execute the task, and the like. If there are multiple suspended tasks, a list can be formed, and the entries in the list can store the above-mentioned information required for each suspended task to be executed. Whenever a suspended task is resumed, the execution of the suspended task can be resumed based on the location where the task was suspended. Resuming the execution of the suspended task can include reading the data required to continue to execute the task based on the id of the suspended task, starting from the address to be executed, and the like.

[0111] The execution of the first task can be resumed according to the suspended position based on the end of the first specific event. As mentioned above, the first specific event can include multiple situations. For example, if the sending unit 213 is blocked from sending the processed data, the first operation unit 212 will suspend the task. If the blockage of sending the data has been eliminated, the suspended task can be resumed. In another situation, the failure of the first storage unit to store the processed data will also cause the first task to be suspended. In this situation, if the storage of the processed data returns to normal, the suspended task can be resumed. In another situation, the hang-up instruction included in the first task indicates that the hang-up period has expired. According to the instruction of the first task, the execution of the first task can be resumed.

[0112] When there are multiple suspended tasks, the execution of the suspended tasks can be resumed in various orders or ways, such as randomly resuming the execution of one of the multiple tasks, or resuming the execution of the task with the highest priority first based on the priorities of the multiple tasks.

[0113] The priority of resuming tasks can also be determined based on the waiting time of the suspended tasks. Preferably, in order to prevent certain tasks from being suspended for too long, the tasks that have been suspended for the longest time can be resumed first; or, a timer can be set for each suspended task. Once the timer expires, the current task is suspended and the task with the expired timer is resumed.

[0114] The above describes the method for inter-chip communication disclosed in the present invention by taking a circuit as a sending role as an example. Figure 4 and Figure 5 To describe the method of task scheduling in a circuit that acts as a receiving role.

[0115] Figure 10a A method for performing inter-chip communication in an inter-chip communication circuit according to another embodiment of the present disclosure is shown.

[0116] Combine Figure 4 and Figure 5 ,like Figure 10a As shown, the inter-chip communication circuit may include a second scheduling unit 421 , a second operation unit 422 and a second storage unit 424 . The method includes: in operation S1010 , the second operation unit 422 receives third task description information from the second scheduling unit 421 .

[0117] Operation S1010 and Figure 9a and Figure 9b The operation S910 in FIG. 4 is the same, that is, the second scheduling unit 421 sends a task descriptor to the second operation unit 422, so that the second operation unit 422 can execute the corresponding task according to the received task descriptor.

[0118] In operation S1020 , the second operation unit 422 extracts the data to be processed from the second storage unit 424 , and performs a third task on the data to be processed according to the third task description information.

[0119] As the inter-chip communication circuit on the receiving side, the data required by the second operation unit 422 to perform the third task can be extracted from the second storage unit 424. The data in the second storage unit 424 can be received from the receiving unit 423.

[0120] by Figure 5 For example, although the numbers are different in this article, Figure 5 The second operation unit 522 in the chip can extract the required data from the second storage unit 523, and the data in the second storage unit 523 can be received by the receiving unit 523 from the sending unit 513 of another chip.

[0121] Therefore, according to one embodiment of the present disclosure, the inter-chip communication circuit further includes a receiving unit 423, which receives data to be processed from outside the chip and sends it to the second storage unit for storage. After receiving the data, the receiving unit 423 can notify the second scheduling unit 421, so that the second scheduling unit 421 can send third task description information capable of processing the received data to the second operation unit 422.

[0122] Next, in operation S1030 , at the second operation unit 422 , execution of the third task is suspended in response to generation of the second specific event.

[0123] The second specific event may include multiple situations. For example, when there is no receivable data for the third task at the receiving unit 423, the second computing unit 422 may suspend the execution of the third task to avoid the second computing unit 422 entering an idle state and wasting computing power.

[0124] In operation S1040 , the fourth task may be executed by the second operation unit in response to suspending the execution of the third task.

[0125] As described above, in the prior art, if a task is suspended, the second computing unit 422 may stop processing and wait for the task to resume. For example, the second computing unit 422 will only release its right to use a task kernel after the task kernel has completed execution. If the task kernel has not yet completed execution, the second computing unit 422 will remain in a waiting state. Obviously, this will result in a waste of computing resources of the second computing unit 422.

[0126] In the technical solution disclosed in the present invention, the second computing unit 422 does not need to stop working after suspending the previous task, but can start executing a new task.

[0127] The new task may be pre-stored in the second operation unit 422 , so that the second operation unit 422 obtains the new task locally and executes it after suspending the previous task. The new task may also be scheduled by the second scheduling unit 421 .

[0128] It should be understood that the scheduling of a new task by the second scheduling unit 421 to the second computing unit 422 is not necessarily dependent on whether the second computing unit 422 has suspended the previous task. For example, before the second computing unit 422 suspends the previous task, the second scheduling unit 421 may first send a new task description to the second computing unit 422. Once the second computing unit 422 suspends the previous task, it may immediately begin executing the newly scheduled task.

[0129] In another embodiment, the second operation unit 422 may notify the second scheduling unit 421 within a time period before being suspended that it will suspend the current task within a specific time period; thus, after receiving the notification, the second scheduling unit 421 may send a new task description information to the second operation unit 422 within the time period, and once the second operation unit 422 suspends the previous task, it may immediately start executing the newly scheduled task.

[0130] In another embodiment, Figure 10b As shown, in this embodiment, operation S1040 may include: in operation S1041, at the second scheduling unit, fourth task description information may be sent to the second operating unit; and in operation S1043, at the second operating unit, in response to receiving the fourth task description information, executing the fourth task described by the fourth task description information.

[0131] In this embodiment, the second scheduling unit 421 can monitor whether the second computing unit 422 is about to suspend a task. Once the second scheduling unit 421 detects that the second computing unit 422 has suspended a task, in order to avoid the second computing unit 422 wasting computing resources due to stopping work, new task information can be scheduled to the second computing unit 422 so that the second computing unit 422 can start executing another task after suspending a task, thereby improving computing efficiency.

[0132] and Figure 9b The same thing is, Figure 10b The dotted line in also indicates that the notification of "suspend" may or may not exist. In addition, in the above operations, the order is not necessarily as shown by the numbers, but may vary according to actual conditions. For example, if the second scheduling unit 421 needs to respond to the second operation unit 422 suspending the execution of the third task before sending the fourth task description information to the second operation unit 422, then operation S1041 is after operation S1030, but if the second scheduling unit 421 sends the fourth task description information to the second operation unit 422 without depending on whether the second operation unit 422 is suspended, then operation S1041 can be before, at the same time or after operation S1030.

[0133] In addition to suspending the execution of the third task in response to the receiving unit 423 having no receivable data, other second specific events may also occur.

[0134] For example, according to one embodiment of the present disclosure, the execution of the third task may be suspended in response to a failure in extracting the to-be-processed data from the second storage unit 424 .

[0135] As can be seen from the above description, when the second computing unit 422 executes a task, it usually needs to retrieve the data required for the task from the second storage unit 424. However, the second storage unit 424 may malfunction, or the network through which the data is retrieved from the second storage unit 424 may be blocked, making it impossible to retrieve the data. In this case, the second computing unit 422 can suspend the currently executing task. After receiving the suspension message from the second computing unit 422, the second scheduling unit 421 sends a new task to the second computing unit 422, thereby preventing the second computing unit 422 from being idle due to the suspended task.

[0136] The second storage unit 424 may include a second memory management subunit 4241, a third memory management subunit 4242 and a second cache subunit 4243; the second memory management subunit 4241 may manage the storage of the data to be processed from the receiving unit 423 to the second cache subunit 4243; and the third memory management subunit 4242 may manage the transmission of the data to be processed from the second cache subunit 4243 to the second computing unit.

[0137] There are various reasons that may cause the second storage unit 424 to fail during data storage or data retrieval. For example, the second memory management subunit 4241 may fail, making it impossible to manage storage in the second cache subunit 4243; the third memory management subunit 4242 may fail, making it impossible to manage data retrieval in the second cache subunit 4243; or the second cache subunit 4243 may fail, making it impossible to store or retrieve data.

[0138] According to one embodiment of the present disclosure, suspending the execution of the third task in response to the generation of the second specific event may further include: suspending the execution of the third task in response to a suspend instruction being included in the third task.

[0139] Combined with the above Figure 9a and Figure 9b The described implementation method is the same. In some cases, the task executed by the second operation unit 422 itself (for example, a task kernel) may contain an instruction that actively instructs the second operation unit 422 to suspend, so that when the task executes the instruction, the second operation unit 422 stops working according to the instruction and suspends the current task. In the prior art, as described above, a new task is executed only after all the kernels are executed. However, in the embodiment of the present disclosure, when the second scheduling unit 421 detects that the second operation unit 422 stops operating or is suspended, it re-schedules a new task for the second operation unit 422 to fully utilize the computing power of the second operation unit 422.

[0140] To facilitate the resumption of suspended tasks, according to one embodiment of the present disclosure, a task execution list may be established at the second computing unit 422 and the second scheduling unit 421 , and the task execution list may include at least the location where the third task is suspended.

[0141] Whenever a task is suspended, a breakpoint for task execution will be generated in the task; whenever a task is suspended, the location where the task is suspended can be stored in the second operation unit 422 and / or the second scheduling unit 421. For example, the information required when the suspended task is to be executed can be saved, including but not limited to the id of the task, the address where the task is to be executed, the data required when continuing to execute the task, etc. If there are multiple suspended tasks, a list can be formed, and the entries in the list can store the above information required for each suspended task to be executed. Whenever a suspended task is resumed, the execution of the suspended task can be resumed according to the location where the task was suspended. Resuming the execution of the suspended task can include reading the data required to continue to execute the task based on the id of the suspended task, starting from the address to be executed, etc.

[0142] The execution of the third task can be resumed according to the suspended position based on the end of the second specific event. As mentioned above, the second specific event can include multiple situations. For example, the receiving unit 413 suspends the execution of the third task due to the lack of receivable data. If receivable data appears, the suspended third task can be resumed. In another situation, the failure to extract the data to be processed from the second storage unit will also suspend the execution of the third task. In this situation, if the data to be processed can be extracted normally, the suspended third task can be resumed. In another situation, the hang-up instruction included in the third task indicates that the hang-up period has expired. According to the instruction of the third task, the execution of the third task can be resumed.

[0143] When there are multiple suspended tasks, the execution of the suspended tasks can be resumed in various orders or ways, such as randomly resuming the execution of one of the multiple tasks, or resuming the execution of the task with the highest priority first based on the priorities of the multiple tasks.

[0144] The priority of resuming tasks can also be determined based on the waiting time of the suspended tasks. Preferably, in order to prevent certain tasks from being suspended for too long, the tasks that have been suspended for the longest time can be resumed first; or, a timer can be set for each suspended task. Once the timer expires, the current task is suspended and the task with the expired timer is resumed.

[0145] The present disclosure also provides a circuit for inter-chip communication, including a first scheduling unit and a first operation unit, wherein the first operation unit is configured to: receive first task description information from the first scheduling unit and execute a first task according to the first task description information; suspend the execution of the first task in response to the generation of a first specific event; and execute a second task in response to the suspension of the execution of the first task.

[0146] According to one embodiment of the present disclosure, the first scheduling unit is configured to, in response to the first operating unit suspending the execution of the first task, send second task description information to the first operating unit; and the first operating unit is further configured to, in response to receiving the second task description information, execute the second task.

[0147] The present disclosure also provides a circuit for inter-chip communication, including a second scheduling unit, a second operation unit and a second storage unit, wherein the second operation unit is configured to receive third task description information from the second scheduling unit; extract data to be processed from the second storage unit, and execute a third task on the data to be processed according to the third task description information; suspend the execution of the third task in response to the generation of a second specific event; and execute a fourth task in response to the suspension of the execution of the third task.

[0148] According to one embodiment of the present disclosure, the second scheduling unit is configured to: send fourth task description information to the second operating unit in response to the second operating unit suspending the execution of the third task; and the second operating unit is further configured to execute the fourth task in response to receiving the fourth task description information.

[0149] The present disclosure also provides a chip comprising the circuit described above.

[0150] The present disclosure also provides a system for inter-chip communication, including a first chip and a second chip.

[0151] The present disclosure also provides an electronic device, comprising the chip or the system as described above.

[0152] Figure 11 An application scenario of hibernating (suspending) and waking up a task in execution in the present disclosure is shown.

[0153] like Figure 11 As shown, the computing unit 20 and the scheduling unit 10 can communicate with each other, and the master roles of the two are also changing. It should be understood that Figure 11 The arithmetic unit in the above can be combined with TC1, TC2 (such as Figure 1 As shown), the first operation unit 212 and 512 and the second operation unit 422 and 522 correspond to each other, and the scheduling unit 10 can be the same as JS1, JS2 (as shown in FIG. Figure 1 As shown), the first scheduling units 211 and 511 and the second scheduling units 421 and 521 correspond to each other.

[0154] like Figure 11As shown, when the computing unit 20 is processing a task, it is in the master state. When a specific event causes the computing unit 20 to suspend a task, it sends a "sleep" notification to the scheduling unit 10 to inform the scheduling unit 10 that the computing unit 20 will suspend the task, causing the task to be dormant. At this time, the computing unit 20 saves the breakpoint when the task is suspended and synchronizes the breakpoint information with the scheduling unit 10.

[0155] At this point, the scheduling unit 10 enters the master control state. In this state, the scheduling unit 10 will schedule a new task to the computing unit 20, so that the computing unit 20 starts executing the new task after suspending the previous task, thereby re-entering the master control state. When the interrupt event of the suspended task ends, the scheduling unit 10 can wake up the suspended task.

[0156] Therefore, it can be seen that in the technical solution disclosed in the present invention, the computing unit 20 will not always be suspended due to the occurrence of specific events, but will often or always be in a running and processing state, which can improve the utilization rate of the computing unit 20 and further improve the computing power of the entire system.

[0157] The technical solution disclosed herein can be applied to the field of artificial intelligence and implemented as or in an artificial intelligence chip. The chip can exist independently or be included in a computing device.

[0158] Figure 7 A combined processing device 700 is shown, which includes the aforementioned computing device 702, a universal interconnection interface 704, and other processing devices 706. The computing device according to the present disclosure interacts with other processing devices to jointly complete user-specified operations. Figure 7 Schematic diagram of the combined processing device.

[0159] Other processing devices include one or more general-purpose or specialized processors such as central processing units (CPUs), graphics processing units (GPUs), and neural network processors. There is no limit on the number of processors included in other processing devices. Other processing devices serve as interfaces between the machine learning computing device and external data and control, including data handling and basic control of the machine learning computing device, such as starting and stopping it. Other processing devices can also collaborate with the machine learning computing device to complete computing tasks.

[0160] A universal interconnect interface is used to transfer data and control instructions between a computing device (including, for example, a machine learning computing device) and other processing devices. The computing device can obtain required input data from other processing devices and write it to its on-chip storage device; obtain control instructions from other processing devices and write them to its on-chip control cache; and read data from the computing device's storage module and transfer it to other processing devices.

[0161] Optionally, the structure may further include a storage device 708, which is connected to the computing device and the other processing device. The storage device is used to store data in the computing device and the other processing device, and is particularly suitable for data that cannot be fully stored in the internal storage of the computing device or other processing device.

[0162] This combined processing device can be used as a system-on-chip (SoC) in devices such as mobile phones, robots, drones, and video surveillance equipment, effectively reducing the core area of ​​the control unit, increasing processing speed, and lowering overall power consumption. In this case, the combined processing device's universal interconnect interface connects to certain components of the device, such as a camera, display, mouse, keyboard, network card, and Wi-Fi interface.

[0163] In some embodiments, the present disclosure also discloses a chip packaging structure, which includes the above-mentioned chip.

[0164] In some embodiments, the present disclosure further discloses a board card, which includes the above chip packaging structure. Figure 8 , which provides an exemplary board card. In addition to the above-mentioned chip 802, the above-mentioned board card may also include other supporting components, which include but are not limited to: a storage device 804, an interface device 806 and a control device 808.

[0165] The memory device is connected to the chip within the chip package structure via a bus for storing data. The memory device may include multiple groups of memory cells 810. Each group of memory cells is connected to the chip via a bus. It is understood that each group of memory cells may be DDR SDRAM (Double Data Rate SDRAM).

[0166] DDR can double the speed of SDRAM without increasing the clock frequency. DDR allows data to be read out on the rising and falling edges of the clock pulse. The speed of DDR is twice that of standard SDRAM. In one embodiment, the storage device may include 4 groups of storage units. Each group of storage units may include multiple DDR4 particles (chips). In one embodiment, the chip may include 4 72-bit DDR4 controllers, 64 bits of the above 72-bit DDR4 controllers are used for data transmission, and 8 bits are used for ECC verification. In one embodiment, each group of storage units includes multiple double-rate synchronous dynamic random access memories arranged in parallel. DDR can transmit data twice in one clock cycle. A controller for controlling DDR is provided in the chip to control the data transmission and data storage of each storage unit.

[0167] The interface device is electrically connected to the chip in the chip packaging structure. The interface device is used to realize data transmission between the chip and an external device 812 (such as a server or a computer). For example, in one embodiment, the interface device can be a standard PCIE interface. For example, the data to be processed is transferred from the server to the chip through the standard PCIE interface to realize data transfer. In another embodiment, the interface device can also be other interfaces. This disclosure does not limit the specific form of expression of the above-mentioned other interfaces. The interface unit can realize the switching function. In addition, the calculation results of the chip are still transmitted back to the external device (such as a server) by the interface device.

[0168] The control device is electrically connected to the chip. The control device is used to monitor the status of the chip. Specifically, the chip and the control device can be electrically connected via an SPI interface. The control device may include a single-chip microcomputer (MCU). For example, the chip may include multiple processing chips, multiple processing cores or multiple processing circuits, which can drive multiple loads. Therefore, the chip can be in different working states such as high load and light load. The control device can realize the regulation of the working states of multiple processing chips, multiple processing and / or multiple processing circuits in the chip.

[0169] In some embodiments, the present disclosure also discloses an electronic device or apparatus, which includes the above-mentioned board.

[0170] Electronic devices or apparatuses include data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, mobile phones, driving recorders, navigation systems, sensors, cameras, servers, cloud servers, cameras, camcorders, projectors, watches, headphones, mobile storage, wearable devices, vehicles, household appliances, and / or medical devices.

[0171] The transportation vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; and the medical equipment include magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs.

[0172] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this disclosure is not limited by the order of the actions described, because according to this disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for this disclosure.

[0173] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0174] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, optical, acoustic, magnetic or other forms.

[0175] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0176] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software program modules.

[0177] If the integrated unit is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, when the technical solution of the present disclosure can be embodied in the form of a software product, the computer software product is stored in a memory, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0178] The above is a detailed introduction to the embodiments of the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, for those skilled in the art, based on the ideas of the present disclosure, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. A method for scheduling tasks in an inter-chip communication circuit, the inter-chip communication circuit comprising a first scheduling unit and a first operation unit, the method comprising: receiving first task description information from the first scheduling unit through the first computing unit, and executing the first task according to the first task description information; At the first computing unit, in response to generation of a first specific event, suspending execution of a first task so that the first task is in a dormant state, and sending a dormancy notification to the first scheduling unit; At the first scheduling unit, in response to the first computing unit suspending execution of the first task, sending second task description information to the first computing unit; The second task is executed by the first computing unit in response to receiving the second task description information.

2. The method according to claim 1, wherein The inter-chip communication circuit further includes a sending unit, at which the processed data is received from the first computing unit and sent to an off-chip. The suspending execution of the first task in response to the generation of the first specific event includes: In response to the sending unit being blocked from sending the processed data, execution of the first task is suspended.

3. The method according to any one of claims 1 to 2, wherein: The inter-chip communication circuit further includes a first storage unit configured to receive processed data from the first operation unit to cache the processed data, wherein suspending execution of the first task in response to generation of the first specific event includes: In response to the first storage unit failing to cache the processed data, execution of the first task is suspended.

4. The method according to claim 3, wherein: The first storage unit includes a first memory management subunit and a first cache subunit, wherein the first memory management subunit is configured to manage storage of the processed data on the first cache subunit; in In response to at least one of the first memory management subunit and the first cache subunit failing to cache the processed data, execution of the first task is suspended.

5. The method according to claim 1, wherein Suspending execution of the first task in response to generation of the first specific event includes: Execution of the first task is suspended in response to the first task including a suspend instruction.

6. The method according to claim 1, further comprising: A task execution list is established at the first computing unit and the first scheduling unit, wherein the task execution list at least includes a location where the first task is suspended.

7. The method according to claim 6, further comprising: At the first computing unit, in response to the completion of the first specific event, execution of the first task is resumed according to the suspended position.

8. The method according to claim 6, wherein: When there are multiple suspended tasks, randomly resuming execution of one of the plurality of tasks; Execution of a task with a high priority is resumed based on the priorities of the multiple tasks.

9. A method for scheduling tasks in an inter-chip communication circuit, the inter-chip communication circuit comprising a second scheduling unit, a second operation unit, and a second storage unit, the method comprising: receiving third task description information from the second scheduling unit through the second computing unit; extracting the data to be processed from the second storage unit through the second computing unit, and performing a third task on the data to be processed according to the third task description information; At the second operation unit, in response to the generation of the second specific event, suspending execution of the third task so that the third task is in a dormant state, and sending a dormancy notification to the second scheduling unit; At the second scheduling unit, in response to the second computing unit suspending execution of the third task, sending fourth task description information to the second computing unit; The fourth task is executed by the second computing unit in response to receiving the fourth task description information.

10. The method according to claim 9, wherein: The inter-chip communication circuit further includes a receiving unit, at which to-be-processed data is received from outside the chip and sent to the second storage unit for storage; wherein the execution of the third task is suspended in response to the generation of the second specific event: In response to the receiving unit having no receivable data, execution of the third task is suspended.

11. The method according to claim 10, wherein: The second storage unit includes a second memory management subunit, a third memory management subunit and a second cache subunit; Managing the storage of the to-be-processed data from the receiving unit to the second cache subunit by the second memory management subunit; The third memory management subunit manages the transmission of the to-be-processed data from the first cache subunit to the second computing unit.

12. The method according to any one of claims 9 to 11, wherein: Suspending execution of the third task in response to generation of the second specific event includes: In response to failure in extracting the to-be-processed data from the second storage unit, execution of the third task is suspended.

13. The method according to claim 9, wherein: Suspending execution of the third task in response to generation of the second specific event includes: Execution of the third task is suspended in response to the suspend instruction being included in the third task.

14. The method according to claim 9, further comprising: A task execution list is established at the second computing unit and the second scheduling unit, wherein the task execution list at least includes a location where the third task is suspended.

15. The method according to claim 14, further comprising: At the second computing unit, in response to the end of the second specific event, execution of the third task is resumed according to the suspended position.

16. The method according to claim 15, wherein When there are multiple suspended tasks, randomly resume execution of one of the multiple tasks; or Execution of a task with a high priority is resumed based on the priorities of the multiple tasks.

17. A circuit for inter-chip communication, comprising a first scheduling unit and a first operation unit, wherein The first operation unit is configured as follows: receiving first task description information from the first scheduling unit, and executing the first task according to the first task description information; In response to the generation of a first specific event, suspending execution of a first task so that the first task is in a dormant state, and sending a dormancy notification to the first scheduling unit; executing a second task in response to receiving the second task description information; The first scheduling unit is configured to, in response to the first computing unit suspending execution of the first task, send the second task description information to the first computing unit.

18. A circuit for inter-chip communication, comprising a second scheduling unit, a second operation unit, and a second storage unit, wherein: The second operation unit is configured as follows: receiving third task description information from the second scheduling unit; extracting the data to be processed from the second storage unit, and performing a third task on the data to be processed according to the third task description information; In response to the generation of the second specific event, suspending the execution of the third task so that the third task is in a dormant state, and sending a dormancy notification to the second scheduling unit; executing a fourth task in response to receiving the fourth task description information; The second scheduling unit is configured to: in response to the second computing unit suspending execution of the third task, send the fourth task description information to the second computing unit.

19. A chip comprising the circuit according to any one of claims 17-18. 20 . A system for inter-chip communication, comprising a first chip and a second chip, wherein the first chip comprises the circuit according to claim 17 , and the second chip comprises the circuit according to claim 18 .

21. An electronic device comprising the chip according to claim 19 or the system according to claim 20.

Citation Information

Patent Citations

  • Communication device, neural network processing chip, combination device and electronic equipment

    CN111381958A