Task processing device, related die, and processing method
By introducing the solidified interface between the adapter decoupled scheduler and the accelerator in the chiplet, the problem that the scheduler cannot schedule different versions of the accelerator is solved, and the overall performance of the chiplet is improved and the independent evolution of accelerator resources is achieved.
Patent Information
- Application Number
- PCT/CN2024/143539
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
In chiplet products, since each die may be manufactured by different manufacturers according to different processes and specifications, the scheduler cannot schedule different versions of the accelerator normally, which reduces the overall performance of the chiplet.
By introducing an adapter into the chiplet to block the solid interface between the scheduler and other die accelerators, decouple the relationship between the scheduler and the accelerator, so that the scheduler can schedule accelerators that cannot be directly scheduled through the adapter, and directly schedule compatible accelerators to ensure that all accelerator resources can be effectively utilized.
It improves the overall performance of chiplet products, supports the independent evolution of various die-based accelerator resources, and ensures efficient and stable task processing.
Smart Images

Figure CN2024143539_10072025_PF_FP_ABST
Abstract
Description
A task processing device, related dies, and a processing method This application claims the priority of a Chinese patent application with the application number 202410009457.2 and the application title "A task processing device, related dies, and a processing method", which was filed with the China National Intellectual Property Administration on January 3, 2024. The entire content of this Chinese patent application is incorporated herein by reference in its entirety. Technical Field This application relates to the field of chip technology, and particularly to a task processing device, related dies, and a processing method. Background Art After the integration development of system on chip (SoC) has reached the post-Moore era, SoC chips have become larger and larger, with higher and higher manufacturing processes. Problems such as a decline in yield, an increase in design costs, and a longer R & D cycle have also emerged. For this reason, the industry has proposed the chiplet technology, which is expected to become an important way to continuously improve the integration level and chip computing power. The design concept of SoC is to integrate many functional modules into one chip to form a large chip, such as a central processing unit (CPU), a memory, and interfaces. Different from SoC, the design concept of chiplet is to split a large chip with rich functions and a large area into multiple dies with specific functions, and then combine and package these dies with specific functions according to functional requirements to obtain a chiplet that can meet the requirements. Based on this design concept, the chiplet technology is expected to improve the flexibility of chip design, increase the yield, and solve the limitations brought by the physical limits of nanometer processes. Currently, various accelerator resources in chiplet products are distributed on different dies, and the scheduler needs to schedule accelerators with different functions on different dies to process tasks. Since each die in a chiplet product may be manufactured and evolved by different manufacturers according to different processes and specifications, it is very likely that there will be a situation where multiple different versions of dies in a chiplet product are interconnected. This will result in the scheduler needing to interface with various different versions of accelerators, which may cause the scheduler to be unable to normally schedule the accelerators, reducing the overall performance of the chiplet. Therefore, how to provide a chiplet structure that can normally schedule accelerator resources and ensure performance is an urgent problem to be solved. Summary of the Invention Embodiments of this application provide a task processing device, related dies, and a processing method, which can decouple the scheduler from accelerators deployed on other dies, ensure that the scheduler can normally schedule cross-die accelerator resources, and thus improve the overall performance of the task processing device. The present application will be introduced from different aspects below. It should be understood that the implementation manners and beneficial effects of the different aspects below can be referred to each other. In a first aspect, the present application provides a task processing device, where the device includes a first die and a second die; the first die includes a scheduler and M first accelerators, and the second die includes an adapter and N second accelerators; wherein, the scheduler is configured to: obtain a first task set from an upper-layer software, the first task set including a first task and a second task; schedule one or more target first accelerators among the M first accelerators to process the first task; the adapter is configured to: obtain a second task set from the scheduler, the second task set including the second task; schedule one or more target second accelerators among the N second accelerators to process the second task. Different from the existing chiplet products including a scheduler and accelerators, where the scheduler directly schedules various accelerators on different dies in the product, there may be a problem that the scheduler cannot normally schedule the accelerators when docking with different versions of accelerators. In the embodiment of the present application, by adding an adapter to shield the fixed interfaces between the scheduler in the chiplet and various accelerators on other dies, the scheduler and the accelerators on other dies are decoupled. When it is necessary to schedule an accelerator, for those accelerators that may not be normally scheduled by the scheduler, that is, the accelerators that are not compatible with the scheduler, the scheduler can allocate the tasks that need to be processed by these accelerators to the adapter, and then the adapter schedules these accelerators to process the corresponding tasks, that is, the scheduler does not directly schedule these accelerators that may not be normally scheduled; for other accelerators that can be normally scheduled by the scheduler, that is, the accelerators that are compatible with the scheduler, after the scheduler obtains the tasks that need to be processed by these accelerators, it can still directly schedule these accelerators to process the corresponding tasks, so that various accelerator resources of the chiplet product can be scheduled as much as possible to process tasks, thereby improving the overall performance of the chiplet, and can support the independent evolution of accelerator resources on various dies in the chiplet. In a possible implementation manner, the scheduler communicates with the M first accelerators through a first protocol, the adapter communicates with the N second accelerators through a second protocol, and the scheduler communicates with the adapter through a third protocol. In the embodiment of the present application, the scheduler and the first accelerators, the adapter and the second accelerators, and the scheduler and the adapter can communicate through a separate set of interaction protocols respectively, so that the scheduler can normally schedule the first accelerators, and the adapter can communicate with the scheduler normally and schedule the second accelerators to process the tasks sent by the scheduler normally. In a possible implementation, the first protocol is different from the second protocol. The first protocol includes the Advanced eXtensible Interface (AXI) protocol, and the second protocol includes the Advanced High-Performance Bus (AHB) protocol. In a possible implementation, the scheduler includes a first interface, a second interface, and a third interface, and the adapter includes a fourth interface and a fifth interface. The scheduler is connected to the M first accelerators through the first interface, and the first interface complies with the first protocol. The scheduler is connected to the fourth interface of the adapter through the second interface, and the second interface and the fourth interface comply with the third protocol. The scheduler obtains the first task set from the upper-layer software through the third interface. The adapter is connected to the N second accelerators through the fifth interface, and the fifth interface complies with the second protocol. In the embodiments of the present application, various interfaces used by the scheduler and the adapter should be interfaces that comply with the corresponding protocols to avoid performance degradation of the task processing device due to incompatibility between the interfaces and the protocols. In a possible implementation, the scheduler is further configured to: Obtain first configuration information. The first configuration information is used to indicate a first configuration relationship between the first task and the one or more target first accelerators. The first configuration relationship includes one or more of the correspondence between the task type of the first task and the type of the target first accelerator, the number of target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task. Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information. The first task information includes one or more of the task type or task identifier of the first task. In the embodiments of the present application, the scheduler can determine what type of accelerators to schedule, the number of accelerators, and the security level of the accelerators, etc., based on the task information and configuration information of the task to be processed, and can process the tasks accurately, timely, and securely. In a possible implementation, the adapter is further configured to: Obtain second configuration information. The second configuration information is used to indicate a second configuration relationship between the second task and the one or more target second accelerators. The second configuration relationship includes one or more of the correspondence between the task type of the second task and the type of the target second accelerator, the number of target second accelerators corresponding to the second task, or the security level of the target second accelerators corresponding to the second task. Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task. In the embodiments of the present application, the adapter can determine what type of accelerator to schedule, the number of accelerators, and the security level of the accelerator, etc. based on the task information and configuration information of the task to be processed, and can process the task accurately, timely, and securely. In a possible implementation, the scheduler is further configured to: Obtain first execution information; the first execution information is used to indicate the execution status of the first task; When the execution status of the first task includes an error in task execution, send a first reset command; the first reset command is used to indicate resetting the accelerator that made an error in executing the first task; or, When the execution status of the first task includes completion of task execution, send a first release command; the first release command is used to indicate releasing the accelerator that completed the execution of the first task. In the embodiments of the present application, after the scheduler schedules the accelerator to process the task, it can obtain the situation of the accelerator processing the task, such as whether there is an error in task processing, whether the task processing is completed, etc., so as to perform corresponding operations to enable subsequent tasks to proceed smoothly. In a possible implementation, the adapter is further configured to: Obtain second execution information; the second execution information is used to indicate the execution status of the second task; When the execution status of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that made an error in executing the second task; or, When the execution status of the second task includes completion of task execution, send a second release command; the second release command is used to indicate releasing the accelerator that completed the execution of the second task. In the embodiments of the present application, after the adapter schedules the accelerator to process the task, it can obtain the situation of the accelerator processing the task, such as whether there is an error in task processing, whether the task processing is completed, etc., so as to perform corresponding operations to enable subsequent tasks to proceed smoothly. In a possible implementation, the M first accelerators are main accelerators, and the N second accelerators are auxiliary accelerators. In the embodiments of the present application, the main accelerator and the scheduler can be deployed together on the same die, and the scheduler directly schedules the main accelerator, so that when the task processing device completes the tasks of the main scenario, the scheduling path of the scheduler for the main accelerator is reduced, enabling the main accelerator to respond to tasks faster, thereby further improving the overall performance of the task processing device. In a possible implementation, the scheduler includes one or more virtual schedulers, and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators; specifically, the scheduler is configured to: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers. In the embodiments of the present application, one or more virtual schedulers can be abstracted in the scheduler, and the accelerator is scheduled through virtual scheduling resources, which facilitates the scheduler to schedule and manage the tasks to be processed and the accelerators, and improves the scheduling flexibility. In a possible implementation, the scheduler is further configured to: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate adjusting the number of virtual schedulers in the scheduler, and / or is used to indicate adjusting the number of first accelerators corresponding to the target virtual schedulers in the scheduler; Adjust the number of virtual schedulers based on the third configuration information, and / or adjust the number of first accelerators corresponding to the target virtual schedulers; the target virtual scheduler is any one of the one or more virtual schedulers. In the embodiments of the present application, the number of virtual schedulers in the scheduler and the number of accelerators corresponding thereto also support being adjusted through configuration, which can meet the requirements of more scenarios. In a possible implementation, the adapter includes one or more virtual adapters, and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators; specifically, the adapter is configured to: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters. In the embodiments of the present application, one or more virtual adapters can be abstracted from an adapter, and an accelerator is scheduled through virtual scheduling resources, which facilitates the scheduling and management of the tasks to be processed and the accelerator by the adapter, and improves the scheduling flexibility. In a possible implementation manner, the scheduler is further configured to: Obtain fourth configuration information from the upper-layer software and send the fourth configuration information to the adapter; the fourth configuration information is used to indicate adjusting the number of virtual adapters in the adapter, and / or, is used to indicate adjusting the number of second accelerators corresponding to a target virtual adapter in the adapter; The adapter is further configured to: Receive the fourth configuration information and adjust the number of virtual adapters based on the fourth configuration information, and / or, adjust the number of second accelerators corresponding to the target virtual adapter; the target virtual adapter is any one of the one or more virtual adapters. In the embodiments of the present application, the number of virtual adapters in the adapter and the number of accelerators corresponding thereto also support being adjusted through configuration, which can meet the requirements of more scenarios. In a possible implementation manner, the first die and the second die are isomorphic or heteromorphic, and / or, the first die and the second die are homogeneous or heterogeneous. In a second aspect, the present application provides a first die, which includes a scheduler and M first accelerators; wherein, The scheduler is configured to: Obtain a first task set from the upper-layer software, where the first task set includes a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; Send a second task set to the adapter of the second die; the second task set includes the second task. In the embodiments of the present application, an adapter is added to shield the fixed interface between the scheduler in the chiplet and various accelerators on other dies, decoupling the scheduler from the accelerators on other dies. When the scheduler needs to schedule an accelerator to process a task, for those accelerators that may not be scheduled normally by the scheduler, that is, accelerators that are not compatible with the scheduler, the scheduler can assign the tasks that need to be processed by these accelerators to the adapter, and then the adapter schedules these accelerators to process the corresponding tasks, that is, the scheduler does not directly schedule these accelerators that may not be scheduled normally; for other accelerators that can be scheduled normally by the scheduler, that is, accelerators that are compatible with the scheduler, after the scheduler obtains the tasks that need to be processed by these accelerators, it can still directly schedule these accelerators to process the corresponding tasks, so that the accelerator resources of the chiplet product can all be scheduled to process tasks as much as possible, thereby ensuring the overall performance of the chiplet, and can support the independent evolution of accelerator resources on various dies in the chiplet accordingly. In a possible implementation manner, the scheduler communicates with the M first accelerators through a first protocol. In a possible implementation manner, the scheduler includes a first interface, a second interface, and a third interface; the scheduler is connected to the M first accelerators through the first interface, and the first interface complies with the first protocol; the scheduler is connected to a fourth interface of the adapter through the second interface, and the second interface and the fourth interface comply with the third protocol; the scheduler obtains the first task set from the upper-layer software through the third interface. In a possible implementation manner, the scheduler is further configured to: Obtain first configuration information; the first configuration information is used to indicate a first configuration relationship between the first task and the one or more target first accelerators; the first configuration relationship includes one or more of the correspondence between the task type of the first task and the target first accelerator type, the number of target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task; Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information; the first task information includes one or more of the task type or task identifier of the first task. In a possible implementation manner, the scheduler is further configured to: Obtain first execution information; the first execution information is used to indicate the execution status of the first task; When the execution status of the first task includes an error in task execution, send a first reset command; the first reset command is used to indicate resetting the accelerator that makes an error in executing the first task; or, The execution of the first task includes sending a first release command when the task is completed; the first release command is used to indicate the release of the accelerator that has completed the execution of the first task. In a possible implementation, the M first accelerators are master accelerators. In a possible implementation, the scheduler includes one or more virtual schedulers, and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators; specifically, the scheduler is configured to: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers. In a possible implementation, the scheduler is further configured to: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate an adjustment to the number of virtual schedulers in the scheduler and / or to indicate an adjustment to the number of first accelerators corresponding to the target virtual schedulers in the scheduler; Adjust the number of virtual schedulers and / or adjust the number of first accelerators corresponding to the target virtual schedulers based on the third configuration information; the target virtual scheduler is any one of the one or more virtual schedulers. In a third aspect, the present application provides a second die, which includes an adapter and N second accelerators; wherein, The adapter is configured to: Obtain a second task set from the scheduler of the first die, the second task set including second tasks; Schedule one or more target second accelerators among the N second accelerators to process the second tasks. In the embodiments of the present application, by adding an adapter to shield the fixed interface between the scheduler in the chiplet and various accelerators on other dies, the scheduler is decoupled from the accelerators on other dies. When scheduling an accelerator is required, for those accelerators that may not be scheduled normally by the scheduler, that is, accelerators that are not compatible with the scheduler, the scheduler can assign the tasks that need to be processed by these accelerators to the adapter, and then the adapter schedules these accelerators to process the corresponding tasks, that is, the scheduler does not directly schedule these accelerators that may not be scheduled normally by it; for other accelerators that can be scheduled normally by the scheduler, that is, accelerators that are compatible with the scheduler, after the scheduler obtains the tasks that need to be processed by these accelerators, it can still directly schedule these accelerators to process the corresponding tasks, so that the accelerator resources of the chiplet product can be scheduled as much as possible to process tasks, thereby ensuring the overall performance of the chiplet, and can support the independent evolution of accelerator resources on various dies in the chiplet accordingly. In a possible implementation manner, the adapter and the N second accelerators communicate with each other through a second protocol. In a possible implementation manner, the adapter includes a fourth interface and a fifth interface. The adapter is connected to the N second accelerators through the fifth interface, and the fifth interface complies with the second protocol; the adapter is connected to the second interface of the scheduler through the fourth interface, and the second interface and the fourth interface comply with the third protocol. In a possible implementation manner, the adapter is further configured to: Obtain second configuration information; the second configuration information is used to indicate the second configuration relationship between the second task and the one or more target second accelerators; the second configuration relationship includes one or more of the correspondence between the task type of the second task and the target second accelerator type, the number of target second accelerators corresponding to the second task, or the security level of the target second accelerator corresponding to the second task; Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task. In a possible implementation manner, the adapter is further configured to: Obtain second execution information; the second execution information is used to indicate the execution status of the second task; When the execution status of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that makes an error in executing the second task; or, The execution of the second task includes sending a second release command when the task is completed; the second release command is used to indicate the release of the accelerator that has completed the execution of the second task. In a possible implementation, the N second accelerators are auxiliary accelerators. In a possible implementation, the adapter includes one or more virtual adapters, and one virtual adapter in the one or more virtual adapters corresponds to one or more of the N second accelerators; specifically, the adapter is configured to: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters. In a possible implementation, the adapter is further configured to: Obtain fourth configuration information; the fourth configuration information is used to indicate an adjustment to the number of virtual adapters in the adapter and / or to indicate an adjustment to the number of second accelerators corresponding to the target virtual adapters in the adapter; Adjust the number of virtual adapters and / or adjust the number of second accelerators corresponding to the target virtual adapters based on the fourth configuration information; the target virtual adapter is any one of the one or more virtual adapters. In a fourth aspect, the present application provides a processing method applied to a scheduler in a first die, and the first die further includes M first accelerators; the method includes: Obtain a first task set from upper-layer software, where the first task set includes a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; Send a second task set to the adapter of the second die; the second task set includes the second task. In a possible implementation, the method further includes: Obtain first configuration information; the first configuration information is used to indicate a configuration relationship between the first task and the one or more target first accelerators; the first configuration relationship includes one or more of a correspondence between the task type of the first task and the type of the target first accelerator, the number of target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task; Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information; the first task information includes one or more of the task type or task identifier of the first task. In a possible implementation manner, the method further includes: Obtain first execution information; the first execution information is used to indicate the execution situation of the first task; When the execution situation of the first task includes an error in task execution, send a first reset command; the first reset command is used to indicate resetting the accelerator that has an error in executing the first task; or, When the execution situation of the first task includes the completion of task execution, send a first release command; the first release command is used to indicate releasing the accelerator that has completed the execution of the first task. In a possible implementation manner, the scheduler includes one or more virtual schedulers, and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators; the method further includes: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers. In a possible implementation manner, the method further includes: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate adjusting the number of virtual schedulers in the scheduler, and / or, is used to indicate adjusting the number of first accelerators corresponding to the target virtual scheduler in the scheduler; Adjust the number of virtual schedulers based on the third configuration information, and / or, adjust the number of first accelerators corresponding to the target virtual scheduler; the target virtual scheduler is any one of the one or more virtual schedulers. In a fifth aspect, the present application provides a processing method, which is applied to an adapter in a second die, and the second die further includes N second accelerators; the method includes: Obtain a second task set from the scheduler of the first die, and the second task set includes second tasks; Schedule one or more target second accelerators among the N second accelerators to process the second task. In a possible implementation manner, the method further includes: Obtain second configuration information; the second configuration information is used to indicate the configuration relationship between the second task and the one or more target second accelerators; the second configuration relationship includes one or more of the correspondence between the task type of the second task and the target second accelerator type, the number of target second accelerators corresponding to the second task, or the security level of the target second accelerators corresponding to the second task; Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task. In a possible implementation manner, the method further includes: Obtain second execution information; the second execution information is used to indicate the execution situation of the second task; When the execution situation of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that has an error in executing the second task; or, When the execution situation of the second task includes the completion of task execution, send a second release command; the second release command is used to indicate releasing the accelerator that has completed the execution of the second task. In a possible implementation manner, the adapter includes one or more virtual adapters, and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators; the method further includes: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters. In a possible implementation manner, the method further includes: Obtain fourth configuration information; the fourth configuration information is used to indicate adjusting the number of virtual adapters in the adapter, and / or, is used to indicate adjusting the number of second accelerators corresponding to the target virtual adapter in the adapter; Adjust the number of virtual adapters based on the fourth configuration information, and / or, adjust the number of second accelerators corresponding to the target virtual adapter; the target virtual adapter is any one of the one or more virtual adapters. In a sixth aspect, the present application provides a task processing device, which includes a scheduler and an adapter. The scheduler is coupled to a first die, and the adapter is coupled to a second die. The first die includes M first accelerators, and the second die includes N second accelerators. Among them, The scheduler is configured to: Obtain a first task set from upper-layer software, where the first task set includes a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; The adapter is configured to: Obtain a second task set from the scheduler, where the second task set includes the second task; Schedule one or more target second accelerators among the N second accelerators to process the second task. In a seventh aspect, the present application provides a semiconductor chip, which includes the task processing device or die provided in any possible implementation manner of the first aspect, the second aspect, the third aspect, the sixth aspect, or any one of them. In an eighth aspect, the present application provides a computer-readable storage medium, on which program instructions are stored. When it runs, the methods described in any possible implementation manner of the fourth aspect, the fifth aspect, or any one of them are executed. In a ninth aspect, the present application provides a program product including program instructions. When it runs, the methods described in any possible implementation manner of the fourth aspect, the fifth aspect, or any one of them are executed. In a tenth aspect, the present application provides an electronic device, which includes the task processing device or die provided in any possible implementation manner of the first aspect, the second aspect, the third aspect, the sixth aspect, or any one of them. The electronic device further includes a memory for storing necessary program instructions and data when the task processing device or die runs. The electronic device may further include a communication interface for the electronic device to communicate with other devices or communication networks. In an eleventh aspect, the present application provides an electronic device, which has the function of implementing any one of the processing methods in the fourth aspect, the fifth aspect, or any one of them. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In a twelfth aspect, the present application provides a chip system, which includes a task processing device or die provided by any one of the possible implementations of the first, second, third, sixth aspects or any one of them. In a possible design, the chip system further includes a memory, which is used to store program instructions and data necessary or related to the task processing device or die. The chip system may be composed of chips or may include chips and other discrete devices. The technical effects achieved in the above aspects can be mutually referred to or refer to the beneficial effects in the method embodiments shown below, and will not be elaborated here. Description of the Drawings To more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the drawings required for use in the embodiments of the present application or the background art will be described below. FIG. 1 is a schematic structural diagram of a chiplet. FIG. 2 is a schematic structural diagram of a task processing device provided by an embodiment of the present application. FIG. 3 is a schematic structural diagram of another task processing device provided by an embodiment of the present application. FIG. 4 is a schematic structural diagram of yet another task processing device provided by an embodiment of the present application. FIG. 5 is a schematic flowchart of a processing method provided by an embodiment of the present application. FIG. 6 is a schematic flowchart of a task processing provided by an embodiment of the present application. FIG. 7 is a schematic flowchart of another task processing provided by an embodiment of the present application. FIG. 8 is a schematic diagram of information interaction of a task processing device provided by an embodiment of the present application. Detailed Embodiments The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. In the description of the present application, words such as "first" and "second" are only used to distinguish different objects, and do not limit the quantity and execution order, and the words such as "first" and "second" do not necessarily limit to be different. For example, the first message and the second message are only used to distinguish different information, and do not limit their order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device, etc. that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices, etc. In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a relational expression describing associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one (item)", "one (or more) of the following items (or items)" or similar expressions refer to any combination of these items, including any combination of a single item (or item) or multiple items (or items). For example, at least one (item) of a, b, or c can mean: a, b, c; a and b; a and c; b and c; or a, b, and c. Here, a, b, and c can be single or multiple. In the description of this application, words such as "exemplary", "exemplarily", or "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in this application should not be construed as more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner. It can be understood that in the description of this application, "when...", "if", and "in case" all refer to the situation where the device will perform corresponding processing under certain objective circumstances, not limited to time, and do not require the device to have a judgment action when implemented, nor does it mean there are other limitations. The "simultaneously" in this application can be understood as at the same time point, can also be understood as within a period of time, and can also be understood as within the same cycle. Specifically, it can be understood in combination with the context. In this application, elements represented in the singular are intended to mean "one or more", rather than "one and only one", unless otherwise specified. It can be understood that in the embodiments of this application, "A and B correspond" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A. B can also be determined based on A and / or other information. It can be understood that in the embodiments of the present application, "for indicating" and "indicating" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. When describing "a certain indication information is used to indicate A" or "the indication information of A", it may include that the indication information directly indicates A or indirectly indicates A, and does not mean that A must be carried in the indication information. The information indicated by a certain information is called the information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated. For example, but not limited to, it can directly indicate the information to be indicated, such as the information to be indicated itself or the index of the information to be indicated, etc. It can also indirectly indicate the information to be indicated by indicating other information, where there is an association relationship between the other information and the information to be indicated. It can also only indicate a part of the information to be indicated, while the other parts of the information to be indicated are known or pre-agreed. For example, it can also rely on the arrangement order of each information pre-agreed (such as protocol regulations) to achieve the indication of specific information, thereby reducing the indication overhead to a certain extent. At the same time, it can also identify the common parts of each information and indicate them uniformly to reduce the indication overhead caused by separately indicating the same information. In addition, the specific indication method can also be various existing indication methods, such as but not limited to, the above indication methods and their various combinations, etc. The specific details of various indication methods can refer to the prior art and will not be elaborated herein. As can be seen from the above description, for example, when it is necessary to indicate multiple information of the same type, there may be a situation where the indication methods of different information are different. In the specific implementation process, the required indication method can be selected according to specific needs. The embodiments of the present application do not limit the selected indication method. In this way, the indication methods involved in the embodiments of the present application should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated. The information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending periods and / or sending times of these sub-information can be the same or different. The specific sending method is not limited in the present application. To better understand the technical solutions of the embodiments of the present application, several terms or nouns related to the present application are briefly introduced below for the convenience of those skilled in the art to understand. I. Chiplet Chiplets, also known as small chip sets, can package multiple dies (chips) that meet specific functions together with the underlying basic chip through die-to-die internal interconnection technology. Among them, these multiple dies can be made by different manufacturers using different manufacturing processes. Compared with SoC, chiplets can achieve higher yields and lower costs. When designing chiplets, they can be decomposed according to different computing units or functional units, and then the most suitable manufacturing process can be selected for each unit (die) respectively. These dies with different functions and manufactured by different processes are interconnected with each other through advanced packaging technology and packaged into a chiplet to achieve a new form of IP reuse. This application provides a chiplet structure including a scheduler and an adapter. By adding an adapter to shield the fixed interface between the scheduler and various accelerators deployed on other dies, the scheduler is decoupled from the accelerators on other dies. The scheduler can schedule those accelerators that may not be normally schedulable by the scheduler through the adapter; the scheduler can directly schedule other accelerators that can be normally scheduled by it, so that the accelerator resources on the chiplet product can be as much as possible scheduled to process tasks, thereby ensuring the overall performance of the chiplet. II. Die A die is a small piece of integrated circuit body made of semiconductor material, and the established function of the integrated circuit is realized on this small piece of semiconductor. Usually, integrated circuits are made on large semiconductor wafers in large quantities through multiple steps such as lithography, and then partitioned into square small pieces, and this small piece is called a die. At this time, each die is a replica of an integrated circuit. The semiconductor material used for the wafer is usually single crystal of electronic-grade silicon (EGS) or other semiconductors (such as gallium arsenide, GaAs). Generally, integrated circuit dies will be packaged in packages such as ceramics or plastics and lead pins will be led out. Due to the requirement of circuit miniaturization, sometimes some integrated circuit dies will not be packaged and directly given to downstream users for use. At this time, the die can also be called a bare chip. In this application, the task processing device can be a chiplet product. The first die and the second die in the task processing device can be used as components of the chiplet. Among them, the scheduler and the first accelerator can be deployed on the first die, and the adapter and the second accelerator can be deployed on the second die. The scheduler can schedule the first accelerator to process tasks and schedule the second accelerator to process tasks through the adapter, as much as possible to ensure that the accelerator resources can be scheduled, thereby ensuring the overall performance of the chiplet. III. Hardware Accelerator (HAC) A hardware accelerator, or accelerator for short, is a specialized hardware device / device designed to accelerate computing through hardware processing. Currently, the more common hardware accelerator is the graphics processing unit (GPU), which is a hardware processor specifically used for graphics computing and is used to accelerate complex calculations in graphics processing. In addition, hardware accelerators include CPUs, tensor processing units (TPUs), direct memory access (DMA), input / output (IO), artificial intelligence (AI) accelerators, and some specific domain accelerators (DSA, also known as dedicated accelerators). For chiplet technology, various resources on the chip are often distributed on different dies. For example, some accelerators (CPU, AI, etc.) are distributed on different dies, so there may be dedicated CPU dies, IO dies, AI dies, Memory dies, etc. These dies may be flexibly composed of chip products according to different product specifications. 4. Intellectual Property (IP) IP core is a general term for integrated circuit cores with intellectual property cores. It is a macro module (logic or functional unit) of integrated circuit design that is gradually separated from the chip design link, repeatedly verified, has specific functions, can be reused, and contains specific core elements (instruction set, function description, code, etc.). IP cores are generally divided into soft cores (stronger flexibility), solid cores (stronger reliability) and hard cores (stronger performance). If divided by major categories, they can be roughly divided into processor and microcontroller IP, memory IP, peripheral and interface IP, analog and mixed circuit IP, communication IP, image and media IP, etc. The role of IP core is mainly to reduce the cost of redundant design in the chip design link, reduce the risk of errors, and improve chip design efficiency. The first accelerator and the second accelerator in the embodiment of the present application can be any one of the above IP cores, such as image and media IP (corresponding to GPU accelerator), peripheral and interface IP (corresponding to DMA accelerator) or processor and microcontroller IP (corresponding to CPU accelerator), etc. First, analyze and propose the specific technical problems to be solved by this application. Refer to FIG. 1. FIG. 1 is a schematic structural diagram of a chiplet. The chiplet includes die 1 and die 2. Various resources are distributed on different dies in the chiplet. For example, a scheduler and one or more accelerators are deployed on die 1, and one or more accelerators are deployed on die 2. The scheduler needs to schedule accelerators with various different functions on die 1 and die 2 to process tasks. Since each die in the chiplet can be manufactured by different manufacturers according to different processes and specifications, each die may have been independently evolved before being packaged together. In this case, there may be a situation where multiple different versions of dies in a chiplet are docked with each other, and then there may be a situation where the scheduler cannot normally schedule the accelerators. For example, version 1.0 of die 1 corresponds to (is adapted to) version 1.0 of die 2, and the accelerator on die 2 can be normally scheduled by the scheduler on die 1. A certain company may need to combine and package the original die 1 (version 1.0) with another die 2 (version x.0, the version has evolved / fallen behind) according to its own needs. At this time, the version of die 1 is fixed (the versions of the scheduler and the accelerator on it remain unchanged). Because of the version change of die 2, the version, function, quantity, or interface interaction mechanism of the accelerator may also change. At this time, the scheduler is docking with the accelerator on die 2 whose version has evolved (or fallen behind), which may cause the scheduler on die 1 to be unable to normally schedule the accelerator on die 2, and then lead to a decrease in the overall performance of the chiplet. Therefore, this application proposes a task processing device, related dies, and a processing method. By adding an adapter to shield the fixed interfaces between the scheduler in the chiplet and various accelerators on other dies, the scheduler is decoupled from the accelerators on other dies. When an accelerator needs to be scheduled, for those accelerators that may not be normally scheduled by the scheduler, that is, accelerators that are not adapted to the scheduler, the scheduler can assign the tasks that need to be processed by these accelerators to the adapter, and then the adapter schedules these accelerators to process the corresponding tasks. That is, the scheduler does not directly schedule these accelerators that may not be normally scheduled by it; for other accelerators that can be normally scheduled by the scheduler, that is, accelerators that are adapted to the scheduler, after the scheduler obtains the tasks that need to be processed by these accelerators, it can still directly schedule these accelerators to process the corresponding tasks, so that the accelerator resources of the chiplet product can be scheduled as much as possible to process tasks, thereby ensuring the overall performance of the chiplet, and it can support the independent evolution of resources on various dies in the chiplet. For easy understanding, the technical solutions provided by this application will be described below with reference to more drawings. In this application, unless otherwise specified, the same or similar parts among various embodiments or implementation manners can be referred to each other. In various embodiments of this application, as well as in each implementation manner / implementation method / realization method in each embodiment, if there is no special specification and logical conflict, the terms and / or descriptions among different embodiments, as well as among each implementation manner / implementation method / realization method in each embodiment, are consistent and can be referenced to each other. The technical features in different embodiments, as well as in each implementation manner / implementation method / realization method in each embodiment, can be combined to form new embodiments, implementation manners, implementation methods, or realization methods according to their internal logical relationships. The implementation manners of this application described below do not constitute a limitation on the protection scope of this application. Referring to FIG. 2, FIG. 2 is a schematic structural diagram of a task processing device provided by an embodiment of this application. This task processing device
[0010] can be a chiplet product or a partial structure in a chiplet product. As shown in FIG. 2, this task processing device
[0010] can include a first die
[0100] and a second die
[0101] . The first die
[0100] can include a scheduler
[1000] and M first accelerators [1001 - 100x]. Optionally, the scheduler
[1000] and these M first accelerators [1001 - 100x] can communicate through a first protocol. The second die
[0101] can include an adapter
[1010] and N second accelerators [1011 - 101x]. Optionally, the adapter
[1010] and these N second accelerators [1011 - 101x] can communicate through a second protocol. Optionally, the scheduler
[1000] of the first die
[0100] and the adapter
[1010] of the second die
[0101] can communicate through a third protocol. Both M and N are integers greater than 0. Among them, the scheduler
[1000] can be used to: obtain a first task set from upper-layer software. This first task set can include a first task and a second task; schedule some or all of the above-mentioned M first accelerators [1001 - 100x] (that is, one or more target first accelerators) to process the above-mentioned first task. The adapter
[1010] can be used to: obtain a second task set from the scheduler
[1000] . This second task set includes the above-mentioned second task; schedule some or all of the above-mentioned N second accelerators [1011 - 101x] (that is, one or more target second accelerators) to process the above-mentioned second task. Understandably, the first task set obtained by the scheduler from the upper-layer software (SW) may include a first task and a second task. The above-mentioned first task is processed by the first accelerator in the first die, and the second task is processed by the second accelerator in the second die. Therefore, the scheduler needs to schedule the first accelerator to process the first task and send the second task to the adapter. During the actual operation of the task processing device, the first task set obtained by the scheduler from the upper-layer software may not necessarily include both the first task and the second task. For example, the first task set may include the first task but not the second task. At this time, the tasks obtained by the scheduler can all be processed by the first accelerator on the first die, and the adapter and the second accelerator on the second die may not participate in the processing of the first task. Similarly, the first task set may also include the second task but not the first task. At this time, the tasks obtained by the scheduler are all processed by the second accelerator on the second die. The adapter can first obtain the second task from the scheduler and then schedule the second accelerator to process the second task, and the first accelerator on the first die may not participate in the processing of the second task. In addition, the above-mentioned first task and second task can be two different tasks, which are respectively processed by accelerators on different dies. The above-mentioned first task and second task can also be the same task (or subtasks of the same task), that is, the same task can be completed by accelerators on different dies in cooperation, and no specific limitation is made here. In a possible implementation manner, the scheduler may include a first interface, a second interface, and a third interface, and the adapter may include a fourth interface and a fifth interface. Among them, the scheduler can be connected to M first accelerators through the first interface. The first interface needs to comply with (meet) the first protocol. Since the first interface is docked with the accelerator, it can also be called an accelerator interface. The scheduler can be connected to the fourth interface of the adapter through the second interface. The second interface and the fourth interface need to comply with the third protocol for the interaction between the scheduler and the adapter. The second interface and the fourth interface can be called interconnection interfaces; when the scheduler and the adapter are deployed on different dies, the second interface and the fourth interface can also be called die-to-die interconnection interfaces; the interface types of these two interfaces can be the same or different. The scheduler can obtain the above-mentioned first task set from the upper-layer software through the third interface. The third interface can also be called a software-hardware interface (or a task acquisition interface). The adapter is connected to N second accelerators through the fifth interface. The fifth interface complies with the second protocol, and the fifth interface is also called an accelerator interface. Among them, the first interface and the fifth interface are both accelerator interfaces. They can be of the same type of interface or different types of interface. They can comply with the same protocol or different protocols. Optionally, the first protocol and the second protocol may be the same protocol, such as the Advanced eXtensible Interface (AXI) protocol, the Advanced High-performance Bus (AHB) protocol, the Advanced Peripheral Bus (APB) protocol, the Peripheral Component Interconnect (PCI) protocol, the Peripheral Component Interconnect Express (PCIe) protocol, or the Compute Express Link (CXL) protocol, etc. Alternatively, the first protocol and the second protocol may also be different protocols. Different protocols may refer to two different protocols, such as one being AXI and the other being AHB, or one being PCIe and the other being CXL. Different protocols may also refer to different versions of the same type of protocol, such as one being AXI 4.0 and the other being AXI 4.0-Lite, or one being PCIe 1.0 and the other being PCIe 4.0. Specific limitations are not made here. In addition, the third protocol may be any one of the above protocols, or may be other unlisted protocols, or may be a predefined private protocol. Specific limitations are not made here. If the first protocol is different from the second protocol, it can be understood that the first accelerator and the second accelerator support different interaction protocols. When the scheduler can normally schedule the first accelerator, there may be a situation where the second accelerator cannot be normally scheduled because the second protocol supported by the second accelerator may not be compatible with the first protocol supported by the first accelerator. Additionally, even if the first accelerator and the second accelerator support the same interaction protocol (i.e., the first protocol is the same as the second protocol), there may still be a situation where the scheduler cannot normally schedule the second accelerator. For example, when the first accelerator and the second accelerator need to obtain different types or quantities of information to be scheduled, and the information sent by the scheduler during the scheduling of the second accelerator is still the same as the type or quantity of information for the first accelerator, the second accelerator may not be normally scheduled. Exemplarily, for the case where the first protocol is different from the second protocol and the scheduler cannot normally schedule the second accelerator, the scheduler can send the tasks that need to be processed by the second accelerator to the adapter, and then the adapter schedules the second accelerator to process these tasks. Since the second accelerator supports the second protocol and the adapter and the second accelerator can interact through the second protocol, it can be ensured that the second accelerator can be normally scheduled, thereby ensuring the overall performance of the task processing device. Exemplarily, for the case where the first protocol is the same as the second protocol and the scheduler cannot normally schedule the second accelerator, the scheduler can send the tasks that need to be processed by the second accelerator to the adapter, and the adapter can analyze the scheduling requirements of the scheduler and send various types of information required during scheduling to the second accelerator, thereby scheduling the second accelerator. In a possible implementation, the above-mentioned first protocol is different from the second protocol. The first protocol can be the AXI protocol, and the second protocol can be the AHB protocol. Optionally, the specific protocols of the first protocol and the second protocol can be determined according to the protocol requirements of the first accelerator and the second accelerator respectively. For example, the first accelerator is a processor (such as a CPU or GPU), and its requirements for the protocol are high performance, high bandwidth, low latency, etc. The AXI protocol can meet the requirements of the first accelerator, and the first protocol can be the AXI protocol. At this time, the first accelerator can communicate with the scheduler through the AXI protocol; the second accelerator is a DMA accelerator or an AI accelerator, and its requirements for the protocol are high frequency and high efficiency, etc. The AHB protocol can meet the requirements of the second accelerator, and the second protocol can be the AHB protocol. At this time, the second accelerator can communicate with the adapter through the AHB protocol. In addition, for some low-speed and low-power accelerators, the protocol used can be the APB protocol,... and no more examples will be given here. It should be noted that an adapter can be understood as a special type of scheduler. The scheduler can be understood as the main scheduler, and the adapter can be understood as the slave scheduler. The adapter (slave scheduler) can have the same task scheduling capabilities as the scheduler (main scheduler). In different designs, the roles of the main scheduler and the slave scheduler can be swapped. That is to say, the scheduler can be designed as the main scheduler or as the slave scheduler (referred to as an adapter at this time); the above-mentioned adapter can be designed as the slave scheduler or as the main scheduler (referred to as a scheduler at this time). For example, in the structural design shown in Figure 2 above, the adapter
[1010] as the slave scheduler can obtain tasks not directly from the upper-layer software, but can obtain tasks from the upper-layer software uniformly by the scheduler
[1000] as the main scheduler, and the adapter
[1010] then obtains tasks from the scheduler
[1000] for scheduling. Optionally, in some possible designs, the adapter of the second die can also be used as the main scheduler, obtain tasks from the upper-layer software, and send some of the tasks to the scheduler (as the slave scheduler at this time) of the first die. Exemplarily, the protocol conversion performed by the adapter between the scheduler and the accelerator can include providing level conversion, signal conversion, and other functions between the scheduler and the accelerator (such as the second accelerator) on different dies, so that the scheduler and the accelerator on different dies can communicate and cooperate, thereby improving the overall compatibility and performance of the task processing device. In other words, the adapter design focuses more on compatibility and interconnectivity, enabling the task processing device to connect the scheduler and the accelerator deployed on different dies through the adapter while allowing normal exchange of data, signals, and control information between them. Therefore, the design of the adapter generally needs to give priority to electrical characteristics, signal characteristics, and protocol characteristics, etc., to ensure the stability and reliability of the connection. Optionally, the above-mentioned M first hardware accelerators are main accelerators, and the above-mentioned N second hardware accelerators are auxiliary accelerators. Exemplarily, the M first accelerators on the first die can be used as main accelerators to meet the requirements in the main application scenarios of the chiplet product and are more likely to be scheduled to process tasks; the N second accelerators on the second die are used as auxiliary accelerators to meet the other functional requirements related to the main application scenarios. For example, assuming that the task processing device in FIG. 2 is mainly applied to the AI scenario, AI accelerators can be deployed on the first die as main accelerators, while some other accelerators (such as DMA and DSA accelerators) can be deployed on the second die as auxiliary accelerators; or, assuming that the task processing device in FIG. 2 is mainly applied to the graphics processing scenario, GPUs can be deployed on the first die as main accelerators, while some other accelerators (such as DMA, DSA accelerators, and AI accelerators) can be deployed on the second die as auxiliary accelerators. In other words, what kind of accelerators the main accelerator and the auxiliary accelerator in the task processing device are and what the specific uses of the accelerators are mainly depend on what functions need to be met in the scenario targeted by the chiplet product. The main accelerator is used to meet the main functional requirements in this scenario, and the auxiliary accelerator can be used to meet the other auxiliary functional requirements related to this scenario. Further optionally, the scheduler and the main accelerator can be deployed together (both are deployed on the first die, as shown in FIG. 2); or the scheduler and the main accelerator are not deployed on the same die, but the scheduler can directly schedule the main accelerator (as shown in FIG. 3 below). In this way, when the chiplet product completes the tasks in the main scenario, the scheduling path of the scheduler for the main accelerator can be reduced, prompting the main accelerator to respond to tasks faster, thereby further improving the overall performance of the chiplet. Optionally, the main accelerator can be deployed on the same die as the scheduler or on the same die as the adapter, and the main accelerator is scheduled by the scheduler or the adapter. For example, when the area of the first die is limited and the number of main accelerators on the first die is limited, and it is necessary to further increase the number of main accelerators to meet more requirements, one way is to update and replace the first die with a first die that includes more main accelerators, and another way is to deploy some additional main accelerators while deploying the auxiliary accelerators on the second die, and these accelerators can be scheduled by the adapter. Understandably, if the scheduler and the main accelerators deployed on the second die meet the condition requirements of cross-die interconnection, the scheduler can also be interconnected with these main accelerators and can directly schedule these main accelerators without relying on the adapter to schedule them. The above Figure 2 only shows a structural example in which the task processing device includes a first die and a second die. In a possible implementation manner, the task processing device provided by the embodiments of the present application may also include multiple first dies and multiple second dies. Among them, each first die may respectively include a scheduler and M first accelerators. The scheduler and the M first accelerators may communicate with each other through a first protocol. The number of first accelerators on different first dies may be the same or different, and the first protocols followed between the scheduler and the first accelerators on different first dies may be the same or different. Each second die may respectively include an adapter and N second accelerators. The adapter and the N second accelerators may communicate with each other through a second protocol. The number of second accelerators on different second dies may be the same or different, and the second protocols followed between the adapter and the second accelerators on different second dies may be the same or different. The scheduler of the first die and the adapter of the second die may communicate with each other through a third protocol. The third protocols followed by different schedulers and adapters may be the same or different; both M and N are integers greater than 0. Optionally, the multiple first dies may be isomorphic or heterogeneous, and may be homogeneous or heterogeneous; further optionally, the multiple second dies may be isomorphic or heterogeneous, and may be homogeneous or heterogeneous; further optionally, the first die and the second die may be isomorphic or heterogeneous, and may be homogeneous or heterogeneous. For example, the above task processing device includes 2 first dies and 3 second dies. These 2 first dies may be prepared using the same process (corresponding to being isomorphic, such as both using 14nm, 7nm, 5nm, etc.), or may be prepared using different processes respectively (corresponding to being heterogeneous, such as one using 14nm and the other using 7nm); further, these 2 first dies may be prepared using the same material (or material quality) (corresponding to being homogeneous, such as both using silicon Si, gallium nitride GaN, indium phosphide InP, etc.), or may be prepared using different materials (corresponding to being heterogeneous, such as one using Si and the other using GaN). Similarly, the 3 second dies may be isomorphic or heterogeneous, and may be homogeneous or heterogeneous; and the first die and the second die may be isomorphic or heterogeneous, and may be homogeneous or heterogeneous, which are not specifically limited herein. The above structure in which the scheduler and the first accelerator are deployed on the same die, and the adapter and the second accelerator are deployed on the same die is taken as an example to briefly describe the task processing device provided by the present application, and it should not constitute a limitation on the structure of the task processing device provided by the present application. For ease of understanding, the following will take the structure in which the scheduler and the first accelerator are separately deployed, and the adapter and the second accelerator are separately deployed as an example to describe the task processing device provided by the present application. Referring to FIG. 3, FIG. 3 is a schematic structural diagram of another task processing device provided by an embodiment of the present application. The task processing device
[0020] may include a scheduler
[0201] and an adapter
[0202] . The scheduler
[0201] is coupled to the first die
[0021] , and the adapter
[0202] is coupled to the second die
[0022] . The first die
[0021] may include M first accelerators [211-21x], and the second die
[0022] may include N second accelerators [221-22x]. Both M and N are integers greater than 0. Optionally, the scheduler
[0201] and the M first accelerators [211-21x] on the first die
[0021] may communicate through a first protocol; the adapter
[0202] and the N second accelerators [221-22x] on the second die may communicate through a second protocol; and the scheduler
[0201] and the adapter
[0202] may communicate through a third protocol. Among them, the scheduler
[0201] may be used to: obtain a first task set from the upper-layer software. The first task set may include a first task and a second task; then, based on the first protocol, schedule some or all of the above-mentioned M first accelerators [211-21x] (i.e., one or more target first accelerators) to process the above-mentioned first task. The adapter
[0202] may be used to: obtain a second task set from the scheduler
[0201] based on the third protocol. The second task set includes the above-mentioned second task; then, based on the second protocol, schedule some or all of the above-mentioned N second accelerators [221-22x] (i.e., one or more target second accelerators) to process the above-mentioned second task. Optionally, in the task processing device shown in FIG. 3 above, the scheduler and the adapter may be deployed on the same die or on different dies, which is not specifically limited herein. It should be noted that the task processing device shown in FIG. 3 may also include the above-mentioned first die and second die, which will not be elaborated herein. In a possible implementation, for the task processing device shown in FIGS. 2 and 3 above, the adapter may include abstracted virtual resources (slots, i.e., one or more virtual adapters). As shown in FIG. 4, the adapter
[3010] may include one or more virtual adapters (such as slot1, slot2, and slot3, etc.). Each of these one or more virtual adapters may correspond to one or more accelerators among N second accelerators. For example, slot1 corresponds to the second accelerator
[3011] and the second accelerator
[3012] , slot2 corresponds to the second accelerator
[3013] , and slot3 corresponds to the second accelerator [301x]. When the scheduler
[3000] performs task scheduling through the adapter
[3010] , it can first select an idle virtual adapter for the task, and then the virtual adapter schedules the task to the corresponding one or more accelerators for processing. In other words, a virtual adapter can implement the functions of a physical adapter for scheduling and managing tasks, support scheduling tasks to real physical accelerators. The scheduler can flexibly call the second accelerators on the second die through the virtual adapters in the adapter, and can configure these accelerators to execute one or more subtasks. Optionally, the scheduler may also include abstracted virtual resources (i.e., one or more virtual schedulers, not shown in FIG. 4). Each of these one or more virtual schedulers may correspond to one or more accelerators among M first accelerators. The scheduler can schedule the M first accelerators through these one or more virtual schedulers to perform task processing. Above, through virtual scheduling resources (including virtual schedulers and virtual adapters), the task processing device can more conveniently schedule and manage the tasks to be processed and the accelerators, making the scheduling more flexible. Above, the structure of the task processing device provided by the embodiments of the present application has been exemplarily described. Next, the processing method provided by the embodiments of the present application will be described. For ease of understanding, the processing method provided by the embodiments of the present application will be briefly described below based on the task processing device shown in FIG. 2 above in combination with the task processing process. It can be understood that the processing method provided by the embodiments of the present application can also be applied to the task processing device shown in FIG. 3 above, and various other types of task processing devices obtained by deformation based on the task processing devices shown in FIGS. 2 and 3 above, which will not be specifically limited here. As shown in FIG. 5, FIG. 5 is a schematic flowchart of a processing method provided by an embodiment of the present application. Taking the application of this method to the task processing device shown in FIG. 2 or FIG. 4 above as an example, it includes the following steps S500 - S503: S500: The scheduler obtains a first task set, and the first task set includes a first task and a second task. Among them, the scheduler can obtain the first task and the second task from the upper-layer software. Exemplarily, the upper-layer software can be an operating system, or other task management software or application management software. The first task and the second task can be various requests initiated by various applications. For example, the first task can be a computing task that needs to be processed by the accelerator CPU, or a transfer task that needs to be processed by the accelerator DMA, or other tasks that need to be processed by the dedicated accelerator DSA. No specific limitation is made here. Optionally, after obtaining the first task and the second task, the scheduler can first determine the accelerators required to process the first task and the second task, and then schedule these accelerators to process the first task and the second task. Exemplarily, the scheduler can determine the accelerators required to process the task based on the task information and configuration information of the task to be processed. Among them, the task information can be obtained by the scheduler when obtaining the task. The task information can indicate information such as the task type and task identifier (ID) of the task to be processed, and can further indicate information such as the security level requirement of the task, the storage address of the data required when processing the task, and the storage address of the data obtained after the task is executed. Simply put, when finally giving the task information to the accelerator, it is necessary to let the accelerator know where the input data required to process the task is stored, what operations to perform on these input data (such as addition and subtraction operations), and where the output data obtained after the operation is stored, and so on. The configuration information can be pre-configured by the upper-layer software and sent to the scheduler. Specifically, it can be sent before the scheduler obtains the task, or can be sent when the scheduler obtains the task, or can be sent within a certain period of time after the scheduler obtains the task. No specific limitation is made here. The configuration information can indicate the configuration relationship between various types of tasks and various types of accelerators. At this time, the configuration information can be called full-scale configuration information, which can generally be obtained by the scheduler before obtaining the task; or, the configuration information can also only indicate the configuration relationship between a certain task and various types of accelerators. At this time, the configuration information can be called partial configuration information, which can generally be obtained by the scheduler when obtaining the task or within a period of time after obtaining the task. Exemplarily, taking the case where the scheduler determines one or more target first accelerators to process the first task as an example, the scheduler may obtain first configuration information, where the first configuration information is used to indicate a first configuration relationship between the first task and one or more target first accelerators; the scheduler may then determine one or more target first accelerators based on the first task information of the first task and the first configuration information, and the one or more target first accelerators are part or all of the above-mentioned M first accelerators. Among them, the first configuration relationship indicated by the first configuration information may include one or more of the correspondence between the task type of the first task and the accelerator type of the target first accelerator, the number of target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task; further, the first configuration information may also indicate the execution priority of the first task. If the priority is high, it can be processed preferentially. For example, if the first task is a computing task, the target first accelerator corresponding to the first task should be a CPU or a type of accelerator similar to the CPU for computing; the number of target first accelerators required for the first task is 2 (it can also be other numbers, not limited); the security level requirement of the first task is high, such as the required security level is 2 (the lower the value, the higher the level), and the corresponding target first accelerator should be able to meet the security level requirement of the first task, that is, the security level of the accelerator should also be 2 or a higher level. Assume that the number of first accelerators available for computing among the above-mentioned M first accelerators is 6 (serial numbers 1-6), where the level of 1 / 2 accelerators is 1, the level of 3 / 4 accelerators is 2, and the level of 5 / 6 accelerators is 3. The scheduler can select 2 idle accelerators from accelerators numbered 1-4 to process the first task, or can preferentially select the 3 / 4 accelerators with a security level of 2 to process the first task. In summary, the scheduler can determine the accelerator type, accelerator number, and accelerator security level required to process a task based on the configuration information and the task information of a certain task, so as to determine which accelerators can be scheduled to process the task according to the current status of many accelerators. S501: The scheduler schedules one or more target first accelerators to process the first task. Among them, after determining one or more target first accelerators from the above-mentioned M first accelerators, the scheduler schedules them based on the first protocol followed by these accelerators to process the first task, or in other words, schedules the first task to these accelerators for processing. Optionally, the scheduler may include abstracted virtual resources (i.e., one or more virtual schedulers), and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators. Therefore, the scheduler can schedule the above-mentioned one or more target first accelerators through these one or more virtual schedulers to process the first task. Optionally, when one virtual scheduler corresponds to multiple first accelerators, when the virtual scheduler schedules a task to an accelerator for processing, it can simultaneously schedule the task to these multiple first accelerators for processing, or only schedule the task to some of these multiple first accelerators for processing, or it can also split the task into multiple subtasks corresponding to the number of accelerators and schedule the subtasks to these multiple first accelerators for processing respectively, which is not specifically limited herein. In addition, after the first task is processed, the above-mentioned one or more target first accelerators can output the processing result to the memory or register of the task processing device. Optionally, for the case where one or more virtual schedulers including abstractions are included in the scheduler, the number of virtual schedulers in the scheduler and the number of first accelerators corresponding to each virtual scheduler can be expanded by means of configuration issued by the upper-layer software. Exemplarily, the scheduler can first obtain third configuration information from the upper-layer software, where the third configuration information is used to indicate an adjustment to the number of virtual schedulers in the scheduler, and / or is used to indicate an adjustment to the number of first accelerators corresponding to a target virtual scheduler in the scheduler; the scheduler then adjusts the number of virtual schedulers based on the third configuration information, and / or adjusts the number of first accelerators corresponding to the target virtual scheduler; the target virtual scheduler is any one of the one or more virtual schedulers. For example, when the number of virtual schedulers needs to be adjusted, the number of virtual schedulers and the memory address segments that each virtual scheduler can have can be configured in the third configuration information. For example, the address space is 0-1G, the number of virtual schedulers is 4 (sequence numbers 1-4), virtual scheduler 1 corresponds to 0-250M, virtual scheduler 2 corresponds to 251-500M, …, to adjust the number of virtual schedulers in the scheduler in this way. For another example, when the number of first accelerators corresponding to a virtual scheduler needs to be adjusted, the identifier (ID) of the target virtual scheduler and the identifiers (IDs) of multiple first accelerators corresponding to the target virtual scheduler can be configured in the third configuration information to establish a mapping relationship between the target virtual scheduler and the multiple first accelerators. For example, virtual scheduler 1 corresponds to first accelerators 1, 2, and 3. When there are a large number of tasks to be processed in the scheduler and multiple accelerators need to be scheduled to process multiple tasks in parallel, the number of virtual schedulers can be increased, and then multiple virtual schedulers can be used to schedule accelerators respectively to process different tasks in parallel, so as to quickly reduce the amount of tasks in the task queue of the scheduler and improve the performance of the task processing device; for some large tasks to be processed in the scheduler that need to occupy multiple accelerators to process together, the number of accelerators corresponding to a certain virtual scheduler can be increased, and this virtual scheduler can be used to schedule multiple accelerators to process this large task at the same time, accelerating the task processing efficiency and improving the performance of the virtual device. S502: The scheduler sends the second task set to the adapter; the second task set includes the second task. Among them, when the scheduler determines that the accelerator that can be scheduled to process the second task is the accelerator among the above N second accelerators, the scheduler can send the second task to the adapter through the third protocol (it can be sent in the form of a task set, such as the second task set). Correspondingly, the adapter obtains the second task set from the scheduler. Optionally, when the scheduler dispatches the second task to the adapter, it can also indicate to the adapter which second accelerator or accelerators to specifically schedule to process the second task. That is to say, after obtaining the second task, when the scheduler determines that the second task requires processing by a second accelerator, it can directly determine one or more target second accelerators from the above N second accelerators for processing the second task, and instruct the adapter to schedule these one or more target second accelerators to process the second task. For example, the scheduler can indicate to the adapter the one or more target second accelerators that specifically need to be scheduled through the accelerator ID. Optionally, the accelerator used to process the second task may not be indicated by the scheduler, but the adapter can, after obtaining the second task, independently determine from the above N second accelerators the accelerators required to process the second task, and then schedule these accelerators to process the task. Exemplarily, the manner in which the adapter determines the accelerator can be similar to the manner in which the scheduler determines the accelerator. For example, the adapter can determine the accelerator required to process the task based on the task information and configuration information of the task to be processed. Among them, the task information can be sent by the scheduler to the adapter together when dispatching the task. The task information can indicate information such as the task type and task identifier (ID) of the task to be processed, and can also indicate the security level requirements of the task, the storage address of the data required when processing the task, and the storage address of the data obtained after the task is executed. The configuration information can be the full configuration information or partial configuration information obtained by the scheduler from the upper-layer software and then forwarded to the adapter, or can be partial configuration information separately generated by the scheduler for the second task and the full configuration information and sent to the adapter. For example, the scheduler selects the configuration information related to the second task from the full configuration information, obtains the partial configuration information and sends it to the adapter. The configuration information can be sent before the adapter obtains the task, or can be sent when the adapter obtains the task, or can be sent within a certain period of time after the scheduler obtains the task. There is no specific limitation here. Exemplarily, the adapter can determine one or more target second accelerators from the above N second accelerators based on the second task information and the second configuration information of the second task. The relevant description can refer to the description of the scheduler determining one or more target first accelerators above, and will not be elaborated here. It should be noted that when the first configuration information and the second configuration information are both full configuration information, they can be the same information. S503: The adapter schedules one or more target second accelerators to process the second task. Among them, after determining one or more target second accelerators from the above N second accelerators, the adapter schedules them to process the second task based on the second protocol followed by these accelerators. Optionally, the adapter may include abstracted virtual resources (i.e., one or more virtual adapters), and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators. Therefore, the adapter can schedule the above one or more target second accelerators to process the second task through these one or more virtual adapters. In addition, after the second task is processed, the above one or more target second accelerators can output the processing result to the memory or register of the task processing device. Optionally, for the case where the adapter includes one or more abstracted virtual adapters, the number of virtual adapters in the adapter and the number of second accelerators corresponding to each virtual adapter can be extended by means of configuration issued by the upper-layer software. Exemplarily, the scheduler first obtains fourth configuration information from the upper-layer software and sends the fourth configuration information to the adapter; the fourth configuration information is used to indicate an adjustment to the number of virtual adapters in the adapter, and / or is used to indicate an adjustment to the number of second accelerators corresponding to the target virtual adapter in the adapter. Accordingly, the adapter receives the above fourth configuration information and adjusts the number of virtual adapters based on this information, and / or adjusts the number of second accelerators corresponding to the target virtual adapter. The configuration method of the fourth configuration information can refer to the relevant description of the above third configuration information and will not be elaborated here. For ease of understanding, the following associates the structure of the task processing device shown in FIG. 4 with the processing flow involved in the above processing method shown in FIG. 5. As shown in FIG. 6, the processing flow may include but is not limited to the following steps: 1. The scheduler
[3000] first obtains a first task set from the upper-layer software (SW), and the first task set includes a first task and a second task. 2. The scheduler
[3000] schedules the first accelerator [300x] to process the first task. 3. The scheduler
[3000] issues a second task set to the adapter
[3010] , and the second task set includes a second task. 4. The adapter
[3010] schedules the second accelerator
[3012] to process the second task. Optionally, the above processing flow may further include steps for the scheduler
[3000] and the adapter
[3010] to obtain configuration information (such as the first configuration information and the second configuration information). In a possible implementation, after the scheduler and the adapter schedule the accelerator to process a task, they can also obtain the situation of the accelerator processing the task. As shown in FIG. 7, after the scheduler
[3000] schedules one or more target first accelerators [300x] to process the first task (i.e., step S501), the scheduler
[3000] can also receive the task execution information (i.e., the first execution information) fed back by the one or more target first accelerators [300x]. The first execution information is used to indicate the execution situation of the first task. The execution situation of the task may include execution errors (such as interruptions) and execution completion. Optionally, when the first task has an execution error, the scheduler
[3000] can send a reset command to the first accelerator [300x] where the task execution error occurs, to indicate that the first accelerator [300x] is reset to avoid the task from not being completed. Or, when the first task is executed successfully, the scheduler
[3000] can send a release command to the first accelerator [300x] where the task is executed successfully, to release the first accelerator [300x], which facilitates the subsequent scheduler
[3000] to schedule it to process other tasks. Further, the above reset command and release command can be determined by the scheduler based on the task execution information fed back by the accelerator, or the scheduler can send the obtained task execution information to the upper-layer software, and the upper-layer software makes a decision and then instructs the scheduler to send a reset command or a release command, which is not specifically limited here. Understandably, the task execution information can be a part of the data fields that are agreed upon and mutually adapted by the scheduler, the adapter, and the upper-layer software. The task execution information should include information about the accelerator (such as the accelerator ID) and task information about the task executed by the accelerator (such as the task ID). In a possible implementation, as shown in FIG. 7, after the adapter
[3010] schedules one or more target second accelerators
[3012] to process the second task (i.e., step S503), the adapter
[3010] may further receive task execution information (i.e., second execution information) fed back by the one or more target second accelerators
[3012] , and the second execution information is used to indicate the execution status of the second task. Similar to the scheduler
[3000] , for different task execution statuses, the adapter
[3010] may send different commands. For example, when the second task execution encounters an error, the adapter
[3010] may send a reset command to the second accelerator
[3012] where the task execution error occurs; or, when the second task execution is completed, the adapter
[3010] may send a release command to the second accelerator
[3012] where the task execution is completed. Among them, the above reset command and release command may be determined whether to be sent by the adapter based on the task execution information fed back by the accelerator; or the adapter forwards the task execution information to the scheduler, and the scheduler determines whether to send it based on the task execution information; or, it may also be that the scheduler sends the obtained task execution information to the upper-layer software, and the upper-layer software makes a decision to determine whether to send it, which is not specifically limited herein. It should be noted that when the above task processing device includes multiple first dies and multiple second dies, each first die includes a scheduler and M first accelerators respectively, and each second die includes an adapter and N second accelerators respectively, there may be a situation where the schedulers on different first dies need to send different tasks to the same second accelerator on the same second die for processing. Since an accelerator can only be occupied by one task at the same time, then the different tasks sent by multiple schedulers can queue in the queue and wait to be processed in turn when the accelerator is idle. In summary, different from the chiplet structure including a scheduler in the prior art, the present application provides a task processing device structure including a scheduler and an adapter. By adding an adapter, the fixed interfaces between the scheduler and various accelerators deployed on other dies are shielded, and the scheduler is decoupled from the accelerators on other dies, so that independent evolution between different dies in the chiplet product can be supported. After each die evolves independently, the scheduler can schedule through the adapter accelerators that may not be schedulable by the scheduler normally (accelerators with advanced or backward versions); the scheduler can directly schedule other accelerators that can be scheduled by it normally, so that all accelerator resources of the chiplet product can be scheduled as much as possible to process tasks, thereby ensuring the overall performance of the chiplet. Taking the task processing device shown in FIG. 4 above as an example, if the second accelerator [3011-301x] (IP core) on the second die
[0301] is a hard core purchased externally and the chip designer cannot modify the IP core independently. As shown in FIG. 8, before the second die is updated, the initial interaction mode of the IP core (the second accelerator
[3012] ) is to send or receive messages through the bus. After the second die is updated, the IP core (the second accelerator
[3012] ) may change to report task execution information (such as the second execution information) through an interrupt line. Then, the adapter
[3010] can convert the interrupt into information recognizable by the scheduler
[3000] and send it to the scheduler
[3000] . At this time, if there is no adapter
[3010] to decouple the scheduler
[3000] from the IP core (the second accelerator
[3012] ), the scheduler
[3000] may not recognize the task execution information, resulting in the scheduler
[3000] being unable to schedule the updated IP core (the second accelerator
[3012] ) normally. The present application also provides a semiconductor chip, which includes the task processing device provided in all the above embodiments of the present application. It can be understood that the functions and roles of each part in the task processing device can be correspondingly referred to the specific implementation manners in the above embodiments of FIGS. 2 to 4, and will not be elaborated here. The present application also provides an electronic device, which includes the task processing device provided in all the above embodiments of the present application. It can be understood that the functions and roles of each part in the task processing device can be correspondingly referred to the specific implementation manners in the above embodiments of FIGS. 2 to 4, and will not be elaborated here. Optionally, the electronic device may further include a communication interface for the electronic device to communicate with other devices or communication networks. The present application also provides an electronic device, which has the function of implementing the processing method of any one of the above task processing devices. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The present application provides a computer storage medium storing a computer program, which, when executed, enables the above task processing device to perform the functions involved in the above processing method flow. The present application provides a computer program comprising instructions, which, when executed, enables the above task processing device to perform the functions involved in the above processing method flow. The present application provides a chip system comprising any one of the above task processing devices. In a possible design, the chip system further comprises a memory for storing program instructions and data necessary or relevant to the task processing device. The chip system may be composed of chips or may comprise chips and other discrete devices. In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In several embodiments provided by the present application, the couplings, direct couplings or communication connections shown or discussed among each other may be indirect couplings or communication connections through some interfaces, devices or units, or may be electrical, mechanical or other forms of connection. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A task processing device, characterized in that, The device includes a first die and a second die; the first die includes a scheduler and M first accelerators, and the second die includes an adapter and N second accelerators; wherein, The scheduler is configured to: Obtain a first task set from the upper-layer software, where the first task set includes a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; The adapter is configured to: Obtain a second task set from the scheduler, where the second task set includes the second task; Schedule one or more target second accelerators among the N second accelerators to process the second task.
2. The device according to claim 1, characterized in that, The scheduler communicates with the M first accelerators through a first protocol, the adapter communicates with the N second accelerators through a second protocol, and the scheduler communicates with the adapter through a third protocol.
3. The device according to claim 2, characterized in that, The first protocol is different from the second protocol. The first protocol includes the Advanced eXtensible Interface (AXI) protocol, and the second protocol includes the Advanced High-Performance Bus (AHB) protocol.
4. The device according to any one of claims 2-3, characterized in that The scheduler includes a first interface, a second interface, and a third interface, and the adapter includes a fourth interface and a fifth interface; the scheduler is connected to the M first accelerators through the first interface, and the first interface complies with the first protocol; the scheduler is connected to the fourth interface of the adapter through the second interface, and the second interface and the fourth interface comply with the third protocol; the scheduler obtains the first task set from the upper-layer software through the third interface; The adapter is connected to the N second accelerators through the fifth interface, and the fifth interface complies with the second protocol.
5. The device according to any one of claims 1-4, characterized in that, The scheduler is further configured to: Obtain first configuration information; the first configuration information is used to indicate a first configuration relationship between the first task and the one or more target first accelerators; the first configuration relationship includes one or more of the correspondence between the task type of the first task and the type of the target first accelerator, the number of the target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task; Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information; the first task information includes one or more of the task type or task identifier of the first task.
6. The device according to any one of claims 1-5, characterized in that, The adapter is further configured to: Obtain second configuration information; the second configuration information is used to indicate a second configuration relationship between the second task and the one or more target second accelerators; the second configuration relationship includes one or more of the correspondence between the task type of the second task and the type of the target second accelerator, the number of the target second accelerators corresponding to the second task, or the security level of the target second accelerators corresponding to the second task; Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task.
7. The device according to any one of claims 1-6, characterized in that, The scheduler is further configured to: Obtain first execution information; the first execution information is used to indicate the execution status of the first task; When the execution status of the first task includes an error in task execution, send a first reset command; the first reset command is used to indicate resetting the accelerator that makes an error in executing the first task; or, When the execution status of the first task includes completion of task execution, send a first release command; the first release command is used to indicate releasing the accelerator that completes the execution of the first task.
8. The device according to any one of claims 1-7, characterized in that, The adapter is further configured to: Obtain second execution information; the second execution information is used to indicate the execution status of the second task; When the execution status of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that makes an error in executing the second task; or, When the execution status of the second task includes completion of task execution, send a second release command; the second release command is used to indicate releasing the accelerator that completes the execution of the second task.
9. The device according to any one of claims 1-8, characterized in that, The M first accelerators are main accelerators, and the N second accelerators are auxiliary accelerators.
10. The device according to any one of claims 1-9, characterized in that, The scheduler includes one or more virtual schedulers, and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators; specifically, the scheduler is configured to: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers.
11. The device according to claim 10, wherein The scheduler is further configured to: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate adjusting the number of virtual schedulers in the scheduler and / or to indicate adjusting the number of first accelerators corresponding to the target virtual schedulers in the scheduler; Adjust the number of virtual schedulers and / or adjust the number of first accelerators corresponding to the target virtual schedulers based on the third configuration information; the target virtual scheduler is any one of the one or more virtual schedulers.
12. The device according to any one of claims 1-11, characterized in that, The adapter includes one or more virtual adapters, and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators; specifically, the adapter is configured to: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters.
13. The device according to claim 12, characterized in that, The scheduler is further configured to: Obtain fourth configuration information from the upper-layer software and send the fourth configuration information to the adapter; the fourth configuration information is used to indicate adjusting the number of the virtual adapters in the adapter, and / or, is used to indicate adjusting the number of the second accelerators corresponding to the target virtual adapter in the adapter; The adapter is further configured to: Receive the fourth configuration information and adjust the number of the virtual adapters based on the fourth configuration information, and / or, adjust the number of the second accelerators corresponding to the target virtual adapter; The target virtual adapter is any one of the one or more virtual adapters.
14. The device according to any one of claims 1 to 13, characterized in that, The first die and the second die are isomorphic or heteromorphic, and / or, the first die and the second die are homogeneous or heterogeneous.
15. A first crystal grain, characterized in that, The first die includes a scheduler and M first accelerators; wherein, The scheduler is configured to: Obtain a first task set from the upper-layer software, the first task set including a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; Send a second task set to the adapter of the second die; the second task set includes the second task.
16. The crystal grain as described in claim 15, characterized in that, The scheduler communicates with the M first accelerators through a first protocol.
17. The crystal grain according to claim 16, wherein, The scheduler includes a first interface, a second interface, and a third interface; the scheduler is connected to the M first accelerators through the first interface, and the first interface complies with the first protocol; the scheduler is connected to a fourth interface of the adapter through the second interface, and the second interface and the fourth interface comply with a third protocol; the scheduler obtains the first task set from the upper-layer software through the third interface.
18. The crystal grains according to any one of claims 15-17, characterized in that, The scheduler is further configured to: Obtain first configuration information; the first configuration information is used to indicate a first configuration relationship between the first task and the one or more target first accelerators; the first configuration relationship includes one or more of a correspondence between the task type of the first task and the type of the target first accelerator, the number of the target first accelerators corresponding to the first task, or the security level of the target first accelerators corresponding to the first task; Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information; the first task information includes one or more of the task type or task identifier of the first task.
19. The crystal grains according to any one of claims 15-18, characterized in that, The scheduler is further configured to: Obtain first execution information; the first execution information is used to indicate the execution situation of the first task; When the execution situation of the first task includes a task execution error, send a first reset command; the first reset command is used to indicate resetting the accelerator that makes an error in executing the first task; or, When the execution situation of the first task includes a task execution completion, send a first release command; the first release command is used to indicate releasing the accelerator that completes the execution of the first task.
20. The crystal grains according to any one of claims 15-19, characterized in that, The M first accelerators are main accelerators.
21. The crystal grains according to any one of claims 15-20, characterized in that, The scheduler includes one or more virtual schedulers, and one of the one or more virtual schedulers corresponds to one or more of the M first accelerators; specifically, the scheduler is configured to: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers.
22. The crystal grain according to claim 21, characterized in that, The scheduler is further configured to: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate adjusting the number of virtual schedulers in the scheduler, and / or, is used to indicate adjusting the number of first accelerators corresponding to the target virtual schedulers in the scheduler; Adjust the number of virtual schedulers based on the third configuration information, and / or, adjust the number of first accelerators corresponding to the target virtual schedulers; the target virtual scheduler is any one of the one or more virtual schedulers.
23. A second grain, characterized in that, The second die includes an adapter and N second accelerators; wherein, The adapter is configured to: Obtain a second task set from the scheduler of the first die, and the second task set includes second tasks; Schedule one or more target second accelerators among the N second accelerators to process the second tasks.
24. The crystal grain according to claim 15, wherein, Communication between the adapter and the N second accelerators is performed through a second protocol.
25. The crystal grain according to claim 24, wherein The adapter includes a fourth interface and a fifth interface, the adapter is connected to the N second accelerators through the fifth interface, and the fifth interface complies with the second protocol; The adapter is connected to the second interface of the scheduler through the fourth interface, and the second interface and the fourth interface comply with a third protocol.
26. The crystal grains according to any one of claims 23-25, characterized in that, The adapter is further configured to: Obtain second configuration information; the second configuration information is used to indicate a second configuration relationship between the second tasks and the one or more target second accelerators; the second configuration relationship includes one or more of a correspondence between the task type of the second task and the type of the target second accelerator, the number of target second accelerators corresponding to the second task, or the security level of the target second accelerator corresponding to the second task; Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task.
27. The crystal grains according to any one of claims 23-26, characterized in that, The adapter is further configured to: Obtain second execution information; the second execution information is used to indicate the execution status of the second task; When the execution status of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that has an error in executing the second task; or, When the execution status of the second task includes completion of task execution, send a second release command; the second release command is used to indicate releasing the accelerator that has completed the execution of the second task.
28. The crystal grains according to any one of claims 23-27, characterized in that, The N second accelerators are auxiliary accelerators.
29. The grain according to any one of claims 23-28, characterized in that, The adapter includes one or more virtual adapters, and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators; specifically, the adapter is configured to: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters.
30. The crystal grain according to claim 29, wherein The adapter is further configured to: Obtain fourth configuration information; the fourth configuration information is used to indicate an adjustment to the number of virtual adapters in the adapter, and / or is used to indicate an adjustment to the number of second accelerators corresponding to the target virtual adapters in the adapter; Adjust the number of virtual adapters based on the fourth configuration information, and / or adjust the number of second accelerators corresponding to the target virtual adapters; The target virtual adapter is any one of the one or more virtual adapters.
31. A processing method, characterized in that, Applied to a scheduler in a first die, the first die further includes M first accelerators; the method includes: Obtain a first task set from upper-layer software, where the first task set includes a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; Send the second task set to the adapter of the second die; the second task set includes the second task.
32. The method according to claim 31, wherein The method further includes: Obtain first configuration information; the first configuration information is used to indicate a configuration relationship between the first task and the one or more target first accelerators; the first configuration relationship includes one or more of a correspondence between the task type of the first task and the type of the target first accelerator, the number of target first accelerators corresponding to the first task, or the security level of the target first accelerator corresponding to the first task; Determine the one or more target first accelerators based on the first task information of the first task and the first configuration information; the first task information includes one or more of the task type or task identifier of the first task.
33. The method according to claim 31 or 32, characterized in that, The method further includes: Obtain first execution information; the first execution information is used to indicate the execution situation of the first task; When the execution situation of the first task includes an error in task execution, send a first reset command; the first reset command is used to indicate a reset of the accelerator that makes an error in executing the first task; or, When the execution situation of the first task includes completion of task execution, send a first release command; the first release command is used to indicate a release of the accelerator that completes the execution of the first task.
34. The method according to any one of claims 31-33, characterized in that, The scheduler includes one or more virtual schedulers, and one virtual scheduler among the one or more virtual schedulers corresponds to one or more of the M first accelerators; the method further includes: Determine one or more target virtual schedulers from the one or more virtual schedulers; the one or more target virtual schedulers correspond to the one or more target first accelerators; the one or more target virtual schedulers are in an idle state; Schedule the one or more target first accelerators to process the first task through the one or more target virtual schedulers.
35. The method according to claim 30, wherein The method further includes: Obtain third configuration information from the upper-layer software; the third configuration information is used to indicate adjusting the number of virtual schedulers in the scheduler, and / or is used to indicate adjusting the number of first accelerators corresponding to the target virtual schedulers in the scheduler; Adjust the number of virtual schedulers based on the third configuration information, and / or adjust the number of first accelerators corresponding to the target virtual schedulers; the target virtual scheduler is any one of the one or more virtual schedulers.
36. A processing method, characterized in that, Applied to an adapter in a second die, the second die further includes N second accelerators; the method includes: Obtain a second task set from a scheduler of a first die, the second task set including second tasks; Schedule one or more target second accelerators among the N second accelerators to process the second tasks.
37. The method according to claim 36, wherein The method further includes: Obtain second configuration information; the second configuration information is used to indicate a configuration relationship between the second task and the one or more target second accelerators; the second configuration relationship includes one or more of a correspondence between the task type of the second task and the type of the target second accelerator, the number of target second accelerators corresponding to the second task, or the security level of the target second accelerator corresponding to the second task; Determine the one or more target second accelerators based on the second task information of the second task and the second configuration information; the second task information includes one or more of the task type or task identifier of the second task.
38. The method according to claim 36 or 37, characterized in that The method further includes: Obtain second execution information; the second execution information is used to indicate the execution situation of the second task; When the execution situation of the second task includes an error in task execution, send a second reset command; the second reset command is used to indicate resetting the accelerator that makes an error in executing the second task; or, When the execution situation of the second task includes the completion of task execution, send a second release command; the second release command is used to indicate releasing the accelerator that completes the execution of the second task.
39. The method according to any one of claims 36 - 38, characterized in that, The adapter includes one or more virtual adapters, and one virtual adapter among the one or more virtual adapters corresponds to one or more of the N second accelerators; the method further includes: Determine one or more target virtual adapters from the one or more virtual adapters; the one or more target virtual adapters correspond to the one or more target second accelerators; the one or more target virtual adapters are in an idle state; Schedule the one or more target second accelerators to process the second task through the one or more target virtual adapters.
40. The method according to claim 39, characterized in that, The method further includes: Obtain fourth configuration information; the fourth configuration information is used to indicate an adjustment to the number of virtual adapters in the adapter, and / or is used to indicate an adjustment to the number of second accelerators corresponding to a target virtual adapter in the adapter; Adjust the number of virtual adapters and / or adjust the number of second accelerators corresponding to the target virtual adapter based on the fourth configuration information; the target virtual adapter is any one of the one or more virtual adapters.
41. A task processing device, characterized in that, The apparatus includes a scheduler and an adapter. The scheduler is coupled to a first die, and the adapter is coupled to a second die; the first die includes M first accelerators, and the second die includes N second accelerators; wherein, The scheduler is configured to: Obtain a first task set from upper-layer software, the first task set including a first task and a second task; Schedule one or more target first accelerators among the M first accelerators to process the first task; The adapter is configured to: Obtain a second task set from the scheduler, the second task set including the second task; Schedule one or more target second accelerators among the N second accelerators to process the second task.
42. A computer-readable storage medium, characterized in that, A computer program or instruction is stored in the storage medium, and when the computer program or instruction is executed, the method according to any one of claims 31-40 is implemented.
43. A computer program, characterized in that, The computer program includes instructions, and when the computer program is executed, the method according to any one of claims 31-40 is implemented.
44. An electronic device, characterized in that, Including the apparatus according to any one of claims 1-14 or 41, or the die according to any one of claims 15-30, or the apparatus for implementing the method according to any one of claims 31-40.
Citation Information
Patent Citations
Task processing device, related crystal grain and processing method
CN120256078A
Chip
CN114070657A
Method for scheduling hardware accelerator and task scheduler
CN114981776A
Hierarchical task scheduling for accelerators
US20220188155A1
Operation acceleration method and operation accelerator
WO2022261928A1