Data production entity unit, data consumption entity unit, related devices and methods
By introducing a communication control unit and a synchronization instruction execution unit into the data production main unit, the successful writing and use of batch data between multiple target bodies is ensured, and the problem of low data call speed between processing units is solved, and data transmission efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202110813521.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-07-19
AI Technical Summary
In large-scale data processing tasks, data call between processing units is slow, especially when batch data is written, it is difficult to ensure the successful writing and use of data between multiple target bodies.
A data production main unit is designed, and a write request is sent to the peer main unit and the scheduling main unit through the first and second communication control units, and a feedback tracking instruction is generated using the synchronization instruction execution unit to ensure that the write request is successfully completed and the set instruction is sent to the target body to allow data use.
It is realized that in the case of multiple target bodies, data writing is ensured to be successful and allowed to be used correctly, and the efficiency and accuracy of data transmission between processing units is improved.
Smart Images

Figure CN115640142B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of chips, and more particularly, to a data production subject unit, a data consumption subject unit, related devices and methods. Background Art
[0002] In the big data era, due to the large-scale processing tasks, it is usually required that processing units (including CPUs for performing conventional arithmetic processing and GPUs, NPUs, tensor processors, etc. for complex arithmetic processing in graphics processing and deep neural networks) cooperate to process a large-scale task. Each processing unit processes a part of the task and stores the task-related data. Since the parts processed by these processing units may be interrelated, a processing unit often finds that it needs the data stored by other processing units when executing its own part of the task. At this time, other processing units are required to write their data to this processing unit. In the traditional technology, the host-to-host traditional communication network between each host (each computing device) is required to complete this writing. The speed of the host-to-host traditional communication network is very slow, and the data also needs to be transmitted via an internal bus such as PCIe inside the host. The data call speed between processing units is low. To solve this problem, an inter-chip interconnect network (ICN) is proposed. The inter-chip interconnect network directly establishes an interconnection between the processing units of each host (computing device). Through the ICN, data calls between processing units can be quickly realized.
[0003] If the processing unit is an acceleration unit such as an NPU or a GPU, when it executes the write of batch data, the write target body is not necessarily one. For example, some of the batch data need to be written to other peer acceleration units, some need to be written to the scheduling unit (such as a CPU) that assigns tasks to it, and even some need to be written to its own on-chip memory. There may also be multiple peer acceleration units, and there may be data in the batch data that need to be written to each acceleration unit. When the batch data starts to be read before being completely written to their respective target bodies, errors are likely to occur because these data are often interrelated. When using them before they are completely written, the data used is likely to be incorrect. However, it is difficult to monitor that the data to each target body is completely written.
[0004] The above-mentioned difficulties are more obvious during instruction-level writing. During instruction-level writing, writing is usually performed by thread groups (warps). Generally, acceleration units perform repetitive operations on batch data. For example, in the multiplication of matrices, the instructions for the operation rules of multiplying and accumulating elements between matrices are fixed, and only the numerical values of elements in different parts of the matrix need to be substituted into the instructions for repeated operations. Therefore, operations are performed by thread groups. A thread group includes multiple threads. Within a thread group, each thread executes the same code operation on different data it is responsible for. Their executed operations are the same, but only the data they operate on is different. A large computing task command can be decomposed into numerous such thread group operations. When each thread group successfully completes the operation, the entire computing task command is executed. In this kind of instruction-level writing, if each thread in a thread group successfully writes the data it is responsible for to the target body where the data is to be written, it does not mean that each thread in other thread groups has completed a similar task. Therefore, on the target body side, the written data still cannot be used because it may still be incorrect data. Therefore, when there are multiple target bodies for the batch data writing to be performed by a data production main body unit, how to ensure that the writing to multiple target bodies is successfully completed during instruction-level writing, so that the data usage instructions of the target body can use the correct data, becomes a difficult problem. Summary of the Invention
[0005] In view of this, an object of the present disclosure is to enable correct instruction-level writing when there are multiple target bodies for the batch data writing to be performed by a data production main body unit, and to ensure that the use of the written data by the target body does not go wrong.
[0006] According to an aspect of the present disclosure, there is provided a data production main body unit for writing data to a target body by thread groups, where the target body includes at least one of a peer main body unit and a scheduling main body unit, and the data production main body unit includes:
[0007] At least one of a first communication control unit and a second communication control unit, where the first communication control unit is used to send a write request to the peer main body unit, and the second communication control unit is used to send a write request to the scheduling main body unit;
[0008] A synchronous instruction execution unit is used to execute for a thread group: generate at least one of a first feedback trace instruction and a second feedback trace instruction based on a synchronous instruction, and send them to at least one of the first communication control unit and the second communication control unit respectively. Wherein, if the first communication control unit receives the first feedback trace instruction, it counts the first responses to the write requests sent to the peer entity unit. If the number of received first responses is equal to the number of write requests sent to the peer entity unit, the first trace is successful; if the second communication control unit receives the second feedback trace instruction, it sends a first false read instruction to the scheduling entity unit. If it receives a second response from the scheduling entity unit to the first false read instruction, the second trace is successful; if both the first trace and the second trace are successful, a set instruction is sent to the peer entity unit and the scheduling entity unit, so that after receiving the set instruction, the data consumption entity unit allows the use of the written data on the premise of ensuring that all write requests of the data production entity unit received before the set instruction are successfully written.
[0009] Optionally, the data production entity unit further includes: a first thread group execution unit, which is used to, after sending write requests to at least one of the first communication control unit and the second communication control unit for a thread group, receive internal confirmation responses from at least one of the first communication control unit, the second communication control unit, or a cache that caches the write requests in the first multi-level cache. The first multi-level cache is a multi-level cache located between the first thread group execution unit and the first communication control unit or the second communication control unit. The first information indicates whether internal confirmation responses to all write requests have been received. Wherein, after all thread groups receive internal confirmation responses to all write requests, the synchronous instruction execution unit starts to work.
[0010] Optionally, if the write request sent by the thread group is cached in the first multi-level cache, the first multi-level cache performs at least one of a first flush and a second flush based on the synchronization instruction, where the first flush is for the data cached to be sent to the peer main unit, which is flushed out from the first multi-level cache and sent from the first communication control unit to the peer main unit via a higher-level cache, and the third response from the peer main unit is received; the second flush is for the data cached to be sent to the scheduling main unit, which is flushed out from the first multi-level cache and sent from the second communication control unit to the scheduling main unit via a higher-level cache, and a second false read instruction is sent to the scheduling main unit, and the fourth response to the second false read instruction is received from the scheduling main unit; wherein the setting instruction is issued to the peer main unit and the scheduling main unit only when both the first trace and the second trace are successful, the third response is received for all the data cached to be sent to the peer main unit that has been flushed out, and the fourth response is received for the second false read instruction.
[0011] Optionally, the first communication control unit counts the first responses to the write requests sent to the peer main unit during the execution of the thread block for each thread block. If the number of first responses received for the thread block is equal to the number of write requests sent by the thread block to the peer main unit, the first trace is successful. The thread block includes multiple thread groups.
[0012] Optionally, the first communication control unit sets a corresponding counter for a single thread block. The initial value of the counter is 0. Each time a write request is sent to the peer main unit during the execution of the thread block, the counter is incremented by 1. After all the write requests corresponding to the thread block are sent, the counter value reaches the maximum value. Then, each time a first response to a write request is received, the counter is decremented by 1 until the counter value is decremented to 0, and then the first trace is successful.
[0013] Optionally, the synchronization instruction execution unit generates a first feedback trace instruction only when the internal confirmation response from the first communication control unit is received; and generates a second feedback trace instruction only when the internal confirmation response from the second communication control unit is received.
[0014] Optionally, after at least one of the first communication control unit and the second communication control unit respectively executes at least one of the first feedback tracking instruction and the second feedback tracking instruction, a wait feedback tracking instruction is executed to ensure that the set instruction can be issued only after at least one of the first feedback tracking instruction and the second feedback tracking instruction is executed; after the first multi-level cache executes at least one of the first flush and the second flush, a wait flush instruction is executed to ensure that the set instruction can be issued only after at least one of the first flush and the second flush is executed.
[0015] Optionally, after sending a set instruction to the peer entity unit and the scheduling entity unit, the first thread group execution unit further causes the thread group to exchange third information with other thread groups, and the third information indicates that the set instruction has been sent to the peer entity unit and the scheduling entity unit.
[0016] Optionally, the target entity further includes the on-chip memory of the data production entity unit. After the first thread group execution unit transfers a write request for the thread group to the on-chip memory to the on-chip memory, an internal acknowledgment response of the on-chip memory is received, and the set instruction is issued only when both the first tracking and the second tracking are successful and the internal acknowledgment response of the on-chip memory is received for the write request to the on-chip memory.
[0017] Optionally, the data production entity unit is a first processing unit, the data consumption entity unit is a second processing unit, the first processing unit and the second processing unit communicate through an inter-chip interconnect network, and the first processing unit writes the data in the first on-chip memory of the first processing unit into the second on-chip memory of the second processing unit through the inter-chip interconnect network.
[0018] Optionally, the first processing unit is a first acceleration unit for executing tasks assigned by a first scheduling unit; the second processing unit is a second acceleration unit for executing tasks assigned by a second scheduling unit.
[0019] According to one aspect of the present disclosure, a data consumption entity unit is provided, including:
[0020] A third communication control unit, configured to perform, for a thread group: after receiving a set instruction from a data production main body unit, determine whether all the previously received write requests of the data production main body unit have been successfully written, and in the case where it is determined that all have been successfully written, change the specific flag bit to a specific value, and execute a set waiting instruction and a data usage instruction, where the set waiting instruction is used to allow the data usage instruction to be executed after the specific flag bit becomes the specific value, where the set instruction is issued by the data production main body unit in the case where both a first trace of a write request to a peer main body unit and a second trace of a write request to a scheduling main body unit are successful in the data production main body unit, the first trace includes counting first responses to write requests sent to the peer main body unit and determining that the number of received first responses is equal to the number of write requests sent to the peer main body unit; the second trace includes, after determining that a write request sent to the scheduling main body unit has been successfully sent, sending a first false read instruction to the scheduling main body unit and receiving a second response from the scheduling main body unit to the first false read instruction;
[0021] A second thread group execution unit, configured to enable the thread group to exchange second information with other thread groups, where the second information indicates that the data usage instruction is allowed to be executed.
[0022] Optionally, the data consumption main body unit further includes: a second multi-level cache, located between the second thread group execution unit and the third communication control unit, after the specific flag bit becomes the specific value, the second multi-level cache executes a cache invalidation instruction and a wait-for-invalidation-success instruction, the cache invalidation instruction is used to prevent data stored in the second multi-level cache from being read by subsequent read instructions, and the wait-for-invalidation-success instruction is used to wait for the invalidation instruction to be executed and, in the case where the invalidation instruction has been executed, only then allow the data usage instruction to be executed.
[0023] Optionally, determining whether all the previously received write requests of the data production main body unit have been successfully written includes: if the received write request sequence of all ports of the data consumption main body unit is empty, determining that all the previously received write requests of the data production main body unit have been successfully written.
[0024] According to one aspect of the present disclosure, there is provided a system on chip, including the data production main body unit and / or the data consumption main body unit as described above.
[0025] According to one aspect of the present disclosure, there is provided a computing device, including the data production main body unit and / or the data consumption main body unit as described above.
[0026] According to one aspect of the present disclosure, there is provided a data writing method at the data production main body unit for writing data to a target body by thread groups, where the target body includes at least one of a peer main body unit and a scheduling main body unit, and the data production main body unit includes at least one of a first communication control unit and a second communication control unit. The first communication control unit is used to send a write request to the peer main body unit, and the second communication control unit is used to send a write request to the scheduling main body unit. The method includes:
[0027] For a thread group, at least one of a first feedback tracking instruction and a second feedback tracking instruction is generated based on a synchronization instruction and sent to at least one of the first communication control unit and the second communication control unit. Wherein, if the first communication control unit receives the first feedback tracking instruction, the first responses to the write requests sent to the peer main body unit are counted. If the number of received first responses is equal to the number of write requests sent to the peer main body unit, the first tracking is successful; if the second communication control unit receives the second feedback tracking instruction, a first false read instruction is sent to the scheduling main body unit. If the second response to the first false read instruction from the scheduling main body unit is received, the second tracking is successful;
[0028] If both the first tracking and the second tracking are successful, a setting instruction is sent to the peer main body unit and the scheduling main body unit, so that after receiving the setting instruction, the data consumption main body unit allows the use of the written data on the premise of ensuring that all the write requests of the data production main body unit received before the setting instruction are successfully written.
[0029] According to one aspect of the present disclosure, there is provided a data using method at the data consumption main body unit, including:
[0030] For a thread group, after receiving a set instruction from a data production main unit, determine whether all the previous write requests from the data production main unit have been successfully written, and if it is determined that all have been successfully written, change the specific flag bit to a specific value, and execute a set waiting instruction and a data usage instruction. The set waiting instruction is used to allow the data usage instruction to be executed after the specific flag bit becomes the specific value. Wherein, the set instruction is issued by the data production main unit when it is determined that both the first tracking of the write request to the peer main unit and the second tracking of the write request to the scheduling main unit are successful. The first tracking includes counting the first responses to the write requests sent to the peer main unit and determining that the number of received first responses is equal to the number of write requests sent to the peer main unit; the second tracking includes sending a first false read instruction to the scheduling main unit after determining that the write request sent to the scheduling main unit has been successfully sent, and receiving a second response from the scheduling main unit to the first false read instruction.
[0031] Cause the thread group to exchange second information with other thread groups, where the second information indicates that the data usage instruction is allowed to be executed.
[0032] The target bodies for batch data writing to be executed by a data production main body unit may include peer main body units and scheduling main body units. When the target body is a peer main body unit, regardless of the route through which the write request sent by the data production main body unit reaches the peer main body unit, the peer main body unit will send back an acknowledgment. The data production main body unit can monitor whether all the write requests to the peer main body unit in this task are received by counting whether the number of write requests sent to the peer main body unit matches the number of acknowledgments returned. When the target body is a scheduling main body unit, the scheduling main body unit will not give an acknowledgment for the received write request. However, a dummy read instruction can be sent to the scheduling main body unit after all the write requests. The scheduling main body unit will give an acknowledgment for this dummy read instruction. Moreover, once this acknowledgment is received, it is assumed that all the write requests sent before this dummy read instruction have been received. Therefore, for the scheduling main body unit, the write requests to the scheduling main body unit in this task can be determined whether all are received in such a way. Therefore, in the embodiments of the present disclosure, for a thread group, a first feedback tracking instruction and a second feedback tracking instruction will be generated based on the synchronization instruction, which are respectively used to determine whether all the write requests to the peer main body unit in this task are received and whether all the write requests to the scheduling main body unit in this task are received in the above-mentioned manner. Once both are received, a set instruction is sent to the peer main body unit and the scheduling main body unit. When the peer main body unit and the scheduling main body unit receive the set instruction, and determine that all the received write requests have been successfully written, the specific flag bit is changed to a specific value, and a set waiting instruction and a data usage instruction are executed. The set waiting instruction is used to allow the data usage instruction to be executed after the specific flag bit becomes the specific value. In this way, when there are multiple target bodies for batch data writing to be executed by a data production main body unit, the data usage of these multiple target bodies will always occur after the batch data has been successfully written to all the target bodies, ensuring that the use of the written data will not go wrong. In addition, considering that the embodiments of the present disclosure are for instruction-level writing and are performed by thread groups, in order to improve efficiency, one thread group is allowed to execute the above process, and after one thread group executes successfully, all thread groups share this information, avoiding the huge overhead caused by each thread group executing the above process separately. Therefore, the embodiments of the present disclosure enable correct instruction-level writing to be performed when there are multiple target bodies for batch data writing to be executed by a data production main body unit, and ensure that the use of the written data by the target bodies will not go wrong. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Through the description of the embodiments of the present disclosure with reference to the following drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0034] Figure 1It is a structural diagram of a data center to which an embodiment of the present disclosure is applied;
[0035] Figure 2 It is an internal structural diagram of a server (computing device) according to an embodiment of the present disclosure;
[0036] Figure 3 It shows the Figure 2 internal structure and external communication diagram of the acceleration unit as the data production main body unit or data consumption main body unit in the embodiment of the present disclosure.
[0037] Figure 4 It shows the Figure 3 internal hierarchical structure diagram of the data production main body unit, peer main body unit, and scheduling main body unit according to an embodiment of the present disclosure.
[0038] Figure 5 It shows the internal structural diagram of the first communication control unit according to an embodiment of the present disclosure.
[0039] Figure 6 It shows the internal structural diagram of the second communication control unit according to an embodiment of the present disclosure.
[0040] Figure 7 It shows an example of compiling source code into assembly language on the data production main body unit and data consumption main body unit according to an embodiment of the present disclosure.
[0041] Figure 8 It shows the flowchart of the data writing method at the data production main body unit end according to an embodiment of the present disclosure.
[0042] Figure 9 It shows the flowchart of the data usage method at the data consumption main body unit end according to an embodiment of the present disclosure. Detailed implementation manners
[0043] The following describes the present disclosure based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific detail parts are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these detail parts. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes, and procedures are not described in detail. Additionally, the drawings are not necessarily drawn to scale.
[0044] The following terms are used in this article.
[0045] Data production entity unit: The party that generates / sends the computing data. Specifically, it writes the data stored in itself to other entity units (data consumption entity units) for the use of other entity units. This entity unit can be an actual physical unit or an abstract software module, i.e., a software entity unit. When the entity unit is a physical unit, it can be a processing unit in a chip, such as a general-purpose central processing unit, or an acceleration unit for special-purpose computing (such as NPU, GPU, tensor processor, etc.). However, this entity unit can also be a smaller unit or a larger unit. A smaller unit is, for example, a microcontroller unit (MCU), which is equivalent to a scaled-down version of a CPU plus some peripheral circuits integrated on a single chip. A larger unit is, for example, a system-on-chip integrating multiple processing units, i.e., integrating multiple processing units plus peripheral circuits on one chip. An even larger unit can be the entire computing device, or even a communication terminal. In the case of a communication terminal, the data production entity unit is the sending terminal, and the data consumption entity unit is the receiving terminal. When the entity unit is a software module, it can be a software entity running on hardware or a functional part in a software entity. For the sake of simplicity, the following mainly takes the entity unit as an acceleration unit as an example. For the cases of other entity units, those skilled in the art can make analogies based on the teachings of this disclosure.
[0046] Target entity: The party that uses / receives the computing data. Specifically, it needs to use the data of other entity units (data production entity units), so it requires other entity units to write data to it. It can be a part of the data production entity unit itself, such as the on-chip memory of the data production entity unit. It can also be other entity units outside the data production entity unit. When the target entity is other entity units, it can also take various forms similar to those of the data production entity unit above. The following mainly takes the entity unit as an acceleration unit as an example. For the cases of other entity units, those skilled in the art can make analogies.
[0047] Computing device: A device with computing or processing capabilities, which can be embodied in the form of a terminal, such as an Internet of Things device, a mobile terminal, a desktop computer, a laptop computer, etc., or can be embodied in the form of a server or a cluster of servers. In the context of a data center, the computing device is the server in the data center.
[0048] System-on-chip: A unit with processing capabilities and peripheral circuits packaged into a chip, which can be inserted into a computing device or replaced from a computing device.
[0049] Scheduling unit or scheduling main unit: A unit that, in addition to performing traditional processing (not processing for complex operations such as image processing and various deep learning models) in a computing device or system-on-chip, also undertakes the scheduling function for the acceleration unit. It allocates tasks that the acceleration unit needs to undertake to the acceleration unit. The scheduling unit can take various forms such as a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.
[0050] Acceleration unit: A processing unit designed to improve the data processing speed in some specialized application fields (such as processing images, processing various operations of deep learning models, etc.) where traditional processing units are not efficient. The acceleration unit includes a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and a dedicated intelligent acceleration hardware (such as a neural network processor NPU).
[0051] Peer main unit: A main unit that is of the same type and similar nature as the data production main unit from which the written data comes. If the data production main unit is an acceleration unit and its nature is an acceleration unit for neural network computing (such as an NPU, a tensor processor, etc.), then other NPUs, tensor processors, etc. that also perform neural network computing can be the peer main unit of this data production main unit because they often need to cooperate to complete a task and may therefore need to call data generated by other main units during the execution process.
[0052] On-chip memory: A storage unit located inside the processing unit. Compared with off-chip memory, it is used to directly access frequently used data inside the processing unit to improve the access efficiency of frequently used data. Off-chip memory is outside the processing unit and is shared by each processing unit, but it needs to go through an external bus to be called, with relatively low efficiency, and is generally used to store infrequently used data.
[0053] Processing unit: A unit with processing capabilities, including units for performing traditional processing (not processing for complex operations such as image processing and various deep learning models), the above-mentioned scheduling unit with a scheduling function, and the above-mentioned acceleration unit, etc.
[0054] On-chip interconnect network: A communication network that interconnects the various units inside the processing unit, such as the internal bus network that interconnects the cores, on-chip caches, etc. inside the acceleration unit in the acceleration unit.
[0055] Inter-chip Interconnection Network: Different from on-chip interconnection network, it is a network that interconnects different processing units. A processing unit connected to it can directly transfer data to different processing units within the same computing device or processing units within different computing devices through it.
[0056] Command: A set of programming instructions issued when a computer program executes a certain task. To complete a certain task, it requires the cooperation of a series of instructions, so a command contains multiple instructions.
[0057] Instruction: An instruction is the smallest unit for a command to run, generally referring to a statement in computer programming.
[0058] Process: An execution activity of a program in a computer on a certain data set. It is the basic unit for the system to allocate and schedule resources and is the basis of the operating system structure. A program is a description of instructions, data, and their organizational forms, and a process is the entity of a program.
[0059] Thread: It is the smallest unit that the operating system can perform operation scheduling on. It is contained within a process and is the actual operating unit within the process. A thread refers to a single sequential control flow within a process. Multiple threads can be concurrent within a process, and each thread executes different tasks in parallel.
[0060] Thread Group (warp): A collection of multiple threads decomposed from a process. These multiple threads execute the same program code, but only the data they operate on is different. This is particularly common in computing units for large-scale data operations such as neural network computing or graphics computing (such as NPUs, GPUs, etc.). For example, in the multiplication of matrices, the code for the operation rules of multiplying and accumulating elements between matrices is fixed, and only the numerical values of elements in different parts of the matrix need to be substituted into the instruction code for repeated operations. Therefore, different parts of the elements are assigned to different threads, but the instruction codes executed by each thread in the thread group are the same. In this way, the entire process can be completed quickly.
[0061] Thread Block: Composed of multiple thread groups, and multiple thread groups execute in parallel within the thread block. A process is first divided into several thread blocks, and each thread block is divided into several thread groups.
[0062] Accelerator Unit Hierarchical Processing Architecture: Since large-scale data operations usually divide a process into thread blocks downward, and thread blocks are further divided into thread groups for processing. Correspondingly, the hardware instruction part for executing it also often needs to be set up as a hierarchical processing architecture, such as Figure 4As shown in the figure, the instruction execution units in the architecture are hierarchically divided into processing execution units 342 (PU, hereinafter referred to as thread group execution units), computing execution units 341 (CU, hereinafter referred to as synchronization instruction execution units), and computing execution elements 340 (CE, i.e., acceleration unit cores) from low to high.
[0063] Thread group execution unit 342: As Figure 4 shown, the lowest-level instruction execution unit in the acceleration unit hierarchical processing architecture is the processing execution unit (PU). It executes processing at the level of the above-mentioned thread group (warp). It can run one thread group, but it can also sequentially and serially run multiple thread groups. However, the same thread group must be executed in one PU and cannot run across PUs. Because of this feature and to avoid confusion with the concept of the aforementioned processing unit, the PU is named the thread group execution unit in the embodiments of the present disclosure. A thread group generally includes 16 threads.
[0064] Synchronization: When the data production main unit executes the writing of batch data, the target bodies for writing different data in the batch data may be different. For example, some of the batch data are to be written to other peer acceleration units, some are to be written to the scheduling unit (such as the CPU) that assigns tasks to it, and even some are to be written to its own on-chip memory. There may also be multiple peer acceleration units. Different data may need to be written to different peer acceleration units. Although these data are written to different target bodies, they are related. After the writing to multiple target bodies is completed, the written data can be used without errors. The mechanism for ensuring that the writing of batch data to different target bodies is completed so that no error occurs during use is called synchronization.
[0065] Synchronization instruction execution unit 341: As Figure 4 shown, the layer above the thread group execution unit 342 in the acceleration unit hierarchical processing architecture is the synchronization instruction execution unit 341, i.e., CU. It is the main operation module for synchronization instructions in the embodiments of the present disclosure. As described above, synchronization needs to ensure that the writing of batch data to different target bodies is completed. However, at the sending end, only the write requests for batch data to different target bodies can be ensured to be sent from the data production main unit to the target bodies. As for the storage space within the target body after being sent to the target body, it needs to be ensured at the receiving end. The synchronization instruction execution unit 341 is the unit for ensuring that the write requests for batch data to different target bodies are sent from the data production main unit to the target bodies. The process of its guarantee will be described in detail in the specific implementation part later. The synchronization instruction execution unit 341 includes multiple thread group execution units 342 and a first-level cache 343 shared by the thread group execution units 342. Figure 4Among them, a synchronization instruction execution unit 341 includes 4 thread group execution units 342, but may also include other numbers of thread group execution units 342. The data cached in the first-level cache 343 is shared by all the thread group execution units 342 within the synchronization instruction execution unit 341. The data written from one thread group execution unit 342 to the outside can be first cached in the first-level cache 343. Then, once other thread group execution units 342 in the same synchronization instruction execution unit 341 need this data, it can be directly retrieved from this cache, without having to go through multiple levels of caches to retrieve it from the target body where the data was originally written.
[0066] Acceleration unit core 340: As Figure 4 shown, in the acceleration unit hierarchical processing architecture, the layer above the synchronization instruction execution unit 341 is the acceleration unit core 340, i.e., the CE. Generally, there are 16 acceleration unit cores 340 in an acceleration unit 230. An acceleration unit core 340 includes multiple (generally 4, but can also be other numbers) synchronization instruction execution units 341 and a second-level cache 344 shared by multiple synchronization instruction execution units 341. The data written from one synchronization instruction execution unit 341 to the outside can be first cached in the second-level cache 344. Then, once other synchronization instruction execution units 341 in the same acceleration unit core 340 need this data, it can be directly retrieved from this cache, without having to go through multiple levels of caches to retrieve it from the target body where the data was originally written.
[0067] Routing: Data may pass through different intermediate nodes when traveling between two nodes. At each pair of adjacent nodes, different sending and receiving ports may be selected for transmission due to each node having multiple ports. Therefore, different paths are formed for data transmission between two nodes, and each path is a route.
[0068] Feedback tracking instruction: An instruction used to determine whether all the write requests sent to the target body have been received by the target body based on the response of the target body.
[0069] False read: The real purpose is not to read the required data, but to perform a read for testing or other purposes. The data read from the target body during a false read may be data at any storage location in the target body.
[0070] Set instruction: An instruction that sets the flag bit (sets the flag bit to a specific value). In the embodiments of the present disclosure, it cooperates with the set wait instruction of the target body to jointly ensure that the use of data in the target body occurs after the data production main body unit has completed writing the corresponding data in the batch data to each target body, thereby ensuring that there are no errors in using the data.
[0071] Multi-level cache: The caches at all levels in the acceleration unit hierarchical processing architecture, such as Figure 4The first-level cache 343, the second-level cache 344, the last-level cache 345, etc. shown in the figure.
[0072] Internal confirmation response: It is an internal mechanism of the data production main body unit. It refers to the confirmation response sent by an internal component when a write request to the target body circulates within the data production main body unit and reaches a certain internal component. For example, it is stipulated in advance that when a write request sent from the first communication control unit of the data production main body unit to the peer main body unit is sent by the first communication control unit, the first communication control unit returns an internal confirmation response; if the write request is cached by a certain level of cache when passing through multiple levels of caches within the data production main body unit and is no longer sent down, the cache returns an internal confirmation response.
[0073] Flushing: The process of moving the data cached in the cache out of the cache and transferring it to the target body where the data actually goes, so that the cache is finally empty.
[0074] Flushing instruction: An instruction used to perform flushing.
[0075] Wait-for-feedback tracking instruction: An instruction used to wait for the above-mentioned feedback tracking instruction to be executed, and subsequent instructions can only be executed after that.
[0076] Wait-for-flushing instruction: An instruction used to wait for the above-mentioned flushing instruction to be executed, and subsequent instructions can only be executed after that.
[0077] Set-wait instruction: An instruction used to check if the setting is successful (a specific flag bit becomes a specific value), and after the setting is successful, subsequent instructions are allowed to be executed.
[0078] Cache invalidation instruction: An instruction used to make the data stored in the cache not be read by subsequent read instructions.
[0079] Wait-for-invalidation-success instruction: An instruction used to wait for the above-mentioned cache invalidation instruction to be executed, and only when the cache invalidation instruction is executed, subsequent instructions are allowed to be executed.
[0080] Data consumption main body unit: That is, the target body where the data production main body unit writes data.
[0081] Application environment of the present disclosure
[0082] An embodiment of the present disclosure proposes a communication control solution for a data production entity unit and a data consumption entity unit. The entire communication control solution is relatively general and can be used in various scenarios. For example, data communication control between processing units in each server in a data center, data communication control between each MCU in a microcontroller unit (MCU) architecture, data communication control between processing units in an Internet of Things terminal, data communication control between each node in the Internet, data communication control between different software modules within the same hardware, data communication control between different functional parts within the same software module, and so on. Due to the large number of application scenarios, in the following description, only the data center is taken as an example for description. However, those skilled in the art should understand that this solution is equally applicable to other application scenarios and can easily conceive the implementation details in other scenarios based on the description of the data center scenario in the embodiments of the present disclosure.
[0083] Data center
[0084] A data center is a specific network of devices for global collaboration, used to transfer, accelerate, display, calculate, and store data information on the Internet network infrastructure. In future development, the data center will also become an asset for enterprise competition. With the widespread application of data centers, the data security of data centers has been increasingly emphasized.
[0085] In a traditional large data center, the network structure is usually as Figure 1 shown, that is, the hierarchical inter-networking model. This model includes the following parts:
[0086] Server 140: Each server 140 is a processing and storage entity in the data center, and a large amount of data processing and storage in the data center are completed by these servers 140.
[0087] Access switch 130: The access switch 130 is used to allow the server 140 to access the data center. One access switch 130 accesses multiple servers 140. The access switch 130 is usually located at the top of the rack, so they are also called Top of Rack switches, and they are physically connected to the servers.
[0088] Aggregation switch 120: Each aggregation switch 120 connects multiple access switches 130 and provides other services at the same time, such as firewalls, intrusion detection, network analysis, etc.
[0089] Core Switch 110: The core switch 110 provides high-speed forwarding for the packets entering and leaving the data center and provides connectivity for the aggregation switches 120. The network of the entire data center is divided into an L3 layer routing network and an L2 layer routing network. The core switch 110 usually provides a flexible L3 layer routing network for the network of the entire data center.
[0090] Under normal circumstances, the aggregation switch 120 is the demarcation point between the L2 and L3 layer routing networks. Below the aggregation switch 120 is the L2 network, and above is the L3 network. Each group of aggregation switches manages a Point of Delivery (POD). Each POD has an independent VLAN network. When a server migrates within a POD, it does not need to modify its IP address and default gateway because one POD corresponds to one L2 broadcast domain.
[0091] The spanning tree protocol (STP) is usually used between the aggregation switch 120 and the access switch 130. STP makes only one aggregation layer switch 120 available for a VLAN network, and other aggregation switches 120 are used only when a failure occurs (the dotted lines in the figure above). That is to say, at the level of the aggregation switch 120, horizontal expansion cannot be achieved because even if multiple aggregation switches 120 are added, only one is working.
[0092] Computing device
[0093] In the era of artificial intelligence, the servers (computing devices) 140 in the data center need to process a large number of neural network operations. Neural network operations require large-scale parallel computing. The architecture design of traditional central processing units makes the control unit and storage unit occupy a large part of the space in the architecture, while the computing unit occupies insufficient space. Therefore, it is very effective in logical control but not efficient enough in large-scale parallel computing. In this case, servers (computing devices) 140 with a large number of acceleration units 230 are adopted. As Figure 2 shown, such a computing device 140 includes a memory 210, a cluster of scheduling units 270, and a cluster of acceleration units 280 connected by a bus. The cluster of scheduling units 270 includes multiple scheduling units 220. The cluster of acceleration units 280 includes multiple acceleration units 230.
[0094] As described above, traditional central processing units are very effective in logical control but not efficient enough in large-scale parallel computing. The acceleration unit 230 is a unit used to more effectively improve the operation speed for calculations with different functions and in different fields. For example, it is a processing unit (NPU) dedicated to accelerating the operation processing speed of neural network models. It adopts an architecture of data-driven parallel computing and is a processing unit used to process a large number of operations (such as convolution, pooling, etc.) of each neural network node. Since the data and intermediate results in a large number of operations (such as convolution, pooling, etc.) of each neural network node are closely related throughout the calculation process and are often used, with the existing central processing unit architecture, due to the small memory capacity inside the core of the central processing unit, it is necessary to frequently access the off-chip memory in large quantities, resulting in low processing efficiency. By using such an acceleration unit dedicated to accelerating the operation processing speed of neural network models, since the computing unit has on-chip memory with a storage capacity suitable for neural network computing, avoiding frequent access to the off-chip memory can greatly improve the processing efficiency and computing performance. In the embodiments of the present disclosure, the acceleration unit 230 can be embodied as a processing unit (NPU) specifically designed for neural network operation processing, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.
[0095] The scheduling unit 220 is a processing unit that schedules the acceleration unit 230 and allocates the sequence of pending instructions to be executed to each acceleration unit 230. It can take various forms such as a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc. The scheduling unit 220 can also act as a central processing unit for performing operations in traditional logical control. The scheduling unit 220 reads the pending instructions from the memory 210. When it is found that the instruction is a logical control instruction, the scheduling unit 220 itself executes the instruction to complete the corresponding logical control operation. If it is found that the instruction is an instruction that needs to be processed by the acceleration unit 230, it is allocated to the acceleration unit 230 for execution. The acceleration unit 230 has to accept the scheduling of the scheduling unit 220.
[0096] Taking neural network operations as an example. The memory 210 stores various neural network models, including the nodes of these models and the weight data of the nodes, etc. These neural network models are Figure 2One scheduling unit 220 in it is deployed to an acceleration unit 230. That is, the scheduling unit 220 can send the addresses of the parameters in the model (such as the weights of each node) in the memory 210 to the acceleration unit 230 in the form of instructions. When the acceleration unit 230 actually uses the neural network model for calculation, it will directly address these parameters in the memory 210 according to the addresses of these parameters (such as weights) in the memory 210 and cache them in its on-chip memory. When the acceleration unit 230 actually uses the neural network model for calculation, the scheduling unit 220 will also send the input parameters of the model to the acceleration unit 230 in the form of instructions and cache them in the on-chip memory of the acceleration unit 230. In this way, the acceleration unit 230 can perform inference calculations based on these input parameters and the parameters in the model (such as weights).
[0097] Acceleration unit
[0098] Figure 3 Shows the internal structure and external communication diagram of the acceleration unit 230 according to an embodiment of the present disclosure.
[0099] As Figure 3 shown, the acceleration unit 230 includes a bus channel 391, a second communication control unit 390, an on-chip memory 370, each port 310, a core 340, a command processor 360, a last-level on-chip cache 345, a first communication control unit 330 including a switching module 320, and an on-chip interconnection network 380. The core 340, the command processor 360, the on-chip cache 345, and the communication control unit 330 are interconnected and communicate with each other through the on-chip interconnection network 380.
[0100] The bus channel 391 is the channel for instructions to enter and exit the acceleration unit 230 from the bus. The instructions assigned by the scheduling unit 220 to the acceleration unit 230 to execute enter through the bus channel 391.
[0101] The second communication control unit 390 is a module that controls the mutual access (read, write) between the scheduling unit 220 and the acceleration unit 230. In the embodiments of the present disclosure, when each core 340 of the acceleration unit 230 needs to read and write data to the scheduling unit 220, it can communicate through the second communication control unit 390, and the second communication control unit 330 writes the data to the scheduling unit 220 or reads the data stored in the scheduling unit 220. This will be combined with Figure 6Details of its internal structure are described below. It has a Direct Memory Access (DMA) function, enabling data to be directly written from an attached device (acceleration unit 230) to the memory 210 of the computing device 140. This method significantly improves the data access efficiency compared to the way that all data transfers between devices have to go through the scheduling unit. Due to such a mechanism, the core 340 of the acceleration unit 230 can directly access the memory 210 to read parameters (such as the weights of each node) in the neural network model, etc., greatly improving the data access efficiency.
[0102] The command processor 360 is a device that divides the commands sent from the scheduling unit 220 to the acceleration unit 230 into instructions and assigns them to the cores 340 for execution. After the sequence of instructions to be executed assigned by the scheduling unit 220 enters through the bus channel 391, it is cached in the command processor 360. The command processor 360 selects the cores 340 and assigns the instructions to the corresponding cores 340 for execution. The command processor 360 can, for example, according to the load balancing algorithm, consider the processing loads of each core 340 and assign the instructions to the core 340 with a lighter processing load.
[0103] The core 340 is the core component for processing instructions inside the acceleration unit 230. It accepts the instructions assigned by the command processor 360 and executes them according to the thread groups. Its internal structure is as Figure 4 shown in the hierarchical structure. After the instructions are divided and executed by each thread group, each thread group executes step by step upward along the hierarchical structure from the end of the hierarchical structure (i.e., the first thread group execution unit 342) (for example, passing write requests step by step upward). Regarding Figure 4 the hierarchical structure and the execution of instructions, it will be described in detail later in combination with Figure 4
[0104] The on-chip memory 370 is a memory shared by multiple cores 340 inside the acceleration unit 230. The command processor 360 assigns the sequence of instructions to be executed to different cores 340 for execution. When different cores 340 execute different instructions, since each instruction is interrelated, when a certain core 340 executes a certain instruction, it may need to use the results of other instructions executed by other cores 340. Therefore, there needs to be some shared storage spaces inside the acceleration unit 230 to store the data shared by multiple cores 340. The on-chip memory 370 is such a shared storage space. In Figure 3 one example, an acceleration unit 230 has 4 on-chip memories 370, which is just an example. Those skilled in the art should understand that it can also be other numbers of on-chip memories 370.
[0105] In the acceleration unit, the cache is set up in a hierarchical structure as Figure 4 shown, which will be described in detail later in combination with Figure 4Details are as follows. The last-level cache 345 is the highest-level cache in the acceleration unit 230. It can store some data that the cores 340 often use. These data should have been stored in the on-chip memory 370, but since they are frequently used, they can be stored in the last-level cache 345, which is closer to the cores 340 and has a faster read speed, so as to achieve fast access by the cores 340.
[0106] The on-chip interconnect network 380 is an interconnect network that enables high-speed communication among the command processor 360, the cores 340, the last-level cache 345, the first communication control unit 330, and the second communication control unit 390. Through it, the command processor 360 can distribute instructions to the cores 340 more quickly. When the cores 340 need to read and write data to other acceleration units 230, they can also quickly establish communication with the first communication control unit 330 through it. The first communication control unit 330 then writes out the data or reads the data of other acceleration units 230.
[0107] The inter-chip interconnect network 150 is a network that interconnects the acceleration units 230. Through it, one acceleration unit 230 can directly transfer data to other acceleration units 230 within the same computing device 140 or acceleration units 230 within different computing devices 230. Since neural network operations involve large-scale parallel operations, each operation may be executed by different acceleration units 230 of the same computing device 140 or even acceleration units 230 of different computing devices 140. Since there is usually a correlation between instructions, when an acceleration unit 230 executes an instruction, it usually needs the result of another instruction executed by another acceleration unit 230. Therefore, it is necessary to exchange data between the acceleration units 230. In the traditional technology, this data exchange is carried out through the traditional inter-computing device interconnect network 160 between the computing devices 140. This network can only transfer data at the level of the computing devices 140 and cannot directly transfer data at the level of each acceleration unit 230. When an acceleration unit 230 inside a computing device 140 needs to transfer data to another acceleration unit 230 inside another computing device 140, it first sends the data to be transferred to the network card 290 through the bus inside the computing device 140. The network card 290 is the entrance and exit for data to enter and leave the computing device 140. The network card 290 sends the data to be transferred to the traditional inter-computing device interconnect network 160, and the traditional inter-computing device interconnect network 160 sends or writes the data to be transferred to another acceleration unit 230 of the other computing device 140 through the network card of the other computing device 140. It can be seen that this process experiences the network transmission of the traditional inter-computing device interconnect network 160, the bus transmission inside the computing device 140 where the data is to be sent, and the bus transmission inside the computing device 140 where the data is to be received. This process is relatively complex, and the network transmission of the traditional inter-computing device interconnect network 160 is very slow, which greatly affects the efficiency of data transfer between the acceleration units 230 of different computing devices 140. To improve this transfer efficiency, the inter-chip interconnect network 150 that directly interconnects the acceleration units 230 of each computing device 140 is adopted. Since the acceleration units 230 of each computing device 140 are directly interconnected through this network and the network transmission speed is fast, the efficiency of data transfer between the acceleration units 230 of different computing devices 140 is greatly improved.
[0108] The port 310 is the entrance and exit for the acceleration unit 230 to exchange data with the inter-chip interconnect network 150. An acceleration unit 230 can have multiple ports 310. Each port 310 can simultaneously serve as a sending port for sending data to the inter-chip interconnect network 150 and a receiving port for receiving data from the inter-chip interconnect network 150, or a part of them can be dedicated to the sending port for sending data to the inter-chip interconnect network 150, and the other part can be dedicated to the receiving port for receiving data from the inter-chip interconnect network 150.
[0109] The first communication control unit 330 is a unit that controls the transmission of data and instructions to other acceleration units 230 via the inter-chip interconnect network 150 and controls the reception of data and instructions from other acceleration units 230 via the inter-chip interconnect network 150. When controlling the transmission of data and instructions to other acceleration units 230 via the inter-chip interconnect network 150, it needs to determine the routing of data and instruction transmission and control the transmission of data and instructions according to the determined routing. When controlling the reception of data and instructions from other acceleration units 230 via the inter-chip interconnect network 150, it needs to control the correct writing of the received data into the storage space of the current acceleration unit 230, and at the same time ensure that the ordered instructions entering from different ports 310 can be executed according to the order requirements.
[0110] The switching module 320 is a module that exchanges the data to be sent to the inter-chip interconnect network 150 through each port 310 to the corresponding port 310 for sending through this port 310 and aggregates the data received by each port 310 to the communication control unit 330. A routing table is set in the switching module 320. The switching module 320 can look up the routing table according to the routing determined by the communication control unit 330 for the data to be sent, determine the sending port 310 corresponding to this routing, and thus send the corresponding data through the sending port 310.
[0111] Hierarchical structure of instruction execution unit and cache in acceleration unit 230
[0112] As described above, large-scale data operations usually divide the process into thread blocks downward, and the thread blocks are further divided into warps for processing. For large-scale data operations, the operation rules and instructions are relatively fixed, and only a large number of different data need to be repeatedly operated. Different data are allocated to multiple threads that execute exactly the same instructions. Since these multiple threads execute exactly the same instructions, only the data they operate on are different, they form a warp. Multiple warps form a thread block, and many thread blocks complete the entire process. Correspondingly, the hardware instruction execution units that execute them are also hierarchical, such as Figure 4 shown, the execution unit at the lowest layer is the warp execution unit (PU) 342, the execution unit at the upper layer is the synchronization instruction execution unit (CU) 341, and the upper layer is the acceleration unit core (CE) 340.
[0113] As Figure 4As shown, PU 342 is the bottommost and most basic execution unit. Threads within the same thread group must run on the same PU 342 and cannot run across different PU 342s. However, a single PU 342 can sequentially execute multiple thread groups. Above the PU 342 is the CU 341. It includes multiple PU 342s and a first-level cache 343 shared by these multiple PU 342s. A single CU 341 can run one thread block, but it can also run multiple thread blocks. However, threads within the same thread block must run within the same CU 341. Above the CU 341 is the CE 340. A single CE 340 includes multiple CU 341s and a second-level cache 344 shared by these multiple CU 341s. Finally, a last-level cache 345 is shared by multiple CE 340s.
[0114] If a processing unit 230 receives a large-scale data writing task that writes data to different target entities (such as Figure 4 other processing units 230, the scheduling unit 220, and the on-chip memory 370 of the processing unit 230 itself in
[0115] corresponding to Figure 4In the hierarchical structure, a thread group is executed by a PU 342. The PU 342 issues a write request for the corresponding data. If the write request is a cacheable write request, it will be cached in the first-level cache 343 and will not be passed up further. At this time, although it has been stored in the first-level cache 343 and has not reached its destination, it is not passed to the destination temporarily because traditionally it is considered that although it has not reached the destination, it can be found in the cache and the data can still be retrieved from the cache later. In some applications, although the write request has a destination address, it actually does not care where it is really written as long as the data can be retrieved to meet the requirements. In this case, if it is cached in the first-level cache 343, other PUs 342 of the same CU 341 can quickly read the data directly from the first-level cache 343 without having to read it through multiple levels of caches to the destination, greatly reducing the subsequent data reading time. Moreover, for write requests with strong requirements for writing to the destination, a forced write-out flag can be set in the write request. For such write requests, they must not be stored in the intermediate-level caches and must be written to the real destination. In this way, both the real writing of data with strong requirements for the write destination and the fast storage and fast reading of general data can be achieved.
[0116] If the write request is not cached in the first-level cache 343, it enters the second-level cache 344 shared by multiple CUs 341 in the CE 340. If the write request is not cached in the second-level cache 344 either, it is determined whether to send it to the first communication control unit 330, the second communication control unit 390, or its own last-level cache 345 according to whether the data of the write request is to go to other acceleration units 230, the scheduling unit 220, or the on-chip memory 370 of the current acceleration unit 230.
[0117] If the data of the write request is to go to other acceleration units 230, the write request enters the first communication control unit 330 from the second-level cache 344. The first communication control unit 330 sends it to the third communication control unit 335 in other acceleration units 230 and saves it to the on-chip memory 370 in other acceleration units 230 via the last-level cache 345 in other acceleration units 230. If the data of the write request is to go to the scheduling unit 220, the write request enters the second communication control unit 390 from the second-level cache 344. The second communication control unit 390 sends it to the PCIe bus 223 in the scheduling unit 220 and enters the system memory 222 of the scheduling unit 220 via the PCIe bus 223.
[0118] Internal structures of the first and second communication control units
[0119] As described above, the acceleration unit 230 writes data to other acceleration units 230 and the scheduling unit 220 through the first communication control unit 330 and the second communication control unit 390 respectively. The internal structures of the first communication control unit 330 and the second communication control unit 390 are described in detail below.
[0120] Figure 5 The internal structure of the first communication control unit 330 in the data production main body unit 410 is shown. It includes an instruction distributor 430, a receiving controller 440, a plurality of direct memory access modules 420 corresponding one-to-one to a plurality of ports 310, and the switching module 320 introduced above. In Figure 5 it also shows external devices such as the core 340, the on-chip interconnect network 380, and the on-chip memory 370, in order to clearly show the communication relationship between the first communication control unit 330 and the outside world.
[0121] Although Figure 5 it is not shown in Figure 5 the first communication control unit 330 may also include a command distributor for controlling command-level writing and reading. Since the present disclosure mainly relates to instruction-level writing, in order not to obscure the focus of the present disclosure, the command distributor is not drawn. Different from the command distributor reading commands in sequence from a command storage ring (not shown), in instruction-level writing, the commands have already been decomposed and executed by the core 340 according to instructions, and what is delivered from the core 340 to the communication control unit 330 through the on-chip interconnect network 380 is instructions rather than commands. Therefore, the instruction distributor 430 directly distributes the instructions to different direct memory access modules 520 according to the routing table and issues them through the port 310 corresponding to the direct memory access module 420.
[0122] Data transfer between two main units may reach through different intermediate main units. And between every two adjacent main units, since each main unit has multiple ports 310, different sending ports 310 and receiving ports 310 may also be selected for transmission. Therefore, when data is transferred between the data-producing main unit 410 and other acceleration units 230 (peer main units) serving as the data-consuming main unit 420, there will be multiple paths to reach, and each path is a route. The communication control unit 330 stores a pre-configured routing table for recording the routes between the data-producing main unit 410 and various other peer main units. When the required peer main unit is known, the instruction distributor 430 can determine multiple routes to the required peer main unit by looking up the routing table. Each route of these multiple routes corresponds to a port 310, and each port 310 in turn corresponds to a direct memory access module 420. In this way, the instruction distributor 430 can send instructions to different direct memory access modules 420 respectively via the ports 310 corresponding to these direct memory access modules 420 by looking up the routing table.
[0123] Each direct memory access module 420 maintains two sequences. One sequence stores the instructions to be sent from the corresponding port 310, and the other sequence stores the instructions received from the corresponding port 310. The instructions in the instruction sequence to be sent are sent through the corresponding port 310 in sequence according to the order in the sequence. The instructions received from the port 310 are queued in the received instruction sequence in the direct memory access module 420 corresponding to this port 310 in the order of reception, and enter the receiving controller 440 for processing in the queued order.
[0124] The receiving controller 440 processes the instructions entering from the direct memory access module 420 in sequence. For example, for a write request, the data targeted by the write request is written into the on-chip memory 370 of the peer main unit.
[0125] Figure 6 The internal structure of the second communication control unit 390 in the data-producing main unit 410 is shown. It includes a sending controller 530, a receiving controller 540, and a direct memory access module 520. Figure 6 Also shown are external devices such as the core 340, the on-chip interconnect network 380, and the on-chip memory 370, in order to clearly show the communication relationship between the second communication control unit 390 and the outside world.
[0126] Two sequences are maintained within the direct memory access module 520. One sequence stores the instructions to be sent from the core 340 to the scheduling unit 220 via the sending controller 530, and the other sequence stores the instructions received from the scheduling unit 220. The instructions in the sequence of instructions to be sent are sequentially sent to the scheduling unit 220 in the order of the sequence. The instructions received from the scheduling unit 220 are sequentially queued in the received instruction sequence in the direct memory access module 520 in the order of reception, and enter the receiving controller 540 for processing in the queued order.
[0127] The sending controller 530 is a unit that controls the sending of instructions to the scheduling unit 220. It can sort the instructions to be sent and send them into the sequence of instructions to be sent in the direct memory access module 520. It can also, as in the embodiments of the present disclosure, generate some management instructions with management functions (such as the fake read instruction in the embodiments of the present disclosure) and insert them into the sequence, and can count the responses from the scheduling unit 220.
[0128] The receiving controller 540 is a unit that controls the reception of instructions from the scheduling unit 220. It can take out the instructions from the received instruction sequence in the direct memory access module 520 for execution, and at the same time perform some statistical processing related to instruction reception. For example, in the embodiments of the present disclosure, it can determine whether all the previous write requests from the data production main body unit 410 have been successfully written after receiving the set instruction.
[0129] Cross-subject unit write process of embodiments of the present disclosure
[0130] When the data production main body unit 410 performs the writing of batch data, the target of the batch data writing is not necessarily one. For example, some of the batch data are to be written to other peer acceleration units, and some are to be written to the scheduling unit (such as the CPU) that assigns tasks to it. And these data are interrelated. Therefore, it is easy to make mistakes if reading starts when the batch data have not been completely written to their respective target bodies. Therefore, it is necessary to determine that all the batch data have reached their respective target bodies, which is difficult in itself. The above-mentioned difficulty is more obvious in instruction-level writing. In instruction-level writing, writing is usually performed by warp. A warp includes multiple parallel threads. Within a warp, each thread performs the same code operation on different data it is responsible for. The operations they perform are the same, only the data they target are different. A large computing task command can be decomposed into many warp operations. Monitoring that each thread in a warp has successfully written the data it targets to the target body does not mean that each thread in other warps has also completed similar things, and it is even more difficult to determine that each thread in all warps has successfully written the data it targets to the target body.
[0131] In the embodiments of the present disclosure, when there are multiple target bodies for batch data writing to be executed by a data production main body unit, correct instruction-level writing can be executed, and it is ensured that the use of the written data by the target bodies will not go wrong.
[0132] Generally speaking, the target bodies for which the data production main body unit needs to write batch data include peer main body units and scheduling main body units. When the target body is a peer main body unit, no matter what route the write request sent by the data production main body unit takes to reach the peer main body unit, the peer main body unit will send back a response. The data production main body unit can monitor whether all the write requests to the peer main body unit in this task are received by counting whether the number of write requests sent to the peer main body unit matches the number of returned responses. When the target body is a scheduling main body unit, the scheduling main body unit will not give a response to the received write request. However, a fake read instruction can be sent to the scheduling main body unit after all the write requests. The scheduling main body unit will give a response to this fake read instruction. Moreover, once this response is received, it is defaulted that all the write requests sent before this fake read instruction have been received. Therefore, for the scheduling main body unit, the write requests to the scheduling main body unit in this task can be determined whether all are received in such a way. Therefore, in the embodiments of the present disclosure, for a thread group, a first feedback tracking instruction and a second feedback tracking instruction will be generated, which are respectively used to track whether all the write requests to the peer main body unit in this task are received and whether all the write requests to the scheduling main body unit in this task are received in the above-mentioned manner. Once both are received, a set instruction is sent to the peer main body unit and the scheduling main body unit. When the peer main body unit and the scheduling main body unit receive the set instruction, and it is determined that all the received write requests have been successfully written, the specific flag bit is made to become a specific value, and a set waiting instruction is executed. The set waiting instruction is used to allow subsequent data usage instructions to be executed after the specific flag bit becomes the specific value. In this way, when there are multiple target bodies for batch data writing to be executed by a data production main body unit, it is ensured that the data can only be used after all the batch data has been successfully written to their respective target bodies, so that the use of the data will not go wrong. At the same time, considering that the embodiments of the present disclosure are for instruction-level writing and are written in thread groups, in order to improve efficiency, one thread group is allowed to execute the above process, and after one thread group executes successfully, all thread groups share this information, avoiding the huge overhead caused by each thread group executing the above process separately.
[0133] The above process will be described in detail below.
[0134] In one embodiment, for a thread group, the first thread group execution unit 342 in the acceleration unit 230, which is the data production main body unit 410, determines the target body for each piece of data among all the data involved in the threads in the thread group. These target bodies may include at least one of other peer acceleration units, the scheduling unit, and the on-chip memory 370 of this acceleration unit itself. If the target body is another peer acceleration unit 230, the write request for this data is sent from the first thread group execution unit 342 and reaches the first-level cache 343. If the write request is a forced write-out, it does not stay in the first-level cache 343 and reaches the second-level cache 344. Since the write request is a forced write-out, it also does not stay in the second-level cache 344 and reaches the first communication control unit 330 via the second-level cache 344, and is sent by the first communication control unit 330 to the peer acceleration unit 230. In this case, the first communication control unit 330 returns an internal confirmation reply indicating that the write request has been sent to the peer acceleration unit 230. If the target body is the scheduling unit 220, the write request for this data is sent from the first thread group execution unit 342. If the write request is a forced write-out, it reaches the second communication control unit 390 via the first-level cache 343 and the second-level cache 344, and is sent by the second communication control unit 390 to the scheduling unit 220. In this case, the second communication control unit 390 returns an internal confirmation reply indicating that the write request has been sent to the scheduling unit 220. If the target body is the on-chip memory 370 of this acceleration unit 230 itself, the write request for this data is sent from the first thread group execution unit 342. If the write request is a forced write-out, it reaches the on-chip memory 370 via the first-level cache 343, the second-level cache 344, and the last-level cache 345. In this case, the on-chip memory 370 returns an internal confirmation reply indicating that the write request has been received and successfully written. It can be seen that in the above three cases, only in the case of the on-chip memory 370, the internal confirmation reply truly indicates that the write request has been successfully completed and written. The internal confirmation replies in the other two cases only indicate that the write request has flowed out of the acceleration unit 230, which is the data production main body unit 410, and further tracking is required, which will be described in detail below.
[0135] The above processes are all based on the premise that write requests are compulsorily written out. If a write request is not compulsorily written out, it may be cached in any one of the first-level cache 343, the second-level cache 344, and the last-level cache 345 without further transmission. In this case, the cache that caches the write request returns an internal confirmation response indicating that the write request has been cached. In summary, the internal confirmation response can only indicate that the write request has been processed in the acceleration unit 230 that serves as the data production main body unit 410, does not represent that it has truly reached its intended target, nor does it represent that the write has been successful. Reaching its intended target and successful writing requires the processes described below to ensure.
[0136] The above process is only for one thread group and ensures that the write request has been processed in the acceleration unit 230 that serves as the data production main body unit 410, but does not mean that all thread groups have completed the above process. In the embodiments of the present disclosure, after a thread group determines that it has received the internal confirmation responses for all the write requests involved by the threads of this thread group, it generates a first message indicating that it has received the internal confirmation responses for all the write requests of this thread group. Then, this thread group exchanges the first message with other thread groups. If all thread groups have generated the first message indicating that they have received the internal confirmation responses for all the write requests of this thread group, it means that the write requests have been processed in the acceleration unit 230 that serves as the data production main body unit 410, and the further tracking of the write requests reaching their intended targets and successful writing as described below can be performed. As Figure 7 shown, _____syncthreads() is the code statement for thread groups to exchange the first message, and when compiled into assembly language, it is the assembly statement S_BLKSYN on the left.
[0137] Next, the synchronization instruction execution unit 341 executes for a thread group: generating at least one of a first feedback trace instruction and a second feedback trace instruction based on the synchronization instruction, and respectively sending them to at least one of the first communication control unit and the second communication control unit. The synchronization instruction is an instruction used to complete synchronization. Synchronization refers to a mechanism that ensures that after the writing of batch data to different target bodies is completed, each target body can use the data, so as to avoid data usage errors. Since the process of each of the above thread groups exchanging the first information ensures that the write requests have been processed in the acceleration unit 230 which is the data production main body unit 410, the synchronization instruction is mainly used to ensure that the write requests sent from this data production main body unit 410 can be received by different target bodies. For the target body of the on-chip memory 370 of the current acceleration unit 230, the internal acknowledgment response has ensured that the write request has been received by the target body. However, for the peer acceleration unit 230 and the scheduling unit 230, additional mechanisms are required for guarantee, and the guarantee methods are different. Therefore, at least one of the first feedback trace instruction and the second feedback trace instruction is generated based on the synchronization instruction and sent to at least one of the first communication control unit 330 and the second communication control unit 390. As Figure 7 shown, ___threadfence_system() is the synchronization instruction, and when compiled into assembly language, it is the assembly statement FENCE.W.SYS on the left. The generated first feedback trace instruction and second feedback trace instruction are not shown.
[0138] When the target body is the peer acceleration unit 230, regardless of the routing through which the write request sent from the acceleration unit 230 which is the data production main body unit 410 reaches the peer main body unit 230, the peer main body unit 230 will send back an acknowledgment, that is, the first acknowledgment. Therefore, after the first communication control unit 330 receives the first feedback trace instruction, it counts the first acknowledgments for the write requests sent to the peer acceleration unit 230. If the number of received first acknowledgments is equal to the number of write requests sent to the peer main body unit, the first trace is successful. Specifically, as Figure 5As shown, the instruction distributor 430 has a routing table that stores the routes from the acceleration unit 230, which is the data production main body unit 410, to different peer acceleration units 230. When the instruction distributor 430 needs to distribute a write request, it refers to the routing table to find its corresponding route, thereby finding the corresponding port 310 and the corresponding direct memory access module 420 for this route, and distributes the write request to the sequence of write requests to be sent in the direct memory access module 420 to queue, and then exchanges it through the switching module 320 to the corresponding port 310 for sending. In one embodiment, the instruction distributor 430 can count the above first responses by thread block, and can set a corresponding counter for a single thread block inside, and the initial value of the counter is 0. Each time during the execution of the thread block, after a write request is sent from the port 310, the port 310 notifies the instruction distributor 430 to increment the corresponding counter by 1. After all the write requests corresponding to the thread block are sent, the value of the corresponding counter reaches the maximum value. Then, each time the port 310 receives the first response of a write request, it notifies the instruction distributor 430 to decrement the counter by 1 until the value of the counter is decremented to 0, then the first tracking is successful.
[0139] When the target entity is the scheduling unit 220, the scheduling unit 220 will not give a response to the received write request. After the second communication control unit 390 receives the second feedback tracking instruction, it cannot track in the way of counting the received responses as above. However, a fake read instruction can be sent to the scheduling unit 220 after all write requests are sent. The scheduling unit 220 will give a response to this fake read instruction, that is, the second response. Since all write requests are sent before the fake read instruction, and there is generally a unique route between the scheduling unit 220 and the acceleration unit 230, once the second response is received, it is defaulted that all write requests sent before the fake read instruction have been received, that is, the second tracking is successful. Therefore, for the scheduling unit 220, it is possible to determine whether all write requests sent to the scheduling unit 220 have been received in this way. As Figure 6 shown, since the scheduling unit 220 corresponding to the acceleration unit 230, which is the data production main body unit 410, is unique, the transmission controller 530 does not need to store a routing table. It can send all write requests to be sent to the scheduling unit 220 to the sequence of write requests to be sent in the direct memory access module 520, queue in this sequence, and send them to the scheduling unit 220 in the order of queuing. After the transmission controller 530 sends all write requests to the sequence of write requests to be sent in the direct memory access module 520, it generates a fake read instruction and sends the fake read instruction to the sequence of write requests to be sent in the direct memory access module 520, and sends it to the scheduling unit 220 in order. If the second response from the scheduling unit 220 is received, it means that all write requests sent to the scheduling unit 220 before have also been successfully received by the scheduling unit 220, and the second tracking is successful.
[0140] If the synchronization instruction execution unit 341 determines that both the first trace and the second trace are successful (which indicates that the write requests to the peer acceleration unit 230 and the scheduling unit 220 in the thread group are successfully received by the peer acceleration unit 230 and the scheduling unit 220), and during the process of exchanging the first information above, it is confirmed that for the write request to the on-chip memory 370 of the current acceleration unit 230, the internal confirmation response of the corresponding on-chip memory 370 is also received (indicating that the write requests to the on-chip memory 370 of the current acceleration unit 230 are successfully received by the on-chip memory 370), a set instruction can be sent to the peer acceleration unit 230 and the scheduling unit 220 to set a specific flag to a specific value. An example of the set instruction is Figure 7 S_SIGNAL.SYS marker,val in Figure 7 , which means setting the flag bit marker to the value val. For the on-chip memory 370 of the current acceleration unit 230, it can be pre-set to a read-prohibited state, that is, it cannot be used. When both the above first trace and second trace are successful, and during the process of exchanging the first information above, it is confirmed that the internal confirmation response of the on-chip memory 370 is also received, the read-prohibited state is changed to a readable state, allowing the data usage instruction to use the data therein.
[0141] The above first feedback trace and second feedback trace only ensure that the write requests to the peer acceleration unit 230 and the scheduling unit 220 in the thread group are successfully received by the peer acceleration unit 230 and the scheduling unit 220. As for whether they are successfully written, it needs to be controlled at the peer acceleration unit 230 and the scheduling unit 220 end in combination with the above set instruction, which will be further described in detail later.
[0142] In addition, in one embodiment, although the synchronization instruction execution unit 341 is to generate a first feedback tracking instruction and a second feedback tracking instruction based on a synchronization instruction, the first feedback tracking instruction is used to track whether a write request to the peer acceleration unit 230 is successfully received by the other party, and the second feedback tracking instruction is used to track whether a write request to the scheduling unit 220 is successfully received by the other party. However, in some cases, the write requests in a thread group may all be to the peer acceleration unit 230 or all to the scheduling unit 220. Mechanically requiring the synchronization instruction execution unit 341 to generate the first feedback tracking instruction and the second feedback tracking instruction based on the synchronization instruction may cause additional unnecessary overhead. Therefore, in one embodiment, during the process of exchanging the first information mentioned above, when the internal confirmation reply of the first communication control unit 330 is received (indicating that there is indeed a write request to the peer acceleration unit 230), the synchronization instruction execution unit 341 generates the first feedback tracking instruction; when the internal confirmation reply of the second communication control unit 390 is received (indicating that there is indeed a write request to the scheduling unit 220), the synchronization instruction execution unit 341 generates the second feedback tracking instruction. In this way, some unnecessary overhead can be avoided.
[0143] The above description is based on the premise that all write requests are forced write requests, that is, the write requests pass through the first-level cache 343, the second-level cache 344, etc. and will not be cached and stopped from being transmitted. However, in reality, if the write request sent by the first thread group execution unit 342 for a thread group is not a forced write, it may be cached when passing through the first-level cache 343 and will not reach the first communication control unit 330 at all. Although this can improve the reading speed greatly because when other thread group execution units 342 in the same synchronization instruction execution unit 341 need this data, they can directly read the data from the first-level cache 343 instead of reading from the real target body, not flushing out this data is extremely disadvantageous for ensuring the true writing of all data in the target body and the correct use of data by subsequent data usage instructions. Therefore, in one embodiment, the first multi-level cache (specifically, the first-level cache 343 and the second-level cache 344) performs at least one of the first flush and the second flush based on the synchronization instruction.
[0144] The first flush is for the data to be sent to the peer acceleration unit 230 cached in the cache, which is flushed out from the first multi-level cache, sent from the first communication control unit 330 to the peer acceleration unit 230 through a higher-level cache, and the third reply from the peer acceleration unit 230 is received. Specifically, for the write request cached in the first-level cache 343 to the peer acceleration unit 230, it is flushed out from the first-level cache 343 (such as Figure 7 the assembly statement L1 FLUSH), and passes through the second-level cache 344 without staying (such asFigure 7 The assembly statement L2 INVALIDATE of ) is sent into the first communication control unit 330 and the third response of the peer acceleration unit 230 is received; for the write requests to the peer acceleration unit 230 cached in the second-level cache 344, they are flushed out from the second-level cache 344 (such as Figure 7 The assembly statement L2 INVALIDATE of ) is sent into the first communication control unit 330 and the third response of the peer acceleration unit 230 is received.
[0145] The second flushing empty is for the data to be sent to the scheduling unit 220 cached in the cache. It is flushed out from the first multi-level cache and sent to the scheduling unit 220 from the second communication control unit 390 through the higher-level cache. A second false read instruction is sent to the scheduling unit 220 and the fourth response of the scheduling unit 220 to the second false read instruction is received. Specifically, for the write requests to the scheduling unit 220 cached in the first-level cache 343, they are flushed out from the first-level cache 343 (such as Figure 7 The assembly statement L1 FLUSH of ) is sent into the second communication control unit 390 without stopping through the second-level cache (such as Figure 7 The assembly statement L2 INVALIDATE of ), then a second false read instruction is sent to the scheduling unit 220 and the fourth response of the scheduling unit 220 to the second false read instruction is received; for the write requests to the scheduling unit 220 cached in the second-level cache 344, they are flushed out from the second-level cache 344 (such as Figure 7 The assembly statement L2 INVALIDATE of ) is sent into the second communication control unit 390, then a second false read instruction is sent to the scheduling unit 220 and the fourth response of the scheduling unit 220 to the second false read instruction is received.
[0146] In the case where the write requests may be cached in the first-level cache 343 and the second-level cache 344, such as Figure 7 The set instruction S_SIGNAL.SYS marker,val cannot be sent after both the first trace and the second trace are successful. Instead, it should be sent to the peer acceleration unit 230 and the scheduling unit 220 only when both the first trace and the second trace are successful, the third responses for the data to be sent to the peer main unit from the flushed cache are received, and the fourth response for the second false read instruction is received. This is because if some of the write requests are cached in the first-level cache 343 and the second-level cache 344, even if both the first trace and the second trace are successful, the part of the write requests cached in the first-level cache 343 and the second-level cache 344 still do not reach the peer acceleration unit 230 or the scheduling unit 220, and the effect of ensuring that all write requests involved in the thread group are successfully received by the target body is not achieved.
[0147] In addition, in the embodiments of the present disclosure, after at least one of the first communication control unit 330 and the second communication control unit 390 respectively executes at least one of the first feedback tracking instruction and the second feedback tracking instruction, as Figure 7 shown, the wait feedback tracking instruction S_WAIT.FENCE_SYS is executed to ensure that after at least one of the first feedback tracking instruction and the second feedback tracking instruction is executed, the set instruction S_SIGNAL.SYS marker,val can be issued. In fact, the function of the wait feedback tracking instruction S_WAIT.FENCE_SYS is that if at least one of the first feedback tracking instruction and the second feedback tracking instruction is not executed, any subsequent instructions shall not be executed.
[0148] In addition, in the embodiments of the present disclosure, after at least one of the first multi-level cache performs at least one of the first flush and the second flush, as Figure 7 shown, the wait flush instruction S_WAIT.CACHE is executed to ensure that after at least one of the first flush and the second flush is executed, the set instruction S_SIGNAL.SYS marker,val can be issued. In fact, the function of executing the wait flush instruction S_WAIT.CACHE is that if at least one of the first flush and the second flush is not executed, any subsequent instructions shall not be executed.
[0149] In Figure 6 the statement if(tid = 0) indicates that subsequent synchronization instructions such as __threadfence_system() are executed for the 0th thread group (even for the 0th thread of only the 0th thread group). This verifies that the above synchronization instruction execution unit 341 only executes the generation of the first feedback tracking instruction and the second feedback tracking instruction for one thread group, and sends a set instruction and other operations after successful tracking. The overhead of having each thread group execute the above operations is very large. To reduce the execution overhead, the embodiments of the present disclosure let one thread group execute the above operations, and by means of subsequent operations of making this thread group exchange third information (the third information indicates that a set instruction has been sent to the peer acceleration unit 230 and the scheduling unit 220) with other thread groups, to make each thread group know the information that the set instruction has been issued, thereby greatly reducing the execution overhead. As Figure 7 shown, ___syncwarp() is the statement for exchanging the third information between thread groups.
[0150] After the above processing on the side of the acceleration unit 230, which is the data production main body unit 410, it has been ensured that the write requests from the acceleration unit 230, which is the data production main body unit 410, to each target body have successfully reached the target body, but it cannot be guaranteed that the writes are successfully performed on the target body side. The successful writing on the target body side is ensured by the following operations of the target body (data consumption main body unit 420).
[0151] On the side of the peer acceleration unit 230 or the scheduling unit 220, which is the target body, the third communication control unit 335 executes for a thread group: after receiving the setting instruction from the acceleration unit 230, which is the data production main body unit 410, it determines whether all the previously received write requests from the data production main body unit 410 have been successfully written. Specifically, for each port of the target body (peer acceleration unit 230 or scheduling unit 220), a received write request sequence is maintained in the third communication control unit 335. If all the received write request sequences corresponding to all ports are empty, it means that all the currently received write requests of the target body have been written into its on-chip memory. This includes the write requests from the data production main body unit 410 and the write requests from other main bodies. If you want to specifically distinguish the write requests from the data production main body unit 410, a source flag can be preset for each write request. If all the write requests with the source flag of the data production main body unit 410 in the received write request sequences corresponding to all ports no longer exist, it means that all the previously received write requests from the data production main body unit 410 have been successfully written. In this way, it is ensured that the write requests are successfully written on the target body side.
[0152] After ensuring that the write requests are successfully written on the target body side, make the specific flag bit become a specific value ( Figure 7 in which the specific flag bit is marker and the specific value is 1), which means that the target body can freely use the data written by the write request without error. At this time, the third communication control unit 335 executes the set wait instruction and the data use instruction. The set wait instruction is used to allow the data use instruction to execute after the specific flag bit becomes the specific value. In Figure 7 it, the set wait instruction is __acquire_sys(marker)!= 1. Its translated assembly statement is S_WAIT.LD. It means that after the specific flag bit marker becomes 1, it allows the subsequent data use instruction LD…,&Data to execute.
[0153] In addition, inside the data consumption main body unit 420 (target body), similar to the data production main body unit 410, there may be a second multi-level cache similar to the first multi-level cache. The second multi-level cache includes a first-level cache 343, a second-level cache 344, and a last-level cache 345. The data to be written into the on-chip memory 370 of the data consumption main body unit 420 inside the data consumption main body unit 420 may be cached by one of the first-level cache 343, the second-level cache 344, and the last-level cache 345 during the process of propagating layer by layer through the second multi-level cache, and will no longer be sent down to the on-chip memory 370. Therefore, when a data usage instruction is executed in the data consumption main body unit 420, the data usage instruction first looks for cached data in the first-level cache 343, the second-level cache 344, and the last-level cache 345. If not, it reads from the on-chip memory 370. If there is, it reads directly from the cache. However, the process described above uses a complex mechanism to ensure that the data to be written from the data production main body unit 410 into the data consumption main body unit 420 is finally actually written into the on-chip memory 370 of the data consumption main body unit 420. If subsequent data usage instructions read from the second multi-level cache again, the read result may be incorrect, and the mechanism that ensures the data to be written from the data production main body unit 410 into the data consumption main body unit 420 is actually written into the on-chip memory 370 of the data consumption main body unit 420 becomes meaningless. Therefore, in one embodiment, on the side of the data consumption main body unit 420, after the specific flag bit becomes a specific value, the second multi-level cache also needs to execute a cache invalidation instruction and wait for the invalidation success instruction. The cache invalidation instruction is used to prevent the data stored in the second multi-level cache from being read by subsequent read instructions, such as Figure 7 the statements L1.FLUSH and L2.INVALIDATE in Figure 7 , where L1.FLUSH is used to ensure that the data cached in the first-level cache 343 is not read by subsequent instructions, and L2.INVALIDATE is used to ensure that the data cached in the second-level cache 344 is not read by subsequent instructions. The wait for invalidation success instruction is used to wait for the invalidation instruction to execute completely, and only when the invalidation instruction has executed completely, is the data usage instruction allowed to execute, such as the statement S_WAIT.CACHE in . After executing the cache invalidation instruction and the wait for invalidation success instruction, subsequent data usage instructions will no longer read the data cached in the second multi-level cache, but use the actual data successfully written by the data production main body unit 410 into the on-chip memory 370 of the data consumption main body unit 420, avoiding incorrect data reading.
[0154] After the above process is completed, the second thread group execution unit 342 causes the thread group to exchange second information with other thread groups, and the second information indicates that the data usage instruction is allowed to be executed. The purpose of exchanging the second information is as follows: Whether the write requests of the data production main body unit received before the above determination are all successfully written, the setting in the case of successful writing, and the process of executing the set waiting instruction are only executed for one thread group, or even one thread in a thread group. Exchanging the second information can save the huge overhead caused by each thread group executing the above process separately. A thread group has completed the above process, and subsequent data usage instructions can use the data successfully written by the data production main body unit 410. However, other thread groups do not know this situation, which may cause them to still not dare to execute subsequent data usage instructions. In the embodiment of the present disclosure, the second information is transmitted between thread groups. Through the second information, each thread group knows this situation, so that subsequent data usage instructions of each thread group can safely use the data successfully written by the data production main body unit 410. The statement for exchanging the second information between thread groups is as Figure 7 the ___syncthreads() in it, and the compiled assembly language is S_BLKSYN.
[0155] Communication control method of data production subject unit
[0156] As Figure 8 shown, according to an embodiment of the present disclosure, a data writing method at the data production main body unit 410 side is provided, which is executed by the data production main body unit 410, and the method includes:
[0157] Step 610, for a thread group, generate at least one of a first feedback tracking instruction and a second feedback tracking instruction based on a synchronization instruction, and send them to at least one of the first communication control unit 330 and the second communication control unit 390 respectively. If the first communication control unit 330 receives the first feedback tracking instruction, count the first responses to the write requests sent to the peer main body unit. If the number of received first responses is equal to the number of write requests sent to the peer main body unit, the first tracking is successful; if the second communication control unit 390 receives the second feedback tracking instruction, send a first false read instruction to the scheduling main body unit. If the second response to the first false read instruction is received from the scheduling main body unit, the second tracking is successful;
[0158] Step 620, if both the first tracking and the second tracking are successful, send a setting instruction to the peer main body unit and the scheduling main body unit, so that after receiving the setting instruction, the data consumption main body unit allows the use of the written data on the premise that all write requests of the data production main body unit received before the setting instruction are successfully written.
[0159] Since the implementation details of the above process have been described in detail in the device embodiment of the foregoing data production main body unit 410, for the sake of saving space, they will not be elaborated here.
[0160] Communication control method of data consumption subject unit
[0161] As Figure 9 shown, according to an embodiment of the present disclosure, a cross-main body unit writing method on the side of the data consumption main body unit 420 is provided, which is executed by the data consumption main body unit 420 and includes:
[0162] Step 710: For a thread group, after receiving a setting instruction from the data production main body unit 410, determine whether all the write requests of the data production main body unit 410 received previously have been successfully written. If it is determined that all have been successfully written, make the specific flag bit become a specific value, and execute a setting wait instruction and a data usage instruction. The setting wait instruction is used to allow the data usage instruction to be executed after the specific flag bit becomes the specific value. Wherein, the setting instruction is issued by the data production main body unit when it is determined that both the first trace of the write request to the peer main body unit and the second trace of the write request to the scheduling main body unit are successful. The first trace includes counting the first responses to the write requests sent to the peer main body unit and determining that the number of received first responses is equal to the number of write requests sent to the peer main body unit. The second trace includes, after determining that the write request sent to the scheduling main body unit has been successfully sent, sending a first false read instruction to the scheduling main body unit and receiving a second response from the scheduling main body unit to the first false read instruction.
[0163] Step 720: Cause the thread group to exchange second information with other thread groups, and the second information indicates that the data usage instruction is allowed to be executed.
[0164] Since the implementation details of the above process have been described in detail in the device embodiment of the foregoing data production main body unit 420, for the sake of saving space, they will not be elaborated here.
[0165] Commercial value of the present disclosure
[0166] In the embodiments of the present disclosure, when there are multiple targets for batch data writing to be executed by a data production main body unit 410, correct instruction-level writing can be executed, and it is ensured that the use of the written data by the target will not go wrong, so that the writing of data by the data production main body unit 410 and the use of data by the target are carried out in an orderly manner. It can be applied to multi-routing instruction-level writing among various main bodies such as acceleration units, processing units, system-on-chips, computing devices, and communication terminals. Experiments have proved that when it is applied to the instruction-level writing from an acceleration unit to a peer acceleration unit and a scheduling unit, the execution accuracy rate of the usage instructions for the written data can be greatly improved, and it has a broad market prospect.
[0167] Those skilled in the art can understand that the present disclosure can be implemented as a system, a method, and a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely, complete hardware, complete software (including firmware, resident software, microcode), and can also be implemented in the form of a combination of software and hardware. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program code.
[0168] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium is, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium include: an electrical connection of one or more specific wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory, a magnetic memory, or any suitable combination of the above. In this article, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by, or in combination with, a processing unit, apparatus, or device.
[0169] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any other suitable combination. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by, or in combination with, an instruction system, apparatus, or device.
[0170] The program code contained on a computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., and any suitable combination of the above.
[0171] The computer program code for performing the embodiments of the present disclosure can be written in one or more programming languages or combinations. The programming languages include object-oriented programming languages such as JAVA and C++, and can also include conventional procedural programming languages such as C. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0172] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A data production main body unit for writing data to a target body by thread groups, where the target body includes at least one of a peer main body unit and a scheduling main body unit, and the data production main body unit includes: At least one of a first communication control unit and a second communication control unit, where the first communication control unit is used to send a write request to the peer main body unit, and the second communication control unit is used to send a write request to the scheduling main body unit; A synchronization instruction execution unit for executing, for a thread group: generating at least one of a first feedback tracking instruction and a second feedback tracking instruction based on a synchronization instruction, and respectively sending them to at least one of the first communication control unit and the second communication control unit. Wherein, if the first communication control unit receives the first feedback tracking instruction, it counts the first responses to the write requests sent to the peer main body unit. If the number of received first responses is equal to the number of write requests sent to the peer main body unit, the first tracking is successful; If the second communication control unit receives the second feedback tracking instruction, it sends a first false read instruction to the scheduling main body unit. If it receives the second response from the scheduling main body unit to the first false read instruction, the second tracking is successful; If both the first tracking and the second tracking are successful, a setting instruction is sent to the peer main body unit and the scheduling main body unit, so that after receiving the setting instruction, the data consumption main body unit allows the use of the written data on the premise of ensuring that all the write requests of this data production main body unit received before this setting instruction are successfully written.
2. The data production main body unit according to claim 1, further including: A first thread group execution unit, after sending a write request to at least one of the first communication control unit and the second communication control unit for a thread group, receiving an internal confirmation response from at least one of the first communication control unit, the second communication control unit, or a cache that caches the write request in a first multi-level cache, and exchanging first information with other thread groups. The first multi-level cache is a multi-level cache located between the first thread group execution unit and the first communication control unit or the second communication control unit. The first information indicates whether an internal confirmation response to all write requests has been received. Wherein, after all thread groups receive the internal confirmation responses to all write requests, the synchronization instruction execution unit starts to work.
3. The data production main body unit according to claim 2, wherein, If the write request sent by the thread group is cached in the first multi-level cache, the first multi-level cache performs at least one of a first flush and a second flush based on the synchronization instruction, where the first flush is for the data cached to be sent to the peer main unit, which is flushed out from the first multi-level cache and sent to the peer main unit from the first communication control unit via a higher-level cache, and the third response from the peer main unit is received; the second flush is for the data cached to be sent to the scheduling main unit, which is flushed out from the first multi-level cache and sent to the scheduling main unit from the second communication control unit via a higher-level cache, and a second false read instruction is sent to the scheduling main unit, and the fourth response to the second false read instruction from the scheduling main unit is received; where the setting instruction is issued to the peer main unit and the scheduling main unit only when both the first trace and the second trace are successful, the third response is received for the data cached to be sent to the peer main unit that has been flushed out, and the fourth response is received for the second false read instruction.
4. The data production main body unit according to claim 3, wherein, After the synchronization instruction execution unit executes at least one of the first feedback trace instruction and the second feedback trace instruction respectively in at least one of the first communication control unit and the second communication control unit, it executes a wait feedback trace instruction to ensure that the setting instruction can be issued only after at least one of the first feedback trace instruction and the second feedback trace instruction is executed; after the first multi-level cache performs at least one of the first flush and the second flush, it executes a wait flush instruction to ensure that the setting instruction can be issued only after at least one of the first flush and the second flush is executed.
5. The data production main body unit according to claim 2, wherein, After sending the setting instruction to the peer main unit and the scheduling main unit, the first thread group execution unit also causes the thread group to exchange third information with other thread groups, and the third information indicates that the setting instruction has been sent to the peer main unit and the scheduling main unit.
6. The data production main body unit according to claim 2, wherein, The target body further includes the on-chip memory of the data production main unit. After the first thread group execution unit passes the write request of the thread group to the on-chip memory, it receives the internal confirmation response of the on-chip memory. The setting instruction is issued only when both the first trace and the second trace are successful and the internal confirmation response of the on-chip memory is received for the write request to the on-chip memory.
7. The data production main body unit according to claim 1 or 2, wherein The data production main unit is a first processing unit, the data consumption main unit is a second processing unit, the first processing unit and the second processing unit communicate through an inter-chip interconnect network, and the first processing unit writes the data in the first on-chip memory of the first processing unit into the second on-chip memory of the second processing unit through the inter-chip interconnect network.
8. The data production main body unit according to claim 7, wherein, The first processing unit is a first acceleration unit for executing tasks assigned by the first scheduling unit; the second processing unit is a second acceleration unit for executing tasks assigned by the second scheduling unit.
9. A data consumption main unit, comprising: A third communication control unit, configured to perform, for a thread group: after receiving a set instruction from a data production main body unit, determine whether all the write requests of the data production main body unit received previously have been successfully written, and in the case where it is determined that all have been successfully written, change a specific flag bit to a specific value, and execute a set wait instruction and a data usage instruction, where the set wait instruction is used to allow the data usage instruction to be executed after the specific flag bit becomes the specific value. Wherein, the set instruction is issued by the data production main body unit when it is determined that both a first trace of a write request to a peer main body unit and a second trace of a write request to a scheduling main body unit are successful. The first trace includes counting first responses to write requests sent to the peer main body unit and determining that the number of received first responses is equal to the number of write requests sent to the peer main body unit; the second trace includes, after determining that a write request sent to the scheduling main body unit has been successfully sent, sending a first false read instruction to the scheduling main body unit and receiving a second response from the scheduling main body unit to the first false read instruction. A second thread group execution unit, configured to enable the thread group to exchange second information with other thread groups, where the second information indicates that the data usage instruction is allowed to be executed.
10. The data consumption main body unit according to claim 9, further comprising: A second multi-level cache, located between the second thread group execution unit and the third communication control unit. After the specific flag bit becomes the specific value, the second multi-level cache executes a cache invalidation instruction and a wait for invalidation success instruction. The cache invalidation instruction is used to prevent data stored in the second multi-level cache from being read by subsequent read instructions, and the wait for invalidation success instruction is used to wait for the invalidation instruction to be executed completely, and only when the invalidation instruction is executed completely, the data usage instruction is allowed to be executed.
11. A system on chip, comprising the data production main body unit according to any one of claims 1-8 and / or the data consumption main body unit according to any one of claims 9-10.
12. A computing device, comprising the data production main body unit according to any one of claims 1-8 and / or the data consumption main body unit according to any one of claims 9-10.
13. A data writing method at a data production main body unit side, for writing data to a target body by thread group, where the target body includes at least one of a peer main body unit and a scheduling main body unit, and the data production main body unit includes at least one of a first communication control unit and a second communication control unit. The first communication control unit is used to send a write request to the peer main body unit, and the second communication control unit is used to send a write request to the scheduling main body unit. The method includes: For a thread group, at least one of a first feedback tracking instruction and a second feedback tracking instruction is generated based on a synchronization instruction and sent to at least one of the first communication control unit and the second communication control unit. If the first communication control unit receives the first feedback tracking instruction, it counts the first responses to the write requests sent to the peer entity unit. If the number of received first responses is equal to the number of write requests sent to the peer entity unit, the first tracking is successful. If the second communication control unit receives the second feedback tracking instruction, it sends a first false read instruction to the scheduling entity unit. If the second response to the first false read instruction is received from the scheduling entity unit, the second tracking is successful. If both the first tracking and the second tracking are successful, a setting instruction is sent to the peer entity unit and the scheduling entity unit, so that after receiving the setting instruction, the data consumption entity unit allows the use of the written data on the premise that all the write requests of the data production entity unit received before the setting instruction are successfully written.
14. A data usage method at the data consumption entity unit side, comprising: For a thread group, after receiving a setting instruction from the data production entity unit, it determines whether all the write requests of the data production entity unit received before are successfully written. If it is determined that all are successfully written, a specific flag bit is changed to a specific value, and a setting wait instruction and a data usage instruction are executed. The setting wait instruction is used to allow the execution of the data usage instruction after the specific flag bit is changed to the specific value. The setting instruction is sent by the data production entity unit when it determines that both the first tracking of the write requests to the peer entity unit and the second tracking of the write requests to the scheduling entity unit are successful. The first tracking includes counting the first responses to the write requests sent to the peer entity unit and determining that the number of received first responses is equal to the number of write requests sent to the peer entity unit. The second tracking includes sending a first false read instruction to the scheduling entity unit after determining that the write request sent to the scheduling entity unit has been successfully sent and receiving the second response to the first false read instruction from the scheduling entity unit. The thread group exchanges second information with other thread groups, and the second information indicates that the execution of the data usage instruction is allowed.
Citation Information
Patent Citations
Speculative multithreading memory data synchronous execution method under support of compiler and device thereof
CN101833440A
Data storage method and a server applied to distributed server cluster
CN107295080A