Data reduction method, device, medium and training system in distributed training
By setting up an atomic operation module in a multi-core chip, receiving and cacheing data specification instructions, and performing data read and write operations in the order of reception, the problems of logical errors and control overhead in distributed training are solved, and efficient and accurate data specifications are achieved.
Patent Information
- Application Number
- CN202310061723.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-01-18
AI Technical Summary
In distributed training scenarios, when multiple computing cores update data to the same memory address at the same time, logical errors may occur, and the prior art sequential control method leads to greater control overhead and performance impacts.
By setting up an atomic operation module close to the external memory in the multi-core chip, the data specification instructions are received and stored in the instruction buffer area, and the data read and write atomic operations are performed in the receiving order, thereby avoiding sequential control.
It reduces the control overhead of data regulation operations, avoids logical errors, and improves the accuracy and efficiency of data regulation.
Smart Images

Figure CN116243978B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chips, and in particular to a data reduction method, device, medium and training system in distributed training. Background Art
[0002] Artificial intelligence training often includes distributed training scenarios with data reduction, where multiple computing cores independently perform calculations to produce multiple results. Data reduction updates the multiple results to the original data, for example, by summing them to ultimately produce one piece of data.
[0003] When implementing data reduction in distributed training scenarios, existing technologies generally run multiple computing cores in multiple chips of a computing accelerator card, or run multiple computing cores in multiple chips of multiple computing accelerator cards. Multiple computing cores simultaneously update the data stored in the same memory address.
[0004] In the process of implementing the present invention, the inventors found that the existing technology has the following defects: when multiple computing cores update the data stored in the same memory address at the same time, if there is no sequence control between the computing cores, different computing cores each read data from the memory address, update it, and then write it back to the memory address, which may cause logical errors; if the order of data updates between the computing cores is controlled, for example, controlling multiple computing cores to update data serially, it will cause a large control overhead and will also affect the performance of the computing cores. Summary of the Invention
[0005] Embodiments of the present invention provide a data reduction method, device, medium, and training system for distributed training, so as to provide a new data reduction method for distributed training scenarios and reduce the control overhead of data reduction operations.
[0006] In a first aspect, an embodiment of the present invention provides a data reduction method for distributed training, which is executed by an atomic operation module disposed close to an external memory in a multi-core chip, comprising:
[0007] Whenever a data reduction instruction is received from an on-chip computing core or an off-chip computing core via a DMA (Direct Memory Access) module, the data reduction instruction is stored in the instruction cache;
[0008] Each data specification instruction is read from the instruction cache area respectively, and according to the data specification description information included in each data specification instruction, the data read and write atomic operation for the external memory is executed.
[0009] In a second aspect, an embodiment of the present invention provides a data reduction device for distributed training, which is executed by an atomic operation module disposed close to an external memory in a multi-core chip, including:
[0010] A data protocol instruction storage module is used to store the data protocol instruction in the instruction cache whenever receiving a data protocol instruction sent by the on-chip computing core or the off-chip computing core through the DMA module;
[0011] The data reduction instruction atomic execution module is used to read each data reduction instruction from the instruction cache area and execute data reading and writing atomic operations on the external memory according to the data reduction description information included in each data reduction instruction.
[0012] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data reduction method in distributed training described in any embodiment of the present invention when executed.
[0013] In a fourth aspect, an embodiment of the present invention provides a distributed training system, the system comprising:
[0014] At least one multi-core chip, each multi-core chip comprising: multiple computing cores, an external memory, and an atomic operation module disposed close to the external memory;
[0015] Each computing core includes a computing unit, internal memory, and a DMA module. The DMA modules in the computing cores of the same multi-core chip communicate with the atomic operation module on the chip through the on-chip bus, while the DMA modules in the computing cores of different multi-core chips communicate with the atomic operation module off-chip through the off-chip bus.
[0016] The computing core is used to send data reduction instructions to the atomic operation module on the chip or the atomic operation module off the chip through its own DMA module during distributed training;
[0017] The atomic operation module is used to execute the data reduction method in distributed training as described in any embodiment of the present invention.
[0018] The technical solution of the embodiment of the present invention is to execute, through an atomic operation module set close to the external memory in the multi-core chip, whenever a data reduction instruction is received from the on-chip computing core or the off-chip computing core through the DMA module, the data reduction instruction is stored in the instruction cache area, and then each data reduction instruction is read from the instruction cache area respectively, and the data reading and writing atomic operations for the external memory are performed according to the read data reduction instructions. This technical means provides a distributed data reduction implementation method that does not require sequential control, while reducing the control overhead of the data reduction operation, avoiding logical errors in the data reduction process, and improving the accuracy and implementation efficiency of the data reduction operation.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1a This is a schematic diagram of the structure of a data reduction system in distributed training in the prior art;
[0022] Figure 1b This is a flow chart of a data reduction method in distributed training provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a flow chart of a data reduction method in distributed training provided according to the second embodiment of the present invention;
[0024] Figure 3 2 is a schematic diagram of the structure of a data reduction device in distributed training provided in accordance with a third embodiment of the present invention;
[0025] Figure 4 It is a structural diagram of an electronic device for implementing a data reduction method in distributed training according to an embodiment of the present invention;
[0026] Figure 5 2 is a structural diagram of a distributed training system provided according to Embodiment 5 of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] To facilitate understanding of the embodiments of the present invention, the distributed data reduction process in the prior art is first briefly described.
[0030] Specifically, in Figure 1a A schematic diagram of the structure of a data reduction system in a distributed training in the prior art is shown in FIG. Figure 1a As shown, an artificial intelligence chip includes two computing cores, computing core 1 and computing core 2. Computing core 1 and computing core 2 independently execute the set computing tasks in distributed training. After computing core 1 and computing core 2 complete their respective computing tasks, they need to write the calculation results to the same storage address of the chip memory, that is, perform data reduction operations. In the prior art, when computing core 1 and computing core 2 simultaneously perform a data reduction operation of adding 1 (the calculation result of the computing core) to the data a at the same storage address in the memory, without sequential control, if the initial value of data a is 1, the updated value of data a obtained after the data reduction operation is completed will be 2, which is different from the expected value 3, that is, a logical error has occurred in the data reduction operation.
[0031] That is, when computing core 1 and computing core 2 read data a, they both read the initial value 1 of a. After computing core 1 and computing core 2 respectively perform the +1 operation on the initial value 1, they will rewrite the result of a=2 into the memory twice. In other words, if the sequence control between the computing cores is not performed during the distributed data reduction process, logical errors will occur. In order to avoid the occurrence of the above logical errors, sequential control is required between computing core 1 and computing core 2. For example, computing core 1 and computing core 2 are required to update data a in a serial order. The above sequential control will cause a large control overhead and will also affect the performance of the computing core.
[0032] Example 1
[0033] Figure 1b This is a flowchart of a data reduction method in distributed training provided in the first embodiment of the present invention. This embodiment is applicable to the case where multiple computing cores jointly perform data reduction operations on data at the same storage address in a distributed training scenario. This method can be executed by a data reduction device in distributed training. The data reduction device in distributed training can be implemented in the form of hardware and / or software. The data reduction device in distributed training can be configured in an atomic operation module close to the external memory in a multi-core chip. Figure 1b As shown, the method includes:
[0034] S110 , whenever a data reduction instruction is received from an on-chip computing core or an off-chip computing core via a DMA module, the data reduction instruction is stored in an instruction cache.
[0035] As mentioned above, the methods of each embodiment of the present invention are executed by an atomic operation module set close to the external memory in a multi-core chip. That is, for a multi-core chip containing multiple computing cores and an external memory, an additional hardware structure needs to be set up. The hardware structure is an atomic operation module set close to the external memory. The atomic operation module is designed by using various logic gate circuits. The atomic operation module can be used to implement the methods described in each embodiment of the present invention.
[0036] Accordingly, when performing data reduction operations in a distributed training scenario, you can use only multiple computing cores in the same multi-core chip to jointly implement the data reduction operation, and store the data reduction results in the external memory of the multi-core chip; or, you can also use multiple computing cores in multiple multi-core chips at the same time to jointly implement the data reduction operation, and store the data reduction results in the external memory of any multi-core chip.
[0037] For the above two optional application scenarios, the atomic operation module in the embodiment of the present invention can not only receive data protocol instructions sent by the computing core on the multi-core chip (on-chip computing core) where it is located, but also receive data protocol instructions sent by the computing cores on other multi-core chips (off-chip computing cores).
[0038] Specifically, the data reduction instruction refers to an instruction executed jointly by multiple on-chip computing cores or off-chip computing cores, which is used to use multiple copies of data to update the data stored in a set storage address in the external memory of any multi-core chip.
[0039] In this embodiment, the data reduction instructions are sent to the atomic operation module by a DMA module configured on an on-chip or off-chip computing core. That is, the number of DMA modules included in a multi-core chip corresponds one-to-one to the number of computing cores included in the multi-core chip. For example, if a multi-core chip includes computing core 1 and computing core 2, a first DMA module is configured on computing core 1, and a second DMA module is configured on computing core 2.
[0040] Among them, the DMA module is a function provided by some computer bus architectures, which enables data to be sent directly from an attached device to the memory of the computer motherboard; further, the DMA module can be programmed by software to perform data transfer, reducing the data transfer overhead of the computing core.
[0041] In this embodiment, to ensure that the distributed data reduction process does not produce logical errors, the atomic operation module is required to execute each data reduction instruction in sequence, that is, each data reduction instruction is executed in the form of an atomic operation. To ensure that each data reduction instruction is executed one by one, an instruction cache area needs to be set in the atomic operation module.
[0042] The instruction cache stores at least one received data reduction instruction in the form of a queue. The queue depth of the instruction cache can be preset based on actual conditions and is generally determined by the throughput requirements of atomic data read and write operations. Specifically, when caching data reduction instructions, the instruction cache stores the data reduction instructions in the queue in the order in which they were received.
[0043] In other words, in this embodiment, whenever the atomic operation module receives a data reduction instruction, it does not directly execute the data reduction instruction, but instead stores the above-mentioned data reduction instructions in an instruction cache area in the form of a queue in sequence to sort the data reduction instructions and prepare for the subsequent execution of each data reduction instruction in the form of an atomic operation.
[0044] S120 , reading each data reduction instruction from the instruction cache, and executing data read and write atomic operations on the external memory according to the data reduction description information included in each data reduction instruction.
[0045] In this embodiment, the atomic operation module stores each data reduction instruction in the instruction cache in sequence according to the time of receiving the data reduction instruction. That is, the first received data reduction instruction is stored at the front of the queue in the instruction cache.
[0046] Accordingly, when the atomic operation module reads each data reduction instruction from the instruction cache, it performs the read operation in queue order, i.e., the earliest stored data reduction instruction is read and processed first. Furthermore, after executing the data read and write atomic operation on the external memory according to the data reduction description information included in the data reduction instruction, the executed data reduction instruction is deleted from the queue of the instruction cache.
[0047] Among them, the data reduction description information may include: data address, reduction operand and reduction logic. Specifically, the data address defines the storage address of the data required to be updated by the data reduction operation in the set external memory. Among them, the set external memory refers to the external memory in one or more multi-core chips used in the distributed training process. The reduction operand defines the operand that needs to be calculated with the data to be updated in the external memory when performing the data reduction operation. The reduction logic defines the specific data update logic performed by the reduction operand on the data to be updated, such as addition, subtraction, bitwise AND operation, taking the larger value or smaller value between the reduction operand and the data to be updated in the external memory, and other simple data calculation logic.
[0048] For example, in this embodiment, assume that a computing core needs to perform an addition operation on the data at address 0130H in external memory a on multi-core chip A. The data reduction instruction A generated by the computing core includes the following data address: external memory a: 0130H, the reduction operand: 1, and the reduction logic: addition operation.
[0049] Among them, the data read and write atomic operations can be understood as the operations performed by the atomic operation module on the external memory set closely for each data reduction instruction. That is, data is obtained from the data address in the external memory, and after the data and the reduction operand are processed according to the reduction logic to obtain the updated data, the updated data is written back to the data address of the external memory. If there are still data read and write atomic operations on the same data address in the instruction cache, the current updated data can be saved in the atomic operation module and not written back to the external memory for the time being, so as to optimize the atomic operation performance.
[0050] In this embodiment, the atomic operation is an operation that will not be interrupted by the thread scheduling mechanism, that is, once the atomic operation module starts to execute a data read or write operation for a certain data specification instruction, it will continue to run until the data read or write operation is completed, and will not be switched to other threads during the execution of the data read or write operation.
[0051] In an optional implementation of this embodiment, reading each data reduction instruction from the instruction cache and performing the data read and write atomic operation on the external memory according to the data reduction description information included in each data reduction instruction may include:
[0052] A first data reduction instruction is read from an instruction cache, and a first data address, a first reduction operand, and a first reduction logic are extracted from the first data reduction instruction; first external memory data matching the first data address is read from an external memory, and the first external memory data and the first reduction operand are processed according to the first reduction logic to obtain a first processing result; and the first processing result is written back to the first data address.
[0053] In a specific example, assuming that the atomic operation module disposed near the external memory B sequentially reads the first data protocol instruction from the queue of the instruction cache in sequence, the first data address included in the instruction is external memory B:1032H, the first protocol operand is 50, and the first protocol logic is an addition operation, then the data read and write operations performed by the atomic operation module for the first data protocol instruction specifically include:
[0054] The first external memory data, for example, 8, is read from the address 1032H of the external memory B. After calculating 8+50=58, 58 is rewritten into the address 1032H of the external memory B.
[0055] Optionally, the first data address corresponds to a continuous address range or a discrete address range; and / or the first reduced operand is integer data or floating-point data. The continuous address range may include multiple continuous addresses. For example, the continuous address range may include nine addresses, 1-8 bits, of an external memory. Furthermore, the discrete address range may include one or more discontinuous addresses. For example, the discrete address range may include three addresses, 1, 3, and 5 bits, of an external memory.
[0056] In this embodiment, the on-chip computing core or the off-chip computing core sends a data reduction instruction through the DMA module. Since the DMA bus protocol for sending data by the DMA module is flexible and configurable, the instruction form that can support the first data reduction instruction is also more diverse.
[0057] Specifically, if the first data address corresponds to a continuous address range, the first and last addresses can be directly specified in the first data reduction instruction, and the reduction operands corresponding to each address in the address space defined by the first and last addresses can be sequentially specified. Accordingly, the instruction format of the first data reduction instruction can be: start addr, end addr, data1, data2, ...
[0058] If the first data address corresponds to a discrete address range, each discrete address can be specified in sequence in the first data instruction, and the protocol operand corresponding to each discrete address can be specified. The instruction format of the first data protocol instruction can be: addr1, data1, addr2, data2, ...
[0059] Of course, it is understandable that the first data address may also correspond to both a continuous address range and a discrete address range, and this embodiment does not limit this.
[0060] The integer data may be numerical data that does not include a decimal part, and the floating-point data may be numerical data that has an integer part and a decimal part.
[0061] Similarly, because the DMA bus protocol supports a larger number of configurable data bits, the first reduction operand can be not only integer data but also floating-point data with a wider data bit width. Furthermore, the technical solutions of the various embodiments of the present invention are applicable to a wider range of distributed training scenarios and are more versatile.
[0062] It should be noted that in this embodiment, the DMA modules in different computing cores operate independently, and no software coordination of execution order is required. In this embodiment, because the execution order of different data protocol instructions is specified in the atomic operation module on each multi-core chip, there is no need to set the execution order for DMA modules in different computing cores as in the prior art. As a result, the DMA modules in different computing cores can operate independently of each other, eliminating the need for complex software control logic to coordinate the execution order of each computing core.
[0063] The technical solution of the embodiment of the present invention is to execute, through an atomic operation module set close to the external memory in the multi-core chip, whenever a data reduction instruction is received from the on-chip computing core or the off-chip computing core through the DMA module, the data reduction instruction is stored in the instruction cache area, and then each data reduction instruction is read from the instruction cache area respectively, and the data reading and writing atomic operations for the external memory are performed according to the read data reduction instructions. This technical means provides a distributed data reduction implementation method that does not require sequential control, while reducing the control overhead of the data reduction operation, avoiding logical errors in the data reduction process, and improving the accuracy and implementation efficiency of the data reduction operation.
[0064] Example 2
[0065] Figure 2 A flowchart of a data reduction method in distributed training provided in the second embodiment of the present invention is provided. This embodiment is a refinement of the above embodiments. In this embodiment, each data reduction instruction is read from the instruction cache area respectively, and according to the data reduction description information included in each data reduction instruction, the steps of performing data read and write atomic operations on the external memory are specifically as follows: reading a second data reduction instruction from the instruction cache area, and extracting a second data address, a second reduction operand and a second reduction logic from the second data reduction instruction; reading second external memory data matching the second data address from the external memory, and processing the second external memory data and the second reduction operand according to the second reduction logic to obtain a second processing result; detecting whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache area; if so, executing each associated data reduction instruction to update the second processing result at least once, and then writing the updated second processing result back to the second data address.
[0066] Correspondingly, such as Figure 2 As shown, the method may specifically include:
[0067] S210. Whenever a data instruction is received from an on-chip computing core or an off-chip computing core via a DMA module, detect whether the data instruction includes an atomic operation identifier: if so, execute S220; if not, execute S230.
[0068] The data instruction may be any data instruction that can be generated by the computing core to read and write data to the external memory.
[0069] It is understandable that the DMA module itself has the regular data reading and writing functions for the external memory. The computing core can perform regular data reading and writing operations on the external memory based on the configured DMA module, and does not need the atomic operation module to execute it on its behalf. This will reduce the execution efficiency of the atomic operation module and greatly weaken the processing function of the DMA module itself.
[0070] Accordingly, in this embodiment, the data instructions issued by the DMA module are distinguished. If the above data instructions are data reduction instructions, the data reduction instructions are processed by the atomic operation module; if some of the above data instructions are not data reduction instructions, the external memory can directly respond to the data instructions to reduce the execution pressure of the atomic operation module and improve the execution efficiency of the atomic operation module.
[0071] Optionally, whether a data instruction is a data reduction instruction can be identified by inserting or not inserting an atomic operation identifier into the data instruction.
[0072] In this embodiment, based on the DMA bus protocol, an atomic operation identification bit can be configured in the data instruction sent by DMA. If the atomic operation identification bit in a data instruction is 1, the data instruction can be directly identified as a data reduction instruction.
[0073] S220 , storing the received data instruction as a data protocol instruction in the instruction cache, and executing S240 .
[0074] The data instruction is a data instruction with an atomic operation identifier, that is, a data specification instruction that requires an atomic operation.
[0075] S230: If not, directly forward the received data instruction to the external memory.
[0076] When the data instruction does not include an atomic operation identifier, the atomic operation module can directly transmit the data instruction to a nearby external memory, and the external memory can directly respond to the data instruction.
[0077] S240: Read a second data reduction instruction from the instruction cache, and extract a second data address, a second reduction operand, and a second reduction logic from the second data reduction instruction.
[0078] The second data address may correspond to a continuous address range or a discrete address range, and the second reduction operand may be integer data or floating-point data.
[0079] S250: Read second external memory data matching the second data address from the external memory, and process the second external memory data and the second reduction operand according to the second reduction logic to obtain a second processing result.
[0080] For example, assuming that the atomic operation module disposed near the external memory B sequentially reads the second data reduction instruction b from the queue of the instruction cache in sequence, the second data address included in the second data reduction instruction b is the external memory B:1032H, the first reduction operand is 50, and the second reduction logic is an addition operation, then the data read and write operations performed by the atomic operation module for the second data reduction instruction b specifically include:
[0081] The second external memory data, for example, 8, is read from the address 1032H of the external memory B; and 8+50=58 is calculated as the current second processing result.
[0082] S260: Detect whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache.
[0083] Illustratively, in this embodiment, by checking data addresses in other data reduction instructions currently stored in the instruction cache, it can be detected whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache.
[0084] Continuing with the previous example, the atomic operation module reads the data addresses included in each data reduction instruction currently stored in the instruction cache. If it is determined that the data address included in the data reduction instruction c is also the external memory B:1032H, which is the same as the second data address included in the second data reduction instruction b in the previous example, then the data reduction instruction c can be used as an associated data reduction instruction of the second data reduction instruction b.
[0085] S270 , after executing each associated data reduction instruction to update the second processing result at least once, write the updated second processing result back to the second data address.
[0086] Specifically, executing each associated data reduction instruction to update the second processing result at least once may include: updating the second processing result in sequence according to the cache order of the associated data reduction instructions in the instruction cache queue.
[0087] For example, based on S260, the reduction operand of the associated data reduction instruction detected in S260 is 100, the reduction logic is an addition operation, and the current second processing result obtained in step S250 is 58, then according to the reduction logic, the second processing result is updated according to the associated data reduction instruction, and the updated second processing result is 58+100=158;
[0088] Furthermore, after all associated data reduction instructions in the instruction cache area update the second processing result in accordance with the cache order, the final update result 158 can be written and overwritten to the original second external memory data 8 at the 1032H address of the external memory B as the new second external memory data of the second data address.
[0089] It should be noted that there may be more than one associated data reduction instruction matching the second data address in the instruction cache. In this case, the associated data instruction updates the second processing result in the order of storage in the queue. It is easy to understand that the number of updates of the second processing result is the same as the number of associated data reduction instructions in the instruction cache. After completing all the updates of the second processing result, the final updated second processing result is written into and overwrites the original second external memory data in the second data address as the new second external memory data of the current second data address.
[0090] The embodiment of the present invention reduces the number of times the reduction result is written to the external memory by detecting whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache, and updating and outputting the second processing result based on the result. While improving the reduction calculation speed, it avoids data errors and logical errors caused by repeated writing.
[0091] The technical solution of the embodiment of the present invention is executed by an atomic operation module set close to the external memory in the multi-core chip. Whenever the atomic operation module receives a data reduction instruction sent by the on-chip computing core or the off-chip computing core through the DMA module, the data reduction instruction is stored in the instruction cache area, and then a second data reduction instruction is read from the instruction cache area, and the second external memory data matching the second data address is read from the external memory. According to the second data reduction instruction, the second external memory data and the second reduction operand are processed to obtain a second processing result. Finally, it is detected whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache area, and the second processing result is updated and output according to the result, thereby reducing the number of times the reduction result is written to the external memory and the control overhead of the data reduction, while avoiding logical errors generated in the process of data reduction, improving the accuracy of the data reduction, and improving the working efficiency of the data reduction.
[0092] Example 3
[0093] Figure 3 This is a schematic diagram of the structure of a data reduction device in distributed training provided by the third embodiment of the present invention. Figure 3 As shown, the device includes:
[0094] The data protocol instruction storage module 310 is configured to store the data protocol instruction in the instruction cache whenever receiving the data protocol instruction sent by the on-chip computing core or the off-chip computing core via the DMA module;
[0095] The data reduction instruction atomic execution module 320 is used to read each data reduction instruction from the instruction cache, and execute data read and write atomic operations on the external memory according to the data reduction description information included in each data reduction instruction.
[0096] The technical solution of the embodiment of the present invention is to execute, through an atomic operation module set close to the external memory in the multi-core chip, whenever a data reduction instruction is received from the on-chip computing core or the off-chip computing core through the DMA module, the data reduction instruction is stored in the instruction cache area, and then each data reduction instruction is read from the instruction cache area respectively, and the data reading and writing atomic operations for the external memory are performed according to the read data reduction instructions. This technical means provides a distributed data reduction implementation method that does not require sequential control, while reducing the control overhead of the data reduction operation, avoiding logical errors in the data reduction process, and improving the accuracy and implementation efficiency of the data reduction operation.
[0097] Based on the above embodiment, the data reduction instruction atomic execution module 320 may include:
[0098] A first data reading unit is configured to read a first data reduction instruction from an instruction cache, and extract a first data address, a first reduction operand, and a first reduction logic from the first data reduction instruction;
[0099] a first processing result obtaining unit, configured to read first external memory data matching the first data address from the external memory, and process the first external memory data and the first reduction operand according to the first reduction logic to obtain a first processing result;
[0100] The first rewriting unit is used to rewrite the first processing result back to the first data address.
[0101] Based on the above embodiment, the data reduction instruction atomic execution module 320 may further include:
[0102] A second data reading unit is used to read a second data protocol instruction from the instruction cache, and extract a second data address, a second protocol operand and a second protocol logic from the second data protocol instruction;
[0103] a second processing result obtaining unit, configured to read second external memory data matching the second data address from the external memory, and process the second external memory data and the second reduction operand according to the second reduction logic to obtain a second processing result;
[0104] a detection unit, configured to detect whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache;
[0105] The second rewriting unit is configured to execute each associated data reduction instruction to update the second processing result at least once, and then rewrite the updated second processing result back to the second data address.
[0106] Based on the above embodiment, the data reduction instruction storage module 310 may include:
[0107] an identification detection unit, configured to detect whether an atomic operation identification is included in a data instruction sent by an on-chip computing core or an off-chip computing core via a DMA module whenever the data instruction is received;
[0108] The storage unit is used to store the received data instruction as a data protocol instruction in the instruction cache.
[0109] Based on the above embodiment, the storage unit further includes:
[0110] The forwarding unit is used to forward the received data instructions directly to the external memory.
[0111] The data reduction device in distributed training provided by the embodiment of the present invention can execute the data reduction method in distributed training provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0112] Example 4
[0113] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0114] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0115] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0116] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data reduction method in distributed training.
[0117] Correspondingly, the method is executed by an atomic operation module set close to the external memory in the multi-core chip, and the method includes: whenever a data protocol instruction is received from the on-chip computing core or the off-chip computing core through the DMA module, the data protocol instruction is stored in the instruction cache area; each data protocol instruction is read from the instruction cache area respectively, and according to the data protocol description information included in each data protocol instruction, the data read and write atomic operation for the external memory is executed.
[0118] In some embodiments, the data reduction method in distributed training can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data reduction method in distributed training described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data reduction method in distributed training by any other appropriate means (for example, by means of firmware).
[0119] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0123] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0124] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0125] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0126] Example 5
[0127] Figure 5 This is a structural diagram of a distributed training system provided by the fifth embodiment of the present invention. Figure 5 As shown, the system includes: at least one multi-core chip 510, each multi-core chip 510 includes: multiple computing cores 5110, an external memory 5140 and an atomic operation module 5130 arranged close to the external memory 5140;
[0128] Each computing core 5110 includes a computing unit 5111, an internal memory 5112, and a DMA module 5113. The DMA module 5113 in each computing core 5110 in the same multi-core chip 510 communicates with the on-chip atomic operation module 5130 via an on-chip bus 5120. The DMA module 5113 in each computing core 5110 in different multi-core chips 510 communicates with the off-chip atomic operation module via an off-chip bus 5150.
[0129] The computing core 5110 is used to send data reduction instructions to the atomic operation module on the chip or the atomic operation module off the chip through its own DMA module 5113 during distributed training;
[0130] The atomic operation module 5130 is used to execute the data reduction method in distributed training described in any embodiment.
[0131] like Figure 5 As shown, in order to achieve distributed training, at least one multi-core chip is required. Figure 5 In this example, two multi-core chips are used. Each multi-core chip includes two computing cores, an external memory, and a DMA module configured in each computing core. The data reduction method in distributed training mainly includes the following operation process:
[0132] 1. The computing core performs the computing task and places the data address of the data to be updated in the external memory, the reduction logic to be executed for the updated data, and the calculated reduction operands in the internal memory, wherein the data address and can support a continuous address range or a discrete address range, and the reduction operands can be integer data or floating-point data.
[0133] 2. The DMA module reads the internal memory to obtain the data address, protocol logic and protocol operand.
[0134] 3. The DMA module constructs a data reduction instruction based on the data address, reduction logic and reduction operand, and sends the data reduction instruction to the atomic operation module corresponding to the external memory matching the data address. The atomic operation module stores the data reduction instruction in the instruction cache.
[0135] 4. When the atomic operation module obtains the data specification instruction from the instruction cache and executes it, it reads the external memory data that matches the data address from the external memory.
[0136] 5. The atomic operation module generates processing results based on the specification logic, external storage data and specification operands;
[0137] 6. If the atomic operation module detects that at least one associated data reduction instruction matching the data address is stored in the instruction cache, it executes each associated data reduction instruction to update the processing result at least once, and then writes the final updated processing result back to the data address.
[0138] 7. If the instruction cache does not store any associated data reduction instruction matching the second data address, the processing result is directly written back to the data address.
[0139] In the above workflow, the DMA modules in different computing cores run independently, and there is no need for software to coordinate the execution order.
[0140] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A data reduction method in distributed training, characterized in that: The method is executed by an atomic operation module disposed close to an external memory in a multi-core chip, and includes: Whenever a data reduction instruction is received from an on-chip computing core or an off-chip computing core via a direct memory access (DMA) module, the data reduction instruction is stored in an instruction cache; Reading each data specification instruction from the instruction cache area respectively, and executing the data read and write atomic operation for the external memory according to the data specification description information included in each data specification instruction; The method comprises: reading each data reduction instruction from the instruction cache, and executing data read and write atomic operations on the external memory according to the data reduction description information included in each data reduction instruction, including: reading a first data reduction instruction from the instruction cache, and extracting a first data address, a first reduction operand, and a first reduction logic from the first data reduction instruction; reading first external memory data matching the first data address from the external memory, and processing the first external memory data and the first reduction operand according to the first reduction logic to obtain a first processing result; and writing the first processing result back to the first data address; Among them, each data reduction instruction is read from the instruction cache area respectively, and according to the data reduction description information included in each data reduction instruction, the data read and write atomic operations for the external memory are performed, including: reading the second data reduction instruction from the instruction cache area, and extracting the second data address, the second reduction operand and the second reduction logic from the second data reduction instruction; reading the second external memory data matching the second data address from the external memory, and processing the second external memory data and the second reduction operand according to the second reduction logic to obtain a second processing result; detecting whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache area; if so, executing each associated data reduction instruction to update the second processing result at least once, and then writing the updated second processing result back to the second data address.
2. The method according to claim 1, characterized in that Whenever a data reduction instruction is received from an on-chip computing core or an off-chip computing core via a DMA module, the data reduction instruction is stored in the instruction cache, including: Whenever a data instruction is received from an on-chip computing core or an off-chip computing core via a DMA module, detecting whether the data instruction includes an atomic operation identifier; If so, the received data instruction is stored in the instruction cache as a data protocol instruction.
3. The method according to claim 2, characterized in that After detecting whether the data instruction includes an atomic operation identifier, the method further includes: If not, the received data instruction is directly forwarded to the external memory.
4. The method according to claim 1, wherein The first data address or the second data address corresponds to a continuous address range or a discrete address range; and / or The first reduction operand or the second reduction operand is integer data or floating-point data.
5. A data reduction device for distributed training, characterized in that: The method is executed by an atomic operation module provided close to an external memory in a multi-core chip, and the device includes: A data protocol instruction storage module is used to store the data protocol instruction in the instruction cache whenever receiving a data protocol instruction sent by the on-chip computing core or the off-chip computing core through the direct memory access (DMA) module; A data reduction instruction atomic execution module is used to read each data reduction instruction from the instruction cache and execute data read and write atomic operations on the external memory according to the data reduction description information included in each data reduction instruction; The data reduction instruction atomic execution module includes: a first data reading unit, configured to read the first data reduction instruction from the instruction cache, and extract the first data address, the first reduction operand, and the first reduction logic from the first data reduction instruction; a first processing result acquisition unit, configured to read first external memory data matching the first data address from the external memory, and process the first external memory data and the first reduction operand according to the first reduction logic to obtain a first processing result; and a first rewriting unit, configured to rewrite the first processing result back to the first data address. Among them, the data reduction instruction atomic execution module also includes: a second data reading unit, used to read the second data reduction instruction from the instruction cache area, and extract the second data address, second reduction operand and second reduction logic from the second data reduction instruction; a second processing result acquisition unit, used to read the second external memory data matching the second data address from the external memory, and process the second external memory data and the second reduction operand according to the second reduction logic to obtain a second processing result; a detection unit, used to detect whether at least one associated data reduction instruction matching the second data address is stored in the instruction cache area; a second rewrite unit, used to execute each associated data reduction instruction respectively to update the second processing result at least once, and then write the updated second processing result back to the second data address.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data reduction method in distributed training according to any one of claims 1 to 4 when executed.
7. A distributed training system, characterized in that: include: At least one multi-core chip, each multi-core chip comprising: multiple computing cores, an external memory, and an atomic operation module disposed close to the external memory; Each computing core includes a computing unit, internal memory, and a direct memory access (DMA) module. The DMA modules in the computing cores of the same multi-core chip communicate with the atomic operation module on the chip through an on-chip bus, while the DMA modules in the computing cores of different multi-core chips communicate with the atomic operation module off-chip through an off-chip bus. The computing core is used to send data reduction instructions to the atomic operation module on the chip or the atomic operation module off the chip through its own DMA module during distributed training; The atomic operation module is used to execute the data reduction method in distributed training according to any one of claims 1 to 4.
8. The distributed training system according to claim 7, wherein: The DMA modules in different computing cores run independently, and there is no need for software to coordinate the execution order.
Citation Information
Patent Citations
Techniques for efficiently performing data reductions in parallel processing units
CN112241290A
High-performance parallel implementation device for K-NN on GPU processor
CN112380003A