Data processing method and related equipment
By sending the initial result data to the data exchange device for accumulation operation when the core computing processing unit completes the subtask, the problem of low data calculation efficiency of multiple core computing processing units is solved, and the effect of improving data calculation efficiency is achieved.
Patent Information
- Application Number
- CN202510035814.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
The data calculation efficiency of multiple core computing processing units needs to be improved. In the prior art, computing and communication are severely isolated, which affects the data calculation efficiency.
When the core computing processing unit completes the subtask, the initial result data is sent to the data exchange device, and a preset accumulation operation is performed to obtain the target result data, and the target result data is fed back to the core computing processing unit.
The data calculation efficiency of the core computing processing unit is improved, allowing the data exchange device to perform further accumulation operations based on the initial result data when the core computing processing unit that has not completed the calculation.
Smart Images

Figure CN119988005A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a data processing method and related equipment. Background Art
[0002] With the continuous advancement of artificial intelligence and big data computing technologies, the demand for high-performance computing is growing. The computing power of a single core computing processing unit can no longer meet the growing computing needs. Therefore, distributed computing technology has emerged to meet this challenge by combining multiple core computing processing units into a large system.
[0003] However, the data computing efficiency of multiple core computing processing units needs to be improved. Summary of the invention
[0004] In view of this, an embodiment of the present application provides a data processing method and related equipment to improve data computing efficiency.
[0005] To achieve the above objectives, the embodiments of the present application provide the following technical solutions.
[0006] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a core computing processing system, wherein the core computing processing system includes a plurality of core computing processing units, and the method includes:
[0007] Each core computing processing unit processes the assigned subtasks of the task to be processed in parallel to obtain initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks and an accumulation task for performing a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process one subtask;
[0008] When any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data;
[0009] The target result data fed back by the data exchange device is obtained as the result data of the task to be processed.
[0010] Optionally, each core computing processing unit is configured with the same virtual address, wherein the virtual address corresponds to different physical addresses in different core computing processing units.
[0011] Optionally, the step of sending the obtained initial result data to the data exchange device when any core computing processing unit completes the corresponding subtask includes:
[0012] When any core computing processing unit completes the corresponding subtask, the obtained initial result data is written into the virtual address;
[0013] Determining whether the virtual address is the same virtual address configured for each core computing processing unit;
[0014] If yes, the obtained initial result data is sent to the data exchange device.
[0015] Optionally, each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate a mapping relationship between a virtual address and a physical address.
[0016] Optionally, after the step of obtaining the target result data fed back by the data exchange device as the result data of the task to be processed, the step further includes:
[0017] Based on the page table, determining a physical address corresponding to the virtual address in each core computing processing unit;
[0018] Based on the physical address, the result data is stored in a corresponding physical address in each core computing processing unit.
[0019] In a second aspect, an embodiment of the present application provides a data processing method, which is applied to a data exchange device, and the method includes:
[0020] receiving initial result data, wherein the initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask;
[0021] Performing a preset accumulation operation on the received initial result data and obtaining the target result data;
[0022] The target result data is fed back to each core computing processing unit.
[0023] In a third aspect, an embodiment of the present application provides a data processing device, which is applied to a core computing processing system, wherein the core computing processing system includes a plurality of core computing processing units, and the device includes:
[0024] The initial result data acquisition module is used to acquire the subtasks of the task to be processed assigned by each core computing processing unit for parallel processing, and obtain the initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks, and an accumulation task for performing a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process one subtask;
[0025] An initial result data sending module is used to send the obtained initial result data to the data exchange device when any core computing processing unit completes the corresponding subtask, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data;
[0026] A target result data acquisition module, used to acquire the target result data fed back by the data exchange device as the result data of the task to be processed;
[0027] Among them, each core computing processing unit is configured with the same virtual address, and in different core computing processing units, the virtual address corresponds to different physical addresses; each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate the mapping relationship between the virtual address and the physical address.
[0028] Optionally, also include:
[0029] A determination module, configured to determine, based on the page table, a physical address corresponding to the virtual address in each core computing processing unit;
[0030] The storage module is used to store the result data to the corresponding physical address in each core computing processing unit based on the physical address.
[0031] In a fourth aspect, an embodiment of the present application provides a data processing device, which is applied to a data exchange device, and the device includes:
[0032] A receiving module, used for receiving initial result data, wherein the initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask;
[0033] A calculation module is used to perform a preset accumulation operation on the received initial result data and obtain target result data;
[0034] The feedback module is used to feed back the target result data to each core computing processing unit.
[0035] In a fifth aspect, an embodiment of the present application provides an electronic device comprising at least one memory and at least one processor, wherein the memory stores one or more computer executable instructions, and the processor calls the one or more computer executable instructions to execute the data processing method as described in the first aspect above, or executes the data processing method as described in the second aspect above.
[0036] In a sixth aspect, an embodiment of the present application provides a storage medium, which stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the data processing method described in the first aspect above is implemented, or the data processing method described in the second aspect above is implemented.
[0037] In the seventh aspect, an embodiment of the present application provides a computer program product, comprising one or more computer executable instructions, which, when executed, implement the data processing method as described in the first aspect above, or implement the data processing method as described in the second aspect above.
[0038] An embodiment of the present application provides a data processing method and related equipment, wherein the method is applied to a core computing processing system including multiple core computing processing units, comprising: each core computing processing unit processes the subtasks of the assigned task to be processed in parallel to obtain initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks, and an accumulation task that performs a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process a subtask; when any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to a data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains target result data; and the target result data fed back by the data exchange device is obtained as the result data of the task to be processed.
[0039] It can be seen that the data processing method provided in the embodiment of the present application, when any core computing processing unit completes the corresponding subtask, sends the obtained initial result data to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data. In this way, when the core computing processing unit that has not completed the calculation performs the calculation of the corresponding subtask, the data exchange device can perform further preset accumulation operations based on the initial result data that has been received, thereby improving the data calculation efficiency of the core computing processing unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0041] Figure 1 It is a data processing method;
[0042] Figure 2 It is an optional flow chart of the data processing method provided in the embodiment of the present application;
[0043] Figure 3 It is an optional structural diagram of the interaction between the core computing processing system and the data exchange device provided in the embodiment of the present application;
[0044] Figure 4 is an optional flow chart of step S200 provided in an embodiment of the present application;
[0045] Figure 5 is another optional flow chart of the data processing method provided in the embodiment of the present application;
[0046] Figure 6 It is a schematic diagram of the data processing flow of the core computing processing unit provided in the embodiment of the present application;
[0047] Figure 7 is a schematic diagram of an optional structure of a GPU provided in an embodiment of the present application;
[0048] Figure 8 is an optional structural diagram of a data processing device provided in an embodiment of the present application;
[0049] Fig. 9 is another optional structural diagram of the data processing device provided in the embodiment of the present application;
[0050] Fig.10 It is an optional block diagram of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0052] As described in the background technology, with the continuous advancement of artificial intelligence and big data computing technologies, the demand for high-performance computing is growing. The computing power of a single core computing processing unit can no longer meet the growing computing needs. Therefore, distributed computing technology has emerged, which combines multiple core computing processing units into a large system to meet this challenge. However, the data computing efficiency of multiple core computing processing units needs to be improved.
[0053] The inventor analyzed that, since the data communication between core computing processing units is currently based on a communication library (such as the NCCL communication library) centered on the core computing processing unit (such as the GPU), it is necessary to write programs according to the paradigm of sequential execution of computing and communication. This means that in the program, the order of computing tasks (i.e., the calculations performed on the core computing processing unit) and communication tasks (i.e., the transmission of data between core computing processing units through the communication library) needs to be clearly specified. For example, a core computing processing unit first executes a computing task to generate data, and then transmits the generated data to another core computing processing unit through the communication library. After confirming that the data has been received, the other core computing processing unit executes the computing task of the core computing processing unit. This method of isolating computing and communication seriously affects the data computing efficiency of multiple core computing processing units.
[0054] For easier understanding, refer to Figure 1 A data processing method shown in FIG. Figure 1 As shown, a core computing processing unit GPU 0 first performs a computing task (such as Figure 1 The producer kernel in Figure 2 generates data and then transmits the generated data to another core computing processing unit GPU 1 through the NCCL communication library. After the other core computing processing unit GPU 1 confirms that the data has been received, it starts to execute the computing task of the core computing processing unit (as shown in Figure 2). Figure 1 The producer-consumer model is implemented as shown in the consumer kernel in Figure 1. GPU 0 is the producer responsible for generating data, while GPU 1 is the consumer responsible for receiving data and performing calculations.
[0055] It can be seen that this paradigm of executing in the order of computing and communication isolates computing and communication. Specifically, when the core computing processing unit is computing, the communication between the core computing processing units cannot be carried out, and when the core computing processing units are communicating, the core computing processing units cannot perform computing. For example, each core computing processing unit starts to perform its own computing task to generate corresponding data, but because the communication between the core computing processing units cannot be carried out when the core computing processing units are calculating, it is necessary to wait for the slowest core computing processing unit to complete the calculation before the communication between the core computing processing units can be carried out. When all core computing processing units have completed the calculation, data communication between the core computing processing units can be carried out through the communication library. In the process of communication based on the communication library, data will be transmitted from each core computing processing unit to the central node (for example, one of the core computing processing units), and then further calculations will be performed. This way of isolating computing and communication seriously affects the data computing efficiency of multiple core computing processing units.
[0056] In view of this, an embodiment of the present application provides a data processing method and related equipment, wherein the method is applied to a core computing processing system including multiple core computing processing units, comprising: each core computing processing unit processes the assigned subtasks of the task to be processed in parallel to obtain initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks, and an accumulation task that performs a preset accumulation operation on the initial result data of multiple subtasks; a core computing processing unit is used to process a subtask; when any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to a data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains target result data; and the target result data fed back by the data exchange device is obtained as the result data of the task to be processed.
[0057] It can be seen that the data processing method provided in the embodiment of the present application, when any core computing processing unit completes the corresponding subtask, sends the obtained initial result data to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data. In this way, when the core computing processing unit that has not completed the calculation performs the calculation of the corresponding subtask, the data exchange device can perform further preset accumulation operations based on the initial result data that has been received, thereby improving the data calculation efficiency of the core computing processing unit.
[0058] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0059] refer to Figure 2 , Figure 2 1 is an optional flow chart of the data processing method provided in the embodiment of the present application. Figure 2 As shown, the data processing method may include the following steps:
[0060] Step S100 , each core computing processing unit in the core computing processing system processes the assigned subtasks of the task to be processed in parallel to obtain initial result data corresponding to the subtasks.
[0061] Each core computing processing unit is configured with the same virtual address, wherein the virtual address corresponds to different physical addresses in different core computing processing units. Accordingly, each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate the mapping relationship between the virtual address and the physical address. Each core computing processing unit is configured with the same virtual address so that the initial result data obtained by each subsequent core computing processing unit can be processed centrally.
[0062] The task to be processed includes a plurality of subtasks, and an accumulation task for performing a preset accumulation operation on the initial result data of the plurality of subtasks. The accumulation operation can be understood as a reduction operation for generating a single value from a set of data. Specifically, the operation type of the accumulation operation may include, but is not limited to, finding the maximum value, finding the minimum value, and finding the sum.
[0063] In the embodiment of the present application, the initial result data can be directly processed on the data exchange device through the accumulation operation without transmitting all the data back to the core computing processing unit for processing. This reduces the delay and bandwidth occupation of data transmission and improves the overall computing efficiency. In addition, the accumulation operation can be performed in parallel on the data exchange device, making full use of the computing resources of the device.
[0064] refer to Figure 3 , Figure 3 Schematic diagram of an optional structure of the interaction between the core computing processing system and the data exchange device provided in the embodiment of the present application. Figure 3 As shown, the core computing processing system includes multiple core computing processing units (such as Figure 3 In the figure, (as shown in GPU 0, GPU 1 ... GPU n), each core computing processing unit can process subtasks of the assigned task to be processed in parallel, wherein one core computing processing unit is used to process one subtask.
[0065] In some optional embodiments, the core computing processing unit may also be a GPU accelerator card (also referred to as a GPU accelerator or GPU card). A GPU accelerator card is a hardware device built on a powerful GPU chip, which can be inserted into a PCIe slot of a computer to accelerate various computing-intensive tasks, such as graphics rendering, scientific simulation, deep learning training, data mining, and big data analysis.
[0066] refer to Figure 2 , continue to execute step S200, when any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to the data exchange device.
[0067] When any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data.
[0068] like Figure 3 As shown, among GPU 0, GPU 1, ..., GPU n, if GPU 0 completes the corresponding subtask, GPU 0 may first send the initial result data obtained to the data exchange device (such as Figure 3 ); if GPU 1 completes the corresponding subtask, GPU 1 can first send the initial result data it obtains to the data exchange device, and so on. That is to say, as long as the core computing processing unit completes the corresponding subtask, it can send the initial result data it obtains to the data exchange device without waiting for other core computing processing units that have not completed the calculation. Therefore, when the core computing processing unit that has not completed the calculation performs the calculation of the corresponding subtask, the data exchange device can perform further preset accumulation operations based on the initial result data that has been received, thereby improving the data calculation efficiency of the core computing processing unit.
[0069] Among them, the data exchange device Switch (SHARP enabled) in the embodiment of the present application is a switch that can support SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) technology. SHARP is a network communication algorithm designed to optimize and accelerate cluster communications in high-performance computing and large-scale data center environments. It reduces the amount of data that needs to be transmitted between nodes by implementing data aggregation and reduction operations at the network hardware level (such as switches), thereby reducing communication delays and network congestion.
[0070] In addition, SHARP can support multiple collective communication operations, such as Allreduce, Allgather, Broadcast, etc., and is widely used in scenarios that require complex collective communication modes.
[0071] For further reference, Figure 4 , Figure 4 is an optional flow chart of step S200 provided in the embodiment of the present application. Figure 4 As shown, step S200, when any core computing processing unit completes the corresponding subtask, the step of sending the obtained initial result data to the data exchange device may include:
[0072] Step S201, when any core computing processing unit completes the corresponding subtask, the obtained initial result data is written into the virtual address.
[0073] Step S202, determining whether the virtual address is the same virtual address configured for each core computing processing unit.
[0074] If yes, execute step S203 to send the obtained initial result data to the data exchange device.
[0075] refer to Figure 2 , continue to execute step S300, the data exchange device receives the initial result data.
[0076] The initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask.
[0077] After receiving the initial result data, the data exchange device can perform further preset accumulation operations based on the received initial result data, thereby improving the data calculation efficiency of the core calculation processing unit.
[0078] Step S400: the data exchange device performs a preset accumulation operation on the received initial result data and obtains target result data.
[0079] The accumulation operation can be understood as a reduction operation, which is used to generate a single value from a set of data. Specifically, the operation type of the accumulation operation may include, but is not limited to, finding the maximum value, finding the minimum value, and summing. In practice, a suitable accumulation operation can be selected according to the needs, and the embodiments of the present application are not limited to this.
[0080] The data exchange device adopts a "streaming processing" method, that is, it gradually receives the initial result data and gradually performs the preset accumulation operation, rather than waiting for all the initial result data to arrive before starting processing. In this way, even if the calculation progress of each core computing processing unit is not exactly the same, the data exchange device can start processing the initial result data that has arrived, so that while the core computing processing unit that has not completed the calculation is performing the calculation of the corresponding subtask, the data exchange device can perform the preset accumulation operation on the received initial result data, thereby improving the data calculation efficiency of the core computing processing unit.
[0081] Step S500: the data exchange device feeds back the target result data to each core computing processing unit.
[0082] The data exchange device can feed back the target result data to each core computing processing unit through a high-speed network interface, direct memory access (DMA) or other data transmission mechanisms.
[0083] Among them, the high-speed network interface is an important component in the data exchange device for connecting to the network and transmitting data. It can support high-speed data transmission and ensure that data can flow quickly between the data exchange device and each core computing processing unit. Through the high-speed network interface, the data exchange device can send the target result data to each core computing processing unit in the form of a network data packet.
[0084] Direct memory access (DMA) is a data transfer method that allows peripherals to directly access system memory without CPU intervention. In data exchange devices, DMA technology can be used to transfer target result data from peripherals (such as storage devices, network interface cards, etc.) directly to the memory of each core computing processing unit, or from memory to peripherals.
[0085] Step S600: the core computing processing system obtains the target result data fed back by the data exchange device as the result data of the task to be processed.
[0086] Each core computing processing unit in the core computing processing system obtains the target result data fed back by the data exchange device, and uses the target result as the result data of the task to be processed.
[0087] It can be seen that the data processing method provided in the embodiment of the present application, when any core computing processing unit completes the corresponding subtask, sends the obtained initial result data to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data. In this way, when the core computing processing unit that has not completed the calculation performs the calculation of the corresponding subtask, the data exchange device can perform further preset accumulation operations based on the initial result data that has been received, thereby improving the data calculation efficiency of the core computing processing unit.
[0088] For further reference, Figure 5 , Figure 5 FIG. 1 is another optional flow chart of the data processing method provided in the embodiment of the present application. Figure 5 As shown, after the step of obtaining the target result data fed back by the data exchange device as the result data of the task to be processed, the following may also be included:
[0089] Step S700: Based on the page table, determine the physical address corresponding to the virtual address in each core computing processing unit.
[0090] The page table is used to indicate the mapping relationship between the virtual address and the physical address. Based on the page table in each core computing processing unit, the physical address corresponding to the virtual address in each core computing processing unit can be determined.
[0091] Step S800: based on the physical address, storing the result data to the corresponding physical address in each core computing processing unit.
[0092] It is understandable that, in the embodiment of the present application, all core computing processing units are set to use the same virtual address, but at the physical level, each core computing processing unit has its own independent physical address and corresponding page table. That is to say, although the virtual addresses of the core computing processing units are the same, the result data is actually stored in the local video memory of each core computing processing unit.
[0093] Take the GPU as an example, refer to Figure 6 The schematic diagram of the data processing flow of the core computing processing unit is shown as an example. Figure 6As shown, all GPUs use the same virtual address (also called collective address). When the GPU stores data to the same virtual address, the GPU will recognize that this is a collective address and bypass the path directly stored in the local HBM (high bandwidth memory). Send the data to a switch that supports the SHARP function. The switch performs corresponding calculations on the data based on the recorded communication operation information (the switch will record the communication operation information corresponding to the virtual address when initialized). For example, the switch performs a sum operation on data from different GPUs. After the calculation is completed, the calculated target result data is forwarded back to all GPUs. Each GPU can store the target result data in the correct location of the local video memory according to its own page table.
[0094] like Figure 6 As shown, GPU 0 stores the target result data at the GPU 0 physical address corresponding to the virtual address, GPU 1 stores the target result data at the GPU 1 physical address corresponding to the virtual address, ..., and so on, GPU n stores the target result data at the GPU n physical address corresponding to the virtual address.
[0095] In a specific implementation, this solution can be implemented by combining an API (Application Programming Interface), as follows:
[0096] The API used to apply for a virtual address can be:
[0097] CMalloc(addr, size, comm, rankid, opdescp)
[0098] Wherein, addr is the virtual address in the collective memory (CM) allocated by CMalloc; size is the capacity of the virtual address; comm is the communicator (the communicator in the embodiment of the present application can be compatible with nccl and MPI); rankid is the number of the current core computing processing unit in comm; opdescp is the communication operator corresponding to the virtual address. The communication operator may include, for example, Reduce, Allreduce, Reduce_scatter, Allgather, Broadcast, Scatter, and P2P.
[0099] Among them, Reduce is a parallel computing operation used to reduce a set of data, such as sum, maximum value (max), minimum value (min), etc. In the Reduce operation, only the root rank will allocate enough video memory to store the final result.
[0100] Allreduce is a parallel computing operation used to aggregate data between multiple computing nodes. The Allreduce operation adds the data on each node between all nodes or performs other types of reduction operations (such as sum, average, maximum, etc.) to ensure that each node will get the same result in the end. When executing the Allreduce operation, since each node needs to receive data from other nodes, reduce the data, and then send the results to other nodes, each node will be allocated the same amount of video memory to ensure that it can store and process the required data.
[0101] Reduce_scatter is a data distribution operation that is used to distribute a large data set to multiple computing nodes for parallel computing and reduce the data on each node (such as summing, averaging, etc.). When executing the Reduce_scatter operation, each computing node (or rank) will allocate video memory based on the ratio of the communication scale (comm size) and the total data volume (size).
[0102] Allgather is a data communication operation that collects data scattered across multiple computing nodes and ensures that each node can get a complete set of data. When performing the Allgather operation, each computing node (or rank) is allocated the same size of video memory (size) so that it can receive and store data from all other nodes.
[0103] Broadcast is a data communication operation used to copy data from one node (usually the master node or root node) to all other nodes. When performing a Broadcast operation, each computing node (or rank) is allocated the same size of video memory to receive and store the sent data.
[0104] Scatter is a data distribution operation used to distribute data from one node (usually the master node or root node, also known as the root rank) to multiple other nodes. When performing a Scatter operation, the size of the video memory allocated to each slave node is the quotient of the total data volume (size) divided by the communication size (comm size) (i.e., size / (commsize)). Among them, "comm size" refers to the number of nodes participating in the Scatter operation.
[0105] P2P is a data exchange operation that allows different nodes (or ranks) to exchange data directly. When performing a P2P operation, the source node (src rank) will allocate enough memory (size) to store the data to be sent. The destination node (dst rank) will also allocate enough memory (size) to receive and store the data sent from the source node.
[0106] When allocating video memory, CMalloc will synchronize the corresponding communication operation information to the data exchange device, and complete a global synchronization within the comm range when the API returns (for P2P type communication operations, only the src rank and dst rank will be synchronized).
[0107] The API used to confirm that the communication operation corresponding to the virtual address has been completed can be:
[0108] WaitCMemory(addr, stream)
[0109] Among them, addr is the virtual address in the collective memory (CM) allocated by CMalloc, and stream is the queue of the core computing processing unit (such as GPU).
[0110] The API used to convert memory can be:
[0111] DisableCMemory(addr, stream)
[0112] Where addr is the virtual address in the collective memory (CM) allocated by CMalloc, and stream is the queue of the core computing processing unit (such as GPU). This API can convert the collective memory area located at addr into regular memory. After the conversion, the memory will follow the regular storage rules and can be accessed by other parts of the system or computing units in a regular manner. The original page table remains unchanged, but store forwarding will not be triggered. Store forwarding refers to the forward propagation of unfinished write operations to other caches or memory locations. This API will complete a global synchronization within the comm range when returning.
[0113] The API used to release video memory can be:
[0114] freeCMemory(addr)
[0115] Where addr is the virtual address in the collective memory (CM) allocated by CMalloc. This API can release the collective memory specified by addr. This API will complete a global synchronization within the comm range when returning.
[0116] Take the GPU as an example, refer to Figure 7 The optional structural diagram of the GPU provided in the embodiment of the present application is shown as an example. Figure 7 As shown in Figure 1, the GPU includes multiple processors, LLC and HBM. The LLC is equipped with a collective memory management CM agent. The GPU will identify the virtual address in the CM and confirm the Store forwarding behavior on the address (such as Figure 7 Store to HBM and Store to Data Switch Device in ).
[0117] Among them, processor refers to the basic unit inside the GPU for performing computing tasks; LLC refers to the last-level cache, which is used to cache frequently accessed data to reduce access latency to lower levels (such as HBM); HBM refers to high-bandwidth memory, which can provide higher bandwidth and lower latency than traditional DRAM; collective memory management CM agent is responsible for managing collective memory in the GPU and coordinating communications between different processors or GPUs. Store forwarding behavior refers to the behavior of storing data (through Store instructions) to a specified location (such as HBM or data exchange device).
[0118] The Store forwarding behavior corresponding to each communication operator in this application may be as follows:
[0119] For the communication operator Reduce, the Store instruction bypasses the local storage (such as HBM) and is directly forwarded to the data exchange device, which can write the target result data calculated by it to the locally allocated HBM. Only the master node or the root rank can send a storage instruction load to the collective memory CM. The storage instruction load of other node ranks will cause hardware errors.
[0120] For the communication operator Allreduce, the Store instruction bypasses the local storage (such as HBM) and is directly forwarded to the data exchange device, which can write the target result data calculated by it to the locally allocated HBM. All node ranks can send storage instructions load to the collective memory CM.
[0121] For the communication operator Reduce_scatter, the Store instruction bypasses the local storage (such as HBM) and is directly forwarded to the data exchange device, which can write the target result data calculated by it to the locally allocated HBM according to the rankid (that is, the identifier of each core computing unit). All node ranks can send storage instructions to the collective memory CM.
[0122] For the communication operator Allgather, the Store instruction will be written to the local storage (such as HBM) and forwarded to the data exchange device, which can write the target result data calculated by it to the locally allocated HBM. All node ranks can send storage instructions to the collective memory CM.
[0123] For the communication operator Broadcast, only the root rank can send a Store instruction to the collective memory CM. The Store instruction will be written to the local storage (such as HBM) and forwarded to the data exchange device. The data exchange device can write the target result data calculated by it to the locally allocated HBM. All node ranks can send a storage instruction load to the collective memory CM.
[0124] For the communication operator Scatter, only the root rank can send a Store instruction to the collective memory CM. The Store instruction will be written to the local storage (such as HBM) and forwarded to the data exchange device. The data exchange device can write the target result data calculated by it to the locally allocated HBM. All node ranks can send a storage instruction load to the collective memory CM.
[0125] For the communication operator P2P, only the source node src rank can send a Store instruction to the collective memory CM. The Store instruction will be written to the local storage (such as HBM), forwarded to the data exchange device, and written to the target node dst rank by the data exchange device. The source node src rank and the target node dst rank can send a storage instruction load to the collective memory CM.
[0126] It should be noted that the above-mentioned API for applying for a virtual address, the API for confirming that the communication operation corresponding to the virtual address has been completed, the API for converting memory, and the API for releasing video memory are only optional examples given in the embodiments of the present application. In practice, the API can be designed accordingly according to the needs, and the embodiments of the present application do not limit this.
[0127] The following is an introduction to the data processing device provided in the embodiment of the present application. The data processing device described below can be considered as a software function module or a hardware function module required to implement the data processing method provided in the embodiment of the present application. The content of the data processing device described below can be referenced to the content of the method described above.
[0128] In an optional implementation, Figure 8 An optional structural diagram of a data processing device provided in an embodiment of the present application is exemplarily shown, wherein the data processing device is used to implement the data processing method applied to a core computing processing system provided in an embodiment of the present application, wherein the core computing processing system includes multiple core computing processing units, such as Figure 8 As shown, the data processing device may include:
[0129] The initial result data acquisition module 11 is used to acquire the subtasks of the task to be processed assigned by each core computing processing unit for parallel processing, and obtain the initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks, and an accumulation task for performing a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process one subtask;
[0130] The initial result data sending module 12 is used to send the obtained initial result data to the data exchange device when any core computing processing unit completes the corresponding subtask, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data;
[0131] A target result data acquisition module 13 is used to acquire the target result data fed back by the data exchange device as the result data of the task to be processed;
[0132] Among them, each core computing processing unit is configured with the same virtual address, and in different core computing processing units, the virtual address corresponds to different physical addresses; each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate the mapping relationship between the virtual address and the physical address.
[0133] Optionally, also include:
[0134] A determination module 14, configured to determine a physical address corresponding to the virtual address in each core computing processing unit based on the page table;
[0135] The storage module 15 is used to store the result data to the corresponding physical address in each core computing processing unit based on the physical address.
[0136] In another alternative implementation, Fig. 9Another optional structural diagram of the data processing device provided in the embodiment of the present application is exemplarily shown, wherein the data processing device is used to implement the data processing method applied to the data exchange device provided in the embodiment of the present application, such as Fig. 9 As shown, the data processing device may include:
[0137] A receiving module 21, configured to receive initial result data, wherein the initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask;
[0138] The calculation module 22 is used to perform a preset accumulation operation on the received initial result data and obtain target result data;
[0139] The feedback module 23 is used to feed back the target result data to each core computing processing unit.
[0140] An embodiment of the present application also provides an electronic device, which may include at least one memory and at least one processor, wherein the memory stores one or more computer-executable instructions, and the processor calls the one or more computer-executable instructions to execute the data processing method as described above.
[0141] As an optional implementation, refer to Fig.10 , Fig.10 is an optional block diagram of an electronic device provided in an embodiment of the present application. Fig.10 As shown, the electronic device may include: at least one processor 10 , at least one communication interface 20 , at least one memory 30 and at least one communication bus 40 .
[0142] In the embodiment of the present application, the number of the processor 10 , the communication interface 20 , the memory 30 and the communication bus 40 is at least one, and the processor 10 , the communication interface 20 , and the memory 30 communicate with each other through the communication bus 40 .
[0143] Optionally, the processor 10 may be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural-network Processing Unit), an FPGA (Field Programmable Gate Array), a TPU (Tensor Processing Unit), an AI chip, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement an embodiment of the present application.
[0144] Optionally, the communication interface 20 may be an interface of a communication module for performing network communication.
[0145] The memory 30 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 30 stores one or more computer executable instructions, and the processor 10 calls the one or more computer executable instructions to execute the data processing method as described above.
[0146] An embodiment of the present application further provides a storage medium, wherein the storage medium stores one or more computer executable instructions, and when the one or more computer executable instructions are executed, the data processing method as described above is implemented.
[0147] An embodiment of the present application also provides a computer program product, which may include one or more computer executable instructions. When the one or more computer executable instructions are executed, the data processing method described above is implemented.
[0148] The above describes multiple implementation schemes provided by the embodiments of the present application. The various optional methods introduced in each implementation scheme can be combined and cross-referenced with each other without conflict, thereby extending a variety of possible implementation schemes, which can all be considered as implementation schemes disclosed and open in the embodiments of the present application.
[0149] Although the embodiments of the present application are disclosed above, the present application is not limited thereto. Any person skilled in the art may make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be subject to the scope defined by the claims.
Claims
1. A data processing method, characterized in that: Applied to a core computing processing system, the core computing processing system includes a plurality of core computing processing units, the method includes: Each core computing processing unit processes the assigned subtasks of the task to be processed in parallel to obtain initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks and an accumulation task for performing a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process one subtask; When any core computing processing unit completes the corresponding subtask, the obtained initial result data is sent to the data exchange device, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data; The target result data fed back by the data exchange device is obtained as the result data of the task to be processed.
2. The data processing method according to claim 1, characterized in that: Each core computing processing unit is configured with the same virtual address, wherein the virtual address corresponds to different physical addresses in different core computing processing units.
3. The data processing method according to claim 2, characterized in that: The step of sending the obtained initial result data to the data exchange device when any core computing processing unit completes the corresponding subtask includes: When any core computing processing unit completes the corresponding subtask, the obtained initial result data is written into the virtual address; Determining whether the virtual address is the same virtual address configured for each core computing processing unit; If yes, the obtained initial result data is sent to the data exchange device.
4. The data processing method according to claim 2, characterized in that: Each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate the mapping relationship between the virtual address and the physical address.
5. The data processing method according to claim 4, characterized in that: After the step of obtaining the target result data fed back by the data exchange device as the result data of the task to be processed, the method further includes: Based on the page table, determining a physical address corresponding to the virtual address in each core computing processing unit; Based on the physical address, the result data is stored in a corresponding physical address in each core computing processing unit.
6. A data processing method, characterized in that: Applied to data exchange equipment, the method comprises: receiving initial result data, wherein the initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask; Performing a preset accumulation operation on the received initial result data and obtaining the target result data; The target result data is fed back to each core computing processing unit.
7. A data processing device, characterized in that: Applied to a core computing processing system, the core computing processing system includes a plurality of core computing processing units, and the device includes: The initial result data acquisition module is used to acquire the subtasks of the task to be processed assigned by each core computing processing unit for parallel processing, and obtain the initial result data corresponding to the subtask; wherein the task to be processed includes multiple subtasks, and an accumulation task for performing a preset accumulation operation on the initial result data of the multiple subtasks; a core computing processing unit is used to process one subtask; An initial result data sending module is used to send the obtained initial result data to the data exchange device when any core computing processing unit completes the corresponding subtask, so that the data exchange device performs a preset accumulation operation on the received initial result data and obtains the target result data; A target result data acquisition module, used to acquire the target result data fed back by the data exchange device as the result data of the task to be processed; Among them, each core computing processing unit is configured with the same virtual address, and in different core computing processing units, the virtual address corresponds to different physical addresses; each core computing processing unit is configured with a corresponding page table, and the page table is used to indicate the mapping relationship between the virtual address and the physical address.
8. The data processing device according to claim 7, characterized in that: Also includes: A determination module, configured to determine, based on the page table, a physical address corresponding to the virtual address in each core computing processing unit; The storage module is used to store the result data to the corresponding physical address in each core computing processing unit based on the physical address.
9. A data processing device, characterized in that: Applied to data exchange equipment, the device comprises: A receiving module, used for receiving initial result data, wherein the initial result data is data corresponding to the subtask obtained when the core computing processing unit completes the corresponding subtask; A calculation module is used to perform a preset accumulation operation on the received initial result data and obtain target result data; The feedback module is used to feed back the target result data to each core computing processing unit.
10. An electronic device, characterized in that: It comprises at least one memory and at least one processor, wherein the memory stores one or more computer executable instructions, and the processor calls the one or more computer executable instructions to execute the data processing method as described in any one of claims 1 to 5, or executes the data processing method as described in claim 6.
11. A storage medium, characterized in that: The storage medium stores one or more computer-executable instructions, and when the one or more computer-executable instructions are executed, the data processing method according to any one of claims 1 to 5 is implemented, or the data processing method according to claim 6 is implemented.
12. A computer program product, characterized in that The method comprises one or more computer executable instructions, which, when executed, implement the data processing method according to any one of claims 1 to 5, or implement the data processing method according to claim 6.
Citation Information
Cited By
Geometric processing method, graphics processor and computer equipment
CN120765448A