Data processing method and device, equipment and storage medium

By storing data items continuously in a distributed buffer and processing them with sub-buffers of uniform size, the problems of load imbalance and high communication costs in distributed data processing are solved, thereby improving data processing efficiency.

CN122044880APending Publication Date: 2026-05-15BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, distributed data processing suffers from problems such as load imbalance and high communication costs, resulting in low data processing efficiency.

Method used

By creating a distributed buffer, each data item is ensured to be stored contiguously in the buffer, and data blocks are stored in sub-buffers of uniform size, thus utilizing multiple computing units to perform data processing.

Benefits of technology

It effectively eliminates the load imbalance of computing units, improves the efficiency of data item location and access, reduces the consumption of video memory bandwidth, avoids the communication cost across computing units, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122044880A_ABST
    Figure CN122044880A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, equipment and a storage medium. The method provided herein comprises: obtaining a plurality of data items, each data item comprising at least one data block; a distributed buffer area is created, the distributed buffer area comprises a plurality of sub-buffer areas corresponding to the plurality of computing units, and the plurality of sub-buffer areas correspond to the same first size; a plurality of data items are written into a distributed buffer area, so that each data item is continuous in the distributed buffer area, and each data block is located in a single sub-buffer area; and executing a data processing process based on the data in the plurality of sub-buffers by using the plurality of computing units. In this way, load balance of all the computing units can be guaranteed, all the data items continuously stored in the distributed buffer area can be rapidly accessed, video memory bandwidth consumption is reduced, cross-device communication can be avoided due to the fact that a single data block is located in a single sub-buffer area, and data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The examples in this article generally relate to the field of computers, and in particular to methods, apparatuses, devices and computer-readable storage media for data processing. Background Technology

[0002] With the rapid growth of data volume, distributed data processing has been widely applied in numerous fields. Distributed data processing refers to breaking down large-scale computing tasks into multiple subtasks, which are then executed in parallel by multiple computing units. Therefore, improving the efficiency of distributed data processing is a key concern. Summary of the Invention

[0003] In a first aspect, a method for data processing is provided. The method includes: acquiring a plurality of data items, each data item comprising at least one data block; creating a distributed buffer, the distributed buffer comprising a plurality of sub-buffers corresponding to a plurality of computing units, the plurality of sub-buffers corresponding to the same first size; writing the plurality of data items into the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; and utilizing the plurality of computing units, performing a data processing procedure based on the data in the plurality of sub-buffers.

[0004] In a second aspect, an apparatus for data processing is provided. The apparatus includes: a first acquisition module configured to acquire a plurality of data items, each data item including at least one data block; a creation module configured to create a distributed buffer, the distributed buffer including a plurality of sub-buffers corresponding to a plurality of computing units, the plurality of sub-buffers corresponding to the same first size; a data writing module configured to write the plurality of data items into the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; and an execution module configured to perform a data processing procedure using the plurality of computing units based on data in the plurality of sub-buffers.

[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] In this way, by storing multiple data items in different sub-buffers of uniform size, the problem of load imbalance among multiple computing units can be effectively eliminated. Since each data item is stored contiguously in the distributed buffer, the efficiency of data item location and access can be improved, thereby effectively reducing the consumption of video memory bandwidth. In addition, all the constituent elements of a single data block are stored in a single sub-buffer, which ensures the integrity of data block storage and avoids the communication costs associated with cross-computing units for processing operations at the data block level.

[0009] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figure 2 Flowcharts illustrating example data processing procedures in several scenarios are shown. Figure 3 Example diagrams are shown for some scenarios of writing data items; Figure 4 Example block diagrams for distributed training frameworks in some scenarios are shown; Figure 5 Schematic block diagrams of example devices for data processing in some scenarios are shown; and Figure 6 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation

[0011] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.

[0012] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.

[0013] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0014] The examples in this document may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and provisions. In the examples presented herein, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.

[0015] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0016] Traditionally, data items are typically stored in arbitrary partitions across devices. This approach can lead to additional communication costs and redundant interactions related to data rearrangement during redistribution, reducing the efficiency of data processing.

[0017] A data processing scheme is provided. The scheme includes: acquiring multiple data items, each data item including at least one data block; creating a distributed buffer, the distributed buffer including multiple sub-buffers corresponding to multiple computing units, the multiple sub-buffers corresponding to the same first size; writing the multiple data items into the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; and using the multiple computing units, performing a data processing procedure based on the data in the multiple sub-buffers.

[0018] In this way, by storing multiple data items in different sub-buffers of uniform size, the problem of load imbalance among multiple computing units can be effectively eliminated. Furthermore, since each data item is stored contiguously in the distributed buffer, the efficiency of data item location and access can be improved, thereby effectively reducing memory bandwidth consumption. Additionally, because all the components of a data block are stored in a single sub-buffer, processing operations at the data block level can avoid the communication costs associated with cross-computing units, thus improving data processing efficiency.

[0019] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.

[0020] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, the example environment 100 may include an electronic device 110 and multiple computing units (including computing unit 120-1, computing unit 120-2... computing unit 120-N, etc.).

[0021] In this example environment 100, the electronic device 110 can divide a data item into multiple data blocks and allocate these data blocks to computing units 120-1, 120-2, ..., 120-N respectively, so that computing units 120-1, 120-2, ..., 120-N can perform data processing operations in parallel based on the allocated partial data blocks. The data processing operations can be any suitable operation, such as processing operations associated with model training tasks, image processing operations, task processing operations associated with embodied intelligence, etc., which will not be elaborated upon here.

[0022] In some cases, electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0023] Electronic device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0024] In some cases, a computing unit can be an execution unit capable of independently performing data processing, computation, or scheduling, which may include, but is not limited to: a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), an accelerator, a process, a thread, or a distributed node, a standalone device, etc.

[0025] Communication connections can be established between each computing unit and the electronic device 110, and between the computing units themselves. These communication connections can be established via wired or wireless means. The communication connections can include, but are not limited to, Bluetooth connections, mobile network connections, Universal Serial Bus (USB) connections, Wireless Fidelity (WiFi) connections, etc. In some cases, the computing units and the electronic device 110 can exchange signaling information through their communication connections.

[0026] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.

[0027] The following description of the example will continue with reference to the accompanying drawings.

[0028] Example process Figure 2 A flowchart of an example process 200 for data processing under certain conditions is shown. Process 200 can be implemented at electronic device 110. (See below for reference.) Figure 1 To describe process 200.

[0029] In box 210, electronic device 110 acquires multiple data items, each data item including at least one data block.

[0030] In some cases, multiple data items can be any suitable structured data unit, which may include, but is not limited to, matrix data, tensor data, etc. In some cases, different data items may have the same size and different types.

[0031] In some cases, multiple data items can be associated with any appropriate data processing task that can support distributed parallel processing, such as, but not limited to, at least one of the following: model inference task, model training task, image processing task, task associated with embodied intelligence, etc.

[0032] Taking the training task of a model as an example, multiple data items may include, but are not limited to, the model parameter matrix, gradient matrix, etc.

[0033] In some cases, a data block can indicate the smallest atomic unit of operation for a data item, which cannot be split into different computational units. In other cases, the data blocks comprising the same data item can be of the same size or different sizes.

[0034] In some cases, for each of these multiple data items, the at least one data block included in that data item can be segmented based on any suitable sharding method. Sharding methods can include, but are not limited to, at least one of the following: full sharding, regular sharding, irregular sharding, structure-aware sharding, etc., which will not be elaborated here.

[0035] In box 220, electronic device 110 creates a distributed buffer, which includes multiple sub-buffers corresponding to multiple computing units, and the multiple sub-buffers correspond to the same first size.

[0036] In some scenarios, a distributed buffer can be a globally unified virtual address space spanning multiple computational units. A distributed buffer can consist of multiple sub-buffers, each corresponding to a computational unit.

[0037] In some cases, the computing unit can be a standalone device or any suitable execution unit capable of independently performing data processing, computation, or scheduling, which may include, but is not limited to: CPU, GPU, NPU, TPU, accelerator, process, thread, or distributed node, etc.

[0038] To ensure load balancing across computing units and to ensure that multiple data items can be stored in the same location across different computing units, in some cases, the multiple sub-buffers included in the distributed buffer can correspond to the same size, i.e., they can correspond to the first size.

[0039] In some cases, in order to ensure that the storage space included in the distributed buffer is sufficient to support the writing of all these data items, the total available capacity of each sub-buffer can be greater than or equal to the amount of data corresponding to the multiple data items when each sub-buffer corresponds to the first size.

[0040] To enable data items to be segmented into data blocks based on a desired granularity, in some cases, the electronic device 110 can determine block size information. In some cases, the block size information can indicate the size of the data block corresponding to each data item. This data block size can indicate a reference size for the data block corresponding to each data item, or in other words, the data block size can be the granularity corresponding to the data blocks segmented for the desired data item. It should be noted that the sizes of the data blocks corresponding to different data items can be the same or different, and can be set according to requirements. As an example, the block size information can include multiple different second sizes. For example, the block size information can indicate the size of the data blocks corresponding to N data items, and this... If there are M data block sizes corresponding to different values ​​among the data items, then the second size information included in this block size information consists of M values. .

[0041] In order to ensure that the segmentation of each data item can follow the size of the data block corresponding to each data item, so as to ensure that each data block corresponding to each data item can be completely written into the sub-buffer, the electronic device 110 can determine the first size of the sub-buffer based on the block size information.

[0042] In some cases, the first dimension can be a common multiple of several different second dimensions. For example, if several different second dimensions include 1024, 2048, and 512, then the first dimension can be... , where 2048 is the least common multiple of the three second dimensions, and .

[0043] As an example, electronic device 110 can determine the least common multiple corresponding to these multiple different second dimensions. Further, electronic device 110 can determine the first dimension based on the product of this least common multiple and a predetermined value. The predetermined value can be any suitable integer value greater than 1.

[0044] It should be noted that, provided that the total available capacity corresponding to each sub-buffer of the first size is greater than or equal to the amount of data corresponding to multiple data items, the smaller the preset value, the higher the utilization rate of the sub-buffer.

[0045] Since there are many possible values ​​for the first size when it corresponds to a common multiple of multiple different second sizes (i.e., many preset values), in order to ensure high utilization of the sub-buffer, in some cases, the electronic device 110 can determine candidate sizes based on block size information. The candidate sizes can be common multiples of multiple different second sizes.

[0046] Since communication hardware itself may have address alignment and length alignment requirements, in order to eliminate the additional communication padding overhead caused by alignment compensation during communication and to improve data processing efficiency, in some cases, electronic device 110 may determine a third size based on configuration information. The configuration information may be a predetermined communication protocol alignment unit.

[0047] Furthermore, the electronic device 110 can determine candidate dimensions based on the block size information and the third dimension. The candidate dimensions can be common multiples of multiple different second and third dimensions. For example, if multiple different second dimensions include 1024 and 2048, and the third dimension is 512, then the first dimension can be... ,in 2048 in the equation is the least common multiple of the two second dimensions and one third dimension, and .

[0048] By incorporating the third dimension (communication alignment unit) into the common multiple calculation in this way, the size of the sub-buffer is naturally an integer multiple of the communication alignment unit, thus avoiding the extra padding operations required to ensure communication alignment.

[0049] To ensure that data blocks of any size among multiple data items can be written completely and aligned into the sub-buffer, as an example, candidate sizes can be common multiples of multiple different second sizes or common multiples of multiple second and third sizes. For instance, if the least common multiple of multiple different second sizes or multiple second and third sizes is... Then the candidate size can be ,in Any one of them.

[0050] Since there are many common multiples corresponding to multiple different second sizes or common multiples corresponding to multiple second and third sizes, in order to ensure both the utilization rate of the sub-buffers and the sufficient capacity of these sub-buffers to write the multiple data items, the electronic device 110 can further determine the first number of computing units required to write the multiple data items based on the candidate sizes.

[0051] In some cases, for each data item, the electronic device 110 can write its data blocks sequentially into the sub-buffer, starting from the beginning position of the data item, without crossing the sub-buffer boundary.

[0052] Specifically, for each data item, the electronic device can perform the following operations in sequence: The electronic device can start from position 0 of the current data item and determine the longest continuous interval that can be completely written within the corresponding sub-buffer. The length of this continuous interval cannot exceed the capacity of this sub-buffer and must ensure the integrity of the data block. And the length of the continuous interval is equal to the size of the data block corresponding to the data item, where This represents the size of the current data item. Furthermore, the electronic device 110 can update the starting position to... This process continues until all data items have been allocated. It should be noted that the longest continuous interval... It can include multiple complete data blocks, and the size of these multiple data blocks can be the size of the data block corresponding to the data item indicated by the block size information. For example, if the size of the data block corresponding to data item 1 is 2, then r in this longest continuous interval needs to satisfy the following condition: It is an integer multiple of 2. For example, if there are two data items, A and B, where data item A contains a total of 8 elements and the size of the data block corresponding to data item A is 2 (meaning each data block in data item A contains 2 elements), and data item B contains a total of 4 elements and the size of the data block corresponding to data item B is 2 (meaning each data block in data item B contains 2 elements), then the current candidate size is 4.

[0053] If data items A and B are allocated sequentially, then data block 0 (index 0-1) and data block 1 (index 2-3) of data item A can be allocated to the first sub-buffer, at which point the first sub-buffer is full, but data item A still has 4 elements that have not been allocated. Further, data block 2 (index 4-5) and data block 3 (index 6-7) of data item A can be allocated to the second sub-buffer, at which point the second sub-buffer is full, and all elements of data item A have been allocated. Further, data block 0 (index 0-1) and data block 1 (index 2-3) of data item B can be allocated to the third sub-buffer, at which point the third sub-buffer is full, and all elements of data item B have been allocated. At this point, electronic device 110 can determine that the first number of computational units required to write multiple data items is 3, where (index xy) can represent the data from the x-th element to the y-th element of the data item.

[0054] Furthermore, the electronic device 110 can determine a first size based on a candidate size in response to a first number being less than or equal to a second number, wherein the second number is the number of available computing units.

[0055] When data item A contains a total of 8 elements and the size of the data block corresponding to data item A is 2, data item B contains a total of 4 elements and the size of the data block corresponding to data item B is 2, the current candidate size is 4, and the first number has been determined to be 3. If the current number of available computing units is 4, then if the size of these 4 computing units is 4, there is enough space to support the writing of multiple data items. Therefore, electronic device 110 can determine the first size as 4.

[0056] In other cases, the electronic device 110 may increase the candidate size in response to the first number being greater than the second number, to further perform the process of determining the first size based on the updated candidate size as described above. It should be noted that the process of determining the first size based on the updated candidate size is the same as the process of determining the first size based on the candidate size before the update, and will not be described in detail here.

[0057] For example, when data item A contains a total of 8 elements and the size of the data block corresponding to data item A is 2, and data item B contains a total of 4 elements and the size of the data block corresponding to data item B is 2, and the current candidate size is 4, and the first number has been determined to be 3, if the current number of available computing units is 2, then if the size of these 2 computing units is 4, there is not enough space to support writing multiple data items, so the candidate size needs to be increased.

[0058] Since there can be multiple possible candidate sizes, for example, the candidate size can be... ,in Any one of the multiple different second dimensions, where the least common multiple of multiple second dimensions and the least common multiple of multiple second dimensions and third dimensions is g, so in order to improve the utilization of the sub-buffer, in some cases, the electronic device 110 can determine the minimum of all multiple candidate dimensions that meet the requirements as the first dimension.

[0059] As an example, if multiple data items include a size of Data item 1 and size Data item 1, where the second size of the data block corresponding to the data item is The second size of the data block corresponding to data item 2 is The third size determined based on the configuration information is Then the electronic device can determine , , Least Common Multiple Furthermore, the electronic device can determine the candidate size ( ,in (any one of them). Furthermore, electronic devices can be from... The smallest candidate size that satisfies the predetermined constraints is selected as the first size S. The predetermined constraints can indicate that all data blocks can be completely written into a single sub-buffer, and when the data items are arranged continuously in the distributed buffer and the sub-buffer corresponds to the first size S, the available capacity of the distributed buffer can be used to write these multiple data items.

[0060] To improve the efficiency of determining the first size, the electronic device 110 can use a binary search method to sequentially determine the corresponding candidate sizes and determine whether the first number determined based on the corresponding candidate size is less than or equal to the second number. Further, if the first number determined based on the corresponding candidate size is less than or equal to the second number, this corresponding candidate size can be updated as the upper limit of the size, and the midpoint between the updated upper and lower limits of the size can be used as the next candidate size to further determine whether the first number determined based on the next candidate size is less than the second number… This process is repeated until the smallest candidate size is determined, wherein the first number determined based on this smallest candidate size is also less than or equal to the second number.

[0061] As another example, if the first number determined based on the corresponding candidate size is greater than the second number, the corresponding candidate size can be updated to the lower limit of the size, and the middle value between the updated lower limit of the size and the upper limit of the size can be used as the next candidate size to further determine whether the first number determined based on the next candidate size is less than the second number... and so on, until the smallest candidate size is determined, wherein the first number determined based on the smallest candidate size is also less than or equal to the second number.

[0062] It should be noted that when searching for each candidate size using the binary search method, the search range corresponding to the candidate size is variable. That is, the upper and lower limits of the range change based on the search process, which will not be elaborated here.

[0063] Furthermore, the electronic device 110 can create a distributed buffer based on the first size of the sub-buffer.

[0064] In some cases, the electronic device 110 may allocate a contiguous sub-buffer corresponding to a first size in the local buffer of each computing unit and map these sub-buffers across the device to a unified, globally addressable logical address space to create a distributed buffer.

[0065] In box 230, electronic device 110 writes multiple data items to a distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer.

[0066] In some cases, electronic device 110 can write multiple data items to corresponding locations in a distributed buffer.

[0067] Since a partitioning strategy for the multiple data items has been determined before the candidate size is determined as the first size, i.e. during the process of determining the first number of computing units required to write the multiple data items based on this candidate size, this partitioning strategy can indicate how the data blocks corresponding to the multiple data items are partitioned and which computing unit each data block is specifically assigned to.

[0068] For example, if there are two data items, A and B, where A contains a total of 8 elements and the size of the data block corresponding to A is 2 (meaning each data block in A contains 2 elements), and B contains a total of 4 elements and the size of the data block corresponding to B is 2 (meaning each data block in B contains 2 elements), then the candidate size (before it is determined as the first size) is 4.

[0069] If data items A and B are allocated sequentially, then data block 0 (index 0-1) and data block 1 (index 2-3) of data item A can be allocated to the first sub-buffer, data block 2 (index 4-5) and data block 3 (index 6-7) of data item A can be allocated to the second buffer, and data block 0 (index 0-1) and data block 1 (index 2-3) of data item B can be allocated to the third sub-buffer. If the distributed buffer includes sub-buffer 1, sub-buffer 2, and sub-buffer 3, then the partitioning strategy determined by electronic device 110 can instruct data block 0 (index 0-1) and data block 1 (index 2-3) of data item A to be allocated to sub-buffer 1, data block 2 (index 4-5) and data block 3 (index 6-7) of data item A to be allocated to sub-buffer 2, and data block 0 (index 0-1) and data block 1 (index 2-3) of data item B to be allocated to sub-buffer 3.

[0070] In some cases, for each sub-buffer, the partitioning strategy can also indicate the write order of the data blocks allocated to that sub-buffer. For example, the write order of different data blocks belonging to the same data item is consecutive, and the write order of different data blocks of the same data item can be determined based on the index information of the data blocks.

[0071] Since some data blocks corresponding to multiple data items may still have free storage space when written to a sub-buffer, in order to ensure the symmetry requirement of set communication, the electronic device 110 can respond to the distributed buffer including unwritten areas by filling the unwritten areas with placeholder data. In some cases, the placeholder data can be blank space that only occupies space, has no semantics, and is not processed.

[0072] To ensure that multiple data blocks belonging to the same data item are stored contiguously but written to different computing units, the electronic device 110 can respond to the fact that the size of each data block written to the first sub-buffer is less than the first size by writing the first data block to the end of the first sub-buffer and writing placeholder data between the first data block and the second data block. The second data block and the first data block correspond to different data items, and the first data block and the third data block in the second sub-buffer are contiguous and correspond to the same data item.

[0073] Figure 3 Example diagrams for writing data items are shown in some scenarios, now focusing on Figure 3 Please provide an explanation.

[0074] like Figure 3As shown, taking multiple data items including data items 333, 310, and 320 as an example, electronic device 110 can write data items 333, 310, and 320 into a distributed buffer. Specifically, it can write data blocks 301, 302, 303, and 304 from data item 333 into a sub-buffer of computing unit 1, and write data block 311 from data item 310 into a sub-buffer of computing unit 1. Additionally, electronic device 110 can also write data blocks 312 and 313 from data item 310, and data blocks 321 and 322 from data item 320 into a sub-buffer of computing unit 2.

[0075] Since the data sizes corresponding to data blocks 301, 302, 303, 304, and 311 are smaller than the total capacity of the sub-buffer of computing unit 1, and since data blocks 312 and 313 of data item 310 are written into the sub-buffer of computing unit 2, in order to ensure the continuity of each data block in data item 310, electronic device 110 can write data block 311 to the tail position of the sub-buffer of computing unit 1, and placeholder data 350 is filled between data block 311 and data block 304.

[0076] In box 240, electronic device 110 utilizes multiple computing units to perform a data processing procedure based on data in multiple sub-buffers.

[0077] In some cases, the data processing procedure can be any appropriate procedure, such as gradient calculation, image feature processing, etc. That is, multiple computing units can process their corresponding data processing procedures in parallel based on the data in their own sub-buffers to obtain their own data processing results. The data processing procedures executed by different computing units can be the same or different, but the data used for their processing procedures will be different.

[0078] In some cases, these multiple data items can be associated with the model's training task.

[0079] In order to aggregate the local computation results corresponding to multiple computing units into a global training signal for training tasks, in some cases, the electronic device 110 can obtain data processing results from at least one of the multiple computing units, the data processing results being related to the data processing process. Furthermore, the electronic device 110 can perform training tasks based on the data processing results.

[0080] As an example, taking electronic device 110 as one of multiple computing units, electronic device 110 can obtain data processing results from other computing units besides itself. Furthermore, electronic device 110 can execute training tasks based on its own determined data processing results and the data processing results of other computing units.

[0081] As another example, taking electronic device 110 as an independent device separate from multiple computing units, electronic device 110 can obtain the data processing results of these multiple computing units. Furthermore, electronic device 110 can perform training tasks based on the data processing results of these multiple computing units.

[0082] In some cases, these multiple computing units include at least a first computing unit and a second computing unit. It should be noted that the first computing unit and the second computing unit merely represent different computing units and do not limit the number of these multiple computing units.

[0083] To improve data processing efficiency and avoid redundant communication costs, electronic device 110 can determine a first computing unit from multiple computing units. The first computing unit is configured to receive processing results sent by other computing units and integrate the results; this first computing unit may also be referred to as a root device. As an example, electronic device 110 can determine the first computing unit based on the resource status of the computing resources corresponding to the multiple computing units. The resource status may indicate, but is not limited to, at least one of the following: the remaining amount of computing resources, the occupancy of computing resources, etc. The computing resources can be any suitable resource, such as a CPU, GPU, etc.

[0084] For example, electronic device 110 can identify the computing unit with the most remaining computing resources among these multiple computing units as the first computing unit.

[0085] In some cases, the first computing unit can obtain a first processing result from the second computing unit, whereby the first processing result is generated by the second computing unit performing a data processing procedure. Further, the first computing unit can generate a data processing result from the first and second processing results, whereby the second processing result is generated by the first computing unit performing a data processing procedure. As an example, the first computing unit can generate a global processing result by fusing the first and second processing results.

[0086] In some cases, the first computing unit can obtain the first processing result of the second computing unit based on a predetermined interface. Specifically, the electronic device 110 can store pointer information corresponding to each data block among multiple data items, and the pointer information can indicate the actual storage area of ​​each data block in each sub-buffer. Further, the first computing unit can use this predetermined interface to obtain the first processing result of the second computing unit based on the pointer information.

[0087] In some cases, data processing may include gradient computation. In other cases, the result of data processing may include aggregated gradients generated by aggregating multiple gradients, which are produced by multiple computational units performing gradient computation. That is, each of these computational units can determine its local gradient based on the data blocks written to its own sub-buffer.

[0088] Furthermore, the first computational unit can generate adjustment information based on aggregated gradients. In some cases, the first computational unit can generate adjustment information for each computational unit based on aggregated gradients. In some cases, the adjustment information can be parameter update amounts. The adjustment information for different computational units may be the same or different.

[0089] Furthermore, the first computing unit can send adjustment information to multiple computing units to update the model parameters at those units. In some cases, the first computing unit can send the corresponding adjustment information to the respective computing unit, so that the corresponding computing unit can update its own model parameters based on the received adjustment information.

[0090] Figure 4 Example block diagrams for distributed training frameworks in some scenarios are shown, now focusing on Figure 4 Please provide an explanation.

[0091] In step 410, the electronic device 110 can call multiple data items of the model using a predetermined interface.

[0092] In some cases, multiple data items can be matrix data or tensor data of any appropriate type, such as model parameter matrix, gradient matrix, etc.

[0093] In some cases, these multiple data items are configured to perform model-related training tasks.

[0094] In step 420, the structure-aware optimizer may indicate at least one constraint associated with data processing, which may include, but is not limited to, block integrity constraints, resource capacity constraints, and communication alignment constraints.

[0095] In some cases, block integrity constraints can instruct that each data block must be written completely to a single sub-buffer and must not cross computation units.

[0096] In some cases, resource capacity constraints can instruct that these multiple data items must be fully contained within the total storage space of the available computing units.

[0097] In some cases, communication alignment constraints can indicate that the size of a sub-buffer is an integer multiple of the communication protocol alignment unit, i.e., the size of the sub-buffer is a common multiple of the third size and multiple different second sizes.

[0098] In step 430, the electronic device 110 uses an intermediate adaptation layer to convert at least one constraint into executable instructions.

[0099] In step 440, the electronic device 110 can create a distributed buffer, which may include multiple computing units, and the size of the sub-buffers corresponding to the multiple computing units may all be a first size.

[0100] In some cases, the first dimension can be determined using executable instructions based on information such as block size information, the size of multiple data items, the number of available computing units, and a third dimension. The block size information can indicate the size of the data block corresponding to the multiple data items, and the third dimension indicates the communication alignment unit.

[0101] In step 450, the electronic device 110 can write multiple data items into a distributed buffer, such that each sub-buffer is written with at least a portion of the data of multiple data items, and then multiple computing units can perform corresponding data processing procedures based on the data written in their respective sub-buffers.

[0102] In this way, by storing multiple data items in different sub-buffers of uniform size, the problem of load imbalance among multiple computing units can be effectively eliminated. Furthermore, since each data item is stored contiguously in the distributed buffer, the efficiency of data item location and access can be improved, thereby effectively reducing memory bandwidth consumption. Additionally, because all the components of a data block are stored in a single sub-buffer, processing operations at the data block level can avoid the communication costs associated with cross-computing units, thus improving data processing efficiency.

[0103] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 5A schematic structural block diagram of an example device 500 for data processing according to some scenarios is shown. Device 500 can be implemented as or included in electronic device 110. The various modules / components in device 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0104] like Figure 5 As shown, the apparatus 500 includes: a first acquisition module 510 configured to acquire a plurality of data items, each data item including at least one data block; a creation module 520 configured to create a distributed buffer, the distributed buffer including a plurality of sub-buffers corresponding to a plurality of computing units, the plurality of sub-buffers corresponding to the same first size; a data writing module 530 configured to write the plurality of data items into the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; and an execution module 540 configured to perform a data processing procedure using the plurality of computing units based on the data in the plurality of sub-buffers.

[0105] In some cases, the creation module 520 is also configured to: determine block size information, which indicates the size of the data block corresponding to each data item; determine a first size of the sub-buffer based on the block size information; and create a distributed buffer based on the first size of the sub-buffer.

[0106] In some cases, the block size information includes multiple different second sizes, and the first size is a common multiple of the multiple different second sizes.

[0107] In some cases, the creation module 520 is also configured to: determine a candidate size based on block size information, the candidate size being a common multiple of a plurality of different second sizes; determine a first number of computing units required to write a plurality of data items based on the candidate size; and determine a first size based on the candidate size in response to the first number being less than or equal to a second number, the second number being the number of available computing units.

[0108] In some cases, the creation module 520 is also configured to: determine a third size based on configuration information; and determine candidate sizes based on block size information and the third size, wherein the candidate sizes are common multiples of multiple different second and third sizes.

[0109] In some cases, the data writing module 530 is also configured to: write multiple data items to corresponding positions in the distributed buffer; and, in response to the distributed buffer including unwritten areas, fill the unwritten areas with placeholder data.

[0110] In some cases, multiple data items are associated with the model's training task. The apparatus 500 also includes a second acquisition module configured to: acquire data processing results from at least one of the multiple computing units, the data processing results being related to the data processing process; and an execution module configured to: execute the training task based on the data processing results.

[0111] In some cases, the multiple computing units include at least a first computing unit and a second computing unit. The first computing unit is configured to: obtain a first processing result from the second computing unit, the first processing result being generated by the second computing unit performing a data processing procedure; and generate a data processing result based on the first processing result and the second processing result, the second processing result being generated by the first computing unit performing a data processing procedure.

[0112] In some cases, the data processing procedure includes a gradient calculation procedure, and the data processing result includes an aggregated gradient generated by aggregating multiple gradients, which are generated by multiple computing units performing the gradient calculation procedure.

[0113] In some cases, the execution module is also configured to: generate adjustment information based on aggregated gradients; and send adjustment information to multiple computational units to update model parameters at multiple computational units.

[0114] The modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 500 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0115] Figure 6 A block diagram of an electronic device 600 in which one or more examples may be implemented is shown. It should be understood that... Figure 6The electronic device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 6 The illustrated electronic device 600 can be used to implement the electronic device 110 discussed above.

[0116] like Figure 6 As shown, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 600.

[0117] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.

[0118] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various examples.

[0119] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0120] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0121] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0122] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0123] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0124] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0125] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0126] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A data processing method, comprising: Retrieve multiple data items, each data item including at least one data block; A distributed buffer is created, the distributed buffer comprising multiple sub-buffers corresponding to multiple computing units, the multiple sub-buffers corresponding to the same first size; The plurality of data items are written to the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; as well as Using the multiple computing units, a data processing procedure is performed based on the data in the multiple sub-buffers.

2. The method of claim 1, wherein creating the distributed buffer comprises: Determine the block size information, which indicates the size of the data block corresponding to each data item; Based on the block size information, the first size of the sub-buffer is determined; and The distributed buffer is created based on the first size of the sub-buffer.

3. The method of claim 2, wherein the block size information includes a plurality of different second sizes, and the first size is a common multiple of the plurality of different second sizes.

4. The method according to claim 3, wherein determining the first size of the sub-buffer based on the block size information comprises: Based on the block size information, candidate sizes are determined, wherein the candidate sizes are common multiples of the plurality of different second sizes; Based on the candidate size, a first number of computing units required to write the plurality of data items is determined; as well as In response to the first number being less than or equal to the second number, the first size is determined based on the candidate size, where the second number is the number of available computing units.

5. The method of claim 4, wherein determining the candidate size based on the block size information includes: Based on the configuration information, determine the third size; as well as Based on the block size information and the third size, the candidate size is determined, wherein the candidate size is a common multiple of the plurality of different second sizes and the third size.

6. The method of claim 1, wherein writing the plurality of data items to the distributed buffer comprises: Write the multiple data items to the corresponding positions in the distributed buffer; as well as In response to the distributed buffer including unwritten areas, placeholder data is filled into the unwritten areas.

7. The method of claim 1, wherein the plurality of data items are associated with a model training task, the method further comprising: A data processing result is obtained from at least one of the plurality of computing units, the data processing result being related to the data processing process; as well as Based on the data processing results, the training task is executed.

8. The method of claim 7, wherein the plurality of computing units includes at least a first computing unit and a second computing unit, the first computing unit being configured to: A first processing result is obtained from the second computing unit, the first processing result being generated by the second computing unit executing the data processing procedure; and The data processing result is generated based on the first processing result and the second processing result, wherein the second processing result is generated by the first computing unit executing the data processing procedure.

9. The method of claim 7, wherein the data processing procedure includes a gradient calculation procedure, and the data processing result includes an aggregated gradient generated by aggregating multiple gradients, the multiple gradients being generated by the multiple computing units performing the gradient calculation procedure.

10. The method of claim 9, wherein performing the training task based on the data processing result comprises: Based on the aggregated gradient, adjustment information is generated; as well as The adjustment information is sent to the plurality of computing units to update the model parameters at the plurality of computing units.

11. An apparatus for data processing, comprising: The first acquisition module is configured to acquire multiple data items, each data item including at least one data block; A creation module is configured to create a distributed buffer, the distributed buffer comprising multiple sub-buffers corresponding to multiple computing units, the multiple sub-buffers corresponding to the same first size; The data writing module is configured to write the plurality of data items into the distributed buffer such that each data item is contiguous in the distributed buffer and each data block is located in a single sub-buffer; as well as The execution module is configured to utilize the plurality of computing units to perform a data processing procedure based on the data in the plurality of sub-buffers.

12. An electronic device, comprising: At least one processor; as well as At least one memory, coupled to the at least one processor and storing instructions for execution by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.