Data Processing Method, Data Processing Device, Electronic Device, and Storage Medium
By analyzing training status and operation records in real time, generating data scheduling strategies, combining task queues and state management, the problem of insufficient storage capacity of AI accelerator is solved, and efficient and stable data transfer and model training are achieved.
Patent Information
- Application Number
- CN202510299974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, the growth rate of storage capacity of AI accelerators is much lower than the growth of neural network scale, making it difficult to meet the training needs of large-scale deep neural network models. The existing data scheduling strategies lack effective mechanisms and automatic adjustment capabilities, resulting in insufficient system stability and flexibility.
By obtaining user interface information and operation status information, generating data scheduling strategies, analyzing training status and operation records in real time, dynamically adjusting data scheduling strategies, introducing task queues and state management mechanisms, realizing task-based push and subscription execution modes, and optimizing data transfer operations.
It improves the efficiency and stability of the system, adapts to the training needs of complex models, improves the flexibility and resource utilization of the system, avoids system crashes, and optimizes the scope of application of data scheduling strategies.
Smart Images

Figure CN119902874B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a data processing method, a data processing apparatus, an electronic device, and a storage medium. Background Art
[0002] With the development of technology, Artificial Intelligence (AI) technology has been widely applied in multiple fields. Deep Learning is one of the important technologies of AI technology. Deep learning technology based on artificial neural networks has made great progress in fields such as object classification, text processing, image search, and human-computer dialogue.
[0003] With the increase in problem complexity, the depth and scale of neural networks have also been continuously improved. However, the growth rate of the storage capacity of AI accelerators such as Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), and Field Programmable Gate Array (FPGA) is much lower than the growth of the neural network scale, making it difficult to meet the training requirements of large-scale deep neural network models. Summary of the Invention
[0004] At least one embodiment of the present disclosure provides a data processing method, where the data processing method includes: obtaining a current operation record according to user interface information and operation status information; generating a data scheduling policy according to the historical operation record and the current operation record; and performing a corresponding data transfer operation according to the data scheduling policy.
[0005] For example, in the data processing method provided by at least one embodiment of the present disclosure, the user interface information includes an interface name, and the operation status information includes a current running location. The obtaining a current operation record according to the user interface information and the operation status information includes: combining the interface name and the current running location to obtain the current operation record.
[0006] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: storing the current operation record in an operation record list; and reading the historical operation record from the operation record list.
[0007] For example, in the data processing method provided by at least one embodiment of the present disclosure, the data scheduling policy includes task information, and the task information is used to generate a task. The generating of the data scheduling policy according to the historical operation record and the current operation record includes: generating the task information corresponding to the current operation record according to the relative order of the current operation record corresponding to the historical operation record.
[0008] For example, in the data processing method provided by at least one embodiment of the present disclosure, the performing of the corresponding data transfer operation according to the data scheduling policy includes: in response to the task information corresponding to an offloading task or a prefetching task, performing a validity check on the task information; in response to the task information passing the validity check, reading the task corresponding to the task information from the corresponding task queue and executing it to implement the corresponding data transfer operation.
[0009] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: in response to the execution of the task, generating a new task, setting the state of the new task based on the state of the task; inserting the new task into the corresponding task queue.
[0010] For example, in the data processing method provided by at least one embodiment of the present disclosure, the state of the task includes a ready state, an offloading state, a prefetching state, or a completed state. The generating of a new task in response to the execution of the task and setting the state of the new task based on the state of the task includes: in response to the execution of an offloading task and the state of the task being the ready state, generating a new task and setting the state of the new task to the offloading state; in response to the execution of a prefetching task and the state of the task being the offloading state, generating a new task and setting the state of the new task to the prefetching state; in response to the completion of the execution of the prefetching task and the state of the task being the prefetching state, generating a new task and setting the state of the new task to the completed state.
[0011] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: in response to the user interface information corresponding to the task information indicating ready for offloading, initializing the state of the task to the ready state.
[0012] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: in response to determining whether to return the data object of the task and the state of the task being the ready state, updating the state of the task to the completed state.
[0013] For example, in the data processing method provided by at least one embodiment of the present disclosure, after generating the task information corresponding to the current operation record, the method further includes: in response to the user interface information corresponding to the task information indicating preparation for unloading, inserting the task corresponding to the task information into the corresponding task queue.
[0014] For example, in the data processing method provided by at least one embodiment of the present disclosure, after generating the task information corresponding to the current operation record, the method further includes: in response to the user interface information corresponding to the task information indicating loading, performing a validity check on the task information; in response to the task information passing the validity check, reading the task corresponding to the task information from the corresponding task queue and determining whether to return the data object of the task.
[0015] For example, in the data processing method provided by at least one embodiment of the present disclosure, tasks in different states correspond to different task queues.
[0016] For example, in the data processing method provided by at least one embodiment of the present disclosure, the data transfer operation includes an unloading operation or a prefetch operation for tensors performed between a central processing unit and a graphics processing unit.
[0017] At least one embodiment of the present disclosure provides a data processing device, the data processing device includes: an acquisition module configured to acquire a current operation record according to user interface information and operating status information; a generation module configured to generate a data scheduling policy according to historical operation records and the current operation record; an execution module configured to perform a corresponding data transfer operation according to the data scheduling policy.
[0018] At least one embodiment of the present disclosure provides an electronic device, the electronic device includes: at least one processor; at least one memory storing one or more computer program modules; wherein, the one or more computer program modules are configured to be executed by the at least one processor to execute instructions for implementing the data processing method of any embodiment of the present disclosure.
[0019] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, on which computer-readable instructions are stored, wherein the computer-readable instructions, when executed by at least one processor, execute the data processing method of any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0021] Figure 1A Schematic diagram of a data parallel mode;
[0022] Figure 1B Schematic diagram of a tensor parallel mode;
[0023] Figure 1C Schematic diagram of a pipeline parallel mode;
[0024] Figure 1D Schematic diagram of a micro-batch pipeline parallel mode;
[0025] Figure 1E Schematic diagram of forward and backward calculations for cross-execution of micro-batch data;
[0026] Figure 1F Schematic diagram of recomputation;
[0027] Figure 2 Schematic block diagram of a distributed training system provided by at least one embodiment of the present disclosure;
[0028] Figure 3 Flowchart of a data processing method provided by at least one embodiment of the present disclosure;
[0029] Figure 4A Exemplary schematic diagram of a data processing method provided by at least one embodiment of the present disclosure;
[0030] Figure 4B Exemplary schematic diagram of a data processing method provided by at least one embodiment of the present disclosure;
[0031] Figure 5A Exemplary schematic diagram of a data processing method provided by at least one embodiment of the present disclosure;
[0032] Figure 5B Schematic diagram of state machine transition provided by at least one embodiment of the present disclosure;
[0033] Figure 6A Schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure;
[0034] Figure 6B Schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure;
[0035] Figure 7 Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure;
[0036] Figure 8 Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure; and
[0037] Figure 9 Schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0039] Flowcharts are used in the present disclosure to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations described above or below do not necessarily have to be executed precisely in sequence. On the contrary, various steps can be processed in reverse order or simultaneously as needed. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0040] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0041] The present disclosure will be described below through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one drawing, the component is denoted by the same or similar reference numerals in each drawing.
[0042] There is a wide variety of processors for model training, such as Graphics Processing Units (GPUs), General-Purpose Graphics Processing Units (GPGPUs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), Artificial Intelligence (AI) accelerators or coprocessors, etc. Currently, GPUs are usually used as the main training devices, combined with Central Processing Units (CPUs) for auxiliary tasks and overall process management. For specific scenarios or requirements, ASICs, FPGAs or AI accelerators may be selected for customized acceleration.
[0043] Memory is the main memory in a computing system, usually having a relatively high bandwidth and low latency, with a fast access speed. The capacity of memory is relatively large. For example, the memory capacity of a server can reach the level of hundreds of gigabytes (GB) or even terabytes (TB). While video memory, as the dedicated memory of a GPU, although having a higher bandwidth, usually has a smaller storage capacity. For example, the video memory capacity of a GPU is usually between a dozen GB (gigabytes) and dozens of GB, such as 16GB, 32GB, etc. The storage capacities of other accelerators besides GPUs are also relatively small. For example, the High Bandwidth Memory (HBM) of a TPU can be 16GB, 32GB, 64GB or 128GB, etc. For the convenience of description, here the main memory with a large storage capacity is called "host memory", and the video memory of a GPU or the small-capacity memory of other accelerators is called "device memory".
[0044] Since a large-scale model may have up to hundreds of billions of parameters, a single machine with a single card is no longer capable of handling the training tasks of such large-scale models. Therefore, distributed training is widely used for training large-scale models in artificial intelligence. Distributed training refers to using multiple machines to work together to accelerate the training process of large deep learning models. In distributed training, the original model, dataset, and training process are decomposed and distributed to multiple machines for simultaneous processing, thus effectively utilizing more computing resources and shortening the training time.
[0045] In the distributed training of large models, common parallel technologies include data parallelism and model parallelism. Model parallelism includes tensor parallelism and pipeline parallelism.
[0046] The data parallelism mode means that the training dataset of the model is divided into multiple sub-datasets. The multiple divided sub-datasets are distributed to multiple devices, and each device holds a complete copy of the model, so that it can independently train the sub-dataset allocated to it. For example, each device can independently complete the forward propagation and backward propagation calculations of the sub-dataset.
[0047] Figure 1A It is a schematic diagram of a data parallelism mode. As Figure 1AAs shown, a training dataset is split into three parts and assigned to Device 0, Device 1, and Device 2 respectively, and each of Device 0, Device 1, and Device 2 has a complete model. After backpropagation is completed on each device, the gradients calculated on each device need to be aggregated (All Reduce) and the global model parameters are updated.
[0048] Figure 1B It is a schematic diagram of a tensor parallelism mode. Tensor parallelism means splitting a layer (or operator) of a model and placing partial weights of the operator on different devices, thereby reducing the memory occupation on each device (e.g., GPU). A neural network model includes multiple layers ( Figure 1B Four layers are shown as an example in the figure). Each layer can be understood as a function that can perform specific mathematical operations on the input tensor. For example, a fully connected layer can perform a linear transformation operation on the tensor, a convolutional layer can perform a convolution operation on the tensor, a pooling layer can perform a pooling operation on the tensor, etc., to extract or transform features. These operations can change the content and shape of the tensor and achieve a gradual transformation from the original data to a high-level abstract representation. As Figure 1B shown, in the tensor parallelism mode, the tensor of a layer is split into multiple parts, and the multiple split partial tensors are respectively assigned to multiple devices, such as Device 0, Device 1, and Device 2. Each device is only responsible for a part of the model's calculations and exchanges necessary intermediate results through communication.
[0049] Pipeline parallelism means dividing different layers of a model into multiple stages, and the multiple stages are respectively assigned to multiple devices to form a pipeline. Activation values and gradients are passed sequentially between different devices. For example, during the forward propagation process, each device passes the activation value to the device where the next pipeline stage is located. During the backpropagation process, each device passes the gradient back to the device where the previous pipeline stage is located, so that multiple devices can be utilized for calculations at the same time, improving the calculation efficiency.
[0050] Figure 1C It is a schematic diagram of a pipeline parallelism mode. As Figure 1C shown, a model includes four layers. These four layers are divided into three stages, and each stage is assigned to a device. For example, the first layer is assigned to Device 0, the second layer is assigned to Device 1, and the third layer and the fourth layer are assigned to Device 2.
[0051] Figure 1D It is a schematic diagram of a micro-batch pipeline parallelism mode. By dividing the global batch data into micro-batch data and artificially creating a pipeline to solve the problem of device idle time, different devices are allowed to participate in the calculation process simultaneously, which can significantly improve the utilization rate of pipeline parallel devices and reduce the time of the device in the idle state.
[0052] As Figure 1D shown, multiple layers of the target model are split into 4 stages, and each stage includes, for example, 4 micro-batch data. F i,j represents the forward propagation of the j-th micro-batch data in the i-th stage, and B x,y represents the backward propagation of the y-th micro-batch data in the x-th stage. During the forward calculation process, the calculation of the subsequent stage depends on the calculation of the previous stage. For example, F 1,0 must start after F 0,0 is calculated. Similarly, during the backward calculation process, the calculation of the previous stage depends on the calculation of the subsequent stage. For example, B 2,3 must start after B 3,3 is calculated.
[0053] In the pipeline parallel mode, data is transmitted between adjacent devices through a communication link. For example, during the forward calculation process, the input data first obtains an intermediate result through the calculation of the first layer on device 0, and transmits the intermediate result to device 1. Then, the output of the second layer is calculated on device 1, and the output result of the second layer of the model is transmitted to device 2. The final result of the forward calculation is obtained through the calculation of the last layer on device 2. The backward propagation process is similar. Finally, the network layers on each device will update the parameters using the gradients calculated during the backward propagation process.
[0054] In Figure 1D the pipeline parallel mode shown, the backward calculation cannot start until all devices have completed the forward calculation. This results in the intermediate result of the forward calculation in the first stage needing to be retained until all stages of the forward calculation are completed before it can be released, which leads to a large number of pipeline bubbles (idle cycles or gaps that appear in the pipeline). Excessive bubbles will cause the processor performance to decline. Therefore, reducing pipeline bubbles becomes a key goal for optimizing processor performance.
[0055] Figure 1E It is a schematic diagram of cross-executing the forward calculation and backward calculation of micro-batch data. As Figure 1E shown, the forward calculation and backward calculation of different micro-batch data are alternately performed in a cross manner, which reduces pipeline bubbles to a certain extent.
[0056] As Figure 1E shown, the previous stage saves the activation values of more micro-batch data compared to the subsequent stage. For example, stage 3 only needs to save the activation values of 1 micro-batch data, while stage 0 needs to save the activation values of 4 micro-batch data, which leads to an imbalance in the device memory between different stages.
[0057] Figure 1F It is a schematic diagram of recomputation. As Figure 1F shown, since backpropagation needs to rely on the intermediate results generated during forward propagation, when recomputation is not used, the activation values generated during forward propagation need to be stored in the device memory all the time. Using recomputation can save only the activation values at several recovery points and release the activation values at non-recovery points, thereby reducing the demand for video memory. During the reverse calculation process, by using the saved activation values at the recovery points as input to perform forward calculation again, the activation values required for reverse calculation can be obtained. Thus, although recomputation reduces the video memory occupancy, it will increase the additional computational amount, resulting in an increase in the calculation time.
[0058] As the model scale gradually increases, simply using a certain parallel mode often cannot meet the requirements of both device memory limitations and high computational efficiency at the same time. Therefore, for large-scale models, it is usually necessary to combine multiple parallel technologies such as data parallelism, model parallelism, and pipeline parallelism for distributed training.
[0059] Since the device memory resources are relatively precious, some data in the device memory can be transferred to the host memory first, and then transferred back to the device memory when these data are needed, thereby reducing the device memory occupancy. Specifically, the following operations are involved:
[0060] 1. Offload: The operation of transferring some data from a device (such as an accelerator like GPU, GPGPU, or TPU) to a host (such as a CPU) during training. By offloading temporarily unnecessary data, the accelerator memory space can be released to make room for other computational tasks. For example, when the model scale exceeds the GPU video memory, by offloading some data to the CPU, it is possible to support larger-scale model training.
[0061] 2. Load: The operation of loading some data from the host to the device. Through the load operation, it can be ensured that the device has the data required to execute the current computational task. For example, during training, some data may only be needed at specific stages, and they can be loaded on demand through the load operation, thereby effectively utilizing resources.
[0062] 3. Prefetch: Also known as preloading, that is, an operation of preloading. The purpose is to move data from the host to the device before the data is actually used. For example, during the backpropagation stage, prefetching can be performed asynchronously so that when a certain data is needed, the data has already been loaded onto the accelerator, reducing the waiting time for data loading.
[0063] It should be noted that loading is an operation directly serving the current request, while prefetching is an operation taken in advance to improve the response speed of future requests.
[0064] In related technologies, some asynchronous offloading-based methods determine when to perform offloading and prefetching operations by analyzing factors such as memory, video memory, and operator performance. Specifically, it is necessary to analyze the distributed parallel strategy and model structure, use a video memory simulator to estimate the amount of data to be offloaded, and adopt a manual analysis and orchestration scheduling method to optimize the data scheduling strategy, aiming to minimize the performance overhead.
[0065] However, the inventors of the present disclosure have noticed that the above methods do not provide an effective mechanism to automatically generate data scheduling strategies, nor can they achieve automatic adjustment and optimization of these strategies. The above methods adopt a one-way command execution mode and do not fully consider the actual running state of the offloading operation. Due to the lack of an effective monitoring and feedback mechanism for the execution state, once a command execution fails, it may lead to the collapse of the entire system. To avoid the occurrence of system collapse, it is necessary to ensure that the offloading system can adapt to all training scenarios, which usually requires a large amount of human and material resources. As the business requirements grow, the workload of manual management will increase exponentially, greatly limiting the flexibility and maintainability of the system.
[0066] Some other asynchronous offloading-based methods introduce independent offloading streams and loading streams, which means that offloading tasks and loading tasks are arranged to be executed in separate streams, separated from the computation stream.
[0067] However, the inventors of the present disclosure have noticed that the above methods run offloading and loading operations in separate independent streams, which may cause resource competition problems during kernel overlap. This competition will result in additional performance overhead. Moreover, this method requires at least three execution streams (computation stream, offloading stream, and loading stream), which highly depends on the support of hardware resources. However, not all hardware configurations can meet these requirements. For hardware that cannot support multiple concurrent execution streams, such as in cases of bandwidth limitation or limited computing power, this method is not only difficult to achieve the expected performance gain but may also lead to a reduction in system efficiency.
[0068] At least one embodiment of the present disclosure provides a data processing method, which includes: obtaining a current operation record according to user interface information and running state information; generating a data scheduling strategy according to historical operation records and the current operation record; and performing corresponding data transfer operations according to the data scheduling strategy.
[0069] The data processing method provided by at least one embodiment of the present disclosure can be applied to a training system, and automatically generates a data scheduling strategy by analyzing the training status and operation records in real time. This method can dynamically adjust and optimize the data scheduling strategy during the training process, without relying on a customized data scheduling analyzer or a handwritten scheduling program, and has the convenience of plug-and-play. At the same time, since the process of generating the data scheduling strategy does not depend on the distributed parallel strategy and the model structure, its scope of application is wider and it can be applied to any type of tensor. This provides strong support for the training of complex models and improves the efficiency and stability of the system.
[0070] Furthermore, by introducing a task queue and a status management mechanism, this method realizes an execution mode based on task push and subscription. On the one hand, it ensures the stability of the system and avoids the system crash caused by task execution failure. On the other hand, it decouples scheduling and execution, enabling the system to have more flexible expansion capabilities.
[0071] Figure 2 It is a schematic block diagram of a distributed training system provided by at least one embodiment of the present disclosure. For example, Figure 2 The shown distributed training system includes multiple computing nodes 0 to m, and each computing node includes a host and multiple devices 0 to n, where m and n are positive integers and can be set according to actual needs. In each computing node, data can be transmitted between the host and the devices through a communication link. For example, the devices can also be referred to as "slave devices" and can be directly connected to the host through a hardware interface, or work in cooperation with the host through a software-defined manner (such as virtualization technology) over a network.
[0072] For example, in at least one example of the embodiments of the present disclosure, the host may include a central processing unit (CPU), and the devices include accelerators such as a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an AI accelerator, an application-specific integrated circuit (ASIC), and a field-programmable gate array (FPGA). For example, the distributed training system 500 may include 8, 16, or 32 computing nodes, and each computing node may include 4 or 8 GPUs.
[0073] The data processing method provided by at least one embodiment of the present disclosure can be applied to various types of training systems, including but not limited to Figure 2 the shown distributed training system and a training system that includes only a single computing node.
[0074] Figure 3 It is a flowchart of a data processing method provided by at least one embodiment of the present disclosure. For example, as Figure 3 shown, the data processing method provided by at least one embodiment of the present disclosure includes the following steps S101 to S103.
[0075] Step S101: Obtain the current operation record according to the user interface information and the running status information.
[0076] For example, in Step S101, the user interface can be embodied as an Application Programming Interface (API), which can include a Module interface or a Tensor interface, etc. The user can call the user interface.
[0077] For example, the Module interface is used to integrate with an existing deep learning framework (such as PyTorch), and the data transfer (i.e., offloading, loading, etc.) function can be enabled with little modification to the original code. In PyTorch, a Module is the basic class for building a neural network model, which provides the basic functions required for defining, managing, and training a neural network. The above Module interface can be implemented, for example, by extending the Module class, enabling the user to specify that the data of the entire model or part of the model (such as certain modules or layers) needs to be transferred without changing the model definition.
[0078] For example, the Tensor interface allows the user to perform data transfer operations on tensors, which can be applied to any type of tensor, including but not limited to activation values, weights, optimizer states, etc. The user can flexibly embed the Tensor interface into the training system and perform specific tensor loading or unloading according to requirements.
[0079] For example, the user interface information refers to the description related to the user interface, which can include the interface name, interface type, function description, etc. The interface name can be, for example, Offload, Prefetch, Prepareoffload, Load, etc.
[0080] For example, in Step S101, the running status information can refer to the running status information of the training system, which is used to provide the status of the target model during execution, that is, the runtime status. For example, it can include information such as the current running location, the current operation (OP) name, the timestamp, or the current resource usage. For example, tools such as a Runtime profiler can be used to obtain the running status information, and the embodiments of the present disclosure are not limited thereto.
[0081] For example, the current running position is used to indicate the specific position that the target model has reached during the training process. For example, in the case of using the mini-batch technique, the current running position may include a mini-batch index, which is used to indicate the mini-batch that is currently being processed. For example, a mini-batch index of 0 means that the first mini-batch is currently being processed. For example, the target model usually consists of multiple layers, and the current running position may include a layer index, which is used to indicate the layer in the target model that is currently being processed. For example, a layer index of 2 means that the third layer is currently being processed. According to actual needs, the current running position may also include other information, such as a step index and a module index. The step index is used to indicate the current training step (Step), and the module index is used to indicate the module that is currently being processed. It should be noted that the examples given above are only some examples, and various indexes can be in the form of numbers, or in other forms such as names or a combination of the two. The embodiments of the present disclosure do not limit this.
[0082] The target model described above can be a machine learning model to be trained, such as a deep learning model. The deep learning model can be a neural network structure including multiple layers (3 layers, 4 layers, 8 layers or more layers), such as a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), etc. The structures of these neural networks usually include an input layer, a hidden layer, and an output layer. The hidden layer refers to those layers in the neural network that are located between the input layer and the output layer, and the hidden layer is also called the processing layer. The input layer is used to receive the data to be processed, such as the image to be processed, etc. The output layer is used to output the processing result, such as the processed image, etc. The processing layer can include a convolutional layer, a pooling layer, a batch normalization layer, a fully connected layer, etc. According to the different structures of the neural network, the processing layer can include different contents and combination methods. In some examples, the target model can be, for example, a large language model (LLM). The large language model is a deep learning model trained based on a large amount of text data and can understand and generate natural language text. The large language model is usually based on the Transformer architecture and learns the statistical laws and semantic information of the language from a large amount of text data through self-supervised learning, and is usually used for tasks such as text generation, translation, question answering, and summarization.
[0083] For example, in step S101, the current operation record can be obtained according to the user interface information and the running state information. The current operation record (Trace) can be regarded as a kind of log for recording the operations at the current moment. The current operation record not only records the operations performed by the user through the interface, but also records the running state information of the training system.
[0084] In some examples, the user interface information includes the interface name, and the running status information includes the current running location. An example of step S101 can be: combining the interface name and the current running location to obtain the current operation record.
[0085] For example, the interface name and the current running location can be combined into a string or a structured data format as the current operation record.
[0086] Step S102: Generate a data scheduling policy based on the historical operation record and the current operation record.
[0087] For example, in step S102, the historical operation record can be the operation record generated in the previous training step through step S101. The data scheduling policy refers to operations such as prefetching or offloading at what time. In some examples, the historical operation record can also be the operation record generated during the training of a trained model, which is applicable to the scenario of extending from single-machine training to multi-machine training. According to the similarity of each step executed during the training of the target model, the subsequent tasks to be executed can be inferred through the historical operation record and the current operation record, which will be specifically introduced later.
[0088] Step S103: Execute corresponding data transfer operations according to the data scheduling policy.
[0089] For example, in step S103, the data transfer operations can include offloading operations or prefetching operations executed between the host and the device. For example, in some examples, the data transfer operations can include offloading operations or prefetching operations for tensors executed between the CPU and the GPU.
[0090] The data processing method provided by at least one of the above embodiments of the present disclosure automatically generates a data scheduling policy by analyzing the training status and operation records in real time. This method can dynamically adjust and optimize the data scheduling policy during the training process, without relying on a customized data scheduling analyzer or a handwritten scheduling program, and has the convenience of plug-and-play. At the same time, since the process of generating the data scheduling policy does not depend on the distributed parallel policy and the model structure, its application scope is wider and it can be applied to any type of tensor. This provides strong support for the training of complex models and improves the efficiency and stability of the system.
[0091] Each operation record generated in the above manner is unique within a training step. Therefore, an operation record list can be maintained to store all the operation records within this training step.
[0092] The data processing method provided by at least one embodiment of the present disclosure may further include the following steps S104 - S105.
[0093] Step S104: Store the current operation record into the operation record list.
[0094] Step S105: Read the historical operation records from the operation record list.
[0095] For example, the above steps S104 - S105 can be executed after step S101 and before step S102.
[0096] For example, in the initial stage of training (e.g., the first training step Step1), the operation record list can be initialized. The operation record list can be initialized based on the following principle: perform offloading operations on all tensors in the target model. Further, the principle of offloading as early as possible and loading (or prefetching) as late as possible can also be adopted for the initialization of the operation record list. Through the above initialization principle, all tensors are first offloaded to release the accelerator storage resources (such as video memory), and are only reloaded back to the accelerator when they are about to be used, thereby optimizing the computing efficiency and resource utilization.
[0097] For example, assume that the target model contains three tensors t1, t2, and t3. According to the above initialization principle, the operation record list can be initialized as follows:
[0098] {{Offload_t1},{Offload_t2},{Offload_t3},{Load_t1},{Load_t2},{Load_t3}}
[0099] Among them, Offload_t1 is the first record in the operation record list, indicating the offloading of tensor t1; Load_t3 is the last record in the operation record list, indicating the loading of tensor t3.
[0100] In subsequent training steps (e.g., the second training step Step2), each operation record in the above - initialized operation record list can be regarded as a historical operation record.
[0101] According to the data processing method provided by at least one of the above - mentioned embodiments of the present application, operation records are continuously collected during the training process and stored in the operation record list. According to the similarity of each step executed during the training of the target model, the tasks to be executed subsequently can be inferred through the historical operation records and the current operation records. This inference algorithm does not need to rely on distributed strategies or model information, and can be implemented only through simple and general scheduling rules. This is because although the schedulers of different distributed configurations are different, there are high similarities in their scheduling timings. For example, the similarities in the scheduling timings can be specifically reflected in at least the following two aspects:
[0102] 1. Prefetch timing: Always perform prefetch operations one unit distance (e.g., one operation or one layer) in advance. For example, when processing the Nth layer currently, prefetch the data of the N+1th layer. This makes the data ready when the calculation reaches the time point that requires this data, thus reducing the waiting time for data loading.
[0103] 2. Optimization of redundant operations: Reduce redundant offloading or loading operations at the beginning and end stages of the training steps, and make them overlap with the computing tasks as much as possible to reduce unnecessary memory management overhead.
[0104] Therefore, in at least one embodiment of the present disclosure, the scheduling rule is independent of the specific distributed strategy, but is determined by the relative order of the current operation in the overall operation process. This means that a simple and general scheduling rule based on operation records can be applied to achieve efficient scheduling.
[0105] In the data processing method provided by at least one embodiment of the present disclosure, the data scheduling strategy may include task information, and the task information is used to generate tasks. For example, the task information records the detailed information necessary for generating tasks. According to different task information, various different types of tasks can be generated, such as offloading tasks, prefetching tasks, etc.
[0106] An example of step S102 may be: Generate the task information corresponding to the current operation record according to the relative order of the current operation record in the historical operation record.
[0107] For example, the relative position of the current operation record in the whole process can be analyzed and matched with the corresponding position in the historical operation record, so as to obtain the relative order of the current operation record in the historical operation record. By analyzing the entire historical operation record, the currently required operations can be predicted. For example, the prefetch timing can be determined according to the relative order of the current operation record in the historical operation record, or the redundant operations can be optimized.
[0108] In the data processing method provided by the above embodiments of the present disclosure, there is no need to know the specific distributed strategy or model details, but only rely on the general rules between operations (such as prefetch timing and optimization of redundant operations). Therefore, even if the configuration of the system changes, as long as the relative order between operations remains consistent, a general scheduling rule can be applied to achieve efficient scheduling.
[0109] For example, Figure 4A The following historical operation record is shown: Offload_t1, Offload_t2, Offload_t3, Load_t1, Load_t2, Load_t3, that is, in the previous training steps, the operations of offloading tensor t1, offloading tensor t2, offloading tensor t3, loading tensor t1, loading tensor t2, and loading tensor t3 are performed in sequence.
[0110] Taking the relative order of the current operation record corresponding to Offload_t3 in the historical operation record as an example, this means that the current operation is at the same position as Offload_t3 in the historical operation record. By analyzing the historical operation record, the operation that needs to be executed currently can be predicted. Specifically, in the above historical operation record, the operation after Offload_t3 is Load_t1. To optimize performance, a prefetch operation needs to be executed one unit distance in advance before the loading operation so that the data is ready when needed. Therefore, the task information corresponding to the current operation record is to prefetch tensor t1, which can be denoted as Prefetch_t1. A prefetch task for tensor t1 can be generated according to this task information.
[0111] And so on. Assuming that the relative order of the current operation record corresponds to Load_t1 in the historical operation record, the task information corresponding to the current operation record is to prefetch tensor t2, which can be denoted as Prefetch_t2; assuming that the relative order of the current operation record corresponds to Load_t2 in the historical operation record, the task information corresponding to the current operation record is to prefetch tensor t3, which can be denoted as Prefetch_t3.
[0112] For example, more in-depth automatic scheduling optimization can also be achieved based on the historical operation record and the current operation record. For example, the historical operation record can be regarded as nodes arranged in a linear order in the computation graph, and redundant operations can be eliminated through graph optimization techniques. Specifically, through graph optimization techniques, operations that do not save video memory but introduce additional overhead can be identified and eliminated, such as adjacent offload and load operations. Through graph optimization techniques, unmatched offload operations (i.e., the situation where there is only an offload operation without a corresponding load operation) can also be identified and regarded as redundant operations for elimination.
[0113] For example, Figure 4B The following historical operation record is shown: Offload_t1, Offload_t2, Offload_t3, Load_t3, Load_t2, that is, the operations of offloading tensor t1, offloading tensor t2, offloading tensor t3, loading tensor t3, and loading tensor t2 were sequentially executed in the previous training step.
[0114] By analyzing the above historical operation records based on graph optimization technology, it can be found that the adjacent operations of unloading tensor t3 and loading tensor t3 are redundant. This operation not only fails to save video memory, but also introduces additional overhead, so it can be skipped in the current operation. Moreover, for tensor t1, the above historical operation record only contains the operation of unloading tensor t1, and there is no corresponding operation of loading tensor t1, which indicates that the unloading operation is unnecessary and can also be skipped. Therefore, assuming that the relative order of the current operation record corresponds to Offload_t1 in the historical operation record, the task information corresponding to the current operation record can be empty.
[0115] In addition, Figure 4A Similar to the example, Figure 4B In the example, assuming that the relative order of the current operation record corresponds to Load_t3 in the historical operation record, the task information corresponding to the current operation record is the prefetch tensor t2, which can be recorded as Prefetch_t2.
[0116] In the traditional imperative offloading execution scheme, the offloading system sends offloading tasks to the training system unidirectionally. Due to the lack of an effective feedback mechanism for the execution status of offloading tasks, the offloading system cannot know the execution stage of the task, nor can it know whether the task is successfully executed. Once there is any deviation between the offloading system and the training system (such as an instruction error, etc.), the entire training program may crash.
[0117] In the data processing method provided in at least one embodiment of the present disclosure, a task queue and state management mechanism are introduced to implement a task-based push and subscription execution mode, which ensures the stability of the system on the one hand and decouples scheduling and execution on the other hand, so that the system has more flexible expansion capabilities. Compared with the traditional imperative solution, it is more adaptable to complex and changeable actual application scenarios, which will be specifically introduced below.
[0118] In the data processing method provided by at least one embodiment of the present disclosure, tasks in different states correspond to different task queues.
[0119] For example, each task follows a preset state machine, ensuring that the task can proceed smoothly according to the predetermined process and preventing the system from crashing due to instruction errors. The state machine will be described in detail later. Tasks in different states are managed using separate task queues, which not only helps to organize and manage tasks, but also allows real-time tracking of the status of each task, so that potential problems can be discovered and handled in a timely manner.
[0120] In the data processing method provided in at least one embodiment of the present disclosure, an example of step S103 may include the following steps S201~S202.
[0121] Step S201: In response to the task information corresponding to an offloading task or a prefetching task, perform a validity check on the task information.
[0122] For example, in step S201, it is necessary to verify whether the upcoming offloading task or prefetching task is valid to avoid executing incorrect tasks. The validity check may include, but is not limited to, resource availability, status consistency, dependencies, etc. Resource availability is used to confirm whether there are sufficient resources to perform the required operations. For example, whether there is sufficient storage space for the offloading task or whether there is enough bandwidth to support the prefetching task. Status consistency is used to ensure that the current task status matches the expectation, avoiding errors caused by mismatched statuses. Dependencies are used to check whether there are unfinished prerequisite tasks. If a task depends on the results of other tasks, it is necessary to confirm that these prerequisite tasks have been completed.
[0123] Step S202: In response to the task information passing the validity check, read the task corresponding to the task information from the corresponding task queue and execute it to implement the corresponding data transfer operation.
[0124] For example, if the task information passes the validity check, it means that the corresponding task meets the execution conditions. At this time, reading and executing the task from the corresponding task queue can implement the corresponding data transfer operation. For example, if the task information corresponds to an offloading task, read and execute the offloading task corresponding to the task information from the task queue storing offloading tasks, which can implement the data offloading operation; if the task information corresponds to a prefetching task, read and execute the prefetching task corresponding to the task information from the task queue storing prefetching tasks, which can implement the data offloading operation. For example, when executing a task, it can trigger the interaction with the underlying hardware, thereby implementing the data transfer operation related to the specific hardware.
[0125] Through the above mechanism, different types of tasks can be effectively managed and scheduled, ensuring that each task will only be executed after it meets all prerequisite conditions and passes the validity check. This method not only improves the stability and reliability of the system but also optimizes the resource utilization efficiency, making the data offloading and prefetching operations more efficient and orderly.
[0126] In the data processing method provided by at least one embodiment of the present disclosure, the following steps S203A to S203B may further be included.
[0127] Step S203A: In response to the execution of a task, generate a new task and set the status of the new task based on the status of the task.
[0128] Step S203B: Insert the new task into the corresponding task queue.
[0129] It should be noted that the task execution in step S203A refers to the task corresponding to the task information read from the corresponding task queue in the above step S202. For example, after executing an uninstall task or a prefetch task, a new task needs to be generated, and based on the status of the executed uninstall task or prefetch task, the status of the new task is set according to a preset state machine, and the new task with the set status is inserted into the task queue corresponding to its status. When generating a new task, the old task will be destroyed, which can be understood as an update of the task. This step enables the task to correctly advance according to the preset state machine process, which helps to maintain the stability of the system.
[0130] After generating the task information corresponding to the current operation record, the data processing method provided by at least one embodiment of the present disclosure may further include the following step S204.
[0131] Step S204: In response to the user interface information corresponding to the task information indicating preparation for uninstallation, insert the task corresponding to the task information into the corresponding task queue.
[0132] For example, in step S204, when it is detected that the user interface information corresponding to the task information indicates preparation for uninstallation, it means that the user has called the user interface with the interface name "prepare for uninstallation", and the generated task can be directly inserted into the task queue without performing a validity check. This step can be regarded as an initialization operation.
[0133] After generating the task information corresponding to the current operation record, the data processing method provided by at least one embodiment of the present disclosure may further include the following steps S205 to S206.
[0134] Step S205: In response to the user interface information corresponding to the task information indicating loading, perform a validity check on the task information.
[0135] Step S206: In response to the task information passing the validity check, read the task corresponding to the task information from the corresponding task queue and determine whether to return the data object of the task.
[0136] For example, in step S205, when it is detected that the user interface information corresponding to the task information indicates loading, it means that the user has called the user interface with the interface name "loading", and it is necessary to first perform a validity check on the task information and then perform subsequent operations. The relevant content of the validity check is similar to that in step S202 above and will not be elaborated here.
[0137] For example, in step S206, each task is associated with a data object to be processed, such as a tensor that needs to be unloaded or prefetched. When it is detected that the user interface information corresponding to the task information indicates loading, it means that the user has called the user interface with the name "loading", which implies that a request to return the data object corresponding to the current task has been received. Before returning the data object of the task, it is necessary to first determine whether the current is the correct time to return the data object of the task. For example, if the unloading task has not been executed yet, the data object of the task can be returned. The data object can be returned to the user or directly applied to the subsequent training process; if the unloading task or prefetch task is currently being executed, it is necessary to wait until the prefetch task is completed before returning the data object of the task to avoid process errors.
[0138] Figure 5A Exemplary schematic diagram of a data processing method provided by at least one embodiment of the present disclosure. Figure 5A Shows an example of the above steps S201 - S206.
[0139] For example, as Figure 5A shown, when it is detected that the user interface information corresponding to the task information indicates ready to unload, the task generated according to the task information is directly inserted into the task queue without performing a validity check.
[0140] For example, as Figure 5A shown, query the task to be executed according to the data scheduling policy as follows: If the task information corresponds to an unloading task, it is necessary to first perform a validity check on the task information. If the task information passes the validity check, read the unloading task corresponding to the task information from the task queue storing the unloading tasks and execute it to implement the unloading operation performed between the host and the device. When the unloading task is executed, a new task is generated, and the state of the new task is set according to the preset state machine based on the state of the unloading task, and the new task with the set state is inserted into the task queue corresponding to its state.
[0141] If the task information corresponds to a prefetch task, it is necessary to first perform a validity check on the task information. If the task information passes the validity check, read the prefetch task corresponding to the task information from the task queue storing the prefetch tasks and execute it to implement the prefetch operation performed between the host and the device. When the prefetch task is executed, a new task is generated, and the state of the new task is set according to the preset state machine based on the state of the prefetch task, and the new task with the set state is inserted into the task queue corresponding to its state.
[0142] For example, as Figure 5AAs shown, when it is detected that the user interface information corresponding to the task information indicates loading, it is necessary to first perform a validity check on the task information. If the task information passes the validity check, the task corresponding to the task information is read from the corresponding task queue and it is determined whether to return the data object of the task. The specific determination method has been introduced above and will not be elaborated here.
[0143] By introducing a task queue and a validity check mechanism, this method not only improves the robustness and reliability of the system, but also makes resource management more flexible and efficient, and can adapt to complex and changing actual application scenarios.
[0144] In the data processing method provided by at least one embodiment of the present disclosure, the states of the task include a ready state, an unloading state, a prefetch state, and a completion state. The ready state indicates that the data object corresponding to the task is about to be unloaded; the unloading state indicates that the data object corresponding to the task is being unloaded; the prefetch state indicates that the data object corresponding to the task is being prefetched; the completion state indicates that the data object corresponding to the task has been prefetched and is ready to be returned. Tasks in different states can be managed by different task queues. For example, a ready queue, an unloading queue, a prefetch queue, and a completion queue can be set up to manage the tasks in the above different states respectively. In some examples, there is no need to set up a completion queue, which can be selected according to actual needs.
[0145] Correspondingly, the data processing method provided by at least one embodiment of the present disclosure may further include the following step S301.
[0146] Step S301: In response to the user interface information corresponding to the task information indicating ready to unload, initialize the state of the task to the ready state.
[0147] For example, in step S301, when it is detected that the user interface information corresponding to the task information indicates ready to unload, it means that the user has called the user interface with the interface name "ready to unload". At this time, an initialization operation is performed on the execution state of the current task, and its state is set to the ready state.
[0148] In the data processing method provided by at least one embodiment of the present disclosure, the above step S203A may include the following steps S302 to S304.
[0149] Step S302: In response to executing an unloading task and the state of the task being the ready state, generate a new task and set the state of the new task to the unloading state.
[0150] Step S303: In response to executing a prefetch task and the state of the task being the unloading state, generate a new task and set the state of the new task to the prefetch state.
[0151] Step S304: In response to the prefetch task being completed and the status of the task being the prefetch status, generate a new task and set the status of the new task to the completed status.
[0152] For example, in steps S302 - S304, the timing of executing the offloading task and the prefetch task is determined according to the data scheduling policy. When the offloading task is executed and its status is the ready status, a new task is generated and the status of the new task is set to the offloading status, which is transformed from the status of the offloading task (ready status); when the prefetch task is executed and its status is the offloading status, a new task is generated and the status of the new task is set to the prefetch status, which is transformed from the status of the prefetch task (offloading status); when the prefetch task is completed and its status is the prefetch status, a new task is generated and the status of the new task is set to the completed status, which is transformed from the prefetch status.
[0153] Correspondingly, the data processing method provided by at least one embodiment of the present disclosure may further include the following step S305.
[0154] Step S305: In response to determining whether to return the data object of the task and the status of the task being the ready status, update the status of the task to the completed status.
[0155] For example, in step S305, for the description of "determining whether to return the data object of the task", reference can be made to the description in step S206 above, which will not be elaborated here. When determining whether to return the data object of the task, if the status of the current task is the ready status, the offloading status and the prefetch status can be skipped and directly updated to the completed status, and at this time, the data object of the task can be returned.
[0156] Figure 5B It is a state machine conversion schematic diagram provided by at least one embodiment of the present disclosure. Figure 5B Shows an example of the above steps S301 - S305. As Figure 5B shown, the status of the task includes the ready status, the offloading status, the prefetch status, and the completed status. Four task queues (not shown), namely the ready queue, the offloading queue, the prefetch queue, and the completed queue, can be correspondingly set to manage the tasks in the above different statuses.
[0157] Refer to Figure 5A and Figure 5B, first, when it is detected that the user interface information corresponding to the task information indicates preparation for unloading, initialize the status of the task to the preparation state and directly insert it into the preparation queue. Next, query the task to be executed according to the data scheduling policy generated in the previous step, that is, when the task information corresponds to an unloading task and passes the validity check, read the task corresponding to the task information (unloading task) from the preparation queue and execute it, generate a new task, set the status of the new task to the unloading state, and insert the new task into the unloading queue; further, continue to query the task to be executed according to the data scheduling policy, that is, when the task information corresponds to a prefetch task and passes the validity check, read the task corresponding to the task information (prefetch task) from the unloading queue and execute it, generate a new task, set the status of the new task to the prefetch state, and insert the new task into the prefetch queue. Finally, when the prefetch task is completed, take the task out of the prefetch queue, generate a new task, set the status of the new task to the completed state, and insert the new task into the completed queue.
[0158] For example, when it is detected that the user interface information corresponding to the task information indicates loading, it is necessary to first perform a validity check on the task information. If the task information passes the validity check, read the task corresponding to the task information from the corresponding task queue and determine whether to return the data object of the task. When determining whether to return the data object of the task, there are the following two cases according to the different task statuses:
[0159] If the status of the current task is the preparation state, the unloading state and the prefetch state can be skipped and directly updated to the completed state to avoid redundant calculations. When the status of the task is the completed state, the data object of the task can be returned.
[0160] If the status of the current task is the unloading state or the prefetch state, it is necessary to follow the Figure 5B process defined by the state machine shown until the prefetch task is completed and the task status is updated to the completed state, then the data object of the task can be returned to ensure data integrity and consistency.
[0161] Figure 6A The schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure. The data processing device provided by at least one embodiment of the present disclosure can be embedded in various types of training systems, including but not limited to the Figure 2 distributed training system shown and the training system that only includes a single computing node. The data processing device provided by at least one embodiment of the present disclosure can also be independently deployed and interact with the training system through a communication connection, which can be selected according to actual needs.
[0162] For example, as Figure 6AAs shown in the figure, the data processing device 600 provided by the embodiments of the present disclosure includes an acquisition module 601, a generation module 602, and an execution module 603.
[0163] For example, the acquisition module 601 is configured to acquire the current operation record according to the user interface information and the running status information. For the relevant content of the acquisition module 601, reference can be made to the relevant description of step S101 in the embodiments of the above data processing method, which will not be elaborated here.
[0164] For example, the generation module 602 is configured to generate a data scheduling policy according to the historical operation record and the current operation record. For the relevant content of the generation module 602, reference can be made to the relevant description of step S102 in the embodiments of the above data processing method, which will not be elaborated here.
[0165] For example, the execution module 603 is configured to execute the corresponding data transfer operation according to the data scheduling policy. For the relevant content of the execution module 603, reference can be made to the relevant description of step S103 in the embodiments of the above data processing method, which will not be elaborated here.
[0166] For example, in at least one embodiment of the present disclosure, the user interface information includes the interface name, and the running status information includes the current running location. The acquisition module 601 is further configured to: combine the interface name and the current running location to obtain the current operation record.
[0167] For example, in at least one embodiment of the present disclosure, the acquisition module 601 is further configured to: store the current operation record in the operation record list; the generation module 602 is further configured to: read the historical operation record from the operation record list.
[0168] For example, in at least one embodiment of the present disclosure, the data scheduling policy includes task information, and the task information is used to generate a task. The generation module 602 is further configured to: generate the task information corresponding to the current operation record according to the relative order of the current operation record corresponding to the historical operation record.
[0169] For example, in at least one embodiment of the present disclosure, the execution module 603 includes a task management unit, and the task management unit is configured to: in response to the task information corresponding to an uninstall task or a prefetch task, perform a validity check on the task information; and in response to the task information passing the validity check, read the task corresponding to the task information from the corresponding task queue and execute it to implement the corresponding data transfer operation.
[0170] For example, in at least one embodiment of the present disclosure, the data processing device 600 may further include an update module. The update module is configured to: in response to the execution of a task, generate a new task, and set the status of the new task based on the status of the task; and insert the new task into the corresponding task queue.
[0171] For example, in at least one embodiment of the present disclosure, the status of a task includes a ready status, an offloading status, a prefetch status, or a completed status. The update module is further configured to: in response to executing an offloading task and the status of the task being the ready status, generate a new task and set the status of the new task to the offloading status; in response to executing a prefetch task and the status of the task being the offloading status, generate a new task and set the status of the new task to the prefetch status; and in response to the completion of the execution of the prefetch task and the status of the task being the prefetch status, generate a new task and set the status of the new task to the completed status.
[0172] For example, in at least one embodiment of the present disclosure, the update module is further configured to: in response to the user interface information corresponding to the task information indicating ready for offloading, initialize the status of the task to the ready status.
[0173] For example, in at least one embodiment of the present disclosure, the update module is further configured to: in response to determining whether to return the data object of the task and the status of the task being the ready status, update the status of the task to the completed status.
[0174] For example, in at least one embodiment of the present disclosure, the task management unit may be configured to: in response to the user interface information corresponding to the task information indicating ready for offloading, insert the task corresponding to the task information into the corresponding task queue.
[0175] For example, in at least one embodiment of the present disclosure, the task management unit may be configured to: in response to the user interface information corresponding to the task information indicating loading, perform a validity check on the task information; in response to the task information passing the validity check, read the task corresponding to the task information from the corresponding task queue and determine whether to return the data object of the task.
[0176] For example, in at least one embodiment of the present disclosure, tasks in different statuses correspond to different task queues.
[0177] For example, in at least one embodiment of the present disclosure, the data transfer operation includes an offloading operation or a prefetch operation for tensors performed between the CPU and the GPU.
[0178] For example, in at least one embodiment of the present disclosure, the data processing device 600 may further include a low-level execution module, configured to directly interact with the hardware according to the tasks executed by the task management unit to implement data transfer operations related to specific hardware. The low-level execution unit may also implement a stream synchronization function. There may be multiple streams in the system (such as a computing stream, a data transfer stream, etc.), and the low-level execution unit may manage and coordinate the dependency relationships between multiple streams to ensure data consistency and correctness.
[0179] It should be noted that the above various modules and units can be implemented by software, hardware, firmware, or any combination thereof. For example, the acquisition module, generation module, and execution module can be respectively implemented as an acquisition circuit, a generation circuit, and an execution circuit. The embodiments of the present disclosure do not limit their specific implementation manners.
[0180] It should be understood that the data processing device 600 provided by at least one embodiment of the present disclosure can be used to implement the foregoing data processing method, and can also achieve technical effects similar to those of the foregoing data processing method, which will not be elaborated herein.
[0181] It should be noted that in the embodiments of the present disclosure, the data processing device 600 may include more or fewer modules or units, and the connection relationship between the various modules or units is not limited and can be determined according to actual needs. The specific composition manner of each module or unit is not limited and can be composed of analog devices according to circuit principles, or can be composed of digital chips, or in other applicable ways.
[0182] Figure 6B It is a schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure. For example, Figure 6B It is a specific example of the data device provided by the foregoing embodiments of the present disclosure.
[0183] For example, as Figure 6B shown, the data processing device provided by at least one embodiment of the present disclosure is embedded in a training system and includes an acquisition module, an execution module, a generation module, and a bottom-layer execution module. The acquisition module obtains user interface information and operating status information from the training system, and according to the user interface information and operating status information, obtains the current operation record and stores the obtained current operation record in the operation record list. The generation module reads the historical operation record from the operation record list and generates a data scheduling policy according to the historical operation record and the current operation record. The data scheduling policy is embodied as task information. The execution module maintains the operation record list and the task queue and realizes the function of the coordination center. The execution module includes a task management unit. The task management unit receives the task information generated by the generation module and can implement functions such as validity check, updating the task queue, and executing tasks. For specific details, reference can be made to the description of the foregoing embodiments, which will not be elaborated herein. The bottom-layer execution module directly interacts with the hardware according to the tasks executed by the task management unit to implement data transfer operations (prefetching, unloading operations) related to specific hardware. It should be noted that Figure 6B The shown data processing device is only an example, and modules therein can be added or reduced according to actual needs. The embodiments of the present disclosure do not limit this.
[0184] Figure 7 It is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.
[0185] For example, as Figure 7 shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. The one or more computer program modules are stored in the memory 702 and configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for performing the above data processing method. When executed by the at least one processor 701, one or more steps in the data processing method provided by at least one embodiment of the present disclosure can be executed. The memory 702 and the processor 701 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0186] For example, the processor 701 can be a central processing unit (CPU), a digital signal processor (DSP), an image processor (GPU), a general-purpose graphics processing unit (GPGPU), an artificial intelligence (AI) accelerator, or other forms of processing units with data processing capabilities and / or program execution capabilities, such as a field-programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) can be of the X86, ARM, RISC-V architecture, etc. The processor 701 can be a general-purpose processor or a dedicated processor, and can control other components in the electronic device 700 to perform desired functions.
[0187] For example, the memory 702 can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.
[0188] Figure 8 Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.
[0189] The electronic device in at least one embodiment of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8The illustrated electronic device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0190] The electronic device includes at least one processor and a memory. The processor here may be referred to as the processing device 801 described below, and the memory may include at least one of the read-only memory (ROM), random access memory (RAM), and storage device 808 described below. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0191] As Figure 8 shown, the electronic device 800 may include a processing device 801 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) or loaded from the storage device 808 into the random access memory (RAM). In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, ROM 802, and RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface is also connected to the bus 804.
[0192] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a display, a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0193] In particular, according to at least one embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, at least one embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of at least one embodiment of the present disclosure are performed.
[0194] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In at least one embodiment of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In at least one embodiment of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0195] The above computer-readable medium can be included in the above electronic device 800; or it can exist separately without being assembled into the electronic device 800.
[0196] Figure 9 A schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.
[0197] For example, as Figure 9 shown, computer-readable instructions 901 are stored on a non-transitory computer-readable storage medium 900, and when the computer-readable instructions 901 are executed by at least one processor, one or more steps of the above data processing method are executed.
[0198] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, and may also be other applicable storage media. For example, the readable storage medium may also be the Figure 7 memory 702 in, and the related description can refer to the foregoing content and will not be elaborated here.
[0199] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, based on the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure all fall within the scope claimed by the present disclosure.
[0200] For the present disclosure, the following points also need to be noted:
[0201] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0202] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of the layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.
[0203] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0204] The above is only the specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A data processing method, applied to a training system, comprising: Obtaining a current operation record according to user interface information and running status information, wherein the current operation record includes operations performed through the interface and the running status information of the training system; Generating a data scheduling policy according to the historical operation record and the current operation record, wherein the data scheduling policy includes the timing of performing a prefetch operation or an offloading operation on the data; Performing a corresponding data transfer operation according to the data scheduling policy, wherein the data transfer operation includes an offloading operation or a prefetch operation performed between a host and a device.
2. The data processing method according to claim 1, wherein, The user interface information includes an interface name, and the running status information includes a current running location. The obtaining of the current operation record according to the user interface information and the running status information includes: Combining the interface name and the current running location to obtain the current operation record.
3. The data processing method according to claim 1, further comprising: Storing the current operation record in an operation record list; Reading the historical operation record from the operation record list.
4. The data processing method according to claim 1, wherein, The data scheduling policy includes task information for generating a task. The generating of the data scheduling policy according to the historical operation record and the current operation record includes: Generating the task information corresponding to the current operation record according to the relative order of the current operation record corresponding to the historical operation record.
5. The data processing method according to claim 4, wherein, The performing of the corresponding data transfer operation according to the data scheduling policy includes: In response to the task information corresponding to an offloading task or a prefetch task, performing a validity check on the task information; In response to the task information passing the validity check, reading the task corresponding to the task information from the corresponding task queue and executing it to implement the corresponding data transfer operation.
6. The data processing method according to claim 5, further comprising: In response to the execution of the task, generating a new task and setting the status of the new task based on the status of the task; Inserting the new task into the corresponding task queue.
7. The data processing method according to claim 6, wherein, The status of the task includes a ready status, an offloading status, a prefetch status, or a completed status. The generating of a new task in response to the execution of the task and setting the status of the new task based on the status of the task includes: In response to the execution of an offloading task and the status of the task being the ready status, generating a new task and setting the status of the new task to the offloading status; In response to the execution of a prefetch task and the status of the task being the offloading status, generating a new task and setting the status of the new task to the prefetch status; In response to the completion of the execution of the prefetch task and the status of the task being the prefetch status, generating a new task and setting the status of the new task to the completed status.
8. The data processing method according to claim 7, further comprising: In response to the user interface information corresponding to the task information indicating ready for offloading, initializing the status of the task to the ready status.
9. The data processing method according to claim 7, further comprising: In response to determining whether to return the data object of the task and the status of the task being in the ready state, update the status of the task to the completed state.
10. The data processing method according to claim 4, wherein, After generating the task information corresponding to the current operation record, the method further includes: In response to the user interface information corresponding to the task information indicating preparation for unloading, insert the task corresponding to the task information into the corresponding task queue.
11. The data processing method according to claim 4, wherein, After generating the task information corresponding to the current operation record, the method further includes: In response to the user interface information corresponding to the task information indicating loading, perform a validity check on the task information; In response to the task information passing the validity check, read the task corresponding to the task information from the corresponding task queue and determine whether to return the data object of the task.
12. The data processing method according to any one of claims 5-11, wherein, Tasks in different states correspond to different task queues.
13. The data processing method according to claim 1, wherein, The data transfer operation includes an unloading operation or a prefetch operation for tensors performed between a central processing unit and a graphics processing unit.
14. A data processing device, applied to a training system, includes: An acquisition module, configured to acquire a current operation record according to user interface information and running state information, where the current operation record includes operations performed through an interface and the running state information of the training system; A generation module, configured to generate a data scheduling policy according to a historical operation record and the current operation record, where the data scheduling policy includes the timing of performing a prefetch operation or an unloading operation on data; An execution module, configured to perform a corresponding data transfer operation according to the data scheduling policy, where the data transfer operation includes an unloading operation or a prefetch operation performed between a host and a device.
15. An electronic device, includes: At least one processor; At least one memory, storing one or more computer program modules; Wherein, the one or more computer program modules are configured to be executed by the at least one processor for executing the instructions of the data processing method according to any one of claims 1-13.
16. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon, wherein, When the computer-readable instructions are executed by at least one processor, they execute the data processing method according to any one of claims 1-13.
Citation Information
Patent Citations
User operation recording method and device and server
CN110874305A
Unloading training optimization method and device, electronic equipment and storage medium
CN118796514A
Process scheduling method, system and equipment based on heterogeneous multi-core processor
CN119415276A