Data processing apparatus, data processing method, electronic device, and storage medium
The data processing device and method address the challenge of processing large data volumes in heterogeneous hardware by dynamically adjusting the number of threads and executable tasks based on storage capacity, thereby optimizing resource utilization and improving efficiency.
Patent Information
- Application Number
- JP2023206152
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2043-12-06
AI Technical Summary
In heterogeneous hardware platforms, efficiently processing data with large batch sizes is challenging, particularly due to limitations in storage capacity and bandwidth, which hinders the full utilization of resources and affects data processing efficiency.
A data processing device and method that determine an initial number of threads based on the data amount and storage capacity, and adjust the number of executable tasks accordingly, ensuring optimal utilization of storage and processing resources, especially in scenarios where data volume exceeds the capacity of lower-level memory.
This approach enables efficient data processing by fully utilizing the capacity and bandwidth of higher-level storage, improving processing efficiency even when dealing with large data volumes, and effectively managing resources across different levels of memory.
Smart Images

Figure 0007691478000001 
Figure 0007691478000002 
Figure 0007691478000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of chip technology and multi-thread parallel technology. More specifically, the present disclosure provides a data processing device, a data processing method, an electronic device, and a storage medium.
Background Art
[0002] With the development of artificial intelligence technology, model inference or model training tasks can be executed in parallel.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The present disclosure provides a data processing device, a data processing method, an electronic device, and a storage medium.
Means for Solving the Problems
[0004] According to one aspect of the present disclosure, there is provided a data processing device including a processor configured to determine an initial number of threads based on a data amount of target data and a capacity of a first target storage means in response to determining that the data amount of the target data including input data to be processed, weight data to be processed, and output data is equal to or less than the capacity of the first target storage means, and to determine a first number of executable tasks based on the initial number of threads in response to determining that the initial number of threads is equal to or more than a predetermined number of threads.
[0005] According to another aspect of the present disclosure, there is provided a data processing method including determining an initial number of threads based on a data amount of target data and a capacity of a first target storage means in response to determining that the data amount of the target data including input data to be processed, weight data to be processed, and output data is equal to or less than the capacity of the first target storage means, and determining a first number of executable tasks based on the initial number of threads in response to determining that the initial number of threads is equal to or more than a predetermined number of threads.
[0006] According to another aspect of the present disclosure, there is provided an electronic device including the data processing apparatus according to the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, which cause a computer to execute the method provided by the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program that realizes the method provided by the present disclosure when executed by a processor. As will be understood, the content described in this part is not for identifying key features or important features of embodiments of the present disclosure, nor is it for limiting the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
[0010] The drawings are for better understanding of the present invention and do not limit the present disclosure.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, which are understood to be merely exemplary. Therefore, as understood by those skilled in the art, various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of known functions and structures are omitted in the following description.
[0013] In the process of performing inference using a deep learning model, based on the task mapping policy of a directed acyclic graph (DAG), the model inference task can be divided into a plurality of subtasks. Different subtasks can be placed in different task queues, and the plurality of subtasks can be executed in parallel. After the execution of the plurality of subtasks is completed, the execution result of the model inference task can be obtained.
[0014] The model inference task may be executed using a heterogeneous hardware platform including a central processing unit (CPU) and a graphics processing unit (GPU). Based on the execution relationship information of different tasks, a directed acyclic graph can be determined, and the parallel processing ability of the graphics processor can be combined with the directed acyclic graph to improve the usage efficiency of heterogeneous devices. For example, the model inference task includes a large number of matrix operations. The graphics processing unit can significantly reduce the time of matrix calculation and improve the execution efficiency of the model inference task.
[0015] Based on the dependency relationships between different data, the resources of heterogeneous hardware platforms can be scheduled. However, for heterogeneous hardware platforms, it is difficult to further divide model inference tasks. For example, it is difficult to efficiently process data with a large batch size.
[0016] In a heterogeneous hardware platform, the central processing unit may be used as the host, and the graphics processor may be used as the device. Data can be transmitted from the host side to the device side. For example, data can be respectively transmitted to a plurality of graphics processor cores on the device side, and parallel acceleration of matrix calculation can be realized by using the graphics processor. The graphics processor may include different levels of high-speed dynamic random access memory (DRAM). For example, the graphics processor may include level 0 high-speed dynamic random access memory (L0), level 1 high-speed dynamic random access memory (L1), level 2 high-speed dynamic random access memory (L2), level 3 high-speed dynamic random access memory (L3), and level 4 high-speed dynamic random access memory (L4). The level 4 high-speed dynamic random access memory has the largest capacity. All processor cores can read and write data from the level 4 high-speed dynamic random access memory. It is difficult to store the weight data or input data of the deep learning model because the capacities from the level 0 high-speed dynamic random access memory to the level 2 high-speed dynamic random access memory are small.
[0017] The bandwidth of the level 3 high-speed dynamic random access memory may be, for example, twice that of the level 4 high-speed dynamic random access memory. In order to fully utilize the capacity and bandwidth of the level 3 high-speed dynamic random access memory, the present disclosure provides a data processing device, which will be described below.
[0018] FIG. 1 is a schematic block diagram of a data processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 1, the data processing apparatus 100 may include a first target storage means 110 and a processor 120.
[0019] The first target storage means 110 may be a level 3 high-speed dynamic random access memory.
[0020] The processor 120 may be configured to determine an initial number of threads based on the data amount of the target data and the capacity of the first target storage means in response to determining that the data amount of the target data is equal to or less than the capacity of the first target storage means. In response to determining that the initial number of threads is equal to or greater than a predetermined number of threads, the first number of executable tasks is determined based on the initial number of threads.
[0021] In an embodiment of the present disclosure, the processor may include a plurality of processor cores. The processor may be at least one of processors such as a graphics processor and a neural network processor (NeuralnetworkProcessingUnit, NPU).
[0022] In an embodiment of the present disclosure, the target data includes input data to be processed, weight data to be processed, and output data. For example, the input data to be processed may be the input data of a deep learning model. The weight data to be processed may include the weight data of each of a plurality of operators of the deep learning model. The output data may include the output data of the deep learning model.
[0023] In an embodiment of the present disclosure, it is possible to determine whether the data volume of the target data is less than or equal to the capacity of the first target storage means. For example, it is exemplified that the data volume of the target data is 16 megabytes (Mbyte). The capacity of the first target storage means may be 64 megabytes. It can be determined that the data volume of the target data is smaller than the capacity of the first target storage means. Next, based on the data volume of the target data and the capacity of the first target storage means, the initial number of threads can be determined to be 4. That is, the first storage means can store four target data.
[0024] In an embodiment of the present disclosure, the predetermined number of threads may be 1. When the initial number of threads is 4, it can be determined that the initial number of threads is larger than the predetermined number of threads. The initial number of threads can be used as the first number of executable tasks.
[0025] According to an embodiment of the present disclosure, in the model training or inference process, when the data volume is small, if a plurality of target data are stored in the first target storage means, the bandwidth and capacity of the first target storage means can be fully utilized, and the data processing efficiency can be improved.
[0026] After determining the number of executable tasks, the tasks of the number of executable tasks can be executed in parallel, which will be further described below.
[0027] In some embodiments, the processor 120 may further be configured to write the data to be processed in the number of the first executable tasks into the first target storage means. In an embodiment of the present disclosure, the data to be processed includes the input data to be processed and the weight data to be processed. For example, when the number of the first executable tasks is 4, four sets of input data to be processed and weight data to be processed can be written into the first target storage means.
[0028] In some embodiments, the processor 120 may be further configured to execute several first-executable tasks in parallel and obtain several pieces of output data corresponding to the first-executable tasks. In embodiments of the present disclosure, a task may include processing input data to be processed using weight data to be processed. For example, when the number of first-executable tasks is 4, four tasks can be executed in parallel to obtain four pieces of output data.
[0029] In some embodiments, the processor 120 may be further configured to write several pieces of output data corresponding to the first-executable tasks to the first target storage means. For example, four pieces of output data may be written to the first target storage means. In embodiments of the present disclosure, the input data to be processed stored in the first target storage means can be processed in parallel, fully exerting the parallel processing ability of the artificial intelligence chip and improving the data processing efficiency.
[0030] As can be understood, in the above, it is described by taking as an example that the data volume of the target data is less than or equal to the capacity of the first target storage means, but the present disclosure is not limited thereto. This will be further described below.
[0031] FIG. 2 is a schematic diagram of a data processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 2, the processor may be configured to execute at least one instruction to implement operation S201. In operation S201, it is determined whether the total data volume of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. For example, taking the first target data as an example, the data volume of the first target data is 16 megabytes. The capacity of the first target storage means may be 64 megabytes. It may be determined that the sum of the data volume of the input data to be processed and the weight data to be processed of the first target data is smaller than the capacity of the first target storage means, and this will be further described in combination with operation S202 below.
[0032] In some embodiments, the processor may be configured to execute at least one instruction to implement operation S202 in response to determining that the total data volume of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. In operation S202, it is determined whether the data volume of the target data is less than or equal to the capacity of the first target storage means. For example, it can be determined that the data volume of the first target data is smaller than the capacity of the first target storage means.
[0033] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S210 in response to determining that the number of target data is less than or equal to the capacity of the first target storage means. In operation S210, the initial number of threads is determined based on the data volume of the target data and the capacity of the first target storage means. For example, based on the data volume of the first target data and the capacity of the first target storage means, it can be determined that the initial number of threads is 4. The first storage means can store four pieces of first target data.
[0034] Next, in an embodiment of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S221. In operation S221, it is determined whether the initial number of threads is greater than a predetermined number of threads. Taking the case where the predetermined number of threads is 1 as an example, when the initial number of threads is 4, it can be determined that the initial number of threads is greater than the predetermined number of threads.
[0035] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S222 in response to determining that the initial number of threads is greater than a predetermined number of threads. In operation S222, the first number of executable tasks is determined based on the initial number of threads. For example, the initial number of threads may be used as the first number of executable tasks. The first number of executable tasks may be 4.
[0036] As understood above, the method for determining the number of executable tasks has been described. Hereinafter, some methods for executing tasks will be described.
[0037] In an embodiment of the present disclosure, the processor may further be configured to write data to be processed in the number of first executable tasks into the first target storage means. For example, the data to be processed may include input data to be processed and weight data to be processed. When the number of first executable tasks is 4, the input data to be processed and the weight data to be processed for each of the four first target data can be written into the first target storage means respectively.
[0038] In an embodiment of the present disclosure, the processor may further be configured to execute the number of first executable tasks in parallel and obtain the number of output data corresponding to the number of first executable tasks. For example, the task may include processing the input data to be processed by the weight data to be processed. When the number of first executable tasks is 4, four tasks can be executed in parallel to obtain four output data.
[0039] In an embodiment of the present disclosure, the processor may further be configured to write the number of output data corresponding to the number of first executable tasks into the first target storage means. For example, four output data can be written into the first target storage means.
[0040] As can be understood, the present disclosure has been described above by taking the first target data as an example. Hereinafter, the present disclosure will be further described by taking the second target data as an example. The data volume of the second target data may be 64 megabytes.
[0041] In some embodiments, the data processing apparatus may further include a second target storage means, and the capacity of the second target storage means may be larger than the capacity of the first target storage means. For example, the second target storage means may be a Global Memory (GM) unit, or may be the above-mentioned level 4 high-speed dynamic random access memory.
[0042] As shown in FIG. 2, the processor may be configured to execute at least one instruction to implement operation S201. In operation S201, it is determined whether the total data amount of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. For example, the data amount of the second target data may be 64 megabytes, and the capacity of the first target storage means may be 64 megabytes. The second target data may include the input data to be processed, the weight data to be processed, and the output data. The total data amount of the input data to be processed and the output data to be processed of the second target data may be smaller than the capacity of the first target storage means.
[0043] In some embodiments, the processor may further be configured to execute at least one instruction to implement operation S202 in response to determining that the total data amount of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. In operation S202, it is determined whether the data amount of the target data is less than or equal to the capacity of the first target storage means. For example, it can be determined that the data amount of the second target data is equal to the capacity of the first target storage means.
[0044] In an embodiment of the present disclosure, the processor may further be configured to execute at least one instruction to implement operation S210 in response to determining that the number of target data is less than or equal to the capacity of the first target storage means. In operation S210, the initial number of threads is determined based on the data amount of the target data and the capacity of the first target storage means. For example, based on the data amount of the second target data and the capacity of the first target storage means, it can be determined that the initial number of threads is 1. The first storage means may store one second target data.
[0045] Next, in the embodiments of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S221. In operation S221, it is determined whether the initial number of threads is greater than a predetermined number of threads. Taking the case where the predetermined number of threads is 1 as an example, when the initial number of threads is 1, it can be determined that the initial number of threads is equal to the predetermined number of threads.
[0046] In the embodiments of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S231 in response to determining that the initial number of threads is less than or equal to a predetermined number of threads. In operation S231, a first number of tasks is determined based on the amount of resources required for the processor to process the target data. For example, the first input data to be processed and the first weight data to be processed of the first second target data are written into the first target storage means. The second input data to be processed and the second weight data to be processed of the second second target data are written into the second target storage means. The processor processes the first input data to be processed and the second input data to be processed respectively, and the amount of resources required for the processor to process the target data can be determined based on the test execution. Next, the test execution can be performed multiple times. In the i-th test execution, the number of input data to be processed in the second target storage means may be i. In the (i + 1)-th test execution, the number of input data to be processed in the second target storage means may be i + 1. At the time of the I-th test execution, the number of input data to be processed in the second target storage means may be I. When the utilization rate of the processor approaches 100% during the I-th test execution, it can be determined that the number of the first tasks is I. I may be an integer greater than 1, and i may be an integer greater than or equal to 1 and less than I.
[0047] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction so as to implement operation S232. In operation S232, based on the number of first tasks and the number of initial threads, the number of second executable tasks is determined. For example, when the number of first tasks is I and the number of initial threads is 1, it can be determined that the number of second executable tasks is I + 1. According to the embodiment of the present disclosure, when the amount of data is large, a plurality of target data are respectively stored in the first target storage means and the second target storage means, the bandwidth of the first target storage means can be fully utilized, the capacity of the second target storage means can be fully utilized, which is helpful for further improving the data processing efficiency.
[0048] As can be understood, in the above, the method for determining the number of executable tasks has been described. Hereinafter, several methods for executing tasks will be described.
[0049] In an embodiment of the present disclosure, the processor may further be configured to write the data to be processed in the number of initial threads into the first target storage means. The data to be processed includes the input data to be processed and the weight data to be processed. For example, the input data to be processed and the weight data to be processed of one second target data can be written into the first target storage means.
[0050] In an embodiment of the present disclosure, the processor may further be configured to write the data to be processed in the number of first tasks into the second target storage means. For example, the input data to be processed and the weight data to be processed of each of I second target data can be written into the second target storage means.
[0051] In an embodiment of the present disclosure, the processor may further be configured to execute the number of second executable tasks in parallel and obtain the number of output data corresponding to the number of second executable tasks. The task includes processing the input data to be processed using the weight data to be processed. For example, I + 1 tasks can be executed in parallel to obtain I + 1 output data.
[0052] In an embodiment of the present disclosure, the processor may be configured to write a number of output data equal to the initial number of threads to the first target storage means. For example, one output data may be written to the first target storage means, and the output data may correspond to the input data to be processed in the first target storage means.
[0053] In an embodiment of the present disclosure, the processor may further be configured to write a number of output data equal to the number of first tasks to the second target storage means. For example, I output data may be written to the second target storage means, and each of the I output data may correspond to I input data to be processed in the second target storage means. According to an embodiment of the present disclosure, the input data to be processed stored in the first target storage means and the second target storage means can be processed in parallel, further exerting the parallel processing ability of the artificial intelligence chip and improving the data processing efficiency.
[0054] As can be understood, above, the present disclosure has been described by taking the second target data as an example. Hereinafter, the present disclosure will be further described by taking the third target data as an example. The data volume of the third target data may be larger than 64 megabytes.
[0055] As shown in FIG. 2, the processor may be configured to execute at least one instruction so as to implement operation S201. In operation S201, it is determined whether the total data volume of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. For example, the data volume of the third target data may be larger than 64 megabytes, and the capacity of the first target storage means may be 64 megabytes. The third target data includes the input data to be processed, the weight data to be processed, and the output data. The total data volume of the input data to be processed and the output data of the third target data may be smaller than the capacity of the first target storage means.
[0056] In some embodiments, the processor may be configured to execute at least one instruction so as to implement operation S202 in response to determining that the total amount of input data to be processed and the amount of output data to be processed is less than or equal to the capacity of the first target storage means. In operation S202, it is determined whether the amount of target data is less than or equal to the capacity of the first target storage means. For example, it can be determined that the amount of the third target data is greater than the capacity of the first target storage means.
[0057] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction so as to implement operation S240 in response to determining that the number of target data is greater than the capacity of the first target storage means. In operation S240, based on the amount of resources required for the processor to process the target data, the number of third executable tasks is determined. For example, taking the three third target data as an example, the input data to be processed and the weight data to be processed of the three third target data can be written into the second target storage means. The processor processes the input data to be processed of the three third target data respectively, and determines the amount of resources required for the processor to process the third target data. If the processor utilization rate required for the processor to process one third target data is 5%, it can be determined that the number of third executable tasks is 20. As can be understood, the amount of resources required to process the number of third target data of the number of third executable tasks does not exceed the total amount of all resources of the processor. For example, the processor utilization rate required to process the number of third target data of the number of third executable tasks does not exceed 100%. According to the embodiment of the present disclosure, when the data volume is large, if a plurality of target data are stored in the second target storage means, the capacity of the second target storage means can be fully utilized, and the data processing efficiency can be improved.
[0058] As can be understood, in the above, the method for determining the number of executable tasks has been described. In the following, some methods for executing tasks will be described.
[0059] In an embodiment of the present disclosure, the processor may further be configured to write data to be processed for a number of third executable tasks into the second target storage means. The data to be processed includes input data to be processed and weight data to be processed. For example, when the number of third executable tasks is 20, the input data to be processed and the weight data to be processed for each of the 20 third target data can be written into the second target storage means.
[0060] In an embodiment of the present disclosure, the processor may further be configured to execute a number of third executable tasks in parallel and obtain output data for a number of third executable tasks. The tasks include processing input data to be processed using weight data to be processed. For example, 20 tasks can be executed in parallel to obtain 20 output data.
[0061] In an embodiment of the present disclosure, the processor may further be configured to write output data for a number of third executable data into the second target storage means. For example, 20 output data may be written into the second target storage means. According to an embodiment of the present disclosure, the input data to be processed stored in the second target storage means can be processed in parallel, fully exerting the parallel processing ability of the artificial intelligence chip and improving the data processing efficiency.
[0062] As understood, above, the present disclosure has been described by taking the third target data as an example. Hereinafter, the present disclosure will be further described by taking the fourth target data as an example. The total amount of the input data to be processed and the weight data to be processed of the fourth target data may be greater than 64 megabytes.
[0063] As shown in FIG. 2, the processor may be configured to execute at least one instruction to implement operation S201. In operation S201, it is determined whether the total data volume of the input data to be processed and the output data to be processed is less than or equal to the capacity of the first target storage means. For example, the capacity of the first target storage means may be 64 megabytes. The total data volume of the input data and the output data to be processed for the fourth target data may be greater than the capacity of the first target storage means.
[0064] In some embodiments, the processor may be configured to execute at least one instruction to implement operation S202 in response to determining that the total data volume of the input data to be processed and the output data to be processed is greater than the capacity of the first target storage means. In operation S202, it is determined whether the data volume of the target data is less than or equal to the capacity of the first target storage means. For example, it can be determined that the data volume of the third target data is greater than the capacity of the first target storage means.
[0065] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction so as to implement operation S251 in response to determining that the number of target data is greater than the capacity of the first target storage means. In operation S251, the input data to be processed is divided into a plurality of sub-input data to be processed. For example, the input data to be processed for the fourth target data may be input image data to be processed. The scale (shape) of the input image data to be processed may be [n, c, h, w]. n is the batch size and may indicate the number of images in the input image data to be processed. c is the number of channels of the image and may be, for example, 3. h may be the height of the image and w may be the width of the image. Based on the batch size of the input image data to be processed, the input image data to be processed can be divided into a plurality of input image sub-data to be processed. When n is 64, the input image data to be processed may include 64 images. The input image data to be processed may be divided into 16 input image sub-data to be processed. The batch size of each input image sub-data to be processed is 4.
[0066] In an embodiment of the present disclosure, the processor may be configured to execute at least one instruction to implement operation S252. In operation S252, based on the amount of resources required to process the input sub-data to be processed by the processor, the number of fourth executable tasks is determined. For example, the weight data to be processed and three input image sub-data to be processed may be written into the second target storage means. The processor processes the three input image sub-data to be processed respectively to determine the amount of resources required for the processor to process the input image sub-data to be processed. If the processor utilization rate required for the processor to process one input image sub-data to be processed is 6%, it can be determined that the number of fourth executable tasks is 16. As can be understood, the amount of resources required to process the input image sub-data to be processed in the number of fourth executable tasks may not exceed all the resources of the processor. For example, the processor utilization rate required to process the input image sub-data to be processed in the number of fourth executable tasks does not exceed 100%. According to the embodiment of the present disclosure, when the data volume is larger, by dividing the input data to be processed, the capacity of the second target storage means can be fully utilized, and the data processing efficiency can be improved.
[0067] As can be understood, in the above, the method for determining the number of executable tasks has been described. Hereinafter, several methods for executing tasks will be described.
[0068] In an embodiment of the present disclosure, the processor may be configured to write the weight data to be processed and the number of sub-input data to be processed in the number of fourth executable tasks into the second target storage means. For example, the weight data to be processed may be written into the second target storage means, and 16 input image sub-data to be processed may be written into the second target storage means.
[0069] In an embodiment of the present disclosure, the processor may be further configured to execute several tasks of the fourth executable task in parallel and obtain several output sub-data of the fourth executable task. The task includes processing sub-input data to be processed using weight data to be processed. For example, 16 tasks can be executed in parallel to obtain 16 output sub-data.
[0070] In an embodiment of the present disclosure, the processor may be further configured to write several output sub-data of the fourth executable number into the second target storage means. For example, 16 output sub-data may be written into the second target storage means.
[0071] In an embodiment of the present disclosure, the processor may be configured to stitch a plurality of output sub-data into output data. For example, 16 output sub-data may be stitched into output data corresponding to the input data to be processed. According to an embodiment of the present disclosure, when the data volume is larger, the data can be divided based on the batch size of the input image data to be processed, realizing dynamic data splitting and efficiently processing the data in parallel.
[0072] As can be understood, the processor core of the processor may be further configured to execute several tasks of the second task based on the data in the level 1 cache (L1 cache).
[0073] As can be understood, above, the data processing apparatus of the present disclosure has been described. Hereinafter, the data processing method of the present disclosure will be described.
[0074] FIG. 3 is a flowchart of a data processing method according to an embodiment of the present disclosure. As shown in FIG. 3, the method 300 may include operations S310 to S320.
[0075] In operation S310, in response to determining that the data amount of the target data is less than or equal to the capacity of the first target storage means, an initial number of threads is determined based on the data amount of the target data and the capacity of the first target storage means.
[0076] In an embodiment of the present disclosure, the target data includes input data to be processed, weight data to be processed, and output data.
[0077] In operation S320, in response to determining that the initial number of threads is greater than or equal to a predetermined number of threads, a first number of executable tasks is determined based on the initial number of threads.
[0078] As can be understood, the method 300 can be executed by the above-described processor 120.
[0079] In some embodiments, the method 300 may further include writing the data to be processed in the number of the first executable tasks to the first target storage means. For example, the data to be processed includes input data to be processed and weight data to be processed. Execute the number of the first executable tasks in parallel, and obtain the output data in the number of the first executable tasks. For example, the task includes processing the input data to be processed using the weight data to be processed. Write the output data in the number of the first executable tasks to the first target storage means.
[0080] In some embodiments, the capacity of the first target storage means is less than or equal to the capacity of the second target storage means.
[0081] In some embodiments, the method 300 may further include determining a first number of tasks based on the amount of resources required for the processor to process the target data in response to determining that the initial number of threads is equal to a predetermined number of threads. Based on the first number of tasks and the initial number of threads, determine a second number of executable tasks.
[0082] In some embodiments, method 300 may further include writing several initial threads' worth of data to be processed to the first target storage means. For example, the data to be processed includes the input data to be processed and the weight data to be processed. Write several first task's worth of data to be processed to the second target storage means. Execute several second executable tasks in parallel and obtain several second executable tasks' worth of output data. For example, the task includes processing the input data to be processed using the weight data to be processed. Write several initial threads' worth of output data to the first target storage means. Write several first task's worth of output data to the second target storage means.
[0083] In some embodiments, method 300 may further include determining the number of third executable tasks based on the amount of resources required for the processor to process the target data in response to determining that the data volume of the target data is larger than the capacity of the first target storage means.
[0084] In some embodiments, method 300 may further include writing several third executable tasks' worth of data to be processed to the second target storage means. For example, the data to be processed includes the input data to be processed and the weight data to be processed. Execute several third executable tasks in parallel and obtain several third executable tasks' worth of output data. For example, the task includes processing the input data to be processed using the weight data to be processed. Write several third executable tasks' worth of output data to the second target storage means.
[0085] In some embodiments, method 300 may further include splitting the input data to be processed into a plurality of sub-input data to be processed in response to determining that the total data volume of the input data to be processed and the output data is larger than the capacity of the first target storage means. Determine the number of fourth executable tasks based on the amount of resources required for the processor to process the sub-input data to be processed.
[0086] In some embodiments, method 300 may further include writing the weight data to be processed and the sub-input data to be processed of the fourth executable task number to the second target storage means. Execute the fourth executable task number of tasks in parallel and obtain the output sub-data of the fourth executable task number. For example, the task includes processing the sub-input data to be processed using the weight data to be processed. Write the output sub-data of the fourth executable number to the second target storage means. Stitch a plurality of output sub-data into output data.
[0087] FIG. 4 is a block diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 4, the device 40 may include the data processing device 400 according to the present disclosure. For example, the data processing device 400 may be the data processing device 100 described above.
[0088] In the technical solution of the present disclosure, any processing such as the collection, storage, use, processing, transmission, provision, disclosure, and application of such user personal information complies with the provisions of relevant laws and regulations, adopts necessary confidentiality, and does not violate public order and good customs. In the technical solution of the present disclosure, before obtaining or collecting user personal information, the permission or consent of the user is obtained.
[0089] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0090] FIG. 5 shows an exemplary block diagram for implementing an exemplary electronic device 500 of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The members, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed herein.
[0091] As shown in FIG. 5, the electronic device 500 includes computing means 501, which can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage means 508 into a random access memory (RAM) 503. The RAM 503 can further store various programs and data necessary for the operation of the electronic device 500. The computing means 501, the ROM 502, and the RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0092] A plurality of components in the electronic device 500 are connected to the I / O interface 505 and include input means 506 such as a keyboard, a mouse, etc., output means 507 such as various types of displays, speakers, etc., storage means 508 such as magnetic disks, optical disks, etc., and communication means 509 such as a network card, a modem, a wireless communication transceiver, etc. The communication means 509 enables the electronic device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0093] The computing means 501 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing means 501 include, but are not limited to, a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, computing means for various machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing means 501 executes processing with each of the methods described above, such as a data processing method. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage means 508. In some embodiments, part or all of the computer program may be loaded and / or installed in the electronic device 500 via the ROM 1002 and / or the communication means 509. When the computer program is loaded into the RAM 1003 and executed by the computing means 501, one or more steps of the data processing method described above may be executed. Alternatively, in another embodiment, the computing means 501 may be configured to execute the data processing method in any other suitable form (e.g., via firmware).
[0094] The various embodiments of the systems and techniques described in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and which receives data and instructions from, and transmits data and instructions to, a memory system, at least one input device, and at least one output device.
[0095] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing device, such that, when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, partially on the device, partially on the device as an independent software package, and partially on a remote device or entirely on a remote device or server.
[0096] In the context of this disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or electronic device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium include electrical connections made with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0097] To provide for interaction with a user, the computer systems and techniques described herein may be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) display or an LCD (liquid crystal display)) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices may be further provided for interaction with the user; for example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).
[0098] The systems and techniques described herein can be implemented in a computing system that includes background components (such as a data server), or a computing system that includes middleware components (such as an application server), or a computing system that includes front-end components (such as a user computer having a graphical user interface or a web browser, through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (such as a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0099] The computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server is generated by a computer program running on the corresponding computer and having a client-server relationship. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0100] It should be understood that various forms of the flows shown above may be used, and the steps may be sorted, added, or deleted again. For example, each step described in the present invention may be executed in parallel, sequentially, or in a different order, and this specification is not limited herein as long as the desired results of the technical solution of the present disclosure can be achieved.
[0101] The foregoing specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and alternatives can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. A data processing apparatus, a first target storage means, in response to determining that the data volume of target data including input data to be processed, weight data to be processed, and output data is less than or equal to the capacity of the first target storage means, determines an initial number of threads based on the data volume of the target data and the capacity of the first target storage means, a processor configured to determine a first number of executable tasks based on the initial number of threads in response to determining that the initial number of threads is equal to or greater than a predetermined number of threads, a second target storage means having a capacity greater than the capacity of the first target storage means, wherein the processor further in response to determining that the initial number of threads is equal to a predetermined number of threads, determines a first number of tasks based on the amount of resources required for the processor to process the target data, and determines a second number of executable tasks based on the first number of tasks and the initial number of threads configured as a data processing apparatus.
2. wherein the processor further writes the data to be processed including the input data to be processed and the weight data to be processed, in an amount equal to the first number of executable tasks, to the first target storage means, executes in parallel tasks including processing the input data to be processed using the weight data to be processed, in an amount equal to the first number of executable tasks, to obtain the output data in an amount equal to the first number of executable tasks, and writes the output data in an amount equal to the first number of executable tasks to the first target storage means configured as the apparatus according to claim 1.
3. wherein the processor further writes the data to be processed including the input data to be processed and the weight data to be processed, in an amount equal to the initial number of threads, to the first target storage means, writes the data to be processed, in an amount equal to the first number of tasks, to the second target storage means, executes in parallel tasks including processing the input data to be processed using the weight data to be processed, in an amount equal to the second number of executable tasks, to obtain the output data in an amount equal to the second number of executable tasks, writes the output data in an amount equal to the initial number of threads to the first target storage means, and writes the output data in an amount equal to the first number of tasks to the second target storage means configured as the apparatus according to claim 1.
4. wherein the processor further In response to determining that the data amount of the target data is larger than the capacity of the first target storage means, based on the amount of resources required for the processor to process the target data, determine the number of third executable tasks configured as The apparatus according to claim 1
5. The processor further Write the data to be processed, including the input data to be processed and the weight data to be processed, corresponding to the number of third executable tasks, into the second target storage means Parallelly execute tasks including processing the input data to be processed using the weight data to be processed corresponding to the number of third executable tasks, and obtain the output data corresponding to the number of third executable tasks Write the output data corresponding to the number of third executable tasks into the second target storage means configured as The apparatus according to claim 4
6. The processor further In response to determining that the total data amount of the input data to be processed and the output data is larger than the capacity of the first target storage means, divide the input data to be processed into a plurality of sub-input data to be processed Based on the amount of resources required for the processor to process the sub-input data to be processed, determine the number of fourth executable tasks configured as The apparatus according to claim 1
7. The processor Write the weight data to be processed and the sub-input data to be processed corresponding to the number of fourth executable tasks into the second target storage means Parallelly execute tasks including processing the sub-input data to be processed using the weight data to be processed corresponding to the number of fourth executable tasks, and obtain the output sub-data corresponding to the number of fourth executable tasks Write the output sub-data corresponding to the number of fourth executable tasks into the second target storage means Stitch the plurality of output sub-data into output data configured as The apparatus according to claim 6
8. A data processing method executed by a processor, comprising In response to determining that the data amount of the target data including the input data to be processed, the weight data to be processed, and the output data is less than or equal to the capacity of the first target storage means, based on the data amount of the target data and the capacity of the first target storage means, determine the initial number of threads In response to determining that the initial number of threads is greater than or equal to a predetermined number of threads, determining a first number of executable tasks based on the initial number of threads, The capacity of the first target storage means is less than or equal to the capacity of the second target storage means, In response to determining that the initial number of threads is equal to a predetermined number of threads, determining a first number of tasks based on the amount of resources required for the processor to process the target data, Further including determining a second number of executable tasks based on the first number of tasks and the initial number of threads Data processing method.
9. Writing data to be processed, including the input data to be processed and the weight data to be processed, corresponding to the first number of executable tasks, into the first target storage means, Executing in parallel tasks including processing the input data to be processed using the weight data to be processed corresponding to the first number of executable tasks, and obtaining the output data corresponding to the first number of executable tasks, Further including writing the output data corresponding to the first number of executable tasks into the first target storage means The method according to claim 8.
10. Writing data to be processed, including the input data to be processed and the weight data to be processed, corresponding to the initial number of threads, into the first target storage means, Writing the data to be processed corresponding to the first number of tasks into the second target storage means, Executing in parallel tasks including processing the input data to be processed using the weight data to be processed corresponding to the second number of executable tasks, and obtaining the output data corresponding to the second number of executable tasks, Writing the output data corresponding to the initial number of threads into the first target storage means, Further including writing the output data corresponding to the first number of tasks into the second target storage means The method according to claim 8.
11. Further including determining a third number of executable tasks based on the amount of resources required for the processor to process the target data in response to determining that the data amount of the target data is greater than the capacity of the first target storage means The method according to claim 8.
12. Writing data to be processed, including the input data to be processed and the weight data to be processed, corresponding to the third number of executable tasks, into the second target storage means, Parallelly execute tasks including processing the input data to be processed using the weight data to be processed among the number of third executable tasks, and obtain the output data among the number of third executable tasks; Further include writing the output data among the number of third executable tasks to the second target storage means. The method according to claim 11.
13. In response to determining that the total data volume of the input data to be processed and the output data is larger than the capacity of the first target storage means, divide the input data to be processed into a plurality of sub-input data to be processed; Further include determining the number of fourth executable tasks based on the amount of resources required for the processor to process the sub-input data to be processed. The method according to claim 8.
14. Write the weight data to be processed and the sub-input data to be processed among the number of fourth executable tasks to the second target storage means; Parallelly execute tasks including processing the sub-input data to be processed using the weight data to be processed among the number of fourth executable tasks to obtain output sub-data among the number of fourth executable tasks; Write the output sub-data among the number of fourth executable tasks to the second target storage means; Further include stitching a plurality of the output sub-data into output data. The method according to claim 13.
15. Including the data processing device according to any one of claims 1 to 7 An electronic device.
16. Including at least one processor; And a memory communicatively connected to at least one processor, In the memory, instructions executable by the at least one processor are stored, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 8 to 14. An electronic device.
17. A non-transitory computer-readable storage medium storing computer instructions, The computer instructions cause a computer to execute the method according to any one of claims 8 to 14. A non-transitory computer-readable storage medium.
18. A computer program, When executed by a processor, realizes the method according to any one of claims 8 to 14. A computer program.
Citation Information
Patent Citations
Flow control method
JP2007122527A
Virtual computer management method, computer system, and resource management program
JP2011215812A
Application programming interface for data parallel computing in multiprocessor systems
JP2011523141A
Information processing device and information processing system
JP2014174946A
TASK PARALLEL PROCESSING METHOD, DEVICE, SYSTEM, STORAGE MEDIUM, AND COMPUTER DEVICE
JP2020522824A