Operation control method and device
By transferring the weighted data from the first memory to the second memory during the inference process of the artificial intelligence model, the problem of insufficient bandwidth of the multi-level memory is solved and the inference performance is improved.
Patent Information
- Application Number
- CN202510399278.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-24
AI Technical Summary
During the inference process, the artificial intelligence model has a small bandwidth between multi-level memories, which causes time-consuming data transmission and affects the inference performance.
By scheduling the second processing unit to read the target weight data from the first memory and store it in the second memory, it is ensured that the first processing unit can directly read the weight data from the second memory with a larger bandwidth when executing the operation sequence.
The impact of data bandwidth limitations between the first memory and the second memory on the inference performance of the artificial intelligence model is reduced, and data reading efficiency and inference performance are improved.
Smart Images

Figure CN120197707A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an operation control method and apparatus. Background Art
[0002] During the inference process of an artificial intelligence model, data reading between multiple levels of memories is usually involved.
[0003] However, the bandwidth between some of the multiple levels of memories is small, resulting in time-consuming data transmission, thus affecting the inference performance of the artificial intelligence model. Therefore, how to improve the inference performance of the artificial intelligence model is a technical problem that those skilled in the art need to solve. Summary of the Invention
[0004] On the one hand, this application provides an operation control method, including:
[0005] Determine a first processing unit for executing an operation sequence, where the operation sequence includes at least one operation required to process an inference task based on an artificial intelligence model;
[0006] In response to the existence of a target operation in the operation sequence that depends on weight data, before the first processing unit acquires the target weight data associated with the target operation, schedule a second processing unit to read the target weight data from a first memory and store the target weight data in a second memory. The first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory;
[0007] During the process of the first processing unit executing the operation sequence, based on the target operation, read the target weight data from the second memory.
[0008] In a possible implementation manner, the scheduling the second processing unit to read the target weight data from the first memory includes:
[0009] Determine a weight reading task associated with the target operation, and schedule the weight reading task to the second processing unit, where the weight reading task is used to indicate reading the target weight data from the first memory.
[0010] In still another possible implementation manner, it further includes:
[0011] Schedule each operation in the operation sequence to the first processing unit in sequence;
[0012] Or, schedule the operation sequence to the first processing unit;
[0013] Before the first processing unit obtains the target weight data associated with the target operation, scheduling the second processing unit to read the target weight data from the first memory includes:
[0014] Before scheduling the target operation to the first processing unit, scheduling the second processing unit to read the target weight data associated with the target operation from the first memory;
[0015] Or, before the first processing unit executes the target operation, scheduling the second processing unit to read the target weight data associated with the target operation from the first memory.
[0016] In another possible implementation, the scheduling the second processing unit to read the target weight data from the first memory includes:
[0017] Determine the operating status of each processing unit in the set of processing units;
[0018] Schedule the second processing unit in the set of processing units whose operating status meets the requirements to read the target weight data from the first memory.
[0019] In another possible implementation, the storing the target weight data in the second memory includes:
[0020] Store the target weight data in the target data storage area in the second memory, where the target data storage area is the storage area allocated for storing weight data for the target operation during the compilation of the operation code corresponding to the target operation;
[0021] The reading the target weight data from the second memory based on the target operation includes:
[0022] Read the target weight data from the target data storage area based on the address information indicated by the target attribute parameter in the target operation, where the target attribute parameter is the parameter generated for the target operation during the compilation of the operation code corresponding to the target operation to characterize the target data storage area.
[0023] In another possible implementation, the storing the target weight data in the target data storage area in the second memory includes:
[0024] In response to the maximum available storage capacity of the target data storage area being less than the data volume of the target weight data, divide the target weight data into at least two data blocks, and the size of each data block is less than or equal to the maximum available storage capacity;
[0025] Each time, store one of the at least two data blocks in the target data storage area;
[0026] Reading the target weight data from the target data storage area based on the address information indicated by the target attribute parameter in the target operation includes:
[0027] When the first processing unit executes the target operation, based on the address information indicated by the target attribute parameter in the target operation, each data block is sequentially read from the target data storage area;
[0028] The operation control method further includes:
[0029] After the first processing unit reads the data block from the target data storage area, each data block that has not been stored in the target data storage area among the at least two data blocks is sequentially stored in the target data storage area by the second processing unit.
[0030] In another possible implementation, determining the first processing unit for executing the operation sequence includes:
[0031] Determining at least two operation sequences required to be executed by the artificial intelligence model for processing the inference task, where the operation sequence includes at least one operation with a sequential order, and the operation types corresponding to the operations at the same order position in different operation sequences are the same, but the input data processed by different operation sequences is different;
[0032] For each operation sequence, determining the first processing unit for executing the operation sequence. In another possible implementation, the first processing units corresponding to different operation sequences are different;
[0033] The weight data required for the target operations that are in the same order position in different operation sequences and depend on the weight data is the same;
[0034] Responding to the existence of a target operation that depends on weight data in the operation sequence, before the first processing unit obtains the target weight data associated with the target operation, scheduling the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory includes:
[0035] Responding to the existence of a target operation that depends on weight data in the operation sequence, before each first processing unit obtains the target weight data associated with the target operation, scheduling the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory.
[0036] In another aspect, the present application further provides an operation control device, including:
[0037] A sequence allocation unit for determining a first processing unit for executing an operation sequence, the operation sequence including at least one operation required to process an inference task based on an artificial intelligence model;
[0038] A weight pre-reading unit for, in response to the existence of a target operation dependent on weight data in the operation sequence, scheduling a second processing unit to read the target weight data from a first memory and store the target weight data in a second memory before the first processing unit acquires the target weight data associated with the target operation, the first processing unit and the second processing unit being connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit being greater than the data bandwidth between the second memory and the first memory;
[0039] A weight reading unit for reading the target weight data from the second memory based on the target operation during the execution of the operation sequence by the first processing unit.
[0040] In another possible implementation, the weight pre-reading unit includes:
[0041] A task scheduling subunit for, in response to the existence of a target operation dependent on weight data in the operation sequence, determining a weight reading task associated with the target operation and scheduling the weight reading task to a second processing unit before the first processing unit acquires the target weight data associated with the target operation, the weight reading task being used to indicate reading the target weight data from the first memory. Description of the Drawings
[0042] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0043] Figure 1 A schematic flowchart of an operation control method provided by the present application;
[0044] Figure 2 An example diagram of a multi-level storage architecture provided by the present application;
[0045] Figure 3 Another schematic flowchart of an operation control method provided by the present application;
[0046] Figure 4 Another schematic flowchart of an operation control method provided by the present application;
[0047] Figure 5Another schematic flowchart of the operation control method provided by this application;
[0048] Figure 6 Another schematic flowchart of the operation control method provided by this application;
[0049] Figure 7 Another schematic flowchart of the operation control method provided by this application;
[0050] Figure 8 It shows a timing example diagram of the processing unit processing each operation in the operation sequence when the inference task processed by the artificial intelligence model corresponds to an operation sequence;
[0051] Figure 9 It shows an execution logic example diagram of two operation sequences split from the inference task of the artificial intelligence model;
[0052] Figure 10 It shows a timing example diagram of sequentially processing each operation in two operation sequences by a single processing unit when the inference task processed by the artificial intelligence model is split into two operation sequences;
[0053] Figure 11 It shows a timing example diagram of two processing units processing two operation sequences in parallel;
[0054] Figure 12 It shows a comparison example diagram of two implementation logics for executing an inference task based on an artificial intelligence model;
[0055] Figure 13 An implementation logic example diagram of the operation control method of this application;
[0056] Figure 14 It shows a timing example diagram of the scheduling weight reading task and two operation sequences in this application;
[0057] Figure 15 A schematic diagram of a composition structure of the operation control device provided by this application;
[0058] Figure 16 A schematic diagram of a composition architecture of the electronic device provided by this application. Detailed implementation manners
[0059] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application. The terms used in the embodiments of this application are only used to explain the specific embodiments of this application, and are not intended to limit this application. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0060] In the description, claims and the above-mentioned drawings of this application, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0061] To improve the inference performance of the inference task of the artificial intelligence model, this application provides an operation control method. The artificial intelligence model applicable to the operation control method of this application can be any artificial intelligence model with weight parameters. Among them, the weight parameters of the artificial intelligence model can be adjusted by training the artificial intelligence model. For example, the artificial intelligence model applicable to the solution of this application can be a traditional machine learning model, a deep learning model, or a currently commonly used large language model, etc., without specific limitations. The operation control method of the embodiments of this application will be introduced in detail below with reference to the drawings.
[0062] As Figure 1 , a flowchart of an operation control method provided by this application is shown. The method of this embodiment can be applied to an electronic device, which can be a terminal device such as a notebook computer, a tablet computer or a desktop computer, or can also be a server, or a device node in a system such as a cloud platform, etc., without limitation.
[0063] The method of this embodiment may include:
[0064] S101, determine a first processing unit for executing an operation sequence.
[0065] Among them, the operation sequence includes at least one operation required to process an inference task based on the artificial intelligence model. In this application, the inference task based on the artificial intelligence model can be image recognition, image generation or other image processing tasks; it can also be inference tasks such as question answering, text information summarization or data analysis report generation, etc., without specific limitations.
[0066] In this application, the operations required for an artificial intelligence model to process inference tasks may include: different types of computing operations. For example, the computing operations can be computing operations that do not require the aid of weight data, such as adding data, taking the square root of data, or calculating the square of data; they can also be computing operations that require computing and processing based on weight data. Among them, the computing operations that require weight data can be computing operations such as General Matrix Multiplication (GEMM) operations, convolution operations, or attention mechanism operations.
[0067] In this application, the processing unit can be a Neural Network Processing Unit (NPU), a Central Processing Unit (CPU), or other types of processors. The processing unit can also be a processor core. For example, the processing unit can be an NPU processor core or a CPU core, etc.
[0068] In a possible implementation, an electronic device may include a set of processing units and a scheduling unit. Among them, the set of processing units includes multiple processing units for performing inference tasks based on an artificial intelligence model, and the scheduling unit is connected to each processing unit in the set of processing units. The scheduling unit can schedule each operation in the operation sequence required for the artificial intelligence model to process the inference task to the corresponding processing unit, so as to execute the operations in the operation sequence through the processing unit. Correspondingly, the electronic device can determine a first processing unit for executing the operation sequence through the scheduling unit.
[0069] S102, in response to the existence of a target operation that depends on weight data in the operation sequence, before the first processing unit obtains the target weight data associated with the target operation, schedule a second processing unit to read the target weight data from the first memory and store the target weight data in the second memory.
[0070] For the sake of easy distinction, the operations in the operation sequence that require weight data are called target operations. For example, if the operations in the operation sequence include GEMM operations, then the GEMM operations belong to target operations. Of course, there may be one or more target operations that require weight data in the operation sequence, but the processing process for each target operation is the same.
[0071] Among them, the first processing unit and the second processing unit are connected to the second memory. The data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory. Since both the first processing unit and the second processing unit are connected to the second memory, the data bandwidth between the second memory and the first processing unit and the second processing unit is actually the data bandwidth between the second memory and any one of the first processing unit and the second processing unit. Usually, the data storage space of the first memory is greater than that of the second memory.
[0072] For example, the first memory can be a Double Data Rate (DDR) memory, also known as a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM). The second memory can be a secondary cache shared by each processing unit in the processing unit set.
[0073] Usually, the weight data required for the artificial intelligence model to process the inference task will be stored in the first memory first. Based on this, in order to perform the target operation based on the weight data, the first processing unit needs to first read the weight data from the first memory and store it in the second memory, and then read the corresponding weight data from the second memory. However, due to the relatively small data bandwidth between the first memory and the second memory, the data reading rate of reading the weight data from the first memory is relatively low, which will affect the inference performance of the artificial intelligence model in processing the inference task.
[0074] For the sake of easy understanding, taking the first memory as a DDR memory and the second memory as a secondary cache shared by each processing unit as an example, and combined with Figure 2 the multi-level storage architecture shown is described as follows:
[0075] In Figure 2 taking the processing unit as a processor core as an example, for example, Figure 2 in
[0076] taking three processor cores as an example, these three processor cores are processor core 0, processor core 1, and processor core 2 respectively.
[0076] As can be seen from Figure 2 : The storage architecture corresponding to multiple processor cores is a three-level cache architecture. Among them, each processor core has its own level-1 cache inside, and each processor core shares a secondary cache. The DDR memory can be regarded as the third-level cache. Multiple processor cores can be connected to the DDR memory through a bus, etc. Among them, the data bandwidth between each processor core (or the level-1 cache of the processor core) and the secondary cache is relatively large. For example, this data bandwidth can be 64 GB / s. InFigure 2 Taking the case where a secondary cache is shared by several processor cores as an example, the total data bandwidth between the secondary cache and three processor cores can reach . However, the data bandwidth between the secondary cache and the DDR memory is relatively small. For example, the data bandwidth can be 32GB / s. In Figure 2 the architecture shown, the processing data required by the processor cores needs to be first transferred from the DDR memory to the secondary cache, and then read from the secondary cache into the first-level cache of the processor cores so that the processor cores can process the corresponding data.
[0077] From Figure 2 it can be seen that if, when the processor core executes the target operation, the weight data is first read from the DDR memory and then cached in the secondary cache, it will inevitably affect the processing efficiency of the processor core for executing the target operation due to the relatively long data reading time, and further affect the performance of task inference based on the artificial intelligence model.
[0078] Based on this, in order to reduce the influence of the data bandwidth limitation between the first memory and the second memory on the inference performance of the artificial intelligence model inference task, in this application, before the first processing unit obtains the target weight data associated with the target operation, the second processing unit is scheduled to read the target weight data from the first storage unit and store it in the second memory, so that when the first processing unit needs to read the target weight data, it can directly read the target weight data from the second memory.
[0079] S103. During the process of the first processing unit executing the operation sequence, based on the target operation, read the target weight data from the second memory.
[0080] It can be understood that during the process of the first processing unit executing the operation sequence, if the operation to be executed by the first processing unit is the target operation, then the first processing unit reads the target weight data associated with the target operation to perform the calculation corresponding to the target operation based on the target weight data.
[0081] Since the target weight data has been stored in the second memory before the first processing unit needs to read the target weight data, therefore, when the first processing unit executes the target operation, it can directly read the target weight data from the second memory without accessing the first memory again, thus reducing the path length of reading the target weight data and improving the reading efficiency of the target weight data. Moreover, since the data bandwidth between the second memory and the first processing unit is relatively larger, the first processing unit can read the target weight data from the second memory more efficiently, which also improves the reading efficiency of the first processing unit for reading the target weight data.
[0082] It can be understood that after the first processing unit reads the target weight data, when the calculation types corresponding to the target operations are different, the specific calculation processes executed by the first processing unit based on the target operations will also be different. In this application, there is no limitation on the specific calculation process executed by the first processing unit after obtaining the target weight data.
[0083] As can be seen from the above, after determining the first processing unit for executing the operation sequence in this application, in response to the existence of a target operation dependent on weight data in the operation sequence, before the first processing unit obtains the target weight data associated with the target operation, the second processing unit will be scheduled to read the target weight data from the first memory and store it in the second memory. On this basis, when the first processing unit needs to read the target weight data based on the target operation, it can directly read the target weight data from the second memory. Since the data bandwidth between the second memory and the first processing unit is relatively large, the first processing unit can read the target weight data more efficiently, thereby improving the processing efficiency of the first processing unit for processing the target operation, and further improving the inference performance of the inference task based on the artificial intelligence model, and reducing the impact on the inference performance of the inference task of the artificial intelligence model due to the data bandwidth limitation between the first memory and the second memory.
[0084] In the above embodiments of this application, in order to reduce the impact on the inference task based on the artificial intelligence model, this application needs to reasonably select the second processing unit to schedule the second processing unit to read the target weight data. Based on this, this application can first determine the running states of the processing units in the processing unit set. Correspondingly, this application can schedule the second processing unit in the processing unit set whose running state meets the requirements to read the target weight data from the first memory and store the target weight data in the second memory.
[0085] Among them, the processing unit set includes multiple processing units for executing inference tasks based on the artificial intelligence model. For example, the processing unit set is a preset set of various processing units for running the artificial intelligence model. Another example is that the processing unit set can be each processor core on the processing chip for running the artificial intelligence model or processor cores of the same type. For example, the processing unit set can include multiple NPU cores.
[0086] Among them, the running state of the processing unit can represent the load condition of the processing unit. Based on this, it can be determined whether the processing unit is in an idle state, whether the processing unit is running operations related to the artificial intelligence model, and the number of operations to be executed by the processing unit through the running state of the processing unit.
[0087] In this application, the requirements for the operating state can be set according to actual needs. For example, the operating state meeting the requirements can be being in an idle state, with the load being less than a set threshold, or being able to be in an idle state before the first processing unit obtains the target weight data associated with the target operation, etc., without specific limitations.
[0088] It can be understood that in this application, after determining the first processing unit for executing the operation sequence, this application will also schedule the operations in the operation sequence to the first processing unit so that the first processing unit can execute each operation in the operation sequence. Among them, there can be various possibilities for the scheduling method of scheduling the operations in the operation sequence to the first processing unit. Correspondingly, when the scheduling methods for scheduling the operations to the first processing unit are different, the specific implementation of scheduling the second processing unit to read the target weight data will also be different. The following will illustrate with several possible cases.
[0089] In a possible case, this application can schedule each operation in the operation sequence to the first processing unit in sequence. In this case, after the first processing unit finishes executing one operation in the operation sequence, it can schedule the next operation to be scheduled in the operation sequence to the first processing unit. Correspondingly, before scheduling the target operation that depends on the weight data to the first processing unit, the second processing unit can be scheduled to read the target weight data associated with the target operation from the first memory. Based on this, before the first processing unit executes the target operation, the target weight data can be read from the first memory by the second processing unit and stored in the second memory. Then, when the first processing unit executes the target operation, the corresponding target weight data can be read from the second memory.
[0090] In another possible case, this application can also schedule the operation sequence to the first processing unit. In this case, the first processing unit can execute each operation in the operation sequence in sequence according to the allocated operation sequence. Correspondingly, before the first processing unit executes the target operation that depends on the weight data, the second processing unit can be scheduled to read the target weight data from the first memory and store it in the second memory, so that the first processing unit can read the target weight data from the second memory when executing the target operation.
[0091] Of course, the above two possible cases can also be combined with each other. For example, in the case of scheduling an operation sequence to the first processing unit each time, if the operation sequence has not been scheduled to the first processing unit, naturally the target operation has not been scheduled to the first processing unit either. Therefore, before scheduling the target operation to the first processing unit, the second processing unit can also be scheduled to read the target weight data from the first memory and store the target weight data in the second memory.
[0092] It can be understood that there may be one or more target operations in the operation sequence that rely on weight data. In order to obtain the target weight data required for the target operation more accurately and efficiently, the present application may first determine a weight reading task associated with the target operation. On this basis, the weight reading task may be scheduled to a second processing unit, and the weight reading task is used to indicate reading the target weight data from a first memory. Correspondingly, when the second processing unit executes the weight reading task, the target weight data will be read from the first memory and stored in the second memory.
[0093] Among them, the weight reading task may be a pre-defined data reading operation.
[0094] For example, a data reading operation that is developed in advance for the target operation and is solely used to read the target weight data.
[0095] Another example is that when deploying an artificial intelligence model, it is necessary to compile the model data file corresponding to the artificial intelligence model. Through encoding, the model data file of the artificial intelligence model can be compiled into binary code executable on electronic devices such as computers, and thus the operation sequence required for the artificial intelligence model to perform inference tasks is obtained. Based on this, the compilation process is the process of compiling and generating the operation codes of each operation in the operation sequence. On this basis, in the process of compiling the operation code of the target operation, the present application can isolate the operation part in the target operation that is used to execute reading the target weight data from the first memory to form a new data reading operation, and this data reading operation is the weight reading task used to implement reading the target weight data from the first memory to the second memory.
[0096] Of course, the compilation of the model data file only needs to be executed once when deploying the artificial intelligence model on the electronic device. After compiling the model data file into binary code, the operation sequence required for the artificial intelligence model to perform inference tasks can be obtained. Therefore, there is no need to perform the compilation operation again every time the artificial intelligence model is called. In the present application, the specific compilation process is not limited.
[0097] It can be understood that in the case of scheduling the weight reading task to the second processing unit, the determination of the second processing unit and the specific timing of scheduling the weight reading task to the second processing unit and so on can refer to the relevant introduction in the previous embodiments. For the sake of easy understanding, the following takes an implementation manner as an example for illustration. For example Figure 3 , shows another implementation flow diagram of the operation control method provided by the present application. The method of this embodiment may include:
[0098] S301, determine a first processing unit for executing the operation sequence.
[0099] Among them, the operation sequence includes at least one operation required to process the inference task based on the artificial intelligence model.
[0100] S302, schedule the operation sequence to the first processing unit so that the first processing unit sequentially executes each operation in the operation sequence.
[0101] S303, in response to the existence of a target operation that depends on weight data in the operation sequence, before the first processing unit executes the target operation, determine the weight reading task associated with the target operation, and schedule the weight reading task to a second processing unit in the processing unit set whose running state meets the requirements.
[0102] Among them, the second processing unit whose running state meets the requirements can refer to the relevant introduction above. For example, the weight reading task can be scheduled to the second processing unit in the idle state so that the second processing unit can execute the weight reading task in a timely manner.
[0103] It can be understood that if there are multiple target operations in the operation sequence, the weight reading tasks associated with different target operations are different. And the second processing units corresponding to different weight reading tasks can be the same or different. For example, specifically, the second processing units corresponding to different weight reading tasks can be determined by combining the running states of each processing unit in the processing unit set.
[0104] It should be noted that for the sake of understanding, in this embodiment, the entire operation sequence is scheduled to the first processing unit. Therefore, before the first processing unit executes the target operation, the weight reading task associated with the target operation is determined. However, it can be understood that if the scheduling of the operations in the operation sequence to the first processing unit is replaced by other implementation methods, it is also applicable to this embodiment as long as the weight reading task is scheduled to the second processing unit before the first processing unit needs to obtain the target weight data associated with the target operation, and the details will not be elaborated here.
[0105] S304, the second processing unit reads the target weight data associated with the target operation from the first memory based on the weight reading task, and stores the target weight data in the second memory.
[0106] Among them, the first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory.
[0107] In this application, the weight reading task is used to indicate reading the target weight data associated with the target operation from the first memory. On this basis, the second processing unit reads the target weight data from the first memory based on the weight reading task and stores it in the second memory.
[0108]
[0108] In S305, during the process of the first processing unit executing the operation sequence, if the currently to-be-executed operation is the target operation, the first processing unit reads the target weight data from the second memory based on the target operation.
[0109] For the specific implementation of this step S305, reference can be made to the relevant introduction in the previous embodiments, which will not be elaborated here.
[0110] It can be understood that in addition to storing the weight data associated with the target operation, the second memory also stores input data required for other operations, etc. For example, assuming that currently image generation or text character recognition in an image, etc., are to be performed based on an artificial intelligence model, then during the process of the processing unit performing an inference task based on the artificial intelligence model, the second storage unit also stores the image data of the image to be processed, and this image data is the input data that needs to be input into the artificial intelligence model for processing. Based on this, in order to enable the first processing unit to accurately read the target weight data associated with the target operation from the second memory, in this application, during the process of compiling the operation code corresponding to the target operation, the target data storage area allocated for storing the weight data for this target operation can be determined from the second memory.
[0111] Correspondingly, during the process of compiling the operation code corresponding to the target operation, a target attribute parameter for the target operation can also be generated to characterize this target data storage area. The target attribute parameter can indicate the address information of this target data storage area.
[0112] For example, the initial value of the target attribute parameter in the target operation is empty. If the target data storage area is allocated for the target operation during the process of compiling the operation code corresponding to the target operation, then the value of the target attribute parameter in this target operation can be set to the address information corresponding to this target data storage area.
[0113] Based on the above, the second processing unit can be scheduled to store the target weight data associated with the target operation into this target data storage area in the second memory. Correspondingly, when the first processing unit executes the target operation, it can read the target weight data from this target data storage area based on the address information indicated by the target attribute parameter in the target operation.
[0114] It can be understood that if it is not necessary to schedule the second processing unit to read the target weight data from the first memory before the first processing unit needs to obtain the target weight data, then in the compilation phase, naturally, no target data storage area will be allocated for the target operation, and naturally, no target attribute parameter will be generated for the target operation. Based on this, if the target attribute parameter in the target operation is empty, then the first processing unit can read the target weight data from the first memory based on the target operation.
[0115] For ease of understanding, the following uses the example of scheduling a weight reading task to a second processing unit to implement scheduling the second processing unit to read target weight data from a first memory for illustration. As Figure 4 , another schematic flowchart of the operation control method provided by the present application is shown. This embodiment may include:
[0116] S401, determine a first processing unit for executing an operation sequence.
[0117] Among them, the operation sequence includes: at least one operation required to process an inference task based on an artificial intelligence model.
[0118] S402, schedule the operation sequence to the first processing unit so that the first processing unit sequentially executes each operation in the operation sequence.
[0119] S403, in response to a target operation in the operation sequence that depends on weight data, before the first processing unit executes the target operation, determine a weight reading task associated with the target operation, and schedule the weight reading task to a second processing unit in the processing unit set whose running state meets the requirements.
[0120] It can be understood that before the first processing unit executes the target operation can be any moment before the first processing unit executes the target operation. For example, it can be when the operation sequence is scheduled to the first processing unit; or when the operation currently executed by the first processing unit is the previous operation of the target operation; or when the operation currently to be executed by the first processing unit is the target operation. Of course, there can be other possible situations, which will not be elaborated here.
[0121] Among them, the weight reading task can be a pre-generated data reading operation for reading target weight data associated with the target operation, as described above in detail, and will not be elaborated here.
[0122] Among them, the weight reading task is used to indicate reading target weight data associated with the target operation from the first memory.
[0123] In a possible implementation manner, the address attribute parameter of the weight reading task indicates the target storage address of the target weight data in the first memory and the address information corresponding to the target data storage area in the second memory for storing the target weight data.
[0124] For example, during the process of compiling the operation code corresponding to the target operation, the target storage address of the target weight data required by the target operation in the first memory can be determined. Moreover, during the process of compiling the operation code corresponding to the target operation, the target data storage area allocated for the target operation to store weight data can also be determined from the second memory. On this basis, during the compilation phase of the weight reading task, the target storage address of the target weight data in the first memory and the address information corresponding to the target data storage area in the second memory can be added to the address attribute parameter of the weight reading task.
[0125] For example, if the weight reading task is a data reading operation developed in advance for the target operation, then during the process of compiling the operation code of the weight reading task associated with the target operation, two parameter values can be added to the address attribute parameter of the weight reading task. These two parameter values are respectively the target storage address of the target weight data in the first memory and the address information corresponding to the target data storage area in the second memory. Another example is that if the weight reading task is separated during the process of compiling the operation code of the target operation, then the target storage address of the target weight data in the first memory and the address information corresponding to the target data storage area can be added to the address attribute parameter of the weight reading task.
[0126] It should be noted that this embodiment is described by taking the scheduling of the operation sequence to the first processing unit as an example. However, the case of sequentially scheduling each operation in the operation sequence to the first processing unit is also applicable to this embodiment. For example, in the case of sequentially scheduling each operation in the operation sequence to the first processing unit, step S403 can be replaced with: in response to the existence of a target operation that depends on weight data in the operation sequence, before scheduling the target operation to the first processing unit, determine the weight reading task associated with the target operation, and schedule the weight reading task to a second processing unit in the set of processing units whose operating status meets the requirements.
[0127] S404, the second processing unit reads the target weight data associated with the target operation from the first memory based on the weight reading task, and stores the target weight data in the target data storage area in the second memory.
[0128] Wherein, the target data storage area is the storage area allocated for the target operation to store weight data during the process of compiling the operation code corresponding to the target operation.
[0129] In a possible implementation, if the address attribute parameter of the weight reading task indicates the target storage address of the target weight data in the first memory and the address information corresponding to the target data storage area in the second memory for storing the target weight data, then the second processing unit may read the target weight data from the first memory based on the target storage address indicated by the address attribute parameter in the weight reading task, and store the target weight data in the target data storage area in the second memory based on the address information corresponding to the target data storage area indicated by the address attribute parameter.
[0130] S405. During the process of the first processing unit executing the operation sequence, the first processing unit reads the target weight data from the target data storage area of the second memory based on the address information indicated by the target attribute parameter in the target operation.
[0131] Wherein, the target attribute parameter is a parameter generated for the target operation during the compilation of the operation code corresponding to the target operation and used to characterize the target data storage area.
[0132] It can be understood that in any embodiment where a target data storage area is allocated for the target operation during the compilation phase, the storage space of the target data storage area allocated for the target operation in the second memory is also relatively limited. Therefore, it is very likely that the target data storage area is not sufficient to accommodate the target weight data corresponding to the target operation. In the case where the target data storage area is not sufficient to accommodate the target weight data, in order to ensure that the first processing unit can still directly read the target weight data from the target data storage area of the second memory, the present application can also block the target weight data and store each data block of the target weight data in the target data storage area in batches.
[0133] The following is combined with Figure 5 for illustration. As Figure 5 shows another flowchart of the operation control method provided by the present application. The method of this embodiment may include:
[0134] S501. Determine the first processing unit for executing the operation sequence.
[0135] Wherein, the operation sequence includes at least one operation required to process the inference task based on the artificial intelligence model.
[0136] S502. In response to the existence of a target operation that depends on weight data in the operation sequence, before the first processing unit obtains the target weight data associated with the target operation, schedule the second processing unit to read the target weight data from the first memory.
[0137] For example, schedule the weight reading task to the second processing unit, so that the second processing unit can read the target weight data from the first memory and perform subsequent related operations.
[0138] Among them, for the specific implementation of scheduling the second processing unit to read the target weight data, reference can be made to the relevant introduction in any of the previous embodiments, which will not be elaborated here.
[0139] S503. In response to the maximum available storage capacity of the target data storage area in the second memory being less than the data volume of the target weight data, divide the target weight data into at least two data blocks, and the second processing unit stores one of the at least two data blocks into the target data storage area each time.
[0140] Among them, the first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory.
[0141] The target data storage area is a storage area allocated for storing weight data for the target operation during the compilation of the operation code corresponding to the target operation.
[0142] Among them, the maximum available storage capacity of the target data storage area is the data volume that the target data storage area can store at most.
[0143] In this application, dividing the target weight data into at least two data blocks can be executed by the scheduling unit. Of course, it can also be executed by the second processing unit, and there is no specific limitation.
[0144] In this embodiment, the size of each data block divided from the target weight data is less than or equal to the maximum available storage capacity. It can be understood that the second processing unit can only read one data block divided from the target weight data from the first memory and store it into the target data storage area of the second memory each time, and repeat this operation continuously until there is no free storage space in the target data storage area.
[0145] S504. When the first processing unit executes the target operation in the operation sequence, based on the address information indicated by the target attribute parameter in the target operation, sequentially read each data block of the target weight data from the target data storage area.
[0146] S505. After the first processing unit reads the data block from the target data storage area, the second processing unit sequentially stores each data block that has not been stored in the target data storage area among the at least two data blocks into the target data storage area.
[0147] It can be understood that since the maximum available storage capacity of the target data storage area is less than the data volume of the target weight data, in step S503, only some of the data blocks divided from the target weight data can be stored in the target data storage area, and there are still data blocks in at least two data blocks swapped out from the target weight data that have not been stored in the target data storage area. After the first processing unit reads a data block from the target data storage area, free storage space appears in the target data storage area again. Therefore, the second processing unit can continue to read each data block that has not been stored in the target data storage area from the first memory and store the read data blocks in the target data storage area.
[0148] It can be understood that every time the first processing unit reads a data block from the target data storage area, the second processing unit can store a data block of the target weight data into the target data storage area until there are no unread data blocks in the target weight data stored in the first memory.
[0149] It can be understood that the operations required to process the inference task based on the artificial intelligence model involve the computational processing of the input data required for the operation. On this basis, if the first processing unit reads the input data corresponding to each operation from the first memory, due to the small data bandwidth between the first memory and the second memory, the first processing unit cannot efficiently obtain the input data corresponding to the operation, thus affecting the performance of processing the inference task based on the artificial intelligence model.
[0150] Based on this, in order to enable the first processing unit to process each operation in the operation sequence more efficiently, in this application, every time the first processing unit finishes the computational processing corresponding to an operation, it will store the output data generated by executing the operation in the second memory. It can be understood that for the first operation and the second operation adjacent to each other in the operation sequence, the output data generated by the first operation is actually the input data required for the second operation to be processed. On this basis, since the output data of the first operation is cached in the second memory, when the first processing unit executes the second operation, it can directly read the input data required for the second operation to be processed from the second memory, thus eliminating the need to frequently access the first memory.
[0151] It can be seen that in the present application, when the operation to be executed by the first processing unit is the first operation in the operation sequence (i.e., the first operation), the first processing unit can read the input data required for the first operation from the first memory. If the operation to be executed by the first processing unit is not the first operation (i.e., does not belong to the first operation in the operation sequence), the first processing unit can directly read the input data to be processed required for the non-first operation from the second memory, thereby reducing the number of times the first processing unit needs to access the first memory, naturally improving the data reading efficiency of the first processing unit, and then improving the efficiency of the first processing unit in processing the operation sequence, and thus improving the efficiency of processing the inference task based on the artificial intelligence model.
[0152] It can be understood that since the first processing unit is not directly connected to the first memory, the first processing unit actually reads the input data required for the operation from the first memory, caches the input data required for the operation in the second memory, and then reads the input data required for the operation from the second memory.
[0153] Furthermore, considering that usually the storage space of the second memory is smaller than that of the first memory, in order to reduce the situation where the output data generated by the first processing unit during the execution of the operation cannot be cached in the second memory due to insufficient storage space of the second memory, in the present application, the initial operation sequence involved in processing the inference task by the artificial intelligence model can be further split into at least two operation sequences in advance, and the data to be processed required for the initial operation sequence is split into at least two input data required for the at least two operation sequences, and the input data required for different operation sequences is different.
[0154] Correspondingly, the number of operations included in different operation sequences is the same, and the operation types corresponding to the operations in the same order position in different operation sequences are also the same, so as to divide the data to be processed required for the initial operation sequence for processing by the at least two operation sequences.
[0155] It can be understood that during the process of executing the inference task based on the artificial intelligence model, it is necessary to sequentially perform calculation processing on each data in the data set to be processed by the artificial intelligence model through various operations in the artificial intelligence model. Since the operation types included in each operation sequence are the same and all include various operations required for the artificial intelligence model to execute the inference task, when the data set required to be processed by the artificial intelligence model is divided into at least two portions of input data and these two portions of input data are processed in parallel through at least two operation sequences, it can still be ensured that each portion of input data can pass through various operations in the artificial intelligence model for calculation processing. In this way, the result after the output data obtained by processing each portion of input data is fused is the same as the result obtained by only processing the data set through one initial operation sequence.
[0156] For example:
[0157] Suppose the data set to be processed by the artificial intelligence model is the feature matrix of an image. For ease of description, only one operation that requires weight data in the artificial intelligence model, namely the target operation, is taken as an example.
[0158] In the solution of this application, the feature matrix of the image can be divided into multiple sub-matrices, and these multiple sub-matrices can be combined to obtain the feature matrix of the image. Since the target operation is included in each of the multiple operation sequences corresponding to the artificial intelligence model, when the sub-matrices are respectively assigned to different operation sequences for processing, each sub-matrix will perform an operation with the weight matrix (i.e., weight data) in the target operation of the corresponding operation sequence. As a result, each sub-matrix in the feature matrix of the image will perform a calculation with the weight matrix associated with the target operation. In this way, the output result obtained after multiplying each sub-matrix by the weight matrix associated with the target operation and then fusing is actually the output result obtained by directly multiplying the feature matrix of the image by the weight matrix associated with the target operation.
[0159] Based on this, at least two operation sequences required for the artificial intelligence model to process the inference task can be determined in this application. Correspondingly, for each operation sequence, a first processing unit for executing the operation sequence can be determined. As introduced above, the operation types corresponding to the operations at the same sequential position in different operation sequences are the same, but the input data processed by different operation sequences is different.
[0160] For example, suppose the initial operation sequence required for the artificial intelligence model to process the inference task is divided into two operation sequences. Then the input data to be processed by the initial operation sequence can be split into two parts, namely the first input data and the second input data. Suppose the artificial intelligence model needs to perform data addition calculation, data multiplication calculation, and GEMM calculation on the data in the data set in sequence when processing the inference task. Then it is determined that the artificial intelligence model needs to execute operation sequence 1 and operation sequence 2 when processing the task inference. Among them, operation sequence 1 includes in sequence: data addition calculation, data multiplication calculation, and GEMM calculation, and operation sequence 2 also includes data addition calculation, data multiplication calculation, and GEMM calculation in sequence. It's just that based on operation sequence 1, the data addition calculation, data multiplication calculation, and GEMM calculation need to be performed on the first input data in sequence, while based on operation sequence 2, the data addition calculation, data multiplication calculation, and GEMM calculation need to be performed on the second input data in sequence.
[0161] It can be understood that since the operation types corresponding to the operations at the same sequential position in different operation sequences are the same, if the operation at a certain sequential position in a certain operation sequence is a target operation that needs to depend on weight data, then the operations at the corresponding sequential positions in other operation sequences are also the same type of target operation that needs to depend on weight data.
[0162] Specifically, since different operation sequences only have different input data being processed, but the specific calculation operations executed sequentially by each operation in the operation sequence can be the same. Based on this, all the weight data required by the target operations that are at the same sequential position and need to depend on weight data in different operation sequences are the same. For this situation, the following will be described in conjunction with Figure 6 for illustration.
[0163] As Figure 6 , it shows another schematic flowchart of the operation control method provided by the present application. The method of this embodiment may include:
[0164] S601, determine at least two operation sequences required to be executed by the artificial intelligence model for processing the inference task.
[0165] Among them, the operation sequence includes at least one operation with a sequential order. The operation types corresponding to the operations at the same sequential position in different operation sequences are the same, but the input data processed by different operation sequences are different.
[0166] In the present application, there is no limitation on the specific implementation process of splitting the initial operation sequence required to be executed by the artificial intelligence processing model into at least two operation sequences.
[0167] S602, for each operation sequence, determine the first processing unit for executing the operation sequence.
[0168] In the present application, the first processing units corresponding to at least two operation sequences may be the same, completely different, or partially the same, and there is no specific limitation.
[0169] In a possible implementation manner, in order to make full use of the computing resources of the processing units in the electronic device and improve the performance of the artificial intelligence model in processing the inference task, at least two processing units may be used to execute the at least two operation sequences in parallel, thereby improving the performance of the artificial intelligence model in processing the inference task. Based on this, the first processing units corresponding to different operation sequences are different.
[0170] S603, in response to the existence of a target operation that depends on weight data in the operation sequence, before each first processing unit obtains the target weight data associated with the target operation, schedule the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory.
[0171] Among them, the first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory.
[0172] In this application, the weight data required for the target operations that are in the same sequential position and require weight data in different operation sequences is the same. Based on this, for the target operations in the same sequential position in multiple operation sequences, only the second processing unit needs to read the target weight data corresponding to the target operation from the first memory, rather than reading the target weight data from the first memory multiple times.
[0173] It can be understood that before adopting the solution of this application, since each operation sequence includes the same type of target operation, each first processing unit needs to read the target weight data of the target operation from the first memory every time it executes the target operation in an operation sequence. Therefore, the more operation sequences there are, the more times the target weight data needs to be repeatedly read from the first memory. And the more times the target weight data is repeatedly read from the first memory, the more the read bandwidth of the first memory is occupied, which is less conducive to efficient data reading, and thus affects the execution of the operation sequence, and naturally affects the efficiency of the artificial intelligence model in executing the inference task.
[0174] By adopting the solution of this application, the inference task executed by the artificial intelligence model is split into multiple operation sequences, and multiple first processing units can be used to execute multiple operation sequences in parallel to improve the inference efficiency of the inference task. Moreover, before each first processing unit executes the target operation in the operation sequence, the second processing unit is used to read the target weight data associated with the target operation into the second memory in advance, so that each first processing unit can directly read the target weight data from the second memory with a very large data bandwidth when executing the target operation, which can further reduce the impact on the inference efficiency of the inference task due to repeatedly reading the target weight data from the first memory multiple times, and naturally can further improve the inference efficiency of the inference task.
[0175] In this embodiment, the specific implementation of scheduling the second processing unit to read the target weight data from the first memory and store it in the second memory can be the relevant introduction in any previous embodiment, and will not be elaborated here.
[0176] S604. For any operation sequence, during the process of the first processing unit executing the operation sequence, based on the target operation, read the target weight data from the second memory.
[0177] Among them, for the specific implementation of each first processing unit to execute the corresponding operation sequence, reference can also be made to the relevant introduction in the previous embodiments, which will not be elaborated here.
[0178] To facilitate understanding the benefits of the solution of this application, the following takes an example where an artificial intelligence model needs to execute two operation sequences for processing an inference task, and is described in conjunction with Figure 7 a possible implementation manner. As shown in FIG. 7, another schematic flowchart of the operation control method provided by this application is shown. The method of this embodiment may include:
[0179] S701, determine two operation sequences required for the artificial intelligence model to process the inference task.
[0180] Among them, the operation sequence includes at least one operation with a sequential order. The operation types corresponding to the operations at the same order position in different operation sequences are the same, but the input data processed by different operation sequences is different.
[0181] S702, determine two first processing units for executing these two operation sequences, and schedule the two operation sequences to these two first processing units respectively, so that each first processing unit executes each operation in the corresponding operation sequence.
[0182] Among them, different operation sequences are scheduled to different first processing units.
[0183] In this embodiment, scheduling the two operation sequences to two different first processing units can enable the two first processing units to process the operations in these two operation sequences in parallel, thereby improving the processing efficiency of the artificial intelligence model for processing the inference task, and naturally improving the inference performance.
[0184] S703, in response to the existence of a target operation that depends on weight data in the operation sequence, before the two first processing units corresponding to the two operation sequences execute the target operation, determine a weight reading task associated with the target operation, and schedule the weight reading task to a second processing unit in the processing unit set whose running state meets the requirements.
[0185] In this application, the weight data required for the target operations that are in the same order position in the two operation sequences and need to depend on weight data is the same. Therefore, for the same type of target operation in the two operation sequences, therefore, the target operations in these two operation sequences correspond to the same weight reading task, and only need to schedule this weight reading task once.
[0186] For example, assume that the operations in the third position of two operation sequences are both the same type of GEMM operation. Since these two GEMM operations both belong to the computational operations that need to be executed thirdly based on the artificial intelligence model, the weight data on which the GEMM operations in these two sequences depend is the same, and only one weight scheduling task for obtaining the weight data corresponding to the GEMM operation needs to be scheduled.
[0187] Among them, the weight reading task is used to indicate reading the target weight data associated with the target operation from the first memory.
[0188] S704. The second processing unit reads the target weight data associated with the target operation from the first memory based on the weight reading task, and stores the target weight data in the target data storage area in the second memory.
[0189] Among them, the target data storage area is the storage area allocated for storing weight data for the target operation during the process of compiling the operation code corresponding to the target operation.
[0190] S705. For each operation sequence, during the process of the first processing unit executing the operation sequence, if the currently executed operation is a target operation that needs to depend on weight data, the first processing unit reads the target weight data from the target data storage area in the second memory based on the address information indicated by the target attribute parameter in the target operation.
[0191] It can be understood that in this embodiment, for any operation in the operation sequence, when performing the computational processing corresponding to the operation, the input data corresponding to the operation also needs to be obtained. As described above, during the process of the first processing unit executing the operation sequence, if the currently executed operation is the first operation in the operation sequence, the first processing unit can read the input data corresponding to the first operation from the first memory, perform the computational processing corresponding to the first operation based on the input data corresponding to the first operation, and store the output data generated by the first operation in the second memory.
[0192] If the currently executed operation is a non-first operation in the operation sequence, since the input data of the non-first operation is the output data generated by the previous operation of the non-first operation, the input data of the non-first operation can be directly read from the second memory, thereby reducing the number of accesses to the first memory during the process of the first processing unit executing the operation sequence and improving the data reading efficiency. Correspondingly, after performing the computational processing corresponding to the second operation based on the input data corresponding to the non-first operation, the output data generated by the non-first operation can be stored in the second memory.
[0193] Among them, the target operation can be the first operation in the operation sequence or a non-first operation. Among them, for the target operation, the first processing unit can perform the calculation processing corresponding to the target operation based on the input data and target weight data corresponding to the target operation, and store the output result generated by the target operation in the second memory.
[0194] Specifically, if the operation currently processed by the first processing unit is the last operation in the operation sequence, the first processing unit can also transfer the output result generated by the last operation from the second memory to the first memory.
[0195] To facilitate understanding, the benefits of splitting the inference task of the artificial intelligence model into two operation sequences, and the benefits of the first processing unit reading the input data required for the operation from the second memory, are described below in conjunction with Figure 8 , Figure 9 , Figure 10 and Figure 11 are illustrated. For the sake of easy understanding, the storage architecture corresponding to the processing unit is used as an example of the storage architecture shown in Figure 2 for illustration.
[0196] Among them, Figure 8 shows an example diagram of the number of clocks required for the processing unit to process each operation in the operation sequence in the case where the inference task processed by the artificial intelligence model corresponds to one operation sequence.
[0197] For the sake of easy understanding and description, in Figure 8 , taking operation 1 and operation 2 involved in the inference task of the artificial intelligence model as examples, that is, the operation sequence successively includes operation 0 and operation 1. For the sake of easy understanding of the subsequent situation where the inference task is split into two operation sequences, in Figure 8 , it is assumed that the input data required for operation 0 includes data 00 and data 01. Considering the characteristics of data processing by the artificial intelligence model, in Figure 8 and the following examples, the data is usually a data matrix composed of multiple data. Here, data 00 and data 01 are two different data matrices, the input data required for operation 1 is data 10 and data 11, data 10 is actually the output data obtained by the processing unit through calculation processing of data 00 based on operation 0, and data 11 is the output data obtained by the processing unit through calculation processing of data 01 based on operation 0.
[0198] From Figure 8It can be seen that when the processing unit executes Operation 0, it needs to sequentially read Data 00 and Data 01 from the DDR memory, and cache Data 00 and Data 01 into the secondary cache L2 respectively. Then, Data 00 and Data 01 are sequentially read from the secondary cache L2 into the primary cache L1 of the processing unit. After the processing unit reads Data 00 into the primary cache L1, the processing unit performs the calculation processing corresponding to Operation 0 on Data 00, and stores the calculated output data into the DDR memory via the primary cache L1 and the secondary cache L2. Similarly, after reading Data 01 into the primary cache L1, the processing unit performs calculation processing on Data 01 based on Operation 0, and stores the calculated output data into the DDR memory via the primary cache L1 and the secondary cache L2.
[0199] After the processing unit completes the processing of Operation 0, the processing unit will continue to execute Operation 1. During the execution of Operation 1 by the processing unit, the reading and processing of Data 10 and Data 11 are similar to the process of the processing unit executing Operation 0, and will not be elaborated here.
[0200] It can be seen from Figure 8 that when the processing unit executes Operation 0 and Operation 1, it needs to first read the input data required for the operation from the DDR memory and store the calculated output data into the DDR memory, resulting in a relatively large number of accesses by the processing unit to the DDR memory. Due to the bandwidth limitation between the DDR memory and the secondary cache L2, the data access and storage speed is relatively slow, which will affect the efficiency of the processing unit in processing Operation 0 and Operation 1. It can be seen from Figure 8 that it takes 29 clocks for the processing unit to complete the processing of Operation 0 and Operation 1 (i.e., complete the inference task of the artificial intelligence model).
[0201] Among them, Figure 9 shows an example diagram of the execution logic of two operation sequences split from the inference task of the artificial intelligence model.
[0202] At Figure 9 the left side of the arrow in is the initial operation sequence corresponding to the inference task of the artificial intelligence model. The initial operation sequence includes Operation 0 and Operation 1, and the input data required for Operation 0 and Operation 1 need to be read from the DDR memory. The processing process corresponding to the execution of the initial operation sequence can be referred to the relevant introduction above Figure 8 .
[0203] Figure 9On the right side of the arrow in the figure is an example of the implementation logic after splitting the inference task of the artificial intelligence model into two operation sequences. For ease of distinction, these two operation sequences are respectively referred to as operation sequence 0 and operation sequence 1. Each of the split operation sequences still includes two operations. For ease of distinction, the two operations of operation sequence 0 are respectively referred to as operation 00 and operation 10, and the two operations in operation sequence 1 are respectively referred to as operation 01 and operation 11.
[0204] Among them, operation 00 in operation sequence 0 and operation 01 in operation sequence 1 belong to the same type of operation as operation 0 in the initial operation sequence. Therefore, the calculation and processing of the input data by operation 00 and operation 01 are the same as those performed by operation 0 on the input data. For example, if operation 0 is to perform an addition operation on the input data, then operation 00 and operation 01 also perform an addition operation on the input data. Correspondingly, operation 10 in operation sequence 0 and operation 11 in operation sequence 1 are also the same type of operation as operation 1 in the initial operation sequence. However, the input data corresponding to operation sequence 0 is data 00. Therefore, the input data of operation 00 in operation sequence 0 is data 00, and the input data 10 of operation 10 is the output data calculated by operation 00. The input data corresponding to operation sequence 1 is data 01. Therefore, the input data of operation 01 is data 01, and the input data of operation 11 is the data 11 output by operation 01.
[0205] Moreover, from Figure 9 it can be seen that when data 00 is cached in the secondary cache L2 based on operation 00 in operation sequence 0, operation 11 can be executed based on the data 00 cached in the secondary cache L2, and the data 10 output by executing operation 00 is also stored in the secondary cache L2. Based on this, when executing operation 10, data 10 can be directly read from the secondary cache L2. Similarly, after data 01 is cached in the secondary cache L2 based on operation 01 in operation sequence 1, the calculation processing corresponding to operation 01 can be executed based on the data 01 in the secondary cache, and the calculated data 11 is cached in the secondary cache L2. In this way, when executing operation 11, data 11 can be directly read from the secondary cache L2.
[0206] To facilitate understanding of the efficiency of executing the inference task of the artificial intelligence model based on operation sequence 0 and operation sequence 1, it is described in combination with Figure 10 as follows. Figure 10 Fig. shows a timing example diagram in which, when the inference task processed by the artificial intelligence model is split into two operation sequences, each operation in the two operation sequences is sequentially processed by a single processing unit.
[0207] In Figure 10 it, operation sequence 0 and operation sequence 1 can be sequentially scheduled to the same processing unit.
[0208] As shown Figure 10 in the figure, first, the processing unit first reads data 00 from the DDR memory based on operation 00 in operation sequence 0 and caches it in the secondary cache L2, and then reads the data 00 in the secondary cache L2 into the primary cache L1. Then, the processing unit performs computational processing on the input data 00 in the primary cache L1 based on operation 00, and caches the calculated output data (i.e., data 10) from the primary cache L1 into the secondary cache L2.
[0209] When the processing unit executes operation 10 in operation sequence 0, it can directly read the data 10 calculated by operation 00 from the secondary cache L2 and cache the data 10 into the primary cache L1. Then, after the processing unit performs computational processing on the data 10 based on operation 10, the data obtained from the computational processing is stored in the DDR memory through the primary cache L1 and the secondary cache L2, thus completing the processing of operation sequence 0.
[0210] After completing the processing of operation sequence 0, the processing unit can execute operation sequence 1. Among them, the process of the processing unit executing each operation in operation sequence 1 is similar to the processing process of executing each operation in operation sequence 0, as specifically shown Figure 10 in the timing diagram and will not be elaborated here.
[0211] Comparing Figure 8 and Figure 10 it can be seen that in Figure 8 the total number of times the processing unit reads and stores data in the DDR memory is 8 times (i.e., the number of accesses to the DDR memory is 8 times), while in Figure 10 the total number of times the processing unit reads data from and stores data in the DDR memory is 4 times. Therefore, in Figure 10 the number of accesses to the DDR memory required for the processing unit to execute the two operation sequences is reduced. Moreover, in Figure 8 it takes 28 clocks to complete the inference task, while in Figure 10 it only takes 24 clocks to complete the task inference (i.e., the processing unit executes these two operation sequences), thus improving the inference efficiency of task inference based on the artificial intelligence model.
[0212] Furthermore, in the case where the inference task of the artificial intelligence model corresponds to two operation sequences, in order to further improve the efficiency of task inference based on the artificial intelligence model, the present application can also process these two operation sequences in parallel through two processing units. As Figure 11 shown, a timing example diagram of two processing units processing two operation sequences in parallel is shown.
[0213] From Figure 11It can be seen that while the processing unit 0 executes each operation in the operation sequence 0, the processing unit 1 also executes each operation in the operation sequence 1. Among them, the process of the processing unit 0 executing each operation in the operation sequence 0 is similar to the process of a single processing unit executing each operation in the operation sequence 0 before, and the process of the processing unit 1 executing each operation in the operation sequence 1 is also similar to the process of the processing unit executing each operation in the operation sequence 1 before, which will not be elaborated here. Figure 10 The process of a single processing unit executing each operation in the operation sequence 0 before is similar, and the process of the processing unit 1 executing each operation in the operation sequence 1 is also similar to the process of the processing unit executing each operation in the operation sequence 1 before, which will not be elaborated here. Figure 10 The process of the processing unit executing each operation in the operation sequence 1 before is similar, which will not be elaborated here.
[0214] Comparing Figure 10 and Figure 11 It can be seen that by parallelly executing two operation sequences through two processing units, the total time required to execute the two operation sequences can be further reduced. Based on this, compared with Figure 8 the initial operation sequence corresponding to the inference task of the artificial intelligence model processed by a single processing unit before, the two operation sequences split from the inference task of the artificial intelligence model processed in parallel by two processing units can not only reduce the number of accesses of the processing unit to the DDR memory, but also reduce the time required for task inference based on the artificial intelligence model. As Figure 11 can be seen, it only takes 12 clocks for the two processing units to parallelly execute the two operation sequences.
[0215] In the above Figures 8 to 11 in order to facilitate the understanding of the benefits of splitting the inference task of the artificial intelligence model into two operation sequences and the processing unit reading input data from the second memory, the target operations that need to depend on weight data in the operation sequence and the process of reading the weight data corresponding to the target operations are not introduced.
[0216] It can be understood that in the case where the inference task of the artificial intelligence model corresponds to at least two operation sequences, if there are target operations that need to depend on weight data in the operation sequence, then when the processing unit executes the target operations in each operation sequence, it needs to read the operation weight data of the target operation from the first memory. In this way, the same target weight data needs to be repeatedly read from the first memory multiple times, thus affecting the data reading efficiency due to the bandwidth limitation between the first memory and the second memory, and further affecting the performance of executing the inference task based on the artificial intelligence model.
[0217] Especially, in the case where multiple operation sequences are scheduled to be executed by multiple different processing units, since the target operations that depend on weight data in multiple operation sequences are in the same sequential position, therefore, in the case of multiple processing units parallelly processing multiple operation sequences, there will be a situation where multiple processing units simultaneously read the same target weight data from the first memory at the same time, thus increasing the data reading pressure on the first memory and further exacerbating the time required to read the weight data. Illustrated with Figure 12 for explanation.Figure 12 A comparative example diagram showing two implementation logics for performing an inference task based on an artificial intelligence model is shown.
[0218] In Figure 12 the left side of the arrow is the implementation logic for an operation sequence corresponding to the inference task of the artificial intelligence model. In this implementation logic, it is still assumed that the operation sequence includes operation 0 and operation 1. At the same time, it is assumed that operation 1 is the target operation that needs to depend on weight data. Then, it can be seen from the implementation logic on the left side of the arrow that when the processing unit executes operation 1, it needs to first read the weight data from the DDR memory, and then it can perform the calculation processing corresponding to this operation 1.
[0219] From Figure 12 the implementation logic diagram on the right side of the arrow, it can be seen that after splitting the inference task of the artificial intelligence model into two operation sequences, when executing operation 10 in operation sequence 0 and operation 11 in operation sequence 1, the weight data needs to be read from the DDR memory. In this way, the weight data needs to be read from the DDR memory twice, increasing the access burden on the DDR and also increasing the time required for performing task inference based on the artificial intelligence model.
[0220] Moreover, if the two operation sequences 0 and 1 are executed in parallel by two processing units, then these two processing units may simultaneously request to read the same weight data from the DDR memory, further exacerbating the access pressure on the DDR memory and increasing the time required for each processing unit to request this weight data.
[0221] In order to reduce the situation of repeatedly reading the same target weight data from the DDR memory multiple times, the present application can add a process branch to the target operations that need to depend on weight data in the two operation sequences during the process of compiling the model file data of the artificial intelligence model, as Figure 13 shown. Figure 13 This is an example diagram of an implementation logic of the operation control method of the present application.
[0222] In Figure 13 it is assumed that the second operation in operation sequence 0 and operation sequence 1 is a GEMM operation that depends on the same weight data, that is, Figure 13 operation 10 in operation sequence 0 and operation 11 in operation sequence 1 in
[0223] are both GEMM operations. On this basis, during the compilation stage of compiling the operation codes of operation sequence 0 and operation sequence 1, weight reading tasks corresponding to this operation 10 and operation 11 will be generated. Figure 13 Compared with Figure 12 the two operation sequences on the right side of the arrow in Figure 13A weight reading task associated with the operation 10 and the operation 11 is newly added. By executing this weight reading task, the weight data corresponding to the GEMM operation is read from the DDR memory and cached in the secondary cache L2.
[0224] Moreover, attribute parameters are added to the weight reading task, the operation 10, and the operation 11, such as Figure 13 In, the attribute parameter is a weight address parameter, and the weight address parameter may include the address information of the target data storage area allocated from the second memory for storing the weight data corresponding to the GEMM operation.
[0225] In Figure 13 On this basis, the scheduling processing unit 0 and the processing unit 1 respectively process the operation sequence 1 and the operation sequence 2, and before the processing unit 0 executes the operation 10 and the processing unit 1 executes the operation 11, the scheduling processing unit 2 executes the weight reading task, and the weight data read from the DDR memory is stored in the target data storage area of the secondary cache L2 through the processing unit 2.
[0226] To facilitate a more intuitive understanding of the benefits of scheduling the processing unit 2 to execute the weight reading task before executing the operation 10 and the operation 11, it is described in combination with Figure 14 As Figure 14 shows a timing example diagram of scheduling the weight reading task and the two operation sequences in the present application.
[0227] From Figure 14 It can be seen that in the present application, before the processing unit 0 and the processing unit 1 execute each operation in their respective operation sequences, the scheduling processing unit 2 is first scheduled to execute the weight reading task to read the weight data from the DDR memory into the secondary cache L2. On this basis, when the processing unit 0 executes the operation 10 and the processing unit 1 executes the operation 11, the required weight data can be read from the secondary cache L2, and there is no need to separately access the DDR memory to read the weight data, so naturally the situation of repeatedly reading the weight data from the DDR can be reduced.
[0228] Corresponding to an operation control method of the present application, the present application also provides an operation control device.
[0229] As Figure 15 , shows a schematic structural diagram of a composition of the operation control device provided by the present application. The operation control device may include:
[0230] A sequence allocation unit 1501, configured to determine a first processing unit for executing an operation sequence, where the operation sequence includes at least one operation required to process an inference task based on an artificial intelligence model;
[0231] A weight prefetching unit 1502, configured to, in response to a target operation that depends on weight data existing in the operation sequence, schedule a second processing unit to read the target weight data from a first memory and store the target weight data into a second memory before the first processing unit acquires the target weight data associated with the target operation, where the first processing unit and the second processing unit are connected to the second memory, and a data bandwidth between the second memory and the first processing unit and the second processing unit is greater than a data bandwidth between the second memory and the first memory;
[0232] A weight reading unit 1503, configured to, during the process of the first processing unit executing the operation sequence, read the target weight data from the second memory based on the target operation.
[0233] In a possible implementation, the weight prefetching unit includes:
[0234] A task scheduling subunit, configured to, in response to a target operation that depends on weight data existing in the operation sequence, determine a weight reading task associated with the target operation and schedule the weight reading task to the second processing unit before the first processing unit acquires the target weight data associated with the target operation, where the weight reading task is used to indicate reading the target weight data from the first memory.
[0235] In a possible implementation, the operation control device further includes:
[0236] A first scheduling unit, configured to sequentially schedule each operation in the operation sequence to the first processing unit;
[0237] Alternatively, a second scheduling unit, configured to schedule the operation sequence to the first processing unit.
[0238] The weight prefetching unit includes:
[0239] A first prefetching subunit, configured to, in response to a target operation that depends on weight data existing in the operation sequence, schedule the second processing unit to read the target weight data associated with the target operation from the first memory and store the target weight data into the second memory before scheduling the target operation to the first processing unit;
[0240] Alternatively, a second prefetching subunit, configured to, in response to a target operation that depends on weight data existing in the operation sequence, schedule the second processing unit to read the target weight data associated with the target operation from the first memory and store the target weight data into the second memory before the first processing unit executes the target operation.
[0241] In yet another possible implementation, the weight prefetching unit includes:
[0242] A status determination subunit, configured to determine the operating status of each processing unit in the set of processing units;
[0243] A scheduling subunit, configured to schedule a second processing unit in the set of processing units whose operating status meets the requirements to read the target weight data from the first memory.
[0244] In yet another possible implementation, when storing the target weight data into the second memory, the weight prefetching unit is specifically configured to store the target weight data into a target data storage area in the second memory, where the target data storage area is an area allocated for storing weight data for the target operation during the process of compiling the operation code corresponding to the target operation;
[0245] The weight reading unit includes:
[0246] A weight reading subunit, configured to read the target weight data from the target data storage area based on the address information indicated by the target attribute parameter in the target operation, where the target attribute parameter is a parameter generated for the target operation during the process of compiling the operation code corresponding to the target operation and used to characterize the target data storage area.
[0247] In yet another possible implementation, when storing the target weight data into the target data storage area in the second memory, the weight prefetching unit is specifically configured to: in response to the maximum available storage capacity of the target data storage area being less than the data volume of the target weight data, divide the target weight data into at least two data blocks, where the size of each data block is less than or equal to the maximum available storage capacity; and store one of the at least two data blocks into the target data storage area each time;
[0248] The weight reading subunit includes:
[0249] A data block reading subunit, configured to sequentially read each data block from the target data storage area based on the address information indicated by the target attribute parameter in the target operation when the first processing unit executes the target operation;
[0250] The operation control device further includes:
[0251] A block continuous storage unit, configured to, after the first processing unit reads a data block from the target data storage area, sequentially store each data block that has not been stored into the target data storage area among the at least two data blocks into the target data storage area through the second processing unit.
[0252] In yet another possible implementation, the sequence allocation unit includes:
[0253] A sequence determination subunit, configured to determine at least two operation sequences required for the artificial intelligence model to process an inference task, where the operation sequences include at least one operation with a sequential order, and the operation types corresponding to the operations at the same sequential position in different operation sequences are the same, but the input data processed by different operation sequences is different;
[0254] A sequence allocation subunit, configured to, for each operation sequence, determine a first processing unit for executing the operation sequence. In yet another possible implementation, the first processing units corresponding to different operation sequences determined by the sequence determination subunit are different; wherein, the weight data required for the target operations at the same sequential position in different operation sequences and depending on weight data is the same;
[0255] Specifically, the weight pre-reading unit is configured to, in response to the existence of a target operation depending on weight data in the operation sequence, schedule a second processing unit to read the target weight data from a first memory and store the target weight data in a second memory before each first processing unit acquires the target weight data associated with the target operation.
[0256] An embodiment of the present application further provides an electronic device. As Figure 16 shown, it shows a schematic structural diagram of a composition of the electronic device, and the electronic device at least includes: a scheduling unit 1601, multiple processing units 1601, a first memory 1603, and a second memory 1604.
[0257] Wherein, each processing unit is connected to the first memory, and the data bandwidth between the second memory and the processing unit is greater than the data bandwidth between the second memory and the first memory.
[0258] The scheduling unit is configured to determine a first processing unit for executing an operation sequence, where the operation sequence includes: at least one operation required for the artificial intelligence model to process an inference task; in response to the existence of a target operation depending on weight data in the operation sequence, before the first processing unit acquires the target weight data associated with the target operation, schedule a second processing unit to read the target weight data from the first memory and store the target weight data in the second memory.
[0259] The first processing unit is configured to, during the execution of the operation sequence, based on the target operation, read the target weight data from the second memory.
[0260] Wherein, the first processing unit and the second processing unit belong to the multiple processing units.
[0261] Among them, for the specific operations of the scheduling unit, the first processing unit, and the second processing unit, reference can be made to the relevant introductions in the previous operation control method, which will not be elaborated here.
[0262] Of course, the electronic device may also include a display unit, an input unit, etc., without specific limitations.
[0263] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the operation control methods provided by the embodiments of the present application.
[0264] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the operation control methods provided by the embodiments of the present application.
[0265] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present application, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0266] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0267] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0268] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. An operation control method, comprising: Determining a first processing unit for executing a sequence of operations, the sequence of operations comprising: at least one operation required to be performed to process an inference task based on an artificial intelligence model; In response to the existence of a target operation that depends on weight data in the operation sequence, before the first processing unit obtains the target weight data associated with the target operation, scheduling the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory, the first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory; During execution of the operation sequence by the first processing unit, the target weight data is read from the second memory based on the target operation.
2. The operation control method according to claim 1, wherein the scheduling second processing unit reads the target weight data from the first memory, comprising: A weight reading task associated with the target operation is determined, and the weight reading task is scheduled to a second processing unit, wherein the weight reading task is used to instruct to read the target weight data from the first memory.
3. The operation control method according to claim 1, further comprising: Dispatching each operation in the operation sequence to the first processing unit in sequence; Alternatively, scheduling the operation sequence to the first processing unit; Before the first processing unit acquires the target weight data associated with the target operation, scheduling the second processing unit to read the target weight data from the first memory includes: Before dispatching the target operation to the first processing unit, dispatching the second processing unit to read the target weight data associated with the target operation from the first memory; Alternatively, before the first processing unit executes the target operation, the second processing unit is scheduled to read the target weight data associated with the target operation from the first memory.
4. The operation control method according to claim 1, wherein the scheduling the second processing unit to read the target weight data from the first memory comprises: Determining the operating status of each processing unit in the processing unit set; The second processing unit in the processing unit set whose operating status meets the requirements is scheduled to read the target weight data from the first memory.
5. The operation control method according to claim 1, wherein storing the target weight data into a second memory comprises: storing the target weight data in a target data storage area in a second memory, wherein the target data storage area is a storage area for storing weight data allocated to the target operation during the process of compiling the operation code corresponding to the target operation; The step of reading the target weight data from the second memory based on the target operation includes: The target weight data is read from the target data storage area based on the address information indicated by the target attribute parameter in the target operation, wherein the target attribute parameter is a parameter generated for the target operation in the process of compiling the operation code corresponding to the target operation and used to characterize the target data storage area.
6. The operation control method according to claim 5, wherein storing the target weight data in a target data storage area in the second memory comprises: In response to the maximum available storage capacity of the target data storage area being less than the data capacity of the target weight data, dividing the target weight data into at least two data blocks, each of which has a size less than or equal to the maximum available storage capacity; storing one of the at least two data blocks into the target data storage area each time; The step of reading the target weight data from the target data storage area based on the address information indicated by the target attribute parameter in the target operation includes: When the first processing unit executes the target operation, based on the address information indicated by the target attribute parameter in the target operation, each data block is read from the target data storage area in sequence; The operation control method further includes: After the first processing unit reads the data block from the target data storage area, the second processing unit sequentially stores each of the at least two data blocks that has not been stored in the target data storage area into the target data storage area.
7. The operation control method according to claim 1, wherein determining the first processing unit for executing the operation sequence comprises: Determine at least two operation sequences that the artificial intelligence model needs to execute to process the reasoning task, wherein the operation sequence includes at least one operation having a sequential order, and operations at the same sequential position in different operation sequences correspond to the same operation type, but different operation sequences process different input data; For each sequence of operations, a first processing unit for executing the sequence of operations is determined.
8. The operation control method according to claim 7, wherein different operation sequences correspond to different first processing units; The weight data required by target operations at the same sequential position in different operation sequences and requiring dependence on weight data are the same; The step of scheduling the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory before the first processing unit acquires the target weight data associated with the target operation comprises: Before each first processing unit obtains the target weight data associated with the target operation, the second processing unit is scheduled to read the target weight data from the first memory and store the target weight data in the second memory.
9. An operation control device, comprising: A sequence allocation unit, configured to determine a first processing unit for executing an operation sequence, wherein the operation sequence includes: at least one operation required to be performed to process an inference task based on an artificial intelligence model; a weight pre-reading unit, configured to, in response to a target operation that depends on weight data in the operation sequence, schedule the second processing unit to read the target weight data from the first memory and store the target weight data in the second memory before the first processing unit obtains the target weight data associated with the target operation, wherein the first processing unit and the second processing unit are connected to the second memory, and the data bandwidth between the second memory and the first processing unit and the second processing unit is greater than the data bandwidth between the second memory and the first memory; A weight reading unit is used to read the target weight data from the second memory based on the target operation during the process of the first processing unit executing the operation sequence.
10. The operation control device according to claim 9, wherein the weight pre-reading unit comprises: A task scheduling subunit is used to respond to the existence of a target operation that depends on weight data in the operation sequence, and before the first processing unit obtains the target weight data associated with the target operation, determine the weight reading task associated with the target operation, and schedule the weight reading task to the second processing unit, wherein the weight reading task is used to instruct the target weight data to be read from the first memory.