Task execution method and device, electronic equipment and storage medium
By multiplexing the same target instruction code in the storage unit in the calculation unit to execute different instruction flows, the problem of high resource occupation in the model training or inference process of computing units such as graphics processors is solved, and task execution efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202510286532.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-20
AI Technical Summary
In the process of performing model training or inference by computing units such as graphics processors, it is necessary to consume more resources to schedule the code of the instruction stream of data processing tasks, resulting in low task execution efficiency of computing units and high storage space occupancy cost.
By determining the target instruction attribute from the instruction attribute set by using the first computing unit and sending the target address to the second computing unit, the first instruction stream and the second instruction stream are executed to obtain the execution result of the target processing task when the target instruction code is multiplexed based on the execution logic of the first target instruction and the second target instruction.
By multiplexing the same target instruction code in the storage unit to execute different instruction streams, the full loading of the instruction code of the instruction stream is avoided, the storage space occupation of the storage unit and the burden of communication bandwidth is reduced, and the execution efficiency of target processing tasks and the utilization rate of computing resources is improved.
Smart Images

Figure CN120179360A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of deep learning technology and large model technology. More specifically, the present disclosure provides a task execution method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, various data generated in various scenarios can be processed based on a processor such as a Graphics Processing Unit (GPU) according to deep learning algorithms. Summary of the Invention
[0003] The present disclosure provides a task execution method, apparatus, electronic device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a task execution method, including: determining a target instruction attribute from an instruction attribute set by using a first computing unit, where the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream, the first instruction stream and the second instruction stream are used for a target processing task, and the target instruction attribute includes a target address of a target instruction code in a storage unit; sending the target address to a second computing unit by using the first computing unit; reading the target instruction code from the target address of the storage unit by using the second computing unit; and executing the first instruction stream and the second instruction stream by using the second computing unit to obtain an execution result of the target processing task in a case where the target instruction code is reused based on the execution logics of the first target instruction and the second target instruction respectively.
[0005] According to another aspect of the present disclosure, there is provided a task execution apparatus, including: a storage unit; a first computing unit configured to: determine a target instruction attribute from an instruction attribute set, where the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream, the first instruction stream and the second instruction stream are used for a target processing task, and the target instruction attribute includes a target address of a target instruction code in the storage unit; send the target address to a second computing unit; and a second computing unit configured to: read the target instruction code from the target address of the storage unit; and execute the first instruction stream and the second instruction stream to obtain an execution result of the target processing task in a case where the target instruction code is reused based on the execution logics of the first target instruction and the second target instruction respectively.
[0006] According to another aspect of the present disclosure, there is provided a task execution device, including: the apparatus provided according to an embodiment of the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method as described above.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which the task execution method and apparatus according to embodiments of the present disclosure can be applied;
[0013] Figure 2 Schematically shows a flowchart of the task execution method according to embodiments of the present disclosure;
[0014] Figure 3 Schematically shows a schematic diagram of the principle of the task execution method according to embodiments of the present disclosure;
[0015] Figure 4 Schematically shows a schematic diagram of the principle of the task execution method according to another embodiment of the present disclosure;
[0016] Figure 5 Schematically shows a flowchart of the task execution method according to another embodiment of the present disclosure;
[0017] Figure 6 Schematically shows a block diagram of the task execution apparatus according to embodiments of the present disclosure;
[0018] Figure 7 Schematically shows a block diagram of the task execution device according to embodiments of the present disclosure; and
[0019] Figure 8A schematic block diagram of an exemplary electronic device that can be used to implement the task execution method of the embodiments of the present disclosure is shown. Detailed implementation manners
[0020] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the descriptions of well-known functions and structures are omitted below.
[0021] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0022] The inventors found that during the execution of model training or inference by computing units such as graphics processors, a relatively large amount of resources are required to schedule the code of the instruction stream of data processing tasks, which occupies a relatively large amount of storage resources and communication bandwidth of the computing device, resulting in low task execution efficiency of the computing unit and high storage space occupancy cost.
[0023] Embodiments of the present disclosure provide a task execution method, apparatus, electronic device, and storage medium. The task execution method includes: using a first computing unit to determine a target instruction attribute from an instruction attribute set, where the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream, the first instruction stream and the second instruction stream are used for a target processing task, and the target instruction attribute includes a target address of the target instruction code in a storage unit; using the first computing unit to send the target address to a second computing unit; using the second computing unit to read the target instruction code from the target address of the storage unit; and when the target instruction code is reused based on the execution logics of the first target instruction and the second target instruction respectively, using the second computing unit to execute the first instruction stream and the second instruction stream to obtain an execution result of the target processing task.
[0024] According to an embodiment of the present disclosure, by using an instruction attribute set to store the instruction attributes of each instruction code, and using the code address of the instruction code in the storage unit in the instruction attributes, before executing the first instruction stream and the second instruction stream, by sending the target instruction code address to the second computing unit, it is realized that the second computing unit can reuse the same target instruction code stored in the storage unit when executing different instruction streams to execute the instructions in different instruction streams, thereby avoiding the negative impacts such as large storage space occupation of the storage unit and communication bandwidth occupation of code transmission resources caused by full loading of the instruction codes of the instruction stream, improving the execution efficiency of the target processing task including multiple instruction streams, and saving the computing resource occupation amount for executing the target processing task.
[0025] Figure 1 Schematically shows an exemplary system architecture to which the task execution method and apparatus according to an embodiment of the present disclosure can be applied.
[0026] It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0027] As Figure 1 Shown, the system architecture 100 according to this embodiment may include a terminal device 101, a network 102, and a server cluster 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server cluster 103. The network 102 can also be used to provide a medium for a communication link within the server cluster 103. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.
[0028] The user can use the terminal device 101 to interact with the server cluster 103 through the network 102 to receive or send messages, etc. For example, the terminal device 101 can send a request for training a deep learning model to the server cluster 103 through the network 102.
[0029] Various communication client applications can be installed on the terminal device 101, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0030] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. The terminal device 101 can perform reading and writing of code data based on a set computing unit such as a CPU (Central Processing Unit, central processor).
[0031] The server cluster 103 may be a server that provides various services. For example, it may be a background management server (only for example) that supports requests sent by the user using the terminal device 101.
[0032] The server cluster 103 may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0033] The server cluster 103 includes multiple server nodes 1031, 1032, 1033, and each server node includes one or more hardware devices. The computing units such as the CPU (Central Processing Unit) and GPU (Graphics Processing Unit) of the hardware devices in the terminal device 101 can be used as the first computing unit, and the computing units such as the CPU (Central Processing Unit) and GPU (Graphics Processing Unit) of the hardware devices in the server cluster 103 or the server node can be used as the second computing unit to execute the method provided by the present disclosure. Based on the task execution method provided by the present disclosure, a large number of instruction streams in the target processing task can be executed with lower storage resources.
[0034] Figure 2 A flowchart of the task execution method according to an embodiment of the present disclosure is schematically shown.
[0035] As Figure 2 shown, the task execution method includes operations S210 to S240.
[0036] In operation S210, the target instruction attribute is determined from the instruction attribute set using the first computing unit.
[0037] According to an embodiment of the present disclosure, the first computing unit or the second computing unit may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence computing unit. The artificial intelligence computing unit may include at least one of a neural network processing unit (NPU), a tensor processing unit (Tensor Processing Unit, tensor processor), and Kunlun Core. The first computing unit and the second computing unit may be of the same type of computing unit. For example, both the first computing unit and the second computing unit are GPUs. Alternatively, the first computing unit and the second computing unit may be of different types of computing units.
[0038] According to an embodiment of the present disclosure, the instruction attribute set may include the instruction attributes of respective multiple instruction codes. The instruction attributes may include the attribute information of the instruction codes such as the code length, code name, and code address of the instruction code. The instruction code may be used to execute a specified type of data processing instruction. For example, it may execute a convolution calculation instruction, a multiplication calculation instruction, etc. The specific instruction type executed by the instruction code in the embodiment of the present disclosure is not limited.
[0039] According to an embodiment of the present disclosure, the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream. The first instruction stream may include multiple first instructions and the execution logic relationship between the multiple first instructions. The second instruction stream may include multiple second instructions and the execution logic relationship between the multiple second instructions. The first instruction stream and the second instruction stream are used for a target processing task. The first target instruction and the second target instruction may be executed based on the same target instruction code. The target instruction attribute includes the target address of the target instruction code in the storage unit.
[0040] In operation S220, the first computing unit is used to send the target address to the second computing unit.
[0041] For example, based on the bus between the first computing unit and the second computing unit, the target address may be sent to the second computing unit so that the second computing unit can determine based on the target address
[0042] In operation S230, the second computing unit is used to read the target instruction code from the target address of the storage unit.
[0043] In operation S240, when the target instruction code is reused based on the execution logic of the first target instruction and the second target instruction respectively, the second computing unit is used to execute the first instruction stream and the second instruction stream to obtain the execution result of the target processing task.
[0044] According to an embodiment of the present disclosure, a storage unit can be used to store instruction codes respectively corresponding to multiple instruction attributes in an instruction attribute set, and the instruction address in the instruction attribute can indicate the storage address of the instruction code in the storage unit.
[0045] According to an embodiment of the present disclosure, the second computing unit can read the same target instruction code in the storage unit based on the received target address, and respectively execute the first target instruction based on the execution logic of the first instruction stream, and reuse the same target instruction code to execute the second target instruction based on the execution logic of the second instruction stream, so as to realize reusing the target instruction code stored at the same target address in the storage unit to execute the first instruction stream and the second instruction stream, obtain the first execution result of the first instruction stream and the second execution result of the second instruction stream, and further determine the execution result of the target processing task based on the first execution result and the second execution result. In this way, during the execution of the target processing task, the second computing unit can be utilized to reuse the same target instruction code in the storage unit to execute different instruction streams, avoiding the excessive occupation of storage resources caused by the first computing unit loading all the instruction codes of each instruction stream into the storage unit. At the same time, by storing multiple preset instruction codes in the storage unit, the instruction streams to be executed do not need to fully load all the instruction codes of the instruction streams in the storage unit before execution, saving the occupation of communication bandwidth resources between the first computing unit and the second computing unit, and saving the loading waiting time before executing the instruction streams by preloading the instruction codes of multiple instruction attributes into the storage unit respectively, improving the processing efficiency of the target processing task and the computing efficiency of the second computing unit.
[0046] In one example, the first instruction stream or the second instruction stream can be different processes in the target processing task, and the first target instruction or the second target instruction can be understood as the thread instructions used to execute according to the execution logic of the process.
[0047] According to an embodiment of the present disclosure, the target processing task can include any type of data processing task. For example, it can include the training task of a deep learning model, the answer inference task of a large model processing question and answer information, etc. The instruction stream of the target processing task can be an instruction sequence for executing the data processing task. For example, in the inference task of a large model processing question and answer information, the instruction stream can be data multiplication instructions, data accumulation instructions, etc. for executing the multi-head self-attention mechanism algorithm. The embodiments of the present disclosure do not limit the specific data processing types executed by the target processing task and the instruction stream. It should be understood that multiple instructions in the instruction stream, for example, can be multiple data processing instructions for executing the data processing process of the network layer according to the algorithm logic of the same network layer to obtain hidden features.
[0048] According to an embodiment of the present disclosure, the first target instruction or the second target instruction includes at least one of the following: point cloud data processing instruction, graph data processing instruction, and image feature processing instruction.
[0049] According to an embodiment of the present disclosure, the point cloud data processing instruction may be an instruction related to point cloud data processing tasks such as three-dimensional model reconstruction and target object detection based on point cloud data. The point cloud data processing task may include multiple instruction streams for processing point cloud data, and each instruction stream may include multiple point cloud data processing instructions for processing point cloud data.
[0050] In one example, the point cloud data processing instruction may include multiple instructions for processing point cloud data according to the Point2Grid algorithm. For example, the point cloud data processing instruction may be a point cloud coordinate conversion instruction, a point cloud feature fusion instruction, and so on.
[0051] According to an embodiment of the present disclosure, the graph data processing instruction may be a data processing instruction for processing graph data such as a knowledge graph. For example, multiple graph data processing instructions in the graph data processing instruction stream may be multiple graph data processing instructions for fusing and calculating graph data features according to the algorithm logic of the hidden layer in the graph neural network model.
[0052] According to an embodiment of the present disclosure, the image feature processing instruction may be a data processing instruction for processing image feature data such as image texture features and image semantic features. For example, the image feature processing instruction stream may be multiple image feature processing instructions for extracting texture features of image features according to the algorithm logic of the texture feature extraction network layer.
[0053] In one example, the instruction attribute may be a descriptor of the instruction code, and the descriptor may store code attribute information such as the code address, code length, and code name for describing the instruction code. The instruction code may be, for example, a kernel function code for performing a specific function calculation task. The instruction attribute base may store descriptors of multiple instruction codes based on the structure form of the descriptor heap.
[0054] It should be noted that the first target instruction or the second target instruction may also be a data processing instruction for processing any other type of data using the second computing unit. The embodiments of the present disclosure do not limit the specific data type processed by the first target instruction or the second target instruction.
[0055] According to an embodiment of the present disclosure, determining a target instruction attribute from an instruction attribute set using a first computing unit may include using the first computing unit to query the instruction attribute set based on the instruction attribute index of a first target instruction in a first instruction stream, so as to obtain a target instruction attribute having a mapping relationship with the instruction attribute index of the first target instruction. The target instruction attribute may indicate that the target instruction code for executing the first target instruction has been stored in a storage unit for instruction scheduling by a second computing unit. In addition, using the first computing unit to query the instruction attribute set based on the instruction attribute index of a second target instruction in a second instruction stream may obtain the same target instruction attribute associated with the second target instruction. This may indicate that the second computing unit needs to call the same target instruction code as that for executing the first target instruction when executing the second target instruction in the second instruction stream, and the target instruction code can be reused by the second computing unit when executing the second instruction stream in the storage unit, thereby avoiding full preloading of the code of each instruction stream of the target processing task in the storage unit, and by sending the address of the instruction code in the instruction stream to the second computing unit, enabling the second computing unit to reuse the target instruction code with the same target address to execute different instruction streams, reducing the storage space occupation of the storage unit.
[0056] In one example, the instruction attribute index of an instruction in an instruction stream may be obtained by processing the instruction code for executing the instruction using a preset algorithm such as a digest algorithm. For example, a digest algorithm may be used to calculate the convolution calculation code for executing a convolution calculation instruction to obtain the instruction attribute index of the convolution calculation code. Establishing a mapping relationship between the instruction attribute index of the instruction code and the instruction attribute may enable the first computing unit to query the instruction attribute set according to the instruction attribute indexes of the respective instructions in the instruction stream to obtain the instruction code addresses corresponding to the respective instructions in the instruction stream.
[0057] In one example, the instruction attribute index may be determined based on the code identifier for executing the instruction. For example, a hash calculation may be performed on the convolution calculation instruction code identifier for executing a convolution calculation instruction, and the obtained hash value may be determined as the instruction attribute index of the convolution calculation instruction code.
[0058] According to an embodiment of the present disclosure, the instruction attributes in the instruction attribute set are associated with a count value, and the count value represents the number of times the instruction code address of the instruction attribute has been sent to the second computing unit. For example, in an execution stream, there are multiple instructions to be executed. The instruction attribute index of the i-th instruction is associated with the i-th instruction code address of the i-th instruction attribute in the instruction attribute set. Thus, the count value of the i-th instruction attribute can be incremented by 1, and the i-th instruction code address of the i-th instruction attribute can be sent to the second computing unit.
[0059] It can be understood that in the process of executing a large number of instructions of each execution flow in the target processing task, the frequency of the instruction code stored in the storage unit being called during the execution of the target processing task can be determined based on the count value of the instruction attribute, so that the importance of the associated instruction code can be determined by the count value of the instruction attribute.
[0060] According to an embodiment of the present disclosure, the query order of the first computing unit for the multiple instruction attributes in the instruction attribute set is determined based on the respective count values associated with the multiple instruction attributes.
[0061] According to an embodiment of the present disclosure, multiple instruction attributes in an instruction attribute set can be sorted based on the count values associated with each instruction attribute, and the instruction attribute with a larger count value is sorted higher. In this way, multiple instruction attributes in the instruction attribute set can be stored according to a sorting level determined based on the count value, and the larger the technical value of the instruction attribute, the higher the sorting level, so that the target instruction code in the target processing task that is frequently called by each execution flow to execute instructions can be queried by the first computing unit more quickly and sent to the second computing unit, so as to improve the query efficiency and call efficiency of the target instruction code that is frequently used in the instruction flow, thereby improving the timeliness of the target instruction code being called by the second computing unit, thereby improving the execution efficiency of the instruction flow and the target processing task.
[0062] According to an embodiment of the present disclosure, the count value related to each instruction attribute in the instruction attribute set can also be dynamically adjusted according to the number of times the instruction proxy address corresponding to the instruction attribute is sent to the second computing unit, so that the query order of multiple instruction attributes can be dynamically adjusted during the execution of multiple execution streams of the target processing task, and then the timeliness of multiple instruction addresses being sent to the second computing unit can be dynamically adjusted in the process of executing the instructions of the instruction stream according to the instruction execution logic, so as to dynamically adjust the query order of the instruction address whose instruction code is called more times as the instruction execution logic causes the execution process of the execution stream, so as to dynamically improve the query efficiency and call efficiency of the instruction code whose call times increase according to the instruction execution logic, and then improve the loading efficiency and execution efficiency of the instruction code according to the execution order of the execution stream of the target processing task and the instruction execution logic of the instruction stream, and improve the computing efficiency of the second computing unit.
[0063] In one example, multiple execution stages may also be set according to the execution order of multiple instruction streams based on the algorithm type of each instruction in the multiple instruction streams of the target processing task. At the end of each execution stage, the count value of the instruction attributes in the instruction attribute set is cleared or initialized to a preset value. In the second execution stage, the multiple instruction attributes in the instruction attribute set may be sorted according to the count value of 0 or the initialized preset value, so as to dynamically adjust the query order according to the instruction attributes of the instruction codes that need to be frequently called in the instruction stream in the current execution stage, realize fast query and sending of the instruction addresses required by the instruction stream in the current execution stage, and dynamically improve the query efficiency and execution efficiency of the instruction codes according to the algorithm type of the instruction stream in different execution stages.
[0064] Figure 3 Schematically shows a schematic diagram of the principle of the task execution method according to an embodiment of the present disclosure.
[0065] As Figure 3 shown, the first storage unit is for the first computing unit, and the second storage unit is for the second computing unit. The first storage unit stores an instruction attribute set, which is a plurality of instruction descriptors stored based on a tree structure. The plurality of instruction descriptors are respectively associated with a plurality of preset instruction codes pre-loaded in the second storage unit. For example, the 1st instruction descriptor, 2nd instruction descriptor, 3rd instruction descriptor, 4th instruction descriptor, and 5th instruction descriptor stored in the first storage unit respectively correspond to the 1st instruction code, 2nd instruction code, 3rd instruction code, 4th instruction code, and 5th instruction code stored in the second storage unit. The plurality of instruction descriptors in the first storage unit are arranged in query order based on the count value of the instruction descriptor. For example, the 1st instruction descriptor, 2nd instruction descriptor, 3rd instruction descriptor, 4th instruction descriptor, and 5th instruction descriptor respectively correspond to the count values 12, 9, 9, 3, and 3.
[0066] As Figure 3As shown, the second computing unit is used to execute the first instruction stream C310 and the second instruction stream C320 in the target processing task. The first computing unit can control the second computing unit to execute the target processing task for multiple first instructions in the first instruction stream C310 and multiple second instructions in the second instruction stream C320. The first computing unit queries the instruction attribute set of the first storage unit based on the instruction attribute index of the second first instruction C312 in the first instruction stream C310, and queries the third instruction descriptor as the target instruction attribute according to the query order determined based on the count value. The first computing unit reads the third instruction code address from the third instruction descriptor of the first storage unit and sends the third instruction code address to the second computing unit. The second computing unit reads the third instruction code from multiple instruction codes pre-loaded from the second storage unit according to the third instruction code address based on the execution logic of the first instruction stream C310, and executes the second first instruction of the first instruction stream C310 according to the third instruction code. In this way, the second computing unit can execute multiple first instructions from the instruction code addresses of each first instruction in the first instruction stream C310 obtained from the first computing unit, and obtain the instruction stream execution result of the first instruction stream C310.
[0067] As Figure 3 shown, before the second computing unit needs to execute the third second instruction C323 of the second instruction stream C320, the first computing unit can query the instruction attribute set of the first storage unit based on the instruction attribute index of the third second instruction C323, and query the third instruction descriptor as the target instruction attribute according to the query order determined based on the count value. The first computing unit reads the third instruction code address from the third instruction descriptor of the first storage unit and sends the third instruction code address to the second computing unit. The second computing unit reads the third instruction code from multiple instruction codes pre-loaded from the second storage unit according to the third instruction code address based on the execution logic of the second instruction stream C320, and executes the third second instruction C323 of the second instruction stream C320 according to the third instruction code. In this way, the second computing unit can execute multiple second instructions from the instruction code addresses of each second instruction in the second instruction stream C320 obtained from the first computing unit, and obtain the instruction stream execution result of the second instruction stream C320. The second computing unit can determine the execution result of the target processing task based on the instruction stream execution results of the first instruction stream C310 and the second instruction stream C320 respectively.
[0068] According to an embodiment of the present disclosure, the target instruction attribute is obtained by the first computing unit querying the instruction attribute set based on the query order and the instruction attribute index of the first target instruction, and the instruction attribute index of the first target instruction is determined by processing the instruction code of the first target instruction based on the integrity verification algorithm.
[0069] According to an embodiment of the present disclosure, the integrity verification algorithm may include algorithms such as the MD5 (Message-Digest Algorithm 5) algorithm and the SHA-256 (Secure Hash Algorithm 256) algorithm for verifying data integrity. The instruction code of the first target instruction is processed by the integrity verification algorithm to generate an instruction attribute index of the first target instruction, so that the instruction attribute index can indicate whether the target instruction code in the storage unit can execute the first target instruction completely, thereby avoiding using the instruction code similar to the execution of the first target instruction in the storage unit as the target instruction code to execute the first instruction stream, and further while reducing the communication bandwidth occupation of the computing unit and the storage space occupation of the storage unit, avoiding errors in the execution of the first instruction stream, and ensuring the computing accuracy and computing efficiency of the target processing task.
[0070] It should be understood that, based on the instruction attribute indexes of the multiple instructions in any one instruction stream of the target processing task, the instruction attributes associated with the instruction attribute indexes can be queried in the instruction attribute set, so as to verify the integrity and accuracy of the instructions in the instruction stream, and then send the instruction code address that meets the instruction execution requirements of the instruction stream to the second computing unit, so that the second computing unit can accurately call the instruction code stored in the storage unit according to the instruction execution logic of the instruction stream.
[0071] According to an embodiment of the present disclosure, the task execution method further includes: using the first computing unit to perform the following operations: querying the instruction attribute set based on the instruction attribute index of the third target instruction in the third instruction stream to obtain a query result; in the case where the query result indicates a query failure, determining the instruction attributes related to the instruction code of the third target instruction; and updating the instruction attribute set according to the instruction attributes related to the instruction code of the third target instruction, and writing the instruction code of the third target instruction into the storage unit.
[0072] According to an embodiment of the present disclosure, the data processing task includes a third instruction stream. The third instruction stream may be the same as the first instruction stream or the second instruction stream. Alternatively, the third instruction stream may be an instruction stream different from the first instruction stream and the second instruction stream in the target processing task.
[0073] According to an embodiment of the present disclosure, in the case where the query result indicates a query failure, it may indicate that there is no instruction attribute associated with the instruction attribute index of the third target instruction stored in the instruction attribute set, and there is no instruction code for executing the third target instruction stored in the storage unit. For example, the instruction attribute is represented based on an instruction descriptor, and there is no descriptor for the instruction code of the third target instruction stored in the instruction descriptor tree structure.
[0074] In one example, determining the instruction attributes related to the instruction code of the third target instruction may include constructing an instruction descriptor of the instruction code as the instruction attributes related to the instruction code of the third target instruction based on the code information such as the instruction code name, instruction code address, and instruction code length related to the instruction code of the third target instruction. By adding the instruction attributes related to the instruction code of the third target instruction to the current instruction attribute set and writing the instruction code of the third target instruction into the storage unit for the second computing unit, the instruction code required for the third target instruction of the third instruction stream is loaded. At the same time, the count value of 1 can be added to the instruction attributes related to the instruction code of the third target instruction in the instruction attribute set to synchronously update the count value of the instruction attributes in the instruction attribute set.
[0075] According to an embodiment of the present disclosure, the task execution method further includes: when the available capacity of the code storage area of the storage unit does not meet the preset capacity condition, using the first computing unit to determine the release instruction attributes that do not meet the count value condition from the instruction attribute set based on the count value; and using the first computing unit to delete the preset instruction code related to the release instruction attributes from the code storage area.
[0076] According to an embodiment of the present disclosure, the code storage area is used to store the preset instruction code related to the instruction attributes. The code storage area may be a storage area divided from the storage unit and shared by multiple arithmetic cores of the second computing unit. By setting a preset available capacity threshold for the code storage area, when the available capacity of the code storage area is less than or equal to the preset available capacity threshold, the instruction attributes with a count value less than or equal to the preset count value threshold can be determined as the release instruction attributes. And the first computing unit is used to delete the preset instruction code corresponding to the release instruction attributes to release the storage space of the code storage area, so as to delete the instruction code that is not frequently called by the second computing unit in the code storage area to expand the available capacity space of the code storage area, and further dynamically adjust the pre-loaded preset code instructions according to the execution logic of the instruction stream and the instruction code requirements, improving the utilization rate of storage resources.
[0077] In one example, the second computing unit may be a graphics processing unit, and the cache unit may be the video memory hardware for the graphics processing unit. The second computing unit may read the target instruction code from the video memory hardware based on the target address of the target instruction code in the video memory to reuse the target instruction code to execute the first target instruction and the second target instruction.
[0078] According to an embodiment of the present disclosure, the second computing unit includes a first arithmetic core and a second arithmetic core. The second computing unit may include multiple arithmetic cores. An arithmetic core can be understood as an element or block in the second computing unit for executing instructions. The first arithmetic core or the second arithmetic core can be different arithmetic cores in the second computing unit.
[0079] According to an embodiment of the present disclosure, when reusing the target instruction code based on the execution logics of the first target instruction and the second target instruction respectively, using the second computing unit to execute the first instruction stream and the second instruction stream includes: based on the execution logic of the first instruction stream, using the first arithmetic core to execute the first target instruction based on the target instruction code to obtain a first instruction execution result; and based on the execution logic of the second instruction stream, using the second arithmetic core to execute the second target instruction based on the target instruction code to obtain a second instruction execution result.
[0080] According to an embodiment of the present disclosure, the execution result of the target processing task is determined based on the first instruction execution result and the second instruction execution result. For example, the first instruction execution result is the convolution calculation result of the first convolutional network layer. Based on the first instruction execution result, the next first instruction having an execution logic relationship with the first target instruction can be executed according to the execution logic of the first instruction stream, so as to obtain the execution result of the first instruction stream of the first instruction stream. By reusing the target instruction code with the same target address, the second computing unit can execute the second target instruction according to the execution logic of the second instruction stream to obtain a second instruction execution result, and execute subsequent second instructions according to the second instruction execution result and the execution logic of the second instruction stream to obtain the execution result of the second instruction stream. In this way, the execution result of the target processing task can be determined based on the execution result of the first instruction stream and the execution result of the second instruction stream.
[0081] It should be noted that the execution logic of the instruction stream involved in the embodiments of the present disclosure can be understood as the instruction execution logic of the instruction stream. For example, the execution logic of the first instruction stream can be understood as the instruction execution logic among multiple first instructions in the first instruction stream. The embodiments of the present disclosure will not be elaborated herein.
[0082] In one example, the second computing unit can be a graphics processing unit, and the arithmetic cores in the second computing unit can be streaming multiprocessors (SMs) arranged in an array manner in the graphics processor. Different arithmetic cores can execute the instruction codes corresponding to different data processing instructions to implement the execution of data processing instructions in different instruction streams according to the instruction execution logic of the instruction stream, and further implement the reuse of the target instruction code pre-loaded in the storage unit by a large number of arithmetic cores to improve the storage space utilization rate of the storage unit for the second computing unit.
[0083] In one example, the arithmetic cores in the second computing unit can be CUDA cores (CUDA Core) arranged in an array in a graphics processor. The target instruction code can be a programming language that represents the execution logic of a kernel function. By enabling multiple different arithmetic cores in the second computing unit to reuse the same target instruction code stored in the storage unit based on the instruction execution logics of different instruction streams, and to execute the calculation logics of instructions based on the same kernel function in their respective instruction streams, a large number of CUDA cores can be used as arithmetic cores to reuse the target instruction code stored at the same address in the storage unit to implement parallel execution of multiple instruction streams, thereby saving the storage resources of the second computing unit and improving the computing efficiency of the target processing task.
[0084] Figure 4 Schematically shows a schematic diagram of the principle of a task execution method according to another embodiment of the present disclosure.
[0085] As Figure 4 shown, the second computing unit may include an arithmetic core array K410. The arithmetic core array K410 may include a plurality of arithmetic cores arranged in an array form, and the arithmetic cores can be CUDA cores. The plurality of arithmetic cores include a first arithmetic core K411 and a second arithmetic core K412. The second computing unit may further include a scheduler D410 and a buffer C410. The storage unit may preload a plurality of instruction codes, such as the 1st instruction code, the 2nd instruction code, and so on. For the first target instruction in the first instruction stream, the scheduler reads the 1st instruction code as the target instruction code from the storage unit using the 1st instruction code address as the target address. The 1st instruction code is stored in the buffer C410 of the second computing unit. When the execution logic based on the first instruction stream requires the first arithmetic core K411 to execute the first target instruction, the scheduler sends the 1st instruction code in the buffer C410 to the first arithmetic core K411 to configure the arithmetic core with the kernel function, enabling the first arithmetic core K411 to execute the first target instruction according to the 1st instruction code. After the execution of the first target instruction is completed, the 1st instruction code in the buffer C410 can be released to provide sufficient buffer capacity space for the buffer C410.
[0086] As Figure 4As shown, for the second target instruction in the second instruction stream, the scheduler reads the first instruction code as the target instruction code from the storage unit using the first instruction code address as the target address. The first instruction code is stored in the buffer C410 of the second computing unit. When the execution logic based on the second instruction stream requires the second arithmetic core K412 to execute the second target instruction, the scheduler sends the first instruction code in the buffer C410 to the second arithmetic core K412 to configure the kernel function of this core, enabling the second arithmetic core K412 to execute the second target instruction according to the first instruction code. This can reuse the target instruction code stored at the same target address in the storage unit based on the scheduler to execute instructions in different instruction streams, achieving an improvement in the utilization rate of the storage resources of the storage unit.
[0087] In one example, the task execution method provided by the embodiments of the present disclosure can be executed based on the following embodiments.
[0088] In this embodiment, the first computing unit is a central processing unit, the second computing unit is a graphics processing unit, the storage unit is the video memory unit of the graphics processing unit, and the video memory unit allocates a global code buffer area as the code storage area for multiple CUDA cores of the graphics processing unit. The target processing task is for the large model to process text data to perform the inference task of intelligent question answering. The target processing task may include multiple instruction streams corresponding to different network layers. For the preset instruction code, the first computing unit is used to construct an instruction descriptor corresponding to the instruction code based on the name, instruction code address, and instruction code length of each preset instruction code. The instruction descriptor is stored as an instruction attribute in the memory for the first computing unit, and multiple instruction descriptors are stored based on the tree structure of the descriptor tree. The structural level of each instruction descriptor corresponds to the count value of this instruction descriptor. At the same time, the first computing unit calculates the digest value of each preset instruction code as the instruction attribute index based on the MD5 algorithm, and associates the instruction attribute index with the instruction descriptor stored in the storage unit.
[0089] For the first instruction stream for the target processing task, use the first computing unit to calculate the digest value of each first instruction in the first instruction stream based on the MD5 algorithm as the instruction attribute index. Use the first computing unit to query in the instruction attribute set of the storage unit based on the instruction attribute index of the first instruction. If the query result is a successful query, increment the count value of the associated instruction descriptor, and send the instruction code address in the instruction description associated with the successful query result to the second computing unit. The second computing unit can configure the instruction code address into the code register of the task scheduler, so that the task scheduler of the second computing unit can load the instruction code through the instruction code address to the first arithmetic core to execute the first target instruction of the first instruction stream. This can avoid excessive consumption of video memory resources and communication bandwidth resources caused by loading the instruction code required by the first instruction stream into the video memory.
[0090] If the query result is a failed query, an instruction descriptor for describing the instruction code can be applied for from the instruction attribute set, and a global code segment buffer can be applied for from the video memory unit to store the newly added instruction code corresponding to the first instruction with the failed query, and the count value of the newly added instruction code is set to 1. Use the first computing unit to calculate the instruction attribute index of the newly added instruction code based on the MD5 algorithm. By adding the descriptor of the newly added instruction code to the current instruction attribute tree (or instruction attribute set), and establishing a mapping relationship between the instruction attribute index of the newly added instruction code and the instruction descriptor, it is convenient for instructions in other instruction streams to perform instruction descriptor queries based on the updated instruction attribute set, realizing the sharing of preset instruction codes stored in the code storage area of the video memory unit, and avoiding the second computing unit from re-applying for cache space from the video memory unit to store the instruction codes of other instruction streams.
[0091] During the execution of multiple instruction streams, the count value of the instruction attributes corresponding to each first instruction in the first instruction stream can be reduced after the execution of the first instruction stream, so that other instruction streams can quickly query the instruction attribute set according to the query order of the instruction attribute set based on the instruction codes they need. In addition, the first computing unit can also monitor whether the available capacity space in the code storage area of the video memory unit is greater than a preset video memory space occupancy threshold by running a monitoring process. If it is greater, descriptors with a count value less than the preset count threshold are released, and the release instruction codes corresponding to the descriptors with a count value less than the preset count threshold are released to save the storage space of the video memory unit.
[0092] Figure 5 Schematically shows a flowchart of a task execution method according to another embodiment of the present disclosure.
[0093] As Figure 5 shown, the task execution method in this embodiment includes operation S501 to operation S516.
[0094] In operation S501, the start operation is performed.
[0095] In operation S502, the target instruction attribute is queried from the instruction attribute set by using the first calculation unit. For example, the target instruction attribute corresponding to the instruction can be queried from the instruction attribute set through the instruction attribute indexes of multiple instructions in the instruction stream.
[0096] In operation S503, it is judged whether the query result is a successful query. If the judgment result of operation S503 is yes, operation S504 is performed, the count value of the target instruction attribute in the instruction attribute set is incremented, and the target address is sent to the second calculation unit by using the first calculation unit. If the judgment result of operation S503 is no, operation S506 is performed to judge whether the available capacity space meets the preset capacity condition. If the judgment result of operation S506 is yes, operation S507 is performed, and the instruction attribute set is updated according to the instruction attribute corresponding to the instruction code of the query failure result by using the first calculation unit, and the newly added instruction code is written into the storage unit as the target instruction code. If the judgment result of operation S506 is no, operation S508 is performed to release the instruction code corresponding to the release instruction attribute in the storage unit, and then operations S507 and S508 are performed by using the first calculation unit.
[0097] In operation S509, the target instruction code is called from the target address by using the second calculation unit to execute the instruction stream. In the case that there is still an instruction stream that has not been executed completely in the target processing task, operations S502 to S509 can be cyclically executed by using the first calculation unit and the second calculation unit to execute the target processing task.
[0098] Figure 6 The block diagram of the task execution device according to an embodiment of the present disclosure is schematically shown.
[0099] As Figure 6 shown, the task execution device 600 includes: a storage unit 610, a first calculation unit 620, and a second calculation unit 630.
[0100] The first calculation unit 620 is configured to: determine the target instruction attribute from the instruction attribute set, where the target instruction attribute is associated with the first target instruction in the first instruction stream and the second target instruction in the second instruction stream, the first instruction stream and the second instruction stream are used for the target processing task, and the target instruction attribute includes the target address of the target instruction code in the storage unit 610; send the target address to the second calculation unit 630.
[0101] A second computing unit 630 is configured to: read a target instruction code from a target address of a storage unit 610; and execute a first instruction stream and a second instruction stream to obtain an execution result of a target processing task when the target instruction code is multiplexed based on execution logics of a first target instruction and a second target instruction respectively.
[0102] According to an embodiment of the present disclosure, instruction attributes in an instruction attribute set are associated with count values, and the count value represents the number of times the instruction code address of the instruction attribute is sent to the second computing unit 630. Among them, the query order of multiple instruction attributes in the instruction attribute set by the first computing unit 620 is determined based on the count values respectively related to the multiple instruction attributes.
[0103] According to an embodiment of the present disclosure, the first computing unit 620 is further configured to: when the available capacity of a code storage area of the storage unit 610 does not meet a preset capacity condition, determine a release instruction attribute that does not meet a count value condition from the instruction attribute set based on the count value, where the code storage area is used to store preset instruction codes related to the instruction attribute; and delete the preset instruction codes related to the release instruction attribute from the code storage area.
[0104] According to an embodiment of the present disclosure, a target instruction attribute is obtained by querying the instruction attribute set by the first computing unit 620 based on a query order and an instruction attribute index of a first target instruction, and the instruction attribute index of the first target instruction is determined by processing the instruction code of the first target instruction based on an integrity check algorithm.
[0105] According to an embodiment of the present disclosure, the first computing unit 620 of the task execution device is further configured to: query the instruction attribute set based on an instruction attribute index of a third target instruction in a third instruction stream to obtain a query result, where a data processing task includes the third instruction stream; when the query result indicates a query failure, determine an instruction attribute related to the instruction code of the third target instruction; and update the instruction attribute set according to the instruction attribute related to the instruction code of the third target instruction, and write the instruction code of the third target instruction into the storage unit 610.
[0106] According to an embodiment of the present disclosure, the second computing unit 630 includes a first arithmetic core and a second arithmetic core.
[0107] According to an embodiment of the present disclosure, the second computing unit 630 is configured to perform the following operations to execute the first instruction stream and the second instruction stream when the execution logics of the first target instruction and the second target instruction respectively multiplex the target instruction code: Based on the execution logic of the first instruction stream, use the first arithmetic core to execute the first target instruction based on the target instruction code to obtain a first instruction execution result. Based on the execution logic of the second instruction stream, use the second arithmetic core to execute the second target instruction based on the target instruction code to obtain a second instruction execution result, where the execution result of the target processing task is determined based on the first instruction execution result and the second instruction execution result.
[0108] According to an embodiment of the present disclosure, the first target instruction or the second target instruction includes at least one of the following: point cloud data processing instruction, graph data processing instruction, image feature processing instruction.
[0109] An embodiment of the present disclosure further provides a task execution device, including: a task execution device provided according to an embodiment of the present disclosure.
[0110] Figure 7 A block diagram of a task execution device according to an embodiment of the present disclosure is schematically shown.
[0111] As Figure 7 shown, the device 7000 may include a task execution device 70. The task execution device 70 may be the above-mentioned device 600.
[0112] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0113] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0114] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0115] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0116] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0117] Figure 8FIG. shows a schematic block diagram of an exemplary electronic device that can be used to implement an embodiment of the task execution method of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0118] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0119] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0120] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the task execution method. For example, in some embodiments, the task execution method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the task execution method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the task execution method in any other suitable manner (e.g., by means of firmware).
[0121] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0125] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0126] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0127] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0128] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A task execution method, comprising: Determining a target instruction attribute from the instruction attribute set using a first computing unit, wherein the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream, the first instruction stream and the second instruction stream being used for a target processing task, and the target instruction attribute includes a target address of a target instruction code in a storage unit; using the first computing unit to send the target address to the second computing unit; Reading the target instruction code from the target address of the storage unit using the second computing unit; In a case where the target instruction code is multiplexed based on the respective execution logics of the first target instruction and the second target instruction, the first instruction stream and the second instruction stream are executed by the second computing unit to obtain an execution result of the target processing task.
2. The method according to claim 1, wherein: An instruction attribute in the instruction attribute set is associated with a count value, wherein the count value represents the number of times the instruction code address of the instruction attribute is sent to the second computing unit; The query order of the first calculation unit for the multiple instruction attributes in the instruction attribute set is determined based on the count values associated with each of the multiple instruction attributes.
3. The method according to claim 2, wherein: The method further comprises: In a case where the available capacity of the code storage area of the storage unit does not satisfy a preset capacity condition, determining a release instruction attribute that does not satisfy the count value condition from the instruction attribute set based on the count value by using the first calculation unit, wherein the code storage area is used to store a preset instruction code related to the instruction attribute; and The first calculation unit is used to delete the preset instruction code related to the release instruction attribute from the code storage area.
4. The method according to claim 2, wherein: The target instruction attribute is obtained by querying the instruction attribute set using the first computing unit based on the query order and the instruction attribute index of the first target instruction, and the instruction attribute index of the first target instruction is determined by processing the instruction code of the first target instruction based on an integrity verification algorithm.
5. The method according to claim 1 or 4, wherein: The method further includes: using the first computing unit to perform the following operations: Based on the instruction attribute index of the third target instruction in the third instruction stream, query the instruction attribute set to obtain a query result, and the data processing task includes the third instruction stream; In the case where the query result indicates a query failure, determining an instruction attribute related to the instruction code of the third target instruction; and The instruction attribute set is updated according to the instruction attribute related to the instruction code of the third target instruction, and the instruction code of the third target instruction is written into the storage unit.
6. The method according to claim 1, wherein: The second computing unit includes a first computing core and a second computing core; Wherein, in the case of multiplexing the target instruction code based on the execution logic of each of the first target instruction and the second target instruction, executing the first instruction stream and the second instruction stream by using the second computing unit includes: Based on the execution logic of the first instruction stream, the first computing core is used to execute the first target instruction based on the target instruction code to obtain a first instruction execution result. In the execution logic based on the second instruction stream, the second target instruction is executed based on the target instruction code using the second computing core to obtain a second instruction execution result, wherein the execution result of the target processing task is determined based on the first instruction execution result and the second instruction execution result.
7. The method according to claim 1, wherein: The first target instruction or the second target instruction includes at least one of the following: Point cloud data processing instructions, graph data processing instructions, image feature processing instructions.
8. A task execution device, comprising: Storage unit; The first computing unit is configured as follows: Determining a target instruction attribute from the instruction attribute set, wherein the target instruction attribute is associated with a first target instruction in a first instruction stream and a second target instruction in a second instruction stream, the first instruction stream and the second instruction stream being used for a target processing task, and the target instruction attribute includes a target address of a target instruction code in a storage unit; sending the target address to the second computing unit; and The second computing unit is configured as follows: Read the target instruction code from the target address of the storage unit; In a case where the target instruction code is multiplexed based on the respective execution logics of the first target instruction and the second target instruction, the first instruction stream and the second instruction stream are executed to obtain an execution result of the target processing task.
9. A task execution device, comprising: The device as claimed in claim 8.
10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
Method for preloading descriptors, artificial intelligence chip, computing device, medium and program product
CN121209960A