Data operation method of model and related device

By compiling the calculation graph of the AI ​​model into bytecode instructions and using virtual machines to interpret and execute, the problem of frequent compilation of the AI ​​model during runtime is solved, which improves the operation efficiency and reduces the run time.

CN120104243APending Publication Date: 2025-06-06HUAWEI TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311665038.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When the existing AI computing framework processes AI models with uncertain lengths of input data or intermediate data, it cannot perform kernel function compilation in advance, resulting in kernel function compilation required every run, increasing the running time of the AI ​​model.

Method used

By compiling the calculation graph corresponding to the operations in the AI ​​model into bytecode instructions, and interpreting the running bytecode instructions through a virtual machine pre-configured with corresponding processing functions, the traditional compilation process is avoided and the run time of the AI ​​model is reduced.

Benefits of technology

It effectively avoids the traditional compilation process, improves the operation efficiency of the AI ​​model, and reduces the run time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104243A_ABST
    Figure CN120104243A_ABST
Patent Text Reader

Abstract

The invention relates to a data operation method of a model, which is applied to the operation of an artificial intelligence (AI) model. In the method, based on an input tensor of an operation in an AI model during actual operation, a calculation graph corresponding to the operation in the AI model is compiled into a byte code instruction, and then the byte code instruction is interpreted and operated through a virtual machine which is pre-configured with a corresponding processing function, so that the operation in the AI model is executed. The execution of a traditional tedious compiling process is effectively avoided, and the running time of the AI model is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data calculation method of a model and related devices. Background Art

[0002] In recent years, AI technology represented by deep learning has achieved great development and has achieved good results in computer vision, natural language processing and other fields. In order to improve the development efficiency and computing performance of AI models, the industry often uses AI computing frameworks to express and calculate AI models. AI computing frameworks usually provide hundreds or thousands of different types of operators for users to use. Multiple operators are connected to form a calculation graph, which corresponds to a specific AI model. Different operators use one or more tensors as calculation input parameters when executed, and then call the matching kernel function for corresponding calculations, and finally use one or more tensors as output results.

[0003] In order to improve AI computing performance, the industry's AI computing frameworks generally use operator fusion to improve performance, that is, to merge one or more adjacent operator nodes in the computational graph into a new fusion operator for overall computational execution. Since the number of operator combinations that can be fused in different computational graphs is very large, automatic kernel function compilation technology is currently commonly used to generate kernel functions corresponding to fusion operators. Automatic kernel function compilation refers to automatically generating machine instructions that can be directly executed on the device based on the computational semantics corresponding to the fusion operator and the shapes of the input and output tensors.

[0004] However, for AI models whose lengths of some input data or intermediate data are uncertain (such as Transformer models), since the shape of the input tensor corresponding to the operation in the AI ​​model can only be known when the AI ​​model is specifically executed, it is impossible to perform kernel function compilation for this part of the AI ​​model in advance, resulting in a kernel function compilation stage every time this part of the AI ​​model is run, which leads to a longer running time of the AI ​​model. Summary of the invention

[0005] The present application provides a data operation method for a model. Based on the input tensors of the operations in the AI ​​model during actual runtime, the calculation graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the execution of the traditional cumbersome compilation process, and reducing the running time of the AI ​​model.

[0006] The first aspect of the present application provides a data operation method of a model, which is applied to the operation of an AI model. In this method, a calculation graph and the shape of an input tensor are first obtained. The calculation graph is used to indicate the operation in the AI ​​model, and the input tensor is used to represent the input data corresponding to the operation. The shape of the input tensor is the size of the input tensor. That is, the calculation graph actually indicates how to perform operation operations on tensors, and the input tensor is the input data corresponding to the operation indicated in the calculation graph. Among them, the input tensor is multidimensional data, and the shape of the input tensor is not fixed, and can be determined according to the actual operation of the AI ​​model. In addition, the input tensor may, for example, include one or more tensors, depending on the number of input data indicated by the calculation graph.

[0007] Then, based on the computation graph and the shape of the input tensor, at least one bytecode instruction is generated, and the at least one bytecode instruction is used to indicate the operation performed on the input tensor. The at least one bytecode instruction indicates the operation performed on the above-mentioned input tensor in the form of a binary instruction. In addition, the at least one bytecode instruction is actually a bytecode, that is, an intermediate code, which cannot be directly recognized and executed by hardware, but needs to be interpreted and executed by a software module.

[0008] Finally, at least one bytecode instruction is executed by a virtual machine, and a processing function corresponding to at least one bytecode instruction is configured in the virtual machine. The processing function is used to interpret at least one bytecode instruction and call a machine instruction that can be recognized by the hardware based on the interpretation result (that is, a machine instruction corresponding to the at least one bytecode instruction) to perform an operation on the input tensor. Among them, the processing function interprets the bytecode instruction means that the processing function analyzes the content of the bytecode instruction, and calls the machine instruction based on the content of the bytecode instruction to implement the operation on the tensor data. Specifically, a plurality of processing functions are configured in the virtual machine, and different processing functions are used to process different types of bytecode instructions to ensure that the bytecode instructions indicating different operation operations are processed by corresponding processing functions. Specifically, the virtual machine in this scheme can also be called a kernel function, which is the execution entity of the operator at runtime, and can interpret and execute the bytecode instruction to realize the operation of the operator. Among them, the specific implementation of the virtual machine can be a binary instruction code that can be directly executed on the device hardware.

[0009] In this solution, the computational graph corresponding to the operation in the AI ​​model is compiled into bytecode instructions based on the input tensors of the operations in the AI ​​model during actual runtime, and then the bytecode instructions are interpreted and run by a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the traditional cumbersome compilation process and reducing the running time of the AI ​​model.

[0010] In addition, compared with the model compilation process performed on the AI ​​model in the existing related technologies, this solution only needs to generate simple bytecode instructions based on the calculation graph and the actual input tensor, without the need to perform a complete compilation process (i.e. preprocessing, syntax parsing, instruction generation, assembly, linking, and output files), and the generated bytecode instructions can be interpreted and executed by the processing functions in the pre-implemented virtual machine, ensuring that the bytecode instructions are generated and executed at a faster speed, effectively improving the operating efficiency of the AI ​​model.

[0011] In a possible implementation, at least one bytecode instruction is also used to instruct to divide the input tensor into multiple parts to perform the operation separately. In this way, when interpreting and executing at least one bytecode instruction through a virtual machine, at least one bytecode instruction can be processed in parallel by multiple virtual machine instances located in different processor cores, and different virtual machine instances in the multiple virtual machine instances are used to process different data in the input tensor. That is, each virtual machine instance in the multiple virtual machine instances is responsible for processing a part of the data in the input tensor, thereby realizing that multiple virtual machine instances process the input tensor in parallel.

[0012] In this solution, by instructing the splitting of tensors in the bytecode instructions, the tensors can be split and allocated to multiple processor cores for parallel processing when interpreting and executing the bytecode instructions, thereby realizing parallel processing of tensor operations and improving the efficiency of AI operations.

[0013] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0014] In a possible implementation, the number of splits is greater than or equal to the number of multiple virtual machine instances. Generally speaking, one virtual machine instance runs on one processor core, so the number of splits is actually greater than or equal to the number of processor cores used to execute bytecode instructions.

[0015] In this solution, by indicating the number of times the tensor is split in the bytecode instruction, the input tensor can be quickly split into multiple parts based on the tensor splitting situation when the bytecode instruction is interpreted and executed, and the parts are given to the corresponding processor cores for processing. There is no need for the virtual machine to additionally determine how to split the tensor, thereby improving the efficiency of AI operations.

[0016] In one possible implementation, the method further includes: generating a first bytecode instruction and a second bytecode instruction based on a computation graph and an input tensor, wherein the first bytecode instruction is used to instruct the input tensor to be moved from a global memory to a local memory, and the second bytecode instruction is used to instruct the output tensor obtained by processing the input tensor to be moved from a local memory to a global memory.

[0017] In the execution phase of the bytecode instructions, specifically, the first bytecode instruction, the at least one bytecode instruction and the second bytecode instruction may be executed in sequence by the virtual machine.

[0018] It should be noted that this solution introduces the generation of a first bytecode instruction for moving an input tensor from global memory to local memory based on a computation graph and a second bytecode instruction for moving an output tensor obtained by processing the input tensor from local memory to global memory. In some special cases, such as when the input tensor is a random tensor, it is not necessary to generate a bytecode instruction for moving the input tensor from global memory to local memory, but to directly generate a random tensor in local memory.

[0019] In this solution, by additionally generating corresponding data movement instructions according to the execution of bytecode instructions on hardware during the bytecode instruction generation stage, it is possible to move tensors between global memory and local memory, ensuring that different processor cores on the hardware can smoothly perform operations on tensors, and ensuring that multiple processor cores can perform tensor operations in parallel, thereby improving the feasibility of the solution.

[0020] In one possible implementation, after generating at least one bytecode instruction, the method further includes: moving at least one bytecode instruction to a memory space accessed by AI hardware, where the AI ​​hardware is used to run a virtual machine. Exemplarily, the AI ​​hardware is hardware specifically used to implement AI computing, such as a graphics processing unit (GPU), a neural network processor (NPU), or a tensor processing unit (TPU), which can accelerate AI computing.

[0021] In one possible implementation, at least one bytecode instruction is generated based on the computation graph and the shape of the input tensor, including: obtaining a first meta-operator graph based on the computation graph conversion, wherein the first meta-operator graph includes multiple meta-operators. Among them, the computation graph is used to indicate part or all of the computing operations in the AI ​​model, and multiple meta-operators are used to indicate basic computing operations. That is, the meta-operator is the most basic unit for performing computing operations, and the meta-operator can no longer be obtained by combining other more basic operators. Then, based on the shape of the input tensor and the first meta-operator graph, at least one bytecode instruction is generated, and multiple meta-operators correspond to the at least one bytecode instruction mentioned above. In addition, a meta-operator may correspond to one or more bytecode instructions.

[0022] In this scheme, since bytecode instructions are processed by processing functions configured in the virtual machine, and all operators can be obtained by combining meta-operators, generating bytecode instructions according to the meta-operator granularity can reduce the types of generated bytecode instructions as much as possible, thereby reducing the number of pre-configured processing functions in the virtual machine and reducing the implementation complexity of the virtual machine.

[0023] In a possible implementation, the first meta-operator graph is obtained based on the conversion of the computation graph, specifically including: converting each operator in the computation graph into one or more meta-operators to obtain a converted computation graph. That is, some operators in the computation graph may be composite operators composed of multiple meta-operators, so all operators in the computation graph can be represented by meta-operators, thereby obtaining a converted computation graph composed of meta-operators. Then, the converted computation graph is divided into a plurality of continuous meta-operator graphs, the plurality of meta-operator graphs include the first meta-operator graph, and the plurality of meta-operator graphs each include a plurality of meta-operators. That is, for the converted computation graph, the converted computation graph can be divided into a plurality of parts according to the execution order of the meta-operators, and each part includes a plurality of adjacent meta-operators. In this way, by fusing the plurality of adjacent meta-operators in each part, a meta-operator graph can be obtained, thereby realizing the division of the converted computation graph into a plurality of meta-operator graphs, and each meta-operator graph is obtained by fusing a plurality of meta-operators.

[0024] In this solution, by splitting the converted computation graph into multiple meta-operator graphs for processing, it can ensure that the virtual machine executes a single meta-operator graph each time, avoiding the situation where the device running the virtual machine is insufficient in memory due to too many meta-operators being processed continuously. In addition, by fusing multiple meta-operators into one meta-operator graph, the intermediate data obtained by processing multiple meta-operators in the same meta-operator graph can be stored in the local memory of the device with faster reading and writing speeds, without the need to frequently read and write data from the global internal memory of the device with slower reading and writing speeds, thereby improving the processing efficiency of the operator.

[0025] In one possible implementation, at least one bytecode instruction includes an instruction identifier, which is used to uniquely identify the type of the at least one bytecode instruction. The virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction. For example, a bytecode instruction indicating that the calculation operation to be performed is an addition operation can be represented by an instruction identifier 00; a bytecode instruction indicating that the calculation operation to be performed is a subtraction operation can be represented by an instruction identifier 01. In this solution, by setting the instruction identifier in the bytecode instruction, the type of the bytecode instruction can be uniquely identified, thereby facilitating the virtual machine to quickly call the corresponding processing function to process the bytecode instruction according to the instruction identifier, thereby improving the interpretation and execution efficiency of the bytecode instruction.

[0026] In one possible implementation, at least one bytecode instruction includes a data type identifier, which is used to indicate the data type of the input tensor. For example, the data type identifier is fp32, which means that the data type of the input tensor is a 32-bit floating point number, so the operation performed on the input tensor is actually a 32-bit floating point calculation. Exemplarily, in bytecode instructions, different data type identifiers can be used to indicate different data types, such as 32-bit floating point numbers, 16-bit floating point numbers, 32-bit integers, and other data types, and the data types are not specifically limited here.

[0027] In a possible implementation, at least one bytecode instruction is further used to indicate a storage address of an input tensor and a storage address of an output tensor. Moreover, the storage address of the input tensor and the storage address of the output tensor are both addresses in a local memory.

[0028] It should be noted that when the input tensor is a random variable, since the input tensor can actually be randomly generated when used, the storage address of the input tensor may not be indicated in the bytecode instruction, but only the storage address of the output tensor.

[0029] In one possible implementation, obtaining a computational graph specifically includes: obtaining an operator call instruction, the operator call instruction is used to indicate the execution of an operation corresponding to a target operator, and the operator call instruction includes an input tensor; based on the target operator indicated in the operator call instruction, a computational graph is generated. The computational graph is generated based on the target operator, and is used to indicate the execution of the operation indicated by the target operator on the input tensor by means of operator nodes and directed edges. In addition, the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic operations. That is, the computational graph itself is composed of multiple meta-operators, and no longer includes a composite operator composed of multiple meta-operators.

[0030] The second aspect of the present application provides a data operation device of a model, including: an acquisition module, used to obtain a calculation graph and the shape of an input tensor, the calculation graph is used to indicate the calculation operation in the artificial intelligence AI model, and the input tensor is used to represent the input data corresponding to the calculation operation; a processing module, used to generate at least one bytecode instruction based on the calculation graph and the shape of the input tensor, and the at least one bytecode instruction is used to indicate the calculation operation performed on the input tensor; the processing module is also used to execute at least one bytecode instruction through a virtual machine, and the virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to interpret the at least one bytecode instruction and call the machine instruction based on the interpretation result to perform the calculation operation on the input tensor.

[0031] In one possible implementation, at least one bytecode instruction is also used to instruct that an input tensor be divided into multiple parts to perform calculation operations separately; the processing module is specifically used to process at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

[0032] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0033] In a possible implementation, the number of splits is greater than or equal to the number of the multiple virtual machine instances.

[0034] In one possible implementation, the processing module is further used to: generate a first bytecode instruction and a second bytecode instruction based on the computation graph and the input tensor, the first bytecode instruction being used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction being used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory; and execute the first bytecode instruction, at least one bytecode instruction, and the second bytecode instruction in sequence through the virtual machine.

[0035] In one possible implementation, after generating at least one bytecode instruction, the processing module is further used to: move the at least one bytecode instruction to a memory space accessed by AI hardware, and the AI ​​hardware is used to run a virtual machine.

[0036] In one possible implementation, the processing module is specifically used to: obtain a first meta-operator graph based on a computational graph conversion, the first meta-operator graph including multiple meta-operators, wherein the computational graph is used to indicate part or all of the computing operations in the AI ​​model, and multiple meta-operators are used to indicate basic computing operations; generate at least one bytecode instruction based on the shape of the input tensor and the first meta-operator graph, and multiple meta-operators correspond to at least one bytecode instruction.

[0037] In one possible implementation, the processing module is further used to: convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; divide the converted computation graph into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include a first meta-operator graph, and each of the plurality of meta-operator graphs includes a plurality of meta-operators.

[0038] In one possible implementation, the acquisition module is also used to acquire an operator call instruction, where the operator call instruction is used to instruct the execution of an operation corresponding to a target operator, and the operator call instruction includes an input tensor; the processing module is also used to generate a computational graph based on the target operator indicated in the operator call instruction, where the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic operations.

[0039] In a possible implementation, at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

[0040] In one possible implementation, at least one bytecode instruction includes a data type identifier, which is used to indicate the data type of the input tensor.

[0041] In a possible implementation, at least one bytecode instruction is further used to indicate a storage address of an input tensor and a storage address of an output tensor, where the output tensor is a tensor obtained after performing a calculation operation on the input tensor. Furthermore, the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

[0042] The third aspect of the present application provides a data operation device of a model, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For the steps in each possible implementation of the first aspect executed by the processor, the first aspect can be specifically referred to, and no further description is given here.

[0043] The fourth aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method of any implementation manner of the first aspect.

[0044] A fifth aspect of the present application provides a circuit system, the circuit system includes a processing circuit, and the processing circuit is configured to execute a method in any implementation manner of the above-mentioned first aspect.

[0045] The sixth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute a method implemented in any one of the first aspects.

[0046] The seventh aspect of the present application provides a chip system, which includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in any implementation of the first aspect, for example, sending or processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of chips, or it can include chips and other discrete devices.

[0047] The beneficial effects of the second to seventh aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0049] Figure 2 A schematic diagram of a system architecture of an electronic device is provided for an embodiment of the present application;

[0050] Figure 3 A schematic diagram of a flow chart of a data calculation method of a model provided in an embodiment of the present application;

[0051] Figure 4 A schematic diagram of a flow chart of a data calculation method of another model provided in an embodiment of the present application;

[0052] Figure 5 A schematic diagram of a system architecture provided for an embodiment of the present application;

[0053] Figure 6 A schematic diagram of a flow chart of bytecode instruction generation provided in an embodiment of the present application;

[0054] Figure 7 A system architecture diagram for an actual application scenario provided by an embodiment of the present application;

[0055] Fig. 8A A schematic diagram of the execution flow of an AI computing framework provided in an embodiment of the present application;

[0056] Figure 8B A schematic diagram of a flow chart of a virtual machine interpreting and executing bytecode instructions provided in an embodiment of the present application;

[0057] Fig. 9 A system architecture diagram for another practical application scenario provided by an embodiment of the present application;

[0058] Fig.10 A schematic diagram of the structure of a data computing device of a model provided in an embodiment of the present application;

[0059] Fig.11 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0060] Fig.12 A schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0061] Fig.13 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those of ordinary skill in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0063] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchanged where appropriate, so that the embodiments can be implemented in a sequence other than that illustrated or described in the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of the steps that appear in the present application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. There may be other division methods when it is implemented in actual applications. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. In addition, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed in multiple circuit units, and some or all of the units may be selected according to actual needs to achieve the purpose of the present application.

[0064] To facilitate understanding, some technical terms involved in the embodiments of the present application are first introduced below.

[0065] (1) Neural Network

[0066] A neural network can be composed of neural units. Specifically, it can be understood as a neural network with an input layer, a hidden layer, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (DNN). The work of each layer in the neural network can be expressed mathematically. To describe, from a physical level, the work of each layer in the neural network can be understood as completing the transformation from input space to output space (i.e., from the row space to the column space of the matrix) through five operations on the input space (the set of input vectors). These five operations include: 1. Dimension increase / reduction; 2. Enlargement / reduction; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are represented by Completed, operation 4 is completed by "+b", and operation 5 is implemented by "a()". The word "space" is used here because the classified object is not a single thing, but a class of things. Space refers to the collection of all individuals of this class of things, where W is the weight matrix of each layer of the neural network, and each value in the matrix represents the weight value of a neuron in the layer. The matrix W determines the spatial transformation from the input space to the output space described above, that is, the W of each layer of the neural network controls how to transform the space. The purpose of training a neural network is to eventually obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of a neural network is essentially about learning how to control spatial transformation, or more specifically, learning the weight matrix.

[0067] (2) Self-Attention Network

[0068] The self-attention network is a neural network that uses the self-attention mechanism. Typical self-attention networks include the Transformer model. The self-attention mechanism is actually an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. Self-attention networks are often used in fields such as machine reading, abstract summarization, or image description.

[0069] (3) AI Model

[0070] AI models refer to mathematical models that learn and predict data with certain regularity and predictability. At present, AI models are generally composed of neural networks. During the operation of AI models, the computational process of learning data is called training, and the prediction of results of input data is called inference.

[0071] (4) AI computing framework

[0072] AI computing frameworks are software platforms used to express and process AI models, such as TensorFlow, PyTorch, MindSpore, and other platforms.

[0073] (5) Operator

[0074] Operators are the basic computing units of AI models, and each operator represents specific computing semantics. Common operators represent computing semantics such as convolution operations, pooling operations, and activation function operations.

[0075] (6) Meta-operator

[0076] A meta-operator is an operator used to indicate the most basic operation, that is, a meta-operator cannot be expressed by combining other more basic operators. For example, some common meta-operators are used to indicate basic operations such as addition, subtraction, multiplication, and division.

[0077] (7) Composite Operator

[0078] A composite operator refers to an operator that can be expressed by combining multiple meta-operators, such as the convolution operator, the pooling operator, and other operators.

[0079] (8) Kernel function

[0080] The kernel function is the execution entity of the operator at runtime, that is, the kernel function is actually a binary instruction code that can be directly executed on the device hardware.

[0081] (9) Tensor

[0082] A tensor is a multidimensional data type that is usually used to represent input data or output data when an operator is running.

[0083] (10)shape

[0084] The shape is a representation of the dimensions of the tensor data. For example, [3,4] represents a 3*4 two-dimensional tensor.

[0085] (11) Computational Graph

[0086] A computational graph is a directed acyclic graph consisting of operators as nodes and tensors as edges. Different AI models can usually be abstracted into corresponding computational graph structures for compilation and execution in the AI ​​model running platform.

[0087] (12) Bytecode

[0088] Bytecode is a binary file that contains an execution program and consists of a sequence of op code / data pairs. Compared with the machine instruction code that can be directly executed by the hardware, bytecode is actually an intermediate code, that is, an instruction code that needs to be interpreted and executed by the software code, and cannot be directly executed on the hardware.

[0089] (13) Virtual Machine

[0090] In this embodiment, the virtual machine is an important tool in the programming language, which is essentially a software module that can convert high-level language code (such as bytecode) into low-level machine instructions, so that the code can run on different operating systems and hardware platforms. For example, the Java virtual machine for the Java language can compile Java code into bytecode, and then convert the bytecode into machine instructions for execution through an interpreter or a just-in-time compiler.

[0091] In addition, a virtual machine instance refers to a virtual machine running on a hardware device (such as a processor), that is, a virtual machine instance refers to a running virtual machine. Based on the same virtual machine software code, multiple virtual machine instances with the same configuration can be quickly created.

[0092] (14) Machine Instructions

[0093] Machine instructions are instructions that can be directly recognized and executed by computer hardware (such as the Central Processing Unit (CPU)). They are expressed in binary code. Machine instructions usually consist of two parts: an operation code and an operand. The operation code indicates the operation to be completed by the machine instruction, that is, the function of the machine instruction; the operand indicates the object involved in the operation and the location where the operation result is stored.

[0094] (15) Global Memory

[0095] Global memory refers to the memory space in device hardware (such as GPU) used to store global variables and static variables. Variables stored in global memory can be accessed and modified by all objects running in the device hardware. In other words, global memory is shared by multiple cores and can be accessed by all processor cores in the device hardware and can also be accessed by the host hardware.

[0096] (16) Local Memory

[0097] Local memory refers to the private memory space allocated to each processor core in the device hardware, which can only be accessed by the corresponding processor core and cannot be accessed by other processor cores.

[0098] At present, traditional AI computing frameworks often use static computational graphs to run AI models. That is, the shape of the input data of the AI ​​model is constant, for example, the input data of the AI ​​model is an image of fixed size. In this way, the traditional AI computing framework can generate the corresponding static computational graph based on the structure of the AI ​​model and the shape of the input data, and compile the kernel function corresponding to the static computational graph in advance, so that the kernel function compiled in advance can be used to process the actual input data of the AI ​​model during the model execution phase.

[0099] However, with the development of AI technology, there are more and more types of AI models, and AI models with uncertain input data or intermediate data length continue to emerge, such as the Transformer model for processing natural language sequences. For these AI models with uncertain input data or intermediate data length, the calculation graphs corresponding to these AI models are actually dynamic calculation graphs, that is, the shape of the input data of the calculation graph is not fixed, but can change. Therefore, for these AI models corresponding to dynamic calculation graphs, since the shape of the input tensor corresponding to the AI ​​model can only be known when the AI ​​model is specifically executed, the traditional AI computing framework cannot perform the kernel function compilation stage for this part of the AI ​​model in advance, resulting in each time this part of the AI ​​model is run, it is necessary to perform the kernel function compilation stage according to the actual input data, resulting in a longer running time of the AI ​​model.

[0100] Based on this, a data operation method of a model is provided in this embodiment. Based on the input tensor of the operation in the AI ​​model during actual operation, the calculation graph corresponding to the operation in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the execution of the traditional cumbersome compilation process and reducing the running time of the AI ​​model.

[0101] To facilitate understanding, the following first introduces the system architecture used in the data calculation method of the model provided in the embodiment of the present application.

[0102] See also Figure 1 , the present application embodiment provides a schematic diagram of a system architecture 100. Figure 1 As shown, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices. The execution device 110 can be arranged at a physical site, or distributed at multiple physical sites. The execution device 110 can use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the data operation method of the model provided in the embodiment of the present application, thereby realizing the operation of the AI ​​model.

[0103] Users can operate their respective user devices (e.g., local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car or other type of cellular phone, a media consumption device, a wearable device, a set-top box, a game console, etc.

[0104] The local device of each user can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0105] In one implementation, the execution device 110 is used to implement the data operation method of the model provided in the embodiment of the present application, and send the obtained operation results to the local device 101 and the local device 102 through the communication network, so that the local device 101 and the local device 102 can obtain the operation results of the AI ​​model, such as image classification results or translation results output by the AI ​​model.

[0106] In another implementation, one or more aspects of the execution device 110 may be implemented by each local device. For example, the local device 101 may provide local data or feedback calculation results to the execution device 110, or execute the data operation method of the model provided in the embodiment of the present application. That is, the execution device 110 may send the AI ​​model or the calculation graph corresponding to the AI ​​model to the local device 101 through the communication network, and the local device 101 executes the data operation method of the model provided in the embodiment of the present application.

[0107] It should be noted that all functions of the execution device 110 may also be implemented by the local device. For example, the local device 101 implements the functions of the execution device 110 and provides services to its own user, or provides services to the user of the local device 102.

[0108] In general, the training method of the model provided in the embodiment of the present application can be applied to an electronic device, such as the execution device 110, the local device 101 or the local device 102 mentioned above. Exemplarily, the electronic device can be, for example, a server, a wireless electronic device in industrial control, a smart phone (mobile phone), a personal computer (personal computer, PC), a laptop, a tablet computer, an autonomous driving car, a smart camera and other devices. For ease of understanding, the method provided in the embodiment of the present application will be applied to a server as an example to introduce the method.

[0109] See also Figure 2 , Figure 2A schematic diagram of a system architecture of an electronic device is provided for an embodiment of the present application. Figure 2 The system architecture shown includes host hardware and device hardware. The host hardware includes a processor and memory, which are used to execute the relevant functions of the host software modules such as the AI ​​model and the AI ​​computing framework, and cache intermediate data. The device hardware includes a processor and memory, which are used to execute the code of the device software, thereby realizing the functions of the virtual machine and caching intermediate data. Among them, the host hardware and the device hardware can be deployed on the same electronic device, for example, the host hardware and the device hardware are deployed on the same server. Among them, the host hardware includes the CPU and memory on the server, and the device hardware includes, for example, a graphics processing unit (GPU), a neural network processor (NPU) or a tensor processor (TPU) on the server.

[0110] The AI ​​computing framework running on the host hardware can obtain the AI ​​model and build the corresponding calculation graph based on the AI ​​model, and then perform bytecode compilation based on the calculation graph and the actual input tensor of the AI ​​model to obtain bytecode instructions. Then, the host hardware sends the compiled bytecode instructions to the virtual machine on the device hardware, and the virtual machine executes the bytecode instructions to execute the calculation operations indicated in the AI ​​model and complete the operation of the AI ​​model.

[0111] It should be noted that in some scenarios (for example, the above-mentioned device hardware is not available or there is no idle device hardware), the host hardware can also assume the function of the device hardware, that is, a virtual machine is run on the host hardware to interpret and execute bytecode instructions.

[0112] The above introduces the system architecture applied by the method provided in this embodiment. The following will describe in detail the specific execution process of the method provided in the embodiment of the present application in conjunction with the accompanying drawings.

[0113] See also Figure 3 , Figure 3 A flow chart of a data calculation method of a model provided in an embodiment of the present application. Figure 3 As shown, the data operation method of the model includes the following steps 301-303.

[0114] Step 301, obtain the shape of the computation graph and the input tensor, the computation graph is used to indicate the calculation operations in the AI ​​model, and the input tensor is used to represent the input data corresponding to the calculation operations.

[0115] In this embodiment, the computation graph is a directed acyclic graph consisting of operators as nodes and tensors as edges, which is used to indicate how to perform computational operations on tensors. In addition, the computation graph is obtained based on the AI ​​model and is used to indicate computational operations in the AI ​​model, such as convolution operations, pooling operations, and the like in the AI ​​model.

[0116] Optionally, the computational graph may be used to indicate computing operations in the entire AI model, or may be used to indicate partial computing operations in the AI ​​model (for example, indicating the operations of a certain neural network layer or a certain operator in the AI ​​model). This embodiment does not specifically limit this.

[0117] Since the calculation graph actually indicates how to perform calculation operations on tensors, this embodiment also obtains the shape of the input tensor corresponding to the calculation graph. That is, the input tensor is the input data corresponding to the calculation operation indicated in the calculation graph. Exemplarily, the input tensor is, for example, the actual input data of the entire AI model; for example, when the AI ​​model is a natural language processing model, the input tensor is, for example, the tensor corresponding to the text that needs to be input into the natural language processing model. Alternatively, the input tensor is, for example, the intermediate data generated during the operation of the AI ​​model, such as the tensor data output by a neural network layer in the AI ​​model. Among them, the input tensor is multidimensional data, and the shape of the input tensor is not fixed, and can be determined according to the actual operation of the AI ​​model. In addition, the input tensor may include, for example, one or more tensors, depending on the number of input data indicated by the calculation graph, and this embodiment does not specifically limit this.

[0118] It should be noted that the shape of the input data indicated in the calculation graph obtained in this embodiment is unknown. Only after the shape of the above-mentioned input tensor is obtained can the shape of the input data processed based on the calculation graph be determined. In addition, the shape of the input tensor can be determined after the actual input tensor is obtained, or it can be obtained in advance during the generation process of the actual input tensor.

[0119] In this embodiment, there are multiple ways to obtain the calculation graph.

[0120] In one possible implementation, the computation graph may be generated based on the AI ​​model. That is, in the process of running the AI ​​model, the computation graph corresponding to the AI ​​model is first generated, and then the operation of the AI ​​model is realized by executing the computation graph.

[0121] For example, after obtaining the AI ​​model that needs to be run, all or part of the operations in the AI ​​model can be converted into a computational graph based on the computational operations of each neural network layer in the AI ​​model, thereby describing the structure of the AI ​​model in the form of a graph.

[0122] In another possible implementation, during the running of the AI ​​model, for the tensor calculations encountered during the running of the AI ​​model, the operator call interface of the AI ​​computing framework is called to trigger the calculation of a certain operator, and then the calculation graph corresponding to the operator is generated. That is, during the running of the AI ​​model, the calculation graph corresponding to the entire AI model is no longer generated, but the operators in the AI ​​model are processed by calling the operator interface and generating the corresponding calculation graph.

[0123] Exemplarily, during the operation of the AI ​​model, an operator call instruction can be obtained, which is used to instruct the execution of the operation corresponding to the target operator, and the operator call instruction includes the above-mentioned input tensor. Specifically, the target operator is, for example, one or more operators indicated in the AI ​​model, such as a convolution operator or a pooling operator. The input tensor included in the operator call instruction is the input data of the target operator.

[0124] Then, based on the target operator indicated in the operator call instruction, a computation graph is generated. The computation graph is generated based on the target operator and is used to indicate the operation indicated by the target operator on the input tensor by means of operator nodes and directed edges. In addition, the computation graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are used to indicate basic operation. That is, the computation graph itself is composed of multiple meta-operators, and no longer includes a composite operator composed of multiple meta-operators.

[0125] Step 302: Generate at least one bytecode instruction based on the computation graph and the shape of the input tensor, where the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor.

[0126] In this embodiment, after obtaining the shape of the computation graph and the input tensor, the type of operation to be performed and the shape of the specific input data corresponding to the operation can be determined. Therefore, based on the computation graph and the shape of the input tensor, at least one bytecode instruction can be generated, and the at least one bytecode instruction indicates the operation performed on the above-mentioned input tensor in the form of a binary instruction. In addition, at least one bytecode instruction is actually a bytecode that needs to be interpreted and executed by a software module, and cannot be directly executed by hardware.

[0127] In addition, since the shape of the input tensor is known, how to store the input tensor and the output tensor can be determined, that is, at least one bytecode instruction can also be used to indicate the storage address of the input tensor and the storage address of the output tensor, where the output tensor is the tensor obtained after performing an operation on the input tensor. In this way, when executing at least one bytecode instruction, the input tensor that needs to perform the operation, the type of operation performed on the input tensor, and the storage address of the operation result obtained after performing the operation on the input tensor can be obtained based on at least one bytecode instruction, ensuring that the corresponding operation in the running AI model can be realized by executing at least one bytecode instruction.

[0128] Optionally, in the process of generating bytecode instructions based on the computation graph and the input tensor, the bytecode instructions may be generated according to the granularity of the meta-operator, that is, one meta-operator corresponds to one or more bytecode instructions.

[0129] Exemplarily, in the case where the computation graph is constructed based on all or part of the operations in the AI ​​model, a first meta-operator graph can be obtained based on the computation graph conversion, and the first meta-operator graph includes multiple meta-operators, wherein the multiple meta-operators are used to indicate basic computational operations. That is, the meta-operator is the most basic unit for performing computational operations, and the meta-operator can no longer be obtained by combining other more basic operators. In other words, the operator indicated in the computation graph may be one or more composite operators, which are composed of the most basic meta-operators. The process of converting a computation graph into a meta-operator graph is actually to split the composite operator in the computation graph into multiple meta-operators for representation, thereby realizing the computational logic of the computation graph represented by the most basic meta-operator. For example, assuming that the operator indicated in the calculation graph is SqrtGrad(x,y), the operator can be split into two element operators, namely SqrtGrad(x,y)=Div(Square(x),y); wherein SqrtGrad() represents the root variance, Div() represents the integer division operation, and Square() represents the square operation.

[0130] Then, based on the shape of the input tensor and the first meta-operator graph, multiple bytecode instructions are generated, and the multiple bytecode instructions include at least one bytecode instruction mentioned above. The multiple bytecode instructions correspond to the multiple meta-operators indicated by the meta-operator graph, that is, each bytecode instruction corresponds to a meta-operator. Generally speaking, a meta-operator may correspond to one bytecode instruction; when the shape of the tensor operated by some meta-operators is too large, one bytecode instruction may be difficult to complete the representation of one meta-operator, so one meta-operator may correspond to multiple bytecode instructions.

[0131] Since bytecode instructions are processed by processing functions configured in the virtual machine, and all operators can be obtained by combining meta-operators, generating bytecode instructions according to the meta-operator granularity can minimize the types of generated bytecode instructions, thereby reducing the number of pre-configured processing functions in the virtual machine and reducing the implementation complexity of the virtual machine.

[0132] Optionally, in the process of obtaining the first meta-operator graph based on the conversion of the computation graph, each operator in the computation graph can be converted into one or more meta-operators to obtain a converted computation graph. That is, some operators in the computation graph may be composite operators composed of multiple meta-operators, so all operators in the computation graph can be represented by meta-operators to obtain a converted computation graph composed of meta-operators. Then, the converted computation graph is divided into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include the first meta-operator graph, and the plurality of meta-operator graphs each include a plurality of meta-operators. That is, for the converted computation graph, the plurality of adjacent meta-operators can be fused to obtain a meta-operator graph. In this way, by dividing the meta-operators in the converted computation graph into a plurality of parts, the plurality of meta-operators in each part can be fused to form a meta-operator graph.

[0133] In this solution, by splitting the converted computation graph into multiple meta-operator graphs for processing, it can ensure that the virtual machine executes a single meta-operator graph each time, avoiding the situation where the device running the virtual machine is short of memory due to too many meta-operators being processed continuously.

[0134] Step 303, execute at least one bytecode instruction through a virtual machine to obtain an output tensor, the virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

[0135] In this embodiment, the virtual machine is a pre-implemented software module for implementing the interpretation and execution of bytecode instructions. Specifically, the virtual machine is configured with multiple processing functions, and different processing functions are used to process different types of bytecode instructions to ensure that bytecode instructions indicating different operation operations are processed by corresponding processing functions. Therefore, after the virtual machine obtains at least one bytecode instruction, the virtual machine can call the corresponding processing function to process the at least one bytecode instruction according to the type of the at least one bytecode instruction.

[0136] Specifically, the virtual machine in this solution can also be called a kernel function, which is the execution entity of the operator at runtime and can interpret and execute bytecode instructions to realize the operation of the operator. The specific implementation of the virtual machine can be a binary instruction code that can be directly executed on the device hardware (such as GPU, NPU or TPU).

[0137] Exemplarily, at least one bytecode instruction includes an instruction identifier, which is used to uniquely identify the type of the at least one bytecode instruction. That is, different types of bytecode instructions are marked by different instruction identifiers. For example, a bytecode instruction indicating that the operation to be performed is an addition operation can be represented by instruction identifier 00; a bytecode instruction indicating that the operation to be performed is a subtraction operation can be represented by instruction identifier 01; a bytecode instruction indicating that the operation to be performed is a multiplication operation can be represented by instruction identifier 10; and a bytecode instruction indicating that the operation to be performed is a division operation can be represented by instruction identifier 11. In this way, the virtual machine can call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier in the at least one bytecode instruction to process the at least one bytecode instruction.

[0138] That is to say, by setting the instruction identifier in the bytecode instruction, the type of the bytecode instruction can be uniquely identified, thereby facilitating the virtual machine to quickly call the corresponding processing function to process the bytecode instruction according to the instruction identifier, thereby improving the efficiency of interpreting and executing the bytecode instruction.

[0139] Exemplarily, the types of bytecode instructions may include the following types: move class and calculation class. Among them, the move class may specifically include the Load type, which means moving tensor data from global memory to local memory; the Store type means writing tensor data from local memory back to global memory. The calculation class may include algebraic calculations (such as Add, Sub, Mul, Div, Sqrt, Abs, Exp or Pow), reduction calculations (such as Sum (sum), Max (maximum value of tensor data) or Min (maximum value of tensor data)), and comparison calculations (such as Greater, Less or Equal).

[0140] The processing function configured in the virtual machine is also a pre-implemented software code. In the process of processing the bytecode instruction, the processing function reads the tensor that needs to be operated from the bytecode instruction, and based on the operation indicated by the bytecode instruction, calls the corresponding machine instruction to perform the operation on the read tensor, and finally stores the operation result obtained by performing the operation to the storage address specified by the bytecode instruction.

[0141] In general, based on the input tensors of the operations in the AI ​​model during actual runtime, the computational graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with the corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the traditional cumbersome compilation process and reducing the running time of the AI ​​model.

[0142] Specifically, compared with the model compilation process performed on the AI ​​model in the existing related technologies, the present embodiment only needs to generate simple bytecode instructions based on the calculation graph and the actual input tensor, without the need to perform a complete compilation process (i.e., preprocessing, syntax parsing, instruction generation, assembly, linking, and output files), and the generated bytecode instructions can be interpreted and executed by the processing functions in the pre-implemented virtual machine, ensuring that the bytecode instructions are generated and executed at a faster speed, effectively improving the operating efficiency of the AI ​​model.

[0143] In addition, the model compilation process performed on the AI ​​model in the related art ultimately generates machine instructions that can be directly executed by the hardware, and the compilation granularity is relatively small. For example, for the common addition operation between two tensors, it is necessary to generate multiple addition instructions for indicating the addition of two integers. Therefore, it is often necessary to generate many machine instructions to fully express a complete operator calculation logic. In this embodiment, large-granularity bytecode instructions corresponding to specific tensors are used, and an operator calculation logic can be expressed based on a small number of instructions. For example, for the addition operation between two tensors, only one bytecode instruction can indicate the addition of the two tensors.

[0144] Optionally, when the above step 302 (i.e., the bytecode instruction compilation process) is implemented by host hardware (e.g., CPU), and the above step 303 (i.e., the bytecode instruction interpretation and execution process) is implemented by dedicated device hardware (e.g., GPU), after generating at least one bytecode instruction, at least one bytecode instruction can also be moved to a memory space accessed by AI hardware, and the AI ​​hardware is used to run a virtual machine. In this way, the interpretation and execution of at least one bytecode instruction is implemented by running a virtual machine on the AI ​​hardware. Exemplarily, the AI ​​hardware is hardware specifically used to implement AI computing, such as a GPU, NPU, or TPU, which can accelerate AI computing.

[0145] For ease of understanding, the process of generating bytecode instructions and interpreting and executing bytecode instructions through a virtual machine will be described in detail below.

[0146] For example, see Figure 4 and Figure 5 , Figure 4 A schematic diagram of a flow chart of a data calculation method of another model provided in an embodiment of the present application; Figure 5 A schematic diagram of a system architecture provided in an embodiment of the present application. Figure 4 The method shown can be applied to Figure 5 In the system architecture shown in Figure 4 As shown, the execution process of the data operation method of the model may include the following steps 401-408.

[0147] Step 401: convert the computation graph into a meta-operator graph.

[0148] In this embodiment, since the granularity of the processing function in the virtual machine when executing bytecode instruction processing corresponds to the meta-operator, in order to facilitate the generation of bytecode instructions, the calculation graph corresponding to the AI ​​model can be first represented as the corresponding meta-operator graph.

[0149] For example, see Figure 6 , Figure 6 A schematic diagram of a bytecode instruction generation process provided in an embodiment of the present application. Figure 6 As shown in FIG. 1 , the operator nodes included in the computation graph are nodes used to represent composite operators. By splitting the composite operator represented in the computation graph into multiple meta-operators, the operator nodes in the computation graph can be converted into multiple meta-operator nodes connected in sequence, thereby realizing the conversion of the computation graph into a meta-operator graph. Specifically, in Figure 6 In the calculation graph shown, the SqrtGrad operator is the operator to be executed, and the input tensors of the SqrtGrad operator are A and B respectively, and the output tensor is C; and the shapes corresponding to A, B and C are all two-dimensional, and the specific dimension values ​​are unknown. By splitting the SqrtGrad operator as a composite operator into square meta-operators and div meta-operators, the conversion from the calculation graph to the meta-operator graph can be achieved. That is, C = SqrtGrad (A, B) = Div (Square (A), B).

[0150] Step 402, representing the meta-operator graph as a meta-operator instruction graph.

[0151] When the actual input tensor is obtained, the meta-operator graph can be represented as a meta-operator instruction graph. Specifically, based on the meta-operator graph, the shapes and storage addresses of the input tensors, intermediate tensors, and output tensors in the meta-operator graph are updated based on the shape of the actual input tensor, so as to facilitate the generation of bytecode instructions based on the meta-operator instruction graph. In short, the meta-operator instruction graph can represent the content required to generate bytecode instructions in the form of a graph, that is, the shape and storage address of the tensor and the operation performed on the tensor.

[0152] In addition, when executing bytecode instructions through a virtual machine in the device hardware, the input tensor and the output tensor are often stored in the global memory of the device, while the intermediate tensor obtained by performing operations on the input tensor is stored in the local memory of the device. That is, when the processor core in the device hardware processes the bytecode instructions, it will store the generated intermediate variables in the local memory corresponding to the processor core to facilitate the rapid execution of the calculation; when the processor core executes the bytecode instructions in the meta-operator instruction graph, the calculated operation results (i.e., output tensors) can be saved to the global memory to facilitate the execution of subsequent other calculations. Based on this, in the process of representing the meta-operator graph as a meta-operator instruction graph, the Load meta-operator and the Store meta-operator can also be added to the boundary of the meta-operator graph. Among them, the Load meta-operator is used to load the input tensor from the global memory of the device to the local memory of the device, and the Store meta-operator is used to load the output tensor from the local memory of the device to the global memory of the device.

[0153] like Figure 6 As shown, in the process of converting the meta-operator graph into the meta-operator instruction graph, the shapes of the input tensor A, input tensor B, intermediate tensors, and output tensor C in the meta-operator graph are first refreshed based on the shapes and storage addresses of the actual input tensors A and B, and the storage addresses of the input tensors A, B, and output tensors C are refreshed. In addition, the Load meta-operator and the Store meta-operator are added to the boundary of the meta-operator graph (i.e., the upper and lower sides of the meta-operator graph) to realize the loading of the input tensors A and B into the local memory of the device hardware and the storage of the output tensor into the global memory of the device hardware.

[0154] Specifically, the shapes of input tensor A, input tensor B, intermediate tensor, and output tensor C are all [10, 200], and the storage address of input tensor A is 0x1000, the storage address of input tensor B is 0x2000, and the storage address of input tensor C is 0x3000. The intermediate tensor refers to the tensor obtained after the operations corresponding to the meta-operators are performed on the input tensor A and the input tensor B, for example, the tensor obtained after the operations corresponding to the meta-operators such as Load meta-operator, Square meta-operator, Div meta-operator, etc. are performed on the input tensor A.

[0155] Step 403, split the shape of the tensor of the elementary operator instruction graph to split the tensor into multiple parts to perform the operation.

[0156] In this embodiment, in order to facilitate the distribution of AI calculations to multiple processor cores for parallel processing, the meta-operator instruction graph can be pre-divided into tensor shapes so that when bytecode instructions are subsequently generated, the generated bytecode instructions can be used to instruct the tensor to be divided into multiple parts to perform operations in parallel.

[0157] Specifically, in the process of dividing the shape of the tensor of the meta-operator instruction graph, it can be determined how to divide the shape of the tensor, that is, how many parts the tensor is divided into, by combining the shape of the input tensor in the meta-operator instruction graph, the local memory constraint of the device hardware, and the number of processor cores of the device hardware. Among them, the local memory constraint of the device hardware means that the local memory corresponding to each processor core in the device hardware is limited, and the processor core needs to store the generated intermediate tensor in the local memory when processing the meta-operator instruction. Therefore, after the tensor is divided, the amount of data of the intermediate tensor generated by each processor core in the process of processing the bytecode instruction cannot be greater than the space size of the local memory corresponding to the processor core. In addition, in order to make the best use of each processor core in the device hardware, when the tensor is divided, the number of parts of the divided tensor can be greater than or equal to the number of processor cores of the device hardware to ensure that each processor core in the device hardware can perform tensor calculations.

[0158] For example, Figure 6 As shown in the figure, the shapes of input tensor A and input tensor B are both [10, 200]. The tensors can be split from the left dimension to split each input tensor into 10 parts. Therefore, both input tensor A and input tensor B are split into 10 tensors, and the shape of each tensor after splitting is [1, 200].

[0159] Step 404 , translate the meta-operators into corresponding bytecode instructions one by one according to the execution dependencies of the meta-operators according to the meta-operators' meta-operators' instruction graph after the execution tensor shape segmentation.

[0160] Since the bytecode instruction processing granularity corresponds to the meta-operator, the meta-operator nodes in the meta-operator instruction graph can be translated into corresponding bytecode instructions in sequence according to the execution dependency order of the meta-operators. In addition, for intermediate tensors, it is necessary to allocate address space for them in the local memory of the device hardware and ensure that they are not overwritten before their life cycle ends.

[0161] For the generated bytecode instruction, the bytecode instruction contains an instruction identifier, which is used to indicate the processing execution function corresponding to the bytecode instruction. The bytecode instruction also includes the storage address of the tensor data to be operated, the type of operation to be performed on the tensor data, and the storage address of the tensor obtained after the operation is performed. The read and write operation objects of different bytecode instructions are multi-dimensional tensor data. When the shape of the input tensor changes, the address used to store the input tensor and the output tensor will inevitably change, which will cause the bytecode instruction to change.

[0162] It should be noted that, when the shape of the tensor is segmented, the bytecode instruction may also indicate that the shape of the tensor is segmented so that the virtual machine can perform tensor operations in parallel through multiple processor cores when interpreting and executing the bytecode instruction.

[0163] Exemplarily, taking at least one bytecode instruction introduced in the above embodiment as an example, at least one bytecode instruction can also be used to instruct the input tensor to be split into multiple parts to perform calculation operations separately. In this way, when the virtual machine interprets and executes at least one bytecode instruction, at least one bytecode instruction can be processed in parallel by multiple virtual machine instances located in different processor cores, and different virtual machine instances in the multiple virtual machine instances are used to process different data in the input tensor. In other words, by indicating that the tensor is to be split in the bytecode instruction, the tensor can be split and allocated to multiple processor cores for parallel processing when the bytecode instruction is interpreted and executed, thereby realizing parallel processing of tensor operations and improving the efficiency of AI operations.

[0164] Optionally, at least one bytecode instruction includes a split number, which is used to indicate the number of splits of the input tensor. In addition, the split number is greater than or equal to the number of processor cores. In this way, by indicating the number of splits of the tensor in the bytecode instruction, the virtual machine can quickly split the input tensor into multiple parts for processing by the corresponding processor cores based on the split of the tensor when interpreting and executing the bytecode instruction, without the virtual machine having to determine how to split the tensor, thereby improving the efficiency of AI operations.

[0165] In addition, when the meta-operator instruction graph includes the Load meta-operator and the Store meta-operator, when generating bytecode instructions, corresponding data movement instructions also need to be generated.

[0166] Taking the above-mentioned embodiment of generating at least one bytecode instruction as an example, in the process of generating bytecode instructions, a first bytecode instruction and a second bytecode instruction can be generated based on the computation graph and the input tensor. Among them, the first bytecode instruction is used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction is used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory. In this way, when the bytecode instruction is executed by the virtual machine, the first bytecode instruction, at least one bytecode instruction, and the second bytecode instruction can be executed in sequence by the virtual machine.

[0167] In general, by generating additional corresponding data movement instructions according to the execution of bytecode instructions on hardware during the bytecode instruction generation stage, it is possible to move tensors between global memory and local memory, ensuring that different processor cores on the hardware can smoothly perform operations on tensors, and ensuring that multiple processor cores can perform tensor operations in parallel, thereby improving the feasibility of the solution.

[0168] For example, Figure 6As shown, in the process of executing bytecode instruction generation on the meta-operator instruction graph, a corresponding bytecode instruction can be generated for each meta-operator in the meta-operator instruction graph, and the generated bytecode instruction is actually in binary format. For ease of understanding, Figure 6 It is presented in text form. A meta-operator instruction graph corresponds to a bytecode instruction segment, which starts with begin and ends with end. Each bytecode instruction corresponds to a meta-operator and contains context information related to the execution of the instruction. In addition, the beginning of the bytecode instruction segment marks the number of divisions of the tensor shape, that is, Figure 6 The tile=10 shown in , means that the tensor is divided into 10 parts for parallel computing.

[0169] It should be noted that when multiple processor cores are used to process bytecode instructions, multiple processor cores actually process the same bytecode instructions, but the processor core will determine which part of the tensor content in the entire input tensor needs to be operated on based on the tensor segmentation and the core number of the current processor core, so that each processor core processes part of the input tensor separately.

[0170] Step 405, move the generated bytecode instruction segment to the device global memory, and start the virtual machine to execute the bytecode instruction segment.

[0171] Since the bytecode instruction generation process is actually executed on the host hardware (such as the CPU), and the virtual machine used to interpret and execute the bytecode instructions is deployed on the device hardware (such as the GPU), and the virtual machine can only access the memory on the device hardware. Therefore, after the bytecode instructions are generated, the generated bytecode instructions can be moved from the host hardware's memory to the device hardware's global memory, and the virtual machine on the device hardware is started to execute the bytecode instructions.

[0172] Generally speaking, when the host hardware generates bytecode instructions, it often stores the bytecode instructions in a certain order, such as placing multiple bytecode instructions corresponding to a meta-operator graph in a continuous memory area, thereby obtaining a bytecode instruction segment composed of multiple bytecode instructions, that is, each bytecode instruction segment includes a section of continuously stored bytecode instructions. In this way, in the process of executing bytecode instruction movement, each bytecode instruction segment generated by the host hardware can be moved from the host hardware to the device global memory in units of bytecode instruction segments.

[0173] Step 406, loading the bytecode instruction segment from the device global memory.

[0174] In the process of interpreting and executing bytecode instructions by the virtual machine, the virtual machine can assign bytecode instruction segments to one or more processor cores for processing. In order to facilitate the processor core to interpret and execute the bytecode instruction segments, the bytecode instruction segments can be first moved from the device global memory to the local memory corresponding to the processor core, and then the processor core reads each bytecode instruction in the bytecode instruction segment in the local memory. In addition, the processor core can also read the bytecode instruction segment directly from the global memory.

[0175] Step 407: According to the instruction identifier in the bytecode instruction, the corresponding processing function is called to execute each bytecode instruction in the bytecode instruction segment in sequence.

[0176] After the bytecode instruction segment is moved to the local memory of the device, the virtual machine can call the corresponding processing functions in sequence to execute the corresponding bytecode instructions according to the execution order between multiple bytecode instructions in the bytecode instruction segment, thereby realizing the processing of multiple bytecode instructions in sequence according to a certain order.

[0177] In the process of processing any bytecode instruction, the virtual machine can determine the operation type corresponding to the bytecode instruction according to the instruction identifier in the bytecode instruction, and then call the corresponding processing function to interpret and execute the bytecode instruction. Different types of bytecode instructions are represented by different instruction identifiers, and different types of bytecode instructions correspond to different processing functions.

[0178] Step 408: The processing function extracts the instruction field in the bytecode instruction, and calls the corresponding machine instruction according to the instruction field to perform the corresponding operation.

[0179] In the process of processing bytecode instructions, the processing function can extract instruction fields in the bytecode instructions, such as the field indicating the operation, the field indicating the storage address of the input tensor, and the field indicating the storage address of the output tensor. In this way, after the processing function extracts these instruction fields, it can determine the type of operation indicated by the bytecode instruction and the storage address of the tensor data that needs to perform the operation, and then call the corresponding machine instruction to perform the corresponding operation on the tensor.

[0180] The above describes in detail the process of generating bytecode instructions and interpreting and executing bytecode instructions through a virtual machine. For ease of understanding, the following will introduce in detail the process of running an AI model based on the above process with specific examples.

[0181] For example, see Figure 7 , Figure 7 This is a system architecture diagram for an actual application scenario provided by an embodiment of the present application. Figure 7As shown, this embodiment implements automatic operator fusion and execution of dynamic computation graphs in the open source MindSpore AI computing framework to accelerate the running performance of dynamic computation graph scenarios. Since operator fusion requires dynamic generation of different kernel functions, and the specific shape of the tensor can only be known at runtime in the dynamic computation graph, this embodiment can be used to generate and execute fusion operator kernel functions. As a special scenario, this embodiment can also solve the problem of automatic operator compilation of dynamic computation graphs.

[0182] exist Figure 7 In the system architecture shown, the MindSpore AI computing framework runs on a server whose hardware includes a processor, memory, and disk. The virtual machine runs on an AI chip (such as a GPU) that has a dedicated AI processor core and high bandwidth memory (HBM).

[0183] The MindSpore AI computing framework mainly includes three modules: Python graph composition, computational graph compilation, and bytecode compilation. The virtual machine mainly includes three modules: bytecode loading, bytecode instruction distribution, and bytecode instruction execution. The following will introduce the process of how the MindSpore AI computing framework and the virtual machine work together to implement the operation of the AI ​​model.

[0184] See also Fig. 8A , Fig. 8A The following is a schematic diagram of the execution flow of an AI computing framework provided in an embodiment of the present application. Fig. 8A As shown, the execution process of the AI ​​computing framework includes the following steps 801-806.

[0185] Step 801, construct a corresponding calculation graph based on the AI ​​model.

[0186] In this embodiment, step 801 is performed by the Python graphing module in the MindSpore AI computing framework, which mainly builds a corresponding computational graph based on the AI ​​model described by the MindSpore Python interface. In other words, in the process of running the AI ​​model, the MindSpore Python interface can be called to execute the operation of the AI ​​model, thereby triggering the execution of a series of steps mentioned in this embodiment.

[0187] Step 802: Replace the composite operator in the computation graph with multiple element operators.

[0188] In this embodiment, step 802 and step 803 are performed by the computation graph compilation module in the MindSpore AI computing framework. After obtaining the computation graph, since the operator nodes in the computation graph usually indicate composite operators (such as convolution operators or pooling operators), the composite operators in the computation graph can be replaced with multiple meta-operators to facilitate the subsequent generation of bytecode instructions based on the meta-operator granularity.

[0189] Step 803: merge multiple adjacent meta-operators that can be fused into a single meta-operator graph.

[0190] After the composite operator in the computation graph is replaced with multiple meta-operators, the entire computation graph is actually composed of a large number of meta-operators. In order to facilitate the generation of subsequent bytecode instructions, multiple adjacent meta-operators that can be fused in the computation graph can be merged into a single meta-operator graph. In this way, by dividing the multiple meta-operators that can be fused in the computation graph, the computation graph can be split into multiple meta-operator graphs, each of which includes multiple meta-operators. The meta-operators included in each meta-operator graph can be determined according to the actual hardware environment in which the virtual machine that interprets and executes the bytecode instructions runs.

[0191] Specifically, if some meta-operators with large differences in output tensor shapes are fused in the same meta-operator graph, it may be difficult to implement tensor segmentation, which may lead to insufficient local memory when the processor core processes bytecode instructions. Therefore, when fusing meta-operators, the differences in the shapes of the output tensors corresponding to the meta-operators need to be considered. In addition, since the local memory of the processor core is limited, if too many meta-operators are fused in a meta-operator graph, it is also easy to cause insufficient local memory. Therefore, when fusing meta-operators, the size of the local memory of the processor core needs to be considered to avoid fusing too many meta-operators. While meeting the local memory constraints, more meta-operators can be fused in the same meta-operator graph as much as possible, so that the intermediate data obtained by processing multiple meta-operators in the same meta-operator graph can be stored in the local memory of the device with faster read and write speeds, without the need to frequently read and write data from the global internal of the device with slower read and write speeds, thereby improving the processing efficiency of the operator.

[0192] Step 804: bytecode compilation is performed on the meta-operator graph to obtain bytecode instructions.

[0193] After obtaining the meta-operator graph, the meta-operator graph can be converted into a meta-operator instruction graph based on the actual input tensor, and the meta-operator instruction graph can be segmented into tensor shapes to generate corresponding bytecode instructions. Specifically, the process of generating bytecode instructions based on the meta-operator graph can refer to the description of the above embodiment, which will not be repeated here.

[0194] Step 805 , based on the data length of the bytecode instruction, apply for HBM space, and move the bytecode instruction to the HBM space.

[0195] After generating the bytecode instructions, you can apply for the corresponding HBM space from the AI ​​chip based on the data length of the generated bytecode instructions, so that the generated bytecode instructions can be moved to the AI ​​chip for interpretation and execution. After applying for the HBM space, you can move the bytecode instructions to the HBM space.

[0196] Step 806, starting the virtual machine to execute the bytecode instruction using the HBM space address and the data length of the bytecode instruction as parameters.

[0197] After the bytecode instructions are moved, the HBM space address and the data length of the bytecode instructions can be used as input parameters to start the virtual machine on the AI ​​chip through the runtime driver interface of the AI ​​chip to execute the bytecode instructions. When starting the virtual machine to execute the bytecode instructions, the number of processor cores on the AI ​​chip that process the bytecode instructions in parallel can be configured to be the same as the number of tensor shape splits in the bytecode instructions, so as to fully utilize the parallel processing capabilities of multiple processor cores on the AI ​​chip.

[0198] In addition, after the virtual machine completes executing the bytecode instruction, it can release the HBM space corresponding to the bytecode instruction.

[0199] See also Figure 8B , Figure 8B A flow chart of a virtual machine interpreting and executing bytecode instructions provided in an embodiment of the present application. Figure 8B As shown, the process of the virtual machine interpreting and executing bytecode instructions includes the following steps 807-812.

[0200] Step 807, based on the HBM space address and the data length of the bytecode instruction, load the bytecode instruction from the HBM to the local memory.

[0201] During the process of the virtual machine executing bytecode instructions, since the hardware that actually processes the bytecode instructions is the multiple processor cores in the AI ​​chip, the bytecode instructions can be loaded from the HBM to the local memory of the processor core based on the HBM space address and the data length of the bytecode instructions, so that each processor core can process the corresponding bytecode instructions.

[0202] Step 808, point the instruction cursor to the first address of the bytecode instruction in the local memory.

[0203] When the processor core processes bytecode instructions, the bytecode instructions can be executed one by one based on the instruction cursor. That is, when the bytecode instruction is first interpreted and executed, the instruction cursor is first pointed to the first address of the bytecode instruction in the local memory, so as to interpret and execute the first bytecode instruction.

[0204] Step 809, read the instruction identifier from the bytecode instruction pointed to by the current instruction cursor, and call the corresponding processing function.

[0205] Since each bytecode instruction includes a unique instruction identifier, and the instruction identifier is determined based on the type of the meta-operator corresponding to the bytecode instruction, the corresponding processing function can be called based on the instruction identifier to process the bytecode instruction.

[0206] Step 810: The processing function reads other instruction fields from the bytecode instruction and calls machine instructions to complete the operation.

[0207] The process of the processing function processing the bytecode instruction may refer to the above-mentioned embodiment, which will not be described in detail here.

[0208] Step 811, determine whether the last bytecode instruction processed is the last instruction.

[0209] After processing a bytecode instruction, it can be determined whether the last bytecode instruction processed is the last instruction. If the last bytecode instruction processed is the last instruction, the bytecode instruction processing flow is exited.

[0210] Step 812, move the instruction cursor to the next bytecode instruction.

[0211] If the last bytecode instruction processed is not the last instruction, the instruction cursor is pointed to the next bytecode instruction, and the processing of the next bytecode instruction continues.

[0212] For example, see Fig. 9 , Fig. 9 A system architecture diagram for another practical application scenario provided in an embodiment of the present application. Fig. 9 The system architecture shown is similar to Figure 7 The system architecture shown is identical in hardware except that Figure 7 The system architecture shown is based on the AI ​​model to generate a computational graph and perform subsequent processing on the computational graph; Fig. 9 The system architecture shown is that in the process of running the AI ​​model, the operator call interface is used to trigger the construction of the meta-operator graph corresponding to the operator and perform subsequent processing on the meta-operator graph.

[0213] Specifically, the AI ​​model code is executed using the dynamic language of Python, and the operator interface is dynamically called at runtime to issue and execute the operator. Since the operator is dynamically issued when the AI ​​model code is executed, and even if the same operator is called, the shape of the input tensor may be different each time it is called, it is not acceptable for existing automatic operator compilation technology to perform full-process operator compilation of binary files every time. Since this solution uses bytecode compilation, it can actually execute the same processing function on the virtual machine in this scenario, and only needs to regenerate the bytecode instructions each time it is run.

[0214] The MindSpore AI computing framework mainly includes four modules: Python operator call, meta-operator graph expansion, meta-operator graph splitting, and bytecode compilation. The virtual machine mainly includes three modules: bytecode loading, bytecode instruction distribution, and bytecode instruction execution. Among them, the steps executed on the virtual machine are the same as Figure 7 The corresponding embodiments are the same, please refer to the above embodiments for details. The following will introduce the process of generating bytecode instructions based on the issued operator by the MindSpore AI computing framework.

[0215] (I) Python operator call: The code for the AI ​​model is implemented in Python. For tensor calculations in the AI ​​model, the operator call is triggered by calling the Python interface corresponding to MindSpore.

[0216] (ii) Meta-operator graph expansion: If the Python operator calls and issues a composite operator, the composite operator can be expanded into the corresponding meta-operator subgraph so that it can be directly recognized and processed during the bytecode compilation stage.

[0217] (III) Meta-operator graph splitting: For some composite operators, since the structure of the expanded meta-operator subgraph is complex and cannot be executed through a single virtual machine kernel function call, the meta-operator graph corresponding to the composite operator can be split into multiple meta-operator graphs.

[0218] (iv) Bytecode compilation: The obtained one or more meta-operator graphs are byte-compiled and the virtual machine kernel function is started and executed according to the corresponding input tensors.

[0219] The method provided in the embodiment of the present application is described in detail above. Next, the device provided in the embodiment of the present application for executing the above method will be introduced.

[0220] See also Fig.10 , Fig.10 A schematic diagram of a data computing device of a model provided in an embodiment of the present application. Fig.10As shown, the data operation device of the model includes: an acquisition module 1001, which is used to obtain a calculation graph and a shape of an input tensor, the calculation graph is used to indicate the operation in the artificial intelligence AI model, and the input tensor is used to represent the input data corresponding to the operation; a processing module 1002, which is used to generate at least one bytecode instruction based on the calculation graph and the shape of the input tensor, and the at least one bytecode instruction is used to indicate the operation performed on the input tensor; the processing module 1002 is also used to execute at least one bytecode instruction through a virtual machine, and the virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to interpret the at least one bytecode instruction and call the machine instruction based on the interpretation result to perform the operation on the input tensor.

[0221] In one possible implementation, at least one bytecode instruction is also used to instruct that an input tensor be divided into multiple parts to perform calculation operations separately; processing module 1002 is specifically used to process at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

[0222] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0223] In a possible implementation, the number of splits is greater than or equal to the number of the multiple virtual machine instances.

[0224] In one possible implementation, the processing module 1002 is further used to: generate a first bytecode instruction and a second bytecode instruction based on the computation graph and the input tensor, the first bytecode instruction being used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction being used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory; and execute the first bytecode instruction, at least one bytecode instruction, and the second bytecode instruction in sequence through the virtual machine.

[0225] In a possible implementation, after generating at least one bytecode instruction, the processing module 1002 is further used to: move the at least one bytecode instruction to a memory space accessed by the AI ​​hardware, and the AI ​​hardware is used to run the virtual machine.

[0226] In one possible implementation, the processing module 1002 is specifically used to: obtain a first meta-operator graph based on the computational graph conversion, the first meta-operator graph including multiple meta-operators, wherein the multiple meta-operators are used to indicate basic calculation operations; generate at least one bytecode instruction based on the shape of the input tensor and the meta-operator graph, and the multiple meta-operators correspond to the at least one bytecode instruction.

[0227] In one possible implementation, the processing module 1002 is further used to: convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; divide the converted computation graph into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include a first meta-operator graph, and the plurality of meta-operator graphs each include a plurality of meta-operators.

[0228] In one possible implementation, the acquisition module 1001 is also used to obtain an operator call instruction, where the operator call instruction is used to instruct the execution of an operation corresponding to a target operator, and the operator call instruction includes an input tensor; the processing module 1002 is also used to generate a computational graph based on the target operator indicated in the operator call instruction, where the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic operations.

[0229] In a possible implementation, at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

[0230] In one possible implementation, at least one bytecode instruction includes a data type identifier, which is used to indicate the data type of the input tensor.

[0231] In a possible implementation, at least one bytecode instruction is further used to indicate a storage address of an input tensor and a storage address of an output tensor, where the output tensor is a tensor obtained after performing a calculation operation on the input tensor. Furthermore, the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

[0232] See also Fig.11 , Fig.11 This is a schematic diagram of a structure of an execution device provided in an embodiment of the present application. The execution device 1100 can be specifically a server, a personal computer, a smart phone, etc., which is not limited here. Specifically, the execution device 1100 includes: a receiver 1101, a transmitter 1102, a processor 1103 and a memory 1104 (wherein the number of processors 1103 in the execution device 1100 can be one or more, Fig.11 In the example of FIG. 1 , the processor 1103 may include an application processor 11031 and a communication processor 11032. In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 may be connected via a bus or other means.

[0233] The memory 1104 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1103. A portion of the memory 1104 may also include a non-volatile random access memory (NVRAM). The memory 1104 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0234] The processor 1103 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.

[0235] The method disclosed in the above embodiment of the present application can be applied to the processor 1103, or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1103 or an instruction in the form of software. The above processor 1103 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0236] The processor 1103 can implement or execute the methods, steps and logic diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1104, and the processor 1103 reads the information in the memory 1104 and completes the steps of the above method in combination with its hardware.

[0237] The receiver 1101 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1102 can be used to output digital or character information through the first interface; the transmitter 1102 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1102 can also include a display device such as a display screen.

[0238] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the method for determining the model structure described in the above embodiment, or so that the chip in the training device executes the method for determining the model structure described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0239] For details, please refer to Fig.12 , Fig.12 A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 1200, NPU 1200 is mounted on the host CPU (Host CPU) as a coprocessor, and the host CPU assigns tasks. The core part of the NPU is the operation circuit 1203, which is controlled by the controller 1204 to extract matrix data from the memory and perform multiplication operations.

[0240] In some implementations, the operation circuit 1203 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1203 is a two-dimensional systolic array. The operation circuit 1203 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1203 is a general-purpose matrix processor.

[0241] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1202 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1201 and performs matrix operation with matrix B, and the partial result or final result of the matrix is ​​stored in the accumulator 1208.

[0242] The unified memory 1206 is used to store input data and output data. The weight data is directly transferred to the weight memory 1202 through the direct memory access controller (DMAC) 1205. The input data is also transferred to the unified memory 1206 through the DMAC.

[0243] BIU stands for Bus Interface Unit, i.e., bus interface unit 1210 , which is used for the interaction between AXI bus, DMAC and instruction fetch buffer (IFB) 1209 .

[0244] The bus interface unit 1210 (BIU) is used for the instruction fetch memory 1209 to obtain instructions from the external memory, and is also used for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0245] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1206 or to transfer weight data to the weight memory 1202 or to transfer input data to the input memory 1201.

[0246] The vector calculation unit 1207 includes multiple operation processing units, and further processes the output of the operation circuit 1203 when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, upsampling of feature planes, etc.

[0247] In some implementations, the vector calculation unit 1207 can store the processed output vector to the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function; or a nonlinear function to the output of the operation circuit 1203, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values ​​to generate an activation value. In some implementations, the vector calculation unit 1207 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1203, for example, for use in a subsequent layer in a neural network.

[0248] An instruction fetch buffer 1209 connected to the controller 1204 is used to store instructions used by the controller 1204;

[0249] Unified memory 1206, input memory 1201, weight memory 1202 and instruction fetch memory 1209 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0250] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0251] See also Fig.13 , Fig.13 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the above Figure 3 The disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of manufacture.

[0252] Fig.13 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0253] In one embodiment, the computer readable storage medium 1300 is provided using a signal bearing medium 1301. The signal bearing medium 1301 may include one or more program instructions 1302, which when executed by one or more processors may provide the above-mentioned Figure 3 Describes the functionality or part of the functionality.

[0254] In some examples, the signal bearing medium 1301 may include a computer readable medium 1303 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, and the like.

[0255] In some embodiments, the signal bearing medium 1301 may include a computer recordable medium 1304, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, etc. In some embodiments, the signal bearing medium 1301 may include a communication medium 1305, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, etc.). Thus, for example, the signal bearing medium 1301 may be communicated by a wireless form of the communication medium 1305 (e.g., a wireless communication medium that complies with the IEEE 802.11 standard or other transmission protocol).

[0256] The one or more program instructions 1302 may be, for example, computer executable instructions or logic implementation instructions. In some examples, the computing device of the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1302 communicated to the computing device via one or more of the computer readable medium 1303, the computer recordable medium 1304, and / or the communication medium 1305.

[0257] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0258] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the method of each embodiment of the present application.

[0259] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0260] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integration. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state hard disk (SSD)), etc.

Claims

1. A data operation method, It is characterized in that include: Obtaining a computational graph and a shape of an input tensor, wherein the computational graph is used to indicate a computing operation in an artificial intelligence (AI) model, and the input tensor is used to represent input data corresponding to the computing operation; Based on the computation graph and the shape of the input tensor, generate at least one bytecode instruction, where the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor; The at least one bytecode instruction is executed by a virtual machine to obtain an output tensor, and a processing function corresponding to the at least one bytecode instruction is configured in the virtual machine, and the processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

2. The method according to claim 1, It is characterized in that The at least one bytecode instruction is further used to instruct to split the input tensor into a plurality of parts to perform operation operations respectively; The executing the at least one bytecode instruction by a virtual machine includes: The at least one bytecode instruction is processed in parallel by multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

3. The method according to claim 2, It is characterized in that The at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

4. The method according to claim 3, It is characterized in that The number of splits is greater than or equal to the number of the multiple virtual machine instances.

5. The method according to any one of claims 1 to 4, It is characterized in that The method further comprises: Based on the computation graph and the input tensor, generate a first bytecode instruction and a second bytecode instruction, wherein the first bytecode instruction is used to instruct to move the input tensor from the global memory to the local memory, and the second bytecode instruction is used to instruct to move an output tensor obtained by processing the input tensor from the local memory to the global memory; The executing the at least one bytecode instruction by a virtual machine includes: The first bytecode instruction, the at least one bytecode instruction and the second bytecode instruction are executed in sequence by a virtual machine.

6. The method according to any one of claims 1 to 5, It is characterized in that After generating at least one bytecode instruction, the method further comprises: The at least one bytecode instruction is moved to a memory space accessed by AI hardware, and the AI ​​hardware is used to run the virtual machine.

7. The method according to any one of claims 1 to 6, It is characterized in that The generating at least one bytecode instruction based on the computation graph and the shape of the input tensor comprises: A first meta-operator graph is obtained based on the conversion of the computation graph, wherein the first meta-operator graph includes a plurality of meta-operators, wherein the computation graph is used to indicate part or all of the computing operations in the AI ​​model, and the plurality of meta-operators are used to indicate basic computing operations; Based on the shape of the input tensor and the first meta-operator graph, the at least one bytecode instruction is generated, and the multiple meta-operators correspond to the at least one bytecode instruction.

8. The method according to claim 7, It is characterized in that The converting the computation graph to obtain a first-element operator graph includes: Convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; The converted computation graph is divided into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include the first meta-operator graph, and each of the plurality of meta-operator graphs includes a plurality of meta-operators.

9. The method according to any one of claims 1 to 6, It is characterized in that The obtaining of the computation graph includes: Obtain an operator call instruction, where the operator call instruction is used to instruct execution of a calculation operation corresponding to a target operator, and the operator call instruction includes the input tensor; The computation graph is generated based on the target operator indicated in the operator call instruction, wherein the computation graph includes a plurality of meta-operators for representing the target operator, and the plurality of meta-operators are all used to indicate basic calculation operations.

10. The method according to any one of claims 1 to 9, It is characterized in that The at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

11. The method according to any one of claims 1 to 10, It is characterized in that The at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate the data type of the input tensor.

12. The method according to claim 5, It is characterized in that The at least one bytecode instruction is also used to indicate a storage address of the input tensor and a storage address of the output tensor, and the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

13. A data computing device for a model, It is characterized in that include: An acquisition module, used to acquire the shape of a computational graph and an input tensor, wherein the computational graph is used to indicate a computing operation in an artificial intelligence AI model, and the input tensor is used to represent input data corresponding to the computing operation; A processing module, configured to generate at least one bytecode instruction based on the computation graph and the shape of the input tensor, wherein the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor; The processing module is also used to execute the at least one bytecode instruction through a virtual machine to obtain an output tensor. The virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

14. The device according to claim 13, It is characterized in that The at least one bytecode instruction is further used to instruct to split the input tensor into a plurality of parts to perform operation operations respectively; The processing module is specifically used to process the at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

15. The device according to claim 14, It is characterized in that The at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

16. The device according to claim 15, It is characterized in that The number of divisions is greater than or equal to the number of the plurality of processor cores.

17. The device according to any one of claims 14 to 16, It is characterized in that The processing module is further used for: Based on the computation graph and the input tensor, generate a first bytecode instruction and a second bytecode instruction, wherein the first bytecode instruction is used to instruct to move the input tensor from the global memory to the local memory, and the second bytecode instruction is used to instruct to move an output tensor obtained by processing the input tensor from the local memory to the global memory; The first bytecode instruction, the at least one bytecode instruction and the second bytecode instruction are executed in sequence by a virtual machine.

18. The device according to any one of claims 13 to 17, It is characterized in that After generating at least one bytecode instruction, the processing module is further configured to: The at least one bytecode instruction is moved to a memory space accessed by AI hardware, and the AI ​​hardware is used to run the virtual machine.

19. The device according to any one of claims 13 to 18, It is characterized in that The processing module is specifically used for: A first meta-operator graph is obtained based on the conversion of the computation graph, wherein the first meta-operator graph includes a plurality of meta-operators, wherein the computation graph is used to indicate part or all of the computing operations in the AI ​​model, and the plurality of meta-operators are used to indicate basic computing operations; Based on the shape of the input tensor and the first meta-operator graph, the at least one bytecode instruction is generated, and the multiple meta-operators correspond to the at least one bytecode instruction.

20. The device according to claim 19, It is characterized in that The processing module is further used for: Convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; The converted computation graph is divided into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include the first meta-operator graph, and each of the plurality of meta-operator graphs includes a plurality of meta-operators.

21. The device according to any one of claims 13 to 18, It is characterized in that The acquisition module is further used to acquire an operator call instruction, where the operator call instruction is used to instruct to execute a calculation operation corresponding to a target operator, and the operator call instruction includes the input tensor; The processing module is further used to generate the calculation graph based on the target operator indicated in the operator call instruction, wherein the calculation graph includes multiple meta-operators used to represent the target operator, and the multiple meta-operators are all used to indicate basic calculation operations.

22. The device according to any one of claims 13 to 21, It is characterized in that The at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

23. The device according to any one of claims 13 to 22, It is characterized in that The at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate the data type of the input tensor.

24. The device according to claim 17, It is characterized in that The at least one bytecode instruction is also used to indicate a storage address of the input tensor and a storage address of the output tensor, and the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

25. A data computing device for a model, It is characterized in that The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 12.

26. A computer storage medium, It is characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 12.

27. A computer program product, It is characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 12.

Citation Information

Cited By

  • Instruction processing method, programmable processor, electronic equipment and storage medium

    CN122018984A

  • Data operation method for model, and related apparatus

    EP4811128A1

  • Data operation method for model, and related apparatus

    WO2025119128A1