Data operation method for model, and related apparatus

By compiling the calculation graph of the AI ​​model into bytecode instructions and using virtual machines to interpret and run, the problem of large running time of AI model in the existing technology is solved, and a more efficient operation of AI model is achieved.

WO2025119128A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136065
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing AI computing framework cannot compile kernel functions in advance, resulting in a large run-time of AI models, especially for AI models with uncertain lengths of input data or intermediate data, such as the Transformer model.

Method used

By compiling the calculation graph corresponding to the operations in the AI ​​model into bytecode instructions, and interpreting the running bytecode instructions through a virtual machine pre-configured with corresponding processing functions, the traditional cumbersome compilation process is avoided and the run time of the AI ​​model is reduced.

Benefits of technology

It effectively avoids the traditional compilation process, reduces the running time of the AI ​​model, and improves the running efficiency of the AI ​​model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136065_12062025_PF_FP_ABST
    Figure CN2024136065_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A data operation method for a model, and a related apparatus, which are applied to the running of an artificial intelligence (AI) model. In the method, on the basis of an input tensor of an operation in an AI model during actual running, a computation graph corresponding to the operation in the AI model is compiled into a bytecode instruction, and the bytecode instruction is then interpreted and run by means of a virtual machine for which a corresponding processing function is pre-configured, such that the operation in the AI model is executed, thereby effectively avoiding the execution of a traditional tedious compiling process, and shortening the running duration of the AI model.
Need to check novelty before this filing date? Find Prior Art

Description

A data calculation method and related device for a model

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 5, 2023, with application number 202311665038.6 and application name “A model data calculation method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data calculation method for a model and related devices. Background Art

[0003] In recent years, AI technologies, represented by deep learning, have made significant progress, achieving promising results in fields such as computer vision and natural language processing. To improve AI model development efficiency and computing performance, the industry often adopts AI computing frameworks to express and compute AI models. AI computing frameworks typically provide hundreds or even thousands of different types of operators for users to use. Multiple operators are interconnected to form a computation graph, which corresponds to a specific AI model. When executed, different operators take one or more tensors as computational input parameters, then call a matching kernel function to perform the corresponding computation, ultimately producing one or more tensors as output.

[0004] To improve AI computing performance, AI computing frameworks across the industry generally use operator fusion. This involves combining one or more adjacent operator nodes in a computation graph into a new fused operator for overall computational execution. Due to the vast number of operators that can be fused across different computation graphs, automatic kernel compilation is currently widely used to generate kernel functions corresponding to fused operators. Automatic kernel compilation automatically generates machine instructions that can be executed directly on the device based on the computational semantics of the fused operator and the shapes of the input and output tensors.

[0005] However, for AI models with uncertain lengths of some input data or intermediate data (such as Transformer models), the shape of the input tensor corresponding to the operation in the AI ​​model can only be known when the AI ​​model is actually executed. Therefore, it is impossible to perform kernel function compilation for this part of the AI ​​model in advance, resulting in a kernel function compilation stage every time this part of the AI ​​model is run, which leads to a longer running time of the AI ​​model. Summary of the Invention

[0006] The present application provides a data operation method for a model. Based on the input tensors of the operations in the AI ​​model during actual operation, the calculation graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the execution of the traditional cumbersome compilation process, and reducing the running time of the AI ​​model.

[0007] The first aspect of the present application provides a data operation method for a model, which is applied to the operation of an AI model. In this method, a calculation graph and the shape of an input tensor are first obtained. The calculation graph is used to indicate the operation in the AI ​​model, and the input tensor is used to represent the input data corresponding to the operation. The shape of the input tensor is the size of the input tensor. That is, the calculation graph actually indicates how to perform operation operations on tensors, and the input tensor is the input data corresponding to the operation indicated in the calculation graph. Among them, the input tensor is multidimensional data, and the shape of the input tensor is not fixed, and can be determined according to the actual operation of the AI ​​model. In addition, the input tensor may, for example, include one or more tensors, depending on the number of input data indicated by the calculation graph.

[0008] Then, based on the computation graph and the shape of the input tensor, at least one bytecode instruction is generated, the at least one bytecode instruction being used to indicate an operation to be performed on the input tensor. The at least one bytecode instruction is a binary instruction that indicates the operation to be performed on the input tensor. In addition, the at least one bytecode instruction is actually a bytecode, i.e., an intermediate code, which cannot be directly recognized and executed by hardware, but requires interpretation and execution by a software module.

[0009] Finally, at least one bytecode instruction is executed through a virtual machine, and a processing function corresponding to at least one bytecode instruction is configured in the virtual machine. The processing function is used to interpret at least one bytecode instruction and call a machine instruction that can be recognized by the hardware based on the interpretation result (that is, a machine instruction corresponding to the at least one bytecode instruction) to perform an operation on the input tensor. Among them, the processing function interprets the bytecode instruction means that the processing function analyzes the content of the bytecode instruction and calls the machine instruction based on the content of the bytecode instruction to implement the operation on the tensor data. Specifically, a plurality of processing functions are configured in the virtual machine, and different processing functions are used to process different types of bytecode instructions to ensure that bytecode instructions indicating different operation operations are processed by corresponding processing functions. Specifically, the virtual machine in this scheme can also be called a kernel function, which is the execution entity of the operator at runtime, and can interpret and execute bytecode instructions to realize the operation of the operator. Among them, the specific implementation of the virtual machine can be a binary instruction code that can be directly executed on the device hardware.

[0010] In this solution, based on the input tensors of the operations in the AI ​​model during actual runtime, the computational graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run by a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the traditional tedious compilation process and reducing the running time of the AI ​​model.

[0011] In addition, compared with the model compilation process performed on AI models in existing related technologies, this solution only needs to generate simple bytecode instructions based on the calculation graph and the actual input tensor, without the need to perform a complete compilation process (i.e. preprocessing, syntax parsing, instruction generation, assembly, linking and output files), and the generated bytecode instructions can be interpreted and executed by the processing function in the pre-implemented virtual machine, ensuring that the generation and execution speed of bytecode instructions is fast, effectively improving the operating efficiency of the AI ​​model.

[0012] In one possible implementation, at least one bytecode instruction is further used to instruct the input tensor to be split into multiple parts for performing separate computations. Thus, when the at least one bytecode instruction is interpreted and executed by a virtual machine, the at least one bytecode instruction can be processed in parallel by multiple virtual machine instances located on different processor cores, with different virtual machine instances in the multiple virtual machine instances being used to process different data in the input tensor. That is, each of the multiple virtual machine instances is responsible for processing a portion of the data in the input tensor, thereby enabling multiple virtual machine instances to process the input tensor in parallel.

[0013] In this solution, by instructing the splitting of tensors in bytecode instructions, the tensors can be split and distributed to multiple processor cores for parallel processing when interpreting and executing bytecode instructions, thereby realizing parallel processing of tensor operations and improving the efficiency of AI operations.

[0014] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0015] In one possible implementation, the number of splits is greater than or equal to the number of virtual machine instances. Generally, one virtual machine instance runs on one processor core, so the number of splits is actually greater than or equal to the number of processor cores used to execute bytecode instructions.

[0016] In this solution, by indicating the number of times the tensor is split in the bytecode instruction, the input tensor can be quickly split into multiple parts based on the tensor split when interpreting and executing the bytecode instruction, and the parts are given to the corresponding processor cores for processing. There is no need for the virtual machine to further determine how to split the tensor, thereby improving the efficiency of AI operations.

[0017] In one possible implementation, the method further includes: generating a first bytecode instruction and a second bytecode instruction based on the computation graph and the input tensor, wherein the first bytecode instruction is used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction is used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory.

[0018] In the execution phase of the bytecode instructions, specifically, the first bytecode instruction, the at least one bytecode instruction, and the second bytecode instruction may be executed in sequence by the virtual machine.

[0019] It should be noted that this solution introduces a first bytecode instruction based on the computation graph that instructs the movement of input tensors from global memory to local memory, and a second bytecode instruction that instructs the movement of output tensors obtained by processing the input tensors from local memory to global memory. In some special cases, such as when the input tensor is a random tensor, it is not necessary to generate the bytecode instruction to move the input tensor from global memory to local memory. Instead, the random tensor can be generated directly in local memory.

[0020] In this solution, by generating additional corresponding data movement instructions based on the execution of bytecode instructions on hardware during the bytecode instruction generation stage, it is possible to move tensors between global memory and local memory, ensuring that different processor cores on the hardware can smoothly perform operations on tensors, and ensuring that multiple processor cores can perform tensor operations in parallel, thereby improving the feasibility of the solution.

[0021] In one possible implementation, after generating at least one bytecode instruction, the method further includes: moving the at least one bytecode instruction to a memory space accessed by AI hardware, where the AI ​​hardware is used to run a virtual machine. Exemplarily, the AI ​​hardware is hardware specifically designed for AI computing, such as a graphics processing unit (GPU), a neural network processor (NPU), or a tensor processing unit (TPU), which can accelerate AI computing.

[0022] In one possible implementation, at least one bytecode instruction is generated based on the computation graph and the shape of the input tensor, including: obtaining a first meta-operator graph based on the computation graph conversion, wherein the first meta-operator graph includes multiple meta-operators. The computation graph is used to indicate some or all of the computing operations in the AI ​​model, and multiple meta-operators are used to indicate basic computing operations. That is, the meta-operator is the most basic unit for performing computing operations, and the meta-operator cannot be obtained by combining other more basic operators. Then, based on the shape of the input tensor and the first meta-operator graph, at least one bytecode instruction is generated, and multiple meta-operators correspond to the at least one bytecode instruction mentioned above. In addition, a meta-operator may correspond to one or more bytecode instructions.

[0023] In this solution, since bytecode instructions are processed by processing functions configured in the virtual machine, and all operators can be obtained by combining meta-operators, generating bytecode instructions according to the meta-operator granularity can minimize the types of generated bytecode instructions, thereby reducing the number of pre-configured processing functions in the virtual machine and reducing the implementation complexity of the virtual machine.

[0024] In one possible implementation, obtaining a first meta-operator graph based on a computation graph conversion specifically includes: converting each operator in the computation graph into one or more meta-operators to obtain a converted computation graph. That is, some operators in the computation graph may be composite operators composed of multiple meta-operators. Therefore, all operators in the computation graph can be represented by meta-operators, thereby obtaining a converted computation graph composed of meta-operators. Then, the converted computation graph is divided into multiple continuous meta-operator graphs, wherein the multiple meta-operator graphs include the first meta-operator graph, and the multiple meta-operator graphs each include multiple meta-operators. That is, for the converted computation graph, the converted computation graph can be divided into multiple parts according to the execution order of the meta-operators, and each part includes multiple adjacent meta-operators. In this way, by fusing multiple adjacent meta-operators in each part, a meta-operator graph can be obtained, thereby achieving the division of the converted computation graph into multiple meta-operator graphs, and each meta-operator graph is obtained by fusing multiple meta-operators.

[0025] This solution splits the converted computation graph into multiple meta-operator graphs for processing, ensuring that the virtual machine executes a single meta-operator graph at a time. This prevents the machine from running out of memory due to excessive sequential processing of meta-operators. Furthermore, by fusing multiple meta-operators into a single meta-operator graph, the intermediate data generated by processing multiple meta-operators within the same meta-operator graph can be stored in the local memory of the device with faster read / write speeds, eliminating the need for frequent global read / write operations on the device with slower read / write speeds. This improves operator processing efficiency.

[0026] In one possible implementation, at least one bytecode instruction includes an instruction identifier, which is used to uniquely identify the type of the at least one bytecode instruction. The virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction. For example, a bytecode instruction indicating that the operation to be performed is an addition operation can be represented by the instruction identifier 00; a bytecode instruction indicating that the operation to be performed is a subtraction operation can be represented by the instruction identifier 01. In this solution, by setting the instruction identifier in the bytecode instruction, the type of the bytecode instruction can be uniquely identified, thereby facilitating the virtual machine to quickly call the corresponding processing function to process the bytecode instruction according to the instruction identifier, thereby improving the efficiency of interpreting and executing the bytecode instruction.

[0027] In one possible implementation, at least one bytecode instruction includes a data type identifier that indicates the data type of the input tensor. For example, the data type identifier fp32 indicates that the data type of the input tensor is a 32-bit floating point number, so the operation performed on the input tensor is actually a 32-bit floating point calculation. Exemplarily, in a bytecode instruction, different data type identifiers can be used to indicate different data types, such as 32-bit floating point numbers, 16-bit floating point numbers, 32-bit integers, and other data types, and the data types are not specifically limited herein.

[0028] In one possible implementation, the at least one bytecode instruction is further configured to indicate a storage address of an input tensor and a storage address of an output tensor, and both the storage address of the input tensor and the storage address of the output tensor are addresses in local memory.

[0029] It should be noted that when the input tensor is a random variable, since the input tensor can actually be randomly generated when used, the storage address of the input tensor may not be indicated in the bytecode instruction, but only the storage address of the output tensor.

[0030] In one possible implementation, obtaining a computational graph specifically includes: obtaining an operator call instruction, the operator call instruction is used to instruct the execution of an arithmetic operation corresponding to a target operator, and the operator call instruction includes an input tensor; and generating a computational graph based on the target operator indicated in the operator call instruction. The computational graph is generated based on the target operator and is used to indicate the execution of the arithmetic operation indicated by the target operator on the input tensor by means of operator nodes and directed edges. Furthermore, the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic arithmetic operations. That is, the computational graph itself is composed of multiple meta-operators, and no longer includes a composite operator composed of multiple meta-operators.

[0031] The second aspect of the present application provides a data operation device for a model, including: an acquisition module, used to obtain a calculation graph and the shape of an input tensor, the calculation graph is used to indicate the calculation operation in the artificial intelligence AI model, and the input tensor is used to represent the input data corresponding to the calculation operation; a processing module, used to generate at least one bytecode instruction based on the calculation graph and the shape of the input tensor, and the at least one bytecode instruction is used to indicate the calculation operation performed on the input tensor; the processing module is also used to execute at least one bytecode instruction through a virtual machine, and the virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to interpret the at least one bytecode instruction and call a machine instruction based on the interpretation result to perform the calculation operation on the input tensor.

[0032] In one possible implementation, at least one bytecode instruction is further used to instruct the input tensor to be divided into multiple parts to perform calculation operations separately; the processing module is specifically used to process at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

[0033] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0034] In a possible implementation, the number of splits is greater than or equal to the number of virtual machine instances.

[0035] In one possible implementation, the processing module is further used to: generate a first bytecode instruction and a second bytecode instruction based on the computation graph and the input tensor, the first bytecode instruction being used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction being used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory; and execute the first bytecode instruction, at least one bytecode instruction, and the second bytecode instruction in sequence through the virtual machine.

[0036] In one possible implementation, after generating at least one bytecode instruction, the processing module is further used to: move the at least one bytecode instruction to a memory space accessed by AI hardware, where the AI ​​hardware is used to run a virtual machine.

[0037] In one possible implementation, the processing module is specifically used to: obtain a first meta-operator graph based on the computational graph conversion, the first meta-operator graph including multiple meta-operators, wherein the computational graph is used to indicate part or all of the computing operations in the AI ​​model, and multiple meta-operators are used to indicate basic computing operations; generate at least one bytecode instruction based on the shape of the input tensor and the first meta-operator graph, and multiple meta-operators correspond to the at least one bytecode instruction.

[0038] In one possible implementation, the processing module is further used to: convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; divide the converted computation graph into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include a first meta-operator graph, and the plurality of meta-operator graphs each include a plurality of meta-operators.

[0039] In one possible implementation, the acquisition module is further used to obtain an operator call instruction, which is used to instruct the execution of an operation corresponding to a target operator, and the operator call instruction includes an input tensor; the processing module is further used to generate a computational graph based on the target operator indicated in the operator call instruction, wherein the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic computational operations.

[0040] In a possible implementation, at least one bytecode instruction includes an instruction identifier, and the virtual machine is configured to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

[0041] In one possible implementation, at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate a data type of an input tensor.

[0042] In one possible implementation, the at least one bytecode instruction further indicates a storage address of an input tensor and a storage address of an output tensor, where the output tensor is a tensor obtained after performing the operation on the input tensor. Furthermore, the storage address of the input tensor and the storage address of the output tensor are both addresses in local memory.

[0043] In a third aspect, the present application provides a data operation device for a model, which may include a processor coupled to a memory, wherein the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For details of the steps in each possible implementation of the first aspect executed by the processor, please refer to the first aspect and will not be repeated here.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method of any implementation of the first aspect.

[0045] A fifth aspect of the present application provides a circuit system, the circuit system including a processing circuit, and the processing circuit is configured to execute a method of any implementation manner of the above-mentioned first aspect.

[0046] The sixth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method of any implementation manner of the first aspect.

[0047] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in any implementation of the first aspect, for example, sending or processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory for storing program instructions and data necessary for the server or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0048] The beneficial effects of the second to seventh aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] FIG1 is a schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0050] FIG2 is a schematic diagram of a system architecture of an electronic device provided in an embodiment of the present application;

[0051] FIG3 is a flow chart of a data calculation method for a model provided in an embodiment of the present application;

[0052] FIG4 is a flow chart of a data calculation method of another model provided in an embodiment of the present application;

[0053] FIG5 is a schematic diagram of a system architecture provided in an embodiment of the present application;

[0054] FIG6 is a schematic diagram of a flow chart of bytecode instruction generation according to an embodiment of the present application;

[0055] FIG7 is a system architecture diagram of an actual application scenario provided by an embodiment of the present application;

[0056] FIG8A is a schematic diagram of an execution flow of an AI computing framework provided in an embodiment of the present application;

[0057] FIG8B is a schematic diagram of a flow chart of a virtual machine interpreting and executing bytecode instructions according to an embodiment of the present application;

[0058] FIG9 is a system architecture diagram of another practical application scenario provided by an embodiment of the present application;

[0059] FIG10 is a schematic diagram of the structure of a data computing device of a model provided in an embodiment of the present application;

[0060] FIG11 is a schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0061] FIG12 is a schematic structural diagram of a chip provided in an embodiment of the present application;

[0062] FIG13 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0064] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0065] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0066] (1) Neural Network

[0067] A neural network can be composed of neural units. Specifically, it can be understood as a neural network with an input layer, a hidden layer, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (DNN). The work of each layer in the neural network can be expressed mathematically as To describe, from a physical level, the work of each layer in the neural network can be understood as completing the transformation from input space to output space (i.e., from the row space to the column space of the matrix) through five operations on the input space (a set of input vectors). These five operations include: 1. Dimensionality increase / decrease; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are represented by Completed, operation 4 is completed by "+b", and operation 5 is implemented by "a()". The word "space" is used here because the object being classified is not a single thing, but a class of things. Space refers to the collection of all individuals of this type of thing, where W is the weight matrix of each layer of the neural network, and each value in the matrix represents the weight value of a neuron in the layer. The matrix W determines the spatial transformation from the input space to the output space described above, that is, the W of each layer of the neural network controls how to transform the space. The purpose of training a neural network is to eventually obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of a neural network is essentially about learning how to control spatial transformation, and more specifically, about learning the weight matrix.

[0068] (2) Self-Attention Network

[0069] A self-attention network is a neural network that uses the self-attention mechanism. Typical examples include the Transformer model. The self-attention mechanism is an attention mechanism that associates different positions in a sequence to compute a representation of the same sequence. Self-attention networks are commonly used in fields such as machine reading, abstract summarization, and image description.

[0070] (3) AI Model

[0071] An AI model is a mathematical model that learns and predicts data that exhibits certain regularities and predictability. Currently, AI models are generally constructed using neural networks. During the operation of an AI model, the computational process of learning from data is called training, and the process of predicting the results of input data is called inference.

[0072] (4) AI computing framework

[0073] AI computing frameworks are software platforms used to express and process AI models, such as TensorFlow, PyTorch, and MindSpore.

[0074] (5) Operator

[0075] Operators are the basic computing units of AI models. Each operator represents specific computational semantics. Common operators represent computational semantics such as convolution, pooling, and activation functions.

[0076] (6) Meta-operator

[0077] Meta-operators are operators that represent the most basic operations. Meta-operators cannot be combined with other more basic operators to represent operations. For example, some common meta-operators represent basic operations such as addition, subtraction, multiplication, and division.

[0078] (7) Composite operator

[0079] A composite operator refers to an operator that can be expressed by combining multiple primitive operators, such as the convolution operator and the pooling operator.

[0080] (8) Kernel function

[0081] The kernel function is the execution entity of the operator at runtime, that is, the kernel function is actually a binary instruction code that can be directly executed on the device hardware.

[0082] (9) Tensor

[0083] A tensor is a multidimensional data type that is usually used to represent input data or output data during the operation of an operator.

[0084] (10)shape

[0085] The shape represents the dimension of the tensor data. For example, [3,4] represents a 3*4 two-dimensional tensor.

[0086] (11) Computational graph

[0087] A computation graph is a directed acyclic graph consisting of operators as nodes and tensors as edges. Different AI models can be abstracted into corresponding computation graph structures for compilation and execution within the AI ​​model runtime platform.

[0088] (12) Bytecode

[0089] Bytecode is a binary file containing an executable program, consisting of a sequence of opcode / data pairs. Compared to machine instruction code, which can be directly executed by hardware, bytecode is actually an intermediate code—an instruction encoding that needs to be interpreted and executed by software code, and cannot be directly executed by hardware.

[0090] (13) Virtual Machine

[0091] In this embodiment, a virtual machine is an important tool in programming languages. It is essentially a software module that converts high-level language code (such as bytecode) into low-level machine instructions, allowing the code to run on different operating systems and hardware platforms. For example, a Java virtual machine for the Java language can compile Java code into bytecode, which is then converted into machine instructions for execution through an interpreter or just-in-time compiler.

[0092] In addition, a virtual machine instance refers to a virtual machine running on a hardware device (such as a processor), that is, a virtual machine instance refers to a running virtual machine. Based on the same virtual machine software code, multiple virtual machine instances with the same configuration can be quickly created.

[0093] (14) Machine Instructions

[0094] Machine instructions are instructions that computer hardware (such as the Central Processing Unit (CPU)) can directly recognize and execute. They are represented by binary code. Machine instructions typically consist of two parts: an opcode and an operand. The opcode specifies the operation to be performed by the machine instruction, i.e., its function; the operand specifies the object involved in the operation and the location where the result of the operation is stored.

[0095] (15) Global Memory

[0096] Global memory refers to the memory space within device hardware (such as a GPU) used to store global and static variables. Variables stored in global memory can be accessed and modified by all objects running within the device hardware. In other words, global memory is shared by multiple cores and can be accessed by all processor cores within the device hardware as well as the host hardware.

[0097] (16) Local Memory

[0098] Local memory refers to the private memory space allocated to each processor core in the device hardware. It can only be accessed by the corresponding processor core and cannot be accessed by other processor cores.

[0099] Currently, traditional AI computing frameworks often use static computation graphs to run AI models. Specifically, the shape of the AI ​​model's input data remains constant, such as an image of fixed size. Consequently, traditional AI computing frameworks generate a corresponding static computation graph based on the AI ​​model's structure and the shape of the input data. This pre-compiled kernel function can then be used to process the AI ​​model's actual input data during the model execution phase.

[0100] However, with the development of AI technology, the types of AI models are increasing, and AI models with uncertain input data or intermediate data lengths continue to emerge, such as the Transformer model for processing natural language sequences. For these AI models with uncertain input data or intermediate data lengths, the computational graphs corresponding to these AI models are actually dynamic computational graphs, that is, the shape of the input data of the computational graph is not fixed, but can change. Therefore, for these AI models corresponding to dynamic computational graphs, since the shape of the input tensor corresponding to the AI ​​model can only be known when the AI ​​model is specifically executed, the traditional AI computing framework cannot perform the kernel function compilation stage for this part of the AI ​​model in advance, resulting in the need to perform the kernel function compilation stage based on the actual input data each time this part of the AI ​​model is run, resulting in a longer running time for the AI ​​model.

[0101] Based on this, this embodiment provides a data operation method for a model. Based on the input tensors of the operations in the AI ​​model during actual runtime, the calculation graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the execution of the traditional cumbersome compilation process, and reducing the running time of the AI ​​model.

[0102] To facilitate understanding, the following first introduces the system architecture used in the data calculation method of the model provided in the embodiment of the present application.

[0103] Please refer to Figure 1. An embodiment of the present application provides a schematic diagram of a system architecture 100. As shown in Figure 1, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices. The execution device 110 can be arranged at one physical site or distributed across multiple physical sites. The execution device 110 can use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the data calculation method of the model provided in the embodiment of the present application, thereby realizing the operation of the AI ​​model.

[0104] Users can operate their respective user devices (e.g., local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smartphone, a tablet computer, a smart camera, a smart car or other type of cellular phone, a media consumption device, a wearable device, a set-top box, a game console, etc.

[0105] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0106] In one implementation, the execution device 110 is used to implement the data operation method of the model provided in the embodiment of the present application, and send the obtained operation results to the local device 101 and the local device 102 through the communication network, so that the local device 101 and the local device 102 can obtain the operation results of the AI ​​model, such as the image classification results or translation results output by the AI ​​model.

[0107] In another implementation, one or more aspects of the execution device 110 may be implemented by each local device. For example, the local device 101 may provide local data or feedback calculation results to the execution device 110, or execute the data operation method of the model provided in the embodiments of the present application. That is, the execution device 110 may send the AI ​​model or the calculation graph corresponding to the AI ​​model to the local device 101 via the communication network, and the local device 101 may execute the data operation method of the model provided in the embodiments of the present application.

[0108] It should be noted that all functions of the execution device 110 may also be implemented by a local device. For example, the local device 101 implements the functions of the execution device 110 and provides services to its own user, or provides services to the user of the local device 102.

[0109] In general, the training method of the model provided in the embodiments of the present application can be applied to electronic devices, such as the execution device 110, the local device 101, or the local device 102 described above. For example, the electronic device can be, for example, a server, a wireless electronic device in industrial control, a smart phone (mobile phone), a personal computer (PC), a laptop, a tablet computer, an autonomous vehicle, a smart camera, and the like. For ease of understanding, the method provided in the embodiments of the present application will be described below by taking the application of the method provided in the embodiments of the present application to a server as an example.

[0110] Please refer to Figure 2, which is a schematic diagram of the system architecture of an electronic device provided in an embodiment of the present application. In the system architecture shown in Figure 2, host hardware and device hardware are included. The host hardware includes a processor and memory for executing relevant functions of host software modules such as AI models and AI computing frameworks, and caching intermediate data. The device hardware includes a processor and memory for executing the code of the device software, thereby realizing the functions of the virtual machine, and caching intermediate data. Among them, the host hardware and the device hardware can be deployed on the same electronic device, for example, the host hardware and the device hardware are deployed on the same server. Among them, the host hardware includes the CPU and memory on the server, and the device hardware includes, for example, a graphics processing unit (GPU), a neural network processor (NPU) or a tensor processing unit (TPU) on the server.

[0111] The AI ​​computing framework running on the host hardware can obtain the AI ​​model and build a corresponding computational graph based on the AI ​​model. It then performs bytecode compilation based on the computational graph and the actual input tensors of the AI ​​model to obtain bytecode instructions. The host hardware then sends the compiled bytecode instructions to the virtual machine on the device hardware, which executes the bytecode instructions, thereby performing the operations indicated by the AI ​​model and completing the operation of the AI ​​model.

[0112] It should be noted that in some scenarios (for example, the aforementioned device hardware is not available or there is no idle device hardware), the host hardware can also assume the functions of the device hardware, that is, a virtual machine is run on the host hardware to interpret and execute bytecode instructions.

[0113] The above describes the system architecture used by the method provided in this embodiment. The following will describe in detail the specific execution process of the method provided in this embodiment in conjunction with the accompanying drawings.

[0114] Please refer to Figure 3, which is a flow chart of a data calculation method for a model provided in an embodiment of the present application. As shown in Figure 3, the data calculation method for the model includes the following steps 301-303.

[0115] Step 301: Obtain the shape of the computation graph and the input tensor. The computation graph is used to indicate the calculation operations in the AI ​​model, and the input tensor is used to represent the input data corresponding to the calculation operations.

[0116] In this embodiment, a computation graph is a directed acyclic graph consisting of operators as nodes and tensors as edges, which is used to indicate how to perform operations on tensors. Furthermore, the computation graph is derived from an AI model and is used to indicate operations within the AI ​​model, such as convolution operations, pooling operations, and the like.

[0117] Optionally, the computational graph may be used to indicate the computational operations in the entire AI model, or may be used to indicate part of the computational operations in the AI ​​model (for example, indicating the operations of a certain neural network layer or a certain operator in the AI ​​model). This embodiment does not specifically limit this.

[0118] Since the computation graph actually indicates how to perform computational operations on tensors, this embodiment also obtains the shape of the input tensor corresponding to the computation graph. That is, the input tensor is the input data corresponding to the computational operation indicated in the computation graph. Exemplarily, the input tensor is, for example, the actual input data of the entire AI model; for example, in the case where the AI ​​model is a natural language processing model, the input tensor is, for example, the tensor corresponding to the text that needs to be input into the natural language processing model. Alternatively, the input tensor is, for example, the intermediate data generated during the operation of the AI ​​model, such as the tensor data output by a neural network layer in the AI ​​model. Among them, the input tensor is multidimensional data, and the shape of the input tensor is not fixed and can be determined according to the actual operation of the AI ​​model. In addition, the input tensor may include one or more tensors, depending on the number of input data indicated by the computation graph, and this embodiment does not specifically limit this.

[0119] It should be noted that the shape of the input data indicated in the computation graph obtained in this embodiment is unknown. Only after the shape of the input tensor is obtained can the shape of the input data processed based on the computation graph be determined. Furthermore, the shape of the input tensor can be determined after the actual input tensor is obtained, or it can be obtained in advance during the actual input tensor generation process.

[0120] In this embodiment, there are multiple ways to obtain the calculation graph.

[0121] In one possible implementation, a computation graph can be generated based on an AI model. That is, when running the AI ​​model, the computation graph corresponding to the AI ​​model is first generated, and then the AI ​​model is run by executing the computation graph.

[0122] For example, after obtaining the AI ​​model that needs to be run, all or part of the operations in the AI ​​model can be converted into a computational graph based on the computational operations of each neural network layer in the AI ​​model, thereby describing the structure of the AI ​​model in the form of a graph.

[0123] In another possible implementation, during the AI ​​model's execution, for tensor calculations encountered during the AI ​​model's operation, the AI ​​computing framework's operator call interface is called to trigger the calculation of a specific operator, thereby generating a computation graph corresponding to the operator. In other words, during the AI ​​model's execution, the computation graph corresponding to the entire AI model is no longer generated. Instead, the operators in the AI ​​model are processed by calling the operator interface and generating the corresponding computation graph.

[0124] For example, during the operation of the AI ​​model, an operator call instruction can be obtained, which is used to instruct the execution of the operation corresponding to the target operator, and the operator call instruction includes the above-mentioned input tensor. Specifically, the target operator is, for example, one or more operators indicated in the AI ​​model, such as a convolution operator or a pooling operator. The input tensor included in the operator call instruction is the input data of the target operator.

[0125] Then, a computation graph is generated based on the target operator indicated in the operator call instruction. The computation graph is generated based on the target operator and is used to indicate the operation indicated by the target operator on the input tensor through operator nodes and directed edges. Furthermore, the computation graph includes multiple meta-operators used to represent the target operator, each of which is used to indicate a basic operation. In other words, the computation graph itself is composed of multiple meta-operators, and no longer includes composite operators composed of multiple meta-operators.

[0126] Step 302: Generate at least one bytecode instruction based on the computation graph and the shape of the input tensor, where the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor.

[0127] In this embodiment, after obtaining the computation graph and the shape of the input tensor, the type of operation to be performed and the shape of the specific input data corresponding to the operation can be determined. Therefore, based on the computation graph and the shape of the input tensor, at least one bytecode instruction can be generated. The at least one bytecode instruction is a binary instruction that indicates the operation to be performed on the input tensor. In addition, the at least one bytecode instruction is actually a bytecode that needs to be interpreted and executed by a software module and cannot be directly executed by hardware.

[0128] In addition, since the shape of the input tensor is known, how to store the input tensor and the output tensor can be determined, that is, at least one bytecode instruction can also be used to indicate the storage address of the input tensor and the storage address of the output tensor, where the output tensor is the tensor obtained after performing an operation on the input tensor. In this way, when executing at least one bytecode instruction, the input tensor on which the operation needs to be performed, the type of operation performed on the input tensor, and the storage address of the operation result obtained after performing the operation on the input tensor can be obtained based on the at least one bytecode instruction, thereby ensuring that the corresponding operation in the running AI model can be realized by executing at least one bytecode instruction.

[0129] Optionally, in the process of generating bytecode instructions based on the computation graph and the input tensor, the bytecode instructions may be generated according to the granularity of the meta-operator, that is, one meta-operator corresponds to one or more bytecode instructions.

[0130] Exemplarily, when a computation graph is constructed based on all or part of the operations in an AI model, a first meta-operator graph can be obtained based on the computation graph conversion, and the first meta-operator graph includes multiple meta-operators, wherein multiple meta-operators are used to indicate basic computation operations. That is, a meta-operator is the most basic unit for performing computation operations, and a meta-operator cannot be obtained by combining other more basic operators. In other words, the operator indicated in the computation graph may be one or more composite operators, which are composed of the most basic meta-operators. The process of converting a computation graph into a meta-operator graph is actually to split the composite operator in the computation graph into multiple meta-operators for representation, thereby realizing the computational logic of the computation graph represented by the most basic meta-operators. For example, assuming that the operator indicated in the computation graph is SqrtGrad(x,y), the operator can be split into two element operators, namely SqrtGrad(x,y)=Div(Square(x),y); where SqrtGrad() represents root variance, Div() represents integer division operation, and Square() represents square operation.

[0131] Then, based on the shape of the input tensor and the first meta-operator graph, multiple bytecode instructions are generated, and the multiple bytecode instructions include at least one bytecode instruction mentioned above. The multiple bytecode instructions correspond to the multiple meta-operators indicated by the meta-operator graph, that is, each bytecode instruction corresponds to a meta-operator. Generally speaking, a meta-operator can correspond to one bytecode instruction; when the shape of the tensor operated by some meta-operators is too large, one bytecode instruction may not be able to completely represent one meta-operator, so one meta-operator may correspond to multiple bytecode instructions.

[0132] Since bytecode instructions are processed by the processing functions configured in the virtual machine, and all operators can be obtained by combining meta-operators, generating bytecode instructions according to the meta-operator granularity can minimize the types of generated bytecode instructions, thereby reducing the number of pre-configured processing functions in the virtual machine and reducing the implementation complexity of the virtual machine.

[0133] Optionally, in the process of obtaining the first meta-operator graph based on the calculation graph conversion, each operator in the calculation graph can be converted into one or more meta-operators to obtain a converted calculation graph. That is, some operators in the calculation graph may be composite operators composed of multiple meta-operators, so all operators in the calculation graph can be represented by meta-operators to obtain a converted calculation graph composed of meta-operators. Then, the converted calculation graph is divided into multiple continuous meta-operator graphs, the multiple meta-operator graphs include the first meta-operator graph, and the multiple meta-operator graphs each include multiple meta-operators. That is, for the converted calculation graph, multiple adjacent meta-operators can be fused to obtain a meta-operator graph. In this way, by dividing the meta-operators in the converted calculation graph into multiple parts, multiple meta-operators in each part can be fused to form a meta-operator graph.

[0134] In this solution, by splitting the converted computation graph into multiple meta-operator graphs for processing, it can ensure that the virtual machine executes a single meta-operator graph each time, avoiding the situation where the device running the virtual machine is short of memory due to too many meta-operators being processed continuously.

[0135] Step 303: Execute at least one bytecode instruction through a virtual machine to obtain an output tensor. The virtual machine is configured with a processing function corresponding to the at least one bytecode instruction. The processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

[0136] In this embodiment, the virtual machine is a pre-implemented software module that interprets and executes bytecode instructions. Specifically, the virtual machine is configured with multiple processing functions, each of which is used to process different types of bytecode instructions, ensuring that bytecode instructions indicating different computational operations are handled by corresponding processing functions. Therefore, after the virtual machine obtains at least one bytecode instruction, it can call the corresponding processing function based on the type of the at least one bytecode instruction to process the at least one bytecode instruction.

[0137] Specifically, the virtual machine in this solution, also known as a kernel function, is the execution entity of the operator at runtime, capable of interpreting and executing bytecode instructions to implement the operator. The specific implementation of the virtual machine can be binary instruction code that can be directly executed on device hardware (such as a GPU, NPU, or TPU).

[0138] Exemplarily, at least one bytecode instruction includes an instruction identifier, which is used to uniquely identify the type of the at least one bytecode instruction. That is, different types of bytecode instructions are marked by different instruction identifiers. For example, a bytecode instruction indicating that the operation to be performed is an addition operation can be represented by instruction identifier 00; a bytecode instruction indicating that the operation to be performed is a subtraction operation can be represented by instruction identifier 01; a bytecode instruction indicating that the operation to be performed is a multiplication operation can be represented by instruction identifier 10; and a bytecode instruction indicating that the operation to be performed is a division operation can be represented by instruction identifier 11. In this way, the virtual machine can call the processing function corresponding to the at least one bytecode instruction based on the instruction identifier in the at least one bytecode instruction to process the at least one bytecode instruction.

[0139] That is to say, by setting the instruction identifier in the bytecode instruction, the type of the bytecode instruction can be uniquely identified, thereby facilitating the virtual machine to quickly call the corresponding processing function to process the bytecode instruction according to the instruction identifier, thereby improving the efficiency of interpreting and executing the bytecode instruction.

[0140] Exemplarily, the types of bytecode instructions may include the following types: move class and calculation class. Among them, the move class may specifically include the Load type, which means moving tensor data from global memory to local memory; the Store type, which means writing tensor data from local memory back to global memory. The calculation class may include algebraic calculations (such as Add, Sub, Mul, Div, Sqrt, Abs, Exp or Pow), reduction calculations (such as Sum (sum), Max (maximum value of tensor data) or Min (maximum value of tensor data)), and comparison calculations (such as Greater, Less or Equal).

[0141] The processing function configured in the virtual machine is also pre-implemented software code. When processing bytecode instructions, the processing function reads the tensor to be operated on from the bytecode instructions, calls the corresponding machine instructions to perform the operation on the read tensor based on the operation indicated by the bytecode instruction, and finally stores the operation result to the storage address specified by the bytecode instruction.

[0142] In general, based on the input tensors of the operations in the AI ​​model during actual runtime, the computational graph corresponding to the operations in the AI ​​model is compiled into bytecode instructions, and then the bytecode instructions are interpreted and run through a virtual machine pre-configured with corresponding processing functions, thereby executing the operations in the AI ​​model, effectively avoiding the traditional tedious compilation process and reducing the running time of the AI ​​model.

[0143] Specifically, compared with the model compilation process performed on the AI ​​model in the existing related technologies, this embodiment only needs to generate simple bytecode instructions based on the calculation graph and the actual input tensor, without the need to perform a complete compilation process (i.e. preprocessing, syntax parsing, instruction generation, assembly, linking and output files), and the generated bytecode instructions can be interpreted and executed by the processing function in the pre-implemented virtual machine, ensuring that the generation and execution speed of the bytecode instructions is fast, effectively improving the operating efficiency of the AI ​​model.

[0144] In addition, the model compilation process performed on the AI ​​model in the related art ultimately generates machine instructions that can be directly executed by the hardware, and the compilation granularity is relatively small. For example, for the common addition operation between two tensors, it is necessary to generate multiple addition instructions for indicating the addition of two integers. Therefore, it is often necessary to generate many machine instructions to fully express a complete operator calculation logic. In this embodiment, large-granularity bytecode instructions corresponding to specific tensors are used, and an operator calculation logic can be expressed based on a small number of instructions. For example, for the addition operation between two tensors, only one bytecode instruction can indicate the addition of the two tensors.

[0145] Optionally, when the above-mentioned step 302 (i.e., the bytecode instruction compilation process) is implemented by host hardware (e.g., CPU), and the above-mentioned step 303 (i.e., the bytecode instruction interpretation and execution process) is implemented by dedicated device hardware (e.g., GPU), after generating at least one bytecode instruction, the at least one bytecode instruction can also be moved to the memory space accessed by the AI ​​hardware, and the AI ​​hardware is used to run the virtual machine. In this way, the interpretation and execution of at least one bytecode instruction is implemented by running the virtual machine on the AI ​​hardware. Exemplarily, the AI ​​hardware is hardware specifically used for implementing AI computing, such as a GPU, NPU, or TPU, which can achieve acceleration of AI computing.

[0146] For ease of understanding, the following details the process of generating bytecode instructions and interpreting and executing bytecode instructions through a virtual machine.

[0147] For example, please refer to Figures 4 and 5. Figure 4 is a schematic flow chart of a data operation method for another model provided in an embodiment of the present application; Figure 5 is a schematic flow chart of a system architecture provided in an embodiment of the present application. The method shown in Figure 4 can be applied to the system architecture shown in Figure 5. As shown in Figure 4, the execution flow of the data operation method for this model can include the following steps 401-408.

[0148] Step 401: Convert the computation graph into a meta-operator graph.

[0149] In this embodiment, since the granularity of the processing function in the virtual machine when executing bytecode instruction processing corresponds to the meta-operator, in order to facilitate the generation of bytecode instructions, the calculation graph corresponding to the AI ​​model can be first represented as the corresponding meta-operator graph.

[0150] For example, please refer to Figure 6, which is a flow chart of bytecode instruction generation provided by an embodiment of the present application. As shown in Figure 6, the operator nodes included in the calculation graph are nodes for representing composite operators. By splitting the composite operator represented in the calculation graph into multiple meta-operators, the operator nodes in the calculation graph can be converted into multiple meta-operator nodes connected in sequence, thereby realizing the conversion of the calculation graph into a meta-operator graph. Specifically, in the calculation graph shown in Figure 6, the SqrtGrad operator is the operator to be performed, and the input tensors of the SqrtGrad operator are A and B respectively, and the output tensor is C; and the shapes corresponding to A, B and C are all two-dimensional, and the specific dimension values ​​are unknown. By splitting the SqrtGrad operator as a composite operator into square meta-operators and div meta-operators, the conversion of the calculation graph to the meta-operator graph can be realized. That is, C=SqrtGrad(A,B)=Div(Square(A),B).

[0151] Step 402: Represent the meta-operator graph as a meta-operator instruction graph.

[0152] When the actual input tensor is obtained, the meta-operator graph can be represented as a meta-operator instruction graph. Specifically, based on the meta-operator graph, the shapes and storage addresses of the input tensors, intermediate tensors, and output tensors in the meta-operator graph are updated based on the shape of the actual input tensor, thereby facilitating the generation of bytecode instructions based on the meta-operator instruction graph. Simply put, the meta-operator instruction graph can represent the content required to generate bytecode instructions in the form of a graph, namely the shape and storage address of the tensor and the operations performed on the tensor.

[0153] In addition, when executing bytecode instructions through the virtual machine in the device hardware, the input tensors and output tensors are often stored in the device's global memory, while the intermediate tensors obtained by performing operations on the input tensors are stored in the device's local memory. That is, when the processor core in the device hardware processes the bytecode instructions, it will store the generated intermediate variables in the local memory corresponding to the processor core to facilitate the rapid execution of the calculation; when the processor core executes the bytecode instructions in the meta-operator instruction graph, it can save the calculated operation results (i.e., output tensors) to the global memory to facilitate the execution of subsequent other calculations. Based on this, in the process of representing the meta-operator graph as a meta-operator instruction graph, the Load meta-operator and the Store meta-operator can also be added to the boundary of the meta-operator graph. Among them, the Load meta-operator is used to load the input tensor from the device's global memory to the device's local memory, and the Store meta-operator is used to load the output tensor from the device's local memory to the device's global memory.

[0154] As shown in Figure 6, during the conversion of a meta-operator graph into a meta-operator instruction graph, the shapes of input tensors A, B, intermediate tensors, and output tensor C in the meta-operator graph are first refreshed based on the actual shapes and storage addresses of input tensors A and B, and the storage addresses of input tensors A, B, and C are also refreshed. Furthermore, Load and Store meta-operators are added to the boundaries of the meta-operator graph (i.e., the upper and lower sides of the meta-operator graph) to load input tensors A and B into the local memory of the device hardware and store the output tensor into the global memory of the device hardware.

[0155] Specifically, the shapes of input tensor A, input tensor B, intermediate tensor, and output tensor C are all [10, 200]. The storage address of input tensor A is 0x1000, the storage address of input tensor B is 0x2000, and the storage address of input tensor C is 0x3000. Intermediate tensors refer to the tensors obtained by performing operations corresponding to meta-operators on input tensors A and input tensors B. For example, the tensors obtained by performing operations corresponding to meta-operators such as the Load meta-operator, the Square meta-operator, and the Div meta-operator on input tensor A.

[0156] Step 403 , performing tensor shape segmentation on the primitive operator instruction graph to divide the tensor into multiple parts to perform operations.

[0157] In this embodiment, in order to facilitate the distribution of AI calculations to multiple processor cores for parallel processing, the meta-operator instruction graph can be pre-divided into tensor shapes so that when bytecode instructions are subsequently generated, the generated bytecode instructions can be used to instruct the tensor to be divided into multiple parts to perform operations in parallel.

[0158] Specifically, in the process of splitting the shape of the tensor of the meta-operator instruction graph, the shape of the input tensor in the meta-operator instruction graph, the local memory constraint of the device hardware, and the number of processor cores of the device hardware can be combined to determine how to split the tensor shape, that is, how many parts the tensor is split into. Among them, the local memory constraint of the device hardware means that the local memory corresponding to each processor core in the device hardware is limited, and the processor core needs to store the generated intermediate tensors in the local memory when processing the meta-operator instruction. Therefore, after the tensor is split, the amount of data of the intermediate tensors generated by each processor core in the process of processing the bytecode instruction cannot be greater than the space size of the local memory corresponding to the processor core. In addition, in order to make the best use of each processor core in the device hardware, when the tensor is split, the number of parts of the split tensor can be greater than or equal to the number of processor cores in the device hardware to ensure that each processor core in the device hardware can perform tensor calculations.

[0159] For example, as shown in Figure 6, the shapes of input tensors A and B are both [10, 200]. The tensors can be split along the left dimension to split each input tensor into 10 parts. Therefore, both input tensors A and B are split into 10 tensors, and the shape of each tensor after splitting is [1, 200].

[0160] Step 404 : The meta-operators are translated into corresponding bytecode instructions one by one according to the execution dependencies of the meta-operators according to the meta-operators' meta-operator instruction graph after the execution tensor shape is segmented.

[0161] Because the bytecode instruction processing granularity corresponds to the meta-operator, the meta-operator nodes in the meta-operator instruction graph can be translated into corresponding bytecode instructions in the order of their execution dependencies. Furthermore, for intermediate tensors, it is necessary to allocate address space in the local memory of the device hardware and ensure that they are not overwritten before their lifetime ends.

[0162] The generated bytecode instructions contain an instruction identifier that indicates the processing execution function corresponding to the bytecode instruction. The bytecode instructions also include the storage address of the tensor data to be operated on, the type of operation to be performed on the tensor data, and the storage address of the tensor obtained after the operation is performed. Different bytecode instructions read and write multidimensional tensor data. If the shape of the input tensor changes, the addresses used to store the input tensor and the output tensor will inevitably change, resulting in changes to the bytecode instructions.

[0163] It should be noted that, when the shape of the tensor is segmented, the bytecode instruction may also indicate that the shape of the tensor is segmented, so that the virtual machine can perform tensor operations in parallel through multiple processor cores when interpreting and executing the bytecode instruction.

[0164] Exemplarily, taking at least one bytecode instruction introduced in the above embodiment as an example, at least one bytecode instruction can also be used to instruct the input tensor to be split into multiple parts to perform calculation operations separately. In this way, when the virtual machine interprets and executes at least one bytecode instruction, the at least one bytecode instruction can be processed in parallel by multiple virtual machine instances located in different processor cores, and different virtual machine instances in the multiple virtual machine instances are used to process different data in the input tensor. In other words, by instructing the tensor to be split in the bytecode instruction, the tensor can be split and allocated to multiple processor cores for parallel processing when the bytecode instruction is interpreted and executed, thereby realizing parallel processing of tensor operations and improving the efficiency of AI operations.

[0165] Optionally, at least one bytecode instruction includes a split number, which is used to indicate the number of splits of the input tensor. In addition, the split number is greater than or equal to the number of processor cores. In this way, by indicating the number of splits of the tensor in the bytecode instruction, the virtual machine can quickly split the input tensor into multiple parts based on the split of the tensor when interpreting and executing the bytecode instruction, and assign them to the corresponding processor core for processing. The virtual machine does not need to further determine how to split the tensor, thereby improving the efficiency of AI operations.

[0166] In addition, when the meta-operator instruction graph includes the Load meta-operator and the Store meta-operator, corresponding data movement instructions need to be generated when generating bytecode instructions.

[0167] Taking the above-mentioned embodiment of generating at least one bytecode instruction as an example, during the process of generating the bytecode instruction, a first bytecode instruction and a second bytecode instruction can be generated based on the computation graph and the input tensor. The first bytecode instruction is used to instruct the input tensor to be moved from global memory to local memory, and the second bytecode instruction is used to instruct the output tensor obtained by processing the input tensor to be moved from local memory to global memory. Thus, when the bytecode instruction is executed by the virtual machine, the first bytecode instruction, the at least one bytecode instruction, and the second bytecode instruction can be executed sequentially by the virtual machine.

[0168] In general, by generating additional data movement instructions based on the execution of bytecode instructions on hardware during the bytecode instruction generation stage, it is possible to move tensors between global memory and local memory, ensuring that different processor cores on the hardware can smoothly perform operations on tensors, and ensuring that multiple processor cores can perform tensor operations in parallel, thereby improving the feasibility of the solution.

[0169] For example, as shown in FIG6 , in the process of executing bytecode instruction generation on the meta-operator instruction graph, a corresponding bytecode instruction can be generated for each meta-operator in the meta-operator instruction graph, and the generated bytecode instruction is actually in binary format. For ease of understanding, FIG6 is presented in text form. Among them, a meta-operator instruction graph corresponds to a bytecode instruction segment, and the bytecode instruction segment starts with begin and ends with end. Each bytecode instruction corresponds to a meta-operator and contains context information related to the execution of the instruction. In addition, the beginning of the bytecode instruction segment marks the number of divisions of the tensor shape, that is, tile=10 shown in FIG6 , which means that the tensor is divided into 10 parts for parallel computing.

[0170] It should be noted that when multiple processor cores are used to process bytecode instructions, multiple processor cores actually process the same bytecode instructions. The processor cores will determine which part of the tensor content in the entire input tensor needs to be operated on based on the tensor segmentation and the core number of the current processor core, so that each processor core processes part of the input tensor separately.

[0171] Step 405: Move the generated bytecode instruction segment to the device global memory, and start the virtual machine to execute the bytecode instruction segment.

[0172] Since the bytecode instruction generation process is actually executed on the host hardware (such as the CPU), and the virtual machine used to interpret and execute the bytecode instructions is deployed on the device hardware (such as the GPU), and the virtual machine can only access the memory on the device hardware, after the bytecode instructions are generated, the generated bytecode instructions can be moved from the host hardware's memory to the device hardware's global memory, and the virtual machine on the device hardware is started to execute the bytecode instructions.

[0173] Generally speaking, when host hardware generates bytecode instructions, it often stores them in a specific order. For example, multiple bytecode instructions corresponding to a meta-operator graph are placed in a continuous memory area, resulting in a bytecode instruction segment consisting of multiple bytecode instructions. In other words, each bytecode instruction segment includes a continuous section of bytecode instructions. Thus, during the bytecode instruction transfer process, each bytecode instruction segment generated by the host hardware can be transferred from the host hardware to the device global memory on a bytecode instruction segment basis.

[0174] Step 406: Load the bytecode instruction segment from the device global memory.

[0175] During the process of interpreting and executing bytecode instructions, the virtual machine can assign bytecode instruction segments to one or more processor cores for processing. To facilitate the interpretation and execution of the bytecode instruction segments by the processor cores, the bytecode instruction segments can first be moved from the device's global memory to the corresponding local memory of the processor core. The processor core then reads each bytecode instruction in the bytecode instruction segment from the local memory. Alternatively, the processor core can directly read the bytecode instruction segment from global memory.

[0176] Step 407: According to the instruction identifier in the bytecode instruction, the corresponding processing function is called to sequentially execute each bytecode instruction in the bytecode instruction segment.

[0177] After moving the bytecode instruction segment to the local memory of the device, the virtual machine can call the corresponding processing function in sequence to execute the corresponding bytecode instructions according to the execution order between multiple bytecode instructions in the bytecode instruction segment, thereby realizing the processing of multiple bytecode instructions in a certain order.

[0178] When processing any bytecode instruction, the virtual machine can determine the corresponding operation type based on the instruction identifier in the bytecode instruction, and then call the corresponding processing function to interpret and execute the bytecode instruction. Different types of bytecode instructions are represented by different instruction identifiers, and different types of bytecode instructions correspond to different processing functions.

[0179] In step 408 , the processing function extracts the instruction field in the bytecode instruction and calls the corresponding machine instruction according to the instruction field to perform the corresponding operation.

[0180] When processing bytecode instructions, the processing function can extract instruction fields from the bytecode instructions, such as the field indicating the operation, the field indicating the storage address of the input tensor, and the field indicating the storage address of the output tensor. In this way, after extracting these instruction fields, the processing function can determine the type of operation indicated by the bytecode instruction and the storage address of the tensor data to be operated on, and then call the corresponding machine instruction to perform the corresponding operation on the tensor.

[0181] The above details the process of generating bytecode instructions and interpreting and executing them through a virtual machine. For ease of understanding, the following details the process of running an AI model based on this process, using specific examples.

[0182] For example, please refer to Figure 7, which is a system architecture diagram for an actual application scenario provided by an embodiment of the present application. As shown in Figure 7, this embodiment is based on the open source MindSpore AI computing framework to implement automatic operator fusion and execution of dynamic computation graphs to accelerate the operating performance of dynamic computation graph scenarios. Since operator fusion requires the dynamic generation of different kernel functions, and the dynamic computation graph can only know the specific shape of the tensor at runtime, this embodiment can be used to generate and execute fusion operator kernel functions. As a special scenario, this embodiment can also solve the problem of automatic operator compilation of dynamic computation graphs.

[0183] In the system architecture shown in Figure 7, the MindSpore AI computing framework runs on a server whose hardware includes a processor, memory, and disk. Virtual machines run on AI chips (such as GPUs), which have dedicated AI processor cores and high-bandwidth memory (HBM).

[0184] The MindSpore AI computing framework primarily includes three modules: Python graph construction, computational graph compilation, and bytecode compilation. The virtual machine primarily includes three modules: bytecode loading, bytecode instruction distribution, and bytecode instruction execution. The following describes how the MindSpore AI computing framework and virtual machine work together to run AI models.

[0185] Please refer to Figure 8A, which is a schematic diagram of the execution flow of an AI computing framework provided in an embodiment of the present application. As shown in Figure 8A, the execution flow of the AI ​​computing framework includes the following steps 801-806.

[0186] Step 801: Build a corresponding computational graph based on the AI ​​model.

[0187] In this embodiment, step 801 is performed by the Python graphing module in the MindSpore AI computing framework. It primarily constructs a corresponding computational graph based on the AI ​​model described by the MindSpore Python interface. In other words, during the AI ​​model's execution, the MindSpore Python interface can be called to execute the AI ​​model, thereby triggering the execution of the series of steps described in this embodiment.

[0188] Step 802: Replace the composite operator in the computation graph with multiple element operators.

[0189] In this embodiment, steps 802 and 803 are performed by the computation graph compilation module in the MindSpore AI computing framework. After obtaining the computation graph, since the operator nodes in the computation graph generally indicate composite operators (such as convolution operators or pooling operators), the composite operators in the computation graph can be replaced with multiple meta-operators to facilitate the subsequent generation of bytecode instructions based on the meta-operator granularity.

[0190] Step 803: merge multiple adjacent fusionable meta-operators into a single meta-operator graph.

[0191] After replacing composite operators in a computation graph with multiple meta-operators, the entire computation graph is actually composed of a large number of meta-operators. To facilitate the subsequent generation of bytecode instructions, multiple adjacent, fusible meta-operators in the computation graph can be merged into a single meta-operator graph. By partitioning the fusible meta-operators in the computation graph, the computation graph can be split into multiple meta-operator graphs, each containing multiple meta-operators. The meta-operators included in each meta-operator graph can be determined based on the actual hardware environment of the virtual machine that interprets and executes the bytecode instructions.

[0192] Specifically, if some meta-operators with large differences in output tensor shapes are fused into the same meta-operator graph, it may make it difficult to implement tensor splitting, which in turn may cause the processor core to experience insufficient local memory when processing bytecode instructions. Therefore, when fusing meta-operators, the differences in the shapes of the output tensors corresponding to the meta-operators need to be considered. In addition, since the local memory of the processor core is limited, fusing too many meta-operators into a meta-operator graph can also easily lead to insufficient local memory. Therefore, when fusing meta-operators, the size of the local memory of the processor core needs to be considered to avoid fusing too many meta-operators. While meeting the local memory constraints, more meta-operators can be fused into the same meta-operator graph as much as possible, so that the intermediate data obtained by processing multiple meta-operators in the same meta-operator graph can be stored in the local memory of the device with faster read and write speeds, without the need to frequently read and write data from the global internal memory of the device with slower read and write speeds, thereby improving the processing efficiency of the operator.

[0193] Step 804: Bytecode compilation is performed on the meta-operator graph to obtain bytecode instructions.

[0194] After obtaining the meta-operator graph, the meta-operator graph can be converted into a meta-operator instruction graph based on the actual input tensor. After performing tensor shape segmentation on the meta-operator instruction graph, the corresponding bytecode instructions are generated. Specifically, the process of generating bytecode instructions based on the meta-operator graph can be referred to the description of the above embodiment and will not be repeated here.

[0195] Step 805 : Based on the data length of the bytecode instruction, apply for HBM space and move the bytecode instruction to the HBM space.

[0196] After generating bytecode instructions, you can request the corresponding HBM space from the AI ​​chip based on the data length of the generated bytecode instructions, so that the generated bytecode instructions can be moved to the AI ​​chip for interpretation and execution. After requesting HBM space, you can move the bytecode instructions to the HBM space.

[0197] Step 806 , starting the virtual machine to execute the bytecode instruction using the HBM space address and the data length of the bytecode instruction as parameters.

[0198] After the bytecode instructions are moved, the HBM space address and the data length of the bytecode instructions can be used as input parameters to start the virtual machine on the AI ​​chip through the AI ​​chip's runtime driver interface to execute the bytecode instructions. When starting the virtual machine to execute the bytecode instructions, the number of processor cores on the AI ​​chip that process the bytecode instructions in parallel can be configured to be the same as the number of tensor shape partitions in the bytecode instructions, thereby fully utilizing the parallel processing capabilities of the multiple processor cores on the AI ​​chip.

[0199] In addition, after the virtual machine completes executing the bytecode instruction, it can release the HBM space corresponding to the bytecode instruction.

[0200] Please refer to Figure 8B, which is a schematic diagram of a flow chart of a virtual machine interpreting and executing bytecode instructions according to an embodiment of the present application. As shown in Figure 8B, the flow of the virtual machine interpreting and executing bytecode instructions includes the following steps 807-812.

[0201] Step 807 : Load the bytecode instruction from the HBM to the local memory based on the HBM space address and the data length of the bytecode instruction.

[0202] During the process of the virtual machine executing bytecode instructions, since the hardware that actually processes the bytecode instructions is the multiple processor cores in the AI ​​chip, the bytecode instructions can be loaded from the HBM to the local memory of the processor core based on the HBM space address and the data length of the bytecode instructions, so that each processor core can process the corresponding bytecode instructions.

[0203] Step 808: Point the instruction cursor to the first address of the bytecode instruction in the local memory.

[0204] When a processor core processes bytecode instructions, it can execute them one by one based on the instruction cursor. That is, when the bytecode instruction is first interpreted and executed, the instruction cursor is first pointed to the first address of the bytecode instruction in the local memory, thereby interpreting and executing the first bytecode instruction.

[0205] Step 809: read the instruction identifier from the bytecode instruction pointed to by the current instruction cursor, and call the corresponding processing function.

[0206] Since each bytecode instruction includes a unique instruction identifier, and the instruction identifier is determined based on the type of the meta-operator corresponding to the bytecode instruction, the corresponding processing function can be called based on the instruction identifier to process the bytecode instruction.

[0207] In step 810, the processing function reads other instruction fields from the bytecode instruction and calls machine instructions to complete the operation.

[0208] The process of the processing function processing the bytecode instruction can refer to the above embodiment and will not be repeated here.

[0209] Step 811: Determine whether the last bytecode instruction processed is the last instruction.

[0210] After processing a bytecode instruction, it can be determined whether the last bytecode instruction processed is the last instruction. If the last bytecode instruction processed is the last instruction, the bytecode instruction processing flow is exited.

[0211] Step 812: Move the instruction cursor to the next bytecode instruction.

[0212] If the last bytecode instruction processed is not the last instruction, the instruction cursor is pointed to the next bytecode instruction and the processing continues to the next bytecode instruction.

[0213] For example, please refer to Figure 9, which is a system architecture diagram for another practical application scenario provided by an embodiment of the present application. The system architecture shown in Figure 9 is the same as the system architecture shown in Figure 7 in terms of hardware. The difference is that the system architecture shown in Figure 7 generates a calculation graph based on an AI model and performs subsequent processing on the calculation graph; the system architecture shown in Figure 9 triggers the construction of the meta-operator graph corresponding to the operator through the operator call interface during the operation of the AI ​​model and performs subsequent processing on the meta-operator graph.

[0214] Specifically, the AI ​​model code is executed using the dynamic language Python, and the operator interface is dynamically called at runtime to dispatch and execute the operator. Since the operator dispatch is performed dynamically during the execution of the AI ​​model code, and even for the same operator call, the shape of the input tensor may be different each time, existing automatic operator compilation technology cannot accept the full-process operator compilation of the binary file every time. However, since this solution uses bytecode compilation, it can actually execute the same processing function on the virtual machine in this scenario, and only the bytecode instructions need to be regenerated each time it is run.

[0215] The MindSpore AI computing framework primarily includes four modules: Python operator invocation, meta-operator graph expansion, meta-operator graph splitting, and bytecode compilation. The virtual machine primarily includes three modules: bytecode loading, bytecode instruction distribution, and bytecode instruction execution. The steps performed on the virtual machine are identical to those in the embodiment corresponding to Figure 7; please refer to the aforementioned embodiment for details. The following describes how the MindSpore AI computing framework generates bytecode instructions based on issued operators.

[0216] (1) Python operator call: The code for the AI ​​model is implemented in Python. For tensor calculations during AI model execution, operator calls are triggered by calling the corresponding Python interface of MindSpore.

[0217] (2) Meta-operator graph expansion: If the Python operator calls a composite operator, the composite operator can be expanded into the corresponding meta-operator subgraph so that it can be directly recognized and processed during the bytecode compilation stage.

[0218] (3) Meta-operator graph splitting: For some composite operators, since the structure of the expanded meta-operator subgraph is complex and cannot be executed through a single virtual machine kernel function call, the meta-operator graph corresponding to the composite operator can be split into multiple meta-operator graphs.

[0219] (4) Bytecode compilation: The obtained one or more meta-operator graphs are byte-compiled and the virtual machine kernel function is started and executed according to the corresponding input tensors.

[0220] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0221] Please refer to Figure 10, which is a structural diagram of a data operation device of a model provided in an embodiment of the present application. As shown in Figure 10, the data operation device of the model includes: an acquisition module 1001, which is used to obtain a calculation graph and the shape of an input tensor, the calculation graph is used to indicate the operation in the artificial intelligence AI model, and the input tensor is used to represent the input data corresponding to the operation; a processing module 1002, which is used to generate at least one bytecode instruction based on the calculation graph and the shape of the input tensor, and the at least one bytecode instruction is used to indicate the operation performed on the input tensor; the processing module 1002 is also used to execute at least one bytecode instruction through a virtual machine, and the virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to interpret the at least one bytecode instruction and call a machine instruction based on the interpretation result to perform the operation on the input tensor.

[0222] In one possible implementation, at least one bytecode instruction is further used to instruct the input tensor to be divided into multiple parts to perform calculation operations separately; the processing module 1002 is specifically used to process at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

[0223] In one possible implementation, at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

[0224] In a possible implementation, the number of splits is greater than or equal to the number of virtual machine instances.

[0225] In one possible implementation, the processing module 1002 is further used to: generate a first bytecode instruction and a second bytecode instruction based on the computation graph and the input tensor, the first bytecode instruction being used to instruct the input tensor to be moved from the global memory to the local memory, and the second bytecode instruction being used to instruct the output tensor obtained by processing the input tensor to be moved from the local memory to the global memory; and execute the first bytecode instruction, at least one bytecode instruction, and the second bytecode instruction in sequence through the virtual machine.

[0226] In one possible implementation, after generating at least one bytecode instruction, the processing module 1002 is further used to: move the at least one bytecode instruction to a memory space accessed by AI hardware, where the AI ​​hardware is used to run a virtual machine.

[0227] In one possible implementation, the processing module 1002 is specifically used to: obtain a first meta-operator graph based on the computational graph conversion, the first meta-operator graph including multiple meta-operators, wherein the multiple meta-operators are used to indicate basic calculation operations; generate at least one bytecode instruction based on the shape of the input tensor and the meta-operator graph, and the multiple meta-operators correspond to the at least one bytecode instruction.

[0228] In one possible implementation, the processing module 1002 is further used to: convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; divide the converted computation graph into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include a first meta-operator graph, and the plurality of meta-operator graphs each include a plurality of meta-operators.

[0229] In one possible implementation, the acquisition module 1001 is also used to obtain an operator call instruction, which is used to indicate the execution of an operation corresponding to a target operator, and the operator call instruction includes an input tensor; the processing module 1002 is also used to generate a computational graph based on the target operator indicated in the operator call instruction, wherein the computational graph includes multiple meta-operators for representing the target operator, and the multiple meta-operators are all used to indicate basic computational operations.

[0230] In a possible implementation, at least one bytecode instruction includes an instruction identifier, and the virtual machine is configured to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

[0231] In one possible implementation, at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate a data type of an input tensor.

[0232] In one possible implementation, the at least one bytecode instruction further indicates a storage address of an input tensor and a storage address of an output tensor, where the output tensor is a tensor obtained after performing the operation on the input tensor. Furthermore, the storage address of the input tensor and the storage address of the output tensor are both addresses in local memory.

[0233] Please refer to Figure 11, which is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1100 can be specifically manifested as a server, a personal computer, a smart phone, etc., which is not limited here. Specifically, the execution device 1100 includes: a receiver 1101, a transmitter 1102, a processor 1103 and a memory 1104 (wherein the number of processors 1103 in the execution device 1100 can be one or more, and Figure 11 takes one processor as an example), wherein the processor 1103 may include an application processor 11031 and a communication processor 11032. In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103 and the memory 1104 may be connected via a bus or other means.

[0234] The memory 1104 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1103. A portion of the memory 1104 may also include non-volatile random access memory (NVRAM). The memory 1104 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0235] Processor 1103 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all of these buses are referred to as a bus system in the figure.

[0236] The method disclosed in the above embodiment of the present application can be applied to the processor 1103, or implemented by the processor 1103. The processor 1103 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 1103 or the instructions in the form of software. The above-mentioned processor 1103 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0237] The processor 1103 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1104, and the processor 1103 reads the information in the memory 1104 and completes the steps of the above method in combination with its hardware.

[0238] Receiver 1101 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1102 can be used to output digital or character information through the first interface. Transmitter 1102 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1102 can also include a display device such as a display screen.

[0239] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the method for determining the model structure described in the above embodiment, or so that the chip in the training device executes the method for determining the model structure described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0240] Specifically, see Figure 12, which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1200. NPU 1200 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is the arithmetic circuit 1203, which is controlled by the controller 1204 to extract matrix data from the memory and perform multiplication operations.

[0241] In some implementations, the arithmetic circuit 1203 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1203 is a two-dimensional systolic array. The arithmetic circuit 1203 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1203 is a general-purpose matrix processor.

[0242] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1201 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1208.

[0243] Unified memory 1206 is used to store input and output data. Weight data is directly transferred to weight memory 1202 through the Direct Memory Access Controller (DMAC) 1205. Input data is also transferred to unified memory 1206 through the DMAC.

[0244] BIU stands for Bus Interface Unit, i.e., bus interface unit 1210 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1209 .

[0245] The bus interface unit 1210 (BIU) is used for the instruction fetch memory 1209 to obtain instructions from the external memory, and is also used for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0246] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1206 or transfer weight data to the weight memory 1202 or transfer input data to the input memory 1201.

[0247] The vector calculation unit 1207 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1203, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0248] In some implementations, the vector calculation unit 1207 can store the processed output vector in the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function or a nonlinear function to the output of the operation circuit 1203, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1207 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1203, for example, for use in subsequent layers in a neural network.

[0249] An instruction fetch buffer 1209 connected to the controller 1204 is used to store instructions used by the controller 1204;

[0250] Unified memory 1206, input memory 1201, weight memory 1202, and instruction fetch memory 1209 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0251] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0252] Please refer to Figure 13, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the method disclosed in Figure 3 above can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0253] 13 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0254] In one embodiment, computer readable storage medium 1300 is provided using signal bearing medium 1301. Signal bearing medium 1301 may include one or more program instructions 1302 that, when executed by one or more processors, may provide the functionality or portions of the functionality described above with respect to FIG.

[0255] In some examples, the signal bearing medium 1301 may include a computer readable medium 1303 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, and the like.

[0256] In some embodiments, the signal-bearing medium 1301 may include a computer-recordable medium 1304, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1301 may include a communication medium 1305, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1301 may be communicated via a wireless form of the communication medium 1305 (e.g., a wireless communication medium conforming to the IEEE 802.11 standard or other transmission protocol).

[0257] The one or more program instructions 1302 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1302 communicated to the computing device via one or more of computer-readable media 1303, computer-recordable media 1304, and / or communication media 1305.

[0258] It should also be noted that the device embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0259] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0260] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0261] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, training device or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center that includes one or more available media integrations. Available media can be magnetic media, (such as floppy disks, hard disks, tapes), optical media (such as DVDs), or semiconductor media (such as solid-state drives (SSDs)).

Claims

1. A data computing method, characterized in that: include: Obtaining a computational graph and a shape of an input tensor, wherein the computational graph is used to indicate a computing operation in an artificial intelligence (AI) model, and the input tensor is used to represent input data corresponding to the computing operation; Based on the computation graph and the shape of the input tensor, generate at least one bytecode instruction, where the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor; The at least one bytecode instruction is executed by a virtual machine to obtain an output tensor, and a processing function corresponding to the at least one bytecode instruction is configured in the virtual machine, and the processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

2. The method according to claim 1, characterized in that The at least one bytecode instruction is further used to instruct to split the input tensor into a plurality of parts to perform operation operations respectively; The executing the at least one bytecode instruction by a virtual machine includes: The at least one bytecode instruction is processed in parallel by multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

3. The method according to claim 2, characterized in that The at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

4. The method according to claim 3, characterized in that The number of splits is greater than or equal to the number of the multiple virtual machine instances.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Based on the computation graph and the input tensor, generate a first bytecode instruction and a second bytecode instruction, wherein the first bytecode instruction is used to instruct to move the input tensor from the global memory to the local memory, and the second bytecode instruction is used to instruct to move an output tensor obtained by processing the input tensor from the local memory to the global memory; The executing the at least one bytecode instruction by a virtual machine includes: The first bytecode instruction, the at least one bytecode instruction and the second bytecode instruction are executed in sequence by a virtual machine.

6. The method according to any one of claims 1 to 5, characterized in that: After generating at least one bytecode instruction, the method further comprises: The at least one bytecode instruction is moved to a memory space accessed by AI hardware, and the AI ​​hardware is used to run the virtual machine.

7. The method according to any one of claims 1 to 6, characterized in that: The generating at least one bytecode instruction based on the computation graph and the shape of the input tensor comprises: A first meta-operator graph is obtained based on the conversion of the computation graph, wherein the first meta-operator graph includes a plurality of meta-operators, wherein the computation graph is used to indicate part or all of the computing operations in the AI ​​model, and the plurality of meta-operators are used to indicate basic computing operations; Based on the shape of the input tensor and the first meta-operator graph, the at least one bytecode instruction is generated, and the multiple meta-operators correspond to the at least one bytecode instruction.

8. The method according to claim 7, characterized in that The converting the computation graph to obtain a first-element operator graph includes: Convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; The converted computation graph is divided into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include the first meta-operator graph, and each of the plurality of meta-operator graphs includes a plurality of meta-operators.

9. The method according to any one of claims 1 to 6, characterized in that: The obtaining of the computation graph includes: Obtain an operator call instruction, where the operator call instruction is used to instruct execution of a calculation operation corresponding to a target operator, and the operator call instruction includes the input tensor; The computation graph is generated based on the target operator indicated in the operator call instruction, wherein the computation graph includes a plurality of meta-operators for representing the target operator, and the plurality of meta-operators are all used to indicate basic calculation operations.

10. The method according to any one of claims 1 to 9, characterized in that: The at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

11. The method according to any one of claims 1 to 10, characterized in that: The at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate the data type of the input tensor.

12. The method according to claim 5, characterized in that The at least one bytecode instruction is also used to indicate a storage address of the input tensor and a storage address of the output tensor, and the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

13. A data computing device for a model, characterized in that: include: An acquisition module, used to acquire the shape of a computational graph and an input tensor, wherein the computational graph is used to indicate a computing operation in an artificial intelligence AI model, and the input tensor is used to represent input data corresponding to the computing operation; A processing module, configured to generate at least one bytecode instruction based on the computation graph and the shape of the input tensor, wherein the at least one bytecode instruction is used to indicate an operation to be performed on the input tensor; The processing module is also used to execute the at least one bytecode instruction through a virtual machine to obtain an output tensor. The virtual machine is configured with a processing function corresponding to the at least one bytecode instruction, and the processing function is used to call a machine instruction corresponding to the at least one bytecode instruction to perform an operation on the input tensor.

14. The device according to claim 13, characterized in that The at least one bytecode instruction is further used to instruct to split the input tensor into a plurality of parts to perform operation operations respectively; The processing module is specifically used to process the at least one bytecode instruction in parallel through multiple virtual machine instances located in different processor cores, and different virtual machine instances among the multiple virtual machine instances are used to process different data in the input tensor.

15. The device according to claim 14, characterized in that The at least one bytecode instruction includes a split number, where the split number is used to indicate the number of splits of the input tensor.

16. The device according to claim 15, characterized in that The number of divisions is greater than or equal to the number of the plurality of processor cores.

17. The device according to any one of claims 14 to 16, characterized in that: The processing module is further used for: Based on the computation graph and the input tensor, generate a first bytecode instruction and a second bytecode instruction, wherein the first bytecode instruction is used to instruct to move the input tensor from the global memory to the local memory, and the second bytecode instruction is used to instruct to move an output tensor obtained by processing the input tensor from the local memory to the global memory; The first bytecode instruction, the at least one bytecode instruction and the second bytecode instruction are executed in sequence by a virtual machine.

18. The device according to any one of claims 13 to 17, characterized in that: After generating at least one bytecode instruction, the processing module is further configured to: The at least one bytecode instruction is moved to a memory space accessed by AI hardware, and the AI ​​hardware is used to run the virtual machine.

19. The device according to any one of claims 13 to 18, characterized in that: The processing module is specifically used for: A first meta-operator graph is obtained based on the conversion of the computation graph, wherein the first meta-operator graph includes a plurality of meta-operators, wherein the computation graph is used to indicate part or all of the computing operations in the AI ​​model, and the plurality of meta-operators are used to indicate basic computing operations; Based on the shape of the input tensor and the first meta-operator graph, the at least one bytecode instruction is generated, and the multiple meta-operators correspond to the at least one bytecode instruction.

20. The device according to claim 19, characterized in that The processing module is further used for: Convert each operator in the computation graph into one or more meta-operators to obtain a converted computation graph; The converted computation graph is divided into a plurality of continuous meta-operator graphs, wherein the plurality of meta-operator graphs include the first meta-operator graph, and each of the plurality of meta-operator graphs includes a plurality of meta-operators.

21. The device according to any one of claims 13 to 18, characterized in that: The acquisition module is further used to acquire an operator call instruction, where the operator call instruction is used to instruct to execute a calculation operation corresponding to a target operator, and the operator call instruction includes the input tensor; The processing module is further used to generate the calculation graph based on the target operator indicated in the operator call instruction, wherein the calculation graph includes multiple meta-operators used to represent the target operator, and the multiple meta-operators are all used to indicate basic calculation operations.

22. The device according to any one of claims 13 to 21, characterized in that The at least one bytecode instruction includes an instruction identifier, and the virtual machine is used to call a processing function corresponding to the at least one bytecode instruction based on the instruction identifier to process the at least one bytecode instruction.

23. The device according to any one of claims 13 to 22, characterized in that: The at least one bytecode instruction includes a data type identifier, where the data type identifier is used to indicate the data type of the input tensor.

24. The device according to claim 17, characterized in that The at least one bytecode instruction is also used to indicate a storage address of the input tensor and a storage address of the output tensor, and the storage address of the input tensor and the storage address of the output tensor are both addresses in the local memory.

25. A data computing device for a model, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 12.

26. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 12.

27. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data operation method of model and related device

    CN120104243A

  • Object interception method and device, medium and electronic equipment

    CN111913741A

  • Tensor processing method and processing system based on parallel branches and tensor segmentation

    CN113485837A

  • Dynamic graph execution method and device for neural network calculation

    CN114461351A

  • Compiling method and related device

    CN115437637A

Cited By

  • Automatic parallelization method and device for hybrid expert model, equipment and medium

    CN119806829A