An operator compilation method and device

By deploying the executable files of the operator with the model files of the AI model separately, the problems of excessive and repeated compilation of the AI model files are solved, and resource saving and execution speed are improved.

CN117992050BActive Publication Date: 2025-07-18HUAWEI TECH CO LTD

Patent Information

Application Number
CN202211325387.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2025-07-18
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

In the prior art, the model files of the AI model are too large and occupy a lot of memory. When the operator is updated, it is necessary to repeatedly compile the entire AI model and multiple other operators, resulting in low compilation efficiency and waste of resources.

Method used

Deploy the executable file of the operator separately from the model files of the AI model, and obtain and send the executable file of the operator and the binary file of the AI model to the terminal device through the host device. The terminal device parses and executes it to avoid duplicate compilation and waste of resources.

Benefits of technology

It effectively reduces the model file size of the AI model, facilitates the update management of operators and AI models, reduces resource waste, and improves the execution speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117992050B_ABST
    Figure CN117992050B_ABST
Patent Text Reader

Abstract

The present application provides an operator compilation method and apparatus. The method includes: a first device obtains a first execution file and sends the first execution file to a second device, where the first execution file is a binary file obtained by compiling a first operator; and, the first device obtains a model file and sends the model file to the second device, where the model file is a binary file obtained by compiling an AI model, and the execution logic of the first execution file is included in the model file. In this way, the first device sets the executable file of the operator outside the model file of the AI model, which can effectively reduce the size of the model file of the AI model, facilitate the update management of the operator and the AI model, avoid a large amount of repeated compilation, reduce resource waste, and help improve the execution speed of the AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to an operator compilation method and apparatus. Background Art

[0002] Artificial intelligence is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. It was first proposed by John McCarthy in 1956. The purpose of artificial intelligence is to enable machines to think like humans and have intelligence. Today, the connotation of artificial intelligence has been greatly expanded. Driven by factors such as deep reinforcement learning and big data, the progress of convolutional neural network models and parameter training techniques, the breakthrough of hardware computing power beyond Moore's law providing considerable computing power, and the Internet plus massive large datasets, it has been applied in many fields such as text classification, sequence labeling, neural machine translation, relation extraction, event extraction, image classification, visual reasoning, semantic segmentation, etc.

[0003] Developers or researchers can design different artificial intelligence (AI) models, such as face recognition models, speech recognition models, etc. Before each execution of the AI model, it is necessary to compile the AI model and the operators corresponding to the AI model to form an executable binary file. Since an AI model corresponds to multiple operators, the model file obtained after compiling the AI model includes binary files of multiple operators, resulting in the model file being too large and occupying more memory. Moreover, when one operator is updated, it is necessary to compile the entire AI model and multiple other operators in the AI model, resulting in a large amount of repeated compilation and low compilation efficiency.

[0004] Therefore, how to reduce the size of the model file of the AI model, avoid repeated compilation, and reduce unnecessary resource waste is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] This application provides an operator compilation method and apparatus for reducing the size of the model file of the AI model, avoiding repeated compilation, reducing unnecessary resource waste, and helping to improve the speed of executing the model.

[0006] In a first aspect, the present application provides an operator compilation method, which includes: a first device obtains a first execution file and sends the first execution file to a second device, where the first execution file is a binary file obtained by compiling a first operator; the first device obtains a model file and sends the model file to the second device, where the model file is a binary file obtained by compiling an artificial intelligence (AI) model, and the execution logic of the first execution file is included in the model file.

[0007] Among them, the first device can be understood as a host device for obtaining and deploying the model file and the first execution file; the second device can be understood as a terminal device for deploying and executing the model file and the first execution file. The process of the first device and the second device communicating and interacting with the model file and the first execution file can be understood as the process of deploying the model file and the first execution file.

[0008] In an embodiment of the present application, the first operator may include one or more operators.

[0009] In this method, the first device deploys the executable file of the operator and the model file of the AI model to the second device respectively. In this way, the first device sets the executable file of the operator outside the model file of the AI model, which can effectively reduce the size of the model file of the AI model, facilitate the update management of the operator and the AI model, avoid a large amount of repeated compilation, reduce resource waste, and help improve the execution speed of the AI model.

[0010] In a possible implementation manner, the first device receives a user instruction for instructing to put the first execution file into the model file; correspondingly, the first device obtains the model file, including: the first device responds to the user instruction and puts the first execution file into the model file. In this implementation manner, the user can flexibly select whether to put the executable file corresponding to the first operator (i.e., the first execution file) into the model file.

[0011] In a possible implementation manner, the first device receives a second operator input by the user, obtains a second execution file, where the second execution file is a binary file obtained by compiling the second operator; the first device puts the second execution file into the model file. It can be understood that the second operator is a user-defined operator, and the second operator may include one or more operators. That is to say, the first device can receive a user-defined operator, obtain the executable file corresponding to the operator, and put the executable file into the model file. In this way, the user can customize the model file according to actual needs to adapt to its business requirements.

[0012] In a possible implementation, the first device may also receive first request information from the second device, where the first request information is used to request a first executable file. Before the first device (i.e., the host device) deploys the first executable file to the second device (i.e., the terminal device), the terminal device may actively request the first executable file from the host device. For example, the second device may send the first request information to the second device when executing a model file or when the first executable file is not stored in the cache of the second device.

[0013] In a possible implementation, the above method further includes: the first device receives a service instruction from the second device, where the service instruction is used to indicate service parameters of a first operator; the first device updates the first executable file according to the service parameters of the first operator. In this way, the first device updates the executable file of the first operator (i.e., the first executable file) according to the service parameters of the first operator, so that the executable file of the first operator is adapted to the service parameters required by the user, and thus the performance of the first executable file can be effectively improved, enabling related services to be executed faster.

[0014] Further, in a possible implementation, after the first device updates the first executable file, the method further includes: the first device receives second request information from the second device, where the second request information is used to request the updated first executable file; the first device sends the updated first executable file to the second device. The second device may actively request the updated first executable file from the first device.

[0015] In another possible implementation, the method further includes: the first device periodically updates the first executable file and sends the updated first executable file to the second device. In this way, the first device can periodically update the first executable file and send the updated first executable file to the second device, so that the executable file of the first operator in the second device can be updated in a timely manner.

[0016] In a second aspect, the present application provides an operator compilation method, which includes: the second device receives a first executable file from the first device, where the first executable file is a binary file obtained by compiling a first operator; the second device receives a model file from the first device, where the model file is a binary file obtained by compiling an AI model.

[0017] In a possible implementation, the second device parses the model file to obtain the execution logic of the first executable file; the second device executes the first executable file according to the execution logic of the first executable file.

[0018] In a possible implementation, before the second device receives the first execution file from the first device, the method further includes: when the first execution file is not stored in the cache of the second device, the second device sends first request information to the first device, and the first request information is used to request the first execution file.

[0019] In a possible implementation, the second device may send a service instruction to the first device, and the service instruction is used to indicate the service parameters of the first operator, and the service parameters of the first operator are used to update the first execution file.

[0020] In a possible implementation, the method further includes: the second device sends second request information to the first device, and the second request information is used to request the updated first execution file; the second device receives the updated first execution file from the first device.

[0021] In a possible implementation, the method further includes: the second device periodically receives the updated first execution file from the first device.

[0022] In a third aspect, the present application provides an operator compilation device, and the device can be applied to the first device.

[0023] As an example, the device includes:

[0024] A processing module, configured to obtain a first execution file, where the first execution file is a binary file obtained by compiling a first operator;

[0025] A communication module, configured to send the first execution file to the second device;

[0026] The processing module is further configured to obtain a model file, where the model file is a binary file obtained by compiling an AI model, and the execution logic of the first execution file is included in the model file;

[0027] The communication module is further configured to send the model file to the second device.

[0028] In a possible implementation, the communication module is further configured to receive a user instruction, and the user instruction is used to indicate putting the first execution file into the model file; specifically, the processing module is configured to: in response to the user instruction, put the first execution file into the model file.

[0029] In a possible implementation, the communication module is further configured to receive a second operator input by the user; the processing module is further configured to obtain a second execution file, where the second execution file is a binary file obtained by compiling the second operator; specifically, the processing module is configured to: put the second execution file into the model file.

[0030] In a possible implementation, the communication module is further configured to receive first request information from a second device, where the first request information is used to request a first execution file.

[0031] In a possible implementation, the communication module is further configured to receive a service instruction from a second device, where the service instruction is used to indicate service parameters of a first operator; the processing module is further configured to update the first execution file according to the service parameters of the first operator.

[0032] In a possible implementation, the communication module is further configured to receive second request information from a second device, where the second request information is used to request the updated first execution file; the communication module is further configured to send the updated first execution file to the second device.

[0033] In a possible implementation, the processing module is further configured to periodically update the first execution file; the communication module is further configured to send the updated first execution file to the second device.

[0034] Fourthly, the present application provides another operator compilation device, which can be applied to a second device.

[0035] As an example, the device includes:

[0036] A communication module, configured to receive a first execution file from a first device, where the first execution file is a binary file obtained by compiling a first operator;

[0037] The communication module is further configured to receive a model file from the first device, where the model file is a binary file obtained by compiling an AI model.

[0038] Further, the device further includes a processing module, where the processing module is configured to parse the model file to obtain the execution logic of the first execution file; the processing module is further configured to execute the first execution file according to the execution logic of the first execution file.

[0039] In a possible implementation, the device further includes a storage module, and the communication module is further configured to: when the first execution file is not stored in the storage module, send first request information to the first device, where the first request information is used to request the first execution file.

[0040] In a possible implementation, the communication module is further configured to send a service instruction to the first device, where the service instruction is used to indicate service parameters of the first operator, and the service parameters of the first operator are used to update the first execution file.

[0041] In a possible implementation, the communication module is further configured to send second request information to the first device, where the second request information is used to request an updated first execution file; the communication module is further configured to receive the updated first execution file from the second device.

[0042] In a possible implementation, the communication module is further configured to periodically receive the updated first execution file from the first device.

[0043] In a fifth aspect, the present application provides a computing device, including a processor and a communication interface. The communication interface is configured to receive signals from other computing devices outside the computing device and transmit them to the processor, or send signals from the processor to other computing devices outside the computing device. The processor is configured to implement the method according to the first aspect or any method in the first aspect or the method according to the second aspect or any method in the second aspect through logic circuits or by executing code instructions.

[0044] In a sixth aspect, a computing device includes a processor and a memory. The memory is configured to store program code, and the processor is configured to call the program code to execute the method according to the first aspect or any method in the first aspect or the method according to the second aspect or any method in the second aspect.

[0045] In a seventh aspect, the present application provides a computer-readable storage medium, in which computer programs or instructions are stored. When the computer programs or instructions are executed by a computing device, the method according to the first aspect or any method in the first aspect or the method according to the second aspect or any method in the second aspect is implemented.

[0046] In an eighth aspect, the present application provides a computer program product, which includes computer programs or instructions. When the computer programs or instructions are executed by a computing device, the method according to the first aspect or any method in the first aspect or the method according to the second aspect or any method in the second aspect is implemented.

[0047] In a ninth aspect, the present application provides a computing system, including a first device configured to execute the method according to the first aspect or any method in the first aspect, and a second device configured to execute the method according to the second aspect or any method in the second aspect.

[0048] The technical effects that can be achieved by any one of the second to ninth aspects above can be referred to the description of the beneficial effects in the first aspect above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A schematic diagram of a computing graph provided by the present application;

[0050] Figure 2A One of the schematic diagrams of the system architecture involved in model training and inference provided by this application;

[0051] Figure 2B Another schematic diagram of the system architecture involved in model training and inference provided by this application;

[0052] Figure 3 Schematic diagram of the structure of a neural network acceleration engine provided by this application;

[0053] Figure 4 Scenario diagram of an operator compilation method provided by this application;

[0054] Figure 5 One of the flow diagrams of an operator compilation method provided by this application;

[0055] Figure 6 Another flow diagram of an operator compilation method provided by this application;

[0056] Figure 7 Another flow diagram of an operator compilation method provided by this application;

[0057] Figure 8 Flow diagram of the second device executing the model file in the embodiments of this application;

[0058] Figure 9 Flow diagram of updating the first execution file in the embodiments of this application;

[0059] Figure 10 Schematic diagram of the structure of an operator compilation device provided by this application;

[0060] Figure 11 Another schematic diagram of the structure of an operator compilation device provided by this application;

[0061] Figure 12 Schematic diagram of the structure of a computing device provided by this application. Detailed implementation manners

[0062] In order to make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. Among them, in the description of the embodiments of this application, hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0063] For ease of understanding, illustrative descriptions of concepts related to this application are given for reference.

[0064] 1) Computational graphs: By defining an AI model and solving its parameters (which can be referred to as model training), a unique computational logic can be determined. After transformation, this computational logic can be applied to inference calculations (which can also be referred to as model inference or usage). This computational logic can be represented by a graph, which is the computational graph.

[0065] The computational graph is represented as a directed graph, which defines the way data flows, the way data is calculated, and the interdependencies between various calculations, etc. Figure 1 A computational graph provided by this application, the computational graph of the AI model consists of operators (nodes) and edges. Among them, the operator is used to represent the applied mathematical operation (computation), or the starting point of data input (feed in) / the ending point of output (push out), or the ending point of reading / writing persistent variables. The operator is the basic computational unit of the AI model. The edge is used to represent the input / output relationship between operators. The edge can transmit a multi-dimensional data array whose size can be dynamically adjusted. Among them, the multi-dimensional data array whose size can be dynamically adjusted is the tensor. This data structure of the tensor can be used to represent the data in the model, that is, a tensor can correspond to an n-dimensional array or list, where n is an integer greater than or equal to zero. The tensor has two attributes: dimension (shape) and rank. In addition, the tensor can flow between the operators of the computational graph.

[0066] 2) Compilation: Compilation is the process of converting a program written in one programming language (source language) into a program in another language (target language). Among them, the source language can be the language used by the user to write the target program, and the target language can be the language used by the device on which the user hopes to run the target program. For example, compilation can convert the high-level language used to write the source program into a binary language that can be recognized by a machine (such as a computer, an actuator, etc.) so that the machine can recognize and execute it. In the embodiments of this application, it involves the compilation of operators and the compilation of the AI model.

[0067] It should be noted that in the embodiments of this application, the operator can also be referred to as a node, a computational task, an operation (operator, OP), an operation layer, etc., and the data dimension can also be referred to as a dimension, a shape, etc. In addition, the AI model described in the embodiments of this application can be a deep learning model, a neural network model, etc.

[0068] It should be understood that in the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single item (s) or plural items (s). For example, at least one (item) of a, b, or c may represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c may be single or multiple.

[0069] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0070] Figure 2A FIG. 7 is one of the schematic diagrams of the system architecture involved in model training and inference provided exemplarily for the present application. This system architecture can be used to compile an AI model to obtain a corresponding model file. It should be understood that this model file is a binary file.

[0071] In Figure 2A , this system architecture includes a graph compiler 210 (which can also be called a model compiler), an executor 220, and an operator compiler 230. Exemplarily, a computational graph as shown in Figure 1 can be input into the graph compiler 210. Specifically, the graph compiler 210 includes a graph compilation module 211, a graph optimization module 212, and a graph loading module 213. Among them, the graph compilation module 211 converts the computational graph into an intermediate representation (IR) to form a model file of the AI model; and the graph optimization module 212 performs target-independent optimizations on the intermediate representation, such as computational graph fusion, subgraph splitting, etc., and then performs target-related optimizations, such as optimizations related to technologies such as single instruction multiple data (SIMD), tiling, unrolling, and vectorization. The graph loading module 213 loads the optimized intermediate representation into the operator compiler 230, and the operator compiler 230 generates an executable binary file (i.e., the executable file of the operator) in the executor 220.

[0072] Figure 2BFIG. 2 is a schematic diagram of a system architecture involved in model training and inference provided by way of example in this application. The architecture further includes a deployment tool 240. After the graph compiler 201 compiles the computation graph to obtain a model file of the AI model, the deployment tool 240 can install and deploy the model file. In addition, after the operator compiler 230 compiles the operator to obtain an executable file of the operator, the deployment tool 240 can install and deploy the executable file of the operator.

[0073] It should be understood that model training and inference can be applicable to various computing frameworks. Taking the computing framework MindSpore as an example, developers use MindSpore to construct a neural network model and solve the parameters of the neural network model. After obtaining the parameter solution, that is, the trained MindSpore model is obtained. The trained MindSpore model is applied to model inference. The specific use of the computation graph in different computing frameworks is different. In the computation graph corresponding to MindSpore, the nodes used to describe the computation process can be called operators. Equivalently, in MindSpore, the description of a series of operators, the parameters of the operators, and the entire computation logic can be called a computation graph. The computation graph can be represented by a static graph or a dynamic graph. Exemplarily, MindSpore can use a static graph to represent the entire computation process. In the entire neural network construction of MindSpore, the operators form a network structure with different application functions, and the neural network acceleration engine can provide the operator development ability. Developers can write corresponding operators to build various neural network models, and use the code written using the neural network acceleration engine as the input of the operator compiler 230 in the above Figure 2A or Figure 2B above.

[0074] Exemplarily, the structure of the neural network acceleration engine is as shown in Figure 3 FIG. 3, and includes a domain-specific language (DSL) module 301, an engine schedule module 302, an intermediate representation module 303, a compiler pass module 304, and a code generation (codegen) module 305. Specifically, the domain-specific language module 301 is used to provide a writing interface for the operator computation logic. Developers can write the computation process and the scheduling process of the operator through this writing interface. The engine schedule module 302 is used to split the data in the operator according to the scheduling description, specify the data transfer process, and provide the operator fusion and optimization ability. The intermediate representation module 303 is used to generate the intermediate representation. After the operator is processed by the compiler pass module 304, the code generation module 305 generates a temporary file of C-like code. This temporary file is input to the operator compiler 230, and then the operator compiler 230 can generate an executable file of the operator.

[0075] It should be understood that the computing framework serializes the computation graph to obtain a file, which will be parsed, compiled, and executed by the computing framework. Specifically, this file can be parsed and compiled by the host device in the computing framework to obtain a binary file (i.e., the model file), and the model file is executed by the terminal device in the computing framework. Exemplarily, the host device can include, for example, Figure 2B the graph compiler 210 and the operator compiler 230 shown in Figure 2B , and the terminal device can include

[0076] Compilation can also be referred to as graph compilation, which can be divided into dynamic compilation and static compilation. Specifically, dynamic compilation, also known as online compilation, means that the host device compiles the AI model during the execution (model inference) of the program to obtain the model file that runs on the terminal device; static compilation, also known as ahead of time (AOT) compilation, means that the host device compiles the AI model before the program execution to obtain the model file that runs on the terminal device, and then calls the corresponding binary file during runtime. To reduce the time consumption during model inference, graph compilation generally adopts static compilation, that is, before the program execution, the host device compiles the AI model into a model file according to the performance of the terminal device, and then sends the model file to the terminal device, thus avoiding the back-and-forth interaction between the host device and the terminal device. During multiple program executions, the host device needs to compile the AI model for each program execution and send the compiled model file to the terminal device. For example, during face recognition, for each input image, the recognition model needs to be called to recognize the face in the current image. That is, for each image, the host device needs to compile the model file corresponding to the recognition model together with the current image to obtain the binary file executed on the terminal device this time, and send the binary file to the terminal device. The terminal device executes the binary file to obtain the execution result of face recognition. Or during the repeated call of a single operator, the host device needs to compile the single operator for each call and send the compiled binary file to the terminal device. For example, to adjust the brightness of an image by executing a single operator, that is, for each input image, the single operator needs to be compiled together with the current image to obtain the binary file executed on the terminal device this time, and send the binary file to the terminal device. The terminal device executes the binary file to obtain the image with adjusted brightness.

[0077] As can be seen from the above, an AI model corresponds to multiple operators. The model file obtained after compiling the AI model includes the binary files of multiple operators, resulting in the model file being too large and occupying more memory. Moreover, when one of the operators is updated, the host device needs to compile the entire AI model and multiple other operators in the AI model, resulting in a large amount of repeated compilation and low compilation efficiency.

[0078] In view of this, the present application provides an operator compilation method to reduce the size of the model file of the AI model, avoid repeated compilation, reduce unnecessary resource waste, and help improve the speed of executing the model. Since the host device can not compile the operators in the AI model when compiling the AI model, the host device can separately update the executable file of the operator without recompiling the AI model.

[0079] As Figure 4As shown in the figure, the specific implementation of the operator compilation method provided in the embodiments of the present application includes a compilation period, a deployment period, and an execution period. Among them, the compilation period is implemented by the host device. Specifically, the operator compiler in the host device can compile the operators in the AI model to obtain the first execution files of these operators, and the operator compiler (i.e., the graph compiler) in the host device compiles the AI model to obtain the model execution file of the AI model. The deployment period can be implemented by the host device and the terminal device. Specifically, the terminal device receives the model file and the first execution file from the host device, and then the deployment tool in the terminal device stores the model file and the first execution file in the terminal device respectively. The execution period can be implemented by the terminal device. Specifically, the model executor in the terminal device can load the model file and the first execution file and parse and execute them.

[0080] Next, first from the perspectives of the host device and the terminal device, the operator compilation method provided in the embodiments of the present application will be further introduced. Exemplarily, please refer to Figure 5 , Figure 5 which shows a schematic flowchart of the operator compilation method provided in the embodiments of the present application. In Figure 5 , the host device is referred to as the first device, and the terminal device is referred to as the second device. The process includes:

[0081] S501, the first device obtains a first execution file, which is a binary file obtained by compiling a first operator.

[0082] In the embodiments of the present application, the first operator may include one or more operators, and the embodiments of the present application do not make limitations. Among them, the first operator may include operator source code and operator input parameters. The operator input parameters may include one or more of weights, variables, constants (fixed values), and strings. The weights may be understood as the size of the weight space or the shape of a tensor.

[0083] Among them, the first device obtains the first execution file, including but not limited to the following situations:

[0084] Situation 1, the first device has the operator compilation function. The first device can obtain the first operator and compile the first operator to obtain the first execution file.

[0085] Situation 2, the first device does not have the operator compilation function. The first device can request the first execution file from other devices with the operator compilation function.

[0086] S502, the first device sends the first execution file to the second device. Correspondingly, the second device receives the first execution file.

[0087] In S502, the process of the first device sending the first execution file to the second device can be understood as a part of the deployment process of the first execution file. Correspondingly, after receiving the first execution file, the process of the second device storing the first execution file can be understood as another part of the deployment process of the first execution file. Exemplarily, the second device can store the first execution file in the chip of the second device.

[0088] Optionally, before the second device receives the first execution file from the first device, when the second device determines that the first execution file is not stored in its cache, the second device sends a first request message to the first device, and the first request message is used to request the first execution file. Then, when the first device receives the first request message, it can execute S502 and send the first execution file to the second device.

[0089] S503, the first device obtains a model file, which is a binary file obtained by compiling an AI model, and the model file includes the execution logic of the first execution file.

[0090] In the embodiments of the present application, the "execution logic of the first execution file" can be understood as the calculation logic or operation logic of the first operator corresponding to the first execution file.

[0091] Among them, the first device obtaining the first execution file includes but is not limited to the following situations:

[0092] Situation 1, the first device has a model compilation function, and the first device can obtain the source file corresponding to the AI model and compile the source file corresponding to the AI model to obtain a model file.

[0093] Situation 2, the first device does not have a model compilation function, and the first device can request the model file of the AI model from other devices with a model compilation function.

[0094] S504, the first device sends the model file to the second device. Correspondingly, the second device receives the model file.

[0095] In S504, the process of the first device sending the model file to the second device can be understood as a part of the deployment process of the model file. Correspondingly, after receiving the model file, the process of the second device storing the model file can be understood as another part of the deployment process of the model file. Exemplarily, the second device can deploy the model file to the application layer.

[0096] It can be understood that the embodiments of the present application do not limit the sequence of S504 and S502, that is, the sequence of the first device sending the model file and the first execution file to the second device is not limited.

[0097] InFigure 5 In the operator compilation method shown, the first device deploys the executable file of the first operator and the model file of the AI model to the terminal device respectively, which can effectively reduce the size of the model file of the AI model, facilitate the update management of the operator and the AI model, avoid repeated compilation, reduce resource waste, and help improve the execution speed of the AI model.

[0098] As described above, in Figure 5 the deployment process of the model file and the first execution file is mainly introduced. Next, in combination with Figure 6 , the process of the second device executing the model file and the first execution file is introduced. As Figure 6 shown, the operator compilation method provided by the embodiment of the present application further includes:

[0099] S505, the second device parses the model file to obtain the execution logic of the first execution file.

[0100] S506, the second device executes the first execution file according to the execution logic of the first execution file.

[0101] Exemplarily, taking the model file corresponding to the face box detection model as an example, the first operator in the face box detection model includes convolution calculation, and the first execution file can be a binary file obtained by compiling the operator for convolution calculation; correspondingly, the execution logic of the first execution file is the calculation logic of convolution calculation.

[0102] In this way, the model file only includes the execution logic of the first execution file, but does not include the first execution file, which can effectively reduce the size of the model file and reduce the memory occupied by the model file.

[0103] In the embodiment of the present application, the first device can update the first execution file to make the first execution file match the current business of the model file to improve the performance of the first execution file. Exemplarily, the first device can update the first execution file, which can be understood as the first device recompiling the first operator to obtain the updated first execution file. Exemplarily, please refer to Figure 7 , the above method further includes:

[0104] S507, the second device sends a service instruction to the first device, and the service instruction is used to indicate the service parameters of the first operator, and the service parameters of the first operator are used to update the first execution file. Correspondingly, the first device receives the service instruction.

[0105] S508, the first device updates the first execution file based on the service parameters of the first operator.

[0106] Optionally, the service instruction may be determined by the second device according to user input, or may be obtained by the second device after analyzing the historical information of the current service.

[0107] Example 1: Taking the face recognition service as an example of the current service, the first operator in the face recognition model is the operator corresponding to convolutional calculation. The user input information received by the second device indicates that the user needs to perform face recognition on a picture with a tensor shape of 64×64 and a value of value11. Then, the second device generates a corresponding service instruction, and the service parameters related to the convolutional calculation operator indicated by this instruction are associated with the shape and value of the tensor being 64×64 and value11. Subsequently, the first device updates the first execution file based on these service parameters. In this way, the first execution file of the first operator can be flexibly updated according to the user's service requirements, making the updated first execution file more in line with the service requirements of the current service, thereby effectively improving the execution efficiency of the current service.

[0108] Example 2: Taking the face recognition service as an example of the current service, the first operator in the face recognition model is the operator corresponding to convolutional calculation. The second device analyzes the historical input information of the face recognition service and finds that 90% of the shapes and values of the tensors corresponding to the pictures input by the user in the face recognition model are 64×64 and value11. The service instruction generated by the second device is used to indicate that the service parameters related to the convolutional calculation operator are associated with the shape and value of the tensor being 64×64 and value11. Subsequently, the first device updates the first execution file based on these service parameters. In this way, the first device can automatically update the first execution file of the first operator in combination with the parameters analyzed by the second device, making the updated first execution file more in line with the service requirements of the current service, thereby effectively improving the execution efficiency of the current service and, without user operation, effectively enhancing the user experience.

[0109] S509: The first device sends the updated first execution file to the second device. Correspondingly, the second device receives the updated first execution file.

[0110] It can be understood that the first device may execute S509 after executing S508. In other words, the first device can actively send the updated first execution file to the second device.

[0111] Optionally, before receiving the updated first execution file, the second device may further send second request information to the first device, where the second request information is used to request the updated first execution file; correspondingly, the first device receives the second request information and, in response to the second request information, sends the updated first execution file to the second device. In this way, the second device can actively request the updated first execution file from the first device.

[0112] In some possible embodiments, the first device periodically updates the first execution file and periodically sends the updated first execution file to the second device. Correspondingly, the second device periodically receives the updated first execution file from the first device. The period for the first device to update the first execution file may be 1 day, 1 week, 1 month, etc., and the period for the second device to receive the updated first execution file may be 1 day, 1 week, 1 month, etc., which are not limited in the embodiments of the present application.

[0113] In some scenarios, the user still needs to put the operator into the model file to improve the integration degree of the model file. To provide some flexibility for the user to choose, in a possible implementation manner, the first device may further receive a user instruction. If the user instruction is used to indicate putting the first execution file into the model file, the first device may, in response to the user instruction, put the first execution file into the model file. In this way, the user can flexibly choose whether to put the executable file corresponding to the operator (i.e., the first execution file) into the model file.

[0114] In a possible embodiment, the first device may further receive a second operator input by the user and obtain a second execution file, where the second execution file is a binary file compiled from the second operator; and the first device may further put the second execution file into the model file. The second operator may be understood as a user-defined operator, and the second operator may include one or more operators. In this way, the user can customize the operators of the AI model according to business requirements and integrate the executable files corresponding to the user-defined operators into the model file of the AI model, making the AI model more suitable for the user's business requirements and effectively improving the performance of the model file of the AI model. Exemplarily, the second operator may be, for example, an axis elimination operation (i.e., summing the elements in the row direction of a 3×3 matrix to obtain a third-order vector all with 6) and a broadcast operation (i.e., broadcasting the third-order vector back to the original dimension (shape) to obtain a 3×3 matrix all with 6).

[0115] Wherein, the first device obtaining the second execution file includes, but is not limited to, the following situations:

[0116] Situation 1, the first device has an operator compilation function. After receiving the second operator input by the user, the first device compiles the second operator to obtain the second execution file.

[0117] In Case 2, when the first device does not have the operator compilation function, after receiving the second operator input by the user, the first device sends the second operator to other devices with the operator compilation function; the other device compiles the second operator to obtain a second execution file and sends the second execution file to the first device. Correspondingly, the first device receives the second execution file.

[0118] The above introduced the operator compilation method provided by the embodiments of the present application from the perspective of the first device and the second device. Next, the process of the second device executing the model file and the first executable file will be further introduced from the perspective of each component in the second device. Exemplarily, please refer to Figure 8 in Figure 8 the second device includes a model executor and a deployment tool, and the process includes:

[0119] S1, the model executor parses the model file to obtain the execution logic of the first execution file.

[0120] S2, the model executor determines whether the first execution file exists in the model file.

[0121] It should be understood that when the model executor determines that the first execution file does not exist in the model file, it executes S3 and sends request message 1 to the deployment tool.

[0122] S3, the model executor sends request message 1 to the deployment tool, and this request message 1 is used to request the first execution file. Correspondingly, the deployment tool receives request message 1.

[0123] S4, the deployment tool determines that the first execution file exists in the cache of the second device.

[0124] It should be understood that after the deployment tool receives request message 1 and determines that the first execution file exists in the cache of the second device, it executes S5. If after the deployment tool receives request message 1 and determines that the first execution file does not exist in the cache of the second device, the deployment tool can request the first execution file from the first device (such as a server).

[0125] S5, the deployment tool sends the storage path of the first execution file to the model executor. Correspondingly, the model executor receives the storage path of the first execution file.

[0126] Correspondingly, after the model executor receives the storage path of the first execution file, it can load the first execution file and execute S6.

[0127] S6, the model executor executes the first execution file according to the execution logic of the first execution file.

[0128] The process of the first device updating the first executable file will be further introduced from the perspective of components in the first device and the second device. Exemplarily, please refer to Figure 9 , in Figure 9 , the second device includes a model executor and a deployment tool, and the first device includes an operator compiler and a model compiler. The process includes:

[0129] A1. The deployment tool determines the service parameters associated with the first operator.

[0130] Exemplarily, the service parameters associated with the first operator can be a hot operator and a hot tensor, that is, an operator that is frequently used in the current service and the tensor corresponding to the operator.

[0131] A2. The deployment tool generates a service instruction according to the service parameters associated with the first operator. The service instruction is used to update the first executable file corresponding to the first operator and is used to indicate the service parameters associated with the first operator.

[0132] A3. The deployment tool sends the service instruction to the operator compiler. Correspondingly, the operator compiler receives the service instruction.

[0133] A4. The operator compiler compiles the first operator according to the service parameters indicated by the service instruction to obtain an updated first executable file.

[0134] A5. The operator compiler sends the updated first executable file to the deployment tool. Correspondingly, the deployment tool receives the updated first executable file.

[0135] A6. The deployment tool stores the updated first executable file in the cache of the second device.

[0136] For ease of understanding, the operator compilation method provided in the embodiments of the present application will be further described below with specific examples.

[0137] Exemplarily, taking the face bounding box detection model as an example of the AI model, the operators corresponding to the face bounding box detection model are convolution calculation and Local Binary Pattern (LBP); among them, convolution calculation is used to perform convolution calculation on image data, and LBP is used to describe the local features of the image. When the first device receives the compilation instruction from the user, it can compile the source file of the face bounding box detection model to generate a model file of the face bounding box detection model, and send the model file to the second device so that the model file can be deployed to the second device. Moreover, the first device compiles the convolution calculation and LBP to obtain the corresponding first executable file, and sends the first executable file to the second device so that the first executable file can be deployed to the second device. In this way, the executable file of the operator is set outside the model file, which can effectively reduce the size of the model file. Furthermore, when updating the executable file of the operator, there is no need to repeatedly compile the AI model.

[0138] Each of the embodiments described in this application can be an independent solution or can be combined according to the internal logic, and these solutions all fall within the protection scope of this application.

[0139] It can be understood that, in each of the above method embodiments, the methods and operations implemented by the first device can also be implemented by components (such as chips or circuits) available for the first device, and the methods and operations implemented by the second device can also be implemented by components (such as chips or circuits) available for the second device.

[0140] In the above embodiments provided in this application, the methods provided in the embodiments of this application are introduced from the perspective of the interaction between devices. To implement each function in the methods provided in the above embodiments of this application, the first device and the second device may include a hardware structure and / or software module, and implement the above functions in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module. Whether a certain function among the above functions is executed in the form of a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraints of the technical solution.

[0141] The division of modules in the embodiments of this application is illustrative, and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in the various embodiments of this application may be integrated in one processor, may also exist separately physically, or two or more modules may be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0142] Based on the above content and the same concept, Figure 10This is a schematic structural diagram of a possible operator compilation device provided by this application. This operator compilation device can be used to implement the functions of the first device in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. This operator compilation device can be applied to the first device described above.

[0143] Exemplarily, as Figure 10 shown, the operator compilation device 1000 may include:

[0144] A processing module 1001, configured to obtain a first execution file, where the first execution file is a binary file obtained by compiling a first operator;

[0145] A communication module 1002, configured to send the first execution file to a second device;

[0146] The processing module 1001 is further configured to obtain a model file, where the model file is a binary file obtained by compiling an AI model, and the execution logic of the first execution file is included in the model file;

[0147] The communication module 1002 is further configured to send the model file to the second device.

[0148] In a possible implementation manner, the communication module 1002 is further configured to receive a user instruction for instructing to put the first execution file into the model file; specifically, the processing module 1001 is configured to: in response to the user instruction, put the first execution file into the model file.

[0149] In a possible implementation manner, the communication module 1002 is further configured to receive a second operator input by a user; the processing module 1001 is further configured to obtain a second execution file, where the second execution file is a binary file obtained by compiling the second operator; specifically, the processing module 1001 is configured to: put the second execution file into the model file.

[0150] In a possible implementation manner, the communication module 1002 is further configured to receive a first request message from the second device, where the first request message is used to request the first execution file.

[0151] In a possible implementation manner, the communication module 1002 is further configured to receive a service instruction from the second device, where the service instruction is used to indicate service parameters of the first operator; the processing module 1001 is further configured to update the first execution file according to the service parameters of the first operator.

[0152] In a possible implementation manner, the communication module 1002 is further configured to receive a second request message from the second device, where the second request message is used to request the updated first execution file; the communication module 1002 is further configured to send the updated first execution file to the second device.

[0153] In a possible implementation, the processing module 1001 is further configured to periodically update the first execution file; the communication module 1002 is further configured to send the updated first execution file to the second device.

[0154] Based on the above content and the same concept, Figure 11 FIG. is a schematic structural diagram of a possible operator compilation device provided by the present application. The operator compilation device can be used to implement the functions of the second device in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments.

[0155] Exemplarily, as Figure 11 shown, the operator compilation device 1100 includes:

[0156] A communication module 1101, configured to receive a first execution file from a first device, where the first execution file is a binary file obtained by compiling a first operator;

[0157] The communication module 1101 is further configured to receive a model file from the first device, where the model file is a binary file obtained by compiling an AI model.

[0158] Further, the device further includes a processing module 1102, configured to parse the model file to obtain the execution logic of the first execution file; the processing module 1102 is further configured to execute the first execution file according to the execution logic of the first execution file.

[0159] In a possible implementation, the device further includes a storage module 1103, and the communication module 1101 is further configured to: when the first execution file is not stored in the storage module 1103, send a first request message to the first device, where the first request message is used to request the first execution file.

[0160] In a possible implementation, the communication module 1101 is further configured to send a service instruction to the first device, where the service instruction is used to indicate the service parameters of the first operator, and the service parameters of the first operator are used to update the first execution file.

[0161] In a possible implementation, the communication module 1101 is further configured to send a second request message to the first device, where the second request message is used to request the updated first execution file; the communication module 1101 is further configured to receive the updated first execution file from the second device.

[0162] In a possible implementation, the communication module 1101 is further configured to periodically receive the updated first execution file from the first device.

[0163] As Figure 12 shown is the device provided by the embodiment of the present application, Figure 12The device shown can be Figure 10 a hardware circuit implementation of the device shown. The device is applicable to perform the functions of the first device in the above method embodiments.

[0164] Or Figure 12 The device shown can be Figure 11 a hardware circuit implementation of the device shown. The device is applicable to perform the functions of the second device in the above method embodiments.

[0165] For ease of description, Figure 12 only the main components of the device are shown. Figure 12 The device 1200 shown includes at least one processor 1220, which is used to implement any one of the Figure 4 Or Figure 5 methods in.

[0166] The device 1200 may further include at least one memory 1230, which is used to store program instructions and / or data. The memory 1230 is coupled to the processor 1220. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, and is used for information interaction between devices, units or modules. The processor 1220 may cooperate with the memory 1230. The processor 1220 may execute the program instructions stored in the memory 1230. At least one of the at least one memories may be included in the processor.

[0167] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0168] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processing circuit (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0169] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and directrambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0170] The apparatus 1200 may further include a communication interface 1210 for communicating with other devices through a transmission medium, so that the devices in the apparatus 1200 can communicate with other devices. In the embodiments of the present application, the communication interface can be a transceiver, a circuit, a bus, a module, or other types of communication interfaces. In the embodiments of the present application, when the communication interface is a transceiver, the transceiver can include an independent receiver, an independent transmitter; or a transceiver integrating transceiver functions, or an interface circuit.

[0171] The apparatus 1200 may further include a communication line 1240. Among them, the communication interface 1210, the processor 1220, and the memory 1230 may be interconnected through the communication line 1240; the communication line 1240 may be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The communication line 1240 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0172] According to the method provided by the embodiments of the present application, the present application further provides a computer program product, which includes: computer program code. When the computer program code runs on a computer, the computer is caused to execute the method of any one of the above embodiments.

[0173] According to the method provided by the embodiments of the present application, the present application further provides a computer-readable medium, which stores program code. When the program code runs on a computer, the computer is caused to execute the method of any one of the above embodiments.

[0174] According to the method provided by the embodiments of the present application, the present application further provides a system, which includes the foregoing first device and second device.

[0175] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0176] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing in the process Figure One one process or multiple processes and / or blocksFigure One means for functions specified in one or more blocks

[0177] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure One one or more processes and / or blocks Figure One means for functions specified in one or more blocks

[0178] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. An operator compilation method, characterized in that Including: The first device obtains a first execution file, where the first execution file is a binary file obtained by compiling a first operator. The first device sends the first execution file to the second device. The first device obtains a model file, where the model file is a binary file obtained by compiling an artificial intelligence (AI) model. The execution logic of the first execution file is included in the model file, and the first execution file is not included. The first device sends the model file to the second device.

2. The method according to claim 1, characterized in that, The method further includes: The first device receives a user instruction for indicating to place the first execution file into the model file. The first device obtaining the model file includes: in response to the user instruction, the first device places the first execution file into the model file.

3. The method according to claim 1 or 2, characterized in that, The method further includes: The first device receives a second operator input by the user and obtains a second execution file, where the second execution file is a binary file obtained by compiling the second operator. The first device places the second execution file into the model file.

4. The method according to claim 1 or 2, characterized in that, The method further includes: The first device receives a first request message from the second device, where the first request message is for requesting the first execution file.

5. The method according to claim 1 or 2, characterized in that, The method further includes: The first device receives a service instruction from the second device, where the service instruction is for indicating the service parameters of the first operator. The first device updates the first execution file according to the service parameters of the first operator.

6. The method according to claim 1 or 2, characterized in that, The method further includes: The first device receives a second request message from the second device, where the second request message is for requesting the updated first execution file. The first device sends the updated first execution file to the second device.

7. The method according to claim 1 or 2, characterized in that, The method further includes: The first device periodically updates the first execution file and sends the updated first execution file to the second device.

8. An operator compilation method, characterized in that, Including: The second device receives a first execution file from the first device, where the first execution file is a binary file obtained by compiling a first operator. The second device receives a model file from the first device, where the model file is a binary file obtained by compiling an AI model. The execution logic of the first execution file is included in the model file, and the first execution file is not included.

9. The method according to claim 8, wherein The method further includes: The second device parses the model file to obtain the execution logic of the first execution file. The second device executes the first execution file according to the execution logic of the first execution file.

10. The method according to claim 8 or 9, characterized in that, Before the second device receives the first execution file from the first device, the method further includes: The second device determines that the first execution file is not stored in the cache of the second device and sends a first request message to the first device, where the first request message is for requesting the first execution file.

11. The method according to claim 8 or 9, characterized in that The method further includes: The second device sends a service instruction to the first device, where the service instruction is for indicating the service parameters of the first operator, and the service parameters of the first operator are used to update the first execution file.

12. The method according to claim 8 or 9, characterized in that The method further includes: The second device sends second request information to the first device, where the second request information is used to request the updated first execution file; The second device receives the updated first execution file from the first device.

13. The method according to claim 8 or 9, characterized in that, The method further includes: The second device periodically receives the updated first execution file from the first device.

14. An operator compilation device, characterized in that, Applied to the first device, it includes: A processing module, configured to obtain a first execution file, where the first execution file is a binary file compiled from a first operator; A communication module, configured to send the first execution file to the second device; The processing module is further configured to obtain a model file, where the model file is a binary file compiled from an AI model, the execution logic of the first execution file is included in the model file, and the first execution file is not included; The communication module is further configured to send the model file to the second device.

15. The device according to claim 14, wherein The communication module is further configured to receive a user instruction, where the user instruction is used to indicate putting the first execution file into the model file; The processing module is specifically configured to: in response to the user instruction, put the first execution file into the model file.

16. The device according to claim 14 or 15, wherein The communication module is further configured to receive a second operator input by the user; The processing module is further configured to obtain a second execution file, where the second execution file is a binary file compiled from the second operator; The processing module is specifically configured to: put the second execution file into the model file.

17. The device according to claim 14 or 15, wherein The communication module is further configured to receive first request information from the second device, where the first request information is used to request the first execution file.

18. The device according to claim 14 or 15, wherein The communication module is further configured to receive a service instruction from the second device, where the service instruction is used to indicate service parameters of the first operator; The processing module is further configured to update the first execution file according to the service parameters of the first operator.

19. The device according to claim 14 or 15, wherein The communication module is further configured to receive second request information from the second device, where the second request information is used to request the updated first execution file; The communication module is further configured to send the updated first execution file to the second device.

20. The device according to claim 14 or 15, wherein The processing module is further configured to periodically update the first execution file; The communication module is further configured to send the updated first execution file to the second device.

21. An operator compilation device, characterized in that, Applied to the second device, it includes: A communication module, configured to receive a first execution file from a first device, where the first execution file is a binary file compiled from a first operator; The communication module is further configured to receive a model file from the first device. The model file is a binary file obtained by compiling an AI model. The execution logic of the first execution file is included in the model file, and the first execution file is not included.

22. The device according to claim 21, characterized in that, It further includes a processing module. The processing module is configured to parse the model file to obtain the execution logic of the first execution file. The processing module is further configured to execute the first execution file according to the execution logic of the first execution file.

23. The device according to claim 21 or 22, characterized in that, It further includes a storage module. The communication module is further configured to: When the first execution file is not stored in the storage module, send first request information to the first device. The first request information is used to request the first execution file.

24. The device according to claim 21 or 22, characterized in that, The communication module is further configured to: Send a service instruction to the first device. The service instruction is used to indicate the service parameters of the first operator. The service parameters of the first operator are used to update the first execution file.

25. The device according to claim 21 or 22, wherein The communication module is further configured to send second request information to the first device. The second request information is used to request the updated first execution file. The communication module is further configured to receive the updated first execution file from the second device.

26. The device according to claim 21 or 22, wherein The communication module is further configured to periodically receive the updated first execution file from the first device.

27. A computing device, characterized in that, The computing device includes a processor and a memory. The memory is used to store program code, and the processor is used to call the program code to execute the method according to any one of claims 1 to 7 or the method according to any one of claims 8 to 13.

28. A computing system, characterized in that, It includes a first device for executing the method according to any one of claims 1 to 7, and a second device for executing the method according to any one of claims 8 to 13.

Citation Information

Patent Citations

  • Operator compiling method and device

    CN114077432A

Cited By

  • Operator compiling method and apparatus

    EP4607338A1