Model file processing method and device

By running multiple models in the inference engine, obtaining operation information and building target model files, the problem of inefficient multi-model task execution is solved and efficient task execution on different inference engines is achieved.

CN120633849APending Publication Date: 2025-09-12LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510727287.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When multiple models are used to perform a specific task, the inference engine runs inefficiently because each model corresponds to a separate model file.

Method used

By running the first model and the second model in the first inference engine, obtaining operation information, determining the target operator, and constructing the target model file based on the operator execution order relationship, the target model file is output so as to efficiently execute the task on the second inference engine.

Benefits of technology

It improves the task execution efficiency of multiple models on different inference engines, reduces the communication overhead between CPU and GPU, and reduces inference latency and computing cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633849A_ABST
    Figure CN120633849A_ABST
Patent Text Reader

Abstract

The invention provides a model file processing method and device. The method comprises the following steps: running a first model and a second model in a first inference engine to execute a target task; in the execution process of the target task, first operation information for operating the first model and second operation information for operating the second model are obtained respectively; determining a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be executed after being analyzed by a second inference engine; determining a target model file based on the target operators and an execution sequence relationship between the target operators; and outputting the target model file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning technology, and more specifically, to a model file processing method and device. Background Art

[0002] When multiple models are used together to perform a specific task, since each model corresponds to a separate model file, it is inefficient when running through the inference engine. Summary of the Invention

[0003] In view of this, the present disclosure provides a control method and device.

[0004] A first aspect of the present disclosure provides a model file processing method, comprising:

[0005] Running the first model and the second model in the first inference engine to perform the target task;

[0006] During the execution of the target task, first operation information for operating the first model and second operation information for operating the second model are respectively obtained;

[0007] Determining a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by the second inference engine;

[0008] Determine the target model file based on the target operator and the execution order relationship between the target operators;

[0009] Output target model file.

[0010] According to an embodiment of the present disclosure, running the first model and the second model in the first inference engine includes:

[0011] Outputting prediction information through the first model, where the prediction information represents an output prediction result of the second model;

[0012] Verifying the prediction result through the second model information to obtain a verification result;

[0013] Target output information of a second model is determined based on the verification result, wherein the second model has a greater number of parameters than the first model.

[0014] According to an embodiment of the present disclosure, the method further includes:

[0015] Determining an execution order relationship between the first operation information and the second operation information according to an execution process of the target task;

[0016] An execution order relationship between target operators is determined according to an execution order relationship between the first operation information and the second operation information.

[0017] According to an embodiment of the present disclosure, determining a target model file based on a target operator and an execution order relationship between target operators includes:

[0018] Acquire first operation data corresponding to the first operation information and second operation data corresponding to the second operation information;

[0019] Determining target data corresponding to each target operator according to the first operation data and the second operation data;

[0020] The target model file is determined based on the target operator, the execution order relationship between the target operators and the target data.

[0021] According to an embodiment of the present disclosure, determining multiple target operators based on first operation information and second operation information includes:

[0022] Obtaining an operator set, where operators in the operator set can be executed by the second inference engine;

[0023] Acquire first operator sub-information in the first operation information and the second operation information, where the first operator sub-information has a corresponding first operator in the operator set;

[0024] At least part of the target operator is determined according to the first operator corresponding to the first operator information.

[0025] According to an embodiment of the present disclosure, determining multiple target operators based on first operation information and second operation information includes:

[0026] Obtaining the first operation information and the second operation sub-information in the second operation information;

[0027] Determine a target operator combination consisting of multiple second operators from the operator set, wherein when the target operator combination is executed in a target order, an execution result corresponding to the second operator information can be obtained;

[0028] At least part of the target operators is determined according to a target operator combination consisting of a plurality of second operators.

[0029] According to an embodiment of the present disclosure, the method further includes:

[0030] Determine third operation sub-information in the first operation information and the second operation information;

[0031] The target script file is determined according to the third operation sub-information, and the target script file is coordinated with the target model file to execute the target task.

[0032] According to an embodiment of the present disclosure, the first inference engine corresponds to the graphics processing unit, and the second inference engine corresponds to the neural network processing unit.

[0033] According to an embodiment of the present disclosure, determining a target model file based on a target operator and an execution order relationship between target operators includes:

[0034] Determine node information based on the target operator;

[0035] Determine the edge information between target operators based on the execution order relationship between target operators;

[0036] Determine the computational graph file based on node information and edge information;

[0037] Determine the target model file based on the computational graph file.

[0038] A second aspect of the present disclosure provides a model file processing device, comprising:

[0039] A task execution module, configured to run the first model and the second model in the first inference engine to execute the target task;

[0040] An information acquisition module, configured to acquire first operation information for operating the first model and second operation information for operating the second model during the execution of the target task;

[0041] an information processing module, configured to determine a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by the second inference engine;

[0042] The file output module is used to determine the target model file based on the target operator and the execution order relationship between the target operators; and output the target model file.

[0043] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0045] Figure 1 Schematically shows one of the flow charts of a model file processing method according to an embodiment of the present disclosure

[0046] Figure 2 The second flowchart of a model file processing method according to an embodiment of the present disclosure is schematically shown;

[0047] Figure 3 The third flowchart schematically shows a method for processing a model file according to an embodiment of the present disclosure;

[0048] Figure 4 A fourth flowchart of a model file processing method according to an embodiment of the present disclosure is schematically shown;

[0049] Figure 5 The following schematically shows a structural diagram of a model file processing device according to an embodiment of the present disclosure;

[0050] Figure 6 A block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0051] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0052] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0053] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0054] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0055] The model file processing method provided by the present disclosure can be applied to scenarios where a specific task is performed through multiple models (such as a first model and a second model).

[0056] An application scenario corresponding to the embodiment of the present disclosure is described below.

[0057] Speculative decoding is a large language model (LLM) inference acceleration technology. Its core is to improve inference efficiency by predicting subsequent tokens (text units) in parallel, while strictly ensuring that decoding quality is not compromised. Traditional LLM inference typically uses an autoregressive approach, generating tokens one by one. This serial processing method is inefficient. Speculative decoding, on the other hand, introduces a lighter-weight "draft model" to predict the subsequent token sequence in advance. If the main model (target model) passes verification, it is directly adopted; otherwise, it falls back to conventional autoregressive generation. The draft model quickly generates candidate sequences, while the main model is responsible for quality control and error correction only when necessary. Its advantage is that parallel prediction significantly reduces the number of calls to the main model, significantly reducing inference latency and computational cost while maintaining output accuracy.

[0058] In related technologies, developers can implement speculative decoding on various inference engines they build. This can be achieved by simulating a Neural Processing Unit (NPU) device as a Graphics Processing Unit (GPU) device and performing inference. The specific processing steps are as follows:

[0059] First, the data to be processed is copied from the CPU main memory to the GPU memory. This step is the foundation for subsequent GPU computations and ensures that the data can be efficiently accessed by the GPU. The CPU then issues an operator instruction to start the GPU compute kernel. An operator instruction is an instruction that instructs the GPU to perform a specific computational task and defines the operations that the GPU needs to perform. The GPU's CUDA architecture then executes the compute kernel. The compute kernel is the program that actually performs the computation on the GPU, accelerating the model inference process by leveraging the GPU's parallel computing capabilities. After the computation is complete, the resulting result data is copied from the GPU memory back to the CPU main memory, where the CPU further processes or outputs the results.

[0060] In this approach, because GPUs lack a logic control unit, they can only execute one operator at a time, returning the result to the CPU after the calculation is complete. The CPU then controls the next operator and sends it to the GPU for execution. This serial instruction execution method requires frequent data transfer and instruction exchange between the CPU and GPU when handling complex inference tasks, increasing communication overhead and latency.

[0061] The present disclosure provides a model file processing method to solve the problems existing in the related art.

[0062] Figure 1One of the flowcharts of a model file processing method according to an embodiment of the present disclosure is schematically shown.

[0063] Specifically, such as Figure 2 As shown, the method includes operations S101 to S104.

[0064] Operation S101, running a first model and a second model in a first inference engine to perform a target task;

[0065] In the disclosed embodiment, the first inference engine can be constructed based on a parsing layer, a computational optimization layer, a hardware layer, and an interface layer. The model parsing layer is responsible for parsing the model files exported by the training framework (such as the ONNX format or other preset formats) into an executable graph structure within the engine, and performs input data preprocessing and output post-processing. The computational optimization layer improves computational efficiency through techniques such as operator fusion, memory reuse, and quantization compression, while also combining parallel computing strategies to reduce latency. The hardware layer adapts to various architectures through a unified interface. The interface layer provides programming interfaces and deployment tools.

[0066] In the embodiment of the present disclosure, there is no limitation on the specific construction method of the first inference engine. In specific applications, it can be constructed according to the target tasks executed by the first model and the second model.

[0067] In the embodiment of the present disclosure, the second model and the first model are used to execute the same target task in the first inference engine.

[0068] Operation S102: during the execution of the target task, respectively obtaining first operation information for operating the second model and second operation information for operating the first model;

[0069] Specifically, during the execution of the target task, the first operation information of the second model includes: model loading operation, verification operation and reasoning operation of the second model; the second operation information of the first model includes: model loading operation, guessing operation and reasoning operation.

[0070] Exemplarily, for the first model, the model loading operation is used to load the model parameters from the storage into the memory of the first inference engine. Since the scale of the first model is smaller than that of the second model, the loading process is usually faster and occupies less memory. The guessing operation enables the first model to generate a guessing result based on the input context information, wherein the first model quickly generates multiple candidate tokens or sequences based on its own training knowledge and algorithms. In the process of the inference operation enabling the first model to generate a candidate sequence, the first model processes the input based on its internal computing logic, calculates the probability of each possible candidate token, and selects a candidate sequence according to the probability distribution. Compared with the inference operation of the second model, the inference operation process of the first model is relatively simple, and its purpose is to generate as many candidate sequences as possible in a short time, so as to provide more options for the verification of the second model.

[0071] Exemplarily, for the second model, the model loading operation is used by the second model to read a pre-trained model parameter file from a storage device (such as a hard drive or cloud) and load it into the memory of the first inference engine. The verification operation is used when the first model generates candidate sequences. The second model will fully verify these sequences, analyzing the candidate sequences' logical rationality, grammatical correctness, and fit with the context. The inference operation is used by the second model to perform deep inference based on the verification results. If there are any problems with the candidate sequences, the second model will attempt to correct the inference and generate more appropriate output results.

[0072] Operation S103 : determining a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by the second inference engine.

[0073] Specifically, the first operation information and the second operation information are converted into multiple target operators. Each target operator is configured as an executable function or operation unit that receives input data, processes the input data according to predefined logic, and ultimately outputs the processed result. Each target operator can be parsed and executed by the second inference engine.

[0074] For example, when the target task is an inference task related to speculative decoding of a large language model, the first inference engine is built based on a GPU device, and the second inference engine is built based on an NPU device.

[0075] Operation S104 : determining a target model file based on the target operator and the execution order relationship between the target operators; and outputting the target model file.

[0076] In an embodiment of the present disclosure, the second inference engine is another inference engine different from the first inference engine.

[0077] Specifically, the target model file is determined based on the target operators and the execution order relationship between them, and the target model file represents a complete model execution flow chart. It includes model loading, candidate sequence generation, probability calculation, selection optimization to each operation of the final output and its execution order relationship. Therefore, the model running in the second inference engine can determine the target model file that needs to be called, that is, the model configuration file containing all necessary operators and execution order information, and output it to the second inference engine for parsing and execution. Outputting the target model file includes outputting the target model file to a device configured with a second inference engine, and the second inference engine performs the inference task according to the target model file. The above embodiment integrates the model information of the first model, the model information of the second model, and the operations outside the model inference process that need to be performed in the process of the first model and the second model cooperating to perform the target task into the same model file, so that the second inference engine can perform the target task more efficiently.

[0078] The model file processing method provided by the present disclosure can also be applied to the following scenarios.

[0079] In some embodiments, the first model is a speech recognition model, and the second model is an image recognition model, and the first and second models work together to complete the target reasoning task. In some embodiments, the first model is a large language model, and the second model is an image generation diffusion model, and the first and second models work together to complete the target reasoning task. This disclosure is for illustrative purposes only and is not intended to limit the present disclosure.

[0080] In the embodiments of the present disclosure, since the effects of executing a specific reasoning task through multiple models on different reasoning engines are different, during the model reasoning process, the type of reasoning engine running multiple models can be adjusted so that the reasoning results can be better obtained when executing the reasoning task.

[0081] By adopting the above method, the operation information of the first model and the second model when performing the target task is obtained to produce the target operator that can be executed on the second inference engine, so that multiple models can be migrated on different inference engines in the process of performing specific tasks.

[0082] According to an embodiment of the present disclosure, running the first model and the second model in the first inference engine includes operations S1011 to S1013.

[0083] Operation S1011: outputting prediction information through the first model, where the prediction information represents an output prediction result of the second model;

[0084] Specifically, the first model receives input context and, based on its lightweight and fast generation capabilities, outputs predictions about the output of the second model. These predictions reflect the first model's understanding and preliminary reasoning about the task at hand, predicting the candidate sequences or outcomes that the second model might generate when processing a specific input.

[0085] Operation S1012: verifying the prediction result using the second model information to obtain a verification result;

[0086] Specifically, the predictions generated by the first model are compared and verified with the candidate sequences actually generated by the second model. The second model leverages its inference capabilities to generate multiple high-quality candidate sequences. The second model then screens these candidate sequences to find the ones that best meet the model's requirements.

[0087] Operation S1013 : determining target output information of a second model based on the verification result, wherein the number of parameters of the second model is greater than that of the first model.

[0088] Specifically, the target output information of the first model is determined based on the verification results of the second model. The second model will sort and screen the verified candidate sequences and select the target sequence that meets the requirements of the second model as the final output.

[0089] In the disclosed embodiment, the second model is configured as a model with a large parameter scale, complex structure, and high accuracy; the first model is configured as a relatively lightweight draft model. Compared with the second model, the first model has a smaller parameter scale, lower computational complexity, and faster inference speed.

[0090] Exemplarily, the first model is a large language model, the second model is a large language model, and the number of parameters of the second model is greater than the number of parameters of the first model, such as the first model is a 7B parameter model and the second model is a 70B parameter model. In some descriptions, the first model may also be referred to as a small model due to its smaller number of parameters relative to the second model. Exemplarily, the first model or the second model may be a large multimodal model, a model capable of generating text, a model with more than 100 million adjustable parameters, a model that can output corresponding results based on input natural language instructions, a model that can convert input natural language instructions into a feature sequence and infer output results based on the feature sequence, or a model that can convert input information into a feature sequence and infer subsequent features of the feature sequence.

[0091] According to an embodiment of the present disclosure, the method further includes:

[0092] According to the execution process of the target task, the execution order relationship between the first operation information and the second operation information is determined; according to the execution order relationship between the first operation information and the second operation information, the execution order relationship between the target operators is determined.

[0093] Specifically, the first operation information includes multiple first sub-operations in the process of the first model executing the target task; the first operation information includes multiple second sub-operations in the process of the second model executing the target task; wherein, each first sub-operation and each second sub-operation has a corresponding target operator; based on the timing relationship when the two models execute the sub-operations; the execution order relationship between the target operators is determined.

[0094] Figure 2 The second flowchart of a model file processing method according to an embodiment of the present disclosure is schematically shown.

[0095] According to an embodiment of the present disclosure, in operation S104 , a target model file is determined based on a target operator and an execution order relationship between target operators, including operations S201 to S203 .

[0096] Operation S201: Acquire first operation data corresponding to first operation information and second operation data corresponding to second operation information;

[0097] Specifically, the first operation data corresponding to the first operation information includes the operation data and data storage address that the first model needs to call when the first model performs the first sub-operation. The second operation data corresponding to the second operation information includes the operation data and data storage address that the second model needs to call when the second model performs the second sub-operation.

[0098] Operation S202: determining target data corresponding to each target operator according to the first operation data and the second operation data;

[0099] Specifically, for the target operator determined based on the first operation information, the operation data and data storage address required to be called by the first model to perform the first sub-operation are used as the target data corresponding to the target operator. For the target operator determined based on the second operation information, the operation data and data storage address required to be called by the second model to perform the second sub-operation are used as the target data corresponding to the target operator.

[0100] Operation S203 : determining a target model file based on the target operator, the execution order relationship between the target operators, and the target data.

[0101] Specifically, the target model file contains the correspondence between each target operator and target data, as well as the execution order relationship between each target operator.

[0102] Figure 3 The third flowchart of a model file processing method according to an embodiment of the present disclosure is schematically shown.

[0103] like Figure 3As shown, according to an embodiment of the present disclosure, determining multiple target operators based on the first operation information and the second operation information includes operations S301 to S303:

[0104] Operation S301: obtaining an operator set, where operators in the operator set can be executed by a second inference engine;

[0105] Operation S302: Acquire first operator sub-information in the first operation information and the second operation information, where the first operator sub-information has a corresponding first operator in the operator set;

[0106] Operation S303: Determine at least part of the target operator according to the first operator corresponding to the first operator information.

[0107] In an embodiment of the present disclosure, the format type of each operator in the operator set may be determined based on the construction manner of the second inference.

[0108] Specifically, any one of the first operation information and the second operation information may include multiple first operator sub-information. When the first operator sub-information has a corresponding first operator in the operator set, the first operator corresponding to each first operator sub-information in the operator set is obtained and used as part of the target operator.

[0109] Figure 4 A fourth flowchart of a model file processing method according to an embodiment of the present disclosure is schematically shown.

[0110] like Figure 4 As shown, according to an embodiment of the present disclosure, determining multiple target operators based on the first operation information and the second operation information includes operations S401 to S403:

[0111] According to an embodiment of the present disclosure, determining multiple target operators based on the first operation information and the second operation information further includes:

[0112] Operation S401: Acquire the first operation information and the second operation sub-information in the second operation information;

[0113] Operation S402: determining a target operator combination consisting of a plurality of second operators from the operator set, wherein when the target operator combination is executed in a target order, an execution result corresponding to the second operator information can be obtained;

[0114] Operation S403 : determining at least part of the target operators according to the target operator combination composed of the plurality of second operators.

[0115] Specifically, either the first operation information or the second operation information may include multiple pieces of second operator sub-information. Since the second operator sub-information does not have a corresponding operator in the operator set, an operator combination consisting of multiple second operators is obtained from the operator set, so that the second operator sub-information can represent the processing logic when the operator combination is executed in the target order. A target operator combination consisting of the multiple second operators is obtained and used as part of the target operator.

[0116] In an embodiment of the present disclosure, any one of the first operation information and the second operation information may include first operator sub-information and second operator sub-information. In this case, the target operator includes: a target operator combination consisting of the first operator corresponding to each first operator sub-information in the operator set and multiple second operators.

[0117] According to an embodiment of the present disclosure, the method also includes: determining third operation sub-information in the first operation information and the second operation information; determining a target script file based on the third operation sub-information, and the target script file is coordinated with the target model file to execute the target task.

[0118] In an embodiment of the present disclosure, any one of the first operation information and the second operation information may include first operation sub-information, second operation sub-information, and third operation sub-information.

[0119] Specifically, when the operation information includes the first operator sub-information, the second operator sub-information, and the third operator sub-information. Since the first operator sub-information and the second operator sub-information can directly or indirectly determine the corresponding operator in the operator set; wherein, the first operator sub-information corresponds to the first operator in the operator set, and the second operator sub-information corresponds to multiple second operators in the operator set. In contrast, the third operator sub-information does not have one or more corresponding operators in the operator set. Therefore, the logical operation represented by the third operator sub-information is converted into a script program to obtain a target script file, and the logical operation represented by the first operator sub-information and the second operator sub-information is converted into a target model file. By coordinating the script file with the target model file, the target task is executed in the second reasoning model.

[0120] In the embodiments of the present disclosure, operator sets can be constructed using various types of ONNX operators. ONNX operators are an open deep learning model representation format that defines operators such as convolution, pooling, and normalization. ONNX operators support model sharing across different frameworks. Specifically, ONNX operators refer to various basic operations and computational units defined in the Open Neural Network Exchange (ONNX) for building and running deep learning models.

[0121] Exemplarily, the types of each operator in the operator set include but are not limited to: basic operation type, activation function type, convolution and pooling type, normalization and standardization type, tensor operation type and logical comparison type.

[0122] According to an embodiment of the present disclosure, the first inference engine corresponds to the graphics processing unit, and the second inference engine corresponds to the neural network processing unit.

[0123] Specifically, the first inference engine can be built based on a graphics processing device, and the second inference engine can be built based on a neural network processing device.

[0124] Exemplarily, after the target model file is output, the target model file is run on the neural network processing device via the second inference engine to perform the target task. Additionally, based on the third operation sub-information in the first and second operation information, the target script file determined is run on the neural network processing device via the second inference engine to execute the target task in conjunction with the target model file.

[0125] According to an embodiment of the present disclosure, determining a target model file based on a target operator and an execution order relationship between target operators includes:

[0126] Determine the node information based on the target operator; determine the edge information between the target operators based on the execution order relationship between the target operators; determine the calculation graph file based on the node information and edge information; determine the target model file based on the calculation graph file.

[0127] Exemplarily, the target operators include: a convolution operator, an activation function operator, and a maximum pooling operator. The edge information between the target operators is determined based on the execution order between the target operators. For example, the target model file is configured such that input data first passes through the input node, then through the convolution node, activation function node, maximum pooling node, and output node in sequence.

[0128] For example, if the target operator is an ONNX operator, the defined node information is integrated into a computation graph. Each node represents an operator and includes its inputs, outputs, and parameters. The nodes are connected based on the edge information to form a complete computation graph file, ensuring that the output of each node is correctly connected to the input of the next node. Finally, the serialized computation graph file is saved as a target model file in ONNX format. The target model file contains the complete execution order of the target operator and parameter information for each node.

[0129] In the embodiment of the present disclosure, when the second inference engine is built based on the NPU device, since the NPU device and the GPU device have essential differences in architecture and working mode. The NPU device has a logic control unit, and this feature gives it a unique advantage in the inference process. The NPU can issue multiple operator instructions, that is, fused operator instructions, execute multiple calculation steps, and then return the result data to the CPU main memory. The method of obtaining the target operator in the present disclosure can reduce the number of communications between the CPU and the NPU, and can more efficiently utilize the computing resources of the NPU device in the second inference engine, thereby improving its efficiency in inference acceleration.

[0130] By adopting the above method, the present disclosure can implement speculative decoding of large language models on different inference engines; when implementing speculative decoding of LLM on NPU devices using target model files, the advantages of NPU in executing operators can be brought into play.

[0131] Figure 5 The following schematically shows a structural diagram of a model file processing device according to an embodiment of the present disclosure.

[0132] The second aspect of the present disclosure provides a model file processing device, such as Figure 5 As shown, the model file processing device 500 includes:

[0133] A task execution module 510 is configured to execute the first model and the second model in the first inference engine to execute the target task;

[0134] The information acquisition module 520 is used to respectively acquire first operation information of operating the first model and second operation information of operating the second model during the execution of the target task;

[0135] An information processing module 530 is configured to determine a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by the second inference engine;

[0136] The file output module 540 is used to determine the target model file based on the target operator and the execution order relationship between the target operators; and output the target model file.

[0137] According to an embodiment of the present disclosure, the task execution module 510 is specifically used to output prediction information through the first model, where the prediction information represents the output prediction result of the second model; verify the prediction result through the second model information to obtain a verification result; and determine the target output information of the second model based on the verification result, wherein the number of parameters of the second model is greater than that of the first model.

[0138] According to an embodiment of the present disclosure, the information processing module 530 is also used to determine the execution order relationship between the first operation information and the second operation information based on the execution process of the target task; and determine the execution order relationship between the target operators based on the execution order relationship between the first operation information and the second operation information.

[0139] According to an embodiment of the present disclosure, the file output module 540 is specifically used to obtain first operation data corresponding to the first operation information and second operation data corresponding to the second operation information; determine the target data corresponding to each target operator based on the first operation data and the second operation data; and determine the target model file based on the target operator, the execution order relationship between the target operators, and the target data.

[0140] According to an embodiment of the present disclosure, the information processing module 530 includes a first information processing unit;

[0141] The first information processing unit is used to obtain an operator set, where the operators in the operator set can be executed by the second inference engine; obtain first operator sub-information in the first operation information and the second operation information, where the first operator sub-information has a corresponding first operator in the operator set; and determine at least part of the target operator based on the first operator corresponding to the first operator sub-information.

[0142] According to an embodiment of the present disclosure, the information processing module 530 further includes a second information processing unit;

[0143] The second information processing unit is used to obtain the second operator information in the first operation information and the second operation information; determine a target operator combination composed of multiple second operators from the operator set, and when the target operator combination is executed in the target order, an execution result corresponding to the second operator sub-information can be obtained; and determine at least part of the target operators based on the target operator combination composed of multiple second operators.

[0144] According to an embodiment of the present disclosure, the information processing module 530 further includes a third information processing unit;

[0145] The third information processing unit is used to determine the third operation sub-information in the first operation information and the second operation information; determine the target script file according to the third operation sub-information, and the target script file is coordinated with the target model file to execute the target task.

[0146] According to an embodiment of the present disclosure, the first inference engine corresponds to the graphics processing unit, and the second inference engine corresponds to the neural network processing unit.

[0147] According to an embodiment of the present disclosure, the file output module 540 is specifically used to determine node information based on the target operator; determine edge information between target operators based on the execution order relationship between target operators; determine a computational graph file based on the node information and edge information; and determine a target model file based on the computational graph file.

[0148] It should be noted that the model file processing device part in the embodiment of the present disclosure corresponds to the processing method part in the embodiment of the present disclosure, and their specific implementation details are also the same, which will not be repeated here.

[0149] Figure 6 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure is schematically shown. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0150] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory configured for cache purposes. The processor 601 may include a single processing unit or multiple processing units configured to perform different actions of the method flow according to an embodiment of the present disclosure.

[0151] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0152] According to an embodiment of the present disclosure, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0153] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code configured to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.

[0154] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0155] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0156] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603 .

[0157] An embodiment of the present disclosure also includes a computer program product, which includes a computer program containing program code configured to execute the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is configured to enable the electronic device to implement the remote sensing image detection method based on deep neural network provided by the embodiment of the present disclosure.

[0158] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0159] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0160] According to an embodiment of the present disclosure, the program code configured to execute the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions configured to implement the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, and all of these combinations and / or couplings fall within the scope of the present disclosure.

[0162] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A model file processing method, comprising: Running the first model and the second model in the first inference engine to perform the target task; During the execution of the target task, first operation information for operating the first model and second operation information for operating the second model are respectively obtained; Determining a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by a second inference engine; determining a target model file based on the target operator and an execution order relationship between the target operators; Output the target model file.

2. The method according to claim 1, wherein running the first model and the second model in the first inference engine comprises: Output prediction information through the first model, where the prediction information represents an output prediction result of the second model; Verifying the prediction result using the second model information to obtain a verification result; Target output information of the second model is determined based on the verification result, wherein the number of parameters of the second model is greater than that of the first model.

3. The method according to claim 1, further comprising: determining, according to the execution process of the target task, an execution sequence relationship between the first operation information and the second operation information; The execution order relationship between the target operators is determined according to the execution order relationship between the first operation information and the second operation information.

4. The method according to claim 1, wherein determining the target model file based on the target operator and the execution order relationship between the target operators comprises: Acquire first operation data corresponding to the first operation information and second operation data corresponding to the second operation information; Determining target data corresponding to each target operator according to the first operation data and the second operation data; The target model file is determined based on the target operator, the execution order relationship between the target operators, and the target data.

5. The method according to claim 1, wherein determining a plurality of target operators based on the first operation information and the second operation information comprises: Obtaining an operator set, where operators in the operator set can be executed by the second inference engine; Acquire first operator sub-information in the first operation information and the second operation information, where the first operator sub-information has a corresponding first operator in the operator set; At least part of the target operators is determined according to the first operator corresponding to the first operator information.

6. The method according to claim 5, wherein determining a plurality of target operators based on the first operation information and the second operation information comprises: Obtaining the first operation information and the second operation sub-information in the second operation information; Determining a target operator combination consisting of a plurality of second operators from the operator set, wherein the target operator combination, when executed in a target order, can obtain an execution result corresponding to the second operator information; At least part of the target operators is determined according to a target operator combination consisting of the plurality of second operators.

7. The method according to claim 6, further comprising: Determining third operation sub-information in the first operation information and the second operation information; A target script file is determined according to the third operation sub-information, and the target script file is coordinated with the target model file to execute the target task.

8. The method according to claim 1, wherein the first inference engine corresponds to a graphics processing unit, and the second inference engine corresponds to a neural network processing unit.

9. The method according to claim 1, wherein determining the target model file based on the target operator and the execution order relationship between the target operators comprises: Determining node information based on the target operator; Determining edge information between the target operators based on an execution order relationship between the target operators; Determine a computation graph file according to the node information and the edge information; The target model file is determined according to the calculation graph file.

10. A model file processing device, comprising: A task execution module, configured to run the first model and the second model in the first inference engine to execute the target task; an information acquisition module, configured to respectively acquire first operation information for operating the first model and second operation information for operating the second model during the execution of the target task; an information processing module, configured to determine a plurality of target operators based on the first operation information and the second operation information, wherein the target operators can be parsed and executed by a second inference engine; The file output module is used to determine the target model file based on the target operator and the execution order relationship between the target operators; and output the target model file.