Model adaptation method, model adaptation device, electronic device, and storage medium

CN122816697APending Publication Date: 2026-09-25SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272274.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0017]在本公开至少一实施例提供的模型适配方法、模型适配装置、电子设备和存储介质中,通过在适配流程入口处通过实测运行状态自动判断模型的适配状态和原型类别,选择最短执行路径,并根据模型原型类别定制关键路径和优化策略,能够大幅提升模型的适配效率,将已知可用模型的适配时间从小时级缩短至分钟级,有效克服线性模型适配流水线的效率瓶颈。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816697A_ABST
    Figure CN122816697A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a model adaptation method, a model adaptation device, an electronic device and a storage medium. The method comprises: determining a target adaptation mode from a plurality of adaptation modes based on a forward propagation running state of a model on a target hardware device; and deploying the model to the target hardware device based on the target adaptation mode, wherein the operation link lengths of different adaptation modes are different in the plurality of adaptation modes. The method can greatly improve the adaptation efficiency of the model by automatically judging the adaptation state of the model at the entrance of the adaptation process, selecting the shortest execution path, and effectively overcoming the efficiency bottleneck of the linear model adaptation pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of artificial intelligence, specifically to a model adaptation method, a model adaptation device, an electronic device, and a storage medium. Background Technology

[0002] With the continuous development of Artificial Intelligence (AI) technology, AI processors have become crucial hardware devices supporting the training and inference of deep learning models. Operators are the basic units for building deep learning models, and their execution efficiency determines the overall performance of the model. The related underlying hardware architectures exhibit significant diversity and heterogeneity, making efficient cross-hardware adaptation of models a problem that needs to be solved. Summary of the Invention

[0003] This disclosure provides at least one embodiment of a model adaptation method, the method comprising: determining a target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on a target hardware device; and deploying the model to the target hardware device based on the target adaptation mode, wherein the operation link lengths of different adaptation modes are different among the multiple adaptation modes.

[0004] In the model adaptation method provided in at least one embodiment of this disclosure, the plurality of adaptation modes include a direct deployment mode, a fixed-point repair mode, and a full-scale adaptation mode, wherein the operation link length of the full-scale adaptation mode is greater than the operation link length of the fixed-point repair mode, and the operation link length of the fixed-point repair mode is greater than the operation link length of the direct deployment mode.

[0005] In the model adaptation method provided in at least one embodiment of this disclosure, the direct deployment mode includes a deployment operation, the fixed-point repair mode includes an operator detection operation, an operator repair operation, and a deployment operation, and the full adaptation mode includes an operator analysis operation, an operator coverage check operation, an operator generation operation, an operator integration operation, and a deployment operation.

[0006] In the model adaptation method provided in at least one embodiment of this disclosure, determining the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device includes: determining the target adaptation mode as the direct deployment mode in response to the model correctly completing forward propagation on the target hardware device; determining the target adaptation mode as the fixed-point repair mode in response to the operator error existing in the forward propagation process of the model on the target hardware device; and determining the target adaptation mode as the full adaptation mode in response to the abnormal loading of the model on the target hardware device.

[0007] In the model adaptation method provided in at least one embodiment of this disclosure, the deployment operation includes: deploying the model to the target hardware device based on a target deployment sub-mode, wherein the target deployment sub-mode is determined based on the configuration information of the model, and the target deployment sub-mode includes at least one of an inference deployment sub-mode and a training deployment sub-mode.

[0008] In at least one embodiment of the model adaptation method provided in this disclosure, before determining the target adaptation mode from multiple adaptation modes, the method further includes: determining the target deployment sub-mode based on the configuration information of the model.

[0009] In the model adaptation method provided in at least one embodiment of this disclosure, determining the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device includes: determining the target adaptation mode from the multiple adaptation modes based on the forward propagation running state of the model on the target hardware device and in conjunction with user instruction information.

[0010] The model adaptation method provided in at least one embodiment of this disclosure further includes: determining the prototype category of the model based on the model's configuration information; determining the operator optimization strategy of the model based on the prototype category of the model; and performing an operator optimization operation based on the operator optimization strategy in response to completing the deployment operation.

[0011] In at least one embodiment of the model adaptation method provided in this disclosure, determining the prototype category of the model based on the model's configuration information includes: for a first prototype category among a plurality of preset prototype categories, determining the confidence level of the first prototype category based on the classification signal corresponding to the first prototype category in the configuration information; and determining the prototype category of the model as the first prototype category in response to the first prototype category confidence level satisfying a preset condition.

[0012] The model adaptation method provided in at least one embodiment of this disclosure further includes: during the execution of each operation included in the method, in response to the completion of the current operation, writing the execution result of the current operation into a status record file.

[0013] The model adaptation method provided in at least one embodiment of this disclosure further includes: in response to an operation interruption, determining the next operation to be executed based on the operation corresponding to the last execution result written in the state record file.

[0014] At least one embodiment of this disclosure provides a model adaptation device, the device comprising: a determining module configured to determine a target adaptation mode from a plurality of adaptation modes based on the forward propagation running state of the model on a target hardware device; and a deployment module configured to deploy the model to the target hardware device based on the target adaptation mode, wherein the operation link lengths of different adaptation modes are different among the plurality of adaptation modes.

[0015] At least one embodiment of this disclosure provides an electronic device, including at least one processor and at least one memory, wherein the at least one memory stores program code that, when executed by the at least one processor, causes the at least one processor to perform a model adaptation method according to at least one embodiment of this disclosure.

[0016] At least one embodiment of this disclosure provides a non-transitory computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by at least one processor, cause the at least one processor to perform the model adaptation method according to at least one embodiment of this disclosure.

[0017] In the model adaptation method, model adaptation device, electronic device and storage medium provided in at least one embodiment of this disclosure, by automatically determining the model adaptation status and prototype category at the entry point of the adaptation process through actual running status, selecting the shortest execution path, and customizing the critical path and optimization strategy according to the model prototype category, the model adaptation efficiency can be greatly improved, the adaptation time of known available models can be shortened from hours to minutes, and the efficiency bottleneck of linear model adaptation pipeline can be effectively overcome. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0019] Figure 1 A flowchart of a model adaptation method provided for at least one embodiment of this disclosure.

[0020] Figure 2 A flowchart of another model adaptation method provided for at least one embodiment of this disclosure.

[0021] Figure 3 This is a schematic block diagram of a model adaptation system provided for at least one embodiment of the present disclosure.

[0022] Figure 4 This is a schematic block diagram of a model adaptation device provided for at least one embodiment of the present disclosure.

[0023] Figure 5 This is a schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.

[0024] Figure 6 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.

[0025] Figure 7 This is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0027] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0028] The present disclosure will now be described through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, the component is represented by the same or similar reference numerals in each drawing.

[0029] A neural network model can be viewed as a directed acyclic graph (DAG) composed of multiple nodes, where each node corresponds to an operator. Operators in a neural network typically refer to the basic mathematical operations or processes used in the network layers. These operators are used to construct the various layers and components of the network, enabling data transfer, transformation, and computation. Operators are the fundamental building blocks of a network model, defining its structure and computational flow, including inputs, outputs, and intermediate computations. In the model, the connections between operators form a directed graph, reflecting the order of computation for different operations. By combining these operators, complex and powerful neural network models can be constructed to handle various complex tasks and data. Some examples of operators include convolution, pooling, loops, activation functions, normalization, matrix multiplication, and element-wise operations.

[0030] The operator kernel is the underlying implementation of the operator. It can be written in a low-level language (such as C++) and manifested as a device program (for example, it can be executed on hardware devices such as graphics processing units, GPUs) that describes the operator's execution process. It is the logic that drives the operator's function to run efficiently on the specified hardware and is responsible for translating the mathematical definition of the operator into an actual sequence of hardware instructions.

[0031] With the rapid development of AI applications such as Large Language Models (LLM), visual models, and diffusion models, the demand for computing power in AI inference and training is growing exponentially. Currently, the operator ecosystem of mainstream AI frameworks (such as PyTorch) is highly dependent on specific computing libraries and programming models. When migrating existing AI models from general-purpose GPU platforms to heterogeneous hardware platforms (such as dedicated accelerator chips), a severe bottleneck in operator compatibility is generally encountered. Because the operator libraries of target hardware devices often cannot fully cover all computational primitives of the model, some operators are forced to fall back to execution by the central processing unit (CPU), or even fail to run, thus severely restricting the deployment efficiency and performance release of models in heterogeneous computing environments.

[0032] Traditional model adaptation methods employ a fixed linear pipeline, sequentially executing environment checks, operator analysis, coverage checks, operator generation, integration, verification, deployment, and optimization. This requires any model to complete all of these steps. This approach results in several hours of wasted computation time, even if the model is fully supported by the current software stack.

[0033] Meanwhile, traditional model adaptation methods employ a uniform performance optimization strategy for all model categories, such as uniform batch size tuning and quantization schemes. This fails to distinguish the fundamental differences in critical paths, bottleneck operators, deployment frameworks, and optimization strategies during the adaptation process caused by different model architectures, such as dense large language models, hybrid expert models, ResNet vision models, and diffusion models. This results in a large amount of unnecessary computation and waiting.

[0034] Furthermore, the technology stacks for inference deployment and training deployment differ significantly. Traditional model adaptation methods only branch off paths in the later stages of the pipeline, resulting in the reading of a large amount of irrelevant context in the earlier steps, leading to a waste of resources.

[0035] Furthermore, coverage checks typically employ static coverage methods, which merely compare the list of operators used by the model with the chip's supported operator registry. This method does not verify runtime behavior and suffers from the "phantom operator" problem—operators that exist in the registry but crash under specific data shapes or data types. Such false positives cannot be detected by static analysis, leading to inaccurate coverage checks. In addition, traditional model adaptation systems lack incremental recovery capabilities; any failure in the model adaptation process requires re-execution from the beginning, severely impacting iteration efficiency.

[0036] This disclosure provides at least one embodiment of a model adaptation method, a model adaptation device, an electronic device, and a storage medium.

[0037] The model adaptation method provided in at least one embodiment of this disclosure includes: determining a target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device; and deploying the model to the target hardware device based on the target adaptation mode, wherein the operation link lengths of different adaptation modes are different among the multiple adaptation modes.

[0038] The model adaptation method provided in at least one embodiment of this disclosure can significantly improve the model adaptation efficiency by automatically determining the model's adaptation status at the entry point of the adaptation process based on the measured running status and selecting the shortest execution path. This reduces the adaptation time of known available models from hours to minutes, effectively overcoming the efficiency bottleneck of linear model adaptation pipelines.

[0039] Figure 1 A flowchart of a model adaptation method provided for at least one embodiment of this disclosure.

[0040] For example, such as Figure 1 As shown, the model adaptation method provided in at least one embodiment of this disclosure includes the following steps S101 to S102.

[0041] Step S101: Based on the forward propagation running state of the model on the target hardware device, determine the target adaptation mode from multiple adaptation modes, wherein the operation link length of different adaptation modes is different among the multiple adaptation modes.

[0042] Step S102: Deploy the model to the target hardware device based on the target adaptation mode.

[0043] For example, in step S101, the model can be a machine learning model, such as a deep learning model. A deep learning model can be a neural network structure comprising multiple layers (3, 4, 8, or more layers), such as a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM) network. These neural networks typically include an input layer, hidden layers, and an output layer. Hidden layers are those layers located between the input and output layers; they are also called processing layers. The input layer receives the data to be processed, such as the image to be processed, and the output layer outputs the processing result, such as the processed image. Processing layers can include convolutional layers, pooling layers, batch normalization layers, fully connected layers, etc. Depending on the structure of the neural network, the processing layer can include different content and combinations. In some examples, the model can be, for instance, an artificial intelligence model with a large number of parameters built from an artificial intelligence network (also known as a large-scale artificial intelligence model, or simply a "large model"). Large-scale models can be content generation models based on prompt words, such as Large Language Models (LLMs), large-scale visual models, or large-scale multimodal models. Examples include models based on Transformer architectures, recurrent neural networks, attention mechanisms, diffusion models, and multimodal models. This disclosure does not limit the specific type, structure, or implementation of the model.

[0044] For example, the target hardware device refers to the hardware device that the operators in the model need to adapt to. The hardware device may be an artificial intelligence processor, which may include a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), an accelerated processing unit (APU), a neural network processing unit (NPU), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA), etc. The embodiments disclosed herein do not impose any limitations on this.

[0045] For example, in step S101, multiple adaptation modes are all working modes used to implement model adaptation. Among the multiple adaptation modes, the operation chain lengths differ. For example, different adaptation modes may contain varying numbers of operation steps, thus making the operation chain lengths different for each adaptation mode.

[0046] For example, multiple adaptation modes can include adaptation mode 1, adaptation mode 2, ..., adaptation mode N, where N is a positive integer. For example, adaptation mode 1 can include a first number of operation steps; adaptation mode 2 can include a second number of operation steps, and the second number is greater than the first number, and so on.

[0047] It should be noted that the embodiments disclosed herein do not impose specific limitations on the specific types, number, and operational steps of the multiple adaptation modes, and can be set according to the actual application scenario.

[0048] By setting the above multiple adaptation modes, we can flexibly select adaptation modes with different operation link lengths according to the needs of actual application scenarios, thereby achieving a balance between operation link length and processing efficiency.

[0049] For example, in step S101, the forward propagation running state of the model on the target hardware device can be used as a gating condition for the target adaptation mode selection, thereby ensuring the accuracy of routing decisions through actual running state measurements.

[0050] In some examples, the running status can be used to indicate whether the forward propagation is running normally or whether the forward propagation result is correct. It should be noted that the above running status is only illustrative and is not intended to limit the scope of protection of the embodiments of this disclosure.

[0051] For example, a Quick Smoke Test can be used to obtain the forward propagation state of the model on the target hardware device. A Quick Smoke Test, for example, involves loading the model on the target hardware device and performing a minimum forward inference once, to quickly verify the basic usability of the model.

[0052] For example, the target adaptation mode can be the one that matches the model's running state and has the shortest operation chain length. The following example illustrates this: If the model's running state indicates that forward propagation is running normally and the results are correct, then theoretically, the model is already usable on the target hardware device. In this case, there's no need to perform operator analysis or operator generation steps; only deployment operations are required to complete model adaptation. Therefore, for this model, any adaptation mode that includes deployment operations can be considered to match the model's running state. In this case, the adaptation mode with the shortest operation chain length can be selected as the target adaptation mode. For example, an adaptation mode that only includes deployment operations can be selected as the target adaptation mode.

[0053] In other words, the "adaptation mode matching the model's running state" mentioned above means that the model can be successfully adapted based on this adaptation mode, which means that the adapted model can run correctly on the target hardware device.

[0054] For example, in step S102, the operation steps corresponding to the target adaptation mode can be executed according to the target adaptation mode determined in step S101, thereby deploying the model to the target hardware device.

[0055] In the model adaptation method provided in at least one embodiment of this disclosure, the adaptation status of the model is automatically determined by the measured running status at the entry point of the adaptation process, and the shortest execution path is selected, which can greatly improve the model adaptation efficiency and effectively overcome the efficiency bottleneck of the linear model adaptation pipeline.

[0056] In some examples, multiple adaptation modes may include direct deployment mode, targeted repair mode, and full adaptation mode. The operation chain length of full adaptation mode is longer than that of targeted repair mode, and the operation chain length of targeted repair mode is longer than that of direct deployment mode.

[0057] It should be noted that the multiple adaptation modes provided in this disclosure are not limited to the direct deployment mode, targeted repair mode, and full adaptation mode described above. More adaptation modes can be set according to actual needs. For example, the operation link length can be increased or decreased based on existing adaptation modes to obtain new adaptation modes.

[0058] In some examples, the full adaptation mode includes operator analysis operations, operator coverage check operations, operator generation operations, operator integration operations, and deployment operations. For example, if the target adaptation mode is the full adaptation mode, an example of step S102 may include: performing operator analysis operations, operator coverage check operations, operator generation operations, operator integration operations, and deployment operations to deploy the model to the target hardware device.

[0059] For example, operator analysis operations can represent a complete analysis of the model to be adapted, traversing the model computation graph to extract all operator nodes, constructing an operator list and counting the type, input and output shape and attribute parameters of each operator, which is used to identify the full set of operators in the model that need to be adapted, generated or require special processing, and to ensure the integrity of the model adaptation range.

[0060] For example, operator coverage checking can involve statistically analyzing the adaptation status of the entire operator list using at least one of dynamic and static analysis techniques. This calculates the ratio of supported operators to the total number of operators, quantitatively assessing the target hardware's support for the model to be adapted and identifying operator gaps that are not yet adapted. Static analysis, without executing code, involves parsing and pattern matching the model computation graph or operator source code to identify and statistically analyze the registered operator interfaces in the current runtime environment (e.g., deep learning framework or hardware backend), quickly assessing the static support of the operator list. Dynamic analysis, during the actual execution of test cases or simulated model forward propagation, uses runtime data acquisition mechanisms such as hardware performance counters (PFC) or profilers to track and record the actual call trajectory and execution status of operators in real time, identifying operator gaps that are not covered by runtime dependencies or conditional branches.

[0061] For example, the operator generation operation can represent automatically generating corresponding operator kernel code (such as high-performance computing kernel functions) based on operator attributes and target hardware instruction sets for the unfit operators identified in the operator coverage check. This includes memory allocation logic, parallel computing strategies, and data transfer instructions, which are used to fill the operator gaps in the model adaptation process and make the model executable on the target platform.

[0062] For example, operator ensemble operations can represent registering the generated operator kernel code into the operator registry of a deep learning framework, establishing a mapping between model computation graph nodes and underlying operator implementations, and configuring operator scheduling strategies and dependency management. This enables the deep learning framework to correctly call the newly generated operators to participate in the model's inference or training process. As a bridge between software and hardware, the deep learning framework can drive the target hardware device to perform forward or backward propagation operations by accessing its computational resources.

[0063] For example, deployment operations can represent loading models and operators onto target hardware devices through a deep learning framework, configuring corresponding running strategies (such as memory management, parallel scheduling, data transfer, etc.) based on the resource characteristics of the target hardware device, and driving the target hardware device to perform forward inference or backward training tasks of the model through the runtime interface of the deep learning framework.

[0064] In some examples, the full adaptation mode may also include performance verification operations. One example of a performance verification operation is a dataset verification operation. A dataset verification operation can represent dividing the verification dataset into multiple levels, such as functionally correct datasets, accuracy-aligned datasets, and performance stress testing datasets, based on the complexity of the model adaptation task and the verification objectives. Then, according to a pre-defined verification strategy, model inference and result comparison are performed on each level of dataset sequentially. This is used to determine the logical correctness, numerical accuracy consistency, and operational stability of the model on the target hardware in a hierarchical manner while controlling verification costs.

[0065] In some examples, the fixed-point repair mode includes operator probing, operator repair, and deployment. For example, if the target adaptation mode is fixed-point repair mode, an example of step S102 may include performing operator probing, operator repair, and deployment to deploy the model to the target hardware device.

[0066] Compared to the full adaptation mode, the targeted repair mode can skip operator analysis and operator integration operations. Therefore, the operation chain length of the targeted repair mode is shorter than that of the full adaptation mode.

[0067] For example, operator probing can refer to loading and executing an operator in an actual hardware device (e.g., the target hardware device) or a simulation environment of the hardware device, collecting the low-level state of the operator during runtime through a hardware performance counter or performance monitor, diagnosing and classifying operator error types (e.g., hardware abnormal interruption, numerical overflow, or precision deviation), thereby locating the cause of operator execution failure.

[0068] For example, operator repair operations can represent targeted adjustments to the kernel code implementation, memory layout, parallel computing strategy, or hardware instruction configuration of an operator based on the operator error type determined by the operator detection operation, in order to resolve execution anomalies caused by the operator itself and enable the operator to run correctly on the target hardware device.

[0069] In some examples, the direct deployment mode includes a deployment operation. For example, if the target adaptation mode is the direct deployment mode, one example of step S102 may include performing a deployment operation to deploy the model to the target hardware device.

[0070] Compared to the full-scale adaptation mode, the direct deployment mode can skip operator analysis, operator coverage check, operator generation, and operator integration operations; compared to the targeted repair mode, the direct deployment mode can skip operator detection and operator repair operations. Therefore, the operation chain length of the direct deployment mode is shorter than that of the targeted repair mode and the full-scale adaptation mode.

[0071] For example, in the examples of full adaptation mode, targeted repair mode, and direct deployment mode described above, the deployment operation may include: deploying the model to a target hardware device based on a target deployment sub-mode, wherein the target deployment sub-mode is determined based on the model's configuration information, and the target deployment sub-mode includes at least one of an inference deployment sub-mode and a training deployment sub-mode. Specifically, the target deployment sub-mode can be determined based on the task type information in the model configuration information. An example of model configuration information may be a model configuration file (e.g., config.json).

[0072] For example, model configuration information may include a use case field, which specifies the specific task type the model needs to perform. Task types include: performing model training only, performing model inference only, or performing both model training and model inference.

[0073] For example, if the task type is to perform model training, then the model will be deployed to the target hardware device based on the training deployment sub-mode.

[0074] For example, if the task type is to perform model inference, then the model will be deployed to the target hardware device based on the inference deployment sub-pattern.

[0075] For example, if the task type is to perform model training and model inference, then the model is deployed to the target hardware device based on the training deployment sub-mode and the inference deployment sub-mode.

[0076] For example, deployment operations can include training deployment operations and inference deployment operations. The training deployment sub-mode mentioned above corresponds to the training deployment operation and the related configuration information for loading the training deployment operation, while the inference deployment sub-mode corresponds to the inference deployment operation and the related configuration information for loading the inference deployment operation. In other words, based on the model's configuration information, one of the following three operations can be selected: perform only the training deployment operation, perform only the inference deployment operation, or perform both the training deployment operation and the inference deployment operation, thereby deploying the model to the target hardware device.

[0077] For example, training deployment operations may include: configuring distributed training strategies (such as data parallelism, model parallelism, tensor parallelism, etc.); establishing gradient synchronization mechanisms (such as all-reduce or reduced-scatter to ensure the consistency of parameter updates across multiple devices); or performing checkpoint saving and restoring operations to ensure the persistence and continuity of the training state when training is interrupted.

[0078] For example, inference deployment operations may include: using graph patterns to capture and compile the model computation graph to reduce runtime operator scheduling overhead; applying quantization techniques (such as INT8 / INT4) to compress model weights and activation values ​​to reduce memory usage and improve computational throughput; or enabling a key-value cache (KV Cache) mechanism to reuse key-value data from historical attention mechanisms, thereby accelerating the autoregressive generation process.

[0079] It should be noted that the above are only some examples of training deployment operations and inference deployment operations, and the embodiments disclosed herein are not limited thereto.

[0080] As mentioned above, training deployment and inference deployment are essentially different technical systems, involving completely different technology stacks. Therefore, loading all the configuration information required for training deployment and inference deployment in a unified manner would result in loading unnecessary graph pattern configurations, quantization parameters, KV cache size, and other information when only training deployment operations are needed; or loading unnecessary distributed training strategy configurations, gradient synchronization parameters, checkpoint paths, and other information when only inference deployment operations are needed. This would lead to the reading and initialization of a large amount of irrelevant content, reducing model adaptation efficiency.

[0081] In the model adaptation method provided in this embodiment, the following steps can be performed before step S101: determining the target deployment sub-mode based on the configuration information of the model. In this way, different deployment sub-modes are determined at the front end of the model adaptation pipeline based on the task type information (e.g., use case field) in the model configuration information, and each only loads the technology stack it needs, avoiding the reading and initialization of irrelevant content, which can effectively improve the model adaptation efficiency.

[0082] An example of step S101 may include the following steps S1011 to S1013.

[0083] Step S1011: In response to the model correctly completing the forward propagation on the target hardware device, determine the target adaptation mode as the direct deployment mode.

[0084] Step S1012: In response to the existence of operator errors during the forward propagation of the model on the target hardware device, the target adaptation mode is determined to be the fixed-point repair mode.

[0085] Step S1013: In response to an error in model loading on the target hardware device, determine that the target adaptation mode is the full adaptation mode.

[0086] In step S1011, the forward propagation is completed correctly, indicating that the forward propagation is running normally and the output result is correct.

[0087] In step S1012, operator errors may include dimension mismatch, data type incompatibility, index out of bounds, etc., that is, various errors caused by abnormal execution of the operator itself. This embodiment of the disclosure does not limit these errors. For example, if the forward propagation of the model is interrupted due to an operator error, or if the forward propagation output does not meet expectations (e.g., is incorrect), then it can be determined that an operator error exists.

[0088] In step S1012, in response to the existence of operator errors during the forward propagation of the model on the target hardware device, the operators to be repaired can also be recorded.

[0089] In step S1013, loading exceptions may include loading failure (LOAD_FAIL) or memory overflow (OOM), etc.

[0090] An example of step S101 may include step S201 as follows.

[0091] Step S201: Based on the forward propagation running status of the model on the target hardware device and combined with user instruction information, determine the target adaptation mode from multiple adaptation modes.

[0092] For example, user instructions can be entered via command line or read from a configuration file, and this disclosure does not limit this. These user instructions are used to indicate the initial adaptation mode to reflect the user's intent, i.e., to clearly define the adaptation mode the user expects to adopt. In subsequent steps, the initial adaptation mode can be further verified to determine whether it can be adopted.

[0093] In some examples, step S201 may include steps S2011 to S2012.

[0094] Step S2011: In response to the user instruction information indicating that the initial adaptation mode is the full adaptation mode, determine the target adaptation mode as the full adaptation mode.

[0095] Step S2012: In response to the user instruction information indicating that the initial adaptation mode is direct deployment mode or fixed-point repair mode, the target adaptation mode is determined from multiple adaptation modes based on the forward propagation running status of the model on the target hardware device.

[0096] Using the above method, when the user instruction indicates that the initial adaptation mode is the full adaptation mode, the smoke test can be skipped, and the target adaptation mode can be directly determined as the full adaptation mode, thereby improving model adaptation efficiency. When the user instruction indicates the use of the direct deployment mode or the targeted repair mode, the initial adaptation mode needs to be verified through a smoke test to determine whether the initial adaptation mode can be used. If the target adaptation mode determined from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device is different from the initial adaptation mode indicated by the user instruction, the initial adaptation mode verification fails.

[0097] For a specific example of "determining the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device" in step S2012, please refer to steps S1011 to S1013 above.

[0098] In the model adaptation method provided in at least one embodiment of this disclosure, by selecting the target adaptation mode based on the smoke test results and user intent, the accuracy of routing decisions and the efficiency of model adaptation can be further improved.

[0099] The model adaptation method provided in at least one embodiment of this disclosure may further include steps S103 to S105.

[0100] Step S103: Determine the prototype type of the model based on the model's configuration information.

[0101] Step S104: Determine the operator optimization strategy for the model based on the prototype category of the model.

[0102] Step S105: In response to the completion of the deployment operation, perform operator optimization operations based on the operator optimization strategy.

[0103] For example, in step S103, prototype classification represents classifying models into several standard types based on model architecture characteristics, each type having similar adaptation critical paths and optimization strategies. For instance, the prototype category of a model can be determined based on architectural signals (such as the number of model parameters, model architecture type, use case type, etc.) in the model's configuration information. An example classification is as follows: Hybrid Expert Large Language Model (MoE LLM), Visual Language Multimodal Model (VL Multimodal Model), Dense Large Language Model (Dense LLM), Large Dense Large Language Model (Large Dense LLM), Convolutional Neural Network or Computer Vision (CNN / CV) Model, and Diffusion Model. Based on the above classification examples, models can be further classified into training models and inference models according to the use case field in the model's configuration information.

[0104] Table 1 shows an example of determining the prototype category of a model based on architectural signals.

[0105] Table 1. Correspondence between architectural signals and prototype categories

[0106] For example, as shown in Table 1, if the configuration file includes the num_experts field, the prototype category of the model can be determined to be MoE LLM (the default is the inference model); if the configuration file includes the num_experts field and the use case is training, the prototype category of the model can be determined to be MoE LLM+trained model.

[0107] For example, if the configuration file includes the vision_config or image_size field, the prototype category of the model can be determined to be a VL multimodal model (the default is an inference model); if the configuration file includes the vision_config or image_size field and the use case is training, the prototype category of the model can be determined to be a VL multimodal model + training model.

[0108] For example, if the architecture class in the configuration file includes the ForCausalLM field and the number of parameters is less than 70B, the prototype category of the model can be determined to be dense LLM (the default is inference model); if the architecture class in the configuration file includes the ForCausalLM field and the number of parameters is less than 70B and the use case is training, the prototype category of the model can be determined to be dense LLM + training model.

[0109] The correspondence between other architectural signals and prototype categories follows the same pattern, as detailed in Table 1, and will not be elaborated upon here. It should be noted that the above classification method is merely an example; different classification methods and criteria can be adopted according to actual needs, and this disclosure does not impose any limitations on this.

[0110] In other examples, the prototype category of the model can be determined based on the prototype category confidence level.

[0111] For example, one example of step S103 may include steps S1031 to S1032 as follows.

[0112] Step S1031: For the first prototype category among multiple preset prototype categories, determine the confidence level of the first prototype category based on the classification signal corresponding to the first prototype category in the configuration information.

[0113] Step S1032: In response to the first prototype category confidence level meeting the preset condition, determine the prototype category of the model as the first prototype category.

[0114] For example, the multiple preset prototype categories may include MoE LLM, VL multimodal model, dense LLM, large dense LLM, CNN / CV model, diffusion model, and training model as shown in the examples above, and may also include other prototype categories. This disclosure does not limit the scope of these categories. The first prototype category can be any one of the multiple preset prototype categories. In step S1031, the confidence level of the first prototype category can be obtained by traversing the classification signals associated with the first prototype category and calculating their matching degree.

[0115] In some examples, step S1031 can be calculated using the following formula:

[0116] in, Indicates the first prototype category; This represents the confidence level of the first prototype category; n represents the total number of classification signals associated with the first prototype category. This represents the i-th classification signal; Indicates configuration information; This represents the preset weight corresponding to the i-th classification signal; This indicates an indicator function.

[0117] For example, indicator functions This is used to indicate whether the i-th classification signal exists in the configuration information. Specifically, if the i-th classification signal exists in the configuration information, the indicator function takes the first value (e.g., 1); if the i-th classification signal does not exist in the configuration information, the indicator function takes the second value (e.g., 0).

[0118] In step S1032, the preset condition may be that the confidence level of the first prototype category is greater than or equal to a preset threshold. For example, if If so, then the current model is determined to belong to the first prototype category; if If so, it can be determined that the current model does not belong to the first prototype category. It should be noted that the preset conditions can be set according to actual needs, and this embodiment does not limit them. If the model does not belong to any prototype category, its prototype category can be determined as a preset general type, and a general optimization strategy matching the general type can be called in subsequent operator optimization operations.

[0119] For ease of understanding, the following explanation will use MoE LLM as an example, which is the first prototype category.

[0120] For example, suppose the classification signals associated with the MoE LLM include: the number of experts (num_experts), the Top-k value (topk), and the shared expert identifier (shared_expert). In this case, the total number of classification signals is n=3.

[0121] During the calculation, the above three signals are traversed. If the configuration information includes num_experts, the corresponding indicator function value is 1, and its corresponding weight is accumulated. Similarly, if the configuration information includes topk and shared_expert, then the corresponding weights are accumulated respectively. and Finally, the prototype category confidence score of the MoE LLM is calculated. If this confidence score exceeds a preset threshold... If so, the prototype category of the model is confirmed to be MoE LLM.

[0122] In some examples, for each prototype category, information such as the set of operators most likely to be missing for that prototype category, operator bottlenecks (also known as the critical path), recommended deployment frameworks, preferred optimization strategies, and key accuracy verification categories can be predefined. The critical path can refer to a sequence of consecutive operators in the model's computation graph or execution plan topology that takes the longest time from input to output and determines the overall computational latency.

[0123] Table 2 shows an example of predefined optimization information for different prototype categories.

[0124] Table 2 Optimization information for different prototype categories

[0125] For example, in step S104, the operator optimization strategy for the model can be determined based on the model's prototype category. In some examples, a predefined optimization strategy corresponding to the model's prototype category can be used as the model's operator optimization strategy. In other examples, the model's operator optimization strategy can be determined based on a predefined critical path corresponding to the model's prototype category.

[0126] For example, as shown in Table 2, if the prototype category of the model is determined to be dense LLM, its corresponding critical paths can be identified as graph pattern capture and KV cache. Furthermore, for the graph pattern capture path, a graph pattern optimization strategy can be adopted, such as reducing graph capture overhead through operator fusion or static shape derivation; for the KV cache processing path, a KV-INT8 quantization strategy can be adopted, that is, compressing the data precision of the KV cache from floating-point (e.g., FP16) to 8-bit integer (INT8) to reduce memory bandwidth requirements and improve computational efficiency.

[0127] For example, if the prototype category of the model is determined to be a large dense LLM, its corresponding critical paths can be identified as cross-die tensor parallelism (TP) and memory capacity constraints. Further, for the memory capacity-constrained path, a quantization strategy (e.g., W8A16 quantization) can be adopted to reduce the model's memory usage by decreasing the numerical precision of model weights or activation values. For the cross-die TP path, a TP sharding strategy can be adopted to reduce the amount of data and latency in cross-die communication by optimizing tensor partitioning dimensions or employing communication and computation overlap techniques, thereby improving overall efficiency.

[0128] For example, if the prototype category of the model is determined to be MoE LLM, its corresponding critical path can be identified as expert routing and all-to-all communication. Furthermore, for the above critical path, a communication-computation overlap strategy can be adopted.

[0129] For example, if the prototype category of the model is determined to be a VL multimodal model, its corresponding critical path can be identified as a dual-pipeline parallel processing path consisting of a visual encoder and an LLM. Furthermore, for the aforementioned critical path, an optimization strategy using the Vision Transformer (ViT) operator or an image preprocessing optimization strategy can be determined.

[0130] For example, if the prototype category of the model is determined to be a CNN / CV model, its corresponding critical paths can be identified as Convolution-Batch Normalization (Conv-BN) fusion and lack of FP16 precision support. Further, for the Conv-BN fusion path, an operator fusion strategy can be adopted to fuse the convolution operator and the batch normalization operator into a single operator, eliminating redundant computation and data transfer. For the path lacking FP16 precision support, a mixed precision strategy can be adopted. For operators that do not support FP16 precision, full-precision (FP32) computation mode is automatically identified and retained, while half-precision computation mode is used for other operators that support FP16, maximizing computational throughput while ensuring model accuracy.

[0131] For example, if the prototype category of the model is determined to be a diffusion model, its corresponding critical paths can be identified as long sequence attention and multi-core communication. Specifically, when dealing with extremely long contexts, the computational cost of self-attention increases quadratically with the sequence length. Furthermore, in a multi-core architecture, the computation and reduction of the attention matrix involve a large amount of die-to-die communication, leading to a significant increase in computational and communication latency. Further, for the aforementioned critical paths, attention optimization strategies can be adopted to reduce the computational complexity of long sequences and optimize cross-core data interaction. For example, linear attention or sparse attention can be used instead of standard full attention computation to reduce computational complexity; a block-wise attention strategy can be adopted to divide the long sequence into sub-blocks adapted to the memory capacity of a single core, compute local attention in parallel within the core, and only exchange necessary statistics or intermediate results between cores, thereby reducing communication bandwidth pressure.

[0132] For example, if the prototype category of the model is determined to be a training model, its corresponding critical path can be identified as gradient synchronization, optimizer, and checkpointing. Specifically, as the number of model parameters and devices increases, the communication overhead generated by gradient synchronization increases linearly. Optimizer states (such as momentum and variance) consume a large amount of GPU memory, limiting the batch size. Furthermore, disk read / write operations during checkpoint saving block the computation flow. These factors collectively constitute the bottleneck of training efficiency. Further, for the above critical paths, a communication overlap strategy or a ZeRO strategy can be adopted. The ZeRO strategy is used to shard the optimizer state, gradients, and model parameters, distributing the data that was originally redundantly stored on each device evenly across all devices participating in training, thus eliminating GPU memory redundancy. The communication overlap strategy is used to schedule the communication tasks of gradient synchronization or checkpoint saving in parallel with the model's forward / backward computation tasks in time, thus masking communication and input / output delays.

[0133] In some examples, steps S103 and S104 can be performed before step S102 (or step S101), thereby determining whether to use the inference deployment sub-mode or the training deployment sub-mode in subsequent deployment operations while determining the prototype category.

[0134] For example, step S105 can be executed after step S102. In step S105, in response to completing the aforementioned deployment operation, an operator optimization operation can be performed based on the operator optimization strategy determined in step S104.

[0135] In the model adaptation method provided in at least one embodiment of this disclosure, the prototype category of the model is automatically identified based on the model configuration file. This prototype classification is then introduced into the model adaptation process, and different operator optimization strategies are applied to different model categories. This effectively avoids resource redundancy caused by traditional globally unified optimization strategies, allowing optimization actions to be precisely applied to operator bottlenecks, thereby achieving more significant differentiated performance gains. Furthermore, prototype classification knowledge can be reused across models, effectively reducing the additional computational resources required for new model adaptation requests.

[0136] The model adaptation method provided in at least one embodiment of this disclosure may further include the following steps S106 to S107.

[0137] Step S106: During the execution of each operation included in the method, in response to the completion of the current operation, the execution result of the current operation is written to the status log file.

[0138] Step S107: In response to the existence of an operation interruption, determine the next operation to be executed based on the operation corresponding to the last execution result written in the status record file.

[0139] Table 3 shows an example of a data structure for a status log file.

[0140] Table 3 Data Structure of Status Log File

[0141] It should be noted that Table 3 is only an example, and more or fewer fields may be included depending on actual needs. This disclosure does not limit this.

[0142] In the model adaptation method provided in at least one embodiment of this disclosure, a pipeline state persistence mechanism can be employed to ensure the continuity and reliability of the task. For example, after each atomic step in the model adaptation pipeline is executed, its execution result is synchronously written to a predefined structured pipeline state record file. When the pipeline operation is interrupted due to an anomaly or external intervention, this structured state can be read, the breakpoint of any step can be accurately located, and execution can be resumed from the point of interruption. Since the execution results of historical steps have been persistently recorded, the resumed pipeline will directly load the existing results, thereby avoiding redundant repeated calculations, improving iteration efficiency by more than 50%, and significantly enhancing the robustness of the system in complex operating environments.

[0143] Figure 2 A flowchart of another model adaptation method provided for at least one embodiment of this disclosure.

[0144] For example, such as Figure 2As shown, the model adaptation method provided in at least one embodiment of this disclosure may include the following steps S301 to S311.

[0145] Step S301: Obtain the model's configuration information.

[0146] For example, in step S301, the model's configuration information can be obtained by parsing the model's configuration file.

[0147] Step S302: Determine the prototype category of the model based on the model's configuration information.

[0148] Step S303: Determine the operator optimization strategy for the model based on the prototype category of the model.

[0149] For a description of steps S302 and S303, please refer to the description of steps S103 and S104 above, which will not be repeated here.

[0150] Step S304: Perform environment verification operation.

[0151] For example, in step S304, the environment verification operation can represent the adaptation and compliance verification process for the model's runtime environment. Since models of different prototype categories differ in terms of computing resources, dependency library versions, or hardware interfaces, a verification strategy or script matching the prototype category determined in step S302 can be dynamically loaded. During actual environment verification, model information corresponding to the prototype category can be extracted, and targeted environment detection tasks can be executed based on the verification logic indicated by the model information to ensure the model's availability and stability in the target runtime environment. In some examples, target hardware device detection (e.g., target hardware device model, computing power, video memory capacity, etc.), compiler version verification, and deep learning framework availability verification can be performed, using a three-level check to determine if the execution environment is ready.

[0152] Step S305: Based on the forward propagation running state of the model on the target hardware device, determine the target adaptation mode from multiple adaptation modes.

[0153] For a description of step S305, please refer to the description of step S101 above; it will not be repeated here.

[0154] Step S306: If the target adaptation mode is direct deployment mode, proceed to step S307; if the target adaptation mode is targeted repair mode, proceed to step S308; if the target adaptation mode is full adaptation mode, proceed to step S309.

[0155] For example, in steps S305 to S306, the target adaptation mode to be executed can be determined from the three adaptation modes, and only the operation steps corresponding to the target adaptation mode are executed.

[0156] Step S307: Perform the deployment operation.

[0157] Step S308: Perform operator detection operation, operator repair operation and deployment operation.

[0158] Step S309: Perform operator analysis, operator coverage check, operator generation, operator integration, and deployment operations.

[0159] For a description of steps S307 to S309, please refer to the description of step S102 above; it will not be repeated here.

[0160] Step S310: Perform operator optimization operations based on the operator optimization strategy.

[0161] For a description of step S310, please refer to the description of step S105 above; it will not be repeated here.

[0162] Step S311: Generate a model adaptation report.

[0163] In the model adaptation method provided in at least one embodiment of this disclosure, by automatically determining the model's adaptation status and prototype category through actual running status at the entry point of the adaptation process, selecting the shortest execution path, and customizing the critical path and optimization strategy according to the model prototype category, the model adaptation efficiency can be greatly improved, reducing the adaptation time of known available models from hours to minutes, and effectively overcoming the efficiency bottleneck of linear model adaptation pipelines.

[0164] The following example analyzes the improvement in model adaptation efficiency achieved by the model adaptation method provided in at least one embodiment of this disclosure.

[0165] For example, the execution time T_full of the model adaptation pipeline using the full adaptation mode is calculated as follows: T_full = t_env + t_analysis + t_gap + t_gen + t_integrate + t_verify + t_deploy + t_optimize, The execution time T_A of the model adaptation pipeline using the direct deployment mode is calculated as follows: T_A = t_env + t_smoke + t_deploy + t_optimize, The execution time T_B of the model adaptation pipeline using the fixed-point repair mode is calculated as follows: T_B = t_env + t_smoke + t_runtime_probe + t_fix + t_deploy + t_optimize, Wherein, t_env represents the execution time of the environment verification operation, t_smoke represents the execution time of the smoke test operation, t_analysis represents the execution time of the operator analysis operation, t_gap represents the execution time of the operator coverage check operation, t_gen represents the execution time of the operator generation operation, t_integrate represents the execution time of the operator integration operation, t_verify represents the execution time of the performance verification operation, t_deploy represents the execution time of the deployment operation, t_optimize represents the execution time of the operator optimization operation, t_runtime_probe represents the execution time of the operator probing operation, and t_fix represents the execution time of the operator repair operation.

[0166] The time savings compared to the full adaptation mode in the direct deployment mode are as follows: Speedup_A = T_full / T_A, The time savings of the targeted repair mode compared to the full adaptation mode are as follows: Speedup_B = T_full / T_B, In typical scenarios, t_analysis + t_gap + t_gen + t_integrate account for 60-80% of the total time. Therefore, the direct deployment mode can achieve a 3-5x speedup, and the targeted repair mode can achieve a 1.5-2x speedup. That is, the adaptation time for known available models can be reduced from hours to minutes, and the repair efficiency of some adapted models can be improved by 1.5-2 times.

[0167] The model adaptation method provided in at least one embodiment of this disclosure can effectively reduce the manpower required for model adaptation engineers, reducing the cost of adapting a single model by more than 60%. Furthermore, it can shorten the model deployment cycle and accelerate the development of the chip ecosystem.

[0168] It should also be noted that the execution order of the various steps of the model adaptation method in the various embodiments of this disclosure is not limited. Although the execution process of each step has been described in a specific order above, this does not constitute a limitation on the embodiments of this disclosure. The various steps in the model adaptation method can be executed sequentially or in parallel, which can be determined according to actual needs.

[0169] For example, compared to the above description, the model adaptation method provided in at least one embodiment of this disclosure may include more or fewer steps, and the embodiments of this disclosure do not limit this.

[0170] Figure 3 This is a schematic block diagram of a model adaptation system provided for at least one embodiment of the present disclosure.

[0171] For example, such as Figure 3 As shown, the model adaptation system provided in at least one embodiment of this disclosure includes an entry layer, a rapid verification layer, an adaptation layer, a deployment layer, an optimization layer, and a state management layer.

[0172] For example, the entry layer may include a prototype classification module and an adaptation mode routing module. For instance, the prototype classification module can be used to perform step S103 above, automatically determining the prototype category of the model based on the model configuration information, and further determining the critical path for subsequent operator optimization based on the prototype category. In some examples, the prototype classification module can also be configured to determine whether to use the inference deployment sub-mode or the training deployment sub-mode in subsequent deployment operations. The adaptation mode routing module can be used to perform step S101 above, determining the target adaptation mode based on the smoke test results, to determine which modules need to be called after entering the adaptation layer.

[0173] In some examples, the entry layer may also include a requirements gathering module. For instance, the requirements gathering module may be configured to receive model configuration information and user instructions, and provide this information to the prototype classification module or the adaptation mode routing module. The requirements gathering module may also be configured to receive user requests, which may carry model identifiers, usage scenarios, and target hardware unit information.

[0174] For example, the fast verification layer may include an environment verification module and a fast smoke test module. For example, the environment verification module may be configured to perform environment verification operations; the fast smoke test module may be configured to perform fast smoke test operations and provide the smoke test results to the adaptation mode routing module.

[0175] For example, the adaptation layer may include an operator analysis module, a runtime detection module, an operator generation / repair module, and an operator integration module. The operator analysis module is used to perform operator analysis operations; the runtime detection module can be configured to perform operator coverage checks or operator detection operations; the operator generation / repair module can be configured to perform operator generation or operator repair operations; and the operator integration module can be configured to perform operator integration operations.

[0176] For example, when the target adaptation mode is full adaptation mode, it is necessary to call the operator analysis module, runtime detection module, operator generation / repair module and operator integration module in the adaptation layer; when the target adaptation mode is fixed-point repair mode, it is necessary to call the runtime detection module and operator generation / repair module in the adaptation layer; when the target adaptation mode is direct deployment mode, it is not necessary to call the functional modules in the adaptation layer, and the functional modules in the deployment layer can be called directly.

[0177] For example, the deployment layer may include an inference deployment module and a training deployment module. For instance, the inference deployment module may be configured to perform inference deployment operations based on the inference deployment sub-mode; the training deployment module may be configured to perform training deployment operations based on the training deployment sub-mode.

[0178] For example, the optimization layer may include an operator optimization module. This module can be configured to perform operator optimization operations, such as target-driven closed-loop performance optimization. Specifically, the operator optimization module can determine the model's operator optimization strategy based on the critical path corresponding to the model's prototype category, and then perform operator optimization operations based on that strategy.

[0179] For example, the state management layer may include a pipeline state management module. For example, the pipeline state management module can be used to execute the above steps S106 to S107, so that the pipeline state runs through all layers and supports breakpoint recovery.

[0180] Figure 4 This is a schematic block diagram of a model adaptation device provided for at least one embodiment of the present disclosure.

[0181] The model adapter can be implemented as a processor or set in a processor. The processor may include a central processing unit, a graphics processing unit, a general-purpose graphics processing unit, a tensor processor, a deep learning processor, an accelerator processor, a neural network processor, an application-specific integrated circuit, or a field-programmable gate array, etc. Of course, the embodiments disclosed herein are not limited to these, and the processor may also be any other type of processor.

[0182] For example, such as Figure 4 As shown, the model adaptation device 400 provided in at least one embodiment of this disclosure may include a determination module 401 and a deployment module 402.

[0183] For example, the determination module 401 is configured to determine a target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device, wherein the operation link lengths of different adaptation modes are different. An example of the determination module could be... Figure 3 The adaptation mode routing module in the model adaptation system.

[0184] For example, deployment module 402 is configured to deploy the model to the target hardware device based on the target adaptation mode.

[0185] In some examples, multiple adaptation modes include direct deployment mode, targeted repair mode, and full adaptation mode. The operation chain length of full adaptation mode is longer than that of targeted repair mode, and the operation chain length of targeted repair mode is longer than that of direct deployment mode.

[0186] In some examples, the direct deployment mode includes deployment operations, the targeted repair mode includes operator detection operations, operator repair operations, and deployment operations, and the full adaptation mode includes operator analysis operations, operator coverage check operations, operator generation operations, operator integration operations, and deployment operations.

[0187] In some examples, the determination module is further configured to: determine the target adaptation mode as direct deployment mode in response to the model correctly completing the forward propagation on the target hardware device; determine the target adaptation mode as fixed-point repair mode in response to the operator error during the forward propagation of the model on the target hardware device; and determine the target adaptation mode as full adaptation mode in response to the abnormal loading of the model on the target hardware device.

[0188] In some examples, the deployment operation includes: deploying the model to a target hardware device based on a target deployment sub-pattern, wherein the target deployment sub-pattern is determined based on the model's configuration information and includes at least one of an inference deployment sub-pattern and a training deployment sub-pattern.

[0189] In some examples, the model adaptation device may also include a deployment mode determination module, which is configured to determine the target deployment sub-mode based on the model's configuration information.

[0190] In some examples, the determination module is further configured to: determine the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device and in conjunction with user instruction information.

[0191] In some examples, the model adaptation device may also include a prototype classification module and an operator optimization module. The prototype classification module is configured to determine the prototype category of the model based on the model's configuration information; and to determine the operator optimization strategy of the model based on the prototype category. The operator optimization module is configured to perform operator optimization operations based on the operator optimization strategy in response to the completion of the deployment operation.

[0192] In some examples, the prototype classification module is further configured to: for a first prototype category among multiple preset prototype categories, determine the confidence level of the first prototype category based on the classification signal corresponding to the first prototype category in the configuration information; and in response to the first prototype category confidence level meeting preset conditions, determine the prototype category of the model as the first prototype category.

[0193] In some examples, the model adaptation device may also include a pipeline state management module. The pipeline state management module can be configured to write the execution result of the current operation to a state log file in response to the completion of the current operation during the execution of each operation included in the method.

[0194] In some examples, the pipeline state management module can be further configured to: in response to an operation interruption, determine the next operation to be executed based on the operation corresponding to the last execution result written in the state log file.

[0195] It should be noted that the above-mentioned modules can be implemented by software, hardware, firmware or any combination thereof. For example, the determination module can be implemented as a determination circuit, and the deployment module can be implemented as a deployment circuit. The embodiments of this disclosure do not limit their specific implementation methods.

[0196] It should be understood that the model adaptation device provided in at least one embodiment of this disclosure can be used to implement the aforementioned model adaptation method and can also achieve similar technical effects as the aforementioned model adaptation method, which will not be elaborated here.

[0197] It should be noted that in the embodiments of this disclosure, the model adaptation device may include more or fewer modules or units, and the connection relationship between the various modules or units is not limited and can be determined according to actual needs. The specific configuration of each module or unit is not limited and can be constructed from analog devices, digital chips, or other suitable methods according to circuit principles.

[0198] At least one embodiment of this disclosure also provides a model adaptation system, which may include the model adaptation device and target hardware device in the above embodiments.

[0199] Figure 5 This is a schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.

[0200] For example, such as Figure 5 As shown, the electronic device 500 includes at least one processor 501 and at least one memory 502. The at least one memory 502 includes one or more computer program modules. These computer program modules are stored in the memory 502 and configured to be executed by the at least one processor 501. The one or more computer program modules include instructions for performing the model adaptation method described above. When executed by the at least one processor 501, they can perform one or more steps of the model adaptation method provided in at least one embodiment of this disclosure. The memory 502 and the processor 501 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0201] For example, processor 501 can be a central processing unit (CPU), digital signal processor (DSP), graphics processing unit (GPU), general-purpose graphics processing unit (GPGPU), artificial intelligence (AI) accelerator, or other form of processing unit with data processing and / or program execution capabilities, such as a field-programmable gate array (FPGA); for example, the central processing unit (CPU) can be an x86, ARM, or RISC-V architecture. Processor 501 can be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 500 to perform desired functions.

[0202] For example, memory 502 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.

[0203] Figure 6 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.

[0204] The electronic devices in at least one embodiment of this disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0205] The electronic device includes at least one processor and a memory. The processor may be referred to as processing device 601 as described below, and the memory may include at least one of ROM 602, RAM 603, and storage device 608 as described below. The memory is used to store programs for performing the methods described in the various method embodiments above; the processor is configured to execute the programs stored in the memory. The processor may include a central processing unit (CPU) or other forms of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0206] like Figure 6As shown, the electronic device 600 may include a processing unit 601 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in ROM 602 or a program loaded from storage device 608 into RAM 603. RAM 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interfaces are also connected to bus 604.

[0207] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, displays, speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0208] In particular, according to at least one embodiment of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, at least one embodiment of this disclosure includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a processing device 601, it performs the functions defined in the methods of at least one embodiment of this disclosure.

[0209] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In at least one embodiment of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In at least one embodiment of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.

[0210] The aforementioned computer-readable medium may be included in the aforementioned electronic device 600; or it may exist independently and not assembled into the electronic device 600.

[0211] Figure 7 This is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.

[0212] For example, such as Figure 7 As shown, a non-transitory computer-readable storage medium 700 stores computer-readable instructions 701, which, when executed by at least one processor, perform one or more steps of the model adaptation method described above.

[0213] For example, the storage medium may include a memory card for a smartphone, a storage component for a tablet computer, a hard drive for a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other suitable storage media. For example, the readable storage medium may also be... Figure 5 The memory 502 in the memory is described in the foregoing content and will not be repeated here.

[0214] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to the embodiments of the present disclosure, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present disclosure are within the scope of protection claimed by the present disclosure.

[0215] The following points should be noted regarding this disclosure: (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0216] (2) For clarity, the thickness of layers or regions in the drawings used to describe embodiments of the present disclosure is enlarged or reduced, i.e., these drawings are not drawn to actual scale.

[0217] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0218] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.

Claims

1. A model adaptation method, characterized in that, The method includes: Based on the forward propagation running state of the model on the target hardware device, the target adaptation mode is determined from multiple adaptation modes; The model is deployed to the target hardware device based on the target adaptation mode. Among the multiple adaptation modes, the operation link lengths of different adaptation modes are different.

2. The model adaptation method according to claim 1, characterized in that, The multiple adaptation modes include direct deployment mode, targeted repair mode, and full adaptation mode. The operation link length of the full-scale adaptation mode is greater than that of the fixed-point repair mode, and the operation link length of the fixed-point repair mode is greater than that of the direct deployment mode.

3. The model adaptation method according to claim 2, characterized in that, The direct deployment mode includes deployment operations; the fixed-point repair mode includes operator detection operations, operator repair operations, and deployment operations; and the full-scale adaptation mode includes operator analysis operations, operator coverage check operations, operator generation operations, operator integration operations, and deployment operations.

4. The model adaptation method according to claim 2 or 3, characterized in that, The determination of the target adaptation mode from multiple adaptation modes based on the forward propagation operation state of the model on the target hardware device includes: In response to the model correctly completing the forward propagation on the target hardware device, the target adaptation mode is determined to be the direct deployment mode; In response to an operator error occurring during the forward propagation of the model on the target hardware device, the target adaptation mode is determined to be the fixed-point repair mode; In response to an error in loading the model on the target hardware device, the target adaptation mode is determined to be the full adaptation mode.

5. The model adaptation method according to claim 3, characterized in that, The deployment operation includes: The model is deployed to the target hardware device based on the target deployment sub-pattern. The target deployment sub-mode is determined based on the configuration information of the model, and the target deployment sub-mode includes at least one of the inference deployment sub-mode and the training deployment sub-mode.

6. The model adaptation method according to claim 5, characterized in that, Before determining the target adaptation mode from multiple adaptation modes, the method further includes: The target deployment sub-mode is determined based on the configuration information of the model.

7. The model adaptation method according to claim 1, characterized in that, The determination of the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device includes: Based on the forward propagation operation status of the model on the target hardware device and combined with user instruction information, the target adaptation mode is determined from the plurality of adaptation modes.

8. The model adaptation method according to claim 3, characterized in that, The method further includes: The prototype category of the model is determined based on the model's configuration information; The operator optimization strategy for the model is determined based on the prototype category of the model; In response to the completion of the deployment operation, an operator optimization operation is performed based on the operator optimization strategy.

9. The model adaptation method according to claim 8, characterized in that, The prototype category of the model is determined based on the model's configuration information, including: For the first prototype category among multiple preset prototype categories, the confidence level of the first prototype category is determined according to the classification signal corresponding to the first prototype category in the configuration information; In response to the first prototype category confidence level meeting a preset condition, the prototype category of the model is determined to be the first prototype category.

10. The model adaptation method according to claim 1, characterized in that, The method further includes: During the execution of each operation included in the method, in response to the completion of the current operation, the execution result of the current operation is written to a status log file.

11. The model adaptation method according to claim 10, characterized in that, The method further includes: In response to an operation interruption, the next operation to be executed is determined based on the operation corresponding to the last execution result written in the status record file.

12. A model adaptation device, characterized in that, The device includes: The determination module is configured to determine the target adaptation mode from multiple adaptation modes based on the forward propagation running state of the model on the target hardware device; The deployment module is configured to deploy the model to the target hardware device based on the target adaptation mode. Among the multiple adaptation modes, the operation link lengths of different adaptation modes are different.

13. An electronic device, characterized in that, The electronic device includes: At least one processor; At least one memory, including one or more computer program modules; The one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the method according to any one of claims 1-11.

14. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer-readable instructions, wherein the computer-readable instructions, when executed by at least one processor, perform the method according to any one of claims 1-11.