A model deployment method, device, storage medium, and electronic device

By using proxy model and label distribution adjustment technology, the exploration process of model deployment strategy is optimized, the problem of high resource occupation and time-consuming in the existing technology is solved, and more efficient model deployment is achieved.

CN119883295BActive Publication Date: 2025-06-10ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386144.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-10
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art requires exploration of a large number of parameter combinations during model deployment, resulting in high resource utilization, long time and low efficiency.

Method used

By obtaining the initial deployment policy group, the performance distribution information of the deployment policy is determined using the pre-trained proxy model, and the deployment policy is adjusted through the preset label distribution until the setting requirements are met, and the optimal deployment policy is determined.

Benefits of technology

It effectively reduces resource requirements and time loss during the exploration of model deployment strategy, and improves model deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883295B_ABST
    Figure CN119883295B_ABST
Patent Text Reader

Abstract

This specification discloses a model deployment method, apparatus, storage medium, and electronic device. The method includes: obtaining an initial deployment policy group for a service model, where the initial deployment policy group includes two deployment policies; inputting the feature encodings of the respective deployment policies in the initial deployment policy group into a pre-trained proxy model to determine the performance distribution information of the respective deployment policies on a processing device; using a preset label distribution to adjust at least one of the inputs, and obtaining an adjusted deployment policy group when the difference between the output performance distribution information and the label distribution meets a set requirement; determining a target deployment policy in the adjusted deployment policy group, and deploying the service model based on the target deployment policy. This solution reduces the time loss for exploring model deployment policies and improves the model deployment efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a model deployment method, apparatus, storage medium, and electronic device. Background Art

[0002] With the rapid development of deep learning technology, neural network models have achieved remarkable results in many fields. However, how to effectively deploy neural network models into actual production environments faces many challenges. Model deployment not only needs to consider the performance of the model, but also needs to take into account multiple factors such as computing resources, storage space, inference speed, and compatibility of the deployment environment.

[0003] Currently, neural architecture search (NAS) technology is usually used to explore the model architecture, and mapping space explore (MSE) technology is used to explore the configuration between the model and the processing device, so as to find the optimal model deployment strategy for deployment.

[0004] However, since the architecture parameters of the model and the configuration parameters between it and the processing device are usually discrete data, in the process of exploring the model architecture and configuration in the above way, a large number of parameter combinations need to be explored to determine different model deployment strategies, and it is necessary to test the actual model performance brought by each parameter combination to find the optimal model deployment strategy. This results in the model deployment process not only consuming a large amount of computing power resources, but also taking a long time and having low efficiency. Summary of the Invention

[0005] This specification provides a model deployment method, apparatus, storage medium, and electronic device to partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions:

[0007] A model deployment method, the method includes:

[0008] Obtain an initial deployment strategy group for a service model, the initial deployment strategy group includes two deployment strategies, and each deployment strategy contains architecture information of the service model and configuration mapping information between the service model and the configuration of the processing device;

[0009] Input the feature codes of each deployment strategy in the initial deployment strategy group into a pre-trained proxy model to determine the performance distribution information of each deployment strategy on the processing device;

[0010] Adjust at least one of the inputs using a preset label distribution, and obtain an adjusted deployment policy group when the difference between the performance distribution information of the output and the label distribution meets the set requirements;

[0011] Determine a target deployment policy in the adjusted deployment policy group, and deploy the business model based on the target deployment policy.

[0012] Optionally, for each deployment policy, the architecture information included in the deployment policy includes: the topological structure information corresponding to the business model and the operation type information corresponding to each network node in the business model.

[0013] Optionally, adjusting at least one of the inputs using a preset label distribution specifically includes:

[0014] Taking the deviation between the performance distribution information output by the proxy model and the label distribution to be maximized and the size relationship represented by the performance distribution information output by the proxy model to be the same as the size relationship represented by the label distribution as the goal, adjust one of the inputs.

[0015] Optionally, the two deployment policies include: a first deployment policy and a second deployment policy; the preset label distributions include: a first label distribution and a second label distribution; wherein, the size relationship between the performances of the two deployment policies represented by the first label distribution on the processing device is opposite to the size relationship between the performances of the two deployment policies represented by the second label distribution on the processing device;

[0016] Adjusting at least one of the inputs using a preset label distribution specifically includes:

[0017] Adjust the first deployment policy and the second deployment policy for multiple adjustment cycles using the preset label distribution, wherein the label distributions corresponding to two adjacent adjustment cycles are different;

[0018] For each adjustment cycle, if the label distribution corresponding to the adjustment cycle is the first label distribution, then adjust the first deployment policy in the adjustment cycle using the first label distribution until reaching the next adjustment cycle;

[0019] In the next adjustment cycle, adjust the second deployment policy in the next adjustment cycle using the second label distribution.

[0020] Optionally, adjusting at least one of the inputs using a preset label distribution specifically includes:

[0021] Determine the probability distribution data corresponding to the at least one input, and determine the gradient information of the performance distribution information with respect to the initial deployment policy group;

[0022] Adjust the probability distribution data according to the gradient information by using the preset tag distribution to obtain adjusted probability distribution data;

[0023] Based on the adjusted probability distribution data, determine the adjusted at least one input.

[0024] Optionally, adjusting at least one of the inputs by using a preset tag distribution specifically includes:

[0025] Determine the gradient information of the performance distribution information with respect to the initial configuration mapping information included in the initial deployment policy group, and determine the configuration constraint information corresponding to the initial configuration mapping information;

[0026] Under the constraint of the configuration constraint information, use the preset tag distribution to adjust the initial configuration mapping information according to the gradient information, and update the configuration constraint based on the adjusted mapping configuration information.

[0027] Optionally, training the proxy model specifically includes:

[0028] Obtain each historical deployment policy, and construct multiple historical policy groups based on each historical deployment policy; wherein, each historical policy group contains two historical deployment policies;

[0029] Determine the actual performance of each historical deployment policy on the processing device, and determine the actual performance distribution information corresponding to each historical deployment policy group according to the actual performance;

[0030] For each historical policy group, input the historical policy group into the proxy model to be trained, so as to determine the predicted performance distribution information corresponding to the historical policy group through the proxy model;

[0031] Determine the loss value of the proxy model according to the predicted performance distribution information and the actual performance distribution information corresponding to the historical policy group, and train the proxy model according to the loss value.

[0032] This specification provides a model deployment device, including:

[0033] An acquisition module, which acquires an initial deployment policy group for a service model, the initial deployment policy group includes two deployment policies, and each deployment policy contains the architecture information of the service model and the configuration mapping information between the service model and the configuration of the processing device;

[0034] A determination module, which inputs the feature codes of each deployment policy in the initial deployment policy group into a pre-trained proxy model to determine the performance distribution information of each deployment policy on the processing device;

[0035] An adjustment module that adjusts at least one input thereto by using a preset label distribution, and obtains an adjusted deployment policy group when the difference between the output performance distribution information and the label distribution meets the set requirements;

[0036] A deployment module that determines a target deployment policy in the adjusted deployment policy group and deploys the business model based on the target deployment policy.

[0037] This specification provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above model deployment method is implemented.

[0038] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above model deployment method is implemented.

[0039] The above at least one technical solution adopted in this specification can achieve the following beneficial effects:

[0040] In the model deployment method provided in this specification, an initial deployment policy group for a business model is obtained. The initial deployment policy group includes two deployment policies, and each deployment policy contains architecture information corresponding to the business model deployed by this deployment policy and configuration mapping information between the business model and the configuration of the processing device; the initial deployment policy group is input into a pre-trained proxy model to determine the performance distribution information between the deployment policies; according to the performance distribution information, at least one deployment policy in the initial deployment policy group is adjusted, and the above steps are repeated until the preset termination condition is met to determine the target deployment policy group; a target deployment policy is determined in the target deployment policy group, and the business model is deployed based on the target deployment policy.

[0041] As can be seen from the above method, in this solution, when performing model deployment, the proxy model can be used to determine the comparison relationship between the performances brought by multiple deployment policies in a group of deployment policies, and several adjustments are made to the deployment policies therein with this comparison relationship as a reference, so as to find the optimal deployment policy in the final deployment policy group to deploy the model. The entire process does not require separately testing the actual model performance brought by each model deployment policy, effectively reducing the resource requirements and time consumption in the process of exploring model deployment policies, and further improving the model deployment efficiency. Description of the Drawings

[0042] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:

[0043] Figure 1 Schematic diagram of a model deployment method provided in this specification;

[0044] Figure 2 Schematic diagram of the topological structure of a service model provided in this specification;

[0045] Figure 3 Schematic diagram of the feature encoding of architecture information provided in this specification;

[0046] Figure 4 Schematic diagram of the hardware structure of a processing device provided in this specification;

[0047] Figure 5 Schematic diagram of the model structure of an agent model provided in this specification;

[0048] Figure 6 Schematic diagram of the training process of a performance comparison model provided in this specification;

[0049] Figure 7 Schematic diagram of the construction process of sample data provided in this specification;

[0050] Figure 8 Schematic diagram of the determination process of a target deployment strategy provided in this specification;

[0051] Figure 9 Schematic diagram of the process of adjusting architecture information provided in this specification;

[0052] Figure 10 Schematic diagram of the process of adjusting configuration mapping information provided in this specification;

[0053] Figure 11 Schematic diagram of a projection process provided in this specification;

[0054] Figure 12 Schematic diagram of a model deployment device provided in this specification;

[0055] Figure 13 Provided in this specification corresponding to Figure 1 Schematic diagram of an electronic device. Specific implementation manners

[0056] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts belong to the scope protected by this specification.

[0057] The NAS technology aims to automate the network architecture design process to reduce its dependence on expert experience, and it constitutes an important part of Automatic Deep Learning (AutoDL). Among them, NAS usually involves three aspects: search space, search strategy, and evaluation strategy. The current search strategies usually adopt heuristic search strategies such as reinforcement learning, evolutionary algorithms, Bayesian deployment, etc.; the evaluation strategies usually adopt methods such as proxy datasets, proxy training, and proxy models. In terms of architecture evaluation, currently, most use proxy models for evaluation acceleration.

[0058] For the MSE technology, the hardware architecture of the processing device determines how to efficiently execute computationally intensive tasks such as Convolutional Neural Network (CNN). The key architecture components of the accelerator include Processing Elements (PEs), on-chip buffer, and control unit. The processing unit is responsible for parallel computing operations, the buffer is used to store intermediate results and reduce external memory access, and the control unit is responsible for scheduling operations. Effective MSE needs to consider how to map the computational steps of the algorithm to these hardware resources to minimize execution time, energy consumption, and storage bandwidth.

[0059] However, both mapping space search and neural network architecture search are complex combinatorial deployment problems, and their search spaces are extremely large. The neural architecture search space is usually discrete, non-convex, and extremely large in scale. At the same time, due to the discreteness and constraints of hardware resources, the performance surface of the mapping space is also usually non-convex and discrete. Therefore, the coupling between neural network architecture parameters and mapping space parameters will significantly increase the difficulty of space search. Existing black-box deployment (zero-order deployment) methods (such as genetic algorithms, reinforcement learning algorithms, etc.) are inefficient in searching the mapping space and are difficult to effectively find the global optimal mapping strategy.

[0060] A surrogate model or surrogate is a model used to simulate, approximate, or replace a complex, expensive, or infeasible real system or process. The representation of the data space by common surrogate models can be modeled as a regression model (numerical regression), a classification model, or a relational model, where the relational model can be a list type, a binary relation comparison type, etc. The selection of an appropriate surrogate model and method depends on the specific problem and the available data. Surrogate models play an important role in space exploration and deployment, thus effectively exploring and deploying complex system and design spaces.

[0061] If the modeling of a differentiable surrogate model is regarded as an approximate representation or surrogate representation of the original architecture space, then architecture differentiable deployment can be directly performed based on the surrogate model. However, the key difficulties and technical bottlenecks are as follows: The architecture space is a discrete space. Even if a differentiable surrogate model can realize the original space, the discrete architecture variables are still not differentiably deployable. The current differentiable architecture search methods mainly rely on the construction and training of differentiable hypernetworks, which require a large amount of computing power, video memory resources, and time. At the same time, the final performance of their deployment search is also limited.

[0062] Based on this, this specification provides a model deployment method. The following will combine the accompanying drawings to detail the technical solutions provided by each embodiment of this specification.

[0063] Figure 1 The following is a schematic flowchart of a model deployment method provided in this specification, including the following steps:

[0064] S101: Obtain an initial deployment policy group for the service model. The initial deployment policy group includes two deployment policies, and each deployment policy contains the architecture information of the service model and the configuration mapping information between the service model and the configuration of the processing device.

[0065] In this specification, the execution subject for implementing a model deployment method can be a specified device such as a server. For the convenience of description, hereinafter, only the server will be used as an example of the execution subject to illustrate a model deployment method provided in this specification.

[0066] In practical applications, the model architecture of a service model usually can include several optional parameter combinations. Among them, the architecture information corresponding to each parameter combination can include the topology structure information corresponding to the topology structure of the service model (i.e., the connection structure between each network node of the service model) and the operation type information corresponding to each network node in the service model (such as convolution (conv), pooling (pool), activation, up / down sampling, input (Input), output (Output), etc.).

[0067] In this specification, the above business model can be a neural network model for performing different services, such as a graphics recognition model, a large language model, a risk control model, a navigation model, an information recommendation model, etc. Of course, it can also include neural network models for performing other services, and this specification does not make specific limitations on this.

[0068] For ease of understanding, this specification provides a schematic diagram of the topological structure of a business model, as Figure 2 shown.

[0069] Among them, Figure 2 in (a) shows all the optional topological structures of the business model, including the possible connection situations between each network node, Figure 2 in (b) shows one of the topological structures. The operation types corresponding to each network node therein include: conv1x1, conv3x3, Max pool3x3, Input, and Output.

[0070] In addition, during the process of deploying the business model to the processing device, there can also be various mapping relationships between the business model and the configuration of the processing device, and these mapping relationships are used to represent how the business model and its network units are configured and executed on the processing device.

[0071] From the above content, it can be concluded that during the process of deploying the business model, the deployment strategy of the model can be jointly determined by the model architecture (topological structure and operation type) and the mapping relationship with the configuration of the processing device, that is, for each model deployment strategy, the model deployment strategy can include the architecture information corresponding to the business model deployed through this deployment strategy and the configuration mapping information between the business model and the configuration of the processing device.

[0072] During the actual exploration of the model deployment strategy, the server can randomly select two model deployment strategies and construct an initial deployment strategy group based on these two model deployment strategies as the initial value for exploring the model deployment strategy.

[0073] Specifically, for each model deployment strategy in the initial deployment strategy group, the server can perform feature encoding on the topological structure information of the business model in the model deployment strategy to obtain the encoded representation T corresponding to the topological structure information, and, perform encoding on the operation type information corresponding to each network node of the business model in the model deployment strategy to obtain the feature representation O corresponding to the operation type information.

[0074] For ease of understanding, this specification provides a schematic diagram of the feature encoding of the architecture information, as Figure 3 shown.

[0075] Among them, Figure 3Among them, (a) shows the feature encoding corresponding to the operation type in (a) of 2, Figure 3 and (b) shows Figure 2 the feature encoding corresponding to the topological structure in (b) of

[0076] During the process of deploying the business model to the processing device, there can also be various mapping relationships between the business model and the configuration of the processing device, and these mapping relationships are used to represent how the business model and its interaction with network units are configured and executed on the processing device.

[0077] In this specification, the processing device can be a Neural Processing Unit (NPU), and of course, it can also be other processing devices such as CPU, GPU, etc. This specification does not make specific limitations on this.

[0078] Taking the processing device being an NPU as an example, its hardware structure is as Figure 4 shown.

[0079] Figure 4 It is a schematic diagram of the hardware structure of a processing device provided in this specification.

[0080] Among them, the key architecture components of the NPU include PEs, cache, and control unit. The processing unit is responsible for parallel computing operations, the cache is used to store intermediate results and reduce external memory access, and the control unit is responsible for scheduling operations.

[0081] Based on the above structure, the mapping relationships between the business model and the configuration of the NPU can include:

[0082] Buffer Allocation: Buffer Allocation determines how the input, output, and weights are allocated to the on-chip cache to reduce memory bandwidth requirements.

[0083] For example, for the on-chip secondary cache, with 16 buffer numbers for each level of cache, the parameter example is: .

[0084] Tiling Size: The tiling strategy defines how the computation is processed in the accelerator cache and determines the granularity of parallel computing.

[0085] For example, for the on-chip secondary cache, with DRAM as the third-level storage, the 7-layer loop bound of the CNN atomic operation is:

[0086]

[0087] The number of tile divisions for each layer loop bound is the number of storage levels.

[0088] Loop Ordering: Defines the order of traversing the input data, convolution kernel, and output feature map, affecting cache hit rate and data reuse.

[0089] For example, with an on-chip secondary cache and DRAM as the third-level storage, the 7-layer loop reorder sequence of CNN atomic operations is described as follows:

[0090]

[0091] Among them, Q (Output Channels), P (Input Channels), N (Batch Size), C (Height / Width), S (Kernel Height), K (Kernel Width), R (Kernel Depth).

[0092] The server can encode the mapping relationship between the service model and the configuration of the processing device to obtain the feature encoding X corresponding to the configuration mapping information.

[0093] Thus, the server can obtain the feature encoding O corresponding to the topology structure information, the feature encoding X corresponding to the operation type information, and the feature encoding T corresponding to the configuration mapping information. Each combination of [O, T, X] corresponds to a model deployment strategy. Among them, the above feature encodings can be in the form of cell-level one-hot encodings.

[0094] S102: Input the feature encodings of each deployment strategy in the initial deployment strategy group into a pre-trained proxy model to determine the performance distribution information of each deployment strategy on the processing device.

[0095] After determining the initial deployment strategy group, the server can input the initial deployment strategy group into a pre-trained proxy model, so as to determine the performance distribution information among the model deployment strategies through the proxy model. Among them, the performance distribution information is used to characterize the comparison relationship between the performances of the model deployment strategies on the processing device.

[0096] In this specification, the above proxy model can be a binary proxy model, and the performance distribution information output by it can include the probability distribution corresponding to the performance size of each model deployment strategy in the initial deployment strategy group on the processing device. This probability distribution characterizes the comparison relationship between the performances of the service models deployed through the above deployment strategies on the processing device.

[0097] It should be noted that the proxy model does not output the specific performance data of the deployment strategy on the processing device. Its role is to compare the performances of the input deployment strategy pairs on the processing device.

[0098] For example, for the model deployment strategy in the initial deployment strategy group and the model deployment strategy , the performance distribution information output by the performance comparison model can be expressed as: , that is, the performance of B is greater than that of A, and the difference degree between B and A is / .

[0099] The model structure of the above performance comparison model can be as Figure 5 shown.

[0100] Figure 5 It is a schematic diagram of the model structure of a proxy model provided in this specification.

[0101] Among them, the spatial point feature extraction layer can be an end-to-end deep learning model such as LSTM, GRU, MLP, etc., and the proxy model is a differentiable model. The paired spatial point encodings (spatial point 1, spatial point 2) respectively obtain feature layers through the feature extraction layer, and the feature layers are stacked and then input into the prediction branch. The prediction branch outputs a binary vector, which is normalized to obtain a binary probability distribution vector.

[0102] The server can use the binary probability distribution vector corresponding to each metric and the binary relationship label, and use the Kullback-Leibler (KL) divergence to measure the binary distribution distance of the task, and finally realize the supervised training of the proxy model. For the convenience of understanding, this specification provides a schematic diagram of the training process of a performance comparison model, as Figure 6 shown.

[0103] Among them, the training process of the performance comparison model can include the following steps:

[0104] S201: Obtain each historical deployment strategy, and construct multiple historical strategy groups based on each historical deployment strategy; among them, each historical strategy group contains two historical deployment strategies.

[0105] The server can sample a small-scale data set from the collaborative space of the neural network architecture and the mapping scheme, perform verification and evaluation on the spatial points, and make it into data pairs for binary relationship training.

[0106] S202: Determine the actual performance of each historical deployment strategy on the processing device, and determine the actual performance distribution information corresponding to each historical deployment strategy group according to the actual performance.

[0107] The server can further determine the actual performance of the service model deployed by each historical deployment strategy collected on the processing device, and determine the actual performance distribution information corresponding to each historical deployment strategy group according to the actual performance, as the label of the sample data. The construction process of the sample data is asFigure 7 as shown

[0108] Among them, the actual performance distribution information can be determined by comparing information such as hardware-related metrics and model metrics after deploying the business model through each historical model deployment strategy.

[0109] S203: For each historical policy group, input the historical policy group into the proxy model to be trained, so as to determine the predicted performance distribution information corresponding to the historical policy group through the proxy model;

[0110] S204: Determine the loss value of the proxy model according to the predicted performance distribution information and the actual performance distribution information corresponding to the historical policy group, and train the proxy model according to the loss value.

[0111] Among them, the above loss value can be determined by the KL loss function.

[0112] S103: Use the preset label distribution to adjust at least one of the inputs. When the difference between the output performance distribution information and the label distribution meets the set requirements, obtain the adjusted deployment policy group.

[0113] In practical applications, the two model deployment strategies in the initial deployment policy group can include the first deployment strategy and the second deployment strategy. The server can use the preset label distribution to adjust the two inputs input to the proxy model multiple times, and when the difference between the output performance distribution information and the label distribution meets the set requirements, obtain the adjusted deployment policy group. Among them, the set requirements that the difference between the output performance distribution information and the label distribution meets can include: the difference degree between the two reaches the preset difference degree.

[0114] In an embodiment provided in this specification, the above label distribution may only include one label distribution. In this case, the server may aim to maximize the deviation between the performance distribution information output by the proxy model and the label distribution, and make the size relationship between the two inputs represented by the performance distribution information output by the proxy model the same as the size relationship between the two inputs represented by the label distribution, and adjust one of the inputs.

[0115] For example, for the binary group [A, B] of the first deployment strategy A and the second deployment strategy B, and the preset label distribution P[0.4, 0.6], the server may use this probability distribution as a reference, and while fixing the first deployment strategy A, adjust the second deployment strategy B multiple times. During this process, the probability of the performance corresponding to A is always less than that of B, but as the difference degree between the two gradually increases, the output performance distribution information and the label distribution the difference between them also gradually increases.

[0116] In another embodiment provided in this specification, the above-mentioned tag distribution may include a first tag distribution and a second tag distribution, and the magnitude relationship between the performances of the two deployment strategies represented by the first tag distribution on the processing device is opposite to the magnitude relationship between the performances of the two deployment strategies represented by the second tag distribution on the processing device.

[0117] The server may use a preset tag distribution to adjust the first deployment strategy and the second deployment strategy for multiple adjustment cycles, where the tag distributions corresponding to two adjacent adjustment cycles are different.

[0118] For each adjustment cycle, if the tag distribution corresponding to this adjustment cycle is the first tag distribution, the server may use the first tag distribution to adjust the first deployment strategy in this adjustment cycle until the next adjustment cycle is reached. In the next adjustment cycle, the server may use the second tag distribution to adjust the second deployment strategy in the next adjustment cycle.

[0119] After the difference between the performance distribution information and the tag distribution meets the set requirements and the preset adjustment cycle is reached, an adjusted deployment strategy group is obtained.

[0120] For ease of understanding, this specification provides a schematic diagram of the determination process of a target deployment strategy, as Figure 8 shown.

[0121] Among them, the server may use the proxy model to regard the model as a mapping operation from features to outputs. Since the gradient information of the output of the entire model with respect to the input is known, the iterative generation of optimal / dominant input features can be realized, and architecture search and mapping scheme search can be realized. As Figure 8 shown, the server may randomly select two model deployment strategies as initial points, and encode them as A and B respectively as the input of the binary relation model. Supervised training is performed using two different tag distributions. and respectively represent the adjustment cycles for training A and B, and the two can be designed to be: alternating and staged. That is:

[0122] : The first tag distribution P1 [0.4, 0.6], fix A, and iteratively adjust B so that A is less than B;

[0123] : The second tag distribution P2 [0.6, 0.4], that is, fix B and iteratively adjust A so that B is less than A.

[0124] In the later stage of exploration, the server may also perform anti-encoding, training, and hardware deployment verification to obtain the final performance and deployment metrics, and select according to a certain preference strategy.

[0125] During the process of adjusting the model deployment strategy, the server can determine the probability distribution data corresponding to at least one input, and determine the gradient information of the performance distribution information with respect to the initial deployment strategy group; use the preset label distribution to adjust the probability distribution data according to the gradient information to obtain the adjusted probability distribution data. Then, based on the adjusted probability distribution data, determine the adjusted at least one input (that is, convert the adjusted probability distribution data back into discrete feature encodings).

[0126] It should be noted that since the surrogate model is a differentiable surrogate model, it has learned the mapping relationship from input features to output during the training process. Therefore, after the training is completed, the surrogate model can directly provide the gradient information of its output with respect to the input.

[0127] Specifically, for the topological structure of a neural network, the connection relationship between nodes in the network is usually a binary selection problem, that is, either there is a connection between nodes or there is no connection. In traditional architecture search, the selection of connection relationships is discrete, so it is difficult to perform search through gradient deployment. Based on this, the server can parameterize these discrete connection relationships into a continuous probability representation through a reparameterization sampling method. In this way, the selection of topological structure information can determine the possibility of connection by controlling the parameters of the probability representation. Through the deployment search of these parameters, the selection process of connection relationships can be smoothed, so that gradient information can be used for deployment during the search process.

[0128] In addition to the topological connection relationships between network nodes, nodes also need to select the optimal operation from multiple candidate operation types. For example, convolution kernels of different sizes or different types of pooling operations. The selection of these operations is usually also discrete, and traditional architecture search methods are difficult to directly perform continuous deployment on them. To solve this problem, the server can reparameterize the candidate operation type information. By converting the operation type information corresponding to network nodes from discrete options into a probability-based representation, it is possible to perform deployment through gradient information during the search process. For ease of understanding, this specification provides a schematic diagram of the process of adjusting architecture information, as Figure 9 shown.

[0129] Among them, for the probability representation corresponding to the architecture information (i.e., the continuous space variables (α, β)), the server can update it through reparameterization sampling.

[0130] For the mapping relationship between the business model and the configuration of the processing device, the server can determine the gradient information of the performance distribution information with respect to the initial configuration mapping information included in the initial deployment policy group, and determine the configuration constraint information corresponding to the initial configuration mapping information; under the constraint of the configuration constraint information, using the preset label distribution, adjust the initial configuration mapping information according to the gradient information, and update the configuration constraint based on the adjusted mapping configuration information. For ease of understanding, this specification provides a schematic diagram of the process of adjusting the configuration mapping information, as shown in Figure 10 shown.

[0131] Among them, the server can explore the configuration mapping information through the projected gradient method, and this process can include the following steps:

[0132] Initialize parameters: Select an initial parameter value, usually within the feasible region;

[0133] Calculate the gradient: Calculate the gradient of the objective function with respect to the parameters;

[0134] Update and project the parameters based on the gradient;

[0135] After that, repeat the above steps until the stop condition is met, for example: reaching the maximum number of iterations or the gradient is small enough.

[0136] The above projection process is as shown in Figure 11 shown, where the rectangular wireframe represents the feasible region of the spatial parameters:

[0137] Parameter @(t + 1) = Parameter @(t) - Learning rate * Gradient.

[0138] Projecting onto the feasible region can ensure that the parameters still satisfy the constraint conditions after updating the parameters. If the updated parameters do not satisfy the constraint conditions, then it is necessary to change the limit of the projection distance, or change the spatial direction of the projection, and project again until reaching the feasible region.

[0139] During the process of exploring the deployment strategy, the loss function corresponding to each exploration can also be the KL divergence loss, and this loss value can be expressed as:

[0140]

[0141] Among them, this loss value characterizes the distance metric between the binary distribution Q output by the surrogate model and the label distribution P.

[0142] S104: Determine the target deployment strategy in the adjusted deployment policy group, and deploy the business model based on the target deployment strategy.

[0143] After determining the target deployment policy group, the optimal model deployment policy is surely included in the target deployment policy group. The server may use this model deployment policy as the target deployment policy and deploy it to the processing device based on the topology structure information, operation type information of each network node, and configuration mapping information between the processing device included in the target deployment policy, so as to execute services (such as graphic recognition, information recommendation, anomaly detection, etc.) based on this service model.

[0144] As can be seen from the above method, the model deployment method provided by this solution uses the proxy model in combination with the gradient deployment method, making the deployment direction have gradient information and efficient search;

[0145] The binary relation proxy model is used to re - model the original landscape, avoiding "getting stuck in the local optimal area" in the deployment;

[0146] In the exploration stage, there is basically no need to perform actual point verification (based on the real environment and the back - end operating environment), improving the exploration efficiency and further improving the model deployment efficiency.

[0147] The above is one or more implementation model deployment methods of this specification. Based on the same idea, this specification also provides a corresponding model deployment device, as Figure 12 shown.

[0148] Figure 12 It is a schematic diagram of a model deployment device provided by this specification, including:

[0149] An acquisition module 1201, configured to acquire an initial deployment policy group for a service model, where the initial deployment policy group includes two deployment policies, and each deployment policy includes architecture information of the service model and configuration mapping information between the service model and the configuration of the processing device;

[0150] A determination module 1202, configured to input the feature encoding of each deployment policy in the initial deployment policy group into a pre - trained proxy model to determine the performance distribution information of each deployment policy on the processing device;

[0151] An adjustment module 1203, configured to use a preset label distribution to adjust at least one of the inputs, and obtain an adjusted deployment policy group when the difference between the output performance distribution information and the label distribution meets the set requirements;

[0152] A deployment module 1204, configured to determine a target deployment policy in the adjusted deployment policy group and deploy the service model based on the target deployment policy.

[0153] Optionally, for each deployment strategy, the architecture information included in the deployment strategy includes: the topological structure information corresponding to the service model and the operation type information corresponding to each network node in the service model.

[0154] Optionally, the adjustment module 1203 is specifically configured to adjust one of the inputs thereto with the goal of maximizing the deviation between the performance distribution information output by the proxy model and the label distribution, and ensuring that the size relationship represented by the performance distribution information output by the proxy model is the same as the size relationship represented by the label distribution.

[0155] Optionally, the two deployment strategies include: a first deployment strategy and a second deployment strategy; the preset label distributions include: a first label distribution and a second label distribution; wherein, the size relationship between the performances of the two deployment strategies represented by the first label distribution is opposite to the size relationship between the performances of the two deployment strategies represented by the second label distribution on the processing device;

[0156] The adjustment module 1203 is specifically configured to use the preset label distributions to adjust the first deployment strategy and the second deployment strategy for multiple adjustment cycles, where the label distributions corresponding to adjacent adjustment cycles are different; for each adjustment cycle, if the label distribution corresponding to the adjustment cycle is the first label distribution, then use the first label distribution to adjust the first deployment strategy in the adjustment cycle until reaching the next adjustment cycle; in the next adjustment cycle, use the second label distribution to adjust the second deployment strategy in the next adjustment cycle.

[0157] Optionally, the adjustment module 1203 is specifically configured to determine the probability distribution data corresponding to the at least one input, and determine the gradient information of the performance distribution information with respect to the initial deployment strategy group; use the preset label distribution to adjust the probability distribution data according to the gradient information to obtain adjusted probability distribution data; based on the adjusted probability distribution data, determine the adjusted at least one input.

[0158] Optionally, the adjustment module 1203 is specifically configured to determine the gradient information of the performance distribution information with respect to the initial configuration mapping information included in the initial deployment strategy group, and determine the configuration constraint information corresponding to the initial configuration mapping information; under the constraint of the configuration constraint information, use the preset label distribution to adjust the initial configuration mapping information according to the gradient information, and update the configuration constraint based on the adjusted mapping configuration information.

[0159] Optionally, the device further includes: a training module 1205, configured to obtain each historical deployment policy, and construct a plurality of historical policy groups based on each historical deployment policy; wherein, each historical policy group includes two historical deployment policies; determine the actual performance of each historical deployment policy on the processing device, and determine the actual performance distribution information corresponding to each historical deployment policy group according to the actual performance; for each historical policy group, input the historical policy group into the proxy model to be trained, so as to determine the predicted performance distribution information corresponding to the historical policy group through the proxy model; determine the loss value of the proxy model according to the predicted performance distribution information and the actual performance distribution information corresponding to the historical policy group, and train the proxy model according to the loss value.

[0160] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 provided model deployment method.

[0161] This specification also provides Figure 13 a schematic structural diagram of an electronic device corresponding to Figure 1 as shown. As Figure 13 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 3 described model deployment method. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device.

[0162] For an improvement in a technology, it can be clearly distinguished whether it is a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method flows). However, with the development of technology, many improvements in method flows today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user's programming of the device. The designer can program by himself to "integrate" a digital system on a piece of PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0163] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0164] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0165] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0166] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0167] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0170] In a typical configuration, a processing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0171] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0172] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a processing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0173] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0174] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0175] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0176] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0177] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A model deployment method, characterized in that: The method comprises: Acquire an initial deployment strategy group for a business model, the initial deployment strategy group including two deployment strategies, each deployment strategy including architecture information of the business model and configuration mapping information between the business model and a configuration of a processing device; Inputting the feature code of each deployment strategy in the initial deployment strategy group into a pre-trained proxy model to determine the performance distribution information of each deployment strategy on the processing device; wherein the performance distribution information is used to characterize the comparison relationship between the performances of each deployment strategy on the processing device; With the goal of maximizing the deviation between the performance distribution information of each deployment strategy and the preset label distribution, and making the size relationship represented by the performance distribution information of each deployment strategy the same as the size relationship represented by the label distribution, at least one deployment strategy is adjusted multiple times until the difference between the determined performance distribution information and the label distribution reaches a preset difference, thereby obtaining an adjusted deployment strategy group; wherein the label distribution is used to represent the minimum expected value of the performance distribution information of each deployment strategy, and when the performance distribution information of each deployment strategy reaches the minimum expected value, the deployment strategy with the largest corresponding performance distribution value in the performance distribution information is adjusted; A target deployment strategy is determined in the adjusted deployment strategy group, and the business model is deployed based on the target deployment strategy.

2. The method according to claim 1, characterized in that For each deployment strategy, the architecture information included in the deployment strategy includes: topology information corresponding to the business model and operation type information corresponding to each network node in the business model.

3. The method according to claim 1, characterized in that The two deployment strategies include: a first deployment strategy and a second deployment strategy; the preset label distribution includes: a first label distribution and a second label distribution; wherein the magnitude relationship between the performances of the two deployment strategies on the processing device represented by the first label distribution is opposite to the magnitude relationship between the performances of the two deployment strategies on the processing device represented by the second label distribution; Make multiple adjustments to at least one deployment strategy, including: Using a preset label distribution, the first deployment strategy and the second deployment strategy are adjusted for a plurality of adjustment cycles, wherein the label distributions corresponding to two adjacent adjustment cycles are different; For each adjustment period, if the label distribution corresponding to the adjustment period is the first label distribution, adjusting the first deployment strategy in the adjustment period by using the first label distribution until the next adjustment period is reached; In the next adjustment cycle, the second deployment strategy in the next adjustment cycle is adjusted using the second label distribution.

4. The method according to claim 1, characterized in that You can make multiple adjustments to at least one deployment strategy, including: Determine probability distribution data corresponding to the at least one deployment strategy, and determine gradient information of the performance distribution information relative to the initial deployment strategy group; Using the preset label distribution, adjusting the probability distribution data according to the gradient information to obtain adjusted probability distribution data; Based on the adjusted probability distribution data, at least one adjusted deployment strategy is determined.

5. The method according to claim 1, characterized in that Make multiple adjustments to at least one deployment strategy, including: Determining gradient information of the performance distribution information relative to initial configuration mapping information included in the initial deployment strategy group, and determining configuration constraint information corresponding to the initial configuration mapping information; Under the constraints of the configuration constraint information, the initial configuration mapping information is adjusted according to the gradient information using the preset label distribution, and the configuration constraint is updated based on the adjusted mapping configuration information.

6. The method according to claim 1, characterized in that Training the agent model specifically includes: Obtain each historical deployment strategy, and construct multiple historical strategy groups based on each historical deployment strategy; wherein each historical strategy group includes two historical deployment strategies; Determine the actual performance of each historical deployment strategy on the processing device, and determine actual performance distribution information corresponding to each historical deployment strategy group according to the actual performance; For each historical strategy group, input the historical strategy group into the proxy model to be trained, so as to determine the prediction performance distribution information corresponding to the historical strategy group through the proxy model; The loss value of the proxy model is determined according to the predicted performance distribution information and the actual performance distribution information corresponding to the historical strategy group, and the proxy model is trained according to the loss value.

7. A model deployment device, characterized in that: include: An acquisition module is configured to acquire an initial deployment strategy group for a business model, wherein the initial deployment strategy group includes two deployment strategies, each of which includes architecture information of the business model and configuration mapping information between the business model and the configuration of a processing device; A determination module inputs the feature code of each deployment strategy in the initial deployment strategy group into a pre-trained proxy model to determine the performance distribution information of each deployment strategy on the processing device; wherein the performance distribution information is used to characterize the comparison relationship between the performances of each deployment strategy on the processing device; An adjustment module, with the goal of maximizing the deviation between the performance distribution information of each deployment strategy and the preset label distribution, and making the size relationship represented by the performance distribution information of each deployment strategy the same as the size relationship represented by the label distribution, adjusts at least one deployment strategy multiple times until the difference between the determined performance distribution information and the label distribution reaches a preset difference, thereby obtaining an adjusted deployment strategy group; wherein the label distribution is used to represent the minimum expected value of the performance distribution information of each deployment strategy, and when the performance distribution information of each deployment strategy reaches the minimum expected value, the deployment strategy with the largest corresponding performance distribution value in the performance distribution information is adjusted; The deployment module determines a target deployment strategy in the adjusted deployment strategy group and deploys the business model based on the target deployment strategy.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor implements the method described in any one of claims 1 to 6 when executing the program.

Citation Information

Patent Citations

  • Method for block-level NN deployment metric modelling

    US20230316044A1

  • Model construction method and apparatus, and storage medium and electronic device

    WO2024234477A1