Model deployment method and device, electronic equipment and computer readable storage medium

By selecting the appropriate lightweight model on the end-side device and deploying it in order to prioritize time, the problem of low deployment efficiency of semantic communication models in the end-side cloud network is solved, and resource conservation and security improvement is achieved.

CN120276740APending Publication Date: 2025-07-08BEIJING UNIV OF POSTS & TELECOMM +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311841999.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In end-edge cloud networks, it is difficult for the existing technology to efficiently deploy semantic communication models, resulting in increased network resource occupation and risk of user information leakage, and the deployment efficiency is low.

Method used

By selecting the appropriate target lightweight model from the lightweight model generated by multiple lightweight modes, the deployment time is determined, and priority is given based on time, and the model deployment order of the end-side equipment is reasonably arranged.

Benefits of technology

It reduces network resource usage, improves user information security, and improves model deployment efficiency, ensuring timely service of end-side equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276740A_ABST
    Figure CN120276740A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model deployment method and device, electronic equipment and a computer readable storage medium. The method comprises: in response to a received model deployment request of an end-side device, determining a target lightweight model required by the end-side device from at least one lightweight model, each lightweight model being generated by performing lightweight processing on a full-amount model based on different lightweight modes; determining model deployment time required for deploying the target lightweight model on the end side equipment; determining a deployment priority order of the end-side equipment based on the model deployment time consumption; and performing model deployment on the end-side equipment based on the deployment priority ranking. According to the scheme, model deployment on the end side equipment is realized by using the lightweight model, and deployment priority ranking is performed on the end side equipment, so that the deployment sequence of the end side equipment is reasonable, and the model deployment efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of cloud computing, edge extreme, end-edge-cloud network, and semantic communication. Specifically, the present application relates to a model deployment method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Semantic communication is an organic combination of traditional communication and the field of artificial intelligence (AI). It uses deep learning technology in the field of AI to process the information source to achieve the purpose of compressing the information source, and its performance is superior to traditional communication under low signal-to-noise ratio conditions. Semantic communication saves communication resources and expands the communication range, meeting the requirements of 6G, and is an important development direction in the communication field.

[0003] Processing the information source in semantic communication requires support in terms of computing power, memory, and energy consumption. Most terminal devices are limited in battery capacity, computing power, and memory, making it impossible for them to apply semantic communication models. However, in the end-edge-cloud network, the terminal side can offload computing tasks to the edge side and the cloud side, or jointly complete computing tasks with the edge side and the cloud side. Therefore, semantic communication is generally applied to the end-edge-cloud system, and the semantic communication model is deployed on the edge side and the cloud side.

[0004] Although the method of distributing computing tasks to the edge side and the cloud side can meet the needs of end-users, it brings some problems. On the one hand, this method may occupy more network resources; on the other hand, it may lead to the leakage of user information on the terminal side, causing security problems.

[0005] If the semantic communication model can be deployed on the terminal side and the terminal side completes the computing task, it can reduce the information transmission between the terminal side and the edge side or the cloud side, reduce the occupation of network resources, and avoid the upload of user information on the terminal side, ensuring the security of user information on the terminal side.

[0006] In addition, when deploying the semantic communication model on the terminal side, how to improve the deployment efficiency of the model is an important technical problem. Summary of the Invention

[0007] The present application provides a model deployment method, apparatus, electronic device, and computer-readable storage medium, which can solve at least one of the above problems. The technical solutions adopted in the present application are as follows:

[0008] In a first aspect, an embodiment of the present application provides a model deployment method, which includes:

[0009] In response to receiving a model deployment request from a terminal-side device, determine a target lightweight model required by the terminal-side device from at least one lightweight model, where each lightweight model is generated by performing lightweight processing on a full-scale model based on different lightweight methods;

[0010] Determine the model deployment time required for deploying the target lightweight model on the edge device;

[0011] Determine the deployment priority ranking of the edge device based on the model deployment time;

[0012] Deploy the model on the edge device based on the deployment priority ranking.

[0013] In a second aspect, an embodiment of the present application provides a model deployment device, which includes:

[0014] A target lightweight model determination module, configured to determine the target lightweight model required by the edge device from at least one lightweight model in response to receiving a model deployment request from the edge device, where each lightweight model is generated by performing lightweight processing on the full-scale model based on different lightweight methods;

[0015] A model deployment time determination module, configured to determine the model deployment time required for deploying the target lightweight model on the edge device;

[0016] A priority ranking module, configured to determine the deployment priority ranking of the edge device based on the model deployment time;

[0017] A model deployment module, configured to deploy the model on the edge device based on the deployment priority ranking.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory;

[0019] The memory is used to store operation instructions;

[0020] The processor is configured to execute the model deployment method as shown in any implementation manner of the first aspect of the present application by calling the operation instructions.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the model deployment method as shown in any implementation manner of the first aspect of the present application.

[0022] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:

[0023] The solution provided by the embodiments of this application determines the target lightweight model required by the end-side device from at least one lightweight model by responding to the received model deployment request of the end-side device. Each lightweight model is generated by lightweight processing of the full-scale model based on different lightweight methods; determines the model deployment time required to deploy the target lightweight model on the end-side device; determines the deployment priority ranking of the end-side device based on the model deployment time; and deploys the model to the end-side device based on the deployment priority ranking. In this solution, by using lightweight models, model deployment on the end-side device is achieved, and by sorting the deployment priorities of the end-side devices, the deployment order of the end-side devices is made reasonable, which helps to improve the model deployment efficiency. Description of the Drawings

[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for description in the embodiments of this application.

[0025] Figure 1 Flow chart of a model deployment method provided by an embodiment of this application;

[0026] Figure 2 Structural diagram of the end-edge-cloud network in an embodiment of this application;

[0027] Figure 3 Flow chart of an alternative implementation manner of the method provided by an embodiment of this application;

[0028] Figure 4 Structural diagram of a model deployment device provided by an embodiment of this application;

[0029] Figure 5 Structural diagram of an electronic device provided by an embodiment of this application. Detailed Embodiments

[0030] The following details the embodiments of this application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain this application and should not be construed as a limitation of the present invention.

[0031] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of this application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0032] To make the objectives, technical solutions and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings.

[0033] The technical solutions of this application and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0034] Figure 1 The flowchart of a model deployment method provided by an embodiment of this application is shown, as Figure 1 shown, the method mainly may include:

[0035] Step S110: In response to receiving a model deployment request from an end-side device, determine a target lightweight model required by the end-side device from at least one lightweight model, where each lightweight model is generated by lightweight processing of a full-scale model based on different lightweight methods;

[0036] Step S120: Determine the model deployment time required to deploy the target lightweight model on the end-side device;

[0037] Step S130: Determine the deployment priority ranking of the end-side device based on the model deployment time;

[0038] Step S140: Perform model deployment on the end-side device based on the deployment priority ranking.

[0039] Among them, the full-scale model may be a semantic communication model. The lightweight model is obtained by performing lightweight processing on the full-scale model using a preset lightweight method. The lightweight method may include but is not limited to distillation, pruning, etc. Multiple lightweight models in this case are respectively obtained by performing lightweight processing on the full-scale model using various different lightweight methods.

[0040] When receiving a model deployment request from an end - side device, the requirements of the end - side device for the lightweight model can be obtained from the model deployment request, so that the target lightweight model required by the end - side device can be determined from multiple lightweight models.

[0041] After determining the target lightweight model, the time required to deploy the target lightweight model to the end - side device can be calculated, that is, the model deployment time. Then, based on the model deployment time, the deployment priority ranking of the end - side device can be determined, so as to deploy the model to the end - side device based on the deployment priority ranking.

[0042] In the embodiments of the present application, a higher deployment priority ranking can be assigned to the end - side device with a shorter model deployment time, that is, the model deployment for the end - side device with a shorter model deployment time is more prioritized. Thus, the end - side device with a shorter model deployment time can be deployed as soon as possible and provide services to users earlier, thereby avoiding the waiting of a large number of end - side devices caused by the prior model deployment for the end - side device with a longer model deployment time, avoiding blocking the model deployment process, and improving the model deployment efficiency.

[0043] In the embodiments of the present application, by deploying lightweight models, the requirements for the computing power, memory, and energy consumption of the end - side device are reduced, so as to realize the deployment of the model on the end - side device and process tasks on the end - side device. Moreover, it can reduce the load on the cloud - side and edge - side and redundant information in the network, reduce the dependence of the end - side on the cloud - side and edge - side, and enable direct semantic communication between the end - to - end, improving the semantic communication ability of the entire network.

[0044] The method provided in the embodiments of the present application, by responding to the received model deployment request from the end - side device, determines the target lightweight model required by the end - side device from at least one lightweight model, and each lightweight model is generated by lightweight processing of the full - scale model based on different lightweight methods; determines the model deployment time required to deploy the target lightweight model on the end - side device; determines the deployment priority ranking of the end - side device based on the model deployment time; and deploys the model to the end - side device based on the deployment priority ranking. In this solution, by using lightweight models, the model deployment on the end - side device is realized, and by sorting the deployment priorities of the end - side devices, the deployment order of the end - side devices is made reasonable, which helps to improve the model deployment efficiency.

[0045] In an optional manner of the embodiments of the present application, determining the model deployment time required to deploy the target lightweight model on the end - side device includes:

[0046] Determine the model lightweighting time and the corresponding model distribution time corresponding to the target lightweight model. The model lightweighting time is the time required to obtain the target lightweight model by lightweight processing the full - scale model, and the model distribution time is the time required to distribute the target lightweight model to the end - side device;

[0047] Based on the time-consuming of model lightweighting and the time-consuming of model distribution, determine the model deployment time-consuming required for deploying the target lightweight model on the edge device.

[0048] In the embodiments of the present application, the model deployment time-consuming can be composed of the time-consuming of model lightweighting and the time-consuming of model distribution. The time-consuming of model lightweighting is the time required to obtain the target lightweight model through lightweighting processing of the full-scale model, and the time-consuming of model distribution is the time required to distribute the target lightweight model to the edge device.

[0049] As an example, the total time obtained by adding the time-consuming of model lightweighting and the time-consuming of model distribution can be used as the model deployment time-consuming.

[0050] In an alternative embodiment of the present application, determining the time-consuming of model lightweighting corresponding to the target lightweight model includes:

[0051] Determine whether the target lightweight model exists in the preset model library;

[0052] Based on whether the target lightweight model exists in the model library, determine the time-consuming of model lightweighting corresponding to the target lightweight model.

[0053] In the embodiments of the present application, the model library can be set on the cloud side or the edge side. The model library can contain resources of multiple different lightweight models and can be directly used for distribution to the edge device.

[0054] If the target lightweight model exists in the model library, it can be directly used without regeneration. Therefore, based on whether the target lightweight model exists in the model library, determine the time-consuming of model lightweighting corresponding to the target lightweight model.

[0055] In an alternative embodiment of the present application, based on whether the target lightweight model exists in the model library, determining the time-consuming of model lightweighting corresponding to the target lightweight model includes:

[0056] In response to the existence of the target lightweight model in the model library, determine the time-consuming of model lightweighting corresponding to the target lightweight model as zero;

[0057] In response to the non-existence of the target lightweight model in the model library, determine the time-consuming of model lightweighting corresponding to the target lightweight model based on the performance parameters of the target lightweight model.

[0058] In the embodiments of the present application, if the target lightweight model exists in the model library, it can be directly used without regeneration. At this time, the time-consuming of model lightweighting can be determined as zero.

[0059] If the target lightweight model does not exist in the model library, the target lightweight model needs to be generated.

[0060] Among them, the performance parameters of the lightweight model may include at least one of the following:

[0061] The memory amount required by the lightweight model;

[0062] The computing power index of the lightweight model;

[0063] The performance index of the lightweight model.

[0064] The performance parameters of the target lightweight model can characterize the degree of lightweighting of the full-scale model during the process of compressing the full-scale model to obtain the target lightweight model, that is, the compression degree. There is a correlation between the degree of lightweighting and the time-consuming of model lightweighting.

[0065] In an optional manner of the embodiments of the present application, determining the model lightweighting time-consuming corresponding to the target lightweight model based on the performance parameters of the target lightweight model includes:

[0066] Based on the preset corresponding relationship between at least one performance parameter and the model lightweighting time-consuming, and based on the performance parameters of the target lightweight model, determining the model lightweighting time-consuming corresponding to the target lightweight model.

[0067] In the embodiments of the present application, a corresponding relationship between at least one performance parameter and the model lightweighting time-consuming can be pre-established. Specifically, a corresponding relationship can be established between the memory amount required by the target lightweight model, the computing power index of the target lightweight model, and the performance index of the target lightweight model and the model lightweighting time-consuming respectively.

[0068] As an example, this corresponding relationship can be expressed in the form of a function F(m). In the coordinate graph of the function F(m), the horizontal axis can represent the model lightweighting time-consuming, and multiple vertical axes can respectively represent the memory amount required by the target lightweight model, the computing power index of the target lightweight model, and the performance index of the target lightweight model.

[0069] In an optional manner of the embodiments of the present application, determining the model distribution time-consuming corresponding to the target lightweight model includes:

[0070] Inputting the network location of the end-side device into the pre-trained distribution time-consuming learning model to obtain the model distribution time-consuming corresponding to the target lightweight model output by the distribution time-consuming learning model.

[0071] In the embodiments of the present application, the distribution time-consuming learning model can be deployed on the cloud side or the edge side, and based on the distribution time-consuming learning model, the model distribution time-consuming corresponding to the target lightweight model is determined. Taking the network location of the end-side device as the input and inputting it into the distribution time-consuming learning model to obtain the model distribution time-consuming corresponding to the target lightweight model output by the distribution time-consuming learning model.

[0072] In an optional manner of the embodiments of the present application, the distribution time-consuming learning model is a reinforcement learning model, and the distribution time-consuming learning model is trained through the following method:

[0073] Obtain the network environment information of the edge-cloud network to which the edge device belongs, and construct a reinforcement learning environment based on the network environment information;

[0074] In the reinforcement learning environment, use the model distribution time as the reward to perform reinforcement learning on the distribution time learning model.

[0075] In the embodiments of the present application, the distribution time learning model can be a reinforcement learning model.

[0076] During the training of the distribution time learning model, the network environment information of the edge-cloud network to which the edge device belongs can be used to construct a reinforcement learning environment, and states, actions, rewards, etc. can be defined. Use the model distribution time as the reward to train the reinforcement learning model, and deploy the pre-trained reinforcement learning model on the edge side and the cloud side.

[0077] Map information such as the network topology, nodes, links, generation model, and transmission time required for the model of the edge-cloud network to

[0078] The type of the reinforcement learning model can be selected according to network requirements and specific situations, including but not limited to Q-Learning, Deep Q Network (DQN), Soft Actor-Critic (SAC), etc. The model and algorithm can also be designed independently according to requirements.

[0079] When the network topology, nodes, links, etc. of the edge-cloud network change, the reinforcement learning model can be fine-tuned or retrained to adapt to the network changes, ensure the accuracy of model distribution time learning, and ensure the deployment efficiency of lightweight models.

[0080] In an alternative embodiment of the present application, based on the model lightweighting time and the model distribution time, determine the model deployment time required to deploy the target lightweight model on the edge device, including:

[0081] In response to detecting a network congestion situation in the edge-cloud network to which the edge device belongs, determine the network delay duration of the edge device;

[0082] Based on the network delay duration, the model lightweighting time, and the model distribution time, determine the model deployment time required to deploy the target lightweight model on the edge device.

[0083] In the embodiments of the present application, there may be a network congestion situation in the edge-cloud network. At this time, the network congestion will cause an increase in the model deployment time. The network delay duration can be determined, and based on the network delay duration, the model lightweighting time, and the model distribution time, jointly determine the model deployment time.

[0084] As an example, the network delay duration can be determined as the model deployment duration based on the total duration obtained by adding the network delay duration, the model lightweighting duration, and the model distribution duration.

[0085] In an alternative implementation of the present application, determining a target lightweight model required by an edge device from at least one lightweight model includes:

[0086] Obtaining the performance parameter requirement conditions of the edge device from the model deployment request;

[0087] Determining a lightweight model whose performance parameters can meet the performance parameter requirement conditions of the edge device as the target lightweight model required by the edge device.

[0088] In the embodiments of the present application, the performance parameter requirement conditions of the edge device may include at least one of the memory amount required by the lightweight model required by the edge device, the computing power index of the required lightweight model, and the performance index of the required lightweight model. For a lightweight model that meets the performance parameter requirement conditions of the edge device, its memory and computing power are both smaller than those of the edge side, and the model performance index should meet the minimum requirements required by the edge device.

[0089] In some cases, there may be other conditional restrictions. For example, when the power of the edge device is limited, it is also necessary to ensure that the power required for model operation is less than the power of the terminal device.

[0090] In an alternative implementation of the present application, determining the deployment priority ranking of edge devices based on the model deployment duration includes:

[0091] Sorting the model deployment durations corresponding to multiple edge devices in ascending order as the deployment priority ranking of each edge device.

[0092] In the embodiments of the present application, the model deployment durations corresponding to edge devices can be sorted in ascending order as the deployment priority ranking of the edge devices. That is, for an edge device with a lower model deployment duration, its deployment priority ranking is relatively higher, while for an edge device with a higher model deployment duration, its deployment priority ranking is relatively lower. Edge devices with a higher deployment priority ranking will be deployed more preferentially.

[0093] Optionally, the deployment priority ranking can also be adjusted according to the actual requirements in the network. When an edge device needs to complete model deployment within a certain time limit, the model lightweighting duration and the model distribution duration can be shortened to meet the requirements of the terminal device. The model distribution duration can be shortened by increasing the transmission priority of the edge device or selecting a better distribution strategy, etc.

[0094] As an example, Figure 2 shows a schematic structural diagram of the edge-cloud network in the embodiments of the present application.

[0095] As shown in Figure 2

[0096]

[0097] Figure 3 As an example, Figure 3

[0098] Figure 3 As shown in

[0099]

[0100] The network configuration stage includes:

[0101] Determine that the model lightweight method used is the semantic model, and the semantic model required to be deployed by the pruning method is M.

[0102] Use the model pruning algorithm to prune M and fit the corresponding relationship F(m) between the performance parameters such as memory and computing power of the pruned model and the pruning time.

[0102] Deploy the pruning algorithm, model M, and function F(m) on the cloud side and the edge side;

[0103] Select the reinforcement learning algorithm as DQN, train the reinforcement learning model, and deploy the pre-trained model on the edge side and the cloud side.

[0104] The model deployment includes:

[0105] The edge side determines the time t1 required to generate the model according to the requests from the end sides within the served range, and the time t2 required to determine the optimal distribution strategy by the reinforcement learning model;

[0106] According to the actual situation and specific requirements in the network, obtain the total time t required to serve one end-side device;

[0107] Sort the priorities of all end-side devices according to t, the smaller t is, the higher the priority. During the service process, establish a model knowledge base on the cloud side and the edge side;

[0108] When a new demand arrives, first check whether there is an available model that meets the user's needs in the model knowledge base; if so, directly distribute the existing applicable model, if not, a new version of the model needs to be generated.

[0109] The method provided in this case is applicable to the semantic edge-cloud network structure. When distributing models, it can meet the requirements of the most edge devices in the shortest time according to the edge device requirements, complete the deployment of semantic models in the entire network in the shortest possible time, thereby realizing end-to-end semantic communication in the entire network, expanding the applicable scope of semantic communication, and improving the communication efficiency of the entire network.

[0110] Based on the same principle as the method shown in Figure 1 a structure diagram of a model deployment device provided by an embodiment of the present application is shown. As shown in Figure 4 Figure 4 , the model deployment device 40 may include: Figure 4 As shown in

[0111] A target lightweight model determination module 410, configured to determine a target lightweight model required by an edge device from at least one lightweight model in response to receiving a model deployment request from the edge device. Each lightweight model is generated by performing lightweight processing on a full-scale model based on different lightweight methods;

[0112] A model deployment time consumption determination module 420, configured to determine the model deployment time consumption required for deploying the target lightweight model on the edge device;

[0113] A priority sorting module 430, configured to determine the deployment priority sorting of edge devices based on the model deployment time consumption;

[0114] A model deployment module 440, configured to perform model deployment on edge devices based on the deployment priority sorting.

[0115] The device provided by the embodiment of the present application, by responding to receiving a model deployment request from an edge device, determines a target lightweight model required by the edge device from at least one lightweight model. Each lightweight model is generated by performing lightweight processing on a full-scale model based on different lightweight methods; determines the model deployment time consumption required for deploying the target lightweight model on the edge device; determines the deployment priority sorting of edge devices based on the model deployment time consumption; and performs model deployment on edge devices based on the deployment priority sorting. In this solution, by using lightweight models, model deployment on edge devices is realized, and by sorting the deployment priorities of edge devices, the deployment order of edge devices is made reasonable, which helps to improve the model deployment efficiency.

[0116] Optionally, the model deployment time consumption determination module is specifically configured to:

[0117] Determine the model lightweight time consumption and the corresponding model distribution time consumption corresponding to the target lightweight model. The model lightweight time consumption is the required duration for obtaining the target lightweight model by performing lightweight processing on the full-scale model, and the model distribution time consumption is the required duration for distributing the target lightweight model to the edge device;

[0118] Based on the time-consuming for model lightweighting and the time-consuming for model distribution, determine the model deployment time-consuming required for deploying the target lightweight model on the edge device.

[0119] Optionally, when determining the time-consuming for model lightweighting corresponding to the target lightweight model, the model deployment time-consuming determination module is specifically configured to:

[0120] Determine whether the target lightweight model exists in the preset model library;

[0121] Based on whether the target lightweight model exists in the model library, determine the time-consuming for model lightweighting corresponding to the target lightweight model.

[0122] Optionally, when determining the time-consuming for model lightweighting corresponding to the target lightweight model based on whether the target lightweight model exists in the model library, the model deployment time-consuming determination module is specifically configured to:

[0123] In response to the target lightweight model existing in the model library, determine the time-consuming for model lightweighting corresponding to the target lightweight model as zero;

[0124] In response to the target lightweight model not existing in the model library, determine the time-consuming for model lightweighting corresponding to the target lightweight model based on the performance parameters of the target lightweight model.

[0125] Optionally, when determining the time-consuming for model lightweighting corresponding to the target lightweight model based on the performance parameters of the target lightweight model, the model deployment time-consuming determination module is specifically configured to:

[0126] Based on the corresponding relationship between at least one preset performance parameter and the time-consuming for model lightweighting, and based on the performance parameters of the target lightweight model, determine the time-consuming for model lightweighting corresponding to the target lightweight model.

[0127] Optionally, the performance parameters include at least one of the following:

[0128] The memory amount required for the lightweight model;

[0129] The computing power index of the lightweight model;

[0130] The performance index of the lightweight model.

[0131] Optionally, when determining the time-consuming for model distribution corresponding to the target lightweight model, the model deployment time-consuming determination module is specifically configured to:

[0132] Input the network location of the edge device into the pre-trained distribution time-consuming learning model to obtain the time-consuming for model distribution corresponding to the target lightweight model output by the distribution time-consuming learning model.

[0133] Optionally, the distribution time-consuming learning model is a reinforcement learning model, and the distribution time-consuming learning model is trained through the following method:

[0134] Obtain the network environment information of the edge-cloud network to which the edge device belongs, and construct a reinforcement learning environment based on the network environment information;

[0135] In the reinforcement learning environment, perform reinforcement learning on the distribution time-consuming learning model with the model distribution time as the reward.

[0136] Optionally, the model deployment time determination module is specifically used for:

[0137] In response to detecting a network congestion situation in the edge-cloud network to which the edge device belongs, determine the network delay duration of the edge device;

[0138] Based on the network delay duration, the model lightweighting time, and the model distribution time, determine the model deployment time required to deploy the target lightweight model on the edge device.

[0139] Optionally, when the target lightweight model determination module determines the target lightweight model required by the edge device from at least one lightweight model, it is specifically used for:

[0140] Obtain the performance parameter requirement conditions of the edge device from the model deployment request;

[0141] Determine the lightweight model whose performance parameters can meet the performance parameter requirement conditions of the edge device as the target lightweight model required by the edge device.

[0142] Optionally, the priority sorting module is specifically used for:

[0143] Sort the model deployment times corresponding to multiple edge devices in ascending order as the deployment priority sorting of each edge device.

[0144] It can be understood that the above-mentioned modules of the model deployment device in this embodiment have the functions of implementing the corresponding steps of the model deployment method in the embodiment shown in Figure 1 The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated by multiple modules. For the function descriptions of the above modules of the model deployment device, reference can specifically be made to the corresponding descriptions of the model deployment method in the embodiment shown in Figure 1 The corresponding descriptions in the model deployment method in the embodiment shown in, and details are not described herein again.

[0145] An embodiment of the present application provides an electronic device, including a processor and a memory;

[0146] The memory is used to store operation instructions;

[0147] The processor is used to execute the model deployment method provided in any implementation manner of the present application by calling the operation instructions.

[0148] As an example, Figure 5 a schematic structural diagram of an electronic device applicable to an embodiment of the present application is shown, such as Figure 5 shown. The electronic device 2000 includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are connected, such as connected through a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in practical applications, the transceiver 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation to the embodiments of the present application.

[0149] Among them, the processor 2001 is applied to the embodiments of the present application and is used to implement the methods shown in the above method embodiments. The transceiver 2004 may include a receiver and a transmitter. The transceiver 2004 is applied to the embodiments of the present application and is used to implement the function of communicating between the electronic device of the embodiments of the present application and other devices when executed.

[0150] The processor 2001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 2001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0151] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0152] The memory 2003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0153] Optionally, the memory 2003 is used to store the application program code for executing the solution of this application, and is controlled and executed by the processor 2001. The processor 2001 is used to execute the application program code stored in the memory 2003 to implement the model deployment method provided in any embodiment of this application.

[0154] The electronic device provided in the embodiments of this application is applicable to any embodiment of the above method, and will not be elaborated herein.

[0155] The embodiments of this application provide an electronic device. Compared with the prior art, in response to receiving a model deployment request from an end-side device, a target lightweight model required by the end-side device is determined from at least one lightweight model, and each lightweight model is generated by lightweight processing of a full-scale model based on different lightweight methods; the model deployment time required for deploying the target lightweight model on the end-side device is determined; the deployment priority order of the end-side device is determined based on the model deployment time; and the model is deployed to the end-side device based on the deployment priority order. In this solution, by using lightweight models, model deployment on the end-side device is achieved, and by sorting the deployment priorities of the end-side devices, the deployment order of the end-side devices is made reasonable, which helps to improve the model deployment efficiency.

[0156] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the model deployment method shown in the above method embodiments.

[0157] The computer-readable storage medium provided in the embodiments of this application is applicable to any embodiment of the above method, and will not be elaborated herein.

[0158] An embodiment of the present application provides a computer-readable storage medium. Compared with the prior art, in response to receiving a model deployment request from an edge device, a target lightweight model required by the edge device is determined from at least one lightweight model, and each lightweight model is generated by lightweight processing of a full-scale model based on different lightweight methods; the model deployment time required for deploying the target lightweight model on the edge device is determined; the deployment priority order of the edge device is determined based on the model deployment time; and the model is deployed to the edge device based on the deployment priority order. In this solution, by using lightweight models, model deployment on edge devices is achieved, and by sorting the deployment priorities of edge devices, the deployment order of edge devices is made reasonable, which helps to improve the model deployment efficiency.

[0159] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0160] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A model deployment method, characterized in that, Including: In response to receiving a model deployment request from an edge device, determining a target lightweight model required by the edge device from at least one lightweight model, where each lightweight model is generated by lightweight processing of a full-scale model based on different lightweight methods; Determining the model deployment time required to deploy the target lightweight model on the edge device; Determining the deployment priority ranking of the edge device based on the model deployment time; Deploying the model to the edge device based on the deployment priority ranking.

2. The method according to claim 1, wherein The determining of the model deployment time required to deploy the target lightweight model on the edge device includes: Determining the model lightweighting time and the corresponding model distribution time corresponding to the target lightweight model, where the model lightweighting time is the time required to obtain the target lightweight model by lightweight processing the full-scale model, and the model distribution time is the time required to distribute the target lightweight model to the edge device; Based on the model lightweighting time and the model distribution time, determining the model deployment time required to deploy the target lightweight model on the edge device.

3. The method according to claim 2, wherein The determining of the model lightweighting time corresponding to the target lightweight model includes: Determining whether the target lightweight model exists in a preset model library; Based on whether the target lightweight model exists in the model library, determining the model lightweighting time corresponding to the target lightweight model.

4. The method according to claim 3, characterized in that, The determining of the model lightweighting time corresponding to the target lightweight model based on whether the target lightweight model exists in the model library includes: In response to the target lightweight model existing in the model library, determining the model lightweighting time corresponding to the target lightweight model as zero; In response to the target lightweight model not existing in the model library, determining the model lightweighting time corresponding to the target lightweight model based on the performance parameters of the target lightweight model.

5. The method according to claim 4, wherein The determining of the model lightweighting time corresponding to the target lightweight model based on the performance parameters of the target lightweight model includes: Based on a preset correspondence between at least one performance parameter and the model lightweighting time, and based on the performance parameters of the target lightweight model, determining the model lightweighting time corresponding to the target lightweight model.

6. The method according to claim 5, wherein The performance parameters include at least one of the following: The memory amount required by the lightweight model; The computing power index of the lightweight model; The performance index of the lightweight model.

7. The method according to claim 3, wherein Determining the model distribution time corresponding to the target lightweight model includes: Inputting the network location of the edge device into a pre-trained distribution time learning model to obtain the model distribution time corresponding to the target lightweight model output by the distribution time learning model.

8. The method according to claim 7, wherein The distribution time learning model is a reinforcement learning model, and the distribution time learning model is trained in the following manner: Obtaining the network environment information of the edge-cloud network to which the edge device belongs, and constructing a reinforcement learning environment based on the network environment information; In the reinforcement learning environment, performing reinforcement learning on the distribution time learning model with the model distribution time as the reward.

9. The method according to claim 3, characterized in that, Determining the model deployment time required to deploy the target lightweight model on the edge device based on the model lightweighting time and the model distribution time includes: In response to detecting network congestion in the edge-cloud network to which the edge device belongs, determining the network delay duration of the edge device; Based on the network delay duration, the model lightweighting time, and the model distribution time, determining the model deployment time required to deploy the target lightweight model on the edge device.

10. The method according to any one of claims 1-9, characterized in that, Determining the target lightweight model required by the edge device from at least one lightweight model includes: Obtaining the performance parameter requirement conditions of the edge device from the model deployment request; Determining the lightweight model whose performance parameters can meet the performance parameter requirement conditions of the edge device as the target lightweight model required by the edge device.

11. The method according to any one of claims 1-9, characterized in that, Determining the deployment priority ranking of the edge device based on the model deployment time includes: Sorting the model deployment times corresponding to multiple edge devices in ascending order as the deployment priority ranking of each edge device.

12. A model deployment device, characterized in that, Including: A target lightweight model determination module, configured to, in response to receiving a model deployment request from an edge device, determine the target lightweight model required by the edge device from at least one lightweight model, where each lightweight model is generated by lightweighting a full model based on different lightweighting methods; A model deployment time determination module, configured to determine the model deployment time required to deploy the target lightweight model on the edge device; A priority ranking module, configured to determine the deployment priority ranking of the edge device based on the model deployment time; A model deployment module, configured to perform model deployment on the edge device based on the deployment priority ranking.

13. The device according to claim 12, characterized in that, Specifically, the model deployment time determination module is configured to: Determine the model lightweighting time and the corresponding model distribution time of the target lightweight model, where the model lightweighting time is the time required to obtain the target lightweight model by lightweighting the full model, and the model distribution time is the time required to distribute the target lightweight model to the edge device; Based on the model lightweighting time and the model distribution time, determine the model deployment time required to deploy the target lightweight model on the edge device.

14. The device according to claim 13, characterized in that, When determining the model lightweighting time corresponding to the target lightweight model, the model deployment time determination module is specifically configured to: Determine whether the target lightweight model exists in a preset model library; Based on whether the target lightweight model exists in the model library, determine the model lightweighting time corresponding to the target lightweight model.

15. The device according to claim 14, characterized in that, When determining the model lightweighting time corresponding to the target lightweight model based on whether the target lightweight model exists in the model library, the model deployment time determination module is specifically configured to: In response to the target lightweight model existing in the model library, determine the model lightweighting time corresponding to the target lightweight model as zero; In response to the target lightweight model not existing in the model library, determine the model lightweighting time corresponding to the target lightweight model based on the performance parameters of the target lightweight model.

16. The device according to claim 15, characterized in that, When determining the model deployment time consumption based on the performance parameters of the target lightweight model, the model deployment time consumption determination module is specifically configured to: Based on the correspondence between at least one preset performance parameter and the model lightweight time consumption, and based on the performance parameters of the target lightweight model, determine the model lightweight time consumption corresponding to the target lightweight model.

17. The device according to claim 16, wherein The performance parameters include at least one of the following: The memory amount required by the lightweight model; The computing power index of the lightweight model; The performance index of the lightweight model.

18. The device according to claim 14, characterized in that, When determining the model distribution time consumption corresponding to the target lightweight model, the model deployment time consumption determination module is specifically configured to: Input the network location of the edge device into the pre-trained distribution time consumption learning model, and obtain the model distribution time consumption corresponding to the target lightweight model output by the distribution time consumption learning model.

19. The device according to claim 18, characterized in that, The distribution time consumption learning model is a reinforcement learning model, and the distribution time consumption learning model is trained through the following method: Obtain the network environment information of the edge-cloud network to which the edge device belongs, and construct a reinforcement learning environment based on the network environment information; In the reinforcement learning environment, perform reinforcement learning on the distribution time consumption learning model with the model distribution time consumption as the reward.

20. The device according to claim 14, characterized in that, The model deployment time consumption determination module is specifically configured to: In response to detecting a network congestion situation in the edge-cloud network to which the edge device belongs, determine the network delay duration of the edge device; Based on the network delay duration, the model lightweight time consumption, and the model distribution time consumption, determine the model deployment time consumption required for deploying the target lightweight model on the edge device.

21. The device according to any one of claims 13-20, characterized in that, When determining the target lightweight model required by the edge device from at least one lightweight model, the target lightweight model determination module is specifically configured to: Obtain the performance parameter requirement conditions of the edge device from the model deployment request; Determine the lightweight model whose performance parameters can meet the performance parameter requirement conditions of the edge device as the target lightweight model required by the edge device.

22. The device according to any one of claims 13-20, characterized in that, The priority ranking module is specifically configured to: Sort the model deployment time consumptions corresponding to multiple edge devices in ascending order as the deployment priority ranking of each edge device.

23. An electronic device, characterized in that, Comprising a processor and a memory; The memory is used to store operation instructions; The processor is configured to execute the method according to any one of claims 1-12 by calling the operation instructions.

24. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1-12 is implemented.