End-side cloud model compression and deployment method based on split learning

By splitting the deep learning model into multiple sub-models under the end-edge cloud architecture and performing iterative pruning and knowledge distillation, the problems of insufficient adaptability of model compression and deployment and insufficient information migration in traditional methods are solved, and efficient model compression and low-latency end-side reasoning are achieved.

CN119940455AActive Publication Date: 2025-05-06GUANGDONG POLYTECHNIC NORMAL UNIV

Patent Information

Application Number
CN202510127152.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2025-05-06
Estimated Expiration
2045-01-31

AI Technical Summary

Technical Problem

Under the end-edge cloud architecture, the traditional model compression method is insufficiently adaptable to the end-edge cloud architecture, the traditional knowledge distillation method is insufficient to transfer information, and the lack of collaborative optimization strategies for the end-edge cloud architecture, resulting in problems such as large communication overhead and high inference delay.

Method used

The end-edge cloud model compression and deployment method based on split learning is used to split the deep learning model into front model, middle model and rear model, which are deployed on the client, edge side and cloud respectively, and the model is optimized through iterative pruning and knowledge distillation technology, and the lightweight front model is finally deployed to the client for local inference.

Benefits of technology

The effective compression and deployment of the model is realized, the computing burden and communication overhead of the end-side equipment is reduced, the inference efficiency and accuracy are improved, and the end-side local inference capability is ensured with low latency and low communication overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940455A_ABST
    Figure CN119940455A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses an end-side cloud model compression and deployment method based on split learning, comprising the steps of splitting a deep learning model into a front model, a middle model and a rear model, and deploying the models at a client, an edge side and a cloud; the client carries out forward propagation on the front model and sends a forward propagation result to the edge side; the edge side receives a forward propagation result from at least one client, performs forward propagation on the received result by using the middle model, and sends the forward propagation result to the cloud; the cloud receives a forward propagation result from at least one edge side, and completes forward propagation, calculates a loss function and performs back propagation by using the rear model so as to update the front model, the middle model and the rear model; performing iterative pruning on the front model, the middle model and the rear model, and performing fine adjustment on the pruned models by utilizing knowledge distillation; and deploying the lightened front model to a client so as to carry out reasoning locally at the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to an end-edge cloud model compression and deployment method based on split learning. Background Art

[0002] In recent years, deep learning technology has achieved significant breakthroughs in various fields and has been widely used in tasks such as image recognition and natural language processing. However, deploying computationally intensive deep learning models to resource-constrained end-side devices (such as smartphones, IoT devices, etc.) still faces severe challenges. Although the traditional cloud-based inference model has powerful computing resources, it requires uploading terminal data to the cloud for processing, which not only brings significant communication overhead, but also introduces higher inference latency and may cause the risk of user privacy leakage.

[0003] To address these challenges, the edge-cloud collaborative architecture came into being. This architecture aims to leverage edge computing power to share computing pressure on the cloud and reduce data transmission between the edge and the cloud, thereby alleviating latency and communication overhead issues to a certain extent. However, even under the edge-cloud architecture, it is still impractical to directly deploy complete deep learning models to resource-limited edge devices. Therefore, model compression technology has become the key to achieving efficient edge-side reasoning.

[0004] At present, mainstream model compression methods include pruning, quantization, knowledge distillation, etc. Traditional model pruning methods are usually directly applied to the complete model to reduce the model size and computational complexity by removing unimportant connections or neurons in the model. However, under the edge-cloud architecture, if only the cloud model is pruned and deployed directly to the end side, it may still not meet the resource constraints of the end-side device and ignore the distributed characteristics of the edge-cloud architecture. In addition, simple pruning operations can easily lead to a significant decrease in model accuracy.

[0005] As an emerging distributed training method, split learning splits deep learning models to different devices for training, which can effectively protect data privacy. However, existing split learning methods mainly focus on the distributed training process of the model, and pay less attention to model compression and deployment. Directly deploying the split model to the end side may still cause the model to be too large, and the advantages of local inference on the end side cannot be fully utilized. Specifically:

[0006] Traditional model compression methods are not compatible with the edge-cloud architecture: Traditional pruning and quantization methods are usually performed on the complete model without fully considering the distributed characteristics and communication constraints of the edge-cloud architecture. As a result, the compressed model may still be too large or the communication overhead is still high when deployed on the edge.

[0007] Insufficient information transfer in traditional knowledge distillation methods: Traditional knowledge distillation mainly focuses on the knowledge transfer of the output layer, ignoring the rich feature information of the middle layer of the model. Under the edge-cloud architecture, especially when the model is split, this may cause the student model to be unable to fully learn the knowledge of the teacher model, and the accuracy recovery effect is limited.

[0008] Lack of collaborative optimization strategies for the end-edge-cloud architecture: Existing technologies rarely consider model splitting, compression, and deployment as a whole, lack collaborative optimization strategies for the characteristics of the end-edge-cloud architecture, and are unable to fully utilize the advantages of the end-edge-cloud architecture to achieve end-side reasoning with low latency and low communication overhead.

[0009] It is precisely because of the above defects in the existing technology that when deploying deep learning models under the edge-cloud architecture, there are still problems such as high communication overhead and high inference latency. Therefore, there is an urgent need for a technical solution that can fully utilize the characteristics of the edge-cloud architecture, efficiently compress the model and ensure accuracy, so as to achieve low-latency and low-communication overhead local inference on the edge side. Summary of the invention

[0010] The present invention provides an end-edge-cloud model compression and deployment method based on split learning, which solves the technical problems of insufficient adaptability of traditional model compression methods to end-edge-cloud architecture, insufficient information migration of traditional knowledge distillation methods, and lack of collaborative optimization strategies for end-edge-cloud architecture.

[0011] The present invention provides a method for compressing and deploying a terminal-edge-cloud model based on split learning, comprising the following steps:

[0012] S1: Split the deep learning model to be compressed into the front model, middle model and back model, and deploy them on the client, edge side and cloud side respectively;

[0013] S2: The client uses local data to perform forward propagation on the front model deployed on the client, and sends the forward propagation result to the edge side;

[0014] S3: The edge side receives the forward propagation result from at least one client, performs forward propagation on the received result using the central model deployed on the edge side, and sends the forward propagation result to the cloud;

[0015] S4: The cloud receives the forward propagation result from at least one edge side, completes the forward propagation using the back-end model deployed on the cloud, calculates the loss function, and performs back propagation based on the loss function to update the parameters of the front model, the middle model, and the back-end model;

[0016] S5: Iteratively prune the front model deployed on the client, the middle model deployed on the edge, and the back model deployed on the cloud, and use knowledge distillation to fine-tune the pruned models;

[0017] S6: Deploy the lightweight front model to the client to perform inference locally on the client.

[0018] Preferably, in step S1, the deep learning model is manually split into a front model, a middle model and a back model according to resource limitations of the client, the edge side and the cloud.

[0019] 1. Device-edge-cloud resources

[0020] (1) Client-side resource situation

[0021] Assume that there are m different end-side devices, denoted as clients C1, C2, ..., C m For each client, we can describe its resource situation through several indicators.

[0022] 1. The client can accept computing power C comp .use Represents client C k (k=1,2,...,m) Acceptable computing power, measured in floating-point operations per second (FLOPs).

[0023] 2. Client available memory capacity C mem .use Represents client C k (k=1,2,...,m) available memory capacity in bytes.

[0024] (2) Edge-side resource situation

[0025] Assume that there are n different edge devices, denoted as edge E1, E2, ..., E n For each edge side, we can describe its resource situation through several indicators.

[0026] 1. Acceptable computing capacity E on the edge side comp .use Indicates client E k (k=1,2,...,n) Acceptable computing power, measured in floating-point operations per second (FLOPs).

[0027] 2. Available memory capacity E on the edge side mem .use Represents client C kThe available memory capacity for (k = 1, 2,..., n) is in bytes (Byte).

[0028] (3) Cloud resource situation

[0029] The acceptable computing power of the cloud is represented by Cloud comp while the available memory capacity is represented by Cloud mem to represent.

[0030] II. Determination of dynamic splitting points

[0031] (1) Traverse the combination of splitting points

[0032] Suppose the model M = {M1, M2,..., M L} with L layers is split into three parts (front part M 1:x , middle part M x:y , rear part M y:L ) according to the splitting points x and y, and are respectively deployed to the edge side, the edge side, and the cloud side. Among them, 1 < x < y < L. By traversing the splitting points x and y and calculating the scores in each case, the optimal splitting points x and y are determined. Start traversing from x = 2 and end at x = L - 2. For each x value: start traversing from y = x + 1 and end at y = L - 1. For each combination of x and y, the following evaluation is performed.

[0033] (2) Evaluation

[0034] For the current splitting points x and y, the performance indicators of each part of the model are as follows:

[0035] The computing power required by the model M comp . The computing power required for each part is expressed as: front part middle part rear part

[0036] The memory capacity required by the model M mem . The computing power required for each part is expressed as: front part middle part rear part

[0037] During evaluation, we mainly consider the communication cost and training delay during the training process.

[0038] The communication cost can be regarded as the sum of the communication costs between the edge and the cloud under the premise of training once. The communication between the edge and the cloud is closely related to x. We use C x-cost to represent the communication cost between a certain edge and a certain cloud. Then the total communication cost between m different edge devices and the edge side is C ce-cost = m·C x-costSimilarly, using C y-cost To represent the communication cost between the edge and the cloud, the total communication cost between n different edge devices and the cloud is C ec-cost =n·C y-cost .

[0039] In summary, the total communication cost is C cost =C ce-cost +C ec-cost

[0040] The communication cost score is calculated by mapping the communication cost to a specific scoring zone, and the minimum communication cost is set to The maximum value is (can be determined based on actual experience), the communication cost score calculation formula is:

[0041]

[0042] The training latency can be regarded as the sum of the inference latency of each part of the model on the device, edge, and cloud under the premise of training once. We use T client-delay , T edge-delay and T cloud-delay To represent the training delay on the client side, edge side, and cloud side. The total training delay is: T delay =T client-delay +T edge-delay +T cloud-delay .

[0043] The training delay score is calculated by mapping the training delay to a specific scoring zone, assuming that the minimum training delay is The maximum value is (can be determined based on actual experience), the communication cost score calculation formula is:

[0044]

[0045] Taking the above factors into consideration, the total score is calculated in a weighted manner. The weight of the communication cost score is α, and the weight of the training delay score is β (and α+β=1, the weight value can be set according to the emphasis of the actual application scenario on communication cost and inference delay). The calculation formula is as follows:

[0046] score=(α·score c +β·score T )·θ

[0047] Among them, θ is the penalty factor (initialized to 1, the penalty state is a very small number, such as 0.001), and the score is penalized when the following situations occur.

[0048] On the client side: or

[0049] Edge side: or

[0050] Cloud: or

[0051] Finally, the best split points x and y that best match the resource environment are selected based on the traversal.

[0052] Preferably, in step S2, the client processes the local data and performs forward propagation of the front model until a predefined cut layer; the output of the client forward propagation is expressed as:

[0053]

[0054] in, is the ith front model, x (i) is the input data of the i-th client device, are the parameters of the ith front model.

[0055] Preferably, in step S3, the edge side receives the intermediate results sent by all clients And perform aggregation to get h1, and continue to perform forward propagation as the input of the edge side. The output of the forward propagation on the edge side is expressed as:

[0056]

[0057] in is the jth middle model, are the parameters of the j-th middle model.

[0058] Preferably, in step S4, the cloud receives the forward propagation result from at least one edge side, and completes the forward propagation using the back-end model deployed on the cloud, including: the cloud receives the intermediate result sent by the edge side And aggregate to get h2, and complete the final forward propagation to get the model output y pred ; The output of the forward propagation in the cloud is expressed as:

[0059] y pred =f Cloud (h2,θ Cloud )

[0060] where f Cloud is the cloud model, θ Cloud are the parameters of the cloud model.

[0061] Preferably, in step S4, the calculating of the loss function and performing back propagation based on the loss function to update the parameters of the front model, the middle model and the rear model include:

[0062] Cloud based on y pred And the actual label y calculates the loss function L = (y pred ,y), the cloud side calculates the gradient of the loss function with respect to the model parameters and passes it to the edge side;

[0063]

[0064] The edge side updates its model part based on the gradient and calculates the gradient passed to the client;

[0065]

[0066] The client updates its part of the model based on the gradient, and finally completes the model update.

[0067] Preferably, in step S5, the fine-tuning of the pruned model using knowledge distillation technology includes: using the unpruned model part as a teacher and the pruned model part as a student, and fine-tuning using split multi-level knowledge distillation, wherein the multi-level distillation fine-tuning includes intermediate feature loss and soft label loss, and simultaneously introducing the cross entropy loss of the hard label to constitute the total loss of the multi-level distillation fine-tuning, and training the student model.

[0068] Preferably, the intermediate feature loss is expressed as:

[0069]

[0070] in, Represented as a feature map of the student model; is the feature map of its corresponding teacher model A; r(·) is a regression variable composed of a 1×1 convolutional layer and a BN layer; D p It is a measure of the L2 distance between the student and teacher feature maps;

[0071]

[0072] in, is the middle feature of the front model on the end side, It is the intermediate feature of the middle model on the edge side;

[0073] The soft label loss is expressed as:

[0074]

[0075] where x ijrepresents the student model logical output of the jth class of the i-th batch of samples; X ij and They represent the soft outputs of the student model and teacher model A of the jth class of the i-th batch of samples respectively; the temperature parameter T determines the degree of softening of the output;

[0076] The cross entropy loss introduced with hard labels is expressed as:

[0077]

[0078] in, Represents the logical output of the jth class of the i-th batch of samples; Y ij Represents the hard label of the jth category of the i-th batch of samples;

[0079] The total loss of the multi-stage distillation fine-tuning is expressed as:

[0080] L=δl inter +εl output +θl CE

[0081] Among them, δ, ε and θ are weight values ​​representing intermediate feature loss, soft label loss and hard label loss respectively, and δ+ε+θ=1.

[0082] Preferably, in step S6, deploying the lightweight front model to the client includes:

[0083] Combine the front model with the middle model and the back model;

[0084] Deploy the combined model to the client to complete the inference task locally on the client;

[0085] During the local reasoning process of the client, there is no need to transmit data to the edge side or the cloud, so as to reduce communication overhead and reasoning delay.

[0086] Preferably, the combining of the front model with the middle model and the rear model is specifically as follows:

[0087]

[0088] Among them, f ji is the combination of the cloud model, the jth edge model, and the jth client model.

[0089] Compared with the prior art, the present invention has the following beneficial effects:

[0090] The present invention discloses an end-edge cloud model compression and deployment method based on split learning. Model splitting and distributed training reduce the computing burden of a single device; the aggregation effect on the edge side reduces the computing burden and communication overhead on the cloud side; iterative pruning effectively reduces the model size and computational complexity; knowledge distillation technology compensates for the accuracy loss caused by pruning; and local reasoning on the end side achieves low latency and low communication overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Figure 1 It is a flowchart of a method for compressing and deploying an end-edge cloud model based on split learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0092] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0093] like Figure 1 As shown, the present application provides a method for compressing and deploying a terminal-edge-cloud model based on split learning, comprising the following steps:

[0094] S1: Split the deep learning model to be compressed into the front model, middle model and back model, and deploy them on the client, edge side and cloud side respectively;

[0095] S2: The client uses local data to perform forward propagation on the front model deployed on the client, and sends the forward propagation result to the edge side;

[0096] S3: The edge side receives the forward propagation result from at least one client, performs forward propagation on the received result using the central model deployed on the edge side, and sends the forward propagation result to the cloud;

[0097] S4: The cloud receives the forward propagation result from at least one edge side, completes the forward propagation using the back-end model deployed on the cloud, calculates the loss function, and performs back propagation based on the loss function to update the parameters of the front model, the middle model, and the back-end model;

[0098] S5: Iteratively prune the front model deployed on the client, the middle model deployed on the edge, and the back model deployed on the cloud, and use knowledge distillation to fine-tune the pruned models;

[0099] S6: Deploy the lightweight front model to the client to perform inference locally on the client.

[0100] Preferably, in step S1, the deep learning model is manually split into a front model, a middle model and a back model according to resource limitations of the client, the edge side and the cloud.

[0101] In the above scheme, model splitting is the premise of model deployment. The split sub-models need to be deployed on the corresponding computing nodes according to their functions and data dependencies; the front model is responsible for processing the original input data, the middle model is responsible for extracting and aggregating the middle-layer features, and the back model completes the final prediction or classification task. The network structure of the deep learning model to be compressed is first analyzed to identify the sub-networks suitable for running on different computing nodes. For example, a shallow network with a small amount of computation and high real-time requirements can be deployed on the client as a front model; a middle-layer network with a moderate amount of computation and the need to aggregate data from multiple clients can be deployed on the edge as a middle model; a deep network with a large amount of computation and relatively low real-time requirements can be deployed on the cloud as a back model. This splitting method can effectively utilize the distributed computing capabilities of the edge-cloud architecture. Through model splitting, it is initially possible to decompose large models into parts that can be run on resource-constrained clients, laying the foundation for subsequent local reasoning on the end side, and also providing the possibility of using the computing power of the edge side to share the pressure on the cloud.

[0102] Preferably, in step S2, the client processes the local data and performs forward propagation of the front model until a predefined cut layer; the output of the client forward propagation is expressed as:

[0103]

[0104] in, is the ith front model, x (i) is the input data of the i-th client device, are the parameters of the ith front model.

[0105] In the above scheme, the client refers to the end-side device that executes the input layer and some hidden layers of the model. These devices usually have limited computing and storage resources. The client uses the locally collected data as input to drive the front-end model deployed on it to perform calculations and generate intermediate feature representations. The forward propagation result is the key data link between the client and the edge side. After the client collects the local data, it inputs it into the front-end model for layer-by-layer calculation until the output layer of the front-end model. The output result, usually a high-dimensional feature vector, is serialized and transmitted to the edge side through the network. This process reflects the initial data processing capabilities of the end side and the collaborative working mode with the edge side. Using the client's computing resources for preliminary data processing reduces the amount of data that needs to be transmitted to the edge side and reduces communication overhead.

[0106] Preferably, in step S3, the edge side receives the intermediate results sent by all clients And perform aggregation to get h1, and continue to perform forward propagation as the input of the edge side. The output of the forward propagation on the edge side is expressed as:

[0107]

[0108] in is the jth middle model, are the parameters of the j-th middle model.

[0109] In the above scheme, the edge side receives the forward propagation results from multiple clients, which serve as the input of the central model; as a bridge between the client and the cloud, it has stronger computing power than the end-side device and can handle more complex tasks or larger model parts; the edge side, as an intermediate layer, is responsible for aggregating local information from multiple clients. The received forward propagation results are input into the central model for further feature extraction and fusion. The output results of the central model, usually a higher-level feature representation, are serialized and transmitted to the cloud through the network. The edge side plays a connecting role in this process, which not only shares the computing pressure of the client, but also reduces the amount of data that needs to be processed in the cloud. Using the computing power of the edge side to aggregate and process the intermediate layer features further reduces the computing burden and communication overhead of the cloud.

[0110] Preferably, in step S4, the cloud receives the forward propagation result from at least one edge side, and completes the forward propagation using the back-end model deployed on the cloud, including: the cloud receives the intermediate result sent by the edge side And aggregate to get h2, and complete the final forward propagation to get the model output y pred ; The output of the forward propagation in the cloud is expressed as:

[0111] y pred =f Cloud (h2,θ Cloud )

[0112] where f Cloud is the cloud model, θ Cloud are the parameters of the cloud model.

[0113] In the above solution, the cloud has powerful computing and storage resources and is responsible for executing the final part of the model, including the final hidden layer and output layer, and performing the final optimization of the model. After receiving the feature representation from the edge side, the cloud inputs it into the rear model to complete the final prediction or classification task.

[0114] Preferably, in step S4, the calculating of the loss function and performing back propagation based on the loss function to update the parameters of the front model, the middle model and the rear model include:

[0115] Cloud based on y pred And the actual label y calculates the loss function L = (y pred ,y), the cloud side calculates the gradient of the loss function with respect to the model parameters and passes it to the edge side;

[0116]

[0117] The edge side updates its model part based on the gradient and calculates the gradient passed to the client;

[0118]

[0119] The client updates its part of the model based on the gradient, and finally completes the model update.

[0120] In the above scheme, the calculation of the loss function depends on the output of the rear model and the true label. Back propagation uses the gradient information of the loss function to update the parameters of the entire distributed model; based on the model's prediction results and the true label, the loss function is calculated to measure the performance of the model. Then, using the back propagation algorithm, the gradient information of the loss function is transmitted back layer by layer, and the parameters of the front model, middle model, and rear model are updated according to the gradient information. This process realizes distributed training with end-edge-cloud collaboration. The powerful computing power of the cloud is used to complete the final training and optimization of the model to ensure the overall performance of the model. By updating the parameters of the distributed model through back propagation, end-edge-cloud collaborative training is realized, overcoming the problem of data islands.

[0121] Preferably, in step S5, the fine-tuning of the pruned model by using knowledge distillation includes: taking the unpruned model part as a teacher and the pruned model part as a student, and fine-tuning by using split multi-level knowledge distillation, wherein the multi-level distillation fine-tuning includes intermediate feature loss and soft label loss, and simultaneously introducing the cross entropy loss of the hard label to constitute the total loss of the multi-level distillation fine-tuning, and training the student model.

[0122] After the training phase is completed, the original model is obtained and pruned. It is pruned iteratively according to the preset pruning rate.

[0123] For the kth filter in the convolutional layer, its importance is calculated by the L1 norm:

[0124]

[0125] Among them, C in is the number of input channels, K h ×Kw is the convolution kernel size, W k,i,j Represents the weight value of the k-th filter at the i-th input channel and the j-th spatial position.

[0126] The model has L layers in total, and each layer has an independent preset pruning rate:

[0127]

[0128] in, is the initial pruning rate of the lth layer, μ (l) is the attenuation coefficient of the lth layer, which can control the attenuation speed of the pruning rate of each layer.

[0129] The cumulative pruning rate of the lth layer after t rounds of iterations:

[0130]

[0131] Global model cumulative pruning rate:

[0132]

[0133] Among them, Params (l) is the parameter quantity of the lth layer.

[0134] The iteration termination condition is as follows: reaching the target pruning rate.

[0135] GlobalPruneRate t ≥MaxGlobalPruneRate

[0136] Preferably, the intermediate feature loss is expressed as:

[0137]

[0138] in, Represented as a feature map of the student model; is the feature map of its corresponding teacher model A; r(·) is a regression variable composed of a 1×1 convolutional layer and a BN layer; D p It is a measure of the L2 distance between the student and teacher feature maps;

[0139]

[0140] in, is the middle feature of the front model on the end side, It is the intermediate feature of the middle model on the edge side;

[0141] In the above scheme, an adaptive layer consisting of a point convolution (1×1 convolution kernel) and a batch normalization (BN) layer is introduced; the role of this adaptive layer is to map the channels of the student model to the corresponding channels of the teacher model, so as to transfer knowledge more efficiently and reduce the difference in feature mapping between the pruned model and the original model. Through this mapping strategy, it can be ensured that the structural changes of the student model will not have a negative impact on the effect of knowledge distillation.

[0142] The soft label loss is expressed as:

[0143]

[0144] where x ij represents the student model logical output of the jth class of the i-th batch of samples; X ij and They represent the soft outputs of the student model and teacher model A of the jth class of the i-th batch of samples respectively; the temperature parameter T determines the degree of softening of the output;

[0145] In the above scheme, the output soft labels of simulated distillation learning are simulated. In order to learn more from the teacher model, it is also necessary to simulate the softened teacher output. Specifically, the KL divergence loss between the student and teacher outputs is used as the distillation loss of the output simulation. The temperature T softens the output between the student and teacher; enabling the student model to more effectively learn the prediction results of the high-performance teacher model, thereby significantly reducing the classification error rate.

[0146] The cross entropy loss introduced with hard labels is expressed as:

[0147]

[0148] in, Represents the logical output of the jth class of the i-th batch of samples; Y ij Represents the hard label of the jth category of the i-th batch of samples;

[0149] The total loss of the multi-stage distillation fine-tuning is expressed as:

[0150] L=δl inter +εl output +θl CE

[0151] Among them, δ, ε and θ are weight values ​​representing intermediate feature loss, soft label loss and hard label loss respectively, and δ+ε+θ=1.

[0152] Preferably, in step S6, deploying the lightweight front model to the client includes:

[0153] Combine the front model with the middle model and the back model;

[0154] Deploy the combined model to the client to complete the inference task locally on the client;

[0155] During the local reasoning process of the client, there is no need to transmit data to the edge side or the cloud, so as to reduce communication overhead and reasoning delay.

[0156] Preferably, the combining of the front model with the middle model and the rear model is specifically as follows:

[0157]

[0158] Among them, f ji is the combination of the cloud model, the jth edge model, and the jth client model.

[0159] In the above solution, the carefully compressed and optimized model is deployed to the end device to achieve complete integration of the model. This deployment strategy enables subsequent reasoning tasks to be performed completely locally on the client, without the need to upload data to the edge or cloud for processing. This significantly reduces the communication overhead caused by data transmission, and also reduces the time delay in the reasoning process.

[0160] By performing inference on the edge, the system’s responsiveness is improved and user privacy is enhanced, as sensitive data no longer needs to leave the user’s device. In addition, this approach reduces the computational burden on the cloud and edge, allowing them to allocate resources to other tasks, thereby improving the efficiency and scalability of the entire system.

[0161] In practical applications, this localized reasoning strategy is particularly suitable for scenarios with high real-time requirements, such as autonomous driving, augmented reality, and instant speech recognition. In these scenarios, fast decision-making and response are crucial, and any time delay caused by network latency or bandwidth limitations may affect user experience or system performance.

[0162] In short, deploying the compressed model directly to the end side and completing all inference tasks locally not only optimizes the end-to-end system performance, but also provides users with a faster and more reliable service experience.

[0163] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for compressing and deploying a device-edge-cloud model based on split learning, characterized in that: The following steps are involved: S1: Split the deep learning model to be compressed into the front model, middle model and back model, and deploy them on the client, edge side and cloud side respectively; S2: The client uses local data to perform forward propagation on the front model deployed on the client, and sends the forward propagation result to the edge side; S3: The edge side receives the forward propagation result from at least one client, performs forward propagation on the received result using the central model deployed on the edge side, and sends the forward propagation result to the cloud; S4: The cloud receives the forward propagation result from at least one edge side, completes the forward propagation using the back-end model deployed on the cloud, calculates the loss function, and performs back propagation based on the loss function to update the parameters of the front model, the middle model, and the back-end model; S5: Iteratively prune the front model deployed on the client, the middle model deployed on the edge, and the back model deployed on the cloud, and use knowledge distillation to fine-tune the pruned models; S6: Deploy the lightweight front model to the client to perform inference locally on the client.

2. According to the method for compressing and deploying a split learning-based edge-cloud model according to claim 1, it is characterized in that: In step S1, the deep learning model is manually split into a front model, a middle model, and a back model according to the resource constraints of the client, edge side, and cloud side.

3. According to the split learning-based end-edge cloud model compression and deployment method of claim 2, it is characterized in that: In step S2, the client processes the local data and performs the forward propagation of the previous model until a predefined cut layer; the output of the client forward propagation is expressed as: in, is the ith front model, x (i) is the input data of the i-th client device, are the parameters of the ith front model.

4. According to the method of claim 3, the method is characterized in that: In step S3, the edge side receives the intermediate results sent by all clients And perform aggregation to get h1, and continue to perform forward propagation as the input of the edge side. The output of the forward propagation on the edge side is expressed as: in is the jth middle model, are the parameters of the j-th middle model.

5. According to the method for compressing and deploying a split learning-based edge-cloud model according to claim 4, it is characterized in that: In step S4, the cloud receives the forward propagation result from at least one edge side, and uses the back-end model deployed on the cloud to complete the forward propagation, including: the cloud receives the intermediate result sent by the edge side And aggregate to get h2, and complete the final forward propagation to get the model output y pred ; The output of the forward propagation in the cloud is expressed as: y pred =f Cloud (h2,θ Cloud ) where f Cloud is the cloud model, θ Cloud are the parameters of the cloud model.

6. According to the method of split learning-based end-edge cloud model compression and deployment in claim 5, it is characterized in that: In step S4, the calculation of the loss function and the back propagation based on the loss function to update the parameters of the front model, the middle model and the rear model include: Cloud based on y pred And the actual label y calculates the loss function L = (y pred ,y), the cloud side calculates the gradient of the loss function with respect to the model parameters and passes it to the edge side; The edge side updates its model part based on the gradient and calculates the gradient passed to the client; The client updates its part of the model based on the gradient, and finally completes the model update.

7. According to claim 6, a split learning-based end-edge cloud model compression and deployment method is characterized in that: In step S5, the use of knowledge distillation technology to fine-tune the pruned model includes: using the unpruned model part as a teacher and the pruned model part as a student, and using split multi-level knowledge distillation for fine-tuning. The multi-level distillation fine-tuning includes intermediate feature loss and soft label loss, and at the same time introduces the cross entropy loss of the hard label to constitute the total loss of the multi-level distillation fine-tuning, and trains the student model.

8. According to claim 7, a split learning-based end-edge cloud model compression and deployment method is characterized in that: The intermediate feature loss is expressed as: in, Represented as a feature map of the student model; is the feature map of its corresponding teacher model A; r(·) is a regression variable composed of a 1×1 convolutional layer and a BN layer; D p It is a measure of the L2 distance between the student and teacher feature maps; in, is the middle feature of the front model on the end side, It is the intermediate feature of the middle model on the edge side; The soft label loss is expressed as: where x ij represents the student model logical output of the jth class of the i-th batch of samples; X ij and They represent the soft outputs of the student model and teacher model A of the jth class of the i-th batch of samples respectively; the temperature parameter T determines the degree of softening of the output; The cross entropy loss introduced with hard labels is expressed as: in, Represents the logical output of the jth class of the i-th batch of samples; Y ij Represents the hard label of the jth category of the i-th batch of samples; The total loss of the multi-stage distillation fine-tuning is expressed as: L=δl inter +εl output +θl CE Among them, δ, ε and θ are weight values ​​representing intermediate feature loss, soft label loss and hard label loss respectively, and δ+ε+θ=1.

9. According to claim 8, a split learning-based end-edge cloud model compression and deployment method is characterized in that: In step S6, deploying the lightweight front model to the client includes: Combine the front model with the middle model and the back model; Deploy the combined model to the client to complete the inference task locally on the client; During the local reasoning process of the client, there is no need to transmit data to the edge side or the cloud, so as to reduce communication overhead and reasoning delay.

10. The method for compressing and deploying a split learning-based edge-cloud model according to claim 9, characterized in that: The combination of the front model, the middle model and the rear model is specifically as follows: Among them, f ji is the combination of the cloud model, the jth edge model, and the jth client model.

Citation Information

Patent Citations

  • Cloud-edge co-learning power transmission inspection method and system

    CN115272981A

  • Deep learning system for edge device

    CN118467164A

  • Distributed AI training service-oriented 6G computing power network adaptive splitting federated learning method

    CN119031415A

  • Federal learning-based time delay and energy consumption optimization method

    CN119323240A

  • Model training method based on edge-end collaboration and related equipment

    CN119358637A

Cited By

  • Privacy protection method and system based on multi-modal physiological signal calculation

    CN122451953A