End-side cloud model segmentation method based on dynamic adjustment
By dynamically adjusting the segmentation points of the end edge cloud model, the problems of unbalanced resource utilization and limited system performance under the static segmentation strategy are solved, and adaptive resource utilization and system performance optimization are achieved.
Patent Information
- Application Number
- CN202510127154.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-31
- Publication Date
- 2025-05-13
AI Technical Summary
The existing end-edge cloud model segmentation method adopts a static segmentation strategy, which cannot adapt to the dynamic changes in the end-edge cloud environment, resulting in unbalanced resource utilization, limited system performance and insufficient robustness.
The end-edge cloud model segmentation method based on dynamic adjustment is adopted. By evaluating the resource situation of the end-side and cloud-side devices in real time, the segmentation points of the deep learning model are dynamically adjusted to adapt to resource changes, and distributed training is carried out to update the model parameters.
It realizes adaptive utilization of resources, optimizes system performance, improves the adaptability and stability of the model in complex environments, and reduces the computing power and communication burden of the end-side equipment.
Smart Images

Figure CN119988028A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model segmentation technology, and in particular to an end-edge-cloud model segmentation method based on dynamic adjustment. Background Art
[0002] With the increasing popularity of the end-edge-cloud collaborative computing architecture, the deployment model of deep learning models is undergoing profound changes. In order to fully utilize the data generation capabilities of end-side devices, the local computing advantages of the edge side, and the powerful computing resources of the cloud, model segmentation deployment has become a key technical means. This model aims to split a complete deep learning model into multiple parts and deploy them in a targeted manner according to the characteristics of the end, edge, and cloud, thereby optimizing resource utilization and system performance.
[0003] However, there are still many challenges in the field of edge-cloud model segmentation deployment. Early research focused on model compression techniques, such as pruning, quantization, and knowledge distillation. These techniques aim to reduce model size and computational complexity so that they can run on resource-constrained edge devices. Although these technologies have alleviated the pressure of edge-side deployment to a certain extent, they do not essentially solve the problem of resource collaborative utilization under the edge-cloud architecture.
[0004] With the in-depth development of the concept of end-edge-cloud collaborative computing, model segmentation deployment has gradually become a research hotspot. Existing model segmentation methods, such as static segmentation strategies based on model hierarchy or preset rules, determine the segmentation points at the beginning of model design or deployment and keep them unchanged during subsequent operation. Although this static segmentation method is simple to implement, it has obvious limitations in practical applications and has become a key bottleneck restricting the performance improvement of end-edge-cloud systems.
[0005] Low resource utilization and even resource bottlenecks: The prominent features of the end-edge cloud environment are resource heterogeneity and dynamic changes. There are many types of end-side devices, with huge differences in computing power, memory capacity, and network bandwidth; edge-side and cloud-side resources may also be affected by load fluctuations. Static segmentation points are pre-set and cannot perceive and adapt to these dynamic changes, which can easily lead to unbalanced resource allocation. For example, when the end-side device resources are idle, the edge or cloud side may become a performance bottleneck due to excessive load, resulting in inefficient overall system efficiency. The defect is that the static segmentation strategy lacks consideration of the dynamic nature of end-edge cloud environment resources and cannot achieve on-demand resource allocation.
[0006] System performance is limited and difficult to achieve optimality: Fixed split points are not optimal in all scenarios. For example, if the front part of the model is split to the end-side with weaker computing power, it may cause the end-side inference delay to be too long, becoming a system performance bottleneck. Conversely, if most of the model is deployed to the cloud, the end-to-cloud communication overhead and latency will increase. The static split strategy lacks flexibility and cannot be optimized according to the actual resource conditions, making it difficult for the system performance to reach the optimal level. The defect is that the static split strategy lacks a dynamic evaluation and optimization mechanism for the impact of the model split point performance, and cannot adjust the split scheme according to actual conditions.
[0007] Poor environmental adaptability and insufficient robustness: The complexity and dynamics of the edge-cloud environment place higher demands on model deployment. For example, fluctuations in network conditions, dynamic changes in device resources, and changes in user needs may affect model performance. Fixed segmentation points cannot adapt to these environmental changes, resulting in unstable performance and poor robustness of the model in different scenarios. The defect is that the static segmentation strategy lacks the ability to perceive and adapt to environmental changes, and cannot guarantee the stability and reliability of the model in a complex dynamic environment. Summary of the invention
[0008] The present invention provides an end-edge-cloud model segmentation method based on dynamic adjustment, which realizes the adaptive utilization of the heterogeneous characteristics of end-side and cloud-side resources, takes into account both the efficient distributed collaboration in the training phase and the localized and efficient execution in the reasoning phase, significantly reduces the computing power burden and communication burden of the end-side devices, and effectively improves the overall operating performance and stability of the system.
[0009] The present invention provides a method for segmenting a terminal-edge-cloud model based on dynamic adjustment, comprising the following steps:
[0010] S1: Split the deep learning model into at least two parts to be deployed on the end-side device and the cloud-side device respectively;
[0011] S2: Real-time evaluation of the resource status of the end-side devices and cloud-side devices;
[0012] S3: Based on the resource situation assessment results, dynamically adjust the split point of the deep learning model to adapt to the resource situation of the terminal side device and the cloud side device;
[0013] S4: Based on the model after adjusting the split point, distributed training is performed in collaboration between the end-side devices and the cloud-side devices to update the deep learning model parameters.
[0014] Preferably, in step S1, the cloud-side device includes an edge-side device and a cloud-side device; the deep learning model is divided into three parts: a front part, a middle part, and a back part;
[0015] Among them, the front part is deployed on the terminal side device, the middle part is deployed on the edge side device, and the back part is deployed on the cloud side device.
[0016] Preferably, in step S2, the resource conditions of the terminal side device and the cloud side device include the computing power, memory capacity and network bandwidth of the terminal side device, the edge side device and the cloud side device; wherein,
[0017] Computing capability is used to characterize the computing processing capability of a device;
[0018] Memory capacity is used to characterize the available storage space of a device;
[0019] Network bandwidth is used to characterize the data transmission capacity between devices.
[0020] Preferably, the computing power is measured in floating point operations per second; and the memory capacity is measured in bytes.
[0021] Preferably, in step S3, the dynamically adjusting the segmentation points of the deep learning model includes:
[0022] Traverse different combinations of model split points;
[0023] For each combination of split points, evaluate the corresponding system performance indicators and resource consumption;
[0024] Based on the evaluation results, the optimal combination of split points is selected as the dynamically adjusted split points.
[0025] Preferably, the system performance indicators include communication costs and training delays; the resource consumption includes computing power requirements and memory capacity requirements of terminal-side devices, edge-side devices, and cloud-side devices.
[0026] Preferably, the evaluation of the corresponding system performance indicators and resource consumption is performed by calculating the total score of each combination of split points;
[0027] The total score is calculated as follows:
[0028] score=(α·score c +β·score T )·θ
[0029] Among them, score c Score for communication cost; score T is the training delay score; θ is the penalty factor; α and β are the weights of the communication cost score and the training delay score, and α+β=1. The weight value is set according to the emphasis on communication cost and inference delay in the actual application scenario.
[0030] Preferably, the calculation formula of the communication cost score is:
[0031]
[0032] in, is the minimum communication cost, is the maximum communication cost, C cost is the total communication cost under the current combination of split points;
[0033] The calculation formula of the training delay score is:
[0034]
[0035] in, is the minimum training delay, is the maximum training delay, T delay It is the total training delay under the current split point combination.
[0036] Preferably, the penalty factor is initialized to 1, and when the following situations occur, the penalty on the total score becomes 0.01;
[0037] The computing power required by the end device is greater than the computing power that the end device can accept.
[0038] The memory capacity required by the end device is greater than the available memory capacity of the end device.
[0039] The computing power required by edge devices > the computing power acceptable to edge devices;
[0040] The memory capacity required by the edge device is greater than the available memory capacity of the edge device.
[0041] The computing power required by the cloud device > the computing power that the cloud device can accept;
[0042] The memory capacity required by the cloud device > the available memory capacity of the cloud device.
[0043] Preferably, in step S4, the distributed training adopts a federated learning method, including:
[0044] Perform local model training on the client device.
[0045] Upload the trained model parameters to the cloud;
[0046] Aggregate model parameters in the cloud;
[0047] Distribute the aggregated model parameters to the end-side devices.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The present invention discloses an end-edge-cloud model segmentation method based on dynamic adjustment, which introduces an adjustment mechanism for dynamic model segmentation points. The model segmentation points can be adaptively adjusted according to the real-time resource conditions of the end-edge-cloud devices, thereby achieving balanced resource utilization, optimizing system performance, and improving the adaptability of the model in complex environments. The dynamic adjustment strategy breaks through the limitations of traditional fixed segmentation methods and provides a solution to the problems of unbalanced resource utilization and performance bottlenecks under the end-edge-cloud architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flow chart of a method for segmenting an end-edge cloud model based on dynamic adjustment provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] like Figure 1 As shown, the present application provides a method for segmenting an edge-cloud model based on dynamic adjustment, comprising the following steps:
[0053] S1: Split the deep learning model into at least two parts to be deployed on the end-side device and the cloud-side device respectively;
[0054] S2: Real-time evaluation of the resource status of the end-side devices and cloud-side devices;
[0055] S3: Based on the resource situation assessment results, dynamically adjust the split point of the deep learning model to adapt to the resource situation of the terminal side device and the cloud side device;
[0056] S4: Based on the model after adjusting the split point, distributed training is performed in collaboration between the end-side device and the cloud-side device to update the model parameters.
[0057] In the above scheme, the deep learning model is first preliminarily segmented in S1 and divided into at least two parts so as to be distributedly deployed on the end-side device and the cloud-side device. This initial segmentation lays the foundation for the subsequent dynamic adjustment; S2 is the key prerequisite for dynamic adjustment, that is, real-time evaluation of the resource status of the end-side device and the cloud-side device. The evaluation of the resource status is the basis for dynamically adjusting the segmentation point to ensure that the adjustment of the segmentation point can respond to changes in device resources; S3 is the core step of the present invention. Based on the resource evaluation result of step S2, the segmentation point of the deep learning model is dynamically adjusted. The purpose of the adjustment is to match the model segmentation point with the resource status of the current end-side device and the cloud-side device to achieve optimal resource utilization; S4 describes the distributed training performed after the segmentation point is adjusted. This training mode is the basis for the collaborative work of the end-side device and the cloud-side device, and the model training is completed together using their respective computing resources.
[0058] Real-time resource evaluation and dynamic split point adjustment mechanism enable the model splitting scheme to be adaptively adjusted according to resource changes of end-side and cloud-side devices, overcoming the defect that fixed split points cannot adapt to dynamic environments. By dynamically adjusting the split points, computing tasks can be more reasonably allocated to devices with sufficient resources, avoiding resource idleness and bottlenecks, thereby improving the overall resource utilization of the end-edge-cloud system. More reasonable resource allocation can reduce communication overhead and inference latency, and improve the overall system performance, such as reducing latency and increasing throughput. The dynamic adjustment mechanism makes the model deployment scheme more flexible, able to adapt to different end-edge-cloud environments and resource constraints, and improve the versatility and robustness of the model.
[0059] Preferably, in step S1, the cloud-side device includes an edge-side device and a cloud-side device; the deep learning model is divided into three parts: a front part, a middle part, and a back part;
[0060] Among them, the front part is deployed on the terminal side device, the middle part is deployed on the edge side device, and the back part is deployed on the cloud side device.
[0061] In the above solution, the cloud-side devices are refined, and it is clarified that the cloud-side devices include edge-side devices and cloud-side devices. The model is divided into three parts: front, middle and back, which are deployed on the end side, edge side and cloud side respectively. This three-stage split deployment architecture is more in line with the typical end-edge-cloud collaborative computing scenario.
[0062] The three-stage split deployment architecture can make more precise use of the resources of the end-side, edge-side and cloud-side devices, so that devices with different computing capabilities can undertake computing tasks that match their capabilities; deploying the central model on the edge-side device can effectively reduce the direct communication between the end-side device and the cloud-side device, and reduce network communication latency; the local computing capabilities of the edge-side device can accelerate data processing speed and improve the overall response speed of the system; the three-layer architecture can perform load balancing more flexibly and allocate computing tasks more reasonably according to the resource status of the three ends of the end, edge and cloud.
[0063] Preferably, in step S2, the resource conditions of the terminal device and the cloud device include the computing power, memory capacity and network bandwidth of the terminal device, the edge device and the cloud device; wherein,
[0064] Computing capability is used to characterize the computing processing capability of a device;
[0065] Memory capacity is used to characterize the available storage space of a device;
[0066] Network bandwidth is used to characterize the data transmission capacity between devices.
[0067] In the above solution, during the resource evaluation phase, the system needs to collect data such as computing power, memory capacity, and network bandwidth of the end-side devices, edge-side devices, and cloud-side devices. These data will be used to dynamically adjust the split point decision. For example, when the computing power of the end-side device is sufficient, more model parts can be deployed to the end-side; when the network bandwidth is limited, the amount of data transmission between the end-edge or edge-cloud can be reduced.
[0068] Clarifying resource evaluation indicators as computing power, memory capacity, and network bandwidth makes the resource evaluation process more quantitative and operational, and facilitates system automation execution; comprehensively considering the three key resource indicators of computing power, memory capacity, and network bandwidth can more comprehensively reflect the resource status of edge-cloud devices, and provide guarantees for more accurate dynamic segmentation point adjustments; based on more comprehensive resource evaluation results, more accurate dynamic segmentation point adjustments can be made, so that the model segmentation plan is more in line with the actual resource status.
[0069] Preferably, the computing power is measured in floating point operations per second; and the memory capacity is measured in bytes.
[0070] In the above scheme, during the resource evaluation phase, the system needs to collect computing power data in FLOPs and memory capacity data in Bytes. These standardized data can be directly used in the subsequent split point adjustment algorithm for quantitative comparison and decision-making. Clarifying the measurement units of computing power and memory capacity makes the resource evaluation results more standardized and comparable, facilitating resource comparisons between different devices and the formulation of split point adjustment strategies; using standardized measurement units can more accurately reflect the actual resource status of the device and improve the accuracy and effectiveness of resource evaluation; standardized resource evaluation data facilitates the implementation and optimization of the split point adjustment algorithm, such as quantitative comparison, threshold judgment and other operations.
[0071] Preferably, in step S3, the dynamically adjusting the segmentation points of the deep learning model includes:
[0072] Traverse different combinations of model split points;
[0073] For each combination of split points, evaluate the corresponding system performance indicators and resource consumption;
[0074] Based on the evaluation results, the optimal combination of split points is selected as the dynamically adjusted split points.
[0075] In the above scheme, the specific method of dynamically adjusting the split point is the traversal evaluation method; this method traverses different combinations of model split points, evaluates each combination, and finally selects the optimal split point combination as the dynamically adjusted split point; this traversal evaluation method can systematically search for the optimal split point to ensure the effectiveness of the split point adjustment. In the dynamic split point adjustment stage, the system first generates different combinations of model split points based on the preset split point range; then, for each split point combination, the system evaluates the performance indicators and resource consumption. After the evaluation is completed, the system compares the evaluation results of different split point combinations and selects the optimal split point combination as the dynamically adjusted split point.
[0076] By traversing different combinations of model split points, all possible splitting schemes can be systematically searched to avoid missing the optimal split point. By evaluating the performance indicators and resource consumption of each split point combination and selecting the optimal split point combination, it can ensure that the dynamically adjusted split points are relatively optimal, thereby improving system performance. The traversal evaluation method can be executed automatically without human intervention, reducing labor costs and improving the efficiency of split point adjustment.
[0077] Preferably, the system performance indicators include communication costs and training delays; the resource consumption includes computing power requirements and memory capacity requirements of terminal-side devices, edge-side devices, and cloud-side devices.
[0078] In the above scheme, during the evaluation phase, the system needs to calculate the communication cost, training latency, computing power requirements and memory capacity requirements of the end-side device, edge-side device and cloud-side device under each split point combination. These indicator data will be used for subsequent optimal split point selection decisions. For example, the system will tend to choose a split point combination with low communication cost and training latency, and resource requirements within the device capability range.
[0079] Clarifying the specific contents of system performance indicators and resource consumption makes the evaluation process more specific and refined, and can more accurately evaluate the advantages and disadvantages of the combination of split points; comprehensively considering system performance indicators and resource consumption indicators can more comprehensively evaluate the advantages and disadvantages of the combination of split points and avoid one-sided pursuit of a single indicator while ignoring other important factors; based on a more comprehensive evaluation system, better split point selection can be made, so that the final selected split point combination can better balance system performance and resource consumption.
[0080] Preferably, the evaluation of the corresponding system performance indicators and resource consumption is performed by calculating the total score of each combination of split points;
[0081] The total score is calculated as follows:
[0082] score=(α·score c +β·score T )·θ
[0083] Among them, score c Score for communication cost; score T is the training delay score; θ is the penalty factor; α and β are the weights of the communication cost score and the training delay score, and α+β=1. The weight value is set according to the emphasis on communication cost and inference delay in the actual application scenario.
[0084] In the above scheme, the specific method for evaluating system performance indicators and resource consumption is clarified, that is, calculating the total score; the calculation formula of the total score performs a weighted summation of the communication cost score and the training delay score, and introduces a penalty factor, comprehensively considering the system performance and resource constraints; this total score calculation method can integrate multi-dimensional evaluation indicators into a scalar value, which is convenient for comparison and selection of split point combinations.
[0085] In the evaluation phase, the system first calculates the communication cost and training delay of each split point combination, and converts them into communication cost score and training delay score; then, the two scores are weighted and summed according to the preset weight sum; if the split point combination violates the resource constraint, a penalty factor is introduced to reduce the total score; finally, the system sorts different split point combinations according to the total score and selects the combination with the highest score as the optimal split point. The total score calculation method integrates multi-dimensional evaluation indicators into a scalar value, realizes the quantitative evaluation of the advantages and disadvantages of the split point combination, and facilitates automated comparison and selection; the weight sum can flexibly configure the emphasis on communication cost and training delay according to the actual application scenario, making the evaluation method more adaptable; the introduction of the penalty factor can effectively handle the situation of resource overload and avoid selecting a split point combination that exceeds the device resource capacity; based on the sorting and selection of the total score, the optimal split point combination can be found efficiently, improving the efficiency of dynamic split point adjustment.
[0086] Preferably, the calculation formula of the communication cost score is:
[0087]
[0088] in, is the minimum communication cost, is the maximum communication cost, C cost is the total communication cost under the current combination of split points;
[0089] The calculation formula of the training delay score is:
[0090]
[0091] in, is the minimum training delay, is the maximum training delay, T delay It is the total training delay under the current split point combination.
[0092] In the above scheme, the specific calculation formulas for the communication cost score and the training delay score are clarified; the linear mapping method is used to map the actual communication cost and training delay values to the score range of [0,1]; this linear mapping method can unify indicators of different dimensions into the same dimension, which is convenient for weighted summation; the setting of the minimum and maximum values defines the range of scores, making the scores comparable.
[0093] In the stage of calculating the communication cost score and the training delay score, it is first necessary to obtain the preset minimum and maximum values of the communication cost and the minimum and maximum values of the training delay; then, calculate the total communication cost and the total training delay under the current combination of split points; finally, according to the provided formula, map the total communication cost and the total training delay to the communication cost score and the training delay score respectively. Providing a specific score calculation formula makes the score calculation process clearer and more operable, and facilitates the automatic execution of the system; the linear mapping method unifies the indicators of different dimensions into the score interval of [0,1], making the scores comparable and facilitating the weighted summation of different indicators; the linear mapping method can better reflect the linear relationship between the indicator value and the score, and improve the accuracy and effectiveness of the score calculation.
[0094] Preferably, the penalty factor is initialized to 1, and when the following situations occur, the penalty on the total score becomes 0.01;
[0095] The computing power required by the end device is greater than the computing power that the end device can accept.
[0096] The memory capacity required by the end device is greater than the available memory capacity of the end device.
[0097] The computing power required by edge devices > the computing power acceptable to edge devices;
[0098] The memory capacity required by the edge device is greater than the available memory capacity of the edge device.
[0099] The computing power required by the cloud device > the computing power that the cloud device can accept;
[0100] The memory capacity required by the cloud device > the available memory capacity of the cloud device.
[0101] In the above scheme, when calculating the total score, we first determine whether the current segmentation point combination meets the resource constraints, that is, whether the computing power requirements and memory capacity requirements of the terminal, edge and cloud devices are within their acceptable capabilities and available capacities; if the resource constraints are met, the penalty factor remains at 1; if any resource constraint is violated, the penalty factor is set to 0.01; then, the system uses the corresponding penalty factor to calculate the total score.
[0102] The penalty mechanism can effectively avoid selecting a combination of split points that exceeds the device resource capacity, ensuring the feasibility of the model deployment plan; resource excess may cause device overload, system crash or a sharp drop in performance, and the penalty mechanism can effectively avoid this situation; ensuring that the model deployment plan is within the device resource capacity can improve the stability and reliability of the system.
[0103] Preferably, in step S4, the distributed training adopts a federated learning method, including:
[0104] Perform local model training on the client device.
[0105] Upload the trained model parameters to the cloud;
[0106] Aggregate model parameters in the cloud;
[0107] Distribute the aggregated model parameters to the end-side devices.
[0108] In the above solution, federated learning is a distributed machine learning framework that protects data privacy. Its core idea is to perform model training on local devices and only upload model parameters or gradient information to the cloud for aggregation, while keeping the original data locally. In the distributed training phase, federated learning is used for model training; each end-side device uses local data to perform local model training and uploads the trained model parameters to the cloud; the cloud device receives model parameters from multiple end-side devices and performs parameter aggregation, such as average aggregation; the aggregated model parameters are distributed back to the end-side devices for the next round of local model training; this process is iterated until the model converges.
[0109] The federated learning method retains the original data on the end-side device and only uploads model parameters or gradient information, effectively protecting user data privacy and complying with the trend of data security and privacy protection. Federated learning can make full use of local data on the end-side device for model training, improving the generalization ability and performance of the model. Federated learning only transmits model parameters or gradient information, which can significantly reduce data transmission overhead compared to traditional data-centralized training. The federated learning framework can support distributed training of large-scale end-side devices and is suitable for end-edge-cloud collaborative computing scenarios.
[0110] In an embodiment provided in the present application, there are m different end-side devices, which are respectively denoted as clients C1, C2, ..., C m For each client, its resource situation can be described by several indicators:
[0111] The client can accept computing power C comp ,use Represents client C k (k=1,2,...,m) acceptable computing power, measured in floating point operations per second (FLOPs);
[0112] Client available memory capacity C mem ,use Represents client C k (k=1,2,...,m) available memory capacity in bytes.
[0113] There are n different edge devices, denoted as edge sides E1, E2, ..., E n . For each edge side, its resource situation can be described by several indicators:
[0114] Edge-side acceptable computing power E comp , denoted by the acceptable computing power of client E k (k = 1, 2, ..., n), measured in floating-point operations per second (FLOPs);
[0115] Edge-side available memory capacity E mem , denoted by the available memory capacity of client C k (k = 1, 2, ..., n), in bytes (Byte).
[0116] The cloud-acceptable computing power is denoted by Cloud comp , and the available memory capacity is denoted by Cloud mem .
[0117] Suppose the model M = {M1, M2, ..., M L} with L layers is sliced into three parts according to the slicing points x and y, namely the front part M 1:x , the middle part M x:y , and the rear part M y:L , which are respectively deployed to the edge side, the edge side, and the cloud side, where 1 < x < y < L. By traversing the slicing points x and y and calculating the scores in each case, the optimal slicing points x and y are determined. Start traversing from x = 2 and end at x = L - 2. For each x value, start traversing from y = x + 1 and end at y = L - 1. For each combination of x and y, the following evaluation is performed.
[0118] For the current slicing points x and y, the performance indicators of each part of the model are as follows:
[0119] Model required computing power M comp , and the required computing power of each part is expressed as: front part middle part rear part
[0120] Model required memory capacity M mem , and the required computing power of each part is expressed as: front part middle part rear part
[0121] During the evaluation, the communication cost and training delay during the training process are mainly considered;
[0122] The communication cost is regarded as the sum of the communication costs between the end and the edge, and between the edge and the cloud under the premise of one training. The communication between the end and the edge is closely related to x, and is represented by C x-cost To represent the communication cost between a certain end and a certain edge, the total communication cost between m different end-side devices and the edge side is C ce-cost =m·C x-cost Similarly, using C y-cost To represent the communication cost between the edge and the cloud, the total communication cost between n different edge devices and the cloud is C ec-cost =n·C y-cost ; In summary, the total communication cost is C cost =C ce-cost +C ec-cost .
[0123] The communication cost score is calculated by mapping the communication cost to a specific scoring zone, and the minimum communication cost is set to The maximum value is The communication cost score calculation formula is:
[0124]
[0125] The training delay is regarded as the sum of the inference delays of each part of the model on the device, edge, and cloud under the premise of training once. client-delay , T edge-delay and T cloud-delay To represent the training delay on the client side, edge side, and cloud side. The total training delay is: T delay =T client-delay +T edge-delay +T cloud-delay .
[0126] The training delay score is calculated by mapping the training delay to a specific scoring zone, assuming that the minimum training delay is The maximum value is The training delay score calculation formula is:
[0127]
[0128] Taking the above factors into consideration, the total score is calculated in a weighted manner. The weight of the communication cost score is α, the weight of the training delay score is β, and α+β=1. The weight value can be set according to the emphasis on communication cost and inference delay in the actual application scenario. The calculation formula is as follows:
[0129] score=(α·score c +β·score T )·θ
[0130] Among them, θ is the penalty factor, which is initialized to 1. The penalty state is a very small number, such as 0.001. When the following situations occur, the score is penalized.
[0131] On the client side: or
[0132] Edge side: or
[0133] Cloud: or
[0134] Finally, the best split points x and y that best match the resource environment are selected based on the traversal.
[0135] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for segmenting a device-edge-cloud model based on dynamic adjustment, characterized in that: The following steps are involved: S1: Split the deep learning model into at least two parts to be deployed on the end-side device and the cloud-side device respectively; S2: Real-time evaluation of the resource status of the end-side devices and cloud-side devices; S3: Based on the resource situation assessment results, dynamically adjust the split point of the deep learning model to adapt to the resource situation of the terminal side device and the cloud side device; S4: Based on the model after adjusting the split point, distributed training is performed in collaboration between the end-side devices and the cloud-side devices to update the deep learning model parameters.
2. According to the method for segmenting the edge-cloud model based on dynamic adjustment in claim 1, it is characterized in that: In step S1, the cloud-side device includes an edge-side device and a cloud-side device; the deep learning model is divided into three parts: the front part, the middle part and the back part; Among them, the front part is deployed on the terminal side device, the middle part is deployed on the edge side device, and the back part is deployed on the cloud side device.
3. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 2, characterized in that: In step S2, the resource conditions of the terminal device and the cloud device include the computing power, memory capacity and network bandwidth of the terminal device, the edge device and the cloud device; wherein, Computing capability is used to characterize the computing processing capability of a device; Memory capacity is used to characterize the available storage space of a device; Network bandwidth is used to characterize the data transmission capacity between devices.
4. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 3, characterized in that: The computing power is measured in floating point operations per second; the memory capacity is measured in bytes.
5. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 4, characterized in that: In step S3, the dynamically adjusting the segmentation points of the deep learning model includes: Traverse different combinations of model split points; For each combination of split points, evaluate the corresponding system performance indicators and resource consumption; Based on the evaluation results, the optimal combination of split points is selected as the dynamically adjusted split points.
6. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 5, characterized in that: The system performance indicators include communication costs and training delays; the resource consumption includes computing power requirements and memory capacity requirements of terminal devices, edge devices, and cloud devices.
7. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 6, characterized in that: The evaluation of the corresponding system performance indicators and resource consumption is performed by calculating the total score of each combination of split points; The total score is calculated as follows: score=(α·score c +b·score T )·θ Among them, score c Score for communication cost; score T is the training delay score; θ is the penalty factor; α and β are the weights of the communication cost score and the training delay score, and α+β=1. The weight value is set according to the emphasis on communication cost and inference delay in the actual application scenario.
8. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 7, characterized in that: The calculation formula of the communication cost score is: in, is the minimum communication cost, is the maximum communication cost, C cost is the total communication cost under the current combination of split points; The calculation formula of the training delay score is: in, is the minimum training delay, is the maximum training delay, T delay It is the total training delay under the current split point combination.
9. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 8, characterized in that: The penalty factor is initialized to 1. When the following situations occur, the penalty on the total score becomes 0.01; The computing power required by the end device is greater than the computing power that the end device can accept. The memory capacity required by the end device is greater than the available memory capacity of the end device. The computing power required by edge devices > the computing power acceptable to edge devices; The memory capacity required by the edge device is greater than the available memory capacity of the edge device. The computing power required by the cloud device > the computing power that the cloud device can accept; The memory capacity required by the cloud device > the available memory capacity of the cloud device.
10. The method for segmenting a terminal-edge-cloud model based on dynamic adjustment according to claim 9, characterized in that: In step S4, the distributed training adopts a federated learning method, including: Perform local model training on the client device. Upload the trained model parameters to the cloud; Aggregate model parameters in the cloud; Distribute the aggregated model parameters to the end-side devices.