Elastic controllable neural network unloading method based on intelligent connection network system
By adopting a flexible and controllable neural network offload method in the intelligent network system, the neural network model of the "end-edge-cloud" multi-computing node is split and deployed, and the problem of excessive pressure on the "cloud" side and immutable strategy in the traditional architecture is solved, and adaptability and efficiency to complex tasks and environments are achieved.
Patent Information
- Application Number
- CN202510229030.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
AI Technical Summary
When the traditional "end-cloud" two-layer architecture uses multiple SC tasks to share a "cloud" computing power node, the computing tasks and communication pressure of the "cloud" side computing power node is too high, and the neural network cannot change the execution strategy based on floating computing power, storage and network conditions.
The elastic controllable neural network offload method based on the intelligent network system is adopted. By storing the computing power, network and storage information of the multi-computing power node in Redis, the conditions of each node are evaluated, and the multi-output head DNN or CNN model is split according to the "computing-network-storage" conditions, the logical structure of the sub-neural network of the "end-side, edge, and cloud" side is obtained. Then centralized training is performed on a single GPU server and deployed to each node separately.
This method can handle complex multi-output head DNN and CNN tasks, adapt to different networks and lower-end hardware operation environments, reduce the pressure on the computing power nodes on the 'cloud' side, and dynamically adjust the execution strategy according to floating conditions.
Smart Images

Figure CN120128583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of the Internet of Things and distributed artificial intelligence, and specifically to a method for elastic and controllable neural network offloading based on an intelligent connection network system. Background Art
[0002] In the process of the cross-integration development of mobile communication and artificial intelligence, a new basic research platform and innovative ecological application have emerged - the Networking Systems of AI (NSAI). NSAI transforms traditional centralized AI into a distributed large-scale intelligent connection network body, realizing the self-organizing real-time intelligent evolution of the AI network. It is a major innovation in the integration of computing and communication disciplines and is oriented towards various traditional disciplines and industries, thus realizing the new ecosystem of the intelligent connection network system of "the network is the AI system, and AI is the network system". NSAI is the product of the deep integration of AI technology and communication networks and also represents a new direction based on intelligent connection services. In complex scenarios such as smartphone intelligent applications and autonomous driving vehicles, the cooperation mode of "terminal + 5G network + edge cloud + cloud service" will trigger a new round of industrial transformation and upgrading. The application of these technologies not only provides higher privacy protection for users but also makes services intelligent, efficient, and convenient. At the same time, AI deployed on edge clouds and terminal devices can better interact with humans and the real world in real time. In terms of AI algorithms, NSAI using a dual learning mechanism combining offline and online will better improve the adaptability of the next generation of AI and take a solid step towards general strong AI through online evolutionary learning.
[0003] Mobile devices such as smartphones and autonomous driving vehicles are increasingly relying on deep neural networks (DNNs) to perform complex inference tasks such as image classification and speech recognition. However, continuously executing the entire DNN on mobile devices may quickly drain the battery. Although offloading tasks to cloud / edge servers may reduce the computational burden on mobile devices, unstable patterns of channel quality, network, and edge server load may lead to severe delays in task execution. In recent years, methods based on split computing (abbreviated as "split", SC) have been proposed, splitting the deep neural network into "terminal" and "cloud" side models, which are executed on mobile devices and edge servers respectively. Ultimately, this may reduce bandwidth usage and energy consumption. Another method, called early exit (abbreviated as "early exit", EE), trains the model to embed multiple "exits" earlier in the architecture, each providing increasingly higher target accuracy. Therefore, the trade-off between accuracy and delay can be adjusted according to current conditions or application requirements.
[0004] Currently, there are already some SC methods for completing computational offloading between the "edge-cloud" dual-computing power node network architectures. The SC method based on the "edge-cloud" dual-computing power node is generally one of the traditional neural network offloading methods. However, this method has certain limitations and defects. First of all, the main disadvantage of this method is that it is difficult to handle complex tasks and adapt to different networks and lower-end hardware operating environments. When multiple SC tasks share a "cloud" computing power node, the computing tasks and communication pressure on the "cloud" side are too high. Secondly, traditional SC methods also cannot adapt to unknown situations and changes. Because once the neural network is running between the "edge-cloud" dual-computing power nodes, the execution strategy cannot be changed according to the floating computing power, storage, and network conditions. This limits the application scope and adaptability of neural network offloading based on the SC method in practical applications. Although the EE-based method can better adapt to lower-end hardware operating environments and change the execution strategy under the floating computing power, storage, and network conditions of the device, it only considers pure edge-side devices to execute the neural network, and the computing power is severely limited, resulting in the inference result not being able to achieve very good accuracy.
[0005] Compared with the SC or EE-based methods, the neural network offloading method based on the combination of SC and EE is more general and adaptable, can handle complex tasks and environments, and has better generalization ability. However, the neural network offloading method based on elasticity and controllability still faces some challenges. First of all, there are still some challenges in how to accurately split the model on the "edge" side, "edge" side, and "cloud" side of the computing power node in the "edge-edge-cloud" multi-computing power node network architecture when the neural network runs for the first time, and make full use of the computing power, network, and storage conditions of the "edge-edge-cloud" multi-computing power node. Secondly, the method based on the combination of SC and EE needs to process a large number of intermediate results of the neural network (generally, it is necessary to frequently store and forward intermediate results), which will increase the computational complexity of the algorithm. Therefore, this invention patent proposes an elasticity and controllable neural network offloading method based on an intelligent network system for this situation. Summary of the Invention
[0006] (I) Technical problems to be solved Aiming at the deficiencies of the prior art, the present invention provides an elasticity and controllable neural network offloading method based on an intelligent network system, which solves the problems that when the traditional "edge-cloud" two-layer architecture faces multiple SC tasks sharing a "cloud" computing power node, the computing tasks and communication pressure on the "cloud" side computing power node are too high, and after the neural network is running between the "edge-cloud" dual-computing power nodes, the execution strategy cannot be changed according to the floating computing power, storage, and network conditions.
[0007] (II) Technical solutions To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for elastic and controllable neural network offloading based on an intelligent connection network system, and the specific operation steps are as follows: Step 1: Obtain the computing power, network, and storage of the "terminal-edge-cloud" multi-computing power nodes according to the prefabricated computing power device information table, store this information in Redis for calling, and then extract and evaluate the computing power device information of the "terminal-edge-cloud" from Redis to evaluate the computing power, network, and storage conditions of the "terminal-edge-cloud" multi-computing power nodes; Step 2: Split the multi-output head DNN or multi-output head CNN model according to the "computing-network-storage" conditions to obtain the sub-neural network logical structures on the "terminal", "edge", and "cloud" sides; Among them, for the hierarchical segmentation training of the DNN or CNN model splitting, the specific steps include: 2.1 Model initialization and data preparation: First, split the complete deep learning model into three sub-models on the "terminal", "edge", and "cloud" sides, and deploy them on the corresponding nodes respectively; 2.2 Hierarchical training strategy: Find the optimal splitting configuration through 2.1.5 Perform model splitting training, and the training process follows the hierarchical logic; 2.3 Optimal model saving: During the training process, the system continuously evaluates the performance of each sub-model; when the current loss is lower than the historical best value update the best record and save the corresponding weight file; At the same time, for the sub-neural network logical structures obtained by splitting, a chained client distributed inference method based on multiple terminals is used to realize the task collaboration of distributed nodes. The specific steps include: decomposing the task into multiple stages, and each stage is executed by the corresponding sub-model; after each layer of node inference, the system calculates the inference factor to evaluate whether the task needs to be transferred; after the inference is completed, the system records the total inference time , accuracy and the final loss to provide a basis for performance analysis; Step 3: Conduct centralized training on the sub-models on the "terminal", "edge", and "cloud" sides on a single GPU server to obtain the sub-neural network models on the "terminal", "edge", and "cloud" sides; Step 4: Deploy the sub-neural network models on the "terminal", "edge", and "cloud" sides into the "terminal", "edge", and "cloud" computing power devices respectively.
[0008] Preferably, the specific steps of Step 4 include: 4.1. When performing task inference using a multi-output head DNN or multi-output head CNN model, after the sub-model inference on the "terminal" side, calculate the "terminal" side inference factor. When the "terminal" side inference factor is less than the "terminal" inference threshold (at this time, the computing power node resources on the "edge" side and "cloud" side are overly occupied by other applications), use the "terminal" side output head (layer) to obtain the task inference result. When the "terminal" side inference factor is greater than the "terminal" inference threshold, transmit the "terminal" side inference intermediate result to the "edge" side sub-model through the Socket network for continued inference; 4.2. After the sub-model inference on the "edge" side, calculate the "edge" side inference factor. When the "edge" side inference factor is less than the "edge" inference threshold (at this time, the computing power node resources on the "cloud" side are overly occupied by other applications), use the "edge" side output head (layer) to obtain the task inference result. When the "edge" side inference factor is greater than the "edge" inference threshold, transmit the "edge" side inference intermediate result to the "cloud" side sub-model through the Socket network for continued inference; 4.3. After the sub-model inference on the "cloud" side, use the "cloud" side output head (layer) to obtain the task inference result.
[0009] Preferably, the structure of each sub-model in step 2.1 is optimized according to the "computing-network-storage" conditions: 2.1.1. Retrieve from Redis 、 、 , which respectively represent the CPU or GPU utilization rates on the terminal, edge, and cloud sides. To eliminate the influence of different metric dimensions, the algorithm normalizes the resource utilization rates and total inference times on the terminal, edge, and cloud sides: , Normalize 、 、 to ensure that the value ranges of all parameters are within [0, 1]; 2.1.2. Calculate the sum of the normalized resource usages: ; 2.1.3. Assign weights to the resource utilization rates on different sides and calculate the weighted score: , where , , , reflecting the priority weights on the terminal, edge, and cloud sides; 2.1.4. Combine the weighted resource score with the total inference time to obtain the final score: , Among them, $\alpha = 0.5$, representing the weighted balance between resource usage and inference time; 2.1.5. Select the optimal configuration from the scoring results: , Among them, is the number of the model structure or configuration, used to record different model partitioning strategies; at the same time, verify whether the specific resource threshold condition is met and ; if the condition is met, retain the current optimal index ; otherwise, fallback to the alternative index .
[0010] Preferably, the hierarchical training strategy in step 2.2 finds the optimal splitting configuration through step 2.1.6 for model splitting training, and the training process follows a hierarchical logic; in each training iteration, data gradually passes through the "terminal", "edge", and "cloud" sub-models from the input end; at the same time, the gradients of each layer of sub-models are calculated separately and the weights are updated synchronously; the system records the total loss and average loss of each round of training and average loss .
[0011] (III) Beneficial effects The present invention provides a method for elastic and controllable neural network offloading based on an intelligent network system. It has the following beneficial effects: The present invention provides a method for elastic and controllable neural network offloading based on an intelligent network system. The method proposed by the present invention can handle complex multi-output head (layer) DNN and CNN tasks and adapt to different networks and lower-end hardware operating environments; when multiple SC tasks share one "cloud", the computing tasks and communication pressure of the computing nodes on the "cloud" side can be relatively dispersed by the computing nodes on the "edge" side; after the neural network is in operation among multiple computing nodes of "terminal-edge-cloud", the execution strategy can be changed according to the floating computing power, storage, and network conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a schematic diagram of the architecture of the method for elastic and controllable neural network offloading based on an intelligent network system proposed by the present invention; Figure 2 is a schematic diagram of the implementation process of the elastic and controllable neural network offloading combining splitting and early exit proposed by the present invention; Figure 3 is a schematic diagram of an example of splitting a DNN multi-output head neural network proposed by the present invention; Figure 4 is a schematic diagram of an example of splitting a CNN multi-output layer neural network proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] In the present invention, the problems of splitting and elastic offloading of neural networks in a multi-computing power node decentralized network congestion architecture based on the intelligent connection network system are considered. Most of the existing neural network deployment methods for mobile terminals adopt pure mobile device deployment with model lightweighting or simply split the neural network model for dual-computing power node "terminal-cloud" calculation and then jointly perform inference calculation with the "terminal" and "cloud" sub-models respectively. However, model lightweighting is bound to cause a serious decline in model accuracy, and simply splitting the neural network model for "terminal-cloud" cannot adapt to the dynamically changing computing power conditions of "terminal" and "cloud" devices. To solve this problem, the present invention adopts an elastic and controllable neural network offloading method combining splitting and early exit. The goal is to reduce the problems that when the traditional "terminal-cloud" two-layer architecture faces multiple SC tasks sharing one "cloud" computing power node, the computing tasks and communication pressure of the "cloud" side computing power node are too large, and after the neural network is in operation between the "terminal-cloud" dual-computing power nodes, it is unable to change the execution strategy according to the floating computing power, storage, and network conditions; This patent takes the terminal-edge-cloud scenario as an example to describe, but is not limited to the splitting under terminal-edge-cloud (splitting under 3 computing power devices). It can also be splitting under more N devices (for example, splitting the neural network model into 10 segments, deploying on 10 devices, jointly performing inference, or even splitting into more segments). The following is the specific implementation method taking the terminal-edge-cloud scenario as an example: Embodiment: As Figures 1-4 , an elastic and controllable neural network offloading method based on the intelligent connection network system is provided in an embodiment of the present invention, and the specific operation steps are as follows: Step 1: Obtain the computing power, network, and storage of the "terminal-edge-cloud" multi-computing power nodes according to the computing power device information table provided in Table 1, store this information in Redis for calling, and then extract and evaluate the "terminal-edge-cloud" computing power device information from Redis to evaluate the computing power, network, and storage conditions of the "terminal-edge-cloud" multi-computing power nodes; Table 1 Computing power device information table The system first obtains the computing, network, and storage resource information of the "device-edge-cloud" multi-computing power nodes. This information is collected through the device performance monitoring interface (such as the computing power device information table shown in Table 1) and stored in the Redis database for subsequent resource scheduling and model splitting decisions. Before each task execution, the system extracts the latest evaluation data from Redis, analyzes the computing power, network bandwidth, and storage conditions of each layer of nodes to ensure the accuracy of the resource status; Step 2: Split the multi-output head DNN or multi-output head CNN model according to the "computing-network-storage" conditions to obtain the sub-neural network logical structures on the "device", "edge", and "cloud" sides; Among them, for model splitting training: The model training adopts a hierarchical splitting training method, and the specific steps are as follows: 2.1. Model initialization and data preparation: The hierarchical splitting training method splits the complete deep learning model into three sub-models on the "device", "edge", and "cloud" sides, which are respectively deployed on the corresponding nodes. Taking the DNN or CNN multi-output head model as an example (as shown in Figure 3 and Figure 4 ), the structure of each sub-model is optimized according to the "computing-network-storage" conditions: 2.1.1. Retrieve from Redis 、 、 , which respectively represent the CPU or GPU utilization rates on the device, edge, and cloud sides. To eliminate the influence of different metric dimensions, the algorithm normalizes the resource utilization rates and total inference times on the device, edge, and cloud sides: , Similarly, normalize 、 、 to ensure that the value ranges of all parameters are [0, 1].
[0015] 2.1.2. Calculate the sum of the normalized resource usages: .
[0016] 2.1.3. Assign weights to the resource utilization rates on different sides and calculate the weighted scores: , Among them, , , , reflecting the priority weights on the device, edge, and cloud sides.
[0017] 2.1.4. Combine the weighted resource scores with the total inference time to obtain the final score: , Among them, $\alpha = 0.5$, representing the weight balance between resource usage and inference time.
[0018] 2.1.5. Select the optimal configuration from the scoring results: Among them, is the number of the model structure or configuration, used to record different model partitioning strategies. At the same time, verify whether the specific resource threshold condition is met and . If the condition is met, retain the current optimal index ; otherwise, fallback to the alternative index .
[0019] The system selects a unified training set and validation set , and uses standard loss functions (such as cross-entropy or DiceLoss) and optimizers.
[0020] 2.2. Hierarchical training strategy: Find the optimal splitting configuration through 2.1.5 Perform model splitting training, and the training process follows a hierarchical logic. In each training iteration, the data gradually passes through the "edge", "side", and "cloud" sub-models from the input end: 2.2.1. The data input passes through the "edge" side sub-model Complete the calculations of the first few layers and generate intermediate results .
[0021] 2.2.2. The intermediate results are passed through a pipeline to the "side" side sub-model for further inference and generate .
[0022] 2.2.3. Finally, the "cloud" side sub-model completes the full model calculation and outputs the results.
[0023] The gradients of each layer of the sub-model are calculated separately and the weights are updated synchronously to ensure the stability of training optimization. The system records the total loss and average loss of each round of training to judge the model performance and convergence situation.
[0024] 2.3. Optimal model saving: During the training process, the system continuously evaluates the performance of each sub-model. When the current loss is lower than the historical best value , update the best record and save the corresponding weight file for subsequent deployment.
[0025] For the sub - neural network logic structure obtained by splitting, the method of chain - type client distributed inference based on multiple terminals realizes the task collaboration of distributed nodes. The following are the detailed steps: A). Test data transmission and node inference: The chain - type client distributed inference method decomposes the task into multiple stages during the inference phase, and each stage is executed by the corresponding sub - model: A1). Test data First, it is transmitted to the "terminal" - side node. After inference, the intermediate result is output .
[0026] A2). If the inference factor on the "terminal" - side exceeds the threshold, the intermediate result is sent to the "edge" - side node through the Socket network and is further processed.
[0027] A3). Similarly, if the inference factor on the "edge" - side exceeds the threshold, the intermediate result will be uploaded to the "cloud" - side node to complete the final inference.
[0028] B). Dynamic judgment of the inference factor: After each layer of node inference, the system calculates the inference factor to evaluate whether the task needs to be transferred: B1). If the computing power of the current node is sufficient and the inference factor is lower than the threshold, the result is directly output.
[0029] B2). Otherwise, the intermediate result is passed to the next - layer node to ensure that the calculation is completed within the allowable range of resources.
[0030] C). Result summary and performance evaluation: After the inference is completed, the system records the total inference time , accuracy and the final loss to provide a basis for performance analysis.
[0031] Step 3: Conduct centralized training on the "terminal" - side, "edge" - side, and "cloud" - side sub - models on a single GPU server to obtain the "terminal" - side, "edge" - side, and "cloud" - side sub - neural network models; Step 4: Deploy the "terminal" - side, "edge" - side, and "cloud" - side sub - neural network models to the "terminal", "edge", and "cloud" computing power devices respectively: 4.1. When using Figure 3 and 4When performing task inference using a multi-output head DNN or multi-output head CNN model, after the "edge" side sub-model inference, calculate the "edge" side inference factor. When the "edge" side inference factor is less than the "edge" inference threshold (at this time, the computing power node resources on the "cloud" side are overly occupied by other applications), use the "edge" side output head (layer) to obtain the task inference result. When the "edge" side inference factor is greater than the "edge" inference threshold, transmit the "edge" side intermediate inference result to the "cloud" side sub-model through the Socket network for continued inference; 4.2. After the "edge" side sub-model inference, calculate the "edge" side inference factor. When the "edge" side inference factor is less than the "edge" inference threshold (at this time, the computing power node resources on the "cloud" side are overly occupied by other applications), use the "edge" side output head (layer) to obtain the task inference result. When the "edge" side inference factor is greater than the "edge" inference threshold, transmit the "edge" side intermediate inference result to the "cloud" side sub-model through the Socket network for continued inference; 4.3. After the "cloud" side sub-model inference, use the "cloud" side output head (layer) to obtain the task inference result.
[0032] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A flexible and controllable neural network unloading method based on an intelligent network system, characterized in that: The specific steps are: Step 1: Obtain the computing power, network, and storage of the "end-edge-cloud" multi-computing nodes based on the prefabricated computing power equipment information table and store them in Redis for calling. Then extract and evaluate the computing power, network, and storage conditions of the "end-edge-cloud" multi-computing nodes from Redis. Step 2: Split the multi-output head DNN or multi-output head CNN model according to the "computing-network-storage" conditions to obtain the logical structure of the sub-neural networks on the "end" side, "edge" side, and "cloud" side; Among them, for the layered segmentation training of DNN or CNN model splitting, the specific steps include: 2.
1. Model initialization and data preparation: First, split the complete deep learning model into three sub-models: "end", "edge" and "cloud", and deploy them on the corresponding nodes respectively; 2.
2. Layered training strategy: Find the optimal split configuration through 2.1.5 Perform model split training, and the training process follows the hierarchical logic; 2.3 Optimal Model Preservation: During the training process, the system continuously evaluates the performance of each sub-model; when the current loss Below historical best When , update the best record and save the corresponding weight file; At the same time, the logical structure of the split sub-neural network is based on multiple terminals to implement the distributed node task collaboration. The specific steps include: decomposing the task into multiple stages, each stage is executed by the corresponding sub-model; after each layer of node reasoning, the system calculates the reasoning factor to evaluate whether the task needs to be transferred; after the reasoning is completed, the system records the total reasoning time , Accuracy and final loss , providing a basis for performance analysis; Step 3: Centrally train the "end" side, "edge" side, and "cloud" side sub-models on a single GPU server to obtain the "end" side, "edge" side, and "cloud" side sub-neural network models; Step 4: Deploy the "end" side, "edge" side and "cloud" side sub-neural network models to the "end", "edge" and "cloud" computing devices respectively.
2. According to claim 1, a flexible and controllable neural network unloading method based on an intelligent network system is characterized in that: The specific steps of step 4 include: 4.
1. When using a multi-output head DNN or multi-output head CNN model for task reasoning, after the "end" side sub-model reasoning, calculate the "end" side reasoning factor. When the "end" side reasoning factor is less than the "end" reasoning threshold (at this time, the computing power node resources on the "edge" side and the "cloud" side are excessively occupied by other applications), use the "end" side output head (layer) to obtain the task reasoning result; when the "end" side reasoning factor is greater than the "end" reasoning threshold, transmit the "end" side reasoning intermediate result to the "edge" side sub-model through the Socket network to continue reasoning; 4.
2. After the "edge" side sub-model inference, the "edge" side inference factor is calculated. When the "edge" side inference factor is less than the "edge" inference threshold (at this time, the computing power node resources on the "cloud" side are excessively occupied by other applications), the "edge" side output head (layer) is used to obtain the task inference result; when the "edge" side inference factor is greater than the "edge" inference threshold, the "edge" side inference intermediate result is transmitted to the "cloud" side sub-model through the Socket network for further inference; 4.
3. After the sub-model inference on the "cloud" side, the task inference result is obtained using the output head (layer) on the "cloud" side.
3. The method for unloading a flexible and controllable neural network based on an intelligent network system according to claim 1, characterized in that: The structure of each sub-model in step 2.1 is optimized according to the "computing-network-storage" conditions: 2.1.1、Retrieve from Redis , , , which represent the CPU or GPU usage on the end, edge, and cloud sides respectively. To eliminate the impact of different indicator dimensions, the algorithm normalizes the resource usage and total inference time on the end, edge, and cloud sides: right , , Normalize to ensure that all parameters are in the range of [0, 1]; 2.1.
2. Calculate the sum of normalized resource usage: ; 2.1.
3. Assign weights to resource usage on different sides and calculate weighted scores: in, , , , reflects the priority weights of the three sides: end, edge, and cloud; 2.1.
4. Combine the weighted resource score with the total inference time to get the final score: Among them, $\alpha = 0.5$, indicating the weight balance between resource usage and inference time; 2.1.
5. Select the best configuration from the scoring results: , in, The model structure or configuration number is used to record different model partitioning strategies; at the same time, verify whether specific resource threshold conditions are met. and ; If the condition is met, keep the current optimal index ; Otherwise, fall back to the alternative index .
4. The method for unloading a flexible and controllable neural network based on an intelligent network system according to claim 1, characterized in that: The hierarchical training strategy in step 2.2 finds the optimal split configuration through step 2.1.6 Perform model split training, and the training process follows the hierarchical logic; In each training iteration, data is passed from the input end to the "end", "edge" and "cloud" sub-models step by step. At the same time, the gradient of each layer of the sub-model is calculated separately, and the weights are updated synchronously. The system records the total loss of each round of training. and the average loss .