Data-driven learning model combination optimization method based on edge cloud collaboration
By designing collaborative computing of cloud, edge and device components in complex industrial systems, combining microservice architecture and KubeEdge architecture, efficient deployment of distributed proxy-assisted evolution algorithms and multiple rounds of model combination strategies are realized, solving the problem of massive distributed data processing and computing latency, and significantly improving optimization efficiency and resource utilization.
Patent Information
- Application Number
- CN202510089612.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the process of optimizing complex industrial systems, how to effectively process massive distributed data, reduce computing delays, and improve optimization efficiency has become a key issue that needs to be solved urgently.
By designing collaborative computing of cloud, edge and device components, the systematization and efficient deployment of distributed proxy-assisted evolution algorithms are realized, and the microservice design architecture and container deployment solution are adopted, and the open source edge computing architecture KubeEdge is combined to build a collaborative computing framework and multiple rounds of model combination strategies to achieve efficient communication and resource scheduling between cloud and edge nodes.
It significantly improves the solution efficiency of complex distributed optimization problems, improves resource utilization efficiency and data processing intelligence, and solves the performance bottlenecks of traditional technologies in the problems of high computing costs and data transmission delays.
Smart Images

Figure CN120128577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data-driven optimization, and particularly relates to a method for optimizing a combination of data-driven learning models based on edge-cloud collaboration. Background Art
[0002] In recent years, with the rapid development of the Internet of Things and edge computing, data processing and optimization decision-making face many challenges and development opportunities. In the optimization process of complex industrial systems, how to effectively process massive distributed data, reduce computational latency, and improve optimization efficiency has become a key problem to be solved urgently. For example, the complex design systems in the automotive industry usually consist of many subsystems distributed in multiple locations, and these subsystems perform different simulation experiments, such as noise vibration intensity, computational fluid dynamics, and collision durability analysis, etc.; such complex problems usually have bottlenecks such as large computational volume, wide data distribution, and high evaluation costs, and it is difficult to effectively solve them with traditional optimization algorithms.
[0003] As a powerful means of solving complex optimization problems, surrogate-assisted evolutionary computation can significantly reduce the computational cost and improve the optimization efficiency by introducing surrogate models in the evolutionary process. However, traditional surrogate-assisted evolutionary computation methods still have certain limitations when facing practical challenges such as the explosion of data volume and limited computational resources. Therefore, how to train surrogate models in the collaborative framework of edge computing and cloud computing, and combine strategies such as model combination and integration to improve the performance of surrogate models, so as to effectively address complex optimization problems driven by distributed data, has become one of the current research hotspots. Summary of the Invention
[0004] In order to overcome the defects and deficiencies existing in the prior art, the present invention provides a method for optimizing a combination of data-driven learning models based on edge-cloud collaboration. In order to address the challenges of complex optimization problems in a distributed environment, the present invention optimizes the solution process of complex problems in a distributed environment through efficient resource utilization and intelligent data processing. Specifically, by designing the collaborative computing of cloud, edge, and device components, the systematic and efficient deployment of the distributed surrogate-assisted evolutionary algorithm is realized, and the microservice design architecture and container deployment scheme are adopted, thereby realizing the flexibility and scalability of surrogate model deployment. In addition, relying on the cluster built by the open-source edge computing architecture KubeEdge provides strong technical support for the interactive operation, real-time monitoring, and comprehensive management of the platform.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for optimizing a combination of data-driven learning models based on edge-cloud collaboration, including the following steps:
[0007] Build a collaborative computing framework based on cloud nodes, edge nodes, and device nodes, and allocate distributed tasks and computing resources;
[0008] Build multiple microservice interfaces, including a platform service interface, an evaluation service interface for edge nodes, a model training service interface for edge nodes, a model upgrade service interface for cloud nodes, and a candidate solution evaluation service interface for cloud nodes;
[0009] The platform service interface is provided with multiple functional modules, which decompose the process of the distributed agent-assisted evolutionary algorithm into multiple functional modules. The evaluation service interface evaluates the candidate solutions of the evaluation requests of cloud nodes or device nodes and stores the evaluation results in the edge nodes. The model training service interface trains the agent model based on the local data of the edge nodes. The model upgrade service interface obtains the model training results, combines and integrates the models, and distributes the updated agent model version to each edge node. The candidate solution evaluation service interface filters the candidate solutions and conducts real environment evaluations by calling the evaluation service interface of the edge nodes;
[0010] Build a multi-round model combination strategy and a model online enhancement combination training strategy based on ensemble learning. The multiple microservice interfaces are encapsulated into Docker images and saved in the cloud image repository for version management and distribution, and build a communication mechanism between cloud nodes and edge nodes.
[0011] As a preferred technical solution, the cloud node executes the agent-assisted evolutionary algorithm, monitors the states of edge nodes and device nodes, and dynamically adjusts the cloud resource distribution. The edge node obtains the data collected by the device node to perform local training and update of the agent model, executes or unloads computing tasks, and manages edge resources. The device node collects real-time data and manages the computing resources and communication resources of the device.
[0012] As a preferred technical solution, the cloud node adopts Kubernetes technology and uses its API to interact with the CloudCore component. The CloudCore component acts as a scheduling center and dynamically adjusts the cloud resource distribution based on the workload prediction algorithm.
[0013] As a preferred technical solution, the edge node realizes the distribution, execution of containerized tasks, and unified management of edge resources based on the EdgeCore component. The EdgeCore component communicates with the CloudCore component through a two-way encrypted channel to achieve real-time synchronization of tasks and resources.
[0014] As a preferred technical solution, the edge node pulls the Docker image through the KubeEdge framework and instantiates the EdgeCore component. The EdgeCore component starts service containers on the edge node, and uses the monitoring and management tools provided by KubeEdge to monitor the running status of container instances in real time. When a container encounters an exception or crashes, the EdgeCore component automatically restarts the container according to a preset policy for fault recovery.
[0015] As a preferred technical solution, construct a multi-round model combination strategy and an online model enhancement combination training strategy based on ensemble learning, specifically including:
[0016] Train multiple surrogate models based on the local dataset, aggregate the multiple surrogate models to form a global model, calculate the fitness of each optimization solution based on the global model, and update the local dataset through real environment evaluation;
[0017] Collect the surrogate models of multiple optimization rounds to obtain a cross-round model, and aggregate the cross-round model;
[0018] Perform sampling without replacement on the local dataset based on the weight distribution to obtain a training set, train a surrogate model based on the training set, evaluate each sample in the local dataset, calculate the relative error of the sample, and update the sample weight.
[0019] As a preferred technical solution, the local dataset includes the production history, raw material ratio, process parameters, equipment status, and production output of each workshop. Train the surrogate model according to the production optimization goal and evaluate the industrial production formula.
[0020] As a preferred technical solution, the calculation of the fitness of each optimization solution based on the global model specifically includes:
[0021] Randomly generate an initial population, where each individual represents a potential formula solution. Calculate the fitness of each optimization solution based on the global model. The individuals in the population are selected, crossed, and mutated to generate a new generation of formula solutions, and iteratively generate the optimal formula solution.
[0022] As a preferred technical solution, construct a multi-round model combination strategy and an online model enhancement combination training strategy based on ensemble learning, specifically including:
[0023] For each edge node k, the cloud node records the surrogate models collected in the previous m rounds, denoted as Select the current optimal surrogate model:
[0024]
[0025] Among them, the Median(·) operator returns the median of the input set, representing the selection of a stable model measured by the median of the metrics, and the metric of the model is calculated as follows:
[0026]
[0027] where represents the model metric;
[0028] Aggregate the cross-round models using the weights obtained by the boundary judgment strategy:
[0029]
[0030] where the weight ω k (x) is calculated according to the data interval where the data set falls, and lb k and ub k represent the upper and lower limits of the statistical data in the edge node k respectively, and Z t is the normalization factor;
[0031] The data set is expressed as: n k is the total number of samples in the edge node k. Perform sampling without replacement on the data set to obtain the training set
[0032] Based on the training set train the surrogate model
[0033] Use the surrogate model to evaluate each sample in the data set and compare the result with the true label value to calculate the relative error of the sample
[0034]
[0035] where is the maximum absolute error, and the model error rate is calculated as:
[0036]
[0037] The model metric is then defined as:
[0038]
[0039] The weight of each sample in the data set is updated using the following formula:
[0040]
[0041] Among them, is a normalization factor.
[0042] As a preferred technical solution, a communication mechanism between cloud nodes and edge nodes is constructed, specifically including:
[0043] The communication between cloud nodes and edge nodes is implemented through the KubeEdge framework, and a control channel is established between the CloudCore component and the EdgeCore component. The control channel uses the TLS encryption protocol for data transmission;
[0044] Based on the periodic heartbeat mechanism, the CloudCore component monitors the status of the EdgeCore component. If an edge node has an abnormality or drops offline, a preset disaster tolerance strategy is triggered to migrate tasks to a standby node;
[0045] Stream data transmission and chunked concurrent transmission technologies are used for data transmission, and two-way real-time communication is achieved through the WebSocket protocol.
[0046] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0047] (1) Through the collaborative work of the cloud and edge computing, the present invention significantly improves the solution efficiency of complex distributed optimization problems, improves the resource utilization efficiency and data processing intelligence, provides strong technical support for the systematic and efficient deployment of distributed agent-assisted evolutionary algorithms, and realizes the efficient solution of complex optimization problems in a distributed environment.
[0048] (2) By introducing a multi-model combination and integration strategy and using historical data and surrogate models for assistance, the present invention optimizes the model training and prediction processes and solves the performance bottleneck encountered by the prior art in solving distributed optimization problems with high computational costs.
[0049] (3) By combining the powerful computing power of cloud computing with the low-latency characteristics of edge computing, the present invention improves the execution efficiency of optimization algorithms through dynamic resource scheduling and intelligent task allocation, and solves the problems of data transmission latency and uneven computing resources in a distributed environment in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic flow chart of the data-driven learning model combination optimization method based on edge-cloud collaboration of the present invention;
[0051] Figure 2 is a schematic implementation architecture diagram of the data-driven learning model combination optimization method based on edge-cloud collaboration of the present invention;
[0052] Figure 3 Schematic diagram of service module interface choreography for the present invention;
[0053] Figure 4 Schematic diagram of service deployment for the data-driven learning model combination optimization method based on edge-cloud collaboration of the present invention;
[0054] Figure 5 Schematic diagram of the cloud-edge communication mechanism for the data-driven learning model combination optimization method based on edge-cloud collaboration of the present invention. Specific embodiments
[0055] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] As Figure 1 shown, this embodiment provides a data-driven learning model combination optimization method based on edge-cloud collaboration, including the module microservice interface design management and distributed agent-assisted evolutionary algorithm deployment implementation strategy for model combination optimization, which is used to overcome the limitations of traditional centralized systems for distributed heterogeneous data-driven optimization and improve the efficiency and effectiveness of solving complex optimization problems; among them, the module microservice interface design management includes the modular design of the microservice architecture and the definition of service interfaces to ensure the efficient execution and flexible deployment of the algorithm; the distributed agent-assisted evolutionary algorithm deployment implementation strategy covers key steps such as algorithm initialization, evolution, selection, model training and evaluation, as well as the management of model adjustment based on evaluation results and the termination conditions of algorithm iteration, specifically including the following steps:
[0057] S1: Construct a three-layer collaborative computing framework based on "cloud-edge-device", and realize the efficient execution of distributed optimization tasks and the reasonable allocation of computing resources through the cooperation of cloud nodes, edge nodes and device nodes. This architecture aims to make full use of the powerful computing power of cloud computing, the low-latency characteristics of edge computing, and the distributed processing capabilities of device-side resources;
[0058] Among them, the cloud node is used to implement:
[0059] 1) The core process of agent-assisted evolutionary computing. The cloud uses high-performance computing resources to execute the surrogate-assisted evolutionary algorithm (SAEA), which includes the global search of optimization problems, the selection and combination of surrogate models, and the global scheduling of optimization solutions;
[0060] 2) System monitoring and management. Through the integrated control panel, the status of edge nodes and device nodes is monitored in real time, including resource usage, task execution progress, and data transmission stability.
[0061] 3) Resource allocation and scheduling. As Figure 2 shown, Kubernetes (K8s) technology is adopted to interact with the CloudCore core components using its API, realizing dynamic allocation of cloud resources and intelligent scheduling of tasks. The CloudCore component acts as a scheduling center and dynamically adjusts the resource distribution between the cloud and the edge based on the workload prediction algorithm.
[0062] Edge nodes are located at the network edge close to devices, and their core responsibilities include:
[0063] 1) Proxy model training. After collecting data from device nodes, edge nodes use the lightweight training framework TensorFlow Lite to locally train and update the proxy model to adapt to real-time data changes.
[0064] 2) Task offloading and execution. By analyzing the device load and data transmission latency, computationally intensive tasks are dynamically offloaded and optimized. The task offloading mechanism is based on multi-metric decision-making (including device processing capacity, bandwidth conditions, and latency constraints).
[0065] 3) Edge resource management. The EdgeCore component runs in edge nodes and is responsible for the distribution and execution of containerized tasks and the unified management of edge resources. EdgeCore communicates with CloudCore through a secure two-way encrypted channel (TLS protocol) to achieve real-time synchronization of tasks and resources.
[0066] Device nodes consist of Internet of Things devices (sensors, actuators, and embedded terminals) and are mainly responsible for:
[0067] 1) Data collection and transmission. Device nodes send the collected real-time data to edge nodes through the lightweight protocol MQTT.
[0068] 2) Task assistance execution. Device nodes assist edge nodes in executing local optimization tasks or evaluation tasks, such as calculating the fitness value of optimization candidate solutions or performing simple inference calculations.
[0069] 3) Resource integration and scheduling. The device controller is responsible for coordinating the computing resources and communication resources of the device to minimize latency and optimize bandwidth utilization.
[0070] S2: Microservice interface design. Implement each functional component of the distributed proxy-assisted evolutionary algorithm in a lightweight manner, deploy the service to the cloud and the edge in a stateless image manner, and the communication mechanism ensures real-time and reliable data transmission;
[0071] In this embodiment, multiple microservice interfaces are designed to achieve seamless collaboration among the components of the distributed agent-assisted evolutionary algorithm. These interfaces cover the evaluation and model training functions of edge nodes, as well as functions such as model upgrade and combined integration of cloud nodes. As Figure 3 shown, in this embodiment, the flexibility and scalability of the system are achieved through modular design and microservice orchestration;
[0072] S21: Platform service interface orchestration:
[0073] 1) Modular design algorithm process decomposition: The entire distributed agent-assisted evolutionary algorithm process is subdivided into multiple independent modules I = {I 1 , I 2 ,..., I n}, and each module is responsible for a single function, facilitating independent update and deployment between modules. The tasks of each module are as follows:
[0074] 1.1) Initialization module: Used to generate an initial population and provide configurable parameters such as population size and gene representation method.
[0075] 1.2) Evolution module: Executes the evolutionary operators in the algorithm, such as mutation and crossover operations. This module supports multiple evolutionary operator strategies and has the function of adjusting algorithm parameters, supporting dynamic strategy switching.
[0076] 1.3) Selection module: Uses methods such as roulette wheel selection and tournament selection to select excellent individuals from the current population, supporting strategies such as sorting by fitness and random selection;
[0077] 1.4) Candidate evaluation module: Reduces the computational overhead through a surrogate model and improves the evaluation efficiency.
[0078] 1.5) Real evaluation module: This module evaluates candidate solutions in a real environment through cloud-edge collaboration to ensure that the evaluation results have practical application significance.
[0079] 1.6) Model training module: Trains and optimizes the surrogate model using historical data, including data preprocessing, model training, parameter tuning, etc.
[0080] 2) Platform service interface definition: Each module exposes clear RESTful API interfaces for data exchange through the JSON protocol. The interface design follows the principles of high cohesion and low coupling to ensure the independence and scalability of system modules, and interface version management is implemented to facilitate updates as needed at any time.
[0081] S22: Define the evaluation service interface on the edge node;
[0082] Interface name: evaluation;
[0083] Function description: Perform the evaluation task of the high-cost model combination on the Internet of Things device;
[0084] Workflow:
[0085] 1) Receive the evaluation request R from the cloud node or device node e ;
[0086] 2) Conduct a real evaluation E(x) on the candidate solution "x" in the request;
[0087] 3) Store the evaluation result r = E(x) in the local database of the edge node;
[0088] 4) Return the evaluation result r to the requester.
[0089] S23: Define the model training service interface on the edge node;
[0090] Interface name: model_training;
[0091] Function description: Train the proxy model using the local data of the edge node and return the updated model parameters.
[0092] Workflow:
[0093] 1) Sample the required dataset from the historical database, supporting data preprocessing (such as normalization, feature selection).
[0094] 2) Use the sampled data to train the proxy model supporting the Adam and SGD algorithms, and using methods such as cross-validation to optimize hyperparameters during training.
[0095] 3) After training is completed, return the updated model parameters and store them in the local model library.
[0096] S24: Design the model upgrade service interface on the cloud node;
[0097] Interface name: model_upgrade;
[0098] Function description: Manage and distribute the new version of the proxy model to the edge nodes;
[0099] Workflow:
[0100] 1) Collect the model training results from the edge nodes;
[0101] 2) Aggregate and optimize the model parameters, and combine and integrate heterogeneous models;
[0102] 3) Distribute the updated proxy model version to each edge node;
[0103] S25: Design the candidate solution evaluation service interface on the cloud node of the design cloud;
[0104] Interface name: candidate_evaluation;
[0105] Function description: Screen candidate solutions and conduct real - environment evaluation by calling the evaluation interface of the edge node;
[0106] Workflow:
[0107] 1) Select potential candidate solutions C = {c 1 , c 2 ,..., c p} from the population;
[0108] 2) Send the candidate solution c i ∈ C to the evaluation interface of the edge node for evaluation, supporting parallel processing of batch evaluation tasks;
[0109] 3) Receive the evaluation result E(c i ), and adjust the individuals in the population according to the result. The evaluation result is fed back to other modules on the cloud node for subsequent operations such as model upgrade.
[0110] S3: Plan the specific implementation process of the distributed agent - assisted evolutionary algorithm, and construct a multi - round model combination strategy and an online enhanced model combination training strategy based on ensemble learning;
[0111] Taking the industrial production formula optimization problem as an example, industrial production formula optimization is a typical application of distributed data - driven optimization, aiming to improve production efficiency, reduce costs, and optimize the production quality of products. In this process, each production workshop, as an independent and autonomous entity, has its own production data, including raw material ratios, process parameters, equipment status, etc., and can build an agent model based on these local data to evaluate the advantages and disadvantages of different formulas. The cloud server is responsible for integrating the edge agent models from each workshop and continuously improving the production formula through a model - driven evolutionary optimization algorithm. At the same time, each workshop will give real - time feedback and evaluation on the optimization strategy given by the cloud to ensure that the optimization plan matches the actual production environment of the workshop, thus realizing a more accurate and efficient production process.
[0112] S31: Edge model training. Each edge client (edge k, k = 1, 2,..., K) uses the local dataset to train the agent model The model is then uploaded to the cloud.
[0113] In the formula optimization problem, one workshop corresponds to one edge client, and the local dataset contains information such as the production history, raw material ratio, process parameters, equipment status, and production output of each workshop. The surrogate model can predict the impact of different formulations on production efficiency, cost, and product quality. Each workshop trains the surrogate model based on the same optimization goal (such as maximizing product quality, minimizing raw material waste, etc.) to generate a model that can evaluate the advantages and disadvantages of different formulations. After training, the workshop uploads its model to the cloud for global optimization.
[0114] S32: Cloud model aggregation. The cloud aggregates the surrogate models from each edge client to form a global model This aggregation process can be achieved through different model combination operators, such as average, weighted average, etc.
[0115] Global evaluation model In the aggregation, the models of each workshop may have strong adaptability under certain specific production conditions. Therefore, the cloud can assign different weights to their models according to the data quality of the workshop or the importance of the workshop.
[0116] S33: Evolutionary algorithm execution. The initial population is randomly generated, and each individual represents a potential solution. The evolutionary algorithm iteratively evolves the population and uses the global model to predict the quality of individuals and guide the evolution of the population. Eventually, the individuals close to the optimal solution will be broadcast to each edge device.
[0117] The initial population consists of multiple randomly generated potential formulation solutions. Each formulation solution (individual) includes different raw material ratios, process parameters, etc. The evolutionary algorithm calculates the fitness of each individual with the assistance of the global model (predicting each formulation solution and evaluating its impact on production efficiency, cost, and quality). Through evolutionary strategies such as genetic algorithms, the excellent individuals in the population are selected, crossed, and mutated to generate a new generation of formulation solutions. As the iteration progresses, the population will continuously tend towards the optimal solution, and finally, a formulation solution close to the optimal solution will be obtained. This solution will be broadcast to each workshop for execution.
[0118] S34: Dataset and model update. In the online scenario, the edge client updates the local data by performing real evaluations, denoted as local surrogate model and then updates it and re - uploads it to the cloud for subsequent evolution.
[0119] The edge workshop validates the effectiveness of the optimized formula through actual production behaviors and updates the local dataset. Specifically, each workshop will produce according to the optimized formula provided by the cloud, collect relevant production data such as production time, output, and quality indicators, and add these new data points to the dataset. The local proxy model of the workshop will be updated based on the new data to enhance its evaluation ability for future formula optimization. After the update, the proxy model is sent back to the cloud for subsequent optimization iterations. Through this process, the edge workshop and the cloud form a continuous feedback and optimization loop, enabling the production formula to be continuously adjusted to adapt to new production requirements and changing process conditions.
[0120] S35: The multi-round model combination strategy based on ensemble learning is a model combination method that runs in the cloud and is used to aggregate proxy models trained by different edges. This method not only aggregates models trained in a single round but also utilizes the information of cross-round models to improve the overall effectiveness and robustness of the model. The specific steps are as follows:
[0121] 1) Edge cross-round model combination. In the current optimization round (the m-th round), for each edge k (k = 1, 2,..., K), the cloud will record the proxy models collected in the previous m rounds, denoted as The algorithm first obtains the current optimal proxy model according to the designed metrics:
[0122]
[0123] The Median(·) operator returns the median of the input set, representing the selection of a stable model measured by the median of the metrics. The metric of the model is calculated as follows:
[0124]
[0125] The metric The calculation refers to step S36.
[0126] 2) Cloud-edge model combination. In the second step of the algorithm, the cross-round models obtained in 1) are aggregated using the weights obtained by the boundary judgment strategy:
[0127]
[0128] The weight ω k (x) is calculated according to the data interval where the sample falls, and lb k and ub k represent the upper and lower limits of the statistical data in edge k respectively, and Z t is the normalization factor.
[0129] S36: Combining the powerful computing capabilities of cloud computing with the low-latency characteristics of edge computing, in this embodiment, through the edge-cloud collaboration method, the model parameters are dynamically adjusted during the distributed evolution process to continuously improve the performance of the combined model and optimize the algorithm efficiency. The following are the specific steps for constructing the online enhanced combined training strategy for the model:
[0130] 1) Input samples and sample weights. The data set where n k is the total number of samples in the edge device k. The sample weights in the previous optimization process At the beginning (the first round of training) are set to 1 / n k , and all samples have equal weights, and the sum is 1.
[0131] 2) Sampling of the model training set. Based on the weight distribution perform sampling without replacement on the data set to obtain the training set The training set is a subset of, and the higher the weight of the sample, the greater the chance of being selected.
[0132] 3) Training of the surrogate model. Use to train the surrogate model This algorithm does not limit the selection of the surrogate model, and machine learning models such as RBF and GP can be selected;
[0133] 4) Calculate the model error rate and model metric indicators, and use the surrogate model to evaluate each sample in, and compare the result with the true label value to calculate the relative error of the sample
[0134]
[0135] where is the maximum absolute error, and the model error rate (weighted sum of sample relative errors) is calculated as follows:
[0136]
[0137] The model metric indicator is then defined as:
[0138]
[0139] 5) Iteration of sample weights. The weight of each sample in is updated using the following formula:
[0140]
[0141] is a normalization factor to ensure that the sum of weights is 1.
[0142] S37: Service distributed containerized deployment solution;
[0143] The deployment solution is as Figure 4 shown. After the algorithm process is decomposed into multiple independent service interfaces, the interfaces are encapsulated into Docker images and saved in the cloud image repository (Docker Hub) for version management and distribution. K8s is used in the cloud to manage the life cycle of all containerized services. As the core control component in the cloud, CloudCore defines and manages the deployment of each service image through the Pod YAML file. The Pod YAML file includes information such as deployment, network configuration, and resource limits to ensure the efficient scheduling and automated management of service containers. CloudCore is responsible for: 1) managing the startup, stop, scaling, and recovery of service containers; 2) continuously monitoring the status of service containers to ensure the health and stability of container operation; 3) dynamically performing elastic scaling and load balancing based on traffic and load conditions to ensure the high availability and performance of the system. The edge node pulls the Docker image through the KubeEdge framework and instantiates EdgeCore, which is responsible for: 1) starting service containers on the edge node; 2) using the monitoring and management tools provided by KubeEdge to monitor the running status of container instances in real time, including CPU usage, memory consumption, and network traffic, etc.; 3) when a container encounters an exception or crashes, EdgeCore will automatically restart the container according to the preset policy for fault recovery.
[0144] S38: Build a cloud-edge communication mechanism;
[0145] In this embodiment, the core of the design of the communication and data transmission mechanism is to ensure stable low-latency communication and efficient data processing between the cloud and edge computing nodes in the system, while ensuring security and scalability. The communication between the cloud and the edge node is implemented through the KubeEdge framework, as Figure 5 shown, and the specific design is as follows:
[0146] 1) Control channel. A stable control channel is established between CloudCore and EdgeCore for key operations such as task scheduling, resource management, and status monitoring. The control channel uses the TLS encryption protocol to ensure the security of data transmission and prevent data from being tampered with or leaked during transmission.
[0147] 2) Heartbeat mechanism. By sending periodic heartbeat signals, CloudCore can monitor the status of EdgeCore in real time, including its health, load, etc. If an abnormality or disconnection occurs in the edge node, a preset disaster recovery strategy will be triggered to migrate tasks to the standby healthy node to ensure the stable operation of the system.
[0148] 3) Data pipeline. It adopts streaming data transmission and chunked concurrent transmission technologies to support the rapid transmission of large-scale data. It realizes two-way real-time communication through the WebSocket protocol, ensuring low latency and reliability of data transmission and avoiding the latency problem of traditional HTTP requests. All interfaces support full-duplex communication, allowing two-way data transmission between the cloud and edge nodes when processing tasks. The communication between the edge node and the underlying device adopts the lightweight MQTT protocol, which is based on the publish-subscribe mode and supports real-time data collection and task instruction issuance.
[0149] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A data-driven learning model combination optimization method based on edge-cloud collaboration, characterized in that: The steps include: Build a collaborative computing framework based on cloud nodes, edge nodes, and device nodes to allocate distributed tasks and computing resources; Build multiple microservice interfaces, including platform service interface, edge node evaluation service interface, edge node model training service interface, cloud node model upgrade service interface, and cloud node candidate solution evaluation service interface; The platform service interface is provided with a plurality of functional modules, and the process of the distributed agent-assisted evolutionary algorithm is decomposed into a plurality of functional modules. The evaluation service interface evaluates the candidate solutions of the evaluation request of the cloud node or the device node, and stores the evaluation results in the edge node. The model training service interface trains the proxy model based on the local data of the edge node. The model upgrade service interface obtains the model training results, combines and integrates the models, and distributes the updated proxy model version to each edge node. The candidate solution evaluation service interface screens the candidate solutions and performs real environment evaluation by calling the evaluation service interface of the edge node. A multi-round model combination strategy based on ensemble learning and a model online enhanced combination training strategy are constructed. Multiple microservice interfaces are encapsulated into Docker images and saved in the cloud image repository for version management and distribution, and a communication mechanism between cloud nodes and edge nodes is constructed.
2. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 1 is characterized in that: The cloud node executes the agent-assisted evolution algorithm, monitors the status of edge nodes and device nodes, and dynamically adjusts the distribution of cloud resources. The edge node obtains data collected by the device node to perform local training and updates on the agent model, executes or offloads computing tasks, and manages edge resources. The device node collects real-time data and manages the computing resources and communication resources of the device.
3. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 2 is characterized in that: Cloud nodes use Kubernetes technology and use its API to interact with CloudCore components. CloudCore components act as a scheduling center and dynamically adjust the distribution of cloud resources based on workload prediction algorithms.
4. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 2 is characterized in that: The edge node implements the distribution and execution of containerized tasks and unified management of edge resources based on the EdgeCore component. The EdgeCore component and the CloudCore component communicate through a two-way encrypted channel to achieve real-time synchronization of tasks and resources.
5. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 4 is characterized in that: The edge node pulls the Docker image and instantiates the EdgeCore component through the KubeEdge framework. The EdgeCore component starts the service container on the edge node and uses the monitoring and management tools provided by KubeEdge to monitor the running status of the container instance in real time. When the container is abnormal or crashes, the EdgeCore component automatically restarts the container according to the preset strategy to recover from the fault.
6. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 1 is characterized in that: Construct a multi-round model combination strategy based on ensemble learning and a model online enhanced combination training strategy, including: Train multiple proxy models based on local data sets, aggregate multiple proxy models to form a global model, calculate the fitness of each optimization scheme based on the global model, and update the local data set through real environment evaluation; Collect proxy models from multiple optimization rounds to obtain cross-round models, and aggregate the cross-round models; Based on the weight distribution, the local data set is sampled without replacement to obtain a training set. The proxy model is trained based on the training set to evaluate each sample in the local data set, calculate the sample relative error, and update the sample weight.
7. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 6 is characterized in that: The local data set includes the production history, raw material ratio, process parameters, equipment status and production output of each workshop. The agent model is trained according to the production optimization goal to evaluate the industrial production formula.
8. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 6 is characterized in that: The fitness of each optimization scheme is calculated based on the global model, specifically including: The initial population is randomly generated, and each individual represents a potential formulation. The fitness of each optimization solution is calculated based on the global model. The individuals in the population are selected, crossed, and mutated to generate a new generation of formulations, and the optimal formulation is iteratively generated.
9. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 1 is characterized in that: Construct a multi-round model combination strategy based on ensemble learning and a model online enhanced combination training strategy, including: For each edge node k, the cloud node records the proxy model collected in the previous m rounds, recorded as Select the current best proxy model: The Median(·) operator returns the median of the input set, which means that a stable model is selected based on the median of the index. The metric of the model is The calculation is as follows: in, Represents model metrics; The cross-round models are aggregated using the weights obtained by the boundary judgment strategy: Among them, the weight ω k (x) is calculated based on the data interval where the data set falls, lb k andub k Respectively represent the upper and lower limits of the statistical data in edge node k, Z t is the normalization factor; The dataset is represented as: n k is the total number of samples in edge node k, for the dataset Perform sampling without replacement to obtain the training set Based on the training set Training the Agent Model Using Proxy Models Evaluation Dataset For each sample, compare the result with the true label value and calculate the sample relative error in, is the maximum absolute error, model error rate The calculation formula is expressed as: Model Metrics It is defined as: Dataset The weight of each sample in Use the following formula to update: in, is the standardization factor.
10. The data-driven learning model combination optimization method based on edge-cloud collaboration according to claim 1 is characterized in that: Build a communication mechanism between cloud nodes and edge nodes, including: The communication between cloud nodes and edge nodes is implemented through the KubeEdge framework. A control channel is established between the CloudCore component and the EdgeCore component. The control channel uses the TLS encryption protocol for data transmission. Based on the periodic heartbeat mechanism, the CloudCore component monitors the status of the EdgeCore component. If an edge node is abnormal or offline, the preset disaster recovery strategy is triggered to migrate tasks to the backup node. Streaming data transmission and block concurrent transmission technology are used for data transmission, and two-way real-time communication is achieved through the WebSocket protocol.
Citation Information
Patent Citations
Convolutional neural network structure optimization method based on proxy-assisted evolutionary algorithm
CN115879509A
Electric power micro-service layering system based on cloud edge collaboration
CN116506474A
Kubernetes-based service container scheduling method and system under cloud edge collaboration
CN118337786A
Agent model assisted evolutionary generative adversarial network architecture search method and system
CN118821905A
Node-type edge computing system
WO2025001634A1
Cited By
Space science experiment data online collaborative analysis method and system
CN121073180A
A space science experiment data online collaborative analysis method and system
CN121073180B
Collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision
CN121170549A
A collaborative learning method and system based on edge sample intelligent grading and cloud intelligent decision
CN121170549B