MEC joint computing performance improvement method under multi-terminal heterogeneous data scene

By designing the MEC heterogeneous data system model in multi-terminal heterogeneous data scenarios, using Pareto principle and generative adversarial network technologies, a reinforcement learning model is built, and resource scheduling problems in MEC joint computing is solved, achieving more accurate performance improvement and intelligent management.

CN120373418APending Publication Date: 2025-07-25INST OF WAR STUDIES ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372569.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The MEC joint computing performance improvement method in multi-terminal heterogeneous data scenarios has the problem of resource scheduling in MEC multi-terminal isomorphic devices, multi-terminal equipment heterogeneous and multi-terminal collaborative heterogeneous MEC networks, resulting in the inability to optimize and improve system performance.

Method used

By designing the MEC heterogeneous data system model, using Pareto principle to divide the available local sample data sets of edge devices, combining generative adversarial networks, asynchronous thread scheduling, reinforcement learning, Markov decision-making and random forest algorithms, a reinforcement learning model and deep reinforcement model based on Markov decision-making are built to solve resource scheduling problems and realize intelligent management.

Benefits of technology

Accurately capture the local sample data distribution characteristics of edge devices, realize effective processing of data heterogeneity, improve the MEC joint computing performance in multi-terminal heterogeneous data scenarios, enhance the degree of intelligence, solve resource scheduling problems, and ensure the accuracy and dynamicity of monitoring data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373418A_ABST
    Figure CN120373418A_ABST
Patent Text Reader

Abstract

The invention discloses an MEC joint calculation performance improvement method in a multi-terminal heterogeneous data scene, and relates to the technical field of multi-terminal heterogeneous MEC joint performance, and the method comprises the following steps: designing an MEC heterogeneous data system model, recording isomorphic multi-terminal equipment operation data, dividing an available local sample data set of edge equipment according to a Pareto principle, and carrying out the calculation of the available local sample data set; a reinforcement learning model based on Markov decision is constructed by using reinforcement learning, Markov decision and a random forest algorithm, and a deep reinforcement model is constructed by combining a Q-learning algorithm, a function approximator and a neural network algorithm through the reinforcement learning model based on Markov decision. By means of reinforcement learning, the Markov decision, the random forest algorithm, the Q-learning algorithm, the function approximator and the neural network algorithm, a reinforcement learning model and a deep reinforcement model based on the Markov decision are constructed, and the intelligent degree in the multi-terminal heterogeneous MEC joint calculation performance improvement process is remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of MEC joint performance for multi-terminal heterogeneity, and particularly to a method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario. Background Art

[0002] In today's digital age, the Internet of Things applications are booming, and MEC joint computing in a multi-terminal heterogeneous data scenario has become a key research area. With the explosive growth of the number of Internet of Things devices, the data generated by a large number of devices shows significant heterogeneity. Not only is the data distribution uneven, but most of it is non-independent and identically distributed, which poses a huge challenge to MEC joint computing. At the same time, the devices vary greatly in computing, storage, and communication capabilities, and the network environment is complex and changeable, further exacerbating the computational complexity. In terms of data heterogeneity, the training time of large-sample data devices is long, resulting in an increase in the overall training delay of the system and a slow convergence speed; device heterogeneity can cause communication timeouts and model loss, resulting in network congestion and affecting the training effect of the system; network heterogeneity makes it difficult to estimate the device training time and the resource utilization rate is low. Therefore, a method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario has emerged, aiming to improve the overall performance of the system and the resource utilization rate to meet the growing needs of Internet of Things applications.

[0003] Although the existing technologies have made great progress in the direction of improving the MEC joint computing performance for multi-terminal heterogeneity, there are still some problems to be optimized. The existing methods for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario have problems of data heterogeneity of MEC multi-terminal homogeneous devices, multi-terminal device heterogeneity, and resource scheduling in a multi-terminal collaborative heterogeneous MEC network, resulting in the inability to optimize and improve the MEC joint computing performance in a multi-terminal heterogeneous data scenario. Summary of the Invention

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario, including the following steps:

[0005] Step 1: Design an MEC heterogeneous data system model, record the operation data of homogeneous multi-terminal devices, and divide the available local sample data sets of edge devices according to the Pareto principle.

[0006] Step 2: According to the divided available local sample data sets of edge devices, adopt a corresponding training algorithm allocation strategy. Combining with Step 1, the problem of data heterogeneity of MEC multi-terminal homogeneous devices is solved, and a generative adversarial network is proposed, providing data support for local gradient descent of small samples.

[0007] Step 3: Use the asynchronous thread scheduling of the central server to obtain the global model parameters through time-weighted average and weighted asynchronous aggregation. Use the intelligent platform management interface device and Geekbench test tool to obtain the computing power value of the central server, and select the global model update method according to the computing power of the central server;

[0008] Step 4: The central server sends the updated global model and scheduling instructions to the edge devices according to the operating status of the edge devices. The edge devices train the global model according to the scheduling instructions. Combining with Step 3, the problem of multi-terminal device heterogeneity is solved;

[0009] Step 5: Use reinforcement learning, Markov decision-making, and random forest algorithms to construct a reinforcement learning model based on Markov decision-making;

[0010] Step 6: Through the reinforcement learning model based on Markov decision-making, combine the Q-learning algorithm, function approximator, and neural network algorithm to construct a deep reinforcement model;

[0011] Step 7: Deploy the DRL agent, define the DRL agent state, select the CPU actions of multi-terminal devices through the DRL agent and set the reward mechanism, and construct a joint reinforcement learning model through the random forest algorithm;

[0012] Step 8: Based on the joint optimization algorithm of reinforcement learning, combine the reinforcement learning model based on Markov decision-making and the deep reinforcement model for application. Combining with Step 5, Step 6, and Step 7, the problem of resource scheduling in the multi-terminal collaborative heterogeneous MEC network is solved.

[0013] A further improvement of the technical solution of the present invention is that in the above Step 1, when designing the MEC heterogeneous data system model and recording the operation data of homogeneous multi-terminal devices, the process of dividing the available local sample data sets of edge devices according to the Pareto principle includes:

[0014] The MEC heterogeneous data system model includes a central server and k MEC edge nodes. Each MEC edge node includes edge devices and local sample data sets, and a federated learning training model is deployed on each MEC edge node;

[0015] The operation data of the homogeneous multi-terminal devices includes the available local sample data set on the kth edge device as , the local computing rate of the kth edge device, and the communication time between the central server and the kth edge device, and record the above operation data of the homogeneous multi-terminal devices;

[0016] The process of calculating the waiting time required for the t-th round of update of the computing server is as follows:

[0017]

[0018]

[0019] Among them, is the time required to wait for the t-th round of update of the server, is the local update training time of the k-th edge device, is the available local sample data set on the k-th edge device, is the local computing rate of the k-th edge device, is the communication time between the central server and the k-th edge device;

[0020] Taking the edge device number k as the abscissa and the available local sample data set on the k-th edge device as as the ordinate, draw a power-law distribution diagram of data samples;

[0021] Through the power-law distribution diagram of data samples, select the available local minimum sample data set on the k-th edge device and the available local maximum sample data set on the k-th edge device , set the threshold of the available local sample data set on the k-th edge device , satisfying , according to the Pareto principle, regard the available local sample data set of edge devices greater than as large samples, and regard the available local sample data set of edge devices less than as small samples.

[0022] A further improvement of the technical solution of the present invention lies in: in the second step, the process of adopting the corresponding training algorithm allocation strategy according to the divided available local sample data set of edge devices includes:

[0023] When the available local sample data set of the edge device is greater than , determine the gradient descent condition. The corresponding available local sample data set of the edge device is a large sample. In the MEC network based on the FedAvg algorithm, select small data samples for training, perform stochastic gradient descent calculation, judge the improvement of the available local sample size on the target loss function, and adjust the size of the available local sample until the gradient descent condition is satisfied;

[0024] When the available local sample data set of the edge device is less than , the corresponding available local sample data set of the edge device is a small sample. Through the generative adversarial network, obtain synthetic samples similar to the small sample data, combine the small sample and the synthetic samples to obtain combined samples, and use the FedAvg algorithm to perform local gradient descent on the combined samples.

[0025] A further improvement of the technical solution of the present invention lies in that the principles of the FedAvg algorithm and the generative adversarial network include:

[0026] S1. Using the available local sample dataset on the edge device, establish a local available sample subset S, select the size of the small samples in the local available sample subset, and the process of calculating the minimized objective loss function is as follows: The process of calculating the minimized objective loss function is as follows:

[0027]

[0028] Among them, is the convex empirical loss objective function, is the size of the small samples in the local available sample subset, is the local available sample subset, and there is , is the prediction function calculated by gradient descent through the parameter ω;

[0029] Determine the gradient descent condition, analyze the improvement of the minimized objective loss function. If the improvement remains unchanged, select the local available sample data with the same sample size as to perform the next iteration. If the improvement changes, re-select a sample larger than to perform the next iteration until the gradient descent condition is met;

[0030] S2. The generative adversarial network consists of a generator, a discriminator, and a loss function. Initialize the neural network parameters of the generator and the discriminator, extract real samples from the small samples, the generator synthesizes synthetic samples similar to the small samples, input the small samples and the synthetic samples into the discriminator, calculate the loss function of the discriminator, obtain the judgment result of the synthetic samples, update the discriminator parameters through the backpropagation algorithm to make the discriminator better distinguish between small samples and synthetic samples, use the judgment result of the synthetic samples to calculate the loss function of the generator, and update the generator parameters through the backpropagation algorithm to make the synthetic samples generated by the generator more capable of deceiving the discriminator and increase the probability that the synthetic samples are close to the small samples. Repeat the alternating training of the generator and the discriminator, and through multiple rounds of iteration, obtain synthetic samples similar to the small samples.

[0031] A further improvement of the technical solution of the present invention lies in that in the third step, using the asynchronous thread scheduling of the central server, through time-weighted average and weight asynchronous aggregation, to obtain the global model. The process of selecting the way to update the global model according to the computing power of the central server includes:

[0032] A1. The central server consists of a scheduling thread, an update thread, and a coordinator. Among them, the scheduling thread task and the update thread task are controlled by two asynchronous parallel threads. The coordinator periodically sends scheduling instructions and the global model to the MEC edge devices. When the edge devices complete their work, they upload and update the global model parameters. The central server obtains the upload queue of the global model through the coordinator and updates the global model sequentially;

[0033] For the k-th device, the central server records the central server scheduling timestamp and the central server global model update timestamp through the asynchronous thread scheduling process, and calculates the central server waiting delay time using the timestamps:

[0034]

[0035] Among them, is the central server waiting delay time, is the central server global model update timestamp, is the central server scheduling timestamp;

[0036] A2. Using time-weighted average, a delay threshold is set. When the parameters fed back by the edge device are higher than the central server waiting delay time, the global model update is lagged. A mixed hyperparameter is set, and the mixed hyperparameter is used to control the impact of the delay on the global model update. The process of calculating the mixed hyperparameter for the t-th iteration through the hinge function to represent the delay function and using the delay function and the mixed hyperparameter is as follows:

[0037]

[0038]

[0039] Among them, is the delay function, is the mixed hyperparameter for the t-th iteration, a and b are constants, and a > 0, b > 0, is the time delay magnitude exceeding the server's expected waiting time, is the delay threshold;

[0040] In the t-th iteration, when the central server receives the local training model parameters fed back by the k-th edge device, the process of calculating the time-weighted average to update the global model is as follows:

[0041]

[0042] Among them, is the time-weighted average to update the global model in the t-th iteration, is the time-weighted average to update the global model in the (t - 1)-th iteration, is the mixed hyperparameter for the t-th iteration, The central server receives the local training model parameters fed back by the k-th edge device;

[0043] A3. Adopt weighted asynchronous aggregation. According to the local data of the k-th edge device scheduled by the central server in the t-th iteration, set the weight parameter of the k-th edge device scheduled in the t-th iteration. The process of calculating the weighted asynchronous aggregation to update the global model is as follows:

[0044]

[0045]

[0046] Among them, is the global model updated by weighted asynchronous aggregation in the t-th iteration, is the global model updated by weighted asynchronous aggregation in the (t - 1)-th iteration, is the weight hyperparameter in the t-th iteration, is the weight parameter of the k-th edge device scheduled in the t-th iteration, is the mixing hyperparameter in the t-th iteration, is the local training model parameters fed back by the central server from the k-th edge device;

[0047] A4. Deploy an intelligent platform management interface device and a Geekbench test tool on the edge device. Among them, the intelligent platform management interface device measures the performance metrics of the CPU, and the Geekbench test tool uses the performance metrics of the CPU to obtain the computing power value of the central server. Set the computing power threshold of the central server. According to the computing power threshold of the central server, select the way to update the global model. When the computing power value of the central server is less than the computing power threshold of the central server, adopt time-weighted average to update the global model; when the computing power value of the central server is greater than the computing power threshold of the central server, adopt weighted asynchronous aggregation to update the global model.

[0048] A further improvement of the technical solution of the present invention is that in the fourth step, the central server transmits the updated global model and the scheduling instruction to the edge device according to the operating state of the edge device. The process of the edge device training the global model according to the scheduling instruction includes:

[0049] The operating state of the edge device includes an idle state, a running state, and a blocked state. The edge device receives the updated global model and the scheduling instruction;

[0050] When the edge device is in the idle state, the edge device enters the running state, updates the global model using the global model and the available local sample data on the edge device, and transmits the updated global model parameters to the central server;

[0051] When the edge device is in a blocked state, switch the state according to the available local sample data usage of the edge device. During operation, when there are situations such as insufficient available local sample resources, network interruption, and power exhaustion of the edge device, suspend the current training task of the edge device. After waiting for the blocked state of the edge device to switch to the idle state, continue to train the global model; during operation, when there are situations such as the edge device being unavailable and the central server rejecting communication, the central server scheduling thread adjusts edge , and wait until the edge device is available and communicating normally, and then train the global model.

[0052] A further improvement of the technical solution of the present invention lies in: in step five, the process of constructing a reinforcement learning model based on Markov decision using reinforcement learning, Markov decision, and random forest algorithms includes:

[0053] Select time t, and the reinforcement learning agent will represent the observation state, represent the execution action, represent the state transition probability, represent the reward, and represent the discount factor. After executing the action, the reinforcement learning agent switches the state according to the state transition probability, and uses to represent the switched state, and the reinforcement learning agent obtains the reward;

[0054] Use the random forest algorithm to construct a random forest model, set the random forest model parameters, use the switched state and its reinforcement learning agent reward as the data set, divide it into a training set and a test set according to a ratio of 7:3, use the training set data to train the random forest model, and through iterative learning, obtain the non-linear relationship between the switched state and the reinforcement learning agent reward, so as to realize that the random forest model inputs the switched state and outputs the corresponding reinforcement learning agent reward. Input the test set data into the trained random forest model, obtain the output value of the trained random forest, compare the reinforcement learning agent reward corresponding to the actual switched state with the output value of the trained random forest model, adjust the random forest model parameters, optimize the performance of the random forest model, and obtain a reinforcement learning model based on Markov decision.

[0055] A further improvement of the technical solution of the present invention lies in: in step six, the process of constructing a deep reinforcement model by combining a reinforcement learning model based on Markov decision, Q-learning algorithm, function approximator, and neural network algorithm includes:

[0056] ​B1. The process of obtaining the transformed state and its reinforcement learning agent reward through a reinforcement learning model based on Markov decision, establishing a transformed state-reward function, and obtaining the maximum reward value of the initial state through the Q-learning algorithm is as follows:

[0057]

[0058]

[0059] Among them, is the maximum reward value function of the initial state, also known as the maximum expected value function of the cumulative total return obtained starting from state based on policy ; is the transformed state-reward function; is the transformed state-transformed maximum reward value function; E is the mathematical expectation; is the policy mapping from the initial state to the selected action; represents the reward; represents the discount factor; represents the observed state, represents the executed action, and represent the specific values of the observed state and the executed action respectively;

[0060] Use a function approximator to learn the function to approximate the maximum reward value function from the initial state to the transformed state, where represents the policy parameter of the current state;

[0061] B2. Use the neural network algorithm to construct a neural network model. Take the initial state and its maximum reward value and the transformed state and its maximum reward value as the data set, and divide it into a training set and a test set according to a ratio of 7:3. Select MLP as the neural network structure. The input layer includes two neurons, which receive the initial state and the transformed state. The hidden layer configures the MSE function. The output layer includes two neurons, which output the maximum reward value of the initial state and the maximum reward value of the transformed state;

[0062] Input the training set data into the neural network model, set the learning rate to 0.01, and the number of iterative training times to 1000. The training process includes forward propagation and backward propagation. Among them, forward propagation is used to calculate the predicted output data, and backward propagation is used to update the weights and biases of the model. Through repeated iterative training, learn the non-linear relationship between the initial state and the maximum reward value of the initial state and the non-linear relationship between the transformed state and the maximum reward value of the transformed state until the set number of iterative training times is reached, and obtain the trained neural network model;

[0063] Input the test set data into the trained neural network model, use the MSE function to evaluate the error between the output value and the actual value of the neural network model, adjust the parameters of the neural network model according to the evaluation results, optimize the performance of the neural network model, and obtain the deep reinforcement model;

[0064] Use the deep reinforcement model to obtain the maximum reward value of the initial state and the maximum reward value of the transformed state, and update the maximum reward value of the initial state with the maximum reward value of the transformed state until the last state ends.

[0065] A further improvement of the technical solution of the present invention lies in: in step seven, deploy a DRL agent, define the DRL agent state, select the CPU actions of multi-terminal devices through the DRL agent and set a reward mechanism. The process of constructing a joint reinforcement learning model through the random forest algorithm includes:

[0066] Install a network tester on the multi-terminal device to test the real-time network bandwidth of the multi-terminal device. Deploy the DRL agent on the central server to realize the interaction between the DRL agent and the multi-terminal device. Take the network bandwidth of each multi-terminal device as the state space of the DRL agent, denoted as Set a network bandwidth threshold. When the real-time network bandwidth of the multi-terminal device is lower than the network bandwidth threshold, select to increase the CPU frequency of the multi-terminal device; when the real-time network bandwidth of the multi-terminal device is higher than the network bandwidth threshold, select to decrease the CPU frequency of the multi-terminal device;

[0067] After the DRL agent executes an action in the initial state in the t-th iteration, obtain a learning reward. The process of setting the reward mechanism is as follows:

[0068]

[0069]

[0070] wherein, is the reward function, T is the total number of iterations, t is the current iteration number, represents the discount factor, is the maximum device training model time obtained by the DRL agent, is the expected cumulative discounted reward obtained by the DRL agent when in state ;

[0071] A random forest model is constructed using the random forest algorithm, and the parameters of the random forest model are set. The actions in the initial state and their corresponding reward functions in the t-th iteration, along with the expected cumulative discounted rewards of the DRL agent, are used as a dataset and divided into a training set and a test set in a ratio of 7:3. The training set data is used to train the random forest model. Through iterative learning, the non-linear relationship between the actions in the initial state in the t-th iteration, their corresponding reward functions, and the expected cumulative discounted rewards of the DRL agent is obtained, enabling the random forest model to input the actions in the initial state in the t-th iteration and output the non-linear relationship of the corresponding reward functions and the expected cumulative discounted rewards of the DRL agent. The test set data is input into the trained random forest model to obtain the output values of the trained random forest. By comparing the actual actions in the initial state in the t-th iteration, their corresponding reward functions, and the expected cumulative discounted rewards of the DRL agent with the output values of the trained random forest model, the parameters of the random forest model are adjusted to optimize the performance of the random forest model, and a joint reinforcement learning model is obtained.

[0072] Using the joint reinforcement learning model, the reward functions corresponding to the actions in the initial state in the t-th iteration and the expected cumulative discounted rewards of the DRL agent are obtained, and the time cost of multi-terminal computing is minimized using these reward functions and the expected cumulative discounted rewards of the DRL agent.

[0073] A further improvement of the technical solution of the present invention lies in: in step eight, in the joint optimization algorithm based on reinforcement learning, the process of combining the reinforcement learning model based on Markov decision-making and the deep reinforcement model includes:

[0074] Combining the reinforcement learning model based on Markov decision-making and the deep reinforcement model, and using the DRL agent for iterative training in a heterogeneous MEC network with multi-terminal collaboration. The workflow is as follows:

[0075] C1. Initialize the MEC parameters and the DRL agent;

[0076] C2. The multi-terminal devices download the global model from the central server;

[0077] C3. The DRL agent observes the heterogeneous network bandwidth and updates the state space;

[0078] C4. The DRL agent selects an action , executes the action of controlling the CPU frequency. The multi-terminal devices obtain and update the global model parameters according to the deep reinforcement model, update the global model, and upload it to the central server. The central server aggregates and updates the global model;

[0079] C5. The DRL agent enters the next state , through the combined reinforcement learning model, obtain the reward function corresponding to the action of the initial state and the expected cumulative discounted reward of the DRL agent in the t-th iteration;

[0080] C6. After the end of an iteration cycle, repeat steps C2 to C5 until the training iteration rounds are completed and the training expected effect is achieved.

[0081] The beneficial effects of the present invention are as follows: The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario, compared with the traditional method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario, in the method of the present invention, through the technology of dividing the data set by the Pareto principle, the distribution characteristics of the local sample data of the edge device are accurately captured, and the division boundary of the large and small sample data sets is obtained, laying a foundation for the effective processing of data heterogeneity; by using the asynchronous thread scheduling technology, the scheduling timestamp of the central server and the timestamp for updating the global model are obtained in real time, the waiting delay time is calculated, and the interaction process between the central server and the edge device is comprehensively monitored. With the help of reinforcement learning, Markov decision-making, random forest algorithm, Q-learning algorithm, neural network algorithm and function approximator, a reinforcement learning model based on Markov decision-making and a deep reinforcement model are constructed, achieving intelligent management of the resources of the MEC system in a multi-terminal heterogeneous data scenario, solving the problems of data heterogeneity of MEC multi-terminal homogeneous devices, multi-terminal device heterogeneity and resource scheduling in multi-terminal collaborative heterogeneous MEC networks in the existing method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario, ensuring that the method in the present invention can refine the dynamic monitoring standard for the method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario within a more accurate range, making the monitored data a more accurate indicator under the same conditions. The research and application of this method significantly enhance the degree of intelligence in the process of improving the MEC joint computing performance of multi-terminal heterogeneity. Description of the Drawings

[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0083] Figure 1 It is a flowchart of the method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario of the present invention;

[0084] Figure 2 It is a scheduling diagram of edge devices in asynchronous communication;

[0085] Figure 3 It is a reinforcement learning model diagram based on Markov decision-making;

[0086] Figure 4 It is a flowchart of a joint optimization algorithm based on reinforcement learning. Specific implementation manners

[0087] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0088] As Figure 1 shown, the present invention provides a method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario, which consists of the following steps:

[0089] Step 1: Design an MEC heterogeneous data system model, record the operation data of homogeneous multi-terminal devices, and divide the available local sample data sets of edge devices according to the Pareto principle;

[0090] Step 2: According to the divided available local sample data sets of edge devices, adopt corresponding training algorithm allocation strategies, and combine with Step 1 to solve the problem of data heterogeneity of MEC multi-terminal homogeneous devices, and propose a generative adversarial network, which provides data support for local gradient descent of small samples;

[0091] Step 3: Use the asynchronous thread scheduling of the central server, obtain the global model parameters through time-weighted average and weight asynchronous aggregation, use the intelligent platform management interface device and Geekbench test tool to obtain the computing power value of the central server, and select the global model update method according to the computing power of the central server;

[0092] Step 4: The central server sends the updated global model and scheduling instructions to the edge devices according to the operating states of the edge devices, and the edge devices train the global model according to the scheduling instructions. Combining with Step 3, the problem of heterogeneity of multi-terminal devices is solved;

[0093] Step 5: Use reinforcement learning, Markov decision-making and random forest algorithms to construct a reinforcement learning model based on Markov decision-making;

[0094] Step 6: Through the reinforcement learning model based on Markov decision-making, combine the Q-learning algorithm, function approximator and neural network algorithm to construct a deep reinforcement model;

[0095] Step 7: Deploy a DRL agent, define the DRL agent state, select the CPU actions of multi-terminal devices through the DRL agent and set up a reward mechanism, and build a joint reinforcement learning model through the random forest algorithm;

[0096] Step 8: Based on the joint optimization algorithm of reinforcement learning, combine and apply the reinforcement learning model based on Markov decision-making and the deep reinforcement model. Combining Step 5, Step 6, and Step 7, the problem of resource scheduling in a multi-terminal collaborative heterogeneous MEC network is solved.

[0097] Preferably, in Step 1, when designing the MEC heterogeneous data system model and recording the operation data of homogeneous multi-terminal devices, the process of dividing the available local sample data set of edge devices according to the Pareto principle includes:

[0098] The MEC heterogeneous data system model includes a central server and k MEC edge nodes. Each MEC edge node includes edge devices and a local sample data set, and a federated learning training model is deployed on each MEC edge node;

[0099] Among them, the operation data of homogeneous multi-terminal devices includes the available local sample data set on the kth edge device as , the local computing rate of the kth edge device, and the communication time between the central server and the kth edge device, and record the above operation data of homogeneous multi-terminal devices;

[0100] The process of calculating the time required for the server to wait for the t-th update is as follows:

[0101]

[0102]

[0103] Among them, is the time required for the server to wait for the t-th update, is the local update training time of the kth edge device, is the available local sample data set on the kth edge device, is the local computing rate of the kth edge device, is the communication time between the central server and the kth edge device;

[0104] Taking the edge device number k as the abscissa and the available local sample data set on the kth edge device as as the ordinate, draw a power-law distribution diagram of data samples;

[0105] Through the power-law distribution diagram of data samples, select the available local minimum sample data set and the available local maximum sample data set on the kth edge device , set the threshold of the available local sample dataset on the k-th edge device , satisfy , according to the Pareto principle, regard the available local sample datasets of edge devices greater than as large samples, and regard the available local sample datasets of edge devices less than as small samples.

[0106] Preferably, in step two, the process of adopting the corresponding training algorithm allocation strategy according to the divided available local sample datasets of edge devices includes:

[0107] When the available local sample dataset of an edge device is greater than , determine the gradient descent condition. For the corresponding available local sample dataset of the edge device as a large sample, in the MEC network based on the FedAvg algorithm, select small data samples for training, perform stochastic gradient descent calculation, judge the improvement of the available local sample size on the objective loss function, and adjust the size of the available local sample until the gradient descent condition is satisfied;

[0108] When the available local sample dataset of an edge device is less than , for the corresponding available local sample dataset of the edge device as a small sample, obtain synthetic samples similar to the small sample data through a generative adversarial network, combine the small sample and the synthetic samples to obtain combined samples, and use the FedAvg algorithm to perform local gradient descent on the combined samples.

[0109] Preferably, the principles of the FedAvg algorithm and the generative adversarial network include:

[0110] S1. Use the available local sample datasets on the edge device to establish a local available sample subset S, and select the size of the small samples in the local available sample subset, and the process of calculating the minimized objective loss function is as follows:

[0111]

[0112] where, is the convex empirical loss objective function, is the size of the small samples in the local available sample subset, is the local available sample subset, and there is , is the prediction function for gradient descent calculation through the parameter ω;

[0113] Determine the gradient descent condition, analyze 's improvement on the minimized objective loss function. If the improvement remains unchanged, select the one related to Perform the next iteration on the locally available sample data with the same sample size. If the improvement situation changes, reselect a sample larger than and perform the next iteration until the gradient descent condition is met;

[0114] S2. The generative adversarial network consists of a generator, a discriminator, and a loss function. Initialize the neural network parameters of the generator and the discriminator. Extract real samples from small samples. The generator synthesizes synthetic samples similar to the small samples. Input the small samples and the synthetic samples into the discriminator, calculate the loss function of the discriminator, obtain the judgment result of the synthetic samples, update the discriminator parameters through the backpropagation algorithm to make the discriminator better distinguish between small samples and synthetic samples. Utilize the judgment result of the synthetic samples to calculate the loss function of the generator, and update the generator parameters through the backpropagation algorithm to make the synthetic samples generated by the generator more capable of deceiving the discriminator, improving the probability that the synthetic samples are close to the small samples. Repeat the alternating training of the generator and the discriminator, and through multiple rounds of iteration, obtain synthetic samples similar to the small samples.

[0115] Preferably, as Figure 2 shown, in step three, using the asynchronous thread scheduling of the central server, through time-weighted average and weight asynchronous aggregation, obtain the global model. The process of selecting the way to update the global model according to the computing power of the central server includes:

[0116] A1. The central server consists of a scheduling thread, an update thread, and a coordinator. Among them, the tasks of the scheduling thread and the update thread are controlled by two asynchronous parallel threads. The coordinator regularly sends scheduling instructions and the global model to the MEC edge devices. When the edge devices complete their work, they upload the updated global model parameters. The central server obtains the upload queue of the global model through the coordinator and updates the global model in sequence;

[0117] For the kth device, the central server records the central server scheduling timestamp and the central server global model update timestamp through the asynchronous thread scheduling process, and calculates the central server waiting delay time using the timestamps:

[0118]

[0119] where, is the central server waiting delay time, is the central server global model update timestamp, is the central server scheduling timestamp;

[0120] A2. Adopt time-weighted average, set a delay threshold. When the feedback parameter of the edge device is higher than the waiting delay time of the central server, the global model update is lagged. Set a hybrid hyperparameter, which is used to control the impact of delay on the global model update. Represent the delay function through a hinge function, and use the delay function and the hybrid hyperparameter to calculate the hybrid hyperparameter in the t-th round of iteration as follows:

[0121]

[0122]

[0123] Among them, is the delay function, is the hybrid hyperparameter in the t-th round of iteration, a and b are constants, and a > 0, b > 0, is the time delay magnitude exceeding the server's expected waiting time, is the delay threshold;

[0124] In the t-th round of iteration, when the central server receives the local training model parameters feedback by the k-th edge device, the process of calculating the time-weighted average to update the global model is as follows:

[0125]

[0126] Among them, is the time-weighted average to update the global model in the t-th round of iteration, is the time-weighted average to update the global model in the (t - 1)-th round of iteration, is the hybrid hyperparameter in the t-th round of iteration, is the local training model parameters feedback by the k-th edge device received by the central server;

[0127] A3. Adopt weighted asynchronous aggregation. According to the local data of the k-th edge device scheduled by the central server in the t-th round of iteration, set the weight parameter of the k-th edge device scheduled in the t-th round of iteration. The process of calculating the weighted asynchronous aggregation to update the global model is as follows:

[0128]

[0129]

[0130] Among them, is the weighted asynchronous aggregation to update the global model in the t-th round of iteration, is the weighted asynchronous aggregation to update the global model in the (t - 1)-th round of iteration, is the weight hyperparameter in the t-th round of iteration, is the weight parameter of the k-th edge device scheduled in the t-th round of iteration, is the hybrid hyperparameter for the t-th round of iteration. is the local training model parameter fed back by the k-th edge device received by the central server.

[0131] A4. Deploy the intelligent platform management interface device and the Geekbench test tool on the edge device. Among them, the intelligent platform management interface device measures the performance metrics of the CPU, and the Geekbench test tool uses the performance metrics of the CPU to obtain the computing power value of the central server, set the computing power threshold of the central server, and select the global model update method according to the central server computing power threshold. When the computing power value of the central server is less than the central server computing power threshold, use time-weighted average to update the global model; when the computing power value of the central server is greater than the central server computing power threshold, use weighted asynchronous aggregation to update the global model.

[0132] Preferably, in step four, the central server transmits the updated global model and the scheduling instruction to the edge device according to the running state of the edge device. The process of the edge device training the global model according to the scheduling instruction includes:

[0133] Among them, the running state of the edge device includes the idle state, the running state, and the blocked state. The edge device receives the updated global model and the scheduling instruction.

[0134] When the edge device is in the idle state, the edge device enters the running state, updates the global model using the global model and the available local sample data on the edge device, and transmits the updated global model parameters to the central server.

[0135] When the edge device is in the blocked state, switch the state according to the usage of the available local sample data of the edge device. During the running process, when there are situations such as insufficient available local sample resources, network interruption, and power exhaustion of the edge device, suspend the current training task of the edge device, and wait until the blocked state of the edge device switches to the idle state before continuing to train the global model; during the running process, when there are situations such as the edge device being unavailable and the central server rejecting communication, the central server scheduling thread adjusts edge , and wait until the edge device is available and communicating normally before training the global model again.

[0136] Preferably, as Figure 3 shown, in step five, the process of constructing a reinforcement learning model based on Markov decision using reinforcement learning, Markov decision, and random forest algorithms includes:

[0137] Select the t-th moment, and the reinforcement learning agent will represent the observation state, represent the execution action, represent the state transition probability, represent the reward, and represent the discount factor. After executing an action, the reinforcement learning agent transitions states according to the state transition probability, and uses to represent the transitioned state. The reinforcement learning agent obtains a reward;

[0138] Build a random forest model using the random forest algorithm, set the parameters of the random forest model, use the transition state and its reinforcement learning agent reward as the data set, divide it into a training set and a test set according to the ratio of 7:3, use the training set data to train the random forest model, through iterative learning, obtain the non-linear relationship between the transition state and the reinforcement learning agent reward, realize that the random forest model inputs the transition state and outputs the corresponding reinforcement learning agent reward, input the test set data into the trained random forest model, obtain the output value of the trained random forest, compare the reinforcement learning agent reward corresponding to the actual transition state with the output value of the trained random forest model, adjust the parameters of the random forest model, optimize the performance of the random forest model, and obtain the reinforcement learning model based on Markov decision-making.

[0139] Preferably, in step six, the process of constructing a deep reinforcement model by combining the Q-learning algorithm, function approximator, and neural network algorithm through the reinforcement learning model based on Markov decision-making includes:

[0140] B1. Through the reinforcement learning model based on Markov decision-making, obtain the transitioned state and its reinforcement learning agent reward, establish a transition state-reward function, and the process of obtaining the maximum reward value of the initial state through the Q-learning algorithm is as follows:

[0141]

[0142]

[0143] where is the maximum reward value function of the initial state, also known as the maximum expected value function of the cumulative total return obtained starting from state based on policy ; is the transition state-reward function; is the transitioned state-maximum reward value function after transition; E is the mathematical expectation; is the policy mapping from the initial state to the selected action; represents the reward; represents the discount factor; represents the observed state, represents the executed action, and Specific values representing the observation state and execution actions respectively;

[0144] Use a function approximator to learn the function , approximating the maximum reward value function from the initial state to the transformed state , where represents the policy parameters of the current state;

[0145] B2. Using the neural network algorithm, construct a neural network model. Take the initial state and its maximum reward value, and the transformed state and its maximum reward value as the data set, and divide it into a training set and a test set according to the ratio of 7:3. Select MLP as the neural network structure. The input layer includes two neurons, receiving the initial state and the transformed state. The hidden layer configures the MSE function. The output layer includes two neurons, outputting the maximum reward value of the initial state and the maximum reward value of the transformed state;

[0146] Input the training set data into the neural network model, set the learning rate to 0.01, and the number of iterative training times to 1000. The training process includes forward propagation and backward propagation. Among them, forward propagation is used to calculate the predicted output data, and backward propagation is used to update the weights and biases of the model. Through repeated iterative training, learn the non-linear relationship between the initial state and the maximum reward value of the initial state, and the non-linear relationship between the transformed state and the maximum reward value of the transformed state until the set number of iterative training times is reached, and obtain the trained neural network model;

[0147] Input the test set data into the trained neural network model, use the MSE function to evaluate the error between the output value and the actual value of the neural network model, adjust the parameters of the neural network model according to the evaluation results, optimize the performance of the neural network model, and obtain the deep reinforcement model;

[0148] Use the deep reinforcement model to obtain the maximum reward value of the initial state and the maximum reward value of the transformed state, and use the maximum reward value of the transformed state to update the maximum reward value of the initial state until the last state ends.

[0149] Preferably, in step seven, deploy a DRL agent, define the DRL agent state, select the multi-terminal device CPU action through the DRL agent and set the reward mechanism. The process of constructing a joint reinforcement learning model through the random forest algorithm includes:

[0150] Install a network tester on the multi-terminal device to test the real-time network bandwidth of the multi-terminal device. Deploy the DRL agent on the central server to realize the interaction between the DRL agent and the multi-terminal device. Take the network bandwidth of each multi-terminal device as the state space of the DRL agent, expressed as , set a network bandwidth threshold. When the real-time network bandwidth of the multi-terminal device is lower than the network bandwidth threshold, select to increase the CPU frequency of the multi-terminal device; when the real-time network bandwidth of the multi-terminal device is higher than the network bandwidth threshold, select to decrease the CPU frequency of the multi-terminal device;

[0151] After the DRL agent executes an action in the initial state in the t-th iteration, obtain a learning reward. The process of setting the reward mechanism is as follows:

[0152]

[0153]

[0154] Among them, is the reward function, T is the total number of iterations, t is the current number of iterations, represents the discount factor, is the time for the DRL agent to obtain the maximum device training model, is when in the state the expected cumulative discounted reward obtained by the DRL agent;

[0155] Use the random forest algorithm to construct a random forest model, set the random forest model parameters. Take the actions in the initial state in the t-th iteration, their corresponding reward functions, and the DRL agent's expected cumulative discounted rewards as the data set, and divide it into a training set and a test set according to the ratio of 7:3. Use the training set data to train the random forest model. Through iterative learning, obtain the non-linear relationship between the actions in the initial state in the t-th iteration, their corresponding reward functions, and the DRL agent's expected cumulative discounted rewards, and realize that the random forest model inputs the actions in the initial state in the t-th iteration and outputs the non-linear relationship between the corresponding reward functions and the DRL agent's expected cumulative discounted rewards. Input the test set data into the trained random forest model to obtain the output value of the trained random forest. Compare the actual actions in the initial state in the t-th iteration, their corresponding reward functions, and the DRL agent's expected cumulative discounted rewards with the output value of the trained random forest model, adjust the random forest model parameters, optimize the performance of the random forest model, and obtain a joint reinforcement learning model;

[0156] Use the joint reinforcement learning model to obtain the reward function and the DRL agent's expected cumulative discounted rewards corresponding to the actions in the initial state in the t-th iteration. Use this reward function and the DRL agent's expected cumulative discounted rewards to minimize the time cost of multi-terminal computing.

[0157] Preferably, as Figure 4 shown, in step eight, the process of combining and applying the reinforcement learning model based on Markov decision and the deep reinforcement model in the joint optimization algorithm based on reinforcement learning includes:

[0158] Combine the reinforcement learning model based on Markov decision and the deep reinforcement model, and use the DRL agent for iterative training in the heterogeneous MEC network with multi-terminal collaboration. The workflow is as follows:

[0159] C1. Initialize the MEC parameters and the DRL agent;

[0160] C2. The multi-terminal devices download the global model from the central server;

[0161] C3. The DRL agent observes the heterogeneous network bandwidth and updates the state space;

[0162] C4. The DRL agent selects an action , executes the action of controlling the CPU frequency. The multi-terminal devices obtain and update the global model parameters according to the deep reinforcement model, update the global model, and upload it to the central server. The central server aggregates and updates the global model;

[0163] C5. The DRL agent enters the next state , and obtains the reward function corresponding to the action of the initial state and the expected cumulative discounted reward of the DRL agent in the t-th iteration through the joint reinforcement learning model;

[0164] C6. After the end of an iteration cycle, repeat steps C2 to C5 until the training iteration rounds are completed and the training expected effect is achieved.

[0165] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. Method for improving MEC joint computing performance in multi - terminal heterogeneous data scenarios, characterized in that: It includes the following steps: Step 1: Design an MEC heterogeneous data system model, record the operation data of homogeneous multi-terminal devices, and divide the available local sample data sets of edge devices according to the Pareto principle; Step 2: According to the divided available local sample data sets of edge devices, adopt corresponding training algorithm allocation strategies; Step 3: Use the asynchronous thread scheduling of the central server, obtain the global model parameters through time-weighted average and weight asynchronous aggregation, use the intelligent platform management interface device and Geekbench test tool to obtain the computing power value of the central server, and select the global model update method according to the computing power of the central server; Step 4: The central server sends the updated global model and scheduling instructions to the edge devices according to the operation status of the edge devices, and the edge devices train the global model according to the scheduling instructions; Step 5: Use reinforcement learning, Markov decision-making, and random forest algorithms to construct a reinforcement learning model based on Markov decision-making; Step 6: Through the reinforcement learning model based on Markov decision-making, combine the Q-learning algorithm, function approximator, and neural network algorithm to construct a deep reinforcement model; Step 7: Deploy a DRL agent, define the DRL agent state, select the CPU actions of multi-terminal devices through the DRL agent and set a reward mechanism, and construct a joint reinforcement learning model through the random forest algorithm; Step 8: Based on the joint optimization algorithm of reinforcement learning, combine and apply the reinforcement learning model based on Markov decision-making and the deep reinforcement model.

2. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 1, wherein: In the above Step 1, the process of designing an MEC heterogeneous data system model, recording the operation data of homogeneous multi-terminal devices, and dividing the available local sample data sets of edge devices according to the Pareto principle includes: The MEC heterogeneous data system model includes a central server and k MEC edge nodes. Each MEC edge node includes edge devices and local sample data sets, and a federated learning training model is deployed on each MEC edge node; The operation data of the isomorphic multi-terminal devices includes the available local sample data set on the k-th edge device as , the local computing rate of the k-th edge device, and the communication time between the central server and the k-th edge device, and record the operation data of the above-mentioned isomorphic multi-terminal devices; The process of calculating the waiting time required for the t-th round of update of the server is as follows: ; ; Among them, is the time required for the server to wait for the t-th round of update, is the local update training time of the k-th edge device, is the available local sample dataset on the k-th edge device, is the local computing rate of the k-th edge device, is the communication time between the central server and the k-th edge device; Taking the edge device number k as the abscissa and the available local sample data set on the k-th edge device as the ordinate, plot the power-law distribution diagram of data samples; Select the available local minimum sample data set on the k-th edge device through the power-law distribution diagram of data samples and the available local maximum sample data set on the k-th edge device , set the threshold of the available local sample data set on the k-th edge device , satisfying . According to the Pareto principle, the available local sample data sets of edge devices greater than are regarded as large samples, and the available local sample data sets of edge devices less than are regarded as small samples.

3. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 2, wherein: In the above Step 2, the process of adopting corresponding training algorithm allocation strategies according to the divided available local sample data sets of edge devices includes: When the available local sample data set of the edge device is greater than , determine the gradient descent condition. The corresponding available local sample data set of the edge device is a large sample. In the MEC network based on the FedAvg algorithm, select small data samples for training, perform stochastic gradient descent calculation, judge the improvement of the available local sample size on the target loss function, and adjust the size of the available local samples until the gradient descent condition is met; When the available local sample dataset of the edge device is less than At this time, the available local sample dataset of the corresponding edge device is a small sample. Through the generative adversarial network, synthetic samples similar to the small sample data are obtained, the small sample and the synthetic samples are combined to obtain combined samples, and the FedAvg algorithm is used to perform local gradient descent on the combined samples.

4. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 3, characterized in that: The principles of the FedAvg algorithm and the generative adversarial network include: S1. Using the available local sample dataset on the edge device, establish a locally available sample subset S, and select the size of the small samples in the locally available sample subset. The process of calculating the minimized objective loss function is as follows: , as follows: ; Among them, is the convex empirical loss objective function, is the size of the small samples in the local available sample subset, is the local available sample subset, and there is , is the prediction function calculated by gradient descent through the parameter ω; Determine the gradient descent conditions and analyze the improvement of the objective loss function to be minimized. If the improvement remains unchanged, select the local available sample data with the same sample size as that for the next iteration. If the improvement changes, re-select a sample larger than that and perform the next iteration until the gradient descent conditions are met; S2. The generative adversarial network consists of a generator, a discriminator, and a loss function. Initialize the neural network parameters of the generator and the discriminator, extract real samples from small samples, the generator synthesizes synthetic samples similar to the small samples, input the small samples and the synthetic samples into the discriminator, calculate the loss function of the discriminator, obtain the judgment result of the synthetic samples, update the discriminator parameters through the backpropagation algorithm, use the judgment result of the synthetic samples to calculate the loss function of the generator, update the generator parameters through the backpropagation algorithm, repeat the alternating training of the generator and the discriminator, and obtain synthetic samples similar to the small samples through multiple rounds of iteration.

5. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 4, characterized in that: In the above Step 3, the process of using the asynchronous thread scheduling of the central server, obtaining the global model through time-weighted average and weight asynchronous aggregation, and selecting the global model update method according to the computing power of the central server includes: A1. The central server consists of a scheduling thread, an update thread, and a coordinator. Among them, the scheduling thread task and the update thread task are controlled by two asynchronous parallel threads. The coordinator periodically sends scheduling instructions and the global model to the MEC edge devices. When the edge devices complete their work, they upload and update the global model parameters. The central server obtains the upload queue of the global model through the coordinator and updates the global model sequentially; For the k-th device, the central server records the central server scheduling timestamp and the central server global model update timestamp through the asynchronous thread scheduling process, and calculates the central server waiting delay time using the timestamps: ; Among them, is the waiting delay time of the central server, is the global model timestamp updated by the central server, is the scheduling timestamp of the central server; A2. Using time-weighted average, a delay threshold is set. When the edge device feedback parameter is higher than the central server waiting delay time, the global model update is lagged. A mixed hyperparameter is set, and the mixed hyperparameter is used to control the impact of the delay on the global model update. The delay function is represented by a hinge function. The process of calculating the mixed hyperparameter for the t-th iteration using the delay function and the mixed hyperparameter is as follows: ; ; Among them, is a delay function, is the mixed hyperparameter for the t-th round of iteration, where a and b are constants, and a > 0, b > 0, is the delay magnitude exceeding the expected waiting time of the server, is the delay threshold; In the t-th iteration, when the central server receives the local training model parameters fed back by the k-th edge device, the process of calculating the time-weighted average to update the global model is as follows: ; Among them, is the time-weighted average to update the global model in the t-th round of iteration, is the time-weighted average to update the global model in the (t-1)-th round of iteration, is the hybrid hyperparameter in the t-th round of iteration, is the local training model parameter feedback received by the central server from the k-th edge device; A3. Using weighted asynchronous aggregation, according to the local data of the k-th edge device scheduled by the central server in the t-th iteration, the weight parameter of the k-th edge device scheduled in the t-th iteration is set, and the process of calculating the weighted asynchronous aggregation to update the global model is as follows: ; ; Among them, is the global model updated by asynchronous aggregation of weights in the t-th round of iteration, is the global model updated by asynchronous aggregation of weights in the (t - 1)-th round of iteration, is the weight hyperparameter in the t-th round of iteration, is the weight parameter of the k-th edge device scheduled in the t-th round of iteration, is the hybrid hyperparameter in the t-th round of iteration, is the local training model parameter fed back by the central server from the k-th edge device; A4. Deploy an intelligent platform management interface device and a Geekbench test tool on the edge device. Among them, the intelligent platform management interface device measures the performance metrics of the CPU, and the Geekbench test tool uses the performance metrics of the CPU to obtain the computing power value of the central server. A central server computing power threshold is set. According to the central server computing power threshold, the global model update method is selected. When the computing power value of the central server is less than the central server computing power threshold, time-weighted average is used to update the global model; when the computing power value of the central server is greater than the central server computing power threshold, weighted asynchronous aggregation is used to update the global model.

6. The method for improving the MEC joint computing performance in a multi - terminal heterogeneous data scenario according to claim 5, characterized in that: In step four above, the process by which the central server transmits the updated global model and scheduling instructions to the edge device according to the operating state of the edge device, and the edge device trains the global model according to the scheduling instructions includes: The operating states of the edge device include the idle state, the running state, and the blocked state. The edge device receives the updated global model and scheduling instructions; When the edge device is in the idle state, the edge device enters the running state, updates the global model using the global model and the available local sample data on the edge device, and transmits the updated global model parameters to the central server; When the edge device is in a blocked state, switch the state according to the available local sample data usage of the edge device. During operation, when there are situations such as insufficient available local sample resources, network interruption, and power exhaustion of the edge device, suspend the current training task of the edge device and wait until the blocked state of the edge device switches to the idle state, then continue to train the global model; during operation, when there are situations where the edge device is unavailable and the central server refuses to communicate, the central server scheduling thread adjusts edge , and wait until the edge device is available and communicating normally, and then train the global model.

7. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 6, wherein: In step five above, the process of constructing a reinforcement learning model based on Markov decision using reinforcement learning, Markov decision, and random forest algorithms includes: At time t, the reinforcement learning agent uses to represent the observation state, to represent the executed action, to represent the state transition probability, to represent the reward, and to represent the discount factor. After executing the action, the reinforcement learning agent transitions the state according to the state transition probability and uses to represent the transitioned state. After the transition is completed, the reinforcement learning agent obtains the reward; A random forest model is constructed using the random forest algorithm. The parameters of the random forest model are set. The conversion state and its reinforcement learning agent reward are used as the data set, which is divided into a training set and a test set according to a ratio of 7:

3. The training set data is used to train the random forest model. Through iterative learning, the non-linear relationship between the conversion state and the reinforcement learning agent reward is obtained, enabling the random forest model to input the conversion state and output the corresponding reinforcement learning agent reward. The test set data is input into the trained random forest model to obtain the output value of the trained random forest. The reinforcement learning agent reward corresponding to the actual conversion state is compared with the output value of the trained random forest model, and the parameters of the random forest model are adjusted to optimize the performance of the random forest model, thereby obtaining a reinforcement learning model based on Markov decision-making.

8. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 7, wherein: In step six, the process of constructing a deep reinforcement model by combining the Q-learning algorithm, function approximator, and neural network algorithm through the reinforcement learning model based on Markov decision-making includes: B1. Through the reinforcement learning model based on Markov decision-making, the converted state and its reinforcement learning agent reward are obtained, and a conversion state-reward function is established. The process of obtaining the maximum reward value of the initial state through the Q-learning algorithm is as follows: ; ; Among them, is the maximum reward value function in the initial state, also known as the maximum expected value function of the cumulative total reward obtained starting from state based on policy ; is the transition state-reward function; is the maximum reward value function of the state after transition; E is the mathematical expectation; is the policy mapping from the initial state to the selected action; represents the reward; represents the discount factor; represents the observed state, represents the executed action, and respectively represent the specific values of the observed state and the executed action; Use a function approximator to learn a function that approximates the maximum reward value function from the initial state to the transformed state , where represents the policy parameter of the current state; B2. Using the neural network algorithm, a neural network model is constructed. The initial state and its maximum reward value, as well as the converted state and its maximum reward value, are used as the data set, which is divided into a training set and a test set according to a ratio of 7:

3. MLP is selected as the neural network structure. The input layer includes two neurons that receive the initial state and the converted state. The hidden layer is configured with the MSE function. The output layer includes two neurons that output the maximum reward value of the initial state and the maximum reward value of the converted state. The training set data is input into the neural network model. The learning rate is set to 0.01, and the number of iterative training times is 1000. The training process includes forward propagation and backward propagation. Among them, forward propagation is used to calculate the predicted output data, and backward propagation is used to update the weights and biases of the model. Through repeated iterative training, the non-linear relationship between the initial state and the maximum reward value of the initial state, as well as the non-linear relationship between the converted state and the maximum reward value of the converted state, is learned until the set number of iterative training times is reached, and the trained neural network model is obtained. The test set data is input into the trained neural network model. Using the MSE function, the error between the output value of the neural network model and the actual value is evaluated. According to the evaluation results, the parameters of the neural network model are adjusted to optimize the performance of the neural network model, and a deep reinforcement model is obtained. Using the deep reinforcement model, the maximum reward value of the initial state and the maximum reward value of the converted state are obtained. The maximum reward value of the converted state is used to update the maximum reward value of the initial state until the last state ends.

9. The MEC joint computing performance improvement method in the multi-terminal heterogeneous data scenario according to claim 8, wherein: In step seven, the DRL agent is deployed, the DRL agent state is defined, and the DRL agent selects the CPU actions of multi-terminal devices and sets the reward mechanism. The process of constructing a joint reinforcement learning model through the random forest algorithm includes: Install a network tester on the multi - terminal device to test the real - time network bandwidth of the multi - terminal device. Deploy the DRL agent on the central server, and use the network bandwidth of each multi - terminal device as the state space of the DRL agent, denoted as , set a network bandwidth threshold. When the real - time network bandwidth of the multi - terminal device is lower than the network bandwidth threshold, choose to increase the CPU frequency of the multi - terminal device; when the real - time network bandwidth of the multi - terminal device is higher than the network bandwidth threshold, choose to decrease the CPU frequency of the multi - terminal device; After the DRL agent executes an action in the initial state in the t-th iteration, it obtains a learning reward. The process of setting the reward mechanism is as follows: ; ; Among them, is the reward function, T is the total number of iterations, t is the current iteration number, represents the discount factor, is the maximum time for the DRL agent to obtain the device training model, is when in the state the expected cumulative discounted reward obtained by the DRL agent; Use the random forest algorithm to construct a random forest model, set the parameters of the random forest model. Take the actions in the initial state in the t-th iteration, their corresponding reward functions, and the expected cumulative discounted rewards of the DRL agent as the data set, and divide it into a training set and a test set according to the ratio of 7:

3. Use the training set data to train the random forest model. Through iterative learning, obtain the non-linear relationship between the actions in the initial state in the t-th iteration, their corresponding reward functions, and the expected cumulative discounted rewards of the DRL agent, and realize that the random forest model inputs the actions in the initial state in the t-th iteration and outputs the non-linear relationship between the corresponding reward functions and the expected cumulative discounted rewards of the DRL agent. Input the test set data into the trained random forest model to obtain the output values of the trained random forest. Compare the actual actions in the initial state in the t-th iteration, their corresponding reward functions, and the expected cumulative discounted rewards of the DRL agent with the output values of the trained random forest model, adjust the parameters of the random forest model, optimize the performance of the random forest model, and obtain the joint reinforcement learning model; Use the joint reinforcement learning model to obtain the reward function corresponding to the action in the initial state in the t-th iteration and the expected cumulative discounted rewards of the DRL agent. Use this reward function and the expected cumulative discounted rewards of the DRL agent to minimize the time cost of multi-terminal computing.

10. The method for improving the MEC joint computing performance in a multi-terminal heterogeneous data scenario according to claim 9, characterized in that: In step eight, the process of combining and applying the reinforcement learning model based on Markov decision-making and the deep reinforcement model in the joint optimization algorithm based on reinforcement learning includes: Combine the reinforcement learning model based on Markov decision-making and the deep reinforcement model, and use the DRL agent for iterative training in the heterogeneous MEC network with multi-terminal cooperation. Its working process is as follows: C1. Initialize the MEC parameters and the DRL agent; C2. The multi-terminal devices download the global model from the central server; C3. The DRL agent observes the heterogeneous network bandwidth and updates the state space; C4. The DRL agent selects an action , performs an action to control the CPU frequency. The multi-terminal device obtains and updates the global model parameters according to the deep reinforcement model, updates the global model, and uploads it to the central server. The central server aggregates and updates the global model; C5. The DRL agent enters the next state , and through the joint reinforcement learning model, obtain the reward function corresponding to the action in the initial state and the expected cumulative discounted reward of the DRL agent in the t-th iteration; C6. After the end of an iteration cycle, repeat steps C2 to C5 until the training iteration rounds are completed and the training expected effect is achieved.