Method and device for adaptive task migration based on end-to-end cloud segmentation of internet of vehicles
By constructing a multi-master, multi-slave Stackelberg game model and a MADDPG model in the Internet of Vehicles (IoV) to optimize task segmentation and resource allocation, the problems of low latency and high computational load in task migration in IoV are solved, resource utilization and communication efficiency are improved, and user experience is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAQIAO UNIVERSITY
- Filing Date
- 2023-06-15
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, task migration in vehicle-to-everything (V2X) networks faces challenges of low latency and high computational load, and the lack of effective incentive mechanisms leads to underutilization of vehicle resources. Existing research has failed to effectively optimize communication and computing resources.
An edge-cloud architecture is constructed using a multi-master, multi-slave Stackelberg game model. By reserving idle computing resources in resource-rich vehicles and combining deep deterministic policy reinforcement learning (MADDPG model) to optimize task segmentation and resource allocation, adaptive migration of tasks between cloud servers, edge nodes, and vehicles is achieved.
It improves the resource utilization of the edge-cloud architecture, reduces communication latency, can adaptively cope with task requirements of different scales and complexities, and enhances the user experience.
Smart Images

Figure CN116760712B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile edge computing, and more specifically to a migration method and apparatus for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) networks. Background Technology
[0002] With the rapid development of intelligent transportation and the exponential growth of data traffic in the Internet of Vehicles (IoV), vehicle terminal request migration tasks face enormous challenges in terms of low latency and high computational demands. While migrating tasks to edge servers or cloud servers is considered an effective solution, the limited resources of edge servers and the long distances of cloud servers restrict their performance. Some studies suggest that autonomous vehicles with strong computing capabilities have the potential to become service providers, but due to the selfishness of vehicles, there is a lack of incentive mechanisms to reserve resources for them. Modern software typically employs a modular design, and most computational tasks can be broken down into multiple smaller tasks, allowing for distributed processing and potentially reducing communication latency. However, existing research largely focuses on complete task migration and does not consider the joint optimization of communication and computing resources. Summary of the Invention
[0003] In view of the aforementioned technical problems, the purpose of the embodiments of this application is to propose a migration method and apparatus for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) technology, thereby solving the technical problems mentioned in the background section.
[0004] In a first aspect, the present invention provides a migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, comprising the following steps:
[0005] Construct and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers. Obtain the idle computing resources reserved for vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture.
[0006] Obtain the vehicle's computing migration request, divide the vehicle's computing task into several sub-tasks according to the task partitioning ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively.
[0007] The task segmentation ratio, computing resource allocation, and communication resource allocation of vehicles are modeled as a multi-constraint problem. The maximum sum of the maximum latency of each subtask of the vehicle under different migration modes is minimized. A MADDPG model is constructed based on the multi-constraint problem. The MADDPG model is trained in a centralized manner to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent. The environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions. The actions include the best migration mode, communication resources, and computing resources selected by each agent. The reward is the latency performance of the agent. The state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient.
[0008] The vehicle's status is obtained and input into the trained MADDPG model to obtain the optimal action. Based on the optimal action, the migration node is selected, and the migration task ratio is determined.
[0009] As a preferred approach, a multi-master, multi-slave Stackelberg game model is constructed and solved to obtain the idle computing resources reserved by the resource-rich vehicles, specifically including:
[0010] Multiple cloud servers publicly announce their unit computing prices and notify vehicles with abundant computing resources of this strategy. After receiving the unit computing prices from different cloud servers, vehicles with abundant computing resources compete with each other to determine which cloud server provides reserved resources. The iteration time for vehicles with abundant computing resources to reach Nash equilibrium is one scheduling cycle of the cloud server. The final result of the cloud server iteration is that all cloud servers and vehicles with abundant computing resources have reached Nash equilibrium. Under this Nash equilibrium, the vehicles obtain the idle computing resources reserved by the vehicles with abundant computing resources.
[0011] Preferably, the vehicle's computing tasks are divided into several subtasks according to the task partitioning ratio based on the migration request, and these subtasks are migrated to migration nodes in the edge-cloud architecture, specifically including:
[0012] The computational task for the collected i-th vehicle uses ψ i ={H i F i , t i Represented as a triple, H i F represents the amount of data required for the computation task of the i-th vehicle. i Let t represent the computational resources required for the computational task of the i-th vehicle. i Let ψ represent the maximum tolerable latency for the computation task of the i-th vehicle; let ψ represent the computation task of the i-th vehicle. i Divide into n subsets {ψ i,1 , ψ i,2 ,..., ψi,n Each subset contains a certain proportion of subtasks. In the edge-cloud architecture, N is defined. cloud Let the set of cloud servers be D, N edge Let P be the set of edge nodes, and N be the set of edge nodes. crv Let E be a set of resource-rich vehicles with reserved idle computing resources, and let subtask ψ be... i,j The subtask ψ is assigned to a migration node X∈D∪P∪E, and processed by the migration node. i,j .
[0013] Preferably, the task partitioning ratio, computational resource allocation, and communication resource allocation of the vehicle are modeled as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes, specifically including:
[0014] Within each scheduling cycle, the task of the i-th vehicle is split and then distributed among the selected migration nodes X, assuming the allocation scheme is S = {(X1, ψ...} i,1 ), (X2, ψ i,2 ), ..., (X n , ψ i,n If the total computation time for allocation scheme S is , then the total computation time for allocation scheme S is .
[0015] t i =max(t(X1, ψ) i,1 ), t(X2, ψ i,2 )...,t(X n , ψ i,n )) (1)
[0016] The optimization problem can then be expressed by the following formula:
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024]
[0025]
[0026] Where I represents the total number of vehicles in the Internet of Things, constraint (2a) indicates that the sum of task partitioning is 1, and B, A, and R represent the strategies for migration task ratio, computing resources, and communication resources, respectively. These represent the proportions of tasks migrated to cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively. in These represent the computing resources required by cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively. This indicates the channel selection for the cloud server. This represents the channel selection at edge nodes and resource-rich vehicles with reserved idle computing resources; constraint (2b) ensures that the size of the task split is greater than the minimum divisible size G of the task; constraints (2c), (2d), and (2e) represent the maximum computing resource limits of cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively; constraint (2f) indicates that the execution time when migrating to an edge node should be less than the maximum tolerance time of the task and the minimum communication time with it; constraint (2g) indicates that the execution time when migrating to a vehicle with abundant computing resources should be less than the maximum tolerance time of the task and the minimum communication time with it.
[0027] For the communication resource allocation in the edge-cloud architecture, there are M remote wireless heads, and the cloud server corresponds to a resource block set R. c ={1, ..., r c The set of resource blocks R corresponding to edge nodes and resource-rich vehicles with reserved idle computing resources. f ={1, ..., r f}, where each resource block has the same bandwidth W, using binary variables. Let R represent the resource allocation scheme for the i-th vehicle, where j∈R c , k∈R f ;like This indicates that the j-th resource block is allocated to the i-th vehicle. This indicates that the k-th resource block is allocated to the i-th vehicle;
[0028] The latency of the i-th vehicle migrating the task to the cloud server d The calculation formula is:
[0029]
[0030]
[0031] Among them, V cloud,d p represents the migration rate on cloud server d. i This represents the transmission power of the i-th vehicle. Let Ni represent the channel gain between the i-th vehicle and the m-th remote wireless head, and N0 represent the noise spectral density. t represents the proportion of tasks migrated to the cloud server. wired F represents the wired propagation delay. cloud,d It is the computing resources provided by cloud server d;
[0032] The latency for the i-th vehicle to migrate the task to the edge node p The calculation formula is:
[0033]
[0034]
[0035] Among them, V edge,p This represents the migration rate at edge node p. This represents the proportion of tasks that migrate to edge node p. F represents the channel gain between the j-th vehicle and the p-th edge node. edge,p This represents the computing resources provided by the edge node p;
[0036] The latency of the i-th vehicle migrating the task to the computationally abundant vehicle e. The calculation formula is:
[0037]
[0038]
[0039] Among them, V crv,e This represents the migration rate of vehicle e, which has abundant computing resources. F represents the proportion of tasks migrated to vehicle e with reserved idle computing resources. crv,e This indicates the computing resources reserved for the resource-rich vehicle e.
[0040] Preferably, the evaluation network and target network for each agent both include a Critic network and an Actor network.
[0041] The state of the i-th agent at time t is defined as follows:
[0042] s t =(h t , τ t I t-1 O t-1 ,∈,ρ) (9)
[0043] Among them, h t This represents the local channel information at time t, including the channel gain between the i-th agent and the remote wireless head at time t. Channel gain between agent i and edge node Channel gain between the i-th agent and a resource-rich vehicle with reserved idle computing resources. τ t This represents the dwell time at time t, including the dwell time of the i-th agent and the edge node at time t. The dwell time of the i-th agent and the resource-rich vehicle with reserved idle computing resources. Network status information at time t-1 includes wireless interference: Remaining computing resources of cloud servers and edge nodes: ∈ represents the number of iterations, and ρ represents the exploration coefficient;
[0044] The action of the i-th agent is defined as:
[0045]
[0046] The reward obtained by the i-th agent in the environment is defined as:
[0047] r t =α(T) i,max -T i )-βr sat (11)
[0048] Where α and β are the weights of each part, and T i,max T represents the maximum tolerable delay for the i-th agent. i Indicates the actual time delay used, r sat The delay non-satisfaction rate for all agents is calculated by dividing the total number of non-satisfying vehicles by the total number of vehicles.
[0049] As a preferred approach, the MADDPG model is trained intensively, specifically including:
[0050] Initialize the parameter weights θ of the Actor network and Critic network of the agent's evaluation network. μ θ Q The parameter weights θ of the Actor network and Critic network of the target network μ′ θ Q′ And initialize the environment and experience pool R;
[0051] Reserved computing resources in vehicles with abundant computing resources are used to observe the state obtained by the i-th agent from the environment.
[0052] Perform actions based on the state observed by the agent. Obtain the state at the next moment after performing the action. All agents have completed their actions and received rewards from the environment.t ,storage The data is fed into the experience pool R, and the Actor and Critic networks of the evaluation network are trained using the data in the experience pool R. The parameters of the evaluation network are then updated to the Actor and Critic networks of the target network. After the loop is completed, the trained MADDPG model is obtained.
[0053] The Critic network in the MADDPG model minimizes the root mean square loss using gradient descent, with the loss function being:
[0054]
[0055] Where ε is the discount rate;
[0056] The Actor network updates its gradient using the following formula:
[0057]
[0058] The parameters θ of the target network μ′ and θ Q′ Use soft update:
[0059] θ Q′ ←τθ Q +(1-τ)θ Q′ (14)
[0060] θ μ′ ←τθ μ +(1-τ)θ μ′ (15)
[0061] When updating network parameters, batch sampling is performed, and during training, the network state information from the previous time step is recorded and added to the multi-agent state.
[0062] Preferably, a custom ReLU activation function is used in the output layers of the Critic network and the Actor network to filter out tasks below a threshold in the output actions, and the filtered actions are then input into the Softmax function.
[0063] Secondly, the present invention provides a migration device for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, comprising:
[0064] The resource computing module is configured to build and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers, obtain the idle computing resources reserved by the vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture.
[0065] The task splitting module is configured to obtain the vehicle's computing migration request, split the vehicle's computing task into several sub-tasks according to the task splitting ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively.
[0066] The model training module is configured to model the task segmentation ratio, computing resource allocation, and communication resource allocation of the vehicle as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes. Based on the multi-constraint problem, a MADDPG model is constructed, and the MADDPG model is trained centrally to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent, the environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions, the actions include the best migration mode, communication resources, and computing resources selected by each agent, the reward is the latency performance of the agent, and the state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient.
[0067] The execution module is configured to acquire the vehicle's status, input it into the trained MADDPG model, obtain the optimal action, select the migration node based on the optimal action, and determine the migration task ratio.
[0068] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0069] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] (1) The migration method of the edge-cloud adaptive segmentation task based on the Internet of Vehicles proposed in this invention uses a multi-master multi-slave Stackelberg game model to solve the reserved idle computing resources of vehicles with abundant computing resources, and uses them as a supplement to the computing resources in the edge-cloud architecture. Not only cloud servers and edge nodes are used as migration nodes in the edge-cloud architecture, but also vehicles with abundant resources that have reserved idle computing resources are used as migration nodes, thereby improving the resource utilization of the edge-cloud architecture.
[0072] (2) The migration method for adaptive segmentation tasks based on the Internet of Vehicles proposed in this invention reduces the communication latency of the Internet of Vehicles architecture by using deep deterministic strategy reinforcement learning to jointly determine the task segmentation ratio, communication and computing resources.
[0073] (3) The migration method of the edge-cloud adaptive segmentation task based on the Internet of Vehicles proposed in this invention can adaptively allocate task ratios and select migration modes, effectively cope with task requirements of different scales and complexities, and effectively improve the user experience of the terminal. Attached Figure Description
[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0075] Figure 1 This is an exemplary device architecture diagram in which an embodiment of this application can be applied;
[0076] Figure 2 This is a flowchart illustrating the migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) according to an embodiment of this application.
[0077] Figure 3 This is a schematic diagram of a system model for a migration method of edge-cloud adaptive segmentation task based on vehicle-to-everything (V2X) according to an embodiment of this application.
[0078] Figure 4 This is a schematic diagram of a lane in a simulation environment for the migration method of the edge-cloud adaptive segmentation task based on the Internet of Vehicles, which is an embodiment of this application.
[0079] Figure 5 This is a schematic diagram of a migration device for an edge-cloud adaptive segmentation task based on vehicle networking, as an embodiment of this application.
[0080] Figure 6 This is a schematic diagram of the structure of a computer device suitable for implementing the electronic device of the present application. Detailed Implementation
[0081] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0082] Figure 1 An exemplary device architecture 100 is shown, which can be applied to the migration method or apparatus for the adaptive segmentation task based on the Internet of Vehicles (IoV) in the end-edge-cloud model according to the embodiments of this application.
[0083] like Figure 1 As shown, the device architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0084] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications, such as data processing applications and file processing applications, can be installed on terminal devices 101, 102, and 103.
[0085] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0086] Server 105 can be a server that provides various services, such as a background data processing server that processes files or data uploaded by terminal devices 101, 102, and 103. The background data processing server can process the acquired files or data and generate processing results.
[0087] It should be noted that the migration method for the edge-cloud adaptive segmentation task based on the Internet of Vehicles provided in this application embodiment can be executed by the server 105 or by the terminal devices 101, 102, and 103. Correspondingly, the migration device for the edge-cloud adaptive segmentation task based on the Internet of Vehicles can be set in the server 105 or in the terminal devices 101, 102, and 103.
[0088] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Any number of terminal devices, networks, and servers can be included depending on implementation needs. If the data being processed does not need to be retrieved remotely, the above architecture may not include a network, requiring only servers or terminal devices.
[0089] Figure 2 This application illustrates an embodiment of a migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, comprising the following steps:
[0090] S1. Construct and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers. Obtain the idle computing resources reserved by the vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture.
[0091] In a specific embodiment, step S1 specifically includes:
[0092] Multiple cloud servers publicly announce their unit computing prices and notify vehicles with abundant computing resources of this strategy. After receiving the unit computing prices from different cloud servers, vehicles with abundant computing resources compete with each other to determine which cloud server provides reserved resources. The iteration time for vehicles with abundant computing resources to reach Nash equilibrium is one scheduling cycle of the cloud server. The final result of the cloud server iteration is that all cloud servers and vehicles with abundant computing resources have reached Nash equilibrium. Under this Nash equilibrium, the vehicles obtain the idle computing resources reserved by the vehicles with abundant computing resources.
[0093] Specifically, this study solves a Stackelberg game model with multiple cloud servers as leaders and multiple resource-rich vehicles as followers. It identifies the idle computing resources reserved by the resource-rich vehicles (CRVs) and uses these resources as a supplement to the computing resources in the edge-cloud architecture. In other words, the CRVs with reserved idle computing resources act as computing migration nodes in the edge-cloud architecture. Under the Nash equilibrium of all cloud servers and resource-rich vehicles, no player can gain a higher payoff by unilaterally changing their strategy. The reserved computing resources obtained after the game are completed will serve as available computing resources in the edge-cloud architecture. The initial reserved computing resources for the CRVs are 0.1 GHz, the initial price of the cloud servers is 0.1, and the convergence accuracy of the cloud servers is 10. -3 The iteration step size for both CRV and cloud server is 0.1.
[0094] The embodiments of this application use PyTorch to verify the performance of the MADDPG model and the multi-master multi-slave Stackelberg game algorithm. The entire vehicle-to-everything (V2X) edge-cloud environment and the MADDPG model are as follows: Figure 3 As shown. The vehicle-to-everything (V2X) edge-cloud simulation environment is designed based on the urban case defined in Annex A of 3GPP TR 36.885. This simulation environment includes data such as vehicle location, number of lanes, and traffic flow, and the street map is shown below. Figure 4 As shown, the map is 1299m long and 750m wide. In this simulation environment, a single macro base station is located at the center of the street map, with 6 edge devices and 6 RRHs distributed on both sides of the road. First, vehicle positions are generated based on a Poisson distribution, and then randomized positions are generated as shown below. Figure 4 The vehicle directions are shown, where U represents upward, R represents right, D represents downward, and L represents left. To prevent vehicles from exceeding map limits and thus improve training efficiency, a clockwise turn is made when a vehicle's next position is about to exceed the map's boundaries, ensuring all vehicles remain within the map's range during training. Each vehicle has a computational task with input data size, required computing resources, and maximum tolerable latency set to 10 Mbits, 1–3 GHz, and 1 second, respectively. The maximum CPU frequencies for the cloud server, edge devices, and CRV are 30 GHz, 8 GHz, and 3–5 GHz, respectively. Communication between the vehicle and the cloud server uses a V2I link, while communication between the vehicle and the edge devices / CRV uses a V2V link.
[0095] S2, obtain the vehicle's computing migration request, divide the vehicle's computing task into several sub-tasks according to the task partitioning ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively.
[0096] In a specific embodiment, step S2 specifically includes:
[0097] The computational task for the collected i-th vehicle uses ψ i ={H i F i , t i Represented as a triple, H i F represents the amount of data required for the computation task of the i-th vehicle. i Let t represent the computational resources required for the computational task of the i-th vehicle. i Let ψ represent the maximum tolerable latency for the computation task of the i-th vehicle; let ψ represent the computation task of the i-th vehicle. i Divide into n subsets {ψ i,1 , ψ i,2 ,..., ψ i,n Each subset contains a certain proportion of subtasks. In the edge-cloud architecture, N is defined. cloud Let the set of cloud servers be D, N edge Let P be the set of edge nodes, and N be the set of edge nodes. crv Let E be a set of resource-rich vehicles with reserved idle computing resources, and let subtask ψ be... i,j The subtask ψ is assigned to a migration node X∈D∪P∪E, and processed by the migration node. i,j .
[0098] Specifically, this step involves collecting vehicle computing migration requests and migrating computing tasks to migration nodes in the edge-cloud architecture according to a certain ratio.
[0099] S3 models the task segmentation ratio, computational resource allocation, and communication resource allocation of vehicles as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes. Based on the multi-constraint problem, a MADDPG model is constructed, and the MADDPG model is trained centrally to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent, the environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions, the actions include the best migration mode, communication resources, and computational resources selected by each agent, the reward is the latency performance of the agent, and the state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient.
[0100] In a specific embodiment, step S3 models the vehicle's task segmentation ratio, computing resource allocation, and communication resource allocation as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes, specifically including:
[0101] Within each scheduling cycle, the task of the i-th vehicle is split and then distributed among the selected migration nodes X, assuming the allocation scheme is S = {(X1, ψ...} i,1 ), (X2, ψ i,2 ), ..., (X n , ψ i,n If the total computation time for allocation scheme S is , then the total computation time for allocation scheme S is .
[0102] t i =max(t(X1, ψ) i,1 ), t(X2, ψ i,2 )...,t(X n , ψ i,n )) (1)
[0103] The optimization problem can then be expressed by the following formula:
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113] Where I represents the total number of vehicles in the Internet of Things, constraint (2a) indicates that the sum of task partitioning is 1, and B, A, and R represent the strategies for migration task ratio, computing resources, and communication resources, respectively. These represent the proportions of tasks migrated to cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively. in These represent the computing resources required by cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively. This indicates the channel selection for the cloud server. This represents the channel selection at edge nodes and resource-rich vehicles with reserved idle computing resources; constraint (2b) ensures that the size of the task split is greater than the minimum divisible size G of the task; constraints (2c), (2d), and (2e) represent the maximum computing resource limits of cloud servers, edge nodes, and resource-rich vehicles with reserved idle computing resources, respectively; constraint (2f) indicates that the execution time when migrating to an edge node should be less than the maximum tolerance time of the task and the minimum communication time with it; constraint (2g) indicates that the execution time when migrating to a vehicle with abundant computing resources should be less than the maximum tolerance time of the task and the minimum communication time with it.
[0114] For the communication resource allocation in the edge-cloud architecture, there are M remote wireless heads, and the cloud server corresponds to a resource block set R. c ={1, ..., r c The set of resource blocks R corresponding to edge nodes and resource-rich vehicles with reserved idle computing resources. f ={1, ..., r f}, where each resource block has the same bandwidth W, using binary variables. Let R represent the resource allocation scheme for the i-th vehicle, where j∈R c , k∈R f ;like This indicates that the j-th resource block is allocated to the i-th vehicle. This indicates that the k-th resource block is allocated to the i-th vehicle;
[0115] The latency of the i-th vehicle migrating the task to the cloud server d The calculation formula is:
[0116]
[0117]
[0118] Among them, V cloud,dp represents the migration rate on cloud server d. i This represents the transmission power of the i-th vehicle. Let Ni represent the channel gain between the i-th vehicle and the m-th remote wireless head, and N0 represent the noise spectral density. t represents the proportion of tasks migrated to the cloud server. wired F represents the wired propagation delay. cloud,d It is the computing resources provided by cloud server d;
[0119] The latency for the i-th vehicle to migrate the task to the edge node p The calculation formula is:
[0120]
[0121]
[0122] Among them, V edge,p This represents the migration rate at edge node p. This represents the proportion of tasks that migrate to edge node p. F represents the channel gain between the j-th vehicle and the p-th edge node. edge,p This represents the computing resources provided by the edge node p;
[0123] The latency of the i-th vehicle migrating the task to the computationally abundant vehicle e. The calculation formula is:
[0124]
[0125]
[0126] Among them, V crv,e This represents the migration rate of vehicle e, which has abundant computing resources. F represents the proportion of tasks migrated to vehicle e with reserved idle computing resources. crv,e This indicates the computing resources reserved for the resource-rich vehicle e.
[0127] In a specific embodiment, the evaluation network and target network corresponding to each agent both include a Critic network and an Actor network. The state of the i-th agent at time t is defined as follows:
[0128] s t =(h t , τ t I t-1 O t-1 ,∈,ρ) (9)
[0129] Among them, h tThis represents the local channel information at time t, including the channel gain between the i-th agent and the remote wireless head at time t. Channel gain between agent i and edge node Channel gain between the i-th agent and a resource-rich vehicle with reserved idle computing resources. τ t This represents the dwell time at time t, including the dwell time of the i-th agent and the edge node at time t. The dwell time of the i-th agent and the resource-rich vehicle with reserved idle computing resources. Network status information at time t-1 includes wireless interference: Remaining computing resources of cloud servers and edge nodes: ∈ represents the number of iterations, and ρ represents the exploration coefficient;
[0130] The action of the i-th agent is defined as:
[0131]
[0132] The reward obtained by the i-th agent in the environment is defined as:
[0133] r t =α(T) i,max -T i )-βr sat (11)
[0134] Where α and β are the weights of each part, and T i,max T represents the maximum tolerable delay for the i-th agent. i Indicates the actual time delay used, r sat Let be the delay non-satisfaction rate of all agents, which is the sum of all non-satisfying vehicles divided by the total number of vehicles. Here, α is taken as 0.4 and β as 0.6.
[0135] Specifically, embodiments of this application employ the MADDPG model to solve multi-constraint problems, and combine fingerprint and action masking methods for centralized training to obtain a trained MADDPG model.
[0136] In a specific embodiment, step S3 involves centralized training of the MADDPG model, specifically including:
[0137] The centralized training of the MADDPG model specifically includes:
[0138] Initialize the parameter weights θ of the Actor network and Critic network of the agent's evaluation network. μ θ Q The parameter weights θ of the Actor network and Critic network of the target networkμ′ θ Q′ And initialize the environment and experience pool R;
[0139] Reserved computing resources in vehicles with abundant computing resources are used to observe the state obtained by the i-th agent from the environment.
[0140] Perform actions based on the state observed by the agent. Obtain the state at the next moment after performing the action. All agents have completed their actions and received rewards from the environment. t ,storage The data is fed into the experience pool R, and the Actor and Critic networks of the evaluation network are trained using the data in the experience pool R. The parameters of the evaluation network are then updated to the Actor and Critic networks of the target network. After the loop is completed, the trained MADDPG model is obtained.
[0141] The Critic network in the MADDPG model minimizes the root mean square loss using gradient descent, with the loss function being:
[0142]
[0143] Where ε is the discount rate, set to 0.85;
[0144] The Actor network updates its gradient using the following formula:
[0145]
[0146] The parameters θ of the target network μ′ and θ Q′ Use soft update:
[0147] θ Q′ ←τθ Q +(1-τ)θ Q′ (14)
[0148] θ μ′ ←τθ μ +(1-τ)θ μ′ (15)
[0149] When updating network parameters, batch sampling is performed, and during training, the network state information from the previous time step is recorded and added to the multi-agent state.
[0150] In a specific embodiment, a custom ReLU activation function is used in the output layers of the Critic network and the Actor network to filter out tasks below a threshold in the output actions, and the filtered actions are then input into the Softmax function.
[0151] Specifically, in terms of agent design, each agent's Critic network consists of three fully connected hidden layers, containing 256, 128, and 64 neurons respectively. A Rectified Linear Unit (ReLU) is used as the activation function, and an adaptive moment estimation optimizer is used to iteratively train and update the neural network weights. The Actor network also consists of three fully connected hidden layers, but the output layer uses a custom ReLU activation function. This function sets values less than the task threshold to 0, while values greater than or equal to the threshold remain unchanged. The filtered data is then input into a Softmax function to obtain a normalized probability distribution that results in a sum of action values of 1. The result of the Softmax layer in the output layer is the final action vector. A total of 1000 episodes were trained. To facilitate action exploration, a variable Gaussian noise was added to the actions selected by the agent. The exploration coefficients were processed using a linear annealing algorithm, annealing from 1 at the beginning to 0.002 at episode 600, and the search probability remained constant in subsequent training. Table 1 lists the main simulation parameters, and Table 2 lists the channel models for the V2I and V2V links.
[0152] Table 1 Simulation Default Parameters
[0153]
[0154]
[0155] Table 2 Channel Models for V2I and V2V
[0156]
[0157] S4: Obtain the vehicle's state and input it into the trained MADDPG model to obtain the optimal action. Select the migration node based on the optimal action and determine the migration task ratio.
[0158] Specifically, the action decisions output by the trained MADDPG model inform the vehicle to select migration nodes and migration task ratios. In the distributed execution phase, each agent obtains its state based on its local environment information, inputs the state into the trained MADDPG model to obtain action decisions, and selects migration nodes and migration task ratios based on the action decisions.
[0159] The steps S1-S4 above do not represent the order of the steps, but are merely symbolic representations of the steps.
[0160] Further reference Figure 5 As an implementation of the methods shown in the above figures, this application provides an embodiment of a migration device for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication. This device embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0161] This application provides a migration device for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, including:
[0162] Resource computing module 1 is configured to build and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers, obtain the idle computing resources reserved by the vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture.
[0163] Task splitting module 2 is configured to obtain the vehicle's computing migration request, split the vehicle's computing task into several sub-tasks according to the task splitting ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively.
[0164] Model training module 3 is configured to model the task segmentation ratio, computing resource allocation, and communication resource allocation of the vehicle as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes. Based on the multi-constraint problem, a MADDPG model is constructed, and the MADDPG model is trained centrally to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent, the environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions, the actions include the best migration mode, communication resources, and computing resources selected by each agent, the reward is the latency performance of the agent, and the state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient.
[0165] Execution module 4 is configured to acquire the vehicle's status, input it into the trained MADDPG model, obtain the optimal action, select the migration node based on the optimal action, and determine the migration task ratio.
[0166] The following is for reference. Figure 6 It illustrates an electronic device suitable for implementing embodiments of this application (e.g., Figure 1 The diagram shows the structure of a computer device 600 (a server or terminal device). Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0167] like Figure 6 As shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 603 or programs loaded from storage section 609 into random access memory (RAM) 604. The RAM 604 also stores various programs and data required for the operation of the device 600. The CPU 601, GPU 602, ROM 603, and RAM 604 are interconnected via a bus 605. An input / output (I / O) interface 606 is also connected to the bus 605.
[0168] The following components are connected to I / O interface 606: an input section 607 including a keyboard, mouse, etc.; an output section 608 including an LCD, speakers, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card, such as a LAN card or modem. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 may also be connected to I / O interface 606 as needed. A removable medium 612, such as a hard disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 611 as needed so that computer programs read from it can be installed into storage section 609 as needed.
[0169] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 610, and / or installed from removable medium 612. When the computer program is executed by central processing unit (CPU) 601 and graphics processing unit (GPU) 602, the functions defined in the methods of this application are performed.
[0170] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable medium, or any combination thereof. A computer-readable medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, or any combination thereof. More specific examples of a computer-readable medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution device, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than a computer-readable medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution device, apparatus, or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0171] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using dedicated hardware-based means to perform the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0173] The modules described in the embodiments of this application can be implemented in software or hardware. These modules can also be located within a processor.
[0174] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: construct and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple resource-rich vehicles as followers, obtain the reserved idle computing resources of the resource-rich vehicles, and include the resource-rich vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture; obtain the vehicle's computing migration request, divide the vehicle's computing tasks into several sub-tasks according to the task partitioning ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively; and model the vehicle's task partitioning ratio, computing resource allocation, and communication resource allocation into a multi-constraint problem. The problem is to minimize the sum of the maximum latency of each subtask of a vehicle under different migration modes. A MADDPG model is constructed based on the multi-constraint problem. The MADDPG model is trained centrally to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent. The environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions. Actions include the optimal migration mode selected by each agent, communication resources, and computing resources. The reward is the agent's latency performance. The state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient. The state of the vehicle is obtained and input into the trained MADDPG model to obtain the optimal action. The migration node is selected based on the optimal action, and the migration task ratio is determined.
[0175] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, characterized in that, Includes the following steps: Construct and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers. Obtain the idle computing resources reserved for vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture. Obtain the vehicle's computing migration request, divide the vehicle's computing task into several sub-tasks according to the task partitioning ratio based on the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively. The task segmentation ratio, computing resource allocation, and communication resource allocation of the vehicle are modeled as a multi-constraint problem. The maximum sum of the maximum latency of each subtask of the vehicle under different migration modes is minimized. A MADDPG model is constructed based on the multi-constraint problem. The MADDPG model is trained in a centralized manner to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent. The environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions. The actions include the best migration mode, communication resources, and computing resources selected by each agent. The reward is the latency performance of the agent. The state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient. The vehicle's status is obtained and input into the trained MADDPG model to obtain the optimal action. Based on the optimal action, the migration node is selected and the migration task ratio is determined.
2. The migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) as described in claim 1, characterized in that, The construction and solution of a multi-master, multi-slave Stackelberg game model, in which multiple cloud servers act as leaders and multiple resource-rich vehicles act as followers, to obtain the idle computing resources reserved for resource-rich vehicles, specifically includes: Multiple cloud servers publicly announce their unit computing prices and notify vehicles with abundant computing resources of this strategy. After receiving the unit computing prices from different cloud servers, vehicles with abundant computing resources compete with each other to determine which cloud server provides reserved resources. The iteration time for vehicles with abundant computing resources to reach Nash equilibrium is one scheduling cycle of the cloud server. The final result of the cloud server iteration is that all cloud servers and vehicles with abundant computing resources have reached Nash equilibrium. Under this Nash equilibrium, the vehicles obtain the idle computing resources reserved by the vehicles with abundant computing resources.
3. The migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) as described in claim 1, characterized in that, The step of dividing the vehicle's computing task into several sub-tasks according to the task partitioning ratio based on the migration request, and migrating the several sub-tasks to the migration nodes of the edge-cloud architecture respectively, specifically includes: The computational task for the collected i-th vehicle adopts... Represented by triples. This represents the amount of data required for the computation task of the i-th vehicle. This represents the computational resources required for the computational task of the i-th vehicle. This represents the maximum tolerable latency for the computation task of the i-th vehicle; the computation task of the i-th vehicle... Divide into n subsets Each subset contains a certain proportion of subtasks. In the edge-cloud architecture, the following is defined: The collection of cloud servers is , The set of edge nodes is , A set of resource-rich vehicles with reserved idle computing resources is Assign subtasks to migration nodes The migration node then processes the subtasks.
4. The migration method for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) as described in claim 1, characterized in that, Each agent's evaluation network and target network include a Critic network and an Actor network. The output layers of the Critic network and Actor network use a custom ReLU activation function to filter out tasks below a threshold from the output actions, and then input the filtered actions into the Softmax function.
5. A migration device for edge-cloud adaptive segmentation tasks based on vehicle-to-everything (V2X) communication, characterized in that, include: The resource computing module is configured to build and solve a multi-master, multi-slave Stackelberg game model with multiple cloud servers as leaders and multiple vehicles with abundant computing resources as followers, obtain the idle computing resources reserved by the vehicles with abundant resources, and include the vehicles with reserved idle computing resources in the migration nodes of the edge-cloud architecture. The task segmentation module is configured to obtain the vehicle's computing migration request, divide the vehicle's computing task into several sub-tasks according to the migration request, and migrate the several sub-tasks to the migration nodes of the edge-cloud architecture respectively. The model training module is configured to model the task segmentation ratio, computing resource allocation, and communication resource allocation of the vehicle as a multi-constraint problem, minimizing the sum of the maximum latency of each subtask of the vehicle under different migration modes. Based on the multi-constraint problem, a MADDPG model is constructed, and the MADDPG model is trained centrally to obtain a trained MADDPG model. In the MADDPG model, each vehicle is an agent, the environment includes vehicle information in the vehicle network, resource information of the edge-cloud architecture, and channel conditions, the actions include the best migration mode, communication resources, and computing resources selected by each agent, the reward is the latency performance of the agent, and the state includes local channel information, dwell time, fingerprint information, number of iterations, and exploration coefficient. The execution module is configured to acquire the vehicle's status, input it into the trained MADDPG model, obtain the optimal action, select a migration node based on the optimal action, and determine the migration task ratio.
6. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-4.