A Knowledge-Driven Method for On-Vehicle Service Resource Allocation and Computation Offloading
By adopting a knowledge-driven near-end strategy optimization method in vehicle-mounted service resource allocation and computational offloading, combining dual time scales and neural networks for access selection and resource allocation, the problems of insufficient generalization performance and low credibility of the existing technology in complex communication scenarios are solved, and more efficient and trustworthy resource allocation and computational offloading effects are achieved.
Patent Information
- Application Number
- CN202410689551.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-05-30
AI Technical Summary
The existing technology has insufficient generalization performance in complex communication scenarios, low credibility and acceptance of the model, poor timeliness, and cannot effectively respond to the challenges of the integrated network in the air and space in the future.
A knowledge-driven near-end strategy optimization method is adopted to allocate resources and calculate and offload resources through a dual time scale, and a neural network is used to make access selection, power allocation and computing resource allocation decisions, and a security layer is introduced into the near-end strategy optimization algorithm to ensure the credibility and security of decisions.
It improves the interpretability and timeliness of the system, reduces the system cost and overhead, enhances the generalization and credibility of the model, and can more effectively adapt to the complex integrated air-space network, and reduces the total delay of on-board tasks.
Smart Images

Figure CN118474796B_ABST
Abstract
Description
Technical Field:
[0001] The present invention belongs to the technical field of vehicle networking, and particularly relates to a method for allocating vehicle-mounted service resources and computing offloading based on knowledge-driven. Background Art:
[0002] With the continuous development of the vehicle networking field, complex communication scenarios have been continuously proposed, driving the demand for more powerful network support. The application of heterogeneous networks provides a solution to meet this demand. As one of the representative networks in heterogeneous networks, the space-air-ground integrated network provides important support for the development of vehicle networking. The demand for seamless coverage of new vehicle-mounted communication services has become more urgent, and at the same time, the requirement for load balancing is also getting higher. Among them, a large amount of network resources will be consumed during the transmission of vehicle-mounted communication services, while the network resources are limited. Therefore, it is necessary to allocate network resources efficiently. For a large number of communication devices and tasks, traditional model-driven and data-driven methods cannot effectively perform unified resource allocation and management in the case of network heterogeneity and complexity.
[0003] Regarding the problems of vehicle-mounted service resource allocation and computing offloading, the existing wireless network communication resource allocation algorithms mainly have the following two methods. The first is the network access selection and resource allocation algorithm driven by a mathematical model: The document with the application number "201810915604.7" discloses "an intelligent wireless network access method", which intelligently selects the best access base station method, reduces base station handover, improves the connectivity and bandwidth performance of the wireless network, and shortens the network delay. The problem of this document is that the adopted model is limited and relatively single, and its generalization performance is poor in complex communication scenarios. The second is the data-driven network access selection and resource allocation algorithm: The document with the application number "202111670934.2" discloses "a method for allocating multi-node computing resources in a space-ground integrated network based on deep reinforcement learning", which intelligently determines local service nodes and cooperative service nodes from each service point in the space-ground integrated network to minimize the weighted system overhead of satellite energy consumption and task execution delay. The problem is that the entire decision-making is processed through reinforcement learning, and the credibility and acceptability of the model are relatively low in key decision-making scenarios.
[0004] The common problems of the above-mentioned existing technologies mainly lie in the generalization performance, the credibility and acceptability of the model, and the challenges in practical applications. Specifically, they show insufficient generalization performance in complex communication scenarios. At the same time, due to the complexity and black-box nature of the model, the credibility and acceptability may be relatively low in key decision-making scenarios. In addition, they have poor timeliness in practical applications, face challenges such as real-time performance and computing resource requirements, and are insufficient in adapting to future network requirements. The above deficiencies make the above two methods unable to cope with the challenges under the future space-air-ground integrated network. Summary of the Invention:
[0005] The object of the present invention is to provide a knowledge-driven method for allocating on-vehicle service resources and computing offloading, so as to solve the problems of low credibility, poor timeliness and insufficient adaptability to future network requirements existing in the prior art.
[0006] To achieve the above object, the steps of the technical solution adopted by the present invention are as follows: A knowledge-driven method for allocating on-vehicle service resources and computing offloading, including the following steps:
[0007] Step 1. Obtain the application scenarios of the space-air-ground integrated network, and complete the establishment of the on-vehicle service uplink transmission and offloading link models.
[0008] Step 2. Set the optimization objective and constraint conditions for minimizing the total overhead of task processing. The total overhead of task processing includes the overhead caused by processing the transmission load imbalance in the small time scale, and the overhead caused by processing the load imbalance in the large time scale of computing, and establish an optimization model;
[0009] Step 3. Solve the optimization model by using the knowledge-driven proximal policy optimization method, and output the corresponding decisions in the large time scale;
[0010] Step 4. Connect the access channel to the space-air-ground integrated network according to the decision output in Step 3, so as to complete the vehicle access selection, computing offloading and resource allocation processes.
[0011] The proximal policy optimization method in the above Step 3 is: taking the size of on-vehicle service data volume, the maximum tolerable delay of the task, the computing resource requirements of the task, vehicle mobility, the initial position of the vehicle, and the vehicle channel state as the input of the neural network; taking access selection and computing resource allocation as actions; the unbalanced transmission load, computing load overhead, penalty terms for violating computing resource constraints and delay constraint quotas jointly form the reward function;
[0012] The decisions in the above Step 3 include access network access selection, power allocation strategy and computing resource scheduling strategy.
[0013] The reward function in the above Step 3 is as follows:
[0014]
[0015] The access selection in the above Step 3 refers to that the vehicle selects a base station, an unmanned aerial vehicle and a satellite for access in the space-air-ground integrated scenario.
[0016] The power allocation strategy in the above Step 3 refers to that the core network allocates different values to the uplink transmission power of the vehicle;
[0017] The computing resource allocation strategy in step 3 above refers to the core network allocating the values of the computing resources available to the base stations, drones, and satellites.
[0018] In step 3 above, a security layer is added to the proximal policy optimization method.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] 1. Aiming at minimizing the overhead problems caused by unbalanced transmission load and unbalanced computing load in the multi-vehicle transmission and computing offloading process, the present invention decomposes the original problem into problems of large and small time scales according to the difference in the execution time granularity of different decisions of the original problem, and obtains the overall optimal decision by combining knowledge with the proximal policy optimization method. The computing offloading of resource allocation is processed in a dual time scale manner, in which the small time scale uses traditional optimization methods for theoretical analysis, effectively increasing the interpretability of the system, reducing the dimension of the traditional deep reinforcement learning action space, reducing the cost overhead of the system, and at the same time improving the timeliness and credibility of the entire system, and can effectively reduce the total delay of vehicle task processing, including transmission delay and computing offloading delay. In the large time scale, the proximal policy optimization method is used to make decisions on the sub-task segmentation ratio, access selection, and computing resource allocation, enhancing the generalization of the model. The dual time scale method integrates the advantages of large and small time scales, can effectively improve the generalization, timeliness, and credibility of the algorithm, and better adapt to complex space-air-ground integrated networks.
[0021] 2. Aiming at the problem that the proximal policy optimization algorithm in step 3 may make uncertain actions when exploring the environment, which may lead to unsafe or unstable behaviors of the system, the present invention introduces a security layer into the knowledge-driven proximal policy optimization algorithm mentioned in step 3, effectively ensuring that the constraints related to computing resources are never violated during the reinforcement learning training process, enhancing the practical feasibility of the overall network, and having strong adaptability to future network requirements. Description of the Drawings:
[0022] Figure 1 is a flowchart provided by the present invention;
[0023] Figure 2 is a framework diagram of the knowledge-driven proximal policy optimization algorithm provided by the present invention;
[0024] Figure 3 is a model diagram of the vehicle network access system in the space-air-ground integrated scenario provided by the present invention;
[0025] Figure 4 is a convergence diagram of the knowledge-driven dual time scale algorithm provided by the present invention;
[0026] Figure 5Experimental result graph of different maximum transmit power limits and system transmission overheads provided by the present invention; Specific implementation manner:
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0028] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0029] The framework of the present invention is as Figure 1 shown. First, the present invention obtains the integrated space-air-ground network scenario and completes model modeling; then sets the optimization objective and constraint conditions for minimizing the total overhead of task processing; secondly, decomposes the original problem into problems of large and small time scales, and obtains the overall optimal decision by combining the knowledge and proximal policy optimization methods; finally, according to the solved strategy, completes the access selection, computing offloading and resource allocation of multiple vehicles in the integrated space-air-ground network.
[0030] Based on the above ideas, the present invention gives the following steps:
[0031] Step 1: Obtain the application scenario of the integrated space-air-ground network and complete the establishment of the on-vehicle service uplink transmission and offloading link model.
[0032] The system model of the present invention is as Figure 2 shown. In the application scenario of the integrated space-air-ground network in the figure, there are multiple vehicles, multiple unmanned aerial vehicles, multiple base stations and at least one low-earth orbit satellite. Considering the decision time granularity difference between task transmission offloading and core network resource allocation, a dual-time scale method is used to solve this problem. For the large time scale, vehicle link selection and computing resource allocation are performed, and for the small time scale, power allocation is performed. The on-vehicle service arrives at the vehicle and performs uplink transmission. After the service arrives at the access network, it is computed and offloaded in the configured mobile edge server, and after processing, it is downlink transmitted back to the vehicle;
[0033] Step 101: Establish the on-vehicle service uplink transmission and offloading link model in the integrated space-air-ground scenario:
[0034] Considering large-scale and small-scale channel fading, the m-th module of the k-th vehicle selects a sub-channel of the access network for access within the time slot (T, t), and the link signal-to-noise ratio of this link can be calculated as
[0035]
[0036]
[0037] Among them is the large-scale fading coefficient at this small time scale, is the small-scale fading coefficient, is the power allocated to the k-th vehicle in the time slot (T, t), and δ 2 is the power of additive white Gaussian noise.
[0038] Since at most one access link can be connected to a module of a vehicle, so there is
[0039]
[0040] Assume that the channel bandwidth is B n , according to Shannon's formula, the maximum data transmission rate that the k-th vehicle can achieve on access link n at time slot (T, t) is
[0041]
[0042] where μ is the fixed time slot length, is the amount of subtask data arriving at time window T. The above constraints indicate that for any vehicle, the amount of task data unloaded through any vehicle should not be greater than the instantaneous channel capacity between them.
[0043] Step 102, establish a vehicle service load balancing model in the space-air-ground integrated scenario:
[0044] Transmission load: The amount of data to be transmitted divided by the amount of data actually transmitted. Sum the different links accessing the same access network, and only consider the uplink transmission. Transmission load balancing usually refers to a strategy for allocating and managing data traffic in a network, whose goal is to maximize the network throughput, minimize the delay, and prevent any single network link from being overloaded.
[0045]
[0046] The computing load is the amount of computing offloading required on a large time scale divided by the amount of actual computing, and the formula is as follows:
[0047]
[0048] Set as the overhead for handling the unbalanced transmission load, set as the overhead for handling the unbalanced computing load, and set the overall overhead as C:
[0049]
[0050] Step 2, set the optimization goal and constraints for minimizing the total overhead of task processing, and establish an optimization model.
[0051] Step 201: Set the optimization goal of minimizing the total overhead of task processing and the constraints. For specific problems and constraints, refer to Problem P0 and Formulas (9.1 - 9.7).
[0052] Among them, the total overhead of task processing includes the overhead for handling the transmission load imbalance at a small time scale and the overhead for handling the load imbalance at a large time scale in computing. Considering that the entire system adopts a centralized control architecture and is controlled at different time granularities, all base stations and satellites are connected to a unified controller deployed near the edge of the core network to monitor the dynamic network conditions and make decisions on subtask segmentation, offloading link selection, power resource allocation, and computing resource allocation.
[0053] The optimization goal of the present invention is to minimize the total overhead of task processing. The optimization factors are access selection, computing resource allocation, and power resource allocation. The problem can be expressed as P0:
[0054]
[0055]
[0056] Formulas (1.1) - (1.7) are the constraints of Problem P0, and their meanings are as follows: C1 ensures that each vehicle can only select one access network for access; C2 gives the maximum allowable uplink transmission power on each vehicle module; C3 gives the maximum allowable computing frequency of each mobile edge server; C4 indicates that the processing delay (uplink transmission delay and computing delay) of the transmission subtask is within the maximum tolerance delay constraint; C5 ensures that the computing load does not exceed the computing capacity; C6 ensures the transmission energy consumption constraint; C7 ensures that for any vehicle, the amount of subtask data unloaded through any vehicle module m should not be greater than the instantaneous channel capacity between them.
[0057] Among them, the symbol explanations are shown in Table 1.
[0058] Table 1 Symbol Table
[0059]
[0060]
[0061] Step 202: Decompose the established optimization problem into problems at large and small time scales respectively, and give the corresponding problem modeling and constraints.
[0062] Large - time - scale problem modeling:
[0063]
[0064] C1, C3, C5
[0065] Modeling of Small Time Scale Problems:
[0066]
[0067] C2, C4, C6, C7
[0068] Step 3: Use the knowledge-driven proximal policy optimization method to solve the optimization model and output corresponding decisions. The decisions include access selection, power allocation strategy, and computing resource allocation strategy;
[0069] See Figure 3 , and the solution uses a dual time scale method: within the small time scale, traditional knowledge is used to allocate power resources, and within the large time scale, the proximal policy optimization method is used to make decisions on sub-task segmentation ratio, access selection, and computing resource allocation. Here, the large and small time scales are relative concepts, and their physical meanings are represented by large and small time slots. A large time slot is composed of multiple small time slots with a fixed time length, and the length of each large time slot is also fixed. The specific implementation includes two parts:
[0070] 301. Use the proximal policy optimization method to solve the optimal strategy for the large time scale problem: Take the size of in-vehicle service data volume, the maximum tolerable delay of tasks, the computing resource requirements of tasks, vehicle mobility, vehicle initial position, and vehicle channel state as the inputs of the neural network; take access selection and computing resource allocation as actions; jointly form a reward function with unbalanced transmission load, computing load overhead, and penalty terms for violating computing resource constraints and delay constraints; output corresponding large time scale decisions at the output layer. The decisions include access network access selection and computing resource allocation decisions.
[0071] The network access selection refers to the vehicle selecting base stations, unmanned aerial vehicles (UAVs), and satellites for access in the integrated space-air-ground scenario.
[0072] The computing resource allocation strategy refers to the core network allocating the values of computing resources available at base stations, UAVs, and satellites.
[0073] The reward function is designed as follows:
[0074]
[0075] Here, represents the average transmission cost on the small time scale, while represents the computing cost on the large time scale. tran pn and com pn are penalty terms for violating delay constraints and computing capacity constraints respectively. If the delay of the in-vehicle service exceeds the maximum tolerable delay, or the volume of transmitted data exceeds the capacity of the transmission link, then -λ(tran pn + compn ) A penalty is applied to the reward. Our goal is to optimize the allocation of resources to minimize the total cost (including transmission and computing costs) while satisfying the constraints of latency and computing power. The penalty term is used to discourage solutions that violate these constraints. The parameter λ is a weight factor that determines the relative importance of the penalty term in the total cost.
[0076] However, the method of introducing an additional hyperparameter λ to balance load balancing, latency guarantee, and link capacity guarantee can only ensure that the constraints are satisfied after learning converges. Due to the random exploration behavior, it cannot guarantee zero constraint violations during the entire learning process, which harms the online performance. Therefore, the present invention introduces a safety layer by using the model of the system to predict the results of each possible action. If the predicted result violates the constraint conditions, the operation is not allowed to be executed. A safety layer is added to the proximal policy optimization algorithm to meet the latency and link capacity limitations. Specifically, a large penalty is imposed on the behavior that violates the constraints, that is, a value greater than the maximum tolerable latency constraint of the task is given.
[0077] 302. Solve the optimal policy for small time-scale problems: Introduce the knowledge of traditional optimization theory and prove through theoretical derivation that the small time-scale problem is an optimization problem. Utilize traditional knowledge to schedule power resources within a small time scale, and the optimal power allocation policy can be directly solved through calculation under the given constraint conditions. The power allocation policy refers to the core network allocating different values to the uplink transmission power of vehicles;
[0078] The specific access selection, power allocation policy, and computing resource allocation policy obtained from the above two parts will be output through the neural network of the proximal policy optimization algorithm.
[0079] Step 4, according to the decision output in step three, access the channel to the space-air-ground integrated network to complete the vehicle access selection, computing offloading, and resource allocation process.
[0080] To test the method of the present invention, the experimental parameters of the present invention are set as shown in Table 2.
[0081] Table 2 Experimental parameters
[0082] Parameter Value Parameter Value Number of users 10 Base station coverage radius 0.5 km Number of base stations 4 UAV coverage radius 2 km Number of UAVs 2 Satellite coverage radius 800 km Number of satellites 1 Satellite altitude 550 km Total transmission bandwidth 500 kHz Base station altitude 15m Maximum transmission power of the uplink 23 dBm UAV altitude 150m Gaussian white noise power -173 dBm Maximum computing frequency of MECS 2 GHz
[0083] The key simulation parameters of the present invention are set as shown in Table 3
[0084] Table 3 Simulation parameters
[0085] Parameter Numerical value Learning rate of the actor network 3e-4 Learning rate of the critic network 1e-4 Reward discount 0.9 Total number of training times 600 time window
[0086] The basic configuration of the experimental equipment is NVIDIA GeForce RTX 3090.
[0087] As Figure 4 shown, when the selected clip ratio is 0.5, the learning rate of the Actor network is 3e-04, and the learning rate of the Critic network is 1e-04, the KD-PPO, H-PPO, and PPO algorithms can all converge, and the performance of the KD-PPO algorithm is the best.
[0088] As Figure 5 shown, the system changes the upper limit of the maximum transmission power of each vehicle from 0.8 to 1.3 (scaling factor). It can be found from the figure that under different transmission power constraints, the transmission overhead of KD_PPO is the smallest.
[0089] As mentioned above, it is only the preferred embodiment of the present invention, and is not used to limit the protection scope of the present invention. Any equivalent structural changes made by using the description and drawings of the present invention shall be included in the patent protection scope of the invention.
Claims
1. A knowledge-driven vehicle-borne business resource allocation and computation offloading method, characterized in that: The following steps are involved: Step 1. Obtain the application scenario of the air-ground integrated network and complete the establishment of the vehicle-borne service uplink transmission and offload link model; Step 2. Set the optimization goal and constraints of minimizing the total overhead of task processing. The total overhead of task processing includes the overhead caused by the imbalance of transmission load at a small time scale and the overhead caused by the imbalance of load at a large time scale. Establish an optimization model. Step 3. Solve the optimization model by using a knowledge-driven proximal strategy optimization method to output a corresponding large-time-scale decision; Step 4. Connect the access channel to the air-ground integrated network according to the decision output in step 3 to complete the vehicle access selection, computation offloading and resource allocation process; The proximal strategy optimization method in step 3 is: taking the amount of vehicle service data, the maximum tolerable delay of the task, the computing resource requirements of the task, the vehicle mobility, the initial position of the vehicle, and the vehicle channel state as the input of the neural network; taking access selection and computing resource allocation as actions; and the unbalanced transmission load, computing load overhead, violation of computing resource constraints and delay constraint penalty items together constitute the reward function; The reward function in step 3 is as follows: in: tran pn and com pn are the penalty terms for violating the delay constraint and computing power constraint respectively; is the average value of transmission cost on a small time scale; t is the tth hour-scale time slot; T is the Tth large time scale time window; n is the number of access networks; The overhead of handling unbalanced transmission load; The overhead of handling unbalanced computing loads; The computational cost for large time scales; λ is the penalty factor; Con is the total time slot of the small time scale; ρ1 and ρ2 are the ratios of transmission load imbalance and computation load imbalance, respectively.
2. The method for allocating vehicle-borne business resources and offloading calculations based on knowledge-driven according to claim 1, characterized in that: The decision in step 3 includes access network selection, power allocation strategy and computing resource scheduling strategy.
3. The method for allocating vehicle-borne business resources and offloading calculations based on knowledge-driven according to claim 1, characterized in that: The access selection in step 4 refers to the vehicle selecting a base station, a drone and a satellite for access in an air-ground-ground integration scenario.
4. The method for allocating vehicle-borne business resources and computing offloading based on knowledge-driven according to claim 2, characterized in that: The power allocation strategy in step 3 refers to the core network allocating different values of uplink transmission power to the vehicle.
5. The method for allocating vehicle-borne business resources and offloading calculations based on knowledge-driven according to claim 4, characterized in that: The computing resource scheduling strategy in step 3 refers to the core network allocating the value of computing resources possessed by base stations, drones and satellites.
Citation Information
Patent Citations
A smart wireless network access method
CN109041153B
A Multi-Node Computational Resource Allocation Method for Space-Ground Fusion Networks Based on Deep Reinforcement Learning
CN115250142B
Knowledge-driven resource scheduling method in Internet of Vehicles in space-air-ground integrated scene
CN116133127A
Multi-type terminal random access competition solution based on near-end strategy optimization
CN117241409A