Internet of vehicles resource slicing method based on transformer federated learning

Through the Internet of Vehicles resource slicing method based on Transformer dual federated learning, a two-layer controller and multi-head attention mechanism are adopted to solve the problems of personalized slicing window adjustment and model generalization of resource management in the Internet of Vehicles, and realize efficient and flexible resource allocation and decision optimization.

CN119421246BActive Publication Date: 2025-10-17NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411733222.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-17
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The Internet of Vehicles faces difficulties in adjusting personalized slicing windows, insufficient agility and timeliness of resource slicing, and poor model scalability and generalization in road network environments. Existing technologies make it difficult to effectively manage resources in highly dynamic and non-uniform scenarios.

Method used

A vehicle network resource slicing method based on Transformer dual federated learning is adopted, a two-layer controller structure is introduced, and global and local models are combined to optimize resource allocation and management through a window-adaptive flexible slicing framework, traffic-aware personalized slicing window adjustment and multi-head attention mechanism.

Benefits of technology

It achieves flexible resource management in highly dynamic Internet of Vehicles scenarios, improves slice performance isolation, resource allocation efficiency and decision-making effectiveness, adapts to different traffic flow fluctuations, and improves the generalization performance of the model and the overall service level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119421246B_ABST
    Figure CN119421246B_ABST
Patent Text Reader

Abstract

A vehicle networking resource slicing method based on Transformer double-federated learning, in a vehicle networking system, the physical resources of base stations are abstracted into a virtual resource pool, which is managed by a controller; according to the terminal request information fed back by each base station, the controller adjusts the allocation of sliced resources. The controller is a double-layer controller structure; the centralized controller in the cloud integrates a global model, which includes a global traffic prediction model and a global resource slicing decision model; the centralized controller slices the resources of the base stations through the global model; the local controller on the edge side connected with adjacent base stations integrates a local model, which includes a local traffic prediction model, a slicing window optimizer and a local resource slicing decision model; the local controller makes resource slicing decisions through the local model; the slicing steps include: adopting a window self-adaptive slicing method to slice the physical resources of each base station; and optimizing the decision result through local model training and double-federated model method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of Internet of Vehicles, and particularly relates to an Internet of Vehicles resource slicing method based on Transformer dual-federated learning. BACKGROUND

[0002] In the post-5G and 6G era, the Internet of Vehicles is facing complex and diversified application demands, especially in scenarios such as autonomous driving, fleet coordination, and remote driving. In order to meet these key applications, the network needs to have higher bandwidth, lower latency, and greater capacity. Autonomous vehicles have extremely high requirements for real-time control and environmental perception when driving at high speeds, and the network must provide ultra-low latency and high reliability communication support. At the same time, the transmission demand of high-definition maps, vehicle sensor data, and multimedia content such as video is also increasing rapidly, posing a serious challenge to network bandwidth and transmission efficiency. With the emergence of new scenarios such as unmanned aerial vehicle-assisted vehicle-road cooperation and intelligent traffic control, the network must flexibly manage resources to cope with frequent changes in demand.

[0003] Internet of Vehicles slicing technology has emerged as the key means to solve these problems. Internet of Vehicles slicing divides the physical network into multiple virtual network instances (slices) and allocates communication, computing, and storage resources independently according to the needs of different application scenarios, providing flexible and fine-grained Quality of Service (QoS) guarantees for various applications. This slicing technology can effectively manage resources in the Internet of Vehicles, ensuring that scenarios such as autonomous driving, fleet coordination, and remote driving can still obtain optimal service support even in the face of diverse resource demands. In 6G networks, slicing technology combines Ultra-Reliable Low-Latency Communication (URLLC), Enhanced Mobile Broadband (eMBB), and other technologies to further improve the resource utilization efficiency and service level of the Internet of Vehicles.

[0004] In order to further enhance the intelligence of Internet of Vehicles slicing, Deep Reinforcement Learning (DRL) methods are introduced to address complex resource allocation and management problems. DRL can continuously optimize resource allocation decisions through continuous interaction with dynamic environments, adapting to changes in demand in different scenarios, thereby improving the flexibility and adaptability of the network. At the same time, combined with the federated learning framework, DRL models can utilize data from each node for collaborative training in a distributed environment, avoiding the privacy and communication burden problems caused by large-scale centralized data collection. Through federated learning, Internet of Vehicles slicing can support efficient operation of large-scale Internet of Vehicles while ensuring data privacy, further enhancing the service capability and intelligence level of the Internet of Vehicles.

[0005] Although the DRL and federated learning framework are introduced, the intelligentization of resource slicing in V2X still faces the following challenges:

[0006] 1. Personalized slicing window adjustment. There are differences in network situations in different base station coverage areas, and the prosperity of different areas leads to different traffic sizes. Even in the same block, traffic changes due to traffic peaks and valleys. These problems force the adjustment of the slicing window to be personalized. If the slicing window is too short, the cloud controller needs to allocate resources frequently, resulting in a significant increase in control and computing costs. If the slicing window is too long, the fluctuation of task traffic may significantly damage the performance isolation effect of slicing. The existing technology mentions a software-defined network framework with a hierarchical structure, which divides multiple slicing windows of equal length and calculates the optimal resource allocation strategy in each slicing window. The existing technology mentions a hierarchical soft RAN slicing framework for differentiated service provision, which performs network-level and BS-level resource slicing at large and small time scales. However, in the above methods, resources are allocated in fixed slicing windows.

[0007] 2. Dual optimization of resource slicing agility and timeliness. Resource slicing allocation decisions are made after slicing window division, which means that each decision has a certain time interval, and the decision effect usually needs to be verified for a period of time. The existing technology mentions using reinforcement learning to optimize resource slicing decisions. However, under the traditional reinforcement learning framework, resource slicing decisions rely on real-time state information. In the event of a sudden situation, it is still impossible to adjust the slicing length in a timely manner.

[0008] 3. Optimization of model scalability and generalization in road network environment. In the road network environment, even the traffic in adjacent road network areas has significant spatiotemporal differences. For example, during peak hours, adjacent road segments of a two-way lane may experience extreme congestion on one side and sparse traffic on the other. Under the huge differences in road network environment, traditional federated aggregation may cause the model to be biased towards a certain traffic pattern, ignoring other patterns. The existing technology mentions using traditional attention methods to handle differences in data from different base stations. However, in the case of high dynamics and non-uniformity in V2X, the traditional attention method has limited effect. By introducing a multi-head attention mechanism, different weights can be assigned to each node during aggregation, enhancing the generalization of the global model. The existing technology mentions using a deep reinforcement learning model as a local model to participate in federated learning. This approach not only learns the best strategy in different scenarios in the local model, but also aggregates the advantages of these strategies into the global model through federated aggregation, thereby improving the performance of the entire system. However, the above methods do not simultaneously solve the problem of model generalization. SUMMARY

[0009] In view of the challenges and limitations in the prior art, the present application proposes a vehicle networking resource slicing method based on Transformer dual federal learning, in which the physical resources of base stations are abstracted into a virtual resource pool managed by a controller in a vehicle networking system; according to the terminal request information fed back by each base station, the controller adjusts the slicing resource allocation;

[0010] The physical resource blocks (RBs) on each base station are virtualized into multiple slices and a resource sharing area, and when the resources of a certain slice are insufficient, the resources of the sharing area are temporarily occupied;

[0011] The controller is a dual-layer controller structure, and the centralized controller and the local controller in the dual-layer controller are respectively deployed in the cloud and the edge side;

[0012] The centralized controller integrates a global model, which includes a global traffic prediction model and a global resource slicing decision model; the centralized controller performs base station resource slicing through the global model;

[0013] The local controller connected with the adjacent base station integrates a local model, which includes a local traffic prediction model, a slicing window optimizer and a local resource slicing decision model; the local controller makes resource slicing decisions through the local model;

[0014] The slicing step includes:

[0015] (I) adopting a window adaptive slicing method to slice the physical resources of each base station;

[0016] (II) optimizing the decision results through the local model training and dual federal model method.

[0017] The dual federal learning framework based on the Transformer multi-head attention of the present application supports resource slicing in large-scale, diversified and high-dynamic vehicle networking scenarios. The present application has the following beneficial effects:

[0018] First, to cope with the fluctuations and uneven spatio-temporal distribution of task traffic, a flexible resource slicing framework with window personalization is designed. A soft index is used to measure the slicing performance isolation quality. The slicing isolation quality maximization problem is modeled as a joint optimization problem of slicing window adjustment, resource slicing and resource sharing area setting.

[0019] Second, a flow-aware personalized slice window adjustment strategy is designed to coordinate with the DRL (Deep Reinforcement Learning)-driven local resource slicing model. Unlike the "static" or region-centered "overall dynamic" slicing method, the slice window optimizer of each base station can flexibly adjust the window length according to the spatiotemporal characteristics of local traffic flow. In the case of heavy traffic or congestion, the slice window length will be temporarily shortened to speed up resource updating; in the case of sparse traffic or congestion dissipation, the window will be extended to reduce control overhead. The local DRL model effectively alleviates the problem of gradual invalidation of slice decisions in long time scales caused by environmental changes, with the support of slice window length and task flow prediction results.

[0020] Third, the resource slicing model of each base station is unified into a dual-federated learning framework that includes prediction and decision-making. At the prediction level, multiple base stations improve the accuracy of task flow prediction and optimize the adjustment efficiency of the slice window through collaborative training. The optimization in these two dimensions further promotes the effectiveness of resource slicing decisions by the local DRL model. At the decision-making level, the federated DRL paradigm amplifies the advantages of local optimization and improves the performance of overall slice management. The Transformer multi-head attention mechanism and its encoder-decoder are used to aggregate the prediction and decision-making models of adjacent base stations to enhance the generalization of the model. Simulation results show that the invention is superior to existing benchmark learning methods in terms of slice performance isolation, resource allocation efficiency, and decision-making effectiveness. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A scenario diagram representing a vehicular networking system in the specific implementation;

[0022] Figure 2 A resource allocation strategy is represented;

[0023] Figure 3 An edge-side model architecture that integrates prediction and decision-making on the local controller is represented;

[0024] Figures 4(a) to 4(c) Task flow fluctuations for three base stations in the same time period are represented respectively;

[0025] Figures 5(a) to 5(c) Fitting curves of slice window and traffic fluctuations for three base stations are represented respectively;

[0026] Figure 6 A dual-federated learning framework for cloud-edge collaboration is represented;

[0027] Figures 7(a) and 7(b) respectively represent reinforcement learning reward convergence analysis, wherein: Figure 7(a) represents the reward convergence curves corresponding to three learning rates, and Figure 7(b) represents the probability density curves corresponding to three learning rates;

[0028] Figures 8(a) to 8(c) Fig. 8 shows the influence of static policy on performance isolation for three base stations, respectively;

[0029] Figures 9(a) to 9(c) Fig. 9 shows the influence of overall dynamic policy on performance isolation for three base stations, respectively;

[0030] Figures 10(a) to 10(c) Fig. 10 shows the influence of personalized slice window length allocation policy on performance isolation for three base stations, respectively;

[0031] Figures 11(a) to 11(c) Fig. 11 shows the influence of resource block quantity on performance isolation for three base stations, respectively;

[0032] Figure 12 Fig. 12 shows the influence of different slice window policies on decision efficiency. DETAILED DESCRIPTION

[0033] The application proposes a dual-federated learning framework based on Transformer multi-head attention to support efficient management and optimization of resource slicing.

[0034] Firstly, a flexible resource slicing framework with window personalization is designed, which quantifies slice isolation performance through soft indicators and models the maximization of slice isolation quality as a joint optimization problem of slice window adjustment, resource slicing, and resource sharing area setting.

[0035] Secondly, a flow-aware personalized slice window adjustment strategy is proposed, which works with a local deep reinforcement learning (DRL) model at the base station to dynamically adjust the window length to adapt to local traffic fluctuations, improving the timeliness and stability of slice decision-making.

[0036] Finally, to overcome the limitations of traditional federated averaging or attention aggregation methods in effectively dealing with task differences between adjacent base stations, the Transformer multi-head attention mechanism is introduced to improve the generalization performance of the global model. At both the prediction and decision-making federated levels, deep collaboration between traffic prediction and slice optimization is achieved.

[0037] Simulation results show that the proposed scheme outperforms existing methods in terms of slice isolation performance, decision efficiency, and resource utilization.

[0038] 1 System Model

[0039] With network functions virtualization (NFV), the physical resources of a base station are abstracted into a virtual resource pool, which is managed by a controller based on software-defined networking (SDN). According to the terminal request information fed back by each base station, the controller dynamically adjusts the slice resource allocation. As shown in FIG. 1, the present application considers a two-layer controller structure. The centralized controller is deployed in the cloud and uses global information to slice the base station resources; the local controller is connected to the adjacent base stations and makes quick decisions according to local information. This multi-level resource management system can improve the network agility while taking into account global performance optimization. Figure 1

[0040] The parameter settings used herein are shown in Table 1.

[0041] Table 1: Symbol and variable definitions

[0042]

[0043]

[0044] 1.1 Flexible slicing framework with window adaptation

[0045] The physical resource blocks (RBs) on each base station are virtualized into multiple slices and a resource sharing area, each slice supporting a customized service, and the sharing area is used to relieve transient resource pressure. The traffic data received by the base station is divided into different task flows, and the arriving tasks are scheduled to the corresponding slice for processing according to the task type. When the resources of a certain slice are insufficient, the resources of the sharing area can be temporarily occupied.

[0046] A window-adaptive slicing method is used to slice the physical resources of each base station. Taking Figure 2 as an example to explain the slice window adjustment and resource allocation strategy. It is assumed that time is divided into multiple consecutive slice windows, each slice window containing multiple discrete time slots. The set and cardinality of base stations are denoted as and N, respectively. The set and cardinality of the service slices of base station n are denoted as and M n , respectively. The set and cardinality of slice windows are denoted as and H n , respectively. At base station n, the slice window contains a set of time slots and a cardinality, which are denoted as and At the end of slice window h-1, the local controller on base station n re-determines and M n .​

[0047] The total number of RBs held by base station n is recorded as B n In slice window h, the number of RBs allocated to slice m and shared area is expressed as and satisfy

[0048]

[0049] At the beginning of the slice window h, the resources of each slice and shared area resources Within the window h, the number of RBs held by each slice and the shared area remains unchanged.

[0050] 1.2 Soft Isolation Metrics

[0051] Performance isolation is the premise for the coexistence of slice resources, that is, to ensure that the overload of one slice does not affect other slices. The actual resource requirement is The number of RBs in the resource sharing area on base station n occupied by this slice is calculated as

[0052]

[0053] The present invention introduces a soft performance isolation metric. Assume that at the beginning of window h, the number of RBs reserved by base station n for the shared area is Base station n in time slot The performance isolation quality of is normalized to

[0054]

[0055] exist (i.e., occupying all shared areas cannot meet resource requirements), the slice performance isolation is violated, and at this time ρ n,t = 0. (ie: the shared area is sufficient to cope with resource pressure), if The resource blocks in the shared area will be temporarily occupied, ρ n,t ∈[0,1), which means that the performance isolation is “discounted”; when When slice m can meet the current demand without occupying the resources of the shared area, then ρ n,t =1.

[0056] 1.3 Problem Modeling

[0057] For base station n, under the slice window h (given ), the slice performance isolation quality maximization problem is modeled as And put the problem Setting as constraint (1) is

[0058]

[0059] Constraint (1) ensures that the total resource blocks allocated to slices and shared region do not exceed the total resource blocks held by the base station. The problem is formulated as Extending to multiple consecutive slice windows, the long-term optimization problem of performance isolation quality with respect to base station n is modeled as

[0060]

[0061] Problem The essence of the problem is how to adjust the slice window length and allocate resources to slices and shared region to maximize the long-term performance isolation quality. Section 2, Section 3 discuss how to optimize the decision results through local model training and federated model algorithm.

[0062] 2 Edge-side model training

[0063] At the edge side, each local controller integrates a traffic prediction model, a slice window optimizer and a resource slicing model. The traffic prediction and resource slicing use stacked LSTM and DRL algorithm respectively. This traffic prediction model is responsible for predicting the task traffic of the base station. Take base station n in Figure 3 as an example, its traffic prediction result (denoted as X n ) is input to the slice window optimizer to determine Based on X n and The DRL model determines the resource allocation of slices and shared region, and the corresponding decision vector is denoted as

[0064] 2.1 Local task traffic prediction

[0065] The historical traffic data on base station n is used to train a local model based on stacked LSTM to predict the traffic in the future period of time. By increasing the depth of the neural network, the stacked structure enhances the fitting ability of complex linear and nonlinear functions. The bottom layer network captures short-term patterns, while the high layer network captures longer-term trends and high-level features by aggregating these bottom layer information.

[0066] The training round set and cardinality of base station n local prediction model are denoted as and O n , and the local parameters of the prediction model are denoted as At the beginning of training, base station n initializes the local parameters of stacked LSTM Assume the number of layers of stacked LSTM is G. The input of the first layer is the original data sequence X = [x1, x2,..., xO ], the output is where represents the o-th hidden state of the o-th layer in the G-th iteration of the o-th training epoch. The output of the G-th layer is n the vector is the hidden state of the first layer. The output of the G-th layer is The prediction model at base station n is denoted as X(·). The number of RBs required by slice m at base station n in window h is predicted as

[0067]

[0068] According to the predicted value of the model output during the training process and the target value X, the loss function is described as

[0069]

[0070] The Back Propagation Through Time (BPTT) algorithm is used to calculate the gradient of the back propagation process. Based on (6), the gradient with respect to the local parameter is calculated as Using SGD, is updated as

[0071]

[0072] where η n is the learning rate.

[0073] 2.2 Slice window personalized adjustment

[0074] As mentioned earlier, the requests issued by vehicles are unevenly distributed in the space-time dimension and the proportion of each type of task flow continues to fluctuate, even in adjacent road network areas. In the fixed slice window mode, the update of the slice resource cannot be synchronized with the network situation, not only leading to a decline in isolation quality, but also producing unnecessary control overhead. To this end, a slice window personalized strategy centered on the base station is proposed.

[0075] Before designing the strategy, the invention collects and visualizes the task flow of three adjacent base stations in the road network in the same period. As Figures 4(a) to 4(c) shown, although adjacent, the fluctuations in task flow on these base stations are significantly different. The invention has preset different initial window lengths for these base stations, and tentatively reduces the window length to explore the optimal matching between slice window length and flow. The numerical pair about task fluctuation and optimal slice window length is observed and recorded. In order to avoid randomness, multiple initial time points are selected. The window length is expressed as a function of the flow, and their relationship is fitted as

[0076]

[0077] The premise of the best fit is to find c1 and c2 that minimize the residual sum of squares (c1 determines the speed of the window length change with the fluctuation of the task flow; c2 represents the base value of the window length when the fluctuation of the task flow is zero or very small). As shown in FIG. 3, the slice window length gradually decreases with the increase of the fluctuation of the task flow. The more intense the fluctuation of the task flow, the smaller the slice window length, and the faster the resource update frequency. These characteristics are consistent with expectations. At the beginning of the window h, the fitting function is used to predict the length of the next window h+1 (i.e., Lh+1= f (Xh) ). Due to the difference of the base station task flow, it is difficult to determine the window length of each base station with a unified fitting function. Figures 5(a) to 5(c)

[0078] In order to realize the personalized window adjustment, the application sets a window optimizer for each base station to cope with the fluctuation of the task flow of different base stations. Each optimizer receives the prediction result Xhfrom the prediction model and then brings it into formula (7) to obtain the corresponding slice window length result Lh. n

[0079] 2.3 Local resource slicing model

[0080] The division of the slice window length is autonomously determined on each base station. Under a given window, the resource allocation of the slice and the shared area is also independent. Therefore, the long-term optimization problem can be simplified as a short-term optimization under a discrete window, which belongs to a Markov Decision Process (MDP) problem in a finite time domain. In this problem, resource slicing refers to the process in which a local controller divides the resources held by the base station into M n slices and a shared area under a given window length to maximize the slice performance isolation quality.

[0081] Under the DRL framework, resource slicing decisions depend on the current environment state. Unlike typical DRL applications such as path planning, resource slicing is centered on a window. That is, the decision output by DRL will continuously produce benefits for a certain period of time. As time goes by, the decision will inevitably lose its effectiveness due to changes in demand. In this regard, a local resource slicing method based on the fusion of prediction and decision is proposed to "extend" the sustainability of resource slicing decisions, so that the decision has the maximum effectiveness under a given window. In this method, the output results of the flow prediction model in section 2.1 and the window optimizer in section 2.2 are input into the DRL model to help it generate the best resource slicing decision.

[0082] ​​​The decision model of base station n is abstracted as a DRL agent. In the slice window h, the set and cardinality of iteration rounds are denoted as and E n respectively. The iteration round The received state of the environment is denoted as s n,e The resource allocation action facing the slice and shared area is denoted as a n,e The reward given by the environment is denoted as r n,e The transition probability is denoted as Pr(s n,e |s n,e+1 , {a n,e , a n,1 ,...,a n,2}) under the condition of state s n,e The environment state is updated to s n,e+1 The expressions related to training include:

[0083] State space: The prediction result X of the resource slice dependent stack LSTM on base station n n and the time length output by the window optimizer In the iteration round e, the corresponding state vector is denoted as

[0084]

[0085] Action space: In the iteration round e, the allocation action made by the agent in the resource slice is denoted as

[0086]

[0087] In the formula, the decision set of iteration round e is , that is, the controller allocates slice resource blocks to the base station according to the action. The decision variable of each action is 0, 1 or 2, which is determined by the current state.

[0088] Reward: The goal of the system is converted to maximize the reward by maximizing the performance isolation quality, which reflects the pros and cons of action execution.

[0089] The total reward value in the iteration round e is denoted as

[0090]

[0091] Wherein, the ρ n,t originated from formula (3) is used to quantify the reward obtained by completing the task at time slot t.

[0092] In the state space, the agent estimates the reward of each action and stores it into the Q table. The action value function is denoted as The maximum reward for each state in the Q table represents the maximum possible reward in the future. By querying the Q table, the agent determines the action with the maximum reward in each state as

[0093]

[0094] Substituting the Bellman equation into equation (11), the vector in the Q table is expressed as

[0095]

[0096] Where γ n is the learning rate, and ν is the greedy probability.

[0097] After the local model is trained, the window optimizer and decision model finally come into play. n and Assist the DRL agent to generate service slices and the allocation results of shared area RBs

[0098] With the help of the prediction model, the local DRL model training process is summarized as Algorithm 1. At each time step t, the traffic prediction value X at the current moment is n and the predicted value X' for the next time step n Integrate into state s n,e and s' n,e This allows the algorithm to foresee future traffic trends when making decisions. By using future traffic predictions, the algorithm can respond to upcoming traffic fluctuations in advance, allocate base station resources more efficiently, and avoid resource overload or idleness. n The greedy strategy selects action a n,e , with probability ε n Random selection Otherwise select After each action is taken, the state s' is updated by receiving the new traffic forecast value n,e , enabling the next decision to be based on the latest traffic information. It should be noted that the window optimizer and DRL model directly determine the resource slices on base station n, while the results of the traffic prediction model reflect the "duration of the state," optimizing the state vector required by DRL and increasing the benefits of each DRL model decision.

[0099]

[0100] 3. Optimization of Dual Federated Learning for Cloud-Edge Collaboration

[0101] As an enhancement to the edge-side model, this section develops a cloud-edge collaborative dual federated learning framework. Figure 6As shown, the framework contains two sets of federated learning workflows, respectively facing task traffic prediction and resource slicing. On one hand, the federated enhanced prediction model optimizes the decision of the local DRL model. On the other hand, the federated for local DRL further improves the decision performance and the scalability of the system. Considering the difference and complexity of the data of adjacent base stations in the road network, it is deployed on the cloud controller, and the multi-head attention mechanism is used to enhance the generalization and adaptability of the global model (including the prediction and decision two sets of models).

[0102] 3.1 Federated model training for traffic prediction

[0103] The cloud centralized controller adopts a serial federated mode to sequentially receive the parameters uploaded by the local model, aggregate and update the global model. Due to the differences in data volume, data distribution, and initialization model quality, the model parameters uploaded by each base station contribute differently to the construction of the global model. The encoder (Temporal Encoder, TE) and decoder (Temporal Decoder, TD) of the Transformer are integrated on the cloud controller to optimize the uploaded parameters in the aggregation process, as shown in FIG. 3. Figure 6 The TE encodes the parameters uploaded by base station n into an input sequence, which generates Q n , K n , V n after multi-layer stacking and adapts to the multi-head attention mechanism in the Transformer.

[0104] The global prediction model parameters are initialized as are used for local training by base station n and are updated as The similarity function is defined as where Under the multi-head attention mechanism, the weight of this local prediction model is calculated as

[0105]

[0106] The attention weight of the prediction model is optimized as

[0107]

[0108] Using the attention weight, the global prediction model parameters are updated as

[0109]

[0110] Thus, the training parameters of the stacked LSTM model of each base station are sequentially fed back to the global model, so that the global model is continuously optimized, and the optimized parameters are fed back to the local model to improve the performance and adaptability of the local model.

[0111] 3.2 Predictive enhanced federated DRL model training

[0112] Base station The local DRL model parameters of the iteration round e are represented as The global decision model parameters of the base station are initialized as The model is locally trained using to obtain the updated local model parameters Under the attention mechanism, the weight of base station n is calculated as

[0113]

[0114] Substituting equation (16) into the global decision model parameters, the global decision model parameters are updated as

[0115]

[0116] Under the optimization of federated learning, the results of the slice window optimizer on the base station are further improved. After receiving the optimized prediction results X' n and , the global decision model parameter aggregation process is summarized as algorithm 2. Each base station is sequentially updated, and the global model gradually receives updates from each base station. After the local training of each base station is completed, the parameters are updated by weighting the previous aggregation parameters through the multi-head attention mechanism, thereby generating new aggregation parameters. The parameters of the multi-head attention mechanism are aggregated after each round of training, and the final aggregation parameters are used to update the global Q network where the specific content of line (6) of algorithm 2 refers to the local parameter update process of algorithm 1.

[0117]

[0118] In the serial federated mode, the update of the model gradually integrates the information of each base station, dynamically adjusts the weight of each base station using the multi-head attention mechanism, so that the model can better adapt to the diversified environment. The parameter update algorithm of the global prediction model can improve the prediction effect of the local prediction model, thereby improving the decision performance of the local decision model and further optimizing the results of the experiment.

[0119] 4 Experimental design and result analysis

[0120] For performance evaluation, we conduct simulation experiments on a high-performance server. The server is configured with an Intel Core i9-14900K processor, 64 GB DDR5 5200 MHz memory, a 4 TB PCIe 4.0 solid state drive, a ASUS PRIME Z790-P WIFI D5 motherboard, and two Gigabyte RTX 4090 24 GB WindForce graphics cards. To simulate inter-node collaboration, the data and models of each ground base station and cloud controller are encapsulated in independent Docker containers. The ground base station container runs the LSTM, slice window optimizer, and DRL model. The cloud container runs the Transformer model for global optimization. These containers run on the same physical server. Containers interact through a virtual network, simulating message passing and collaboration in a real network environment. Docker Compose is used to dynamically increase the number of base station nodes, making it easy to test the effects of multi-base station collaboration. Such a setup allows us to flexibly adjust the number of nodes and simulate multi-base station road networks.

[0121] To observe the contribution of different strategies in the proposed scheme to overall performance, the proposed scheme is divided into three categories as shown in Table 3. Proposed-1 uses the multi-head attention method under a fixed slice window; Proposed-2 uses the multi-head attention method under an overall adjusted slice window; and Proposed-3 uses the multi-head attention method under a personalized slice window. The ten methods listed in Table 2 are selected as baseline methods. Baselines 1, 2, and 3 use personalized slice window adjustment methods;

[0122] Baseline-4 uses the multi-head attention method based on only using the DRL model; Baselines 5, 6, and 7 use fixed slice window adjustment methods; and Baselines 8, 9, and 10 use overall adjusted slice window adjustment methods. Based on the PEMS dataset, this paper selects several nodes with large traffic differences as experimental objects to simulate the vehicle task flow reaching situation in a real environment. The experimental parameter configuration is shown in Table 4.

[0123] Table 2

[0124]

[0125] Table 3

[0126]

[0127] Table 4 Experimental parameter configuration table

[0128]

[0129] 4.1 Reward convergence analysis

[0130] The first experiment evaluates the convergence of the proposed method under different learning rates by calculating the average cumulative reward of the policy network over multiple time windows. Sampling a large batch of tuples from the experience replay pool can ensure the stability of learning. First, set the sampling batch size Batch Size = 2048.

[0131] As shown in FIG. 7(a), when the learning rate γ is set to 0.005, the convergence curve grows slowly, and the policy network converges to a stable state after 600 updates; when γ = 0.01, the average cumulative reward converges to about 180 relatively quickly, and then slowly converges to 200; when γ is set to 0.025, the reward converges quickly to about 200 and tends to be stable. This shows that the size of γ will affect the convergence of the method proposed in this paper. When γ is set to 0.025, the method can balance the convergence speed and stability, and the average cumulative reward can be stabilized at about 200 after 200 updates. Therefore, γ = 0.025 is taken in the subsequent experiments. FIG. 7(b) is the probability density curve corresponding to the three learning rates.

[0132] 4.2 Slice window strategy analysis

[0133] Generally, the increase in the number of RBs helps to improve the performance isolation quality expectation, but the actual effect is largely determined by the resource allocation strategy. In many cases, once the resource allocation of a certain slice does not match the actual resource demand, the overall performance isolation effect will be destroyed.

[0134] Next, we further investigate the impact of different slice window division strategies on the performance isolation effect during the morning peak period. As shown in FIG. 8(a), the task flow received by adjacent base stations in the same period is very different. The static strategy uniformly divides the time axis into two slice windows. The overall dynamic strategy divides the time axis into three uneven slice windows by capturing the fluctuations in the number of tasks. Specifically, the size of the slice window gradually decreases as the number of vehicles reaching the task increases. As shown in FIG. 8(b), the static strategy uniformly divides the morning peak period into two slice windows. As shown in FIG. 8(c), the overall dynamic strategy divides the morning peak period into three slice windows with gradually decreasing lengths. Figures 4(a) to 4(c) Figures 8(a) to 8(c) Figures 9(a) to 9(c)

[0135] In contrast, the present application obtains a fitting function by analyzing the relationship between the slice window length and the size of the task flow fluctuation, which is used to predict the length of the next slice window The number of slice windows at different base stations during the morning peak period can be dynamically adjusted. Combined with Figures 4(a) to 4(c) Figures 10(a) to 10(c) ​​​​The slice window size division selection and the corresponding slice allocation strategy performance of the three base stations in the morning peak period according to the fitting function can be seen. It can be seen that the task flow growth trend received by base station 1 remains unchanged, so according to the fitting function and the growth of the task flow, it is divided into two slice windows, and the corresponding slice allocation strategy in the window also reflects the performance isolation expectation. It can be seen that as the task flow grows, the performance isolation expectation is decreasing, but the window length is adjusted, so that the performance isolation expectation is always maintained above 86%. The task flow of base station 2 has a significant upward trend in the early morning peak period, and then slowly rises as a whole under the severe fluctuations, so according to the fitting function and the fluctuation of the task flow, the morning peak period is divided into three slice windows, and the window length decreases in turn. It can be seen that the slice performance isolation expectation changes obviously at the window transition, but the overall level is high. The task flow of base station 3 is more severe than that of base station 1 and more gentle than that of base station 2, and there is a significant upward fluctuation in the middle, so it is divided into three slice windows to adapt to the task reception under different traffic.

[0136] From the above results, it can be seen that the adjustment of the slice window length is more obvious at the task flow peak period; when the task flow fluctuates gently, the slice window length can also be appropriately increased according to the adjustment of the slice window optimizer, so as to reduce the slice division times and save the overall overhead of the base station under the premise of ensuring the performance isolation quality expectation.

[0137] 4.3 Slice performance isolation analysis

[0138] This group of experiments investigates the influence of the total number of ground base station RBs on the slice performance isolation quality expectation, and tests 5 base stations. The physical RB number of all ground base stations is determined by the prediction result of the stacked LSTM model, and the vehicle arrival task quantity in the peak period of a day is combined with the historical slice window size to allocate physical RBs and slice windows for the local base station. Figures 11(a) to 11(c) Select 3 base stations as experimental objects. Due to the complex road network structure, the task flow received by different base stations is very different (such as Figures 4(a) to 4(c) ), so the number of RBs is increased in steps. The number of RBs is tested from 145 to 342, and the influence of five different RB total numbers on the slice performance isolation quality expectation of the base station is tested. It can be seen that when the number of RBs of the base station is much smaller or much larger than the task flow received by the base station, the performance isolation quality expectation of the slice will be constant at 0 or 1. When the total number of RBs is at a certain value, the performance of different base stations under different total number of RBs can be seen. When the total number of RBs is 310, the difference in performance isolation quality expectation of base stations 1, 2 and 3 can be clearly seen. In the B n= 310, the performance isolation quality of the method used by base station 1 is all above 0.8, while that of base station 2 is all between 0.1 and 0.2, and that of base station 3 is distributed around 0.6-0.8. This phenomenon is caused by the great difference in the vehicle task flow received by the three base stations.

[0139] 4.4 Decision efficiency analysis

[0140] The influence of different window division strategies (static, overall dynamic or dynamic window) on the decision efficiency of training is observed in a long time scale. The decision efficiency refers to the number of times of triggering resource reallocation in a given period. The number of times is equivalent to the number of slice windows, which is determined by the slice window optimizer. As shown in Figure 12 , under the same performance isolation effect, the long-term decision cost of the dynamic window is lower than that of the static window. Although the decision cost of the dynamic window may be higher than that of the static window in a specific period (such as when the task flow in the coverage area fluctuates greatly), from the perspective of long-term optimization (such as within a day), the decision cost of the proposed dynamic window division strategy is lower than that of the static window. Compared with Proposed-1 which adopts the static window strategy, Proposed-3 has a performance improvement of (8.33%-17.39%) in decision cost. Compared with Proposed-2 which adopts the overall dynamic window strategy, Proposed-3 has a performance improvement of (8.33%-19.04%) in decision cost. Compared with Baseline-4 which adopts the individualized window strategy, Proposed-3 has a performance improvement of (13.63%-19.79%) in decision cost. As can be seen from the results of Figure 12 , the individualized window strategy can enhance the flexibility of RAN slices and balance the control overhead and service provision.

Claims

1. A vehicle network resource slicing method based on Transformer dual federated learning, characterized in that: include: The controller manages the physical resources of the base station; Based on the terminal request information fed back by each base station, the controller adjusts the slice resource allocation; The physical resource blocks (RBs) on each base station are virtualized into multiple slices and a resource sharing area. When the resources of a slice are insufficient, the resources in the shared area are temporarily occupied. The controller is a two-layer controller structure, in which the centralized controller and the local controller are deployed on the cloud and edge sides respectively; The centralized controller integrates a global model, which includes a global traffic prediction model and a global resource slicing decision model. The centralized controller uses the global model to slice base station resources. The local controller connected to the neighboring base station is integrated with a local model, which includes a local traffic prediction model, a slice window optimizer, and a local resource slice decision model; The local controller makes resource slicing decisions based on the local model; The slicing steps include: (1) Adopting a window-adaptive slicing method to slice the physical resources of each base station, using soft indicators to measure the slice isolation quality, and modeling the slice isolation quality maximization problem as a joint optimization problem of slice window adjustment, resource slicing, and resource sharing area setting. (2) Optimizing decision results through local model training and dual federated model training methods; the dual federated model training methods include a federal model training method for traffic prediction and a prediction-enhanced federal resource slicing decision model training method.

2. The vehicle network resource slicing method based on Transformer dual federated learning according to claim 1 is characterized in that: The step (1) specifically includes: For base station n, the number of slice windows is H n , the time slots included in the slice window h are T n (h) At the end of the slice window h-1, the local controller on base station n re-determines T n (h) and the number of serving slices M n ; The total number of resource blocks (RBs) held by base station n is recorded as B n ; In the slice window h, the number of RBs allocated to slice m and the shared area are expressed as and satisfy At the beginning of slice window h, the resources of each slice and the shared area are redistributed; within window h, the number of RBs held by each slice and the shared area remains unchanged; Performance isolation requires that the overload of one slice does not affect other slices; if the actual resource demand of slice m in time slot t is The number of RBs in the resource sharing area on base station n occupied by this slice is calculated as Assume that at the beginning of window h, the number of RBs reserved by base station n for the shared area is The performance isolation quality of base station n at time slot t is normalized to exist That is, when all shared areas cannot meet resource requirements, the slice performance isolation is violated. At this time, ρ n,t =0; That is, if the shared area is sufficient to cope with resource pressure, The resource blocks in the shared area will be temporarily occupied, ρ n,t ∈[0,1); when When slice m can meet the current demand without occupying the resources of the shared area, then ρ n,t =1; For base station n, given under slice window h The slice performance isolation quality maximization problem is modeled as the problem At the same time, the problem Set as constraint (1), Constraint (1) ensures that the resource blocks allocated to the slice and shared area do not exceed the total number of resource blocks held by the base station; Extending to multiple consecutive slice windows, the long-term optimization problem of the performance isolation quality of base station n is modeled as question The essence of this is how to adjust the slice window length and allocate resources between slices and shared areas to maximize the long-term performance isolation quality.

3. The vehicle network resource slicing method based on Transformer dual federated learning according to claim 1 is characterized in that: In step (2), local model training specifically includes: Use the traffic prediction model of the local controller of base station n to predict the task traffic of the base station. The prediction result is expressed as X n ;X n Input to the slice window optimizer to determine the time slot of the slice window h That is, the length of the slice window; based on X n and The local resource slice decision model determines the resource allocation of slices and shared areas, and the corresponding decision vector is expressed as The traffic prediction model specifically includes: The historical traffic flow data at base station n is used to train a local traffic flow prediction model based on a stacked long short-term memory neural network (LSTM) to predict traffic flow in the future. During the training of the local traffic flow prediction model, the set of training rounds o and the cardinality are represented as and O n , the local parameters of the model are expressed as At the beginning of training, base station n initializes the local parameters of the stacked LSTM Assume that the number of stacked LSTM layers is G; the input of the first layer is the original data sequence X = [x1, x2, ..., x O ], the output is in Represents the oth training round vector The hidden state of the first layer; the output of the G layer is The local traffic prediction model is denoted as X(·); the number of resource blocks required for slice m on base station n in window h is predicted as According to the predicted value output by the model during training and target value X, the loss function is described as The gradient of the back propagation process is calculated using the time series back propagation algorithm; about the local parameters The gradient of Using the stochastic gradient descent algorithm SGD, Updated to where η n is the learning rate; A slice window optimizer is set for each base station; each optimizer receives the prediction result X n Then, substitute into the formula Get the corresponding slice window length result The long-term optimization problem of performance isolation quality is simplified to a short-term optimization problem under discrete windows, which is a Markov decision process MDP problem with finite time domain; in this problem, a local controller divides the resources held by base station n into M under a given window length. n slices and a shared area to maximize slice performance isolation quality; The output of the local traffic prediction model and window optimizer is fed into the DRL model to generate the best resource slicing decision; The resource slicing decision model of base station n is abstracted as a DRL agent; under the slicing window h, the set and cardinality of the model iteration rounds are represented as ε n and E n ; Iteration round e∈ε n The receiving state of the environment is recorded as s n,e , the resource allocation action for slices and shared areas is recorded as a n,e , the reward given by the environment is recorded as r n,e ; The model is in state s n,e Under the condition of n,e+1 |s n,e ,{a n,1 ,a n,2 ,...,a n,e }), the environment state is updated to s n,e+1 ; The resource slice on base station n depends on X n and In the iteration round e, the state vector of the state space is expressed as In the iteration round e, the allocation action of the action space made by the agent when slicing resources is recorded as in is the decision set of iteration round e, i.e., the local controller allocates slice resource blocks to the base station according to the action; the decision variable of each action is 0, 1, or 2, determined by the current state; The total reward value in iteration round e is recorded as where ρ n,t is used to quantify the reward obtained by completing the task in time slot t; In the state space, the agent estimates the reward for each action and stores it in the Q table; the action value function is recorded as The maximum reward for each state in the Q table represents the maximum possible reward in the future. By querying the Q table, the agent determines the action with the maximum reward in each state. Among them, a represents all possible actions; Substituting the Bellman equation into the following formula, the vector in the Q table is expressed as Where γ n is the learning rate, ν is the greedy probability; After the local model is trained, X n and Assist the DRL agent to generate service slices and shared area RB allocation results.

4. The vehicle network resource slicing method based on Transformer dual federated learning according to claim 1 is characterized in that: In the step (2), The federated model training method for traffic prediction specifically includes: The centralized controller uses a serial federation mode to sequentially receive parameters uploaded by local models, aggregate them, and update the global model; The centralized controller integrates the Transformer encoder TE and decoder TD to optimize the parameters uploaded during the aggregation process; TE uploads the parameters of base station n Encoded into an input sequence, which is stacked over multiple layers to generate Q n , K n , V n , and adapted to the multi-head attention mechanism in Transformer; The global traffic prediction model parameters are initialized as It is used by base station n to train the local prediction model and is updated to Define the similarity function in Under the multi-head attention mechanism, the weight of the local traffic prediction model is calculated as Substituting the similarity function into the above formula, the attention weight of the prediction model is optimized as Using the attention weights, the global prediction model parameters are weighted and updated as The training parameters of each base station's local traffic prediction model are fed back to the global model in sequence. The global model is continuously optimized, and the optimized parameters are fed back to the local model. The prediction-enhanced federated resource slice decision model training method specifically includes: The parameters of the local resource slicing decision model of base station n in iteration round e are expressed as The global resource slicing decision model parameters of the base station in the system are initialized as Local resource slicing decision model usage Perform local training to obtain updated local model parameters Under the attention mechanism, the weight of base station n is calculated as The global resource slicing decision model parameters are weighted and updated as After receiving the optimized prediction result X' n and After that, the global resource slice decision model parameter aggregation process is as follows: each base station is updated in sequence, and the global resource slice decision model receives updates from each base station; after the local model training is completed at each base station, the parameters are updated. Through the multi-head attention mechanism and the previous aggregation parameters Perform weighted updates to generate new aggregation parameters; after each round of training, aggregate the parameters of the multi-head attention mechanism and use the final aggregation parameters Used to update the global Q network 5. The vehicle network resource slicing method based on Transformer dual federated learning according to claim 3 is characterized in that: The training process of the local resource slicing decision model specifically includes: At each time step t, the current traffic forecast value X n and the predicted value X' for the next time step n Integrate into state s n,e and s' n,e This allows the DRL model to foresee future traffic trends when making decisions; By using future traffic predictions, the local resource slicing decision model can react to upcoming traffic fluctuations in advance and allocate base station resources efficiently. If, according to ε n The greedy strategy selects action a n,e , then with probability ε n Randomly select actions Otherwise, select Action After each action is taken, the state s' is updated by receiving the new traffic prediction value n,e .

Citation Information

Patent Citations

  • Unmanned aerial vehicle small base station cluster RAN slicing method based on hierarchical federated learning

    CN114997737A

  • Internet of vehicles multi-dimensional resource allocation method and system based on federated learning

    CN115080249A