Task Offloading and Resource Allocation Method and System Based on Federated Self-Supervised Learning

By adopting task offloading and resource allocation methods based on federal self-supervised learning in the autonomous driving system, the problem of insufficient resource allocation in the existing technology is solved, efficient task offloading and resource optimization allocation is achieved, and the overall performance and energy efficiency of the system are improved.

CN119396600BActive Publication Date: 2025-05-30JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510009607.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-30
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

When the prior art uses federal self-supervised learning to classify images during autonomous driving, resource allocation issues are not fully considered, resulting in inefficient resource use and affecting the overall performance of the system.

Method used

A task offloading and resource allocation method based on federated self-supervised learning is provided. By obtaining the global model stored by the base station and the allocated transmission power, CPU frequency and task offloading ratio, a local model of the target vehicle is built, and the number of task offloading iterations and untrained iterations are determined based on the CPU frequency and task offloading ratio, efficient offloading of tasks and optimized allocation of resources are achieved.

Benefits of technology

By optimizing task offloading and resource allocation, resource utilization and computing efficiency are improved, system energy consumption is reduced, and the identification accuracy and generalization capabilities of the autonomous driving system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119396600B_ABST
    Figure CN119396600B_ABST
Patent Text Reader

Abstract

The present invention provides a task offloading and resource allocation method and system based on federated self-supervised learning, belonging to the technical field of edge computing in vehicle-to-everything (V2X) networks. The method includes constructing a first local model of a target vehicle; determining the number of task offloading iterations and the number of untrained iterations; determining a second local model of the target vehicle; offloading local data to a roadside unit (RSU) to determine a third local model; determining the system state of the target vehicle, the action and reward of a target time slot to update the Soft Actor-Critic (SAC) network; determining a global model of a base station; and performing resource allocation according to the final SAC network and task offloading according to the final global model. The present invention reduces system energy consumption and improves the offloading efficiency and accuracy of federated self-supervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle network edge computing, and particularly relates to a task offloading and resource allocation method and system based on federated self-supervised learning. Background Art

[0002] With the rapid popularization of Internet of Things (IoT) devices and the growth of intelligent transportation system requirements, Internet of Vehicle (IoV) has become an important part of modern transportation systems. Vehicle Edge Computing (VEC) solves this problem by deploying computing resources at the network edge, especially Road Side Unit (RSU).

[0003] Self-supervised learning can utilize a large amount of unlabeled image data to achieve image recognition and classification in the field of autonomous driving. Through self-supervised learning methods, the system can automatically learn effective feature representations from unlabeled data, reduce the dependence on manual labeling, and thus reduce the data labeling cost. Self-supervised learning is particularly crucial in complex environment perception tasks, which helps to improve the recognition accuracy and generalization ability of autonomous driving systems for various scenarios, providing important support for enhancing the safety and reliability of vehicles in actual road environments. At the same time, this also significantly increases the computing load and resource consumption in IoT scenarios.

[0004] Resource allocation and task offloading are important application scenarios of VEC. By offloading some computing tasks to edge servers, vehicles can effectively reduce their own computing load, optimize resource utilization, and improve computing efficiency without increasing latency. This method is particularly important in environments with low latency and high performance requirements, such as real-time decision-making and complex driving assistance systems. In addition, resource allocation algorithms can dynamically adjust edge computing resources to adapt to changes in vehicle mobility and network conditions, thus ensuring the stability and quality of service of the system.

[0005] Existing technologies usually adopt federated self-supervised learning for image classification during autonomous driving to protect vehicle privacy. However, these methods often do not fully consider the resource allocation problem and also ignore whether the vehicle has the ability to complete all local training tasks within each training cycle. Such limitations may lead to low resource utilization efficiency and affect the overall performance of the system in the case of limited computing and storage resources. Summary of the Invention

[0006] The present invention aims at the deficiencies in the prior art and provides a task offloading and resource allocation method and system based on federated self-supervised learning.

[0007] In a first aspect, the present invention provides a task offloading and resource allocation method based on federated self-supervised learning, including:

[0008] S1. Obtain the global model of federated training stored in the base station, as well as the first transmission power, CPU frequency, and task offloading ratio allocated by the base station to the target vehicle;

[0009] S2. Construct the first local model of the target vehicle according to the parameters and model structure of each layer of the global model;

[0010] S3. Determine the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and task offloading ratio;

[0011] S4. Determine the new local model of the target vehicle based on the first local model and local data, and use it as the second local model;

[0012] S5. According to the number of task offloading iterations, offload the local data to the RSU at the first transmission power to determine the local model of the RSU, and use it as the third local model;

[0013] S6. Upload the second local model, the third local model, the position information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the system state of the target vehicle, the action and reward in the target time slot;

[0014] S7. Update the SAC network according to the system state of the target vehicle, the action and reward in the target time slot;

[0015] S8. Determine the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in federated training;

[0016] S9. Repeat the operations of S1-S8 until the preset number of training rounds of federated training is reached, and obtain the final SAC network and global model, so as to perform resource allocation according to the final SAC network and perform task offloading according to the final global model.

[0017] Optionally, the determining the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and task offloading ratio includes:

[0018] Obtain the total number of training iterations required by vehicle n at time slot t ;

[0019] Determine the expected number of training iterations of the local model and the actual number of training iterations of the local model ;

[0020] Determine the total number of training iterations N on the RSU side R ;

[0021] Calculate the expected number of task offloading training iterations on the RSU side according to the following formula :

[0022] ;

[0023] where is a binary variable indicating whether vehicle n needs to offload part of its tasks to the RSU; when the task offloading ratio of vehicle n is less than the offloading ratio threshold, ; when the task offloading ratio of vehicle n is not less than the offloading ratio threshold, ;

[0024] Take and the minimum value of as the number of task offloading training iterations ;

[0025] Calculate the number of untrained iterations according to the following formula :

[0026] .

[0027] Optionally, determining a new local model of the target vehicle based on the first local model and local data as the second local model includes:

[0028] Calculate the information noise contrastive estimation loss according to the following formula :

[0029] ;

[0030] where exp(·) represents the exponential function with the natural constant as the base; is the augmented local data the anchor sample output after being input into the function χ composed of an encoder, two fully connected neural networks, and a ReLU function; is the augmented local data the positive sample output after being input into the function χ; is the augmented local data the negative sample output after being input into the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i≠j; π 1 (·) is the first data augmentation function; π 2 (·) is the second data augmentation function;

[0031] Calculate the loss of dual temperature according to the following formula :

[0032] ;

[0033] Among them, sg[·] represents the stop gradient; The hyperparameter representing the temperature is τ 2 The information noise contrast estimation loss at this time; The hyperparameter representing the temperature is τ 1 The information noise contrast estimation loss at this time;

[0034] Construct a second local model according to the loss of the dual temperature Expression:

[0035] ;

[0036] Among them, Represents the first local model; η r Is the learning rate; Is the gradient descent algorithm; Z is the total number of local data.

[0037] Optionally, uploading the second local model, the third local model, the position information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the global model of the base station, the system state of the target vehicle, the action and reward in the target time slot, including:

[0038] Calculate the number of iterations for local model training according to the following formula :

[0039] ;

[0040] Calculate the total energy consumption of each vehicle at time slot t according to the following formula:

[0041] ;

[0042] Among them, Is the total energy consumption of vehicle n at time slot t; Is the computing energy consumption of the RSU; Is the local computing energy consumption; Is the transmission energy consumption of the RSU; Is the energy consumed by vehicle n for one training iteration; Is the energy consumed by the RSU for one training iteration;

[0043] Construct the system state s t Expression:

[0044] ;

[0045] Among them, is the fading information, determined by the current location of the vehicle; ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in federated training;

[0046] Construct the action a at time slot t t Expression:

[0047] ;

[0048] where, is the power when all vehicles send data to the RSU; ; is the transmission power of vehicle n to transmit data to the RSU; is the CPU frequency allocated by the base station to all vehicles; ; is the CPU frequency allocated by the base station to vehicle n; is the task offloading ratio of all vehicles; ; is the task offloading ratio of vehicle n;

[0049] Construct the reward r t Expression:

[0050] ;

[0051] where, θ 1 is the weight of ; θ 2 is the weight of the overload amount ; θ 3 is the weight of ;

[0052] Optionally, determining the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in federated training includes:

[0053] Construct the global model of the base station Expression:

[0054] ;

[0055] where, represents the second local model; represents the third local model; N is the total number of vehicles participating in federated training.

[0056] Second aspect, the present invention provides a task offloading and resource allocation system based on federated self-supervised learning, including:

[0057] An acquisition module, configured to acquire the global model of federated training stored in the base station, as well as the first transmission power, CPU frequency, and task offloading ratio allocated by the base station for the target vehicle;

[0058] A construction module, configured to construct the first local model of the target vehicle according to the parameters and model structure of each layer of the global model;

[0059] A first determination module, configured to determine the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and task offloading ratio;

[0060] A second determination module, configured to determine a new local model of the target vehicle based on the first local model and local data, as the second local model;

[0061] A third determination module, configured to offload local data to the RSU at the first transmission power according to the number of task offloading iterations to determine the local model of the RSU, as the third local model;

[0062] A fourth determination module, configured to upload the second local model, the third local model, the location information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the system state of the target vehicle, the action and reward in the target time slot;

[0063] A network update module, configured to update the SAC network according to the system state of the target vehicle, the action and reward in the target time slot;

[0064] A fifth determination module, configured to determine the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in federated training;

[0065] A loop module, configured to repeatedly execute the operations of the acquisition module to the fifth determination module until the preset number of training rounds of federated training is reached, to obtain the final SAC network and global model, so as to perform resource allocation according to the final SAC network and perform task offloading according to the final global model.

[0066] Optionally, the first determination module includes:

[0067] An acquisition unit, configured to acquire the total number of training iterations required by vehicle n in time slot t ;

[0068] A first determination unit, configured to determine the expected number of training iterations of the local model and the actual number of training iterations of the local model ;

[0069] A second determination unit, configured to determine the total number of training iterations N on the RSU side R ;

[0070] A first calculation unit, configured to calculate the expected number of task offloading training iterations on the RSU side according to the following formula :

[0071] ;

[0072] wherein is a binary variable, indicating whether vehicle n needs to offload part of the tasks to the RSU; when the task offloading ratio of vehicle n is less than the offloading ratio threshold ; when the task offloading ratio of vehicle n is not less than the offloading ratio threshold ;

[0073] A third determination unit, configured to use the minimum value of and as the number of task offloading training iterations ;

[0074] A second calculation unit, configured to calculate the number of untrained iterations according to the following formula :

[0075] .

[0076] Optionally, the second determination module includes:

[0077] A third calculation unit, configured to calculate the information noise contrast estimation loss according to the following formula :

[0078] ;

[0079] wherein, exp(·) represents the exponential function with the natural constant as the base; is the enhanced local data the anchor sample output after being input into the function χ composed of an encoder, two fully connected neural networks, and a ReLU function; is the enhanced local data the positive sample output after being input into the function χ; is the enhanced local data the negative sample output after being input into the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i≠j; π 1 (·) is the first data augmentation function; π 2 (·) is the second data augmentation function;

[0080] The fourth calculation unit is used to calculate the loss of the dual temperature according to the following formula :

[0081] ;

[0082] where sg[·] represents stopping the gradient; The hyperparameter representing the temperature is τ 2 when the information noise contrastive estimation loss; The hyperparameter representing the temperature is τ 1 when the information noise contrastive estimation loss;

[0083] The first construction unit is used to construct the second local model according to the loss of the dual temperature Expression:

[0084] ;

[0085] where, represents the first local model; η r is the learning rate; is the gradient descent algorithm; Z is the total number of local data.

[0086] Optionally, the fourth determination module includes:

[0087] The fifth calculation unit is used to calculate the number of iterations of local model training according to the following formula :

[0088] ;

[0089] The sixth calculation unit is used to calculate the total energy consumption of each vehicle at time slot t according to the following formula:

[0090] ;

[0091] where, is the total energy consumption of vehicle n at time slot t; is the computing energy consumption of the RSU; is the local computing energy consumption; is the transmission energy consumption of the RSU; is the energy consumed by vehicle n for one training iteration; is the energy consumed by the RSU for one training iteration;

[0092] The second construction unit is used to construct the system state s t Expression:

[0093] ;

[0094] where, is the fading information, determined by the current location of the vehicle; ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in federated training;

[0095] The third construction unit is used to construct the action a at time slot t t Expression:

[0096] ;

[0097] Among them, is the power when all vehicles send data to the RSU; ; is the transmission power of vehicle n to transmit data to the RSU; is the CPU frequency allocated by the base station to all vehicles; ; is the CPU frequency allocated by the base station to vehicle n; is the task offloading ratio of all vehicles; ; is the task offloading ratio of vehicle n;

[0098] The fourth construction unit is used to construct the reward r t Expression:

[0099] ;

[0100] Among them, θ 1 is 's weight; θ 2 is the overload 's weight; θ 3 is 's weight.

[0101] Optionally, the fifth determination module includes:

[0102] The fifth construction unit is used to construct the global model of the base station Expression:

[0103] ;

[0104] Among them, represents the second local model; represents the third local model; N is the total number of vehicles participating in federated training.

[0105] The present invention provides a task offloading and resource allocation method and system based on federated self-supervised learning. In the method, the SAC algorithm is used to allocate the transmission power of vehicles to base stations and RSUs, the CPU frequency in vehicles, and the computing resource allocation ratio in RSUs. According to the allocated ratio, the computing resources available for each vehicle in the RSU are calculated. Then, according to the allowed computing resources, each local model training iteration task is divided into local training and RSU-assisted training. An offloading ratio threshold is set to optimize the task offloading efficiency. After offloading some iteration tasks to the RSU, the vehicle processes the remaining iteration tasks locally. If there are still unfinished iterations in a round of federated learning, the remaining number of iterations will be stored and wait for the next training, reducing the system energy consumption and improving the offloading efficiency and accuracy of federated self-supervised learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0106] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0107] Figure 1 It is a schematic flowchart of a task offloading and resource allocation method based on federated self-supervised learning provided by an embodiment of the present invention;

[0108] Figure 2 It is an application scenario diagram of the task offloading and resource allocation method based on federated self-supervised learning provided by an embodiment of the present invention;

[0109] Figure 3 It is a comparison result diagram of the SAC algorithm and other DRL methods in terms of energy consumption provided by an embodiment of the present invention;

[0110] Figure 4 Provided by an embodiment of the present invention Figure 3 Details in the round interval [1, 500];

[0111] Figure 5 It is a comparison result diagram of the SAC algorithm and other DRL methods in terms of the number of calculations provided by an embodiment of the present invention;

[0112] Figure 6 It is a comparison result diagram of the SAC algorithm and other DRL methods in terms of the reward function provided by an embodiment of the present invention;

[0113] Figure 7 Provided by an embodiment of the present invention Figure 6 Details in the round interval [1, 500];

[0114] Figure 8 Provided by an embodiment of the present invention Figure 6 Detail diagram in the round interval [2700, 3000];

[0115] Figure 9 Comparison result diagram of the SAC algorithm provided by an embodiment of the present invention and other DRL methods in terms of the number of overloads;

[0116] Figure 10 Comparison result diagram of the SAC algorithm provided by an embodiment of the present invention and other DRL methods in terms of the overload rate;

[0117] Figure 11 Comparison result diagram of the SAC algorithm provided by an embodiment of the present invention and other DRL methods in terms of improving the offloading efficiency;

[0118] Figure 12 Comparison result diagram of the energy consumption of the SAC algorithm provided by an embodiment of the present invention under different numbers of vehicles;

[0119] Figure 13 Comparison result diagram of the reward function of the SAC algorithm provided by an embodiment of the present invention under different numbers of vehicles;

[0120] Figure 14 Comparison result diagram of the test results of different DRL methods provided by an embodiment of the present invention;

[0121] Figure 15 Partial parameter and partial parameter value diagram provided by an embodiment of the present invention;

[0122] Figure 16 Schematic diagram of a task offloading and resource allocation system based on federated self-supervised learning provided by an embodiment of the present invention. Detailed implementation manner

[0123] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0124] The present invention aims to minimize energy consumption, and performs task offloading and resource allocation based on the task offloading ratio threshold and deep reinforcement learning. Considering a vehicle communication network in an urban scenario, the network consists of a connected base station (BaseStation, BS), roadside units (Road Side Unit, RSU) and N vehicles.

[0125] Embodiment 1.

[0126] As shown Figure 1 in the figure, this embodiment provides a task offloading and resource allocation method based on federated self-supervised learning, including:

[0127] S1. Obtain the global model of federated training stored in the base station, as well as the first transmission power, CPU frequency, and task offloading ratio allocated by the base station for the target vehicle.

[0128] As shown Figure 2 in the figure, the scenario is an urban VEC traffic system containing four intersections. In this scenario, a base station is located at a corner, a roadside unit is deployed on the roadside, and there are N vehicles driving on the road. Assume that all four intersections are within the coverage of the base station and the RSU. When vehicle n arrives at the intersection, it will turn left, turn right, or go straight through the intersection with probabilities φ 1 , φ 2 , and φ 3 respectively. The neural network configured on the base station is used to provide decision support. The RSU has fixed computing resources and can process the task offloading from vehicles in parallel.

[0129] The whole process is divided into two stages: offloading-based federated self-supervised learning (SSL) and resource allocation based on deep reinforcement learning (DRL). These two stages are independent of each other and yet influence each other. The offloading-based federated self-supervised learning stage contains K max segments. For each segment e, is divided into R max rounds. Similarly, the resource allocation stage based on DRL, i.e., the SAC (Soft Actor-Critic) algorithm, also contains K max segments, and each segment e is divided into S max time slots. For each round r (r ∈ [1, R max ) and time slot t (t ∈ [1, S max ), the duration is T. Each round of federation corresponds to a time slot. Just when the time slot t reaches S max , the segment e transfers to the next segment e + 1. At this time, the time slot t is reset to zero, while the round r continues to accumulate until the entire algorithm is completed. Therefore, the relationship between the round r and the time slot t in the segment e can be described as:

[0130] .

[0131] S2. Construct the first local model of the target vehicle according to the parameters and model structure of each layer of the global model.

[0132] Before each time slot t starts, the base station stores the global model obtained from round r-1 aggregation and uses DRL to allocate resources, including transmission power for data transmission to RSU and base station and , CPU frequency for local training , and the task allocation ratio of vehicle n When segment e = 1, in the first round / time slot, the base station randomly initializes the parameters of the global model and the DRL network. At the same time, vehicle n stores Z images captured in the previous round, called local data, and has a model with the same structure as the global model on the base station side, called the local model.

[0133] S3, determining the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and the task offloading ratio.

[0134] In each time slot t, vehicle n first downloads the stored data from the BS and distributes the parameters of the BS-side global model to the local model. Next, according to the CPU frequency of the download Calculate the theoretical total number of local training iterations for vehicle n , and calculate the actual total number of training iterations based on the available CPU time T' of vehicle n Finally, vehicle n is assigned a ratio according to the downloaded tasks (task offloading ratio) and task allocation ratio threshold q 0 (offloading ratio threshold) determines whether to offload the computing task to the RSU in time slot t. Specifically, when Greater than q 0 When , vehicle n will offload part of the task to the RSU; otherwise, the vehicle will not offload part of the task to the RSU and Reset to 0. If the decision is made to offload the task, vehicle n will immediately pass the downloaded transmission power (First transmission power) Send an unloading request to the RSU. The request includes the local model, local data, and the task allocation ratio of vehicle n in time slot t.

[0135] In this embodiment, the total task that vehicle n needs to process in time slot t is recorded as (Total number of training iterations), which includes the estimated number of training iterations of the local model in time slot t and the remaining number of training iterations not processed from the previous time slot When vehicle n offloads part of its task to the RSU, the total task The ratio will be allocated according to the assigned tasks The number of iterations that are divided into local model training And the number of training iterations offloaded to RSU. The total number of training iterations N on the RSU side can be calculated according to the following formulaR :

[0136] 。

[0137] Wherein, t max is the longest time for data transmission; is the time required during the process of vehicle n transmitting data to the RSU, i.e., the transmission delay; D is the size of the local model transmitted by each vehicle; in this embodiment, the size of the local model transmitted by each vehicle is the same; is the information transmission rate between the vehicle and the RSU; B n is the bandwidth occupied when vehicle n transmits data to the RSU; N 0 is the noise power; is the channel gain at time slot t.

[0138] The vehicles participating in the training communicate with the RSU using orthogonal frequency division multiple access (OFDM) technology, so the interference between vehicles will not be considered during the communication process. Each vehicle uses the same channel model in the same time slot, and as the time slot t enters time slot t + 1, the channel model will be updated. Therefore, at time slot t, the channel gain can be expressed as:

[0139] 。

[0140] Wherein, dis(R,n) is the distance between vehicle n and the RSU, with the unit of meter; represents small-scale fading and follows an exponential distribution with a unit mean, i.e., ; represents the influence of large-scale fading on the signal, including path loss and shadow fading . Then can be expressed as:

[0141] 。

[0142] Where the shadow fading follows a normal distribution with a mean of 0 and a standard deviation of 8, expressed as , and in a high-density urban environment, the path loss can be expressed as:

[0143] 。

[0144] Exemplarily, this step includes:

[0145] Obtain the total number of training iterations required by vehicle n at time slot t 。

[0146] Determine the expected number of training iterations of the local model and the actual number of training iterations of the local model 。

[0147] Determine the total number of training iterations N on the RSU side R 。

[0148] Calculate the expected number of training iterations for task offloading on the RSU side according to the following formula :

[0149] 。

[0150] Wherein, is a binary variable indicating whether vehicle n needs to offload part of the task to the RSU; when the task offloading ratio of vehicle n is less than the offloading ratio threshold, ; when the task offloading ratio of vehicle n is not less than the offloading ratio threshold, 。

[0151] If exceeds , this indicates that the expected offloaded task is too large and exceeds the total task volume that needs to be processed locally. In this case, the actual number of training iterations for offloading should be recorded as . Therefore, take the minimum value of and as the number of training iterations for task offloading 。

[0152] Calculate the number of untrained iterations according to the following formula :

[0153] 。

[0154] S4. Determine a new local model of the target vehicle based on the first local model and local data as the second local model.

[0155] The training process is federated self-supervised learning. First, input each local data into two different data augmentation methods π 1 (·) and π 2 (·). The remaining images are marked as negative samples. Input the augmented images and into a function χ composed of an encoder, two fully connected neural networks, and a ReLU function. The output of the function χ includes anchor samples , positive samples and encoded negative samples ; i ∈ [1, Z], j ∈ [1, Z].

[0156] Exemplarily, this step includes:

[0157] Calculate the Information Noise Contrastive Estimation (InforNCE) loss according to the following formula :

[0158] .

[0159] where exp(·) represents the exponential function with the natural constant as the base; is the enhanced local data The anchor sample output after being input into the function χ composed of an encoder, two fully connected neural networks, and a ReLU function; is the enhanced local data The positive sample output after being input into the function χ; is the enhanced local data The negative sample output after being input into the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i ≠ j; π 1 (·) is the first data augmentation function; π 2 (·) is the second data augmentation function.

[0160] Calculate the loss of dual temperature according to the following formula :

[0161] .

[0162] where sg[·] represents stopping the gradient; The hyperparameter representing temperature is τ 2 The Information Noise Contrastive Estimation loss at this time; The hyperparameter representing temperature is τ 1 The Information Noise Contrastive Estimation loss at this time.

[0163] Construct the second local model according to the loss of dual temperature Expression:

[0164] .

[0165] where, represents the first local model; η r is the learning rate; is the gradient descent algorithm; Z is the total number of local data.

[0166] S5. According to the number of task offloading iterations, offload the local data to the RSU with the first transmission power to determine the local model of the RSU and use it as the third local model.

[0167] After receiving the task offloading request of the vehicle, the RSU uses the local data and local model uploaded by vehicle n to perform a training session to generate a new model (i.e., the third local model). The training process is similar to the process of vehicle local training. After the training is completed, the RSU sends the trained model back to the vehicle.

[0168] S6. Upload the second local model, the third local model, the location information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the system state of the target vehicle, the actions and rewards in the target time slot.

[0169] After receiving all the trained models sent by the RSU, vehicle n uploads the locally trained model and the model trained by the RSU , the speed in the next time slot t+1, the total energy consumption, and the number of unprocessed training iterations to the base station through the downloaded transmission power.

[0170] During the training process, the energy consumed by vehicle n for one training iteration is calculated as follows:

[0171] .

[0172] Among them, , and respectively represent the power, calculation delay, and energy consumed for one iteration training when vehicle n performs local training locally; is the local CPU calculation frequency of the vehicle, which can be calculated by the dynamic voltage and frequency scaling (DVFS) method, i.e.:

[0173] .

[0174] Among them, is the effective switching capacitance, which depends on the chip architecture, represents the CPU frequency allocated to the local, and , and respectively represent the maximum and minimum CPU calculation frequencies. To calculate the local calculation delay of the vehicle, let c be the number of CPU cycles required to process a unit size of data, and Z be the size of the training data. Then cZ represents the number of CPU cycles required for training. Therefore, the calculation delay is calculated as follows:

[0175] 。

[0176] It can be obtained that:

[0177] 。

[0178] Similarly, the latency of each round of calculation on the RSU can be calculated and the energy consumed. :

[0179] 。

[0180] 。

[0181] Among them, represents the CPU frequency of the RSU calculation.

[0182] If the vehicle offloads part of the calculation task to the RSU, during the transmission process, the transmission energy consumption of vehicle n can be expressed as:

[0183] 。

[0184] Therefore, the energy consumed by vehicle n for transmission to the RSU can be expressed as:

[0185] 。

[0186] When the vehicle is performing local iterative tasks, if all iterative tasks cannot be completed due to limitations of local computing resources, part of the iterative tasks need to be offloaded to the RSU. This offloading process requires the base station to use the DRL algorithm to allocate the task allocation ratio and determine the computing resources allowed for each vehicle by the RSU. The total iterative tasks will be allocated to local training and training on the RSU. This optimization aims to improve resource utilization and task completion efficiency.

[0187] The base station allocates local computing power according to the DRL algorithm. Within a time slot t, vehicle n can obtain the total expected number of training rounds locally , that is:

[0188] 。

[0189] Among them, T represents the total duration of a time slot; T - t maxRepresents the total duration available for computing within a time slot. In practical applications, the CPU usually needs to process multiple tasks simultaneously to improve CPU utilization efficiency, such as system maintenance, the running of other applications, and background processes, etc. These tasks will occupy some CPU resources, resulting in the dispersion of computing power for training. Therefore, the actual time used for local training is less than T, denoted as T', and the total actual number of local training rounds Can be expressed as:

[0190] 。

[0191] Can be obtained:

[0192] 。

[0193] 。

[0194] To achieve efficient task offloading and resource allocation, the base station and the vehicle need to incorporate the long-term impact of their actions in the decision-making process. Therefore, this embodiment aims to minimize the sum of the energy consumption of all vehicles in each time slot. An optimization problem is established:

[0195] 。

[0196] Among them,

[0197] 。

[0198] 。

[0199] 。

[0200] 。

[0201] p min and p max Are the minimum transmission power and the maximum transmission power respectively.

[0202] After the base station allocates the offloading ratio, according to the actual computing power of the RSU, the number of offloading tasks assigned to each vehicle may exceed the number of tasks that the vehicle itself needs to process. The overload amount Measures the excess degree of the offloading training rounds.

[0203] The overload ratio Is used to describe the severity of the overload behavior, that is:

[0204] 。

[0205] As described above, when the offloading ratio Is lower than the specific threshold q 0When the vehicle n will not perform the unloading operation. Therefore, the unloading efficiency is defined as an index to measure the system's unloading task ability, that is:

[0206] .

[0207] The task unloading and resource allocation problem can be solved by the DRL algorithm. For this purpose, DRL is used to find the optimal task unloading and resource allocation scheme. The SAC algorithm is selected because it is robust and efficient in dealing with continuous action spaces and can effectively balance exploration and exploitation. Compared with other DRL algorithms, SAC introduces an entropy term in the objective function to encourage the base station to continuously explore new strategies. This entropy term effectively prevents the base station from prematurely converging to a local optimal solution during the policy iteration process, thus avoiding training failure. The goal of adopting the SAC algorithm is to minimize the overall energy consumption of the system, aiming to find an optimal resource allocation scheme.

[0208] Exemplarily, this step includes:

[0209] For the local computing ability of vehicle n, its actual total number of training iterations is . When does not exceed , vehicle n can process all remaining tasks locally, that is . When exceeds , it means that vehicle n cannot process all remaining tasks locally, and the maximum number of iterations that can be processed locally is . At the same time, the unprocessed part will be placed in the local buffer. Therefore, the number of iterations for local model training is calculated according to the following formula :

[0210] .

[0211] Calculate the total energy consumption of each vehicle at time slot t according to the following formula:

[0212] .

[0213] Among them, is the total energy consumption of vehicle n at time slot t; is the computing energy consumption of the RSU; is the local computing energy consumption; is the transmission energy consumption of the RSU; is the energy consumed by vehicle n for one training iteration; is the energy consumed by the RSU for one training iteration.

[0214] The base station can analyze the state and experience during the decision-making process. The system state consists of two elements: fading information and vehicle speed . Among them, the fading information is determined according to the current position of the vehicle, and the speed , , remains unchanged within each time slot and changes at the start of the next time slot. Among them, and represent the minimum and maximum values of the vehicle speed respectively. Therefore, the system state s t is constructed as an expression:

[0215] .

[0216] Among them, ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in federated training.

[0217] After obtaining the state, the base station makes decisions on task offloading and power allocation for the current time slot. The action is set as three discrete variables: the power when the vehicle sends data to the RSU, the CPU computing frequency allocated by the base station to the vehicle, and the task offloading ratio. t Therefore, the action a at time slot t

[0218] .

[0219] Among them, is the power when all vehicles send data to the RSU; ; is the transmission power of vehicle n to transmit data to the RSU; is the CPU frequency allocated by the base station to all vehicles; ; is the CPU frequency allocated by the base station to vehicle n; is the task offloading ratio of all vehicles; ; is the task offloading ratio of vehicle n.

[0220] The reward r t is constructed as an expression:

[0221] .

[0222] Among them, θ 1 is 's weight; θ 2 is the overload amount The weight of; θ 3 is the weight of.

[0223] S7, update the SAC network according to the system state of the target vehicle, the actions of the target time slot, and the rewards.

[0224] The SAC algorithm, i.e., the Soft Actor-Critic algorithm, is a reinforcement learning method based on the principle of maximum entropy. It consists of five networks: 1 actor , 2 critics and , and 2 targets and .

[0225] The core idea of the SAC algorithm is to balance the relationship between exploration and exploitation by increasing the randomness of the policy. Traditional reinforcement learning algorithms tend to get stuck in local optima. Even if they perform well in the initial stage, they may not be able to find the global optimal solution when facing complex environments. While SAC encourages policy diversity, enabling the agent to try more possible paths during training and enhancing the global search ability of the policy. The objective function of SAC can be expressed as the weighted sum of the expected reward and entropy, i.e.:

[0226] .

[0227] Where, is the expected reward for taking action in state , hereinafter abbreviated as r t . is the entropy of the policy π in state , , representing the discount factor. is the temperature parameter that regulates the randomness of the optimal policy and is used to balance the importance of the expected reward and entropy.

[0228] This way enables the agent to tend to maintain a certain degree of action randomness while pursuing high rewards, thus being able to better explore the environment and learn the policy. The purpose of this embodiment is to maximize , and the corresponding policy π is denoted as . At this time, using can adjust , obtain the optimal solution in state , denoted as , i.e.:

[0229] .

[0230] The base station obtains the initial state from the system, and then according to the actor network The policy selection action . After executing the action , the state is converted to a new state and a corresponding reward is obtained . These interaction data are stored in the replay buffer R in the form of tuples . The replay buffer R allows random sampling of these tuples, thus breaking the correlation between samples during the training process. When the number of tuples stored in the replay buffer R reaches the batch size , the five networks of the SAC algorithm start to be updated.

[0231] Randomly draw M tuples from the replay buffer R, which is called a mini-batch. For each tuple in the mini-batch , where , input into the actor network . According to the policy of the actor network , a new action is obtained. It should be noted that is different from in this tuple. Then the loss function of can be calculated as follows:

[0232] .

[0233] represents the dimension of the action , denoted as .

[0234] Next, update the actor network :

[0235] Input into the two critic networks and respectively, and obtain the corresponding action value functions and . The smaller of these two values is defined as . Therefore, through the Stochastic Gradient Descent (SGD) algorithm, the update process of the actor network is as follows:

[0236] .

[0237] where o is the noise sampled from the multivariate normal distribution; is a parameter for reparameterizing the action function. Finally, update the action network using the stochastic gradient descent method.

[0238] Next, update the two critic networks and :

[0239] To update the critic networks, first input the in the mini-batch M into the two critic networks and to obtain the action value pairs and . At the same time, input into the actor network to obtain the action under the policy . To maintain the randomness of the policy during training and explore a wider range of states and actions, the SAC algorithm introduces an entropy regularization term in the policy function, i.e., . Then, input into the two target networks and to obtain the target action value pairs and . Take the smaller of these two values and denote it as . Then, the target action value can be calculated as follows:

[0240] .

[0241] Then, update the network through the SGD algorithm:

[0242] .

[0243] .

[0244] Based on , , and , use the Adam optimizer to update the actor and critic networks every time slots. After the process of updating the actor and critic networks, enter the next time slot. It should be noted that every time slots, the parameters of the two target networks and also need to be updated as follows:

[0245] .

[0246] .

[0247] Among them, is much less than 1; and are respectively two updated target networks.

[0248] S8. Determine the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in the federated training.

[0249] After the base station receives the local models of all vehicles, it aggregates the received models. Exemplarily, construct the global model of the base station Expression:

[0250] .

[0251] Among them, represents the second local model; represents the third local model; N is the total number of vehicles participating in the federated training.

[0252] S9. Repeat the operations of S1 - S8 until the preset number of training rounds of the federated training is reached, and obtain the final SAC network and global model, so as to perform resource allocation according to the final SAC network and perform task offloading according to the final global model.

[0253] In this embodiment, the values or value ranges of some parameters are as Figure 15 shown.

[0254] As Figure 3 shown, it shows the comparison results of different methods with five vehicles participating in terms of energy consumption during the training process. The training methods include four DRL methods: SAC, DDPG (Deep Deterministic Policy Gradient), TD3 (Twin Delayed Deep Deterministic Policy Gradient), and PPO (Proximal Policy Optimization), and a method of always transmitting at the maximum power. In terms of energy consumption, the PPO algorithm fails to effectively reduce the energy consumption, and its energy consumption value continues to fluctuate between J. In contrast, the energy consumption of the SAC, DDPG, and TD3 algorithms shows a rapid downward trend and gradually converges to a relatively stable value. Specifically, the SAC algorithm has the fastest convergence speed and is more stable than the other two algorithms. The final energy consumption result is stable at J. This is mainly attributed to the fact that the SAC algorithm balances exploration and exploitation during the policy update process, improving the convergence efficiency. As Figure 4 shown, it details the details in the round interval [1, 500]. From Figure 4It can be clearly seen that the fluctuation of TD3 is smaller than that of DDPG, but the DDPG algorithm can explore lower energy consumption values. Finally, the energy consumption result of the DDPG algorithm stabilizes at J, while the energy consumption result of the TD3 algorithm stabilizes at J. This indicates that the TD3 algorithm performs better in dealing with noise and uncertainty, so it is more stable in the initial stage and reaches a lower energy consumption faster.

[0255] As Figure 5 shown, it presents the comparison results of different methods with five vehicles participating in terms of the number of calculations during the training process. The training methods include four DRL methods: SAC, DDPG, TD3, and PPO, as well as a method that always transmits at the maximum power. In terms of the number of calculations, all four algorithms finally converge to a certain stable value. The set goal is to perform as much local training as possible in each round. The results show that the number of calculations of the SAC algorithm gradually stabilizes and rises, and finally stabilizes at times, which is significantly higher than the other two algorithms. The final number of calculations of the DDPG and TD3 algorithms is similar, which are times and times respectively, but the fluctuation of the DDPG algorithm is larger. The number of calculations of the PPO algorithm is the largest, stabilizing at times. Theoretically, the global model trained with more calculations has a higher accuracy.

[0256] As Figure 6 shown, it presents the comparison of different methods with five vehicles participating in terms of the reward function during the training process. It can be seen from Figure 6 that the reward value of the PPO algorithm is the lowest, and the final reward value is the smallest among the four algorithms, stabilizing at . This indicates that it is not a wise choice. At the same time, it can also be seen from Figure 6 that all four algorithms can finally converge to a relatively stable reward value. As Figure 7 shown, it details the details in the round interval [1, 500]. It can be clearly seen that the reward value of the SAC algorithm converges rapidly and remains stable after 76 rounds. After 150 rounds, the DDPG and TD3 algorithms can also reach a relatively stable state. As Figure 8 shown, it presents the details in the round interval [2700, 3000]. It can be seen that the reward values of the SAC, DDPG, and TD3 algorithms finally stabilize at , and respectively. It can also be seen from Figure 8 that the reward values of the DDPG and TD3 algorithms tend to exceed that of the SAC algorithm, but in terms of stability, the SAC algorithm is still the best.

[0257] As Figure 9 shown, the number of offloading times under different deep reinforcement learning (DRL) algorithms is presented. It can be seen from Figure 9 that the number of offloading times of DDPG, TD3, and PPO algorithms shows a fluctuating state, while the number of offloading times of the SAC algorithm shows a trend of stable decline and finally converges.

[0258] As Figure 10 shown, the overload ratio under different DRL algorithms is presented. Except for the PPO algorithm, the overload rates of the other three algorithms show a downward trend. In particular, the overload rate of the SAC algorithm approaches 0 after the 1332nd round. Combining and Figure 9 and Figure 10 it can be seen that the SAC algorithm gradually reduces the number of offloading times during the training process and conducts more training locally on the vehicle, thereby reducing the overload rate and finally achieving a balance between local training on the vehicle and offloading to the RSU for training. Although the DDPG and TD3 algorithms do not significantly reduce the number of offloading times during the training process, they can still reduce the overload rate, indicating that these two algorithms tend to offload the training tasks to the RSU for processing and explore an appropriate offloading ratio. Combining Figure 3 it can be seen that this also explains why the energy consumption of these two algorithms is higher than that of the SAC algorithm. The number of offloading times of the PPO algorithm always remains at a high level among the four algorithms, but its overload rate is low and stable. This shows that the PPO algorithm also tends to offload data to the RSU for processing, and the number of offloading times per time is close to the expected number of processing rounds per time locally on the vehicle. The PPO algorithm tends to offload tasks to the RSU for processing, but the number of tasks offloaded each time does not far exceed the total number of local tasks.

[0259] As Figure 11 shown, the difference in offloading efficiency under the condition of setting a threshold and not setting a threshold is presented. From the overall data, it can be seen that all values are above 0, indicating that the offloading efficiency of setting a threshold is always better than that of not setting a threshold . This phenomenon verifies the threshold setting The method can significantly improve the offloading efficiency. Specifically, among all the tested reinforcement learning algorithms, the TD3 algorithm performs best in improving offloading efficiency, followed by the SAC and DDPG algorithms, while the PPO algorithm has the least improvement effect. The superior performance of the TD3 algorithm may stem from its double-delayed deep deterministic policy gradient mechanism, which effectively reduces the variance during the update of the policy network, thereby improving the accuracy of offloading decisions. The SAC algorithm, by introducing the concept of entropy to encourage exploration diversity, also improves the offloading efficiency to a certain extent. In contrast, although the PPO algorithm has a certain degree of stability and security, its policy update is relatively conservative, which limits its performance in complex environments.

[0260] As Figure 12 shown, the energy consumption of the SAC algorithm under different numbers of vehicles is compared. As Figure 13 shown, the reward function of the SAC algorithm under different numbers of vehicles is compared. From Figure 12 and Figure 13 it can be observed that the energy consumption curve shows a downward trend and gradually stabilizes, while the reward function shows an upward trend and gradually stabilizes. In addition, as the number of vehicles increases, the number of rounds required to reach the stable state also increases. This indicates that in a multi-vehicle scenario, task offloading and resource allocation become more complex, resulting in the need for more rounds to achieve convergence to the stable state. Generally speaking, as the number of vehicles participating in training gradually increases, the total energy consumption of the system also increases, while the reward value decreases.

[0261] As Figure 14 shown, four rounds of tests were conducted, and each round of tests included 50 independent experiments. The results of each round of tests were obtained by calculating the mean of the results of these 50 experiments. Finally, the means of these four rounds of tests were further averaged to obtain the average value of the comprehensive experimental results, and the final experimental results are shown in Figure 14 In Figure 14 's box plot, the three algorithms with better performance: SAC, DDPG, and TD3 algorithms were selected, and the results of these three algorithms were analyzed from three aspects: energy consumption, number of calculations, and reward function. In these three aspects, the test results of the model trained by the SAC algorithm have a lower degree of dispersion, indicating that the data is concentrated and the fluctuation range is small. In contrast, the test results of the model trained by the DDPG algorithm are relatively dispersed, with a certain degree of variability. Specifically, in Figure 14 's (a), although there is an outlier in the SAC algorithm, this outlier is relatively close to the lower whisker, and the overall data values are concentrated in a relatively small range, the overall distribution is relatively compact, and there are few outliers. Therefore, the impact of this outlier on the overall analysis can be ignored. In Figure 14In (b), no outliers appeared in all three algorithms. Moreover, it can be seen that the SAC algorithm obtained the most calculation times, which will contribute to the SAC algorithm achieving a higher classification accuracy in federated self-supervised learning. However, in Figure 14 In (c), outliers appeared in both the DDPG and TD3 algorithms. In particular, the outliers of the DDPG algorithm were far from the lower bound, further illustrating the relative instability of these two algorithms. In summary, the model trained by the SAC algorithm exhibits high stability and excellent performance.

[0262] In summary, the task offloading and resource allocation method based on federated self-supervised learning provided in this embodiment uses the SAC algorithm to allocate the transmission power of vehicles to base stations and RSUs, the CPU frequency in vehicles, and the calculation resource allocation ratio in RSUs. According to the allocated ratio, the available calculation resources for each vehicle in the RSU are calculated. Then, according to the allowed calculation resources, each local model training iteration task is divided into local training and RSU-assisted training. An offloading ratio threshold is set to optimize the task offloading efficiency. After offloading some iteration tasks to the RSU, the vehicle processes the remaining iteration tasks locally. If there are still unfinished iterations in a round of federated learning, the remaining number of iterations will be stored and waiting for the next training, reducing the system energy consumption and improving the offloading efficiency and accuracy of federated self-supervised learning.

[0263] Embodiment 2.

[0264] Based on the same inventive concept as Embodiment 1, this embodiment provides a task offloading and resource allocation system based on federated self-supervised learning. Since the principle of the system to solve problems is similar to the task offloading and resource allocation method based on federated self-supervised learning provided in the foregoing Embodiment 1, the implementation of the system can refer to the implementation of the task offloading and resource allocation method based on federated self-supervised learning provided in Embodiment 1.

[0265] As Figure 16 shown, the task offloading and resource allocation system based on federated self-supervised learning includes:

[0266] An acquisition module 10, configured to acquire the global model of federated training stored in the base station, as well as the first transmission power, CPU frequency, and task offloading ratio allocated by the base station for the target vehicle.

[0267] A construction module 20, configured to construct the first local model of the target vehicle according to the parameters and model structure of each layer of the global model.

[0268] A first determination module 30, configured to determine the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and the task offloading ratio.

[0269] The second determination module 40 is configured to determine a new local model of the target vehicle based on the first local model and local data, and use it as the second local model.

[0270] The third determination module 50 is configured to unload local data to the RSU at a first transmission power according to the number of task offloading iterations, so as to determine the local model of the RSU and use it as the third local model.

[0271] The fourth determination module 60 is configured to upload the second local model, the third local model, the position information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station, so as to determine the system state of the target vehicle, the actions and rewards in the target time slot.

[0272] The network update module 70 is configured to update the SAC network according to the system state of the target vehicle, the actions and rewards in the target time slot.

[0273] The fifth determination module 80 is configured to determine the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in the federated training.

[0274] The loop module 90 is configured to repeatedly execute the operations of the acquisition module to the fifth determination module until the preset number of training rounds of the federated training is reached, obtain the final SAC network and global model, so as to perform resource allocation according to the final SAC network and perform task offloading according to the final global model.

[0275] Exemplarily, the first determination module includes:

[0276] An acquisition unit configured to acquire the total number of training iterations required by vehicle n at time slot t .

[0277] A first determination unit configured to determine the expected number of training iterations of the local model and the actual number of training iterations of the local model ;

[0278] A second determination unit configured to determine the total number of training iterations N on the RSU side R .

[0279] A first calculation unit configured to calculate the expected number of task offloading training iterations on the RSU side according to the following formula :

[0280] .

[0281] Where is a binary variable indicating whether vehicle n needs to offload some tasks to the RSU; when the task offloading ratio of vehicle n is less than the offloading ratio threshold, ; when the task offloading ratio of vehicle n is not less than the offloading ratio threshold, .

[0282] The third determination unit is configured to use and the minimum value among them as the number of task offloading training iterations .

[0283] The second calculation unit is configured to calculate the number of untrained iterations according to the following formula :

[0284] .

[0285] Exemplarily, the second determination module includes:

[0286] The third calculation unit is configured to calculate the information noise contrastive estimation loss according to the following formula :

[0287] .

[0288] where exp(·) represents the exponential function with the natural constant as the base; is the enhanced local data the anchor sample output after being input into the function χ composed of an encoder, two fully connected neural networks, and a ReLU function; is the enhanced local data the positive sample output after being input into the function χ; is the enhanced local data the negative sample output after being input into the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i≠j; π 1 (·) is the first data augmentation function; π 2 (·) is the second data augmentation function.

[0289] The fourth calculation unit is configured to calculate the loss of dual temperature according to the following formula :

[0290] .

[0291] where sg[·] represents the stop gradient; the hyperparameter representing temperature is τ 2 when the information noise contrastive estimation loss; the hyperparameter representing temperature is τ 1 when the information noise contrastive estimation loss.

[0292] The first construction unit is used to construct a second local model according to the loss of dual temperatures Expression:

[0293] .

[0294] Wherein, represents the first local model; η r is the learning rate; is the gradient descent algorithm; Z is the total number of local data.

[0295] Exemplarily, the fourth determination module includes:

[0296] The third determination unit is used to determine the expected number of training iterations of the local model and the actual number of training iterations of the local model .

[0297] The fifth calculation unit is used to calculate the number of training iterations of the local model according to the following formula :

[0298] .

[0299] The sixth calculation unit is used to calculate the total energy consumption of each vehicle at time slot t according to the following formula:

[0300] .

[0301] Wherein, is the total energy consumption of vehicle n at time slot t; is the computing energy consumption of the RSU; is the local computing energy consumption; is the transmission energy consumption of the RSU; is the energy consumed by vehicle n for one training iteration; is the energy consumed by the RSU for one training iteration.

[0302] The second construction unit is used to construct the system state s t Expression:

[0303] .

[0304] Wherein, is the fading information, which is determined by the current position of the vehicle; ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in federated training.

[0305] The third construction unit is used to construct the action a at time slot t t Expression:

[0306] .

[0307] Wherein, is the power when all vehicles send data to the RSU; ; is the transmission power for vehicle n to transmit data to the RSU; is the CPU frequency allocated by the base station to all vehicles; ; is the CPU frequency allocated by the base station to vehicle n; is the task offloading ratio of all vehicles; ; is the task offloading ratio of vehicle n.

[0308] The fourth construction unit is used to construct the reward r t Expression:

[0309] .

[0310] Wherein, θ 1 is 's weight; θ 2 is the weight of the overload ; θ 3 is 's weight.

[0311] Exemplarily, the fifth determination module includes:

[0312] The fifth construction unit is used to construct the global model of the base station Expression:

[0313] .

[0314] Wherein, represents the second local model; represents the third local model; N is the total number of vehicles participating in federated training.

[0315] For the more specific working processes of the above various modules, reference can be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated here.

[0316] Embodiment 3.

[0317] This embodiment provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program saved in the memory, the steps of the task offloading and resource allocation method based on federated self-supervised learning described in Embodiment 1 are implemented.

[0318] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated herein.

[0319] Embodiment 4.

[0320] This embodiment provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the task offloading and resource allocation method based on federated self-supervised learning described in Embodiment 1 are implemented.

[0321] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated herein.

[0322] Embodiment 5.

[0323] This embodiment provides a computer program product, including computer-executable instructions or a computer program, when the computer-executable instructions or the computer program are executed by a processor, the steps of the task offloading and resource allocation method based on federated self-supervised learning described in Embodiment 1 are implemented.

[0324] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated herein.

[0325] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems, devices, storage media, and computer program products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0326] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0327] In some embodiments, the computer-executable instructions can be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0328] By way of example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts stored in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).

[0329] By way of example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one site, or, on multiple electronic devices distributed at multiple sites and interconnected via a communication network.

[0330] The present invention has been described in detail above in conjunction with specific embodiments and exemplary examples, but these descriptions should not be construed as limiting the present invention. Those skilled in the art understand that without departing from the spirit and scope of the present invention, various equivalent substitutions, modifications, or improvements can be made to the technical solutions of the present invention and their implementation manners, and these all fall within the scope of the present invention. The protection scope of the present invention is subject to the appended claims.

Claims

1. A task offloading and resource allocation method based on federated self-supervised learning, characterized in that: include: S1, obtaining a global model of federated training stored in a base station and a first transmission power, CPU frequency, and task offloading ratio allocated by the base station to a target vehicle; S2, build the first local model of the target vehicle according to each layer parameter and model structure of the global model; wherein vehicle n stores Z images captured in the previous round, called local data, and has a model with the same structure as the global model on the base station side, called the local model; S3, determining the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and the task offloading ratio; S4, determining a new local model of the target vehicle according to the first local model and the local data to serve as a second local model; S5, unloading the local data to the RSU at the first transmission power according to the number of task offloading iterations, so as to determine a local model of the RSU and use it as a third local model; S6, uploading the second local model, the third local model, the location information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the system state of the target vehicle, the action and the reward of the target time slot; S7, updating the SAC network according to the system state of the target vehicle, the action and reward of the target time slot; S8, determining a global model of the base station according to the second local models and the third local models corresponding to all vehicles participating in the federated training; S9, repeating the operations of S1-S8 until the preset training rounds of federated training are reached, and the final SAC network and the global model are obtained, so as to allocate resources according to the final SAC network and offload tasks according to the final global model.

2. The task offloading and resource allocation method according to claim 1, characterized in that: The determining the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and the task offloading ratio includes: Get the total number of training iterations required for vehicle n in time slot t ; Determine the expected number of training iterations for the local model And the actual number of training iterations of the local model ; Determine the total number of training iterations N on the RSU side R ; The expected number of task offloading training iterations on the RSU side is calculated according to the following formula : ; in, It is a binary variable, indicating whether vehicle n needs to offload part of its tasks to the RSU; when the task offloading ratio of vehicle n When it is less than the uninstall ratio threshold, ; When the task unloading ratio of vehicle n When the unloading ratio is not less than the threshold, ; Will and The minimum value among them is taken as the number of task offloading training iterations ; The number of untrained iterations is calculated according to the following formula : 。 3. The task offloading and resource allocation method according to claim 1, characterized in that: The step of determining a new local model of the target vehicle according to the first local model and the local data as a second local model includes: The information-noise contrast estimation loss is calculated according to the following formula : ; Among them, exp(·) represents an exponential function with a natural constant as the base; For enhanced local data The anchor point sample is input to the function χ which consists of an encoder, two fully connected neural networks and a ReLU function; For enhanced local data The positive sample output after input into the function χ; For enhanced local data The negative sample output after being input to the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i≠j; π1(·) is the first data enhancement function; π2(·) is the second data enhancement function; The loss of dual temperature is calculated according to the following formula : ; Among them, sg[·] represents the stop gradient; The information-noise contrast estimation loss when the temperature hyperparameter is τ2; The information-noise contrast estimation loss when the temperature hyperparameter is τ1; Constructing a second local model based on the loss of dual temperature expression: ; in, represents the first local model; η r is the learning rate; is the gradient descent algorithm; Z is the total number of local data.

4. The task offloading and resource allocation method according to claim 2, characterized in that: The second local model, the third local model, the location information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle and the number of untrained iterations are uploaded to the base station to determine the global model of the base station, the system state of the target vehicle, the action and reward of the target time slot, including: The number of iterations of local model training is calculated according to the following formula : ; The total energy consumption of each vehicle at time slot t is calculated according to the following formula: ; in, is the total energy consumption of vehicle n in time slot t; is the computing energy consumption of RSU; Calculate energy consumption for local areas; is the transmission energy consumption of RSU; The energy consumed to perform one training iteration for vehicle n; The energy consumed by performing one training iteration for the RSU; Build system status t expression: ; in, is the fading information, which is determined by the current position of the vehicle; ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in the federated training; Construct action a in time slot t t expression: ; in, The power of all vehicles when sending data to RSU; ; is the transmission power of vehicle n transmitting data to RSU; The CPU frequency allocated to all vehicles by the base station; ; The CPU frequency assigned to vehicle n by the base station; mission unloading ratio for all vehicles; ; is the task unloading ratio of vehicle n; Build Rewards t expression: ; Among them, θ1 is The weight of θ2 is the overload The weight of The weight of .

5. The task offloading and resource allocation method according to claim 1, characterized in that: The determining of the global model of the base station according to the second local model and the third local model corresponding to all vehicles participating in the federated training includes: Build a global model of base stations expression: ; in, represents the second local model; represents the third local model; N is the total number of vehicles participating in the federated training.

6. A task offloading and resource allocation system based on federated self-supervised learning, characterized in that: include: An acquisition module, used to acquire a global model of federated training stored in a base station and a first transmission power, a CPU frequency, and a task offloading ratio allocated by the base station to a target vehicle; A construction module is used to construct a first local model of the target vehicle according to each layer parameter and model structure of the global model; wherein the vehicle n stores the Z images captured in the previous round, called local data, and has a model with the same structure as the global model on the base station side, called the local model; A first determination module is used to determine the number of task offloading iterations and the number of untrained iterations according to the CPU frequency and the task offloading ratio; A second determination module, used for determining a new local model of the target vehicle according to the first local model and the local data, to serve as a second local model; A third determination module is used to offload the local data to the RSU at a first transmission power according to the number of task offloading iterations, so as to determine a local model of the RSU as a third local model; a fourth determination module, configured to upload the second local model, the third local model, the location information of the target vehicle, the speed of the target vehicle in the target time slot, the total energy consumption of the target vehicle, and the number of untrained iterations to the base station to determine the system state of the target vehicle, the action and the reward in the target time slot; A network update module, used to update the SAC network according to the system status of the target vehicle, the action and reward of the target time slot; A fifth determination module, configured to determine a global model of the base station according to the second local models and the third local models corresponding to all vehicles participating in the federated training; A loop module is used to repeatedly execute the operations of the acquisition module to the fifth determination module until a preset training round of federated training is reached, so as to obtain a final SAC network and a global model, so as to perform resource allocation according to the final SAC network and perform task unloading according to the final global model.

7. The task offloading and resource allocation system according to claim 6, characterized in that: The first determining module comprises: The acquisition unit is used to obtain the total number of training iterations required for vehicle n in time slot t ; The first determination unit is used to determine the estimated number of training iterations of the local model And the actual number of training iterations of the local model ; The second determination unit is used to determine the total number of training iterations N on the RSU side. R ; The first calculation unit is used to calculate the expected number of task offloading training iterations on the RSU side according to the following formula : ; in, It is a binary variable, indicating whether vehicle n needs to offload part of its tasks to the RSU; when the task offloading ratio of vehicle n When it is less than the uninstall ratio threshold, ; When the task unloading ratio of vehicle n When the unloading ratio is not less than the threshold, ; The third determining unit is used to and The minimum value among them is taken as the number of task offloading training iterations ; The second calculation unit is used to calculate the number of untrained iterations according to the following formula : 。 8. The task offloading and resource allocation system according to claim 6, characterized in that: The second determining module comprises: The third calculation unit is used to calculate the information noise contrast estimation loss according to the following formula : ; Among them, exp(·) represents an exponential function with a natural constant as the base; For enhanced local data The anchor point sample is input to the function χ which consists of an encoder, two fully connected neural networks and a ReLU function; For enhanced local data The positive sample output after input into the function χ; For enhanced local data The negative sample output after being input to the function χ; τ is the hyperparameter of temperature; K is the total number of local data marked as negative samples; i≠j; π1(·) is the first data enhancement function; π2(·) is the second data enhancement function; The fourth calculation unit is used to calculate the loss of dual temperature according to the following formula : ; Among them, sg[·] represents the stop gradient; The information-noise contrast estimation loss when the temperature hyperparameter is τ2; The information-noise contrast estimation loss when the temperature hyperparameter is τ1; The first building unit is used to build a second local model according to the loss of the dual temperature expression: ; in, represents the first local model; η r is the learning rate; is the gradient descent algorithm; Z is the total number of local data.

9. The task offloading and resource allocation system according to claim 7, characterized in that: The fourth determination module comprises: The fifth calculation unit is used to calculate the number of iterations of local model training according to the following formula : ; The sixth calculation unit is used to calculate the total energy consumption of each vehicle at time slot t according to the following formula: ; in, is the total energy consumption of vehicle n in time slot t; is the computing energy consumption of RSU; Calculate energy consumption for local areas; is the transmission energy consumption of RSU; The energy consumed to perform one training iteration for vehicle n; The energy consumed by performing one training iteration for the RSU; The second building unit is used to build the system state s t expression: ; in, is the fading information, which is determined by the current position of the vehicle; ; is the fading information of vehicle n at time slot t; is the vehicle speed; ; is the speed of vehicle n at time slot t; N is the total number of vehicles participating in the federated training; The third construction unit is used to construct an action a in time slot t t expression: ; in, The power of all vehicles when sending data to RSU; ; is the transmission power of vehicle n transmitting data to RSU; The CPU frequency allocated to all vehicles by the base station; ; The CPU frequency assigned to vehicle n by the base station; mission unloading ratio for all vehicles; ; is the task unloading ratio of vehicle n; The fourth building block is used to build reward r t expression: ; Among them, θ1 is The weight of θ2 is the overload The weight of The weight of .

10. The task offloading and resource allocation system according to claim 6, characterized in that: The fifth determining module comprises: The fifth construction unit is used to construct a global model of the base station expression: ; in, represents the second local model; represents the third local model; N is the total number of vehicles participating in the federated training.

Citation Information

Patent Citations

  • Federal learning optimization method based on block chain

    CN117610644A

  • Agent policy learning method with privacy protection in mobile edge computing

    WO2024254892A1