Vehicle edge computing network client scheduling method and system based on reputation model
By adopting dynamic interactive reputation model and asynchronous parallel DDPG algorithm in vehicle edge computing networks, the problem that traditional reputation models are difficult to identify malicious nodes and reinforcement learning methods are poorly performed in large-scale problems is solved, and efficient, secure and efficient resource scheduling and federated learning are achieved.
Patent Information
- Application Number
- CN202410874217.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-02
AI Technical Summary
In vehicle edge computing networks, traditional reputation models are difficult to identify and filter out malicious nodes that perform dirty label attacks. At the same time, reinforcement learning methods such as DQN, A3C and DDPG show limited exploration capabilities and insufficient adaptability in large-scale problems, resulting in inefficient resource scheduling.
A dynamic interactive reputation model is used to assist the drone in selecting vehicle clients, improve the server's ability to resist malicious tampering with data tags, and uses asynchronous and parallel DDPG algorithm to centrally schedule the client's computing resources and communication capabilities, and balance the delay and energy consumption caused by local training.
It improves the reliability and security of the system, improves task execution efficiency and resource utilization, enhances the adaptability and attack resistance of the system, and achieves efficient federated learning.
Smart Images

Figure CN118612743B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing and federated learning technology, and in particular to a vehicle edge computing network client scheduling method and system based on a reputation model. Background Art
[0002] With the rapid development of urban informatization, a large number of intelligent terminal devices, including vehicles, smart phones, monitoring systems, etc., have been popularized in cities. These devices generate a large amount of high-quality data, which needs to be integrated, calculated and analyzed to be fully utilized. However, traditional centralized processing methods encounter various limitations, including wireless network bandwidth, cloud platform computing power, and delays caused by medium and long-distance transmission. Therefore, they are gradually being replaced by mobile edge computing (MEC). As a cutting-edge technology, MEC is gradually penetrating into various fields, among which vehicular edge computing networks (VECN) have become a highly concerned application scenario. At the same time, when multiple edge devices collaborate to learn and share models, the use of federated learning effectively solves problems such as data privacy and data islands. Federated learning effectively meets the information exchange needs in complex traffic environments and provides efficient and secure edge computing for the Internet of Vehicles. The application of federated learning in VECN is becoming more and more extensive, providing a large number of effective solutions for urban information security and data processing.
[0003] The main challenges in achieving efficient federated learning in VECN are as follows. First, the complexity of the urban traffic environment leads to uneven quality of data obtained by clients. Therefore, it is necessary to screen trusted clients participating in federated learning to prevent attacks from malicious nodes and improve the quality of data sources. Secondly, MEC places more complex computing requirements on mobile devices with limited resources. Therefore, it is necessary to reasonably arrange the computing resources required for training by local clients to achieve a balance between cost and benefit. To this end, this patent designs a joint scheduling method for client selection and resource allocation in the edge computing network scenario of drone-assisted vehicles.
[0004] Reputation model is one of the commonly used methods for data quality assessment and customer selection. Reputation model is a comprehensive assessment of the reliability of participants based on factors such as historical behavior and performance indicators. In the field of MEC, reputation model plays an important role in data quality assessment, node reliability assessment and malicious node identification. However, in the process of selecting high-quality data sources, traditional reputation model is difficult to screen out malicious nodes that carry out dirty label attacks.
[0005] In addition, to ensure the latency and energy consumption of the system, efficient client local resources are crucial. With the integration of edge computing in the Industrial Internet of Things, solving the latency and energy consumption problems brought by mobile computing has become an important concern. However, the reinforcement learning methods currently used in resource scheduling, such as Deep Q-Network (DQN), Asynchronous Advantage Actor-Critic (A3C), and Deep Deterministic Policy Gradient (DDPG), show limited exploration capabilities and insufficient adaptability when facing large-scale problems, making it challenging to manage multiple action space dimensions simultaneously. Summary of the invention
[0006] In view of the defects of the prior art, the present invention provides a method and system for scheduling vehicle edge computing network clients based on a reputation model. The dynamic interactive reputation model is used to assist drones in selecting vehicle clients, improve the server's ability to resist malicious tampering of data labels by users, and use the asynchronous parallel DDPG algorithm to centrally schedule the client's computing resources and communication capabilities, balance the delay and energy consumption caused by local training, and ultimately achieve efficient federated learning.
[0007] In order to achieve the above invention object, the technical solution adopted by the present invention is as follows:
[0008] A vehicle edge computing network client scheduling method based on a reputation model comprises the following steps:
[0009] Step 1: Establish an edge computing scenario for drone-assisted vehicles. This scenario covers scenario construction, communication model, computing model, federation model, and optimization goals, including:
[0010] S11. Build a scenario for drone-assisted vehicles in an edge computing network, including determining the network topology and participant role functions;
[0011] S12. Establish a communication model between the vehicle and the UAV, including defining the communication protocol and data transmission mechanism;
[0012] S13, determining the computing capability and resource configuration of the vehicle, including local computing parameters and transmission delay;
[0013] S14. Establish a unified dispatch model for UAVs to vehicles, including global model transmission and local model training management;
[0014] S15. Set optimization goals for the edge computing system, including minimizing resource consumption and improving model accuracy.
[0015] Step 2: Establish a dynamic interactive reputation model by constructing a direct reputation model and an indirect reputation model, including:
[0016] S21. Build a direct reputation model to evaluate the reliability of the client;
[0017] S22. Construct an indirect reputation model that comprehensively considers the historical records of other drones and collaborative evaluation of the vehicle’s reputation.
[0018] Step 3: Propose a complete user recruitment method based on the reputation model and select clients in this way, including:
[0019] S31. Collect historical performance data and credit rating information of the vehicle;
[0020] S32. Integrate direct and indirect reputation models to conduct comprehensive evaluation and reputation scoring of vehicles;
[0021] S33, select appropriate clients to participate in the task according to the reputation score and system requirements;
[0022] S34. Update the vehicle's credit rating based on the task execution results, and provide feedback and record.
[0023] Step 4: Combine the asynchronous superior action criticism algorithm and the deep deterministic policy gradient algorithm to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and use this algorithm for resource scheduling.
[0024] Furthermore, the specific process of step one is as follows:
[0025] S11. Build a scenario of drone-assisted vehicles, including an edge computing scenario consisting of drones and vehicles. The vehicle is responsible for data collection and local model training, and the drone is responsible for scheduling and maintaining the update of the global model. As an edge node, the vehicle has an independent data set and computing storage capabilities, and the drone manages vehicle scheduling as a central server. The system runs in discrete time, and the task generation interval follows a Poisson distribution.
[0026] S12. Establish a communication model for drone-assisted vehicles, describing the uplink channel as a flat Rayleigh fading channel, with each drone connecting and dispatching multiple vehicles. The drone moves to a favorable position according to the road traffic distribution and issues periodic round instructions to the vehicle, including client selection, global model distribution, and resource scheduling. Calculate the channel gain and transmission rate during the vehicle-drone communication process, taking into account obstacles between communication links.
[0027] S13. Establish a computing model for UAV-assisted vehicles, generate tasks for the vehicles, define task models and local computing parameters, calculate the time cost and resource consumption of local model training, and consider the transmission process delay and energy consumption.
[0028] S14. Establish a federated model for UAV-assisted vehicles. The UAV coordinates the scheduling and training of the global federated model, manages the local federated model trained by each device, and defines the model loss of the local model and the global model. The goal is to minimize the loss function of the global model parameters.
[0029] S15. Set the optimization goal of the drone-assisted vehicle, jointly optimize the client selection and resource scheduling, select vehicles with rich computing resources and sufficient data, screen out malicious vehicle attacks, ensure the minimum model loss, and use reasonable resource scheduling methods to reduce the loss of training accuracy and time and resource cost consumption. Model the optimization goals of the edge computing network.
[0030] Furthermore, the specific process of step 2 is as follows:
[0031] S21. Build a direct reputation model to evaluate customer reputation based on dynamic factors such as vehicle performance, capabilities, and costs, taking into account federated model similarity, communication link smoothness, number of data sets, and estimated calculation time, and establish a direct reputation model that takes into account time correlation.
[0032] S22. Construct an indirect reputation model, utilize the historical reputation records of various drones through model similarity evaluation, comprehensively consider the reputation records of other drones, evaluate vehicle reputation through collaborative cooperation, and construct a comprehensive indirect reputation model as the core of the user recruitment method to provide a basis for optimizing client selection.
[0033] Furthermore, the specific process of step three is as follows:
[0034] S31,Data Collection,The UAV collects vehicle client data from historical mission experiences,,including local historical model accuracy, dataset quality, and computing power.
[0035] S32. Reputation modeling: organize all user data within the jurisdiction, establish a direct reputation model based on the dynamic factors of performance, capability and cost, establish an indirect reputation model based on model similarity and historical reputation records, and build a complete reputation model to provide a basis for client selection.
[0036] S33, user selection, using the reputation model to select a group of vehicles to perform the task, vehicles with high reputation values have higher reliability and data quality.
[0037] S34, model training, the UAV transmits the global federated model to the selected vehicle, and the vehicle uses its own resource training data to train the local federated model based on the global model.
[0038] S35, reputation feedback, the drone collects the vehicle's local training model, summarizes and updates the model, records the model accuracy, data set quality and computing power, calculates and updates the client's reputation score based on the vehicle's performance, and realizes the user recruitment cycle.
[0039] Furthermore, the specific process of step 4 is as follows:
[0040] S41: Algorithm selection, combining the asynchronous advantage action criticism algorithm and the deep deterministic policy gradient algorithm to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm.
[0041] S42: Algorithm optimization, optimization of the asynchronous parallel DDPG algorithm, including adjusting the neural network structure, optimizing hyperparameter settings, improving training strategies, etc., to improve the performance and efficiency of the algorithm.
[0042] S43: Resource scheduling, using the optimized asynchronous parallel DDPG algorithm for client scheduling and resource allocation, ensuring the effective selection of vehicles for task processing in the vehicle edge computing network, and dynamically adjusting the resource allocation strategy according to the real-time situation.
[0043] The present invention also discloses a client scheduling system based on a reputation model in a vehicle edge computing network. The system can be used to implement the above-mentioned vehicle edge computing network client scheduling method, specifically including:
[0044] Scenario establishment module: responsible for establishing the computing scenario of drone-assisted vehicles in the vehicle edge computing network, including determining the network topology and participant role functions.
[0045] Communication model establishment module: establish the communication model between the vehicle and the drone, define the communication protocol and data transmission mechanism, and ensure effective information transmission and interaction.
[0046] Computing model building module: Determines the vehicle's computing power and resource configuration, including local computing parameters and transmission delays, providing a basis for task allocation and execution.
[0047] Optimization target setting module: sets the optimization targets of the edge computing system, including minimizing resource consumption and improving model accuracy, and guides the overall optimization direction of the system.
[0048] Reputation model building module: build direct reputation model and indirect reputation model, evaluate the reliability of the client, and provide a reputation evaluation basis for user recruitment and task allocation.
[0049] User recruitment method module: proposes a complete user recruitment method, including collecting historical performance data and reputation evaluation information of vehicles, comprehensively evaluating the reputation of vehicles and selecting appropriate clients to participate in the task.
[0050] Algorithm optimization module: The asynchronous superior action criticism algorithm and the deep deterministic policy gradient algorithm are integrated to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and this algorithm is used for resource scheduling to improve system efficiency and performance.
[0051] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned vehicle edge computing network client scheduling method when executing the program.
[0052] The present invention also discloses a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the above-mentioned vehicle edge computing network client scheduling method is implemented.
[0053] Compared with the prior art, the advantages of the present invention are:
[0054] 1. Improve system reliability: The vehicle edge computing network client scheduling method based on the reputation model can evaluate the reliability and reputation of the vehicle, select vehicles with good reputation to participate in the task, thereby improving the reliability and stability of the system.
[0055] 2. Improve the efficiency of task execution: By selecting vehicles with higher credibility to participate in tasks, the efficiency and accuracy of task execution can be improved, and the time and resource consumption of task execution can be reduced.
[0056] 3. Enhanced security: The reputation model can identify and screen out malicious or low-reputation vehicles, reducing the risk of the system being attacked by malicious forces, thereby enhancing the security of the system.
[0057] 4. Reasonable allocation of resources: By comprehensively considering the vehicle's credit score and system requirements, computing resources can be allocated more reasonably, resource utilization can be improved, and resource waste can be reduced.
[0058] 5. Strong adaptability: The reputation model can be dynamically adjusted based on the vehicle's historical performance data and real-time performance to adapt to changing needs in different scenarios, enhancing the adaptability and flexibility of the system.
[0059] 6. Reduce management costs: Through automated reputation evaluation and client selection processes, the manpower and time costs of system management and maintenance can be reduced, improving system management efficiency.
[0060] 7. Joint optimization of client selection and resource scheduling: This invention realizes the joint optimization of client selection and resource scheduling in the UAV-assisted vehicle edge computing network. By comprehensively considering the reputation score of the vehicle and the system resource situation, the intelligent allocation of computing resources is realized, thereby maximizing the overall performance of the system.
[0061] 8. Effectively resist malicious node attacks: The introduction of the reputation model can effectively identify and filter attacks from malicious nodes and protect the system from malicious behavior. By excluding the participation of low-reputation vehicles, the system's security and anti-attack capabilities are improved.
[0062] 9. More efficient use of resources: By rationally selecting vehicles with higher credibility to participate in tasks and dynamically adjusting resource allocation according to task requirements, more efficient use of resources is achieved. This can reduce idle waste of resources and improve the efficiency and economy of the system.
[0063] 10. Achieve efficient federated learning: By optimizing client selection and resource scheduling through the reputation model, a more efficient environment can be provided for federated learning. Vehicles with higher reputation participate in model training, which improves the quality and efficiency of the training model, thereby achieving a faster federated learning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a flow chart of the vehicle edge computing network client scheduling method proposed in an embodiment of the present invention;
[0065] Figure 2 is an edge computing scene diagram of a drone-assisted vehicle proposed in an embodiment of the present invention;
[0066] Figure 3 4 is a structural diagram of the asynchronous parallel DDPG algorithm proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples.
[0068] like Figure 1 As shown, the present invention provides a vehicle edge computing network client scheduling method based on a reputation model, comprising the following steps:
[0069] Step 1: Establish an edge computing scenario for drone-assisted vehicles. This scenario covers scenario construction, communication model, computing model, federation model, and optimization goals.
[0070] 1.1) Build a UAV-assisted vehicle scenario: This edge computing scenario consists of a UAV and a vehicle. The vehicle is responsible for collecting data and training models locally, while the UAV is responsible for managing vehicle scheduling and maintaining global model updates, such as Figure 2. Vehicles, as edge nodes of federated learning, have independent data sets and computing storage capabilities. Drones, as central servers, are responsible for the transmission, reception, and update of global models, and for unified scheduling of vehicles in the coverage area. Vehicles travel along one-way roads, and drones patrol the area at a fixed altitude. Drones and vehicles establish a connection through a wireless network, but this connection may be interrupted by obstacles such as trees. The entire system runs in discrete time, and the time intervals between user-generated tasks follow a Poisson distribution.
[0071] In a mission cycle, the drone selects vehicles based on the reputation model. The selected clients allocate specific computing resources for local model training according to the drone's scheduling strategy. These clients then transmit the updated model parameters to the drone for aggregation, and the drone then evaluates the clients participating in the training and assigns rewards or penalties. Vehicles as edge nodes can be trained collaboratively without sharing raw data to obtain a comprehensive global model.
[0072] 1.2) Establish the communication model of drone-assisted vehicles: The uplink channel between the vehicle and the drone is a flat Rayleigh fading channel. Each drone will connect, manage and dispatch multiple vehicles at the same time. At the beginning of a mission cycle, the drone will move to a more favorable position according to the road traffic distribution, and then issue instructions for this cycle to the vehicle, including client selection, global model distribution, and resource scheduling. In the three-dimensional Cartesian coordinate system, the vertical height h of any drone u∈{1,2,...,v} remains unchanged, and its plane coordinates are Any vehicle i∈ within the coverage area of drone u
[0073] The running speed of {1,2,...,j} remains unchanged, and the plane coordinates are expressed as For UAV u, it needs to establish a communication relationship with each vehicle i within the coverage area, and the channel gain of the line-of-sight link between the two can be expressed as:
[0074]
[0075] Where α0 is the channel gain when the reference distance is 1m, is the Euclidean distance between UAV u and vehicle i within the coverage area. Assuming the upload link between the vehicle and the UAV is a flat Rayleigh fading channel, according to the Shannon formula, the upload link transmission rate of vehicle i in this time slot can be calculated as:
[0076]
[0077] Where W is the communication bandwidth between the vehicle and the UAV, is the transmission power from vehicle i to UAV u in the upload link, σ2 is the noise power. Considering the existence of obstacles between communication links, the wireless transmission rate can be expressed as:
[0078]
[0079] In the formula, q(i) represents the identifier of whether there is an obstacle between the drone and the vehicle, P NLOS Defined as the average transmission loss.
[0080] 1.3) Establishing the computing model of drone-assisted vehicles: In the vehicle edge computing network, tasks are generated by vehicles. Vehicle i will generate independent tasks Task i , the task will contain several important attributes Where id i Indicates the id of the task. Indicates the data size of the task (in bits), Indicates the number of CPU cycles required to calculate one bit of data. Indicates the total number of CPU cycles for the task (in CPU cycles), and has At the same time, the mobile terminal has the ability to store and calculate, and the maximum CPU frequency that vehicle i can provide is f i , after the computing capacity scheduling coefficient of the UAV χ i The frequency after is df i = x i f i , then the time cost of local model training is:
[0081]
[0082] Usually, the data size of the calculation results provided by the drone to the vehicle is very small, so the delay of the downlink transmission link is ignored. The time cost of each task round includes two parts: transmission delay and calculation delay, where the transmission delay can be expressed as:
[0083]
[0084] In addition, vehicle i performs task id i The power and energy consumed can be expressed as:
[0085]
[0086] Where κ is the effective switch capacitance that depends on the chip architecture. For vehicle i, the energy consumed during the transmission process is:
[0087]
[0088] In the formula, represents the transmission power of the model uploaded by vehicle i dp i is the upload power scheduling coefficient of the UAV to vehicle i.
[0089] At time t, drone u passes the reputation model The selected client set is I u ∈{1,2,...,j}, then in this cycle, the total time cost is the maximum value of the sum of the selected client transmission delay and the calculation delay, and the total energy consumption is the sum of the client resource consumption participating in the calculation, which are expressed as:
[0090]
[0091] 1.4) Establishing a federated model for drone-assisted vehicles: In the vehicle edge computing network, the overall goal of drone coordination is to train a stable global federated model. The model trained for each device under its management scope is called a local federated model. The global model parameters of drone u are expressed as ω u The local model of vehicle i is represented by ω i The training set of vehicle i is represented by D i Indicates that for each input sample sa, the loss function ls(ω,in sa ,ou sa ) to measure the model ω in the input sample vector in sa The expected output sample vector ou sa The performance error between them. Then vehicle i in the training set D i The loss function on is expressed as:
[0092]
[0093] Where |D i | is the dataset D i Since the data set of the global model on the drone is distributed, the loss function of drone u is defined as:
[0094]
[0095] The goal of federated learning is to minimize the loss function of the global model parameters on the UAV.
[0096] 1.5) Setting optimization goals for drone-assisted vehicles: In the federated learning system for drone-assisted vehicles, the ultimate goal is to jointly optimize client selection and resource scheduling. The drone must select vehicles with rich computing resources and sufficient data while screening out attacks from malicious vehicles to ensure minimal model loss. At the same time, a reasonable resource scheduling method must be used to reduce the loss of training accuracy and the time and resource costs consumed. Specifically, the optimization problem of VECN can be described as:
[0097]
[0098] In the formula, λ LS , t , E They represent the weights of the loss function, time cost and resource cost of the global model respectively. Constraint (a) specifies the selected client set I u Reputation Model Constraints (b) and (c) define the scheduling range of transmission power and computing power. Constraint (d) defines the barriers of the communication link. Constraint (e) determines that each round of training must be completed within the deadline τ. max To solve this NP-hard problem, we use a reputation model to select client nodes and use the APDDPG algorithm to optimize transmission power and processing frequency. For this NP-hard problem, we will select client nodes by building a reputation model and schedule transmission power and operation frequency through the asynchronous parallel DDPG algorithm.
[0099] Step 2: Establish a dynamic interactive reputation model by constructing direct reputation model and indirect reputation model.
[0100] 2.1) Constructing a direct reputation model: Direct reputation is a method where the drone evaluates the customer’s reputation based on dynamic factors such as the vehicle’s performance, capabilities, and cost. During the mission cycle τ, the vehicle i selected by the drone u will perform model training locally and set the model parameters Update to drone aggregation Before the model is aggregated, the drone will perform a quality assessment on the model parameters submitted by the vehicle and give the vehicle a performance score for this round of training. in, is the Euclidean distance between the local training model parameters and the global model parameters of the previous round, defined as:
[0101]
[0102] In the formula, the first round of training used The model is the average aggregation model. f(x) represents the normalization function, which normalizes the Euclidean distance:
[0103]
[0104] Where δ is the Euclidean distance normalization coefficient, and ψ is the normalization function coefficient. The original Euclidean distance of the model parameters is normalized and mapped through the Euclidean distance normalization function to obtain the performance evaluation of the drone on the vehicles participating in this round of training. In a complex and ever-changing urban environment, vehicles will encounter many objective factors that affect the assessment of user credibility during movement. Therefore, the assessment of user credibility should be comprehensively assessed based on current performance and objective environment. Therefore, direct credibility also considers three factors: the smoothness of the communication link, the number of data sets collected by the vehicle, and the estimated calculation time of the vehicle. The smoothness of the communication link will affect the success probability of model transmission, and the reward and punishment constraints of the model parameters are defined as:
[0105]
[0106] In the formula, is the task transmission success rate, and α is the transmission failure normalization coefficient. According to the number of data sets collected by the user, the data set constraint can be expressed as:
[0107]
[0108] In the formula, The data size of the task (in bits). Based on the maximum computing resources that the client can provide, the constraint for estimating the user's computing time cost can be expressed as:
[0109]
[0110] In the formula, s n The number of CPU cycles required to calculate one bit of data, f n is the maximum calculation frequency that can be provided. Combined with the dynamic factors considered above, the direct credibility value of UAV u to vehicle i during the mission cycle τ can be expressed as:
[0111]
[0112] In addition, the reputation model also considers the historical performance of the vehicle and evaluates the time relevance of historical reputation data. Over time, the vehicle will accumulate a large amount of historical performance records. When recruiting users to perform tasks, the vehicle reputation score needs to integrate past performance and current credibility. Since the environment in which the vehicle is located changes randomly, over time, historical data will not reflect the user's current real situation well, and its utilization value will gradually decrease. For this reason, experience is given a greater weight, and the direct reputation score of the vehicle at time τ+1 is derived as follows:
[0113]
[0114] Where γ is the historical decay of direct credit rating.
[0115] 2.2) Constructing an indirect reputation model: The indirect reputation model uses the historical reputation records of various drones through model similarity evaluation. When evaluating the reputation of a customer, the drone will not only subjectively evaluate its reputation performance, but also consider the reputation records of other drones. When drone u is performing reputation evaluation, it will take into account the reputation records of other drones, taking drone v as an example. At the mission time τ, the historical record vehicle set of drone u is I, and the vehicle set of drone v is K. Then, drone u has a certain number of vehicles. When making an assessment, you can refer to the reputation record of drone v as the basis for indirect evaluation. The combination of direct and indirect reputation can make fuller use of reputation information in different dimensions, which helps to more accurately evaluate the historical performance and credibility of the vehicle. First, define the similarity of the reputation model between drones:
[0116]
[0117] In the formula, H u,j represents the complete reputation vector of drone u to vehicle j, represents the average value of the reputation vector of drone u to all vehicles. The indirect reputation value of the user at time τ+1 is derived as:
[0118]
[0119] In the formula, indirect reputation also considers the problem of time decay, and the historical decay of indirect reputation evaluation is defined as η. In summary, the reputation value of drone u to vehicle i can be summarized as:
[0120]
[0121] At this point, the drone collects information about all users in the jurisdiction and builds direct and indirect reputation models based on the user's performance and ability. The reputation model is the core of the user recruitment method and provides a basis for optimizing the selection of clients.
[0122] Step 3: Propose a complete user recruitment method based on the reputation model and use this method to select clients.
[0123] 3.1) Data Collection: The drone collects vehicle client data from historical mission experiences, including the accuracy of the local historical model, the quality of the dataset, and the computing power of the local client.
[0124] 3.2) Reputation Modeling: The drone collates the data of all users in the jurisdiction, establishes a direct reputation model based on dynamic factors such as user performance, capabilities and costs, and establishes an indirect reputation model based on the similarity of reputation models and historical reputation records, and constructs a complete reputation model based on this, which is used as the basis for optimizing the client selection method.
[0125] 3.3) User Selection: At the beginning of each mission cycle, the drone uses the reputation model to select a group of vehicles to perform the current mission. Vehicles with high reputation values have higher reliability and data quality and will have a greater probability of being selected.
[0126] 3.4) Model training: The drone broadcasts the global federated model of this round to the selected vehicles. The vehicles use their own computing and storage resources to train the collected data and obtain the local federated model of this round based on the global model.
[0127] 3.5) Reputation feedback: After a round of model training is completed, the drone collects the client model trained locally by the vehicle and updates the model. At the same time, the drone will also record some information, including the accuracy of the model, the quality of the data set, and the computing power of the local client. The drone will calculate and update the client's reputation score based on the vehicle's performance in this round. Finally, the user recruitment process has achieved a complete cycle.
[0128] Step 4: Combine the A3C algorithm and the DDPG algorithm to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and use this algorithm for resource scheduling.
[0129] 4.1) Algorithm environment: Deep neural networks have powerful representation learning capabilities and can provide complex decision-making and control assistance for reinforcement learning. However, the simple combination of online reinforcement learning algorithms and deep neural networks may cause non-smoothness and strong correlation in data sequences. Therefore, this patent considers integrating experience replay into the asynchronous reinforcement learning framework, combining the asynchronous framework of the A3C algorithm with the deterministic strategy of DDPG to increase the training speed of the algorithm while performing complex high-dimensional continuous decision control.
[0130] Firstly, the joint scheduling problem of computing resources and communication capabilities is formulated as a Markov decision process, namely:<S,A,R,P> . Among them, S, A, R and P represent the state space, action space, reward function and state transition function respectively. In the edge computing network scenario of drone-assisted vehicles, the drone acts as an agent, observes the environment and interacts with the environment, and strives to maximize the reward function. For the Markov decision process, asynchronous parallel DDPG is used to solve it. The state space, action space, reward function and state transition function of the Markov decision process are described as follows:
[0131] The first is the state space. During the mission period t, the state space of UAV u is It consists of the following parts: The transmission power from the vehicle to the drone u The processing frequency of the edge device CPU within the coverage area of drone u is i} i∈u , the horizontal position m of the UAV u t (u) = [x t (u),y t (u)] T ; Obstacle sign {q(i)} between drone u and selected vehicle i i∈u and the vehicle's location The state of UAV u at mission cycle t can be expressed as:
[0132]
[0133] The second is the action space. In the task cycle t, the agent u has a certain state. Choose Action action It consists of the following two parts: computing capacity scheduling coefficient {χ i} i∈u and the transmission power scheduling coefficient {ξ i} i∈u The action of UAV u in mission cycle t can be expressed as:
[0134]
[0135] The third is the reward function: for agent u, the reward is the training accuracy loss LS(ω u ), time consumption and resource costs The sum of . It can be expressed as:
[0136]
[0137] The fourth is the state transition function, which represents the action Conditions, status Transition to state probability.
[0138] 4.2) Algorithm Framework: In the asynchronous parallel DDPG, there are multiple intelligent agents, each of which is trained in a parallel environment based on different experiences. Their common goal is to perfect an accurate global model, such as Figure 3. Each intelligent agent consists of three key parts: Actor, Critic, and Learner. In addition, the assistance of Updater and Puller is required in the interaction between the intelligent agent model and the global model. In a parallel environment, each intelligent agent interacts with the environment independently based on the actions taken by its agent. At the same time, the action outputs of the Actors of all intelligent agents are collectively stored in a shared experience pool. Critic evaluates the value based on the current state space in the parallel environment and the corresponding action outputs of its related actors. Learner combines environmental rewards and historical experience to calculate gradients for updating the Critic network. These gradients are then used to update the Critic network and the Actor network. The gradient update strategy generated by Learner is asynchronously applied by Updater to the update of the global model, while Puller is responsible for distributing the latest global model to each intelligent agent.
[0139] Due to the change in the calculation time of the intelligent agent, the Updater and Puller operate on the global model asynchronously. The parallelization of Actor and Critic improves the efficiency of experience replay and expands the scope of exploration, while the asynchronization of Updater and Puller speeds up network iteration and training. These components complement each other and work in coordination to develop more reasonable strategies.
[0140] 4.3) Algorithm Implementation: In a reinforcement learning algorithm where multiple agents run independently and in parallel, the environment is partially observable. In task cycle t, agent i receives state s t , and output actions The environment uses the actions of the n agents to determine the rewards and generate t+1 Agent i has the same experience and the same action space, but different states and uses different parameters for each task.
[0141] First, the agent randomly initializes its critic network and actor networks The parameters are and pass and To initialize the target network and The weight value of . Agent i makes action choices in the following way:
[0142]
[0143] Each agent takes action a (t) Interact with the parallel environment to obtain a reward r(t) and status (t+1) . Then record a message (s (t) ,a (t) ,r (t) ,s (t+1) ) is stored in the experience replay pool D.
[0144] When agent i learns, it updates the target actor network and the target critic network in the following way:
[0145]
[0146] Afterwards, based on the reward r j and the next action choice Agent i calculates an estimate of the actual Q-value by:
[0147]
[0148] Then the loss of the Critic network (C-loss) and the loss of the Actor network (A-loss) are calculated by the following method:
[0149]
[0150] Then, calculate the gradient of the Critic network loss and the Actor network loss:
[0151]
[0152] Finally, the global model is updated using the gradients calculated by each agent to complete this round of experience learning.
[0153] In another embodiment of the present invention, a client scheduling system based on a reputation model in a vehicle edge computing network is provided. The system can be used to implement the above-mentioned vehicle edge computing network client scheduling method based on a reputation model, specifically including:
[0154] Scenario establishment module: responsible for establishing the computing scenario of drone-assisted vehicles in the vehicle edge computing network, including determining the network topology and participant role functions.
[0155] Communication model establishment module: establish the communication model between the vehicle and the drone, define the communication protocol and data transmission mechanism, and ensure effective information transmission and interaction.
[0156] Computing model building module: Determines the vehicle's computing power and resource configuration, including local computing parameters and transmission delays, providing a basis for task allocation and execution.
[0157] Optimization target setting module: sets the optimization targets of the edge computing system, including minimizing resource consumption and improving model accuracy, and guides the overall optimization direction of the system.
[0158] Reputation model building module: build direct reputation model and indirect reputation model, evaluate the reliability of the client, and provide a reputation evaluation basis for user recruitment and task allocation.
[0159] User recruitment method module: proposes a complete user recruitment method, including collecting historical performance data and reputation evaluation information of vehicles, comprehensively evaluating the reputation of vehicles and selecting appropriate clients to participate in the task.
[0160] Algorithm optimization module: The asynchronous superior action criticism algorithm and the deep deterministic policy gradient algorithm are integrated to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and this algorithm is used for resource scheduling to improve system efficiency and performance.
[0161] In another embodiment of the present invention, a terminal device is provided, the terminal device includes a processor and a memory, the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of the vehicle edge computing network client scheduling method based on the reputation model.
[0162] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage.
[0163] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the vehicle edge computing network client scheduling method based on the reputation model in the above-mentioned embodiment; one or more instructions in the computer-readable storage medium are loaded and executed by the processor.
[0164] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0165] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0166] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0168] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the implementation methods of the present invention, and should be understood that the protection scope of the present invention is not limited to such special statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A vehicle edge computing network client scheduling method based on a reputation model, characterized in that: The following steps are involved: Step 1: Establish an edge computing scenario for drone-assisted vehicles. This scenario covers scenario construction, communication model, computing model, federation model, and optimization goals, including: S11. Build a scenario of drone-assisted vehicles, including an edge computing scenario consisting of drones and vehicles. The vehicle is responsible for data collection and local model training, and the drone is responsible for scheduling and maintaining the update of the global model. The vehicle, as an edge node, has an independent data set and computing storage capabilities, and the drone manages vehicle scheduling as a central server. The system runs in discrete time, and the task generation interval follows a Poisson distribution. S12. Establish a communication model for drone-assisted vehicles, describe the uplink channel as a flat Rayleigh fading channel, and each drone connects and dispatches multiple vehicles; drones move to favorable positions according to road traffic distribution and issue periodic round instructions to vehicles, including client selection, global model distribution and resource scheduling; calculate the channel gain and transmission rate during vehicle-drone communication, taking into account obstacles between communication links; S13. Establish a computing model for drone-assisted vehicles, generate tasks for the vehicles, define task models and local computing parameters, calculate the time cost and resource consumption of local model training, and consider the transmission process delay and energy consumption; S14. Establish a federated model of drone-assisted vehicles. The drone is responsible for coordinating and aggregating local models, generating a global model, managing the local models trained by each device, and defining the model loss of local models and global models. The goal is to minimize the loss function of the global model parameters. S15. Set the optimization goal of the drone-assisted vehicle, jointly optimize the client selection and resource scheduling, select vehicles with rich computing resources and sufficient data, screen out malicious vehicle attacks, ensure the minimum loss of the global model, and use reasonable resource scheduling methods to reduce the loss of training accuracy and time and resource cost consumption; model the optimization goals of the edge computing network; Step 2: Establish a dynamic interactive reputation model by constructing a direct reputation model and an indirect reputation model, including: S21. Build a direct reputation model to evaluate the reliability of the client; S22, build an indirect reputation model that comprehensively considers the historical records of other drones and collaborative evaluation of the reputation of the vehicle; Step 3: Propose a complete user recruitment method based on the reputation model, and use the user recruitment method to select clients, including: S31. Collect historical performance data and credit rating information of the vehicle; S32. Integrate direct and indirect reputation models to conduct comprehensive evaluation and reputation scoring of vehicles; S33, select appropriate clients to participate in the task according to the reputation score and system requirements; S34, updating the reputation evaluation of the vehicle according to the task execution results, and providing feedback and recording; Step 4: Combine the asynchronous superior action criticism algorithm and the deep deterministic policy gradient algorithm to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and use the optimized asynchronous parallel DDPG algorithm for resource scheduling, including: S41: Algorithm selection, combining the asynchronous advantage action criticism algorithm and the deep deterministic policy gradient algorithm to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm; S42: Algorithm optimization, which optimizes the asynchronous parallel DDPG algorithm, including adjusting the neural network structure, optimizing hyperparameter settings, and improving training strategies to improve the performance and efficiency of the algorithm; S43: Resource scheduling, using the optimized asynchronous parallel DDPG algorithm for client scheduling and resource allocation, ensuring the effective selection of vehicles for task processing in the vehicle edge computing network, and dynamically adjusting the resource allocation strategy according to the real-time situation.
2. The vehicle edge computing network client scheduling method according to claim 1, characterized in that: The specific process of step 2 is as follows: S21. Construct a direct reputation model to evaluate customer reputation based on dynamic factors of vehicle performance, capability, and cost, taking into account federated model similarity, communication link smoothness, number of data sets, and estimated computation time, and establish a direct reputation model that takes into account time correlation; S22. Construct an indirect reputation model, utilize the historical reputation records of various drones through model similarity evaluation, comprehensively consider the reputation records of other drones, evaluate vehicle reputation through collaborative cooperation, and construct a comprehensive indirect reputation model as the core of the user recruitment method to provide a basis for optimizing client selection.
3. The vehicle edge computing network client scheduling method according to claim 1, characterized in that: The specific process of step three is as follows: S31,Data Collection, the drone collects vehicle client data from historical mission experiences, including local historical model accuracy, dataset quality, and computing power; S32, reputation modeling, sort out all user data in the jurisdiction, establish a direct reputation model based on the dynamic factors of performance, capacity and cost, establish an indirect reputation model based on model similarity and historical reputation records, and build a complete reputation model to provide a basis for client selection; S33, user selection, using the reputation model to select a group of vehicles to perform the task, vehicles with high reputation values have higher reliability and data quality; S34, model training, the UAV transmits the global model to the selected vehicle, and the vehicle uses its own resource training data to train the local federated model based on the global model; S35, reputation feedback, the drone collects the local models trained locally by the vehicle, summarizes and updates the global model, records the accuracy of the local model, the quality of the data set and the computing power, calculates and updates the client's reputation score based on the vehicle's performance, and realizes the user recruitment cycle.
4. A client scheduling system based on a reputation model in a vehicle edge computing network, characterized in that: The system can be used to implement the vehicle edge computing network client scheduling method described in any one of claims 1 to 3, specifically including: Scenario establishment module: responsible for establishing the computing scenario of drone-assisted vehicles in the vehicle edge computing network, including determining the network topology and participant role functions; Communication model building module: establish the communication model between the vehicle and the drone, define the communication protocol and data transmission mechanism, and ensure effective information transmission and interaction; Computational model building module: determines the vehicle's computing power and resource configuration, including local computing parameters and transmission delays, providing a basis for task allocation and execution; Optimization target setting module: sets the optimization target of the edge computing system, including minimizing resource consumption and improving model accuracy, and guides the overall optimization direction of the system; Reputation model building module: build direct reputation model and indirect reputation model, evaluate the reliability of the client, and provide a reputation evaluation basis for user recruitment and task allocation; User recruitment method module: Proposes a complete user recruitment method, including collecting historical performance data and reputation evaluation information of vehicles, comprehensively evaluating the reputation of vehicles and selecting appropriate clients to participate in the task; Algorithm optimization module: The asynchronous superior action criticism algorithm and the deep deterministic policy gradient algorithm are integrated to optimize the asynchronous parallel DDPG algorithm as the client scheduling algorithm, and this algorithm is used for resource scheduling to improve system efficiency and performance.
5. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the vehicle edge computing network client scheduling method described in one of claims 1 to 3 is implemented.
6. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which, when executed by a processor, implements the vehicle edge computing network client scheduling method described in one of claims 1 to 3.
Citation Information
Patent Citations
User data privacy protection method based on mobile edge computing
CN111339554A
Internet of vehicles data security sharing method and system based on block chain and dynamic reputation
CN116233177A