Unmanned aerial vehicle auxiliary edge storage service optimization method based on erasure codes
By adopting erasure coding technology and deep reinforcement learning algorithms in the drone-assisted edge storage system, data encoding and placement strategies are optimized, and D2D link instability and data access delay caused by user random mobility is solved, and efficient data distribution and bandwidth resource allocation are achieved.
Patent Information
- Application Number
- CN202510134926.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
In drone-assisted edge storage systems, users' random mobility leads to intermittent connections in D2D links. Encoded block storage locations are crucial for data access and seamless switching, and data encoding and placement policies affect data access latency and low bandwidth resource allocation efficiency.
A drone-assisted edge storage service optimization method based on erasure coding is proposed. By building a system architecture, the original data is blocked and encoded and stored on the drone and edge server. Combining the trajectory prediction model of CNN and ConvLSTM, user mobility is modeled and data encoding and placement strategies are optimized. The hierarchical deep reinforcement learning algorithm is used to solve the problem process, and data distribution and bandwidth resource allocation are optimized.
It effectively solves the problem of D2D link instability caused by user random mobility, optimizes data access delay, improves bandwidth resource allocation efficiency, and realizes the optimal solution to simultaneously optimize data placement and content transmission.
Smart Images

Figure CN120066409A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of edge computing technology, and particularly to an optimization method for an unmanned aerial vehicle (UAV)-assisted edge storage service based on erasure codes. Background Art
[0002] In recent years, UAV-assisted mobile edge computing has been regarded as a key technology for improving wireless connection and expanding coverage due to its flexible deployment and reliable line-of-sight communication characteristics. UAVs equipped with storage resources can be dynamically deployed through device-to-device (D2D) communication to provide data storage services for mobile users with poor connection quality. Compared with traditional mobile edge computing systems composed only of static servers and other ground devices, this architecture can significantly improve the quality of service.
[0003] Compared with the multi-copy technology, erasure codes can achieve higher data reliability and availability at a lower storage cost. In an erasure code storage scheme, the original data is split into k data blocks, and m parity blocks are obtained through encoding. All data blocks are distributed and stored on k + m storage nodes. A user can retrieve any k data blocks / parity blocks (collectively referred to as encoded blocks) from the storage nodes it can access to reconstruct the original data. Erasure codes have been widely applied to cloud-based storage systems to reduce storage costs. Therefore, applying erasure code technology to UAV-assisted edge storage systems is expected to further improve the quality of service for ground mobile users.
[0004] However, there are still some challenges in using erasure codes to store data in the UAV-assisted edge storage system. The random mobility of users can lead to intermittent connections in D2D links, which is a significant challenge. Considering the user's movement pattern and the non-full connection characteristics of the UAV cluster network, the storage location of the coded blocks is crucial for realizing data access and seamless handover of D2D users. In addition, in the problem of data placement based on erasure codes, different data encoding and placement strategies directly affect the data access latency. If the coded blocks are properly placed, users can easily obtain sufficient coded blocks (k coded blocks) in the UAV cluster close to themselves and then recover the original data through short-distance D2D communication. If the coded blocks are not properly placed, they need to obtain the coded blocks from the remote edge server through the backhaul link, which will significantly increase the data access latency. Secondly, in the UAV-assisted edge storage system based on erasure codes, multiple users may initiate data requests at the same time. Similarly, the edge server and the UAV may also receive multiple requests at the same time. Due to the limited bandwidth resources of the edge server and the UAV, how to effectively implement data distribution and bandwidth resource allocation is a problem worthy of in-depth discussion. It should be noted that the data placement decision based on erasure codes and the content delivery decision affect each other and are coupled, so it is difficult to find the optimal solution that optimizes both decisions at the same time. Summary of the Invention
[0005] In view of this, the embodiments of the present application propose an optimization method for UAV-assisted edge storage services based on erasure codes, which decomposes the joint data placement and content delivery problem into a data encoding and placement sub-problem based on erasure codes, and a content delivery and resource allocation sub-problem, so as to cope with the decision coupling, find the optimal solution that optimizes both decisions at the same time, and then effectively implement data distribution and bandwidth resource allocation.
[0006] To achieve the above object, an embodiment of the present application proposes an optimization method for an unmanned aerial vehicle (UAV)-assisted edge storage service based on erasure codes. The method includes: constructing a system architecture for UAV-assisted edge storage based on erasure codes, in which the original data is chunked and encoded for storage on UAVs and edge servers to provide storage services for each mobile user; modeling the impact of the random mobility of each mobile user on decision-making to obtain a user mobility model, and constructing a trajectory prediction model combining CNN and ConvLSTM to obtain the trajectory prediction result output by the trajectory prediction model; based on the system architecture, constructing an objective function with the aim of minimizing the storage cost and the long-term average service delay of the system architecture under multiple constraint conditions; where the objective function characterizes the joint data placement and content delivery problem; decomposing the joint data placement and content delivery problem into an erasure-code-based data encoding and placement sub-problem and a content delivery and resource allocation sub-problem, and modeling the two sub-problems as a Markov decision process; using a mobile-enhanced hierarchical deep reinforcement learning algorithm to solve the Markov decision process based on the trajectory prediction result output by the trajectory prediction model to obtain an optimal solution that optimizes both sub-problems.
[0007] To achieve the above object, an embodiment of the present application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an optimization method for a UAV-assisted edge storage service based on erasure codes as described above.
[0008] To achieve the above object, an embodiment of the present application also proposes a computer-readable storage medium storing a computer program, which can implement an optimization method for a UAV-assisted edge storage service based on erasure codes as described above when executed by a processor.
[0009] An optimization method for drone-assisted edge storage service based on erasure code proposed in the embodiments of this application constructs a system architecture for drone-assisted edge storage based on erasure code. In the system architecture, the original data is divided into blocks and encoded and stored on drones and edge servers to provide storage services for each mobile user. Then, a user mobility model is modeled for the impact of the random mobility of each mobile user on decision-making, and a trajectory prediction model combining CNN and ConvLSTM is constructed to obtain the trajectory prediction results output by the trajectory prediction model. Based on the system architecture, under multiple constraint conditions, with the aim of minimizing the storage cost and the long-term average service delay of the system architecture, an objective function representing the joint data placement and content delivery problem is constructed. The joint data placement and content delivery problem is decomposed into an erasure code-based data encoding and placement sub-problem and a content delivery and resource allocation sub-problem, and the two sub-problems are modeled as Markov decision processes. A framework for joint data placement and content delivery based on hierarchical deep reinforcement learning is proposed. Finally, using this framework for joint data placement and content delivery based on hierarchical deep reinforcement learning, based on the trajectory prediction results output by the trajectory prediction model, the Markov decision process is solved to obtain the optimal solution that simultaneously optimizes the two sub-problems. This method can well handle decision coupling, find the optimal solution that simultaneously optimizes these two decisions (sub-problems), and thus effectively realize data distribution and bandwidth resource allocation.
[0010] In some alternative embodiments, the constructed system architecture for drone-assisted edge storage based on erasure code consists of an edge server s and U drones hovering in the air, which cooperate to provide data storage services for D mobile users on the ground. The drone set and the mobile user set are respectively denoted as and , , , and the time domain is divided into time slots with equal durations ; In the system architecture, the original data is divided into blocks and encoded and stored on drones and edge servers to provide storage services for mobile users, including: The original data f is equally divided into k data blocks, and m parity check blocks are generated according to the encoding scheme. The total number of encoded blocks is N, , and the size of each encoded block is , denoting the size of the original data f; It is allowed to store at most one encoded block on each drone. At each time slot t, mobile user d has probability of initiating a data request, and mobile users prefer to obtain encoded blocks from the drones covering their areas; Coded chunks are transmitted among neighboring UAVs through the UAV swarm network topology and delivered to mobile users. If insufficient k coded chunks cannot be retrieved from the UAVs, the remaining coded chunks are obtained from the remote edge server via the UAVs. According to the UAV swarm network topology, the neighbor nodes of UAV u are defined as the nodes directly connected to the current node and capable of communication, denoted as the set of neighbor nodes , .
[0011] In some alternative embodiments, a user mobility model is developed by modeling the impact of the random mobility of each mobile user on decision-making, including: Consider a quasi-static scenario where the UAV hovers in the air and the position of the mobile user remains unchanged during data requests, but is allowed to move at different times; At any time slot t, the positions of the edge server s, UAV u, and mobile user d are respectively represented as: ; ; ; where, represents the position of the edge server s at time slot t, , , respectively represent the abscissa, ordinate, and altitude of the edge server s at time slot t, represents the position of the UAV u at time slot t, , respectively represent the abscissa and ordinate of the UAV u at time slot t, and the UAV hovers at a constant altitude , represents the position of the mobile user d, , respectively represent the abscissa and ordinate of the mobile user d at time slot t; It is stipulated that the mobile user d is allowed to move a preset distance in a fixed direction during time slot t, so as to reach the position in the next time slot, , , represents the maximum distance that the user can move in one time slot, which is calculated by the following formula: ; where, , respectively represent the mobile user d at time slot The abscissa and ordinate at a certain time; Define a coverage indicator variable to indicate whether the UAV covers the mobile user, which is expressed as:
[0012] where, takes values of 0 or 1, indicating that mobile user d is within the coverage range of UAV u at time slot t, indicating that mobile user d is outside the coverage range of UAV u at time slot t; Divide the target area into cells of equal size, and represent the position of the mobile user using only two-dimensional coordinates. Let the historical trajectories of D mobile users from time slot 1 to time slot t be expressed as: : ; where, represents the historical trajectory of mobile user d, represents the two-dimensional coordinates of mobile user d at time slot t; Thus, the user movement model is established.
[0013] In some alternative embodiments, a trajectory prediction model combining CNN and ConvLSTM is constructed, and the trajectory prediction result output by the trajectory prediction model is obtained, including: Construct a trajectory prediction model combining CNN and ConvLSTM, and input the historical trajectories of the past t time slots into the trajectory prediction model. The trajectory prediction model extracts spatial features from through CNN, and then ConvLSTM further extracts temporal features based on the spatial features to obtain spatio-temporal features, and finally makes a prediction based on the spatio-temporal features and outputs the trajectory prediction result.
[0014] In some alternative embodiments, based on the system architecture, under multiple constraint conditions, with the aim of minimizing the storage cost and the long-term average service delay of the system architecture, an objective function is constructed, including: Different data encoding and placement strategies will result in different storage costs. First, consider the data placement strategy on the UAV. For the k data blocks and m parity blocks encoded from the original data, use the data block placement decision vector and the parity block placement decision vector to represent whether each UAV stores a data block or a parity block respectively. The data block placement decision vector and the parity block placement decision vector are expressed by the formula as: ; ; where, and both take values of 0 or 1, indicating that the data block is stored on the drone u, indicating that the data block is not stored on the drone u, indicating that the parity block is stored on the drone u, indicating that the parity block is not stored on the drone u; To improve the reliability and availability of storage, it is stipulated that each storage node is allowed to store at most one coded block. Thus, the total number of coded blocks is , ; Therefore, the storage cost is defined as Cost, and Cost is expressed by the formula: ; In the system architecture, the communications involved include A2G communication between the drone and the mobile user, A2A communication between drones, and G2A communication between the edge server and the drone. All devices use FDMA technology to transmit coded blocks. The total downlink bandwidths of the edge server and the drone are respectively and ; Considering the occlusion in the 3D environment, the A2G communication and G2A communication are more likely to encounter NLoS paths. Therefore, the communication channel models between the drone and the mobile user, and between the edge server and the drone are designed to include both LoS and NLoS path loss probabilities, and the communication channel model between drones is designed to include only the LoS path loss probability; For G2A communication, the composite channel gain between the edge server s and the drone u combines the LoS component and the NLoS component, and is expressed by the formula: ; ; ; ; where, represents the composite channel gain between the edge server s and the drone u, represents the distance between the edge server s and the drone u, represents the channel gain at the reference distance at, represents the adjusted LoS probability considering the NLoS channel signal attenuation, represents the LoS probability, represents the path loss exponent, represents the NLoS attenuation factor, and are environmental constants for a specific environment, indicating the elevation angle from the edge server s to the UAV u; In the same time slot, the edge server is allowed to communicate with multiple UAVs simultaneously. Define the bandwidth resource allocation vector of the edge server for all UAVs as , , , indicating the percentage of the spectrum allocated to UAV u in time slot t; Based on and , define the data transmission speed between the edge server s and the UAV u as: ; where represents the transmit power of the edge server s, is the noise spectral density, represents the data transmission speed between the edge server s and the UAV u; For A2G communication, the composite channel gain between the UAV and the mobile user also combines the LoS component and the NLoS component, and is expressed by the formula as: ; ; ; ; where represents the composite channel gain between the UAV u and the mobile user d, represents the distance between the UAV u and the mobile user d, represents the LoS probability adjusted considering the NLoS channel signal attenuation, represents the LoS probability, represents the elevation angle from the UAV u to the mobile user d; In the same time slot, the UAV u is allowed to communicate with multiple mobile users simultaneously. Define the bandwidth resource allocation vector of the UAV u for all mobile users as , , , indicating the percentage of the spectrum allocated to the mobile user d by the UAV u in time slot t; Based on and , define the data transmission speed between the UAV u and the mobile user d as: ; where represents the transmit power of the UAV u, Denote the data transmission speed between the UAV \(u\) and the mobile user \(d\); For A2A communication, the channel gain between the UAV \(u\) and the UAV only contains the LoS component, which is expressed by the formula , allowing the UAV \(u\) to communicate with multiple UAVs simultaneously in the same time slot. Define the bandwidth resource allocation vector of the UAV \(u\) for all UAVs as , , , indicating the percentage of the spectrum allocated by the UAV \(u\) to the UAV in time slot \(t\); Based on and , define the data transmission speed between the UAV \(v\) and the UAV \(u\) as: ; where represents the transmit power of the UAV \(v\), represents the data transmission speed between the UAV \(v\) and the UAV \(u\); In the system architecture, the data request of the mobile user \(d\) requires \(k\) coded blocks to decode the original data. The positions of these coded blocks have three cases, corresponding to different transmission delays; Use the access indicator to indicate whether the mobile user \(d\) obtains the coded block from the UAV \(u\) that directly covers itself, , indicating that the mobile user \(d\) does not obtain the coded block from the UAV \(u\) that directly covers itself, indicating that the mobile user \(d\) obtains the coded block from the UAV \(u\) that directly covers itself, indicating that the coded block is stored on the UAV \(u\), indicating that the mobile user \(d\) is within the coverage of the UAV \(u\); Use the access indicator to indicate whether the mobile user \(d\) obtains the coded block from the neighbor node \(v\) of the UAV \(u\) that directly covers itself, , indicating that \(v\) is a neighbor node of the UAV \(u\); Use the access indicator to indicate that the mobile user \(d\) needs to obtain the coded block from the remote edge server \(s\) through the UAV \(u\) that directly covers itself, which is expressed as: ; When is specified as , indicating that the mobile user \(d\) obtains the coded block from the UAV \(u\) that directly covers itself, the direct access delay is expressed by the formula as: ; Among them, represents the direct access delay; When is specified as 1, it means that the mobile user d obtains the encoded block from the neighbor node v of the drone u that directly covers itself. Then, the indirect access delay consists of two parts: the communication delay between the drone v and the drone u and the communication delay between the drone u and the mobile user d, and is expressed by the formula as: ; Among them, represents the indirect access delay; When it means that the mobile user d obtains the encoded block from the remote edge server through the drone u that directly covers itself. Then, the edge access delay consists of two parts: the communication delay between the edge server s and the drone u and the communication delay between the drone u and the mobile user d, and is expressed by the formula as: ; Among them, represents the edge access delay; The total transmission delay for the mobile user d to access the encoded block from the above three locations is expressed by the formula as: ; Among them, represents the total transmission delay; Under multiple constraint conditions, with the aim of minimizing Cost and a target function is constructed.
[0015] In some alternative embodiments, a target function is constructed and is expressed by the formula as: ; ; ; ; ; ; ; ; ; ; ; ; .
[0016] In some optional embodiments, the joint data placement and content delivery problem is decomposed into a data encoding and placement sub-problem based on erasure coding, and a content delivery and resource allocation sub-problem, and the two sub-problems are modeled as a Markov decision process, including: If no coding blocks are stored on any drone, it will be impossible to determine which drones the mobile user should establish D2D communication with. Therefore, the data placement decision should be determined before the content delivery decision. To this end, the joint data placement and content delivery problem is decomposed into a data encoding and placement sub-problem based on erasure coding, and a content delivery and resource allocation sub-problem. First, the coding block access indicator and bandwidth resource allocation variables are fixed to solve the data encoding and placement sub-problem based on erasure codes. The data encoding and placement sub-problem based on erasure codes is expressed by the formula: ; ; ; In determining and Afterwards, the coding block access indicator and bandwidth resource allocation variables are further optimized through the content delivery and resource allocation sub-problem, which is expressed by the formula: ; ; ; ; ; ; ; ; ; ; ; The original approximate optimal solution is obtained by sequentially solving the data encoding and placement sub-problems based on erasure codes, as well as the content delivery and resource allocation sub-problems. To this end, a hierarchical deep reinforcement learning framework consisting of multiple UAV agents and an edge agent is proposed using the divide-and-conquer idea. The UAV agent is responsible for determining data placement decisions, the edge agent is responsible for determining content delivery decisions, and all agents use the same reward function. Cooperate with each other to minimize storage costs and the long-term service latency of the system. The reward function is defined as: ; Determine the UAV placement state. The mobility of users will cause the coverage vector of the UAV to change dynamically. Therefore, the UAV agent needs to know the positions of the mobile users from time t to future times to calculate the coverage vector. In addition, it also needs to know the set of neighbor nodes , and the UAV placement state is represented as , ; Determine the UAV placement action. The UAV agent makes an action on whether to store data blocks or parity blocks. The UAV placement action is represented as , , when , it means that UAV u does not store any coded blocks, means that UAV u stores data blocks, means that UAV u stores parity blocks; Determine the edge state. The edge agent is responsible for specifying the access locations of coded blocks for each mobile user and allocating bandwidth resources for each mobile user. Therefore, the edge agent needs to know the bandwidth resources of the edge server, the bandwidth resource information of the UAVs, the UAV cluster network topology, the mobile trajectories of the users, and the placement information of the coded blocks. The edge state is represented as , ; Determine the edge action. The action generated by the edge agent is the location node for each mobile user to access coded blocks and the allocated bandwidth resources. The edge action is represented as ; .
[0017] In some alternative embodiments, the mobile-enhanced hierarchical deep reinforcement learning algorithm, based on the trajectory prediction results output by the trajectory prediction model, solves the Markov decision process to obtain the optimal solution that simultaneously optimizes two sub-problems, including: Based on the DDQN algorithm to solve the data encoding and placement sub-problem based on erasure codes. The DDQN algorithm contains two DNNs, namely the Q current network for parameter training and the Q target network for forward propagation to generate the target Q value. The Q value update function is: ; where is the discount factor; The loss function of the DDQN algorithm is: ; To improve the exploration efficiency of DDQN, when making action selections, the UAV agent has probability to select a random action; Since the edge agent needs to collect global complex information to make decisions, the PPO algorithm based on the actor-critic framework is adopted to solve the content delivery and resource allocation sub-problems; For the actor network, the probability ratio is used to quantify the change before and after policy update under the same state and action, which is expressed by the formula: ; On the other hand, the temporal difference residual is used to calculate the advantage function to evaluate the actual return and expected return of the selected action under the current policy, which is expressed by the formula: ; ; where is the state value function; At the same time, to enhance the stability and efficiency of the calculation of the advantage function generalized advantage estimation is introduced into the network. By introducing the weight parameter the advantage function is calculated more smoothly, and the calculated smoothed advantage function is expressed by the formula: ; To prevent the training from being unstable due to excessive policy updates, PPO uses a clipping function to ensure the update range, and the clipping function is expressed by the formula: ; where represents the clipping function; Based on this, the objective of the actor network is expressed as: ; where represents the policy entropy that encourages exploration under the current policy, represents the weight of the policy entropy, represents the objective of the actor network; The critic network is used to evaluate the expected reward in a certain state. It takes the state as the input and outputs the relevant value function , and updates itself by reducing the difference between the predicted value function and the calculated target value. The objective of the critic network is expressed as: ; Among them, represents the target value using GAE, which is calculated by combining discounted rewards and advantage estimation from state and is the target of the critic network. Description of the Drawings
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 is a flowchart of an erasure code-based UAV-assisted edge storage service optimization method provided in an embodiment of the present application; Figure 2 is a schematic diagram of the composition of a system architecture of an erasure code-based UAV-assisted edge storage provided in an embodiment of the present application; Figure 3 is a schematic diagram of the model structure of a trajectory prediction model combining CNN and ConvLSTM provided in an embodiment of the present application; Figure 4 is a schematic diagram of the structure of an electronic device provided in another embodiment of the present application. Detailed Embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will elaborate on each embodiment of the present application with reference to the drawings. In various embodiments of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The following division of each embodiment is only for convenience of description and should not constitute any limitation to the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other without conflict.
[0021] An embodiment of the present application proposes an optimization method for an erasure code-based UAV-assisted edge storage service, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and each of the following embodiments, the electronic device is described by taking the server as an example. The implementation details of an erasure code-based UAV-assisted edge storage service optimization method proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenience of understanding and is not necessary for implementing this solution.
[0022] The specific process of an erasure code-based UAV-assisted edge storage service optimization method proposed in this embodiment can be as Figure 1 shown, including: Step 101, construct a system architecture for erasure code-based UAV-assisted edge storage. In the system architecture, the original data is divided into blocks and encoded and stored on the UAVs and the edge server to provide storage services for each mobile user.
[0023] In specific implementation, the server first needs to construct a system architecture for erasure code-based UAV-assisted edge storage. In this system architecture, the original data is divided into blocks and encoded and stored on the UAVs and the edge server, so as to provide storage services for each mobile user.
[0024] In an example, the system architecture for erasure code-based UAV-assisted edge storage constructed by the server consists of an edge server s and U UAVs hovering in the air, which cooperate to provide data storage services for D mobile users on the ground. The UAV set and the mobile user set are respectively represented as and , , , in this system architecture, the time domain is divided into time slots with equal durations . The specific composition of the system architecture for erasure code-based UAV-assisted edge storage constructed by the server can be as Figure 2 shown.
[0025] In the system architecture, the original data is divided into blocks and encoded and stored on the UAVs and the edge server to provide storage services for mobile users. This system architecture equally divides the original data f into k data blocks and generates m parity check blocks according to the encoding scheme. The total number of encoded blocks is N, , and the size of each encoded block is , represents the size of the original data f.
[0026] Typically, the edge server s has sufficient storage resources to store all the encoded blocks. However, to improve the reliability and availability of storage, the system architecture allows storing at most one encoded block on each drone. In each time slot t, the mobile user d has a probability of initiating a data request, and each mobile user will preferentially obtain the encoded blocks from the drones covering its area.
[0027] In addition, the encoded blocks can be transmitted between adjacent drones through the drone cluster network topology and delivered to the mobile users in need. If enough k encoded blocks cannot be retrieved from the drones, the remaining encoded blocks need to be obtained from the remote edge server through the drones.
[0028] According to the drone cluster network topology, the system architecture defines the neighbor nodes of the drone u as the nodes directly connected to the current node and capable of communication, denoted as the set of neighbor nodes , .
[0029] Step 102: Model the impact of the random mobility of each mobile user on decision-making to obtain a user mobility model, and construct a trajectory prediction model combining CNN and ConvLSTM to obtain the trajectory prediction results output by the trajectory prediction model.
[0030] In a specific implementation, to cooperate with the system architecture, the server models the impact of the random mobility of each mobile user on decision-making to obtain a user mobility model, and constructs a trajectory prediction model combining CNN and ConvLSTM to obtain the trajectory prediction results output by the trajectory prediction model.
[0031] In an example, consider a quasi-static scenario where the drones hover in the air and the positions of the mobile users remain unchanged during data requests but are allowed to move (possibly move) at different times.
[0032] For unified expression, the positions of all devices are represented as variables related to the time slot t. Specifically, at any time slot t, the positions of the edge server s, the drone u, and the mobile user d are respectively represented as: ; ; ; where represents the position of the edge server s at time slot t, , , respectively represent the abscissa, ordinate, and altitude of the edge server s at time slot t, represents the position of the drone u at time slot t, , respectively represent the abscissa and ordinate of the UAV u at time slot t. The UAV u hovers at a constant altitude . represents the position of the mobile user d, , respectively represent the abscissa and ordinate of the mobile user d at time slot t.
[0033] It is stipulated that the mobile user d is allowed to move a preset distance in a fixed direction during time slot t so as to reach the position in the next time slot . . . represents the maximum distance that the user can move in one time slot, which is calculated by the following formula: ; where , respectively represent the abscissa and ordinate of the mobile user d at time slot .
[0034] Due to the random movement of the users and the limited service range of each UAV, the coverage indicator variable is defined to indicate whether the UAV covers the mobile user, which is expressed as:
[0035] where takes values of 0 or 1, indicates that the mobile user d is within the coverage range of the UAV u at time slot t, indicates that the mobile user d is outside the coverage range of the UAV u at time slot t.
[0036] Since the mobile users all move on a two-dimensional plane, to simplify the problem, the target area is divided into cells of equal size, and only two-dimensional coordinates are used to represent the positions of the mobile users. Let the historical trajectories of D mobile users from time slot 1 to time slot t be represented as: : ; where represents the historical trajectory of the mobile user d, represents the two-dimensional coordinates of the mobile user d at time slot t.
[0037] So far, the server has modeled the user movement model.
[0038] In one example, since only predicting the location of a mobile user at a certain moment in the future (time slot and moment have the same meaning) is not enough to support subsequent complete decision making, the trajectory prediction model constructed adopts a sequence-to-sequence prediction method. The server constructs a Figure 3 The trajectory prediction model combining CNN and ConvLSTM shown in the figure converts the historical trajectory of the past t time slots into Input to the trajectory prediction model, the trajectory prediction model uses CNN from The spatial features are extracted from the ConvLSTM, and then the ConvLSTM performs further temporal feature extraction based on the spatial features to obtain further temporal features and spatial features (space-time features). Finally, the prediction is performed based on the space-time features and the trajectory prediction result is output.
[0039] Step 103 , based on the system architecture, under multiple constraints, and with the goal of minimizing storage costs and long-term average service delay of the system architecture, construct an objective function, wherein the objective function represents the joint data placement and content delivery problem.
[0040] In the specific implementation, the next step is to construct the objective function. Based on the system architecture, the server constructs the objective function under multiple constraints with the purpose of minimizing the storage cost and the long-term average service delay of the system architecture. The objective function represents the joint data placement and content delivery problem.
[0041] In an example, different data encoding and placement strategies will lead to different storage costs. First, consider the data placement strategy on the drone. For the k data blocks and m check blocks encoded from the original data, the data block placement decision vector and the check block placement decision vector are used to indicate whether the data block or check block is stored on each drone. The data block placement decision vector and the check block placement decision vector are expressed by the formula: ; ; in, and The value of is 0 or 1. Indicates that the data block is stored on the drone u. Indicates that there is no data block stored on drone u. Indicates that the check block is stored on drone u. Indicates that there is no checksum block stored on drone u.
[0042] In order to improve the reliability and availability of storage, it is stipulated that each storage node is allowed to store at most one coding block. , the total number of coding blocks is , .
[0043] Thus, we can define the storage cost as Cost, which is expressed by the formula as: .
[0044] In the system architecture, the communications involved include A2G communication between the drone and the mobile user, A2A communication between drones, and G2A communication between the edge server and the drone. All devices use FDMA technology to transmit coded blocks. The total downlink bandwidths of the edge server and the drone are respectively and .
[0045] Considering the occlusion in the 3D environment, the A2G communication and G2A communication are more likely to encounter NLoS paths. Therefore, the communication channel models between the drone and the mobile user, and between the edge server and the drone are designed to include both LoS and NLoS path loss probabilities, while the communication channel model between drones is designed to include only the LoS path loss probability.
[0046] The channel gain related to the LoS link can be expressed by , and . The channel gain related to the NLoS link can be expressed by and . represents the channel gain at the reference distance , represents the path loss exponent, represents the NLoS attenuation factor.
[0047] For G2A communication, the composite channel gain between the edge server s and the drone u combines the LoS component and the NLoS component, and is expressed by the formula as: ; ; ; ; where represents the composite channel gain between the edge server s and the drone u, represents the distance between the edge server s and the drone u, represents the channel gain at the reference distance , represents the LoS probability adjusted considering the NLoS channel signal attenuation, represents the LoS probability, represents the path loss exponent, Denotes the NLoS attenuation factor, and is the environmental constant in a specific environment, denotes the elevation angle from the edge server s to the UAV u.
[0048] In the same time slot, the edge server is allowed to communicate with multiple UAVs simultaneously. Define the bandwidth resource allocation vector of the edge server for all UAVs as , , , which represents the percentage of the spectrum allocated to UAV u in time slot t.
[0049] Based on and , define the data transmission speed between the edge server s and the UAV u as: ; wherein, denotes the transmit power of the edge server s, is the noise spectral density, denotes the data transmission speed between the edge server s and the UAV u.
[0050] For A2G communication, the composite channel gain between the UAV and the mobile user also combines the LoS component and the NLoS component, and is expressed by the formula as: ; ; ; ; wherein, denotes the composite channel gain between the UAV u and the mobile user d, denotes the distance between the UAV u and the mobile user d, denotes the LoS probability adjusted considering the NLoS channel signal attenuation, denotes the LoS probability, denotes the elevation angle from the UAV u to the mobile user d.
[0051] In the same time slot, the UAV u is allowed to communicate with multiple mobile users simultaneously. Define the bandwidth resource allocation vector of the UAV u for all mobile users as , , , which represents the percentage of the spectrum allocated by the UAV u to the mobile user d in time slot t.
[0052] Based on and , define the data transmission speed between the drone u and the mobile user d as: ; Among them, represents the transmission power of the drone u, represents the data transmission speed between the drone u and the mobile user d.
[0053] For A2A communication, the channel gain between the drone u and the drone only contains the LoS component, which is expressed by the formula , allowing the drone u to communicate with multiple drones simultaneously in the same time slot. Define the bandwidth resource allocation vector of the drone u for all drones as , , , representing the percentage of the spectrum allocated by the drone u to the drone in time slot t.
[0054] Based on and , define the data transmission speed between the drone v and the drone u as: ; Among them, represents the transmission power of the drone v, represents the data transmission speed between the drone v and the drone u.
[0055] In a storage system using erasure codes, the data reading delay includes the transmission delay of the encoded blocks and the data decoding delay. With the development of erasure code technology, more and more low-complexity coding schemes have been proposed, and the encoding and decoding complexity has gradually decreased. Therefore, compared with the communication delay, the decoding delay can be ignored.
[0056] In the system architecture, for the data request of the mobile user d, k encoded blocks are required to decode the original data. The positions of these encoded blocks have three situations, corresponding to different transmission delays (i.e., direct access delay, indirect access delay, and edge access delay) respectively. The server needs to build a system delay model.
[0057] Use the access indicator to indicate whether the mobile user d obtains the encoded block from the drone u that directly covers itself, , indicates that the mobile user d does not obtain the encoded block from the drone u that directly covers itself, indicates that the mobile user d obtains the encoded block from the drone u that directly covers itself, indicates that the encoded block is stored on the drone u, Indicates that the mobile user d is within the coverage of the drone u.
[0058] Use the access indicator to indicate whether the mobile user d obtains the encoded block from the neighbor node v of the drone u that directly covers itself, , indicating that v is a neighbor node of the drone u.
[0059] Use the access indicator to indicate that the mobile user d needs to obtain the encoded block from the remote edge server s through the drone u that directly covers itself, expressed as: .
[0060] When is specified as , it indicates that the mobile user d obtains the encoded block from the drone u that directly covers itself, and the direct access delay is expressed by the formula as: ; where represents the direct access delay.
[0061] When is specified as 1, it indicates that the mobile user d obtains the encoded block from the neighbor node v of the drone u that directly covers itself, and the indirect access delay consists of two parts: the communication delay between the drone v and the drone u and the communication delay between the drone u and the mobile user d, and is expressed by the formula as: ; where represents the indirect access delay.
[0062] When , it indicates that the mobile user d obtains the encoded block from the remote edge server through the drone u that directly covers itself, and the edge access delay consists of two parts: the communication delay between the edge server s and the drone u and the communication delay between the drone u and the mobile user d, and is expressed by the formula as: ; where represents the edge access delay.
[0063] The total transmission delay for the mobile user d to access the encoded block from the above three locations is expressed by the formula as: ; where represents the total transmission delay.
[0064] Under multiple constraint conditions, with the aim of minimizing Cost and, an objective function is constructed.
[0065] In one example, different data encodings (different k and m), placement strategies (on which drones the encoded blocks are stored), and content delivery strategies (which drones to establish D2D links with and how to allocate bandwidth resources) in the system architecture will result in different storage costs and access latencies. For example, storing an encoded block on all drones will result in high storage costs, although this may reduce the access latency of mobile users. In contrast, choosing a small number of drones to store the encoded blocks can reduce the storage cost, but usually leads to longer transmission delays because more encoded blocks need to be transmitted from remote edge servers. Based on this, the objective function constructed by the server can be expressed by the formula: ; ; ; ; ; ; ; ; ; ; ; ; .
[0066] Step 104, decompose the joint data placement and content delivery problem into an erasure-coding-based data encoding and placement sub-problem, and a content delivery and resource allocation sub-problem, and model the two sub-problems as Markov decision processes.
[0067] In a specific implementation, the server decomposes the joint data placement and content delivery problem into an erasure-coding-based data encoding and placement sub-problem, and a content delivery and resource allocation sub-problem, and models the two sub-problems as Markov decision processes.
[0068] If there are no encoded chunks stored on the drones, it will be impossible to determine which drones the mobile users should establish D2D communication with. Therefore, the data placement decision should be determined before the content delivery decision. To this end, the joint data placement and content delivery problem is decomposed into an erasure-coding-based data encoding and placement sub-problem, and a content delivery and resource allocation sub-problem.
[0069] First, fix the encoded chunk access indicator and the bandwidth resource allocation variable, and solve the erasure-coding-based data encoding and placement sub-problem. The erasure-coding-based data encoding and placement sub-problem is expressed by the formula: ; ; .
[0070] After determining and , further optimize the encoded chunk access indicator and the bandwidth resource allocation variable through the content delivery and resource allocation sub-problem. The content delivery and resource allocation sub-problem is expressed by the formula: ; ; ; ; ; ; ; ; ; ; .
[0071] By successively solving the erasure-coding-based data encoding and placement sub-problem and the content delivery and resource allocation sub-problem, an original approximate optimal solution is obtained. To this end, using the divide-and-conquer idea, a hierarchical deep reinforcement learning framework consisting of multiple drone agents and one edge agent is proposed. Among them, the drone agents are responsible for deciding the data placement decision, and the edge agent is responsible for determining the content delivery decision. All agents cooperate with each other through the same reward function to minimize the storage cost and the long-term service latency of the system. The reward function is defined as: ; Determine the UAV placement status. The mobility of users will cause the coverage vector of the UAV to change dynamically. Therefore, the UAV agent needs to know the position of the mobile user from the current time t to the future time to calculate the coverage vector. In addition, it also needs to know the set of neighbor nodes . The UAV placement status is represented as , ; Determine the UAV placement action. The UAV agent makes an action on whether to store data blocks or parity blocks. The UAV placement action is represented as , . When , it means that UAV u does not store any coded blocks. means that UAV u stores data blocks. means that UAV u stores parity blocks; Determine the edge status. The edge agent is responsible for specifying the access location of coded blocks for each mobile user and for allocating bandwidth resources for each mobile user. Therefore, the edge agent needs to know the bandwidth resources of the edge server, the bandwidth resource information of the UAV, the UAV cluster network topology, the mobile trajectory of the user, and the placement information of the coded blocks. The edge status is represented as , ; Determine the edge action. The action generated by the edge agent is the location node for each mobile user to access the coded blocks and the allocated bandwidth resources. The edge action is represented as ; .
[0072] Step 105: Use the mobile-enhanced hierarchical deep reinforcement learning algorithm. Based on the trajectory prediction results output by the trajectory prediction model, solve the Markov decision process to obtain the optimal solution that simultaneously optimizes the two sub-problems.
[0073] In a specific implementation, the server uses the mobile-enhanced hierarchical deep reinforcement learning algorithm. Based on the trajectory prediction results output by the trajectory prediction model, solve the Markov decision process to obtain the optimal solution that simultaneously optimizes the two sub-problems.
[0074] In an example, since the UAV agent receives simple information to generate discrete actions, the DDQN algorithm can be used to solve the data placement problem. Based on the data encoding and placement sub-problem of erasure codes, the DDQN algorithm contains two DNNs, namely the Q current network for parameter training and the Q target network for forward propagation to generate the target Q value. The Q value update function is: ; where is the discount factor.
[0075] The loss function of the DDQN algorithm is as follows: .
[0076] To improve the exploration efficiency of DDQN, when making action selections, the UAV agent has probability to select a random action.
[0077] Since the edge agent needs to collect global complex information to make decisions, the PPO algorithm based on the actor-critic framework is adopted to solve the content transmission and resource allocation sub-problems.
[0078] For the actor network, the probability ratio is used to quantify the change before and after policy update under the same state and action, which is expressed by the formula as: .
[0079] On the other hand, the temporal difference residual is used to calculate the advantage function , evaluating the actual return and expected return of the selected action under the current policy, which is expressed by the formula as: ; ; where, is the state value function.
[0080] Meanwhile, to enhance the stability and efficiency of the advantage function calculation, generalized advantage estimation is introduced into the network. By introducing the weight parameter to calculate the advantage function more smoothly, the calculated smoothed advantage function is expressed by the formula as: .
[0081] To prevent the training from being unstable due to excessive policy updates, PPO adopts a clipping function to ensure the update range. The clipping function is expressed by the formula as: ; where, represents the clipping function.
[0082] Based on this, the objective of the actor network is expressed as: ; where, represents the policy entropy that encourages exploration under the current policy, Represents the weight of the policy entropy, and represents the target of the actor network.
[0083] The critic network is used to evaluate the expected reward in a certain state, taking the state as input and outputting the relevant value function , and updates itself by reducing the difference between the predicted value function and the calculated target value. The target of the critic network is expressed as: ; where, represents the target value using GAE, which is calculated by combining discounted rewards and advantage estimates from state , and represents the target of the critic network.
[0084] An optimization method for an unmanned aerial vehicle (UAV)-assisted edge storage service based on erasure coding proposed in this embodiment constructs a system architecture for UAV-assisted edge storage based on erasure coding. In the system architecture, the original data is divided into blocks and encoded for storage on UAVs and edge servers, so as to provide storage services for each mobile user. Then, a user mobility model is established by modeling the impact of the random mobility of each mobile user on decision-making, and a trajectory prediction model combining CNN and ConvLSTM is constructed to obtain the trajectory prediction results output by the trajectory prediction model. Based on the system architecture, under multiple constraint conditions, with the aim of minimizing the storage cost and the long-term average service delay of the system architecture, an objective function representing the joint data placement and content delivery problem is constructed. The joint data placement and content delivery problem is decomposed into an erasure-coding-based data encoding and placement sub-problem and a content delivery and resource allocation sub-problem, and the two sub-problems are modeled as Markov decision processes. A framework for joint data placement and content delivery based on hierarchical deep reinforcement learning is proposed. Finally, using this framework and based on the trajectory prediction results output by the trajectory prediction model, the Markov decision process is solved to obtain the optimal solutions that simultaneously optimize the two sub-problems. This method can well handle decision coupling, find the optimal solutions that simultaneously optimize these two decisions (sub-problems), and thus effectively realize data distribution and bandwidth resource allocation.
[0085] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of the algorithm and process, are all within the protection scope of this application.
[0086] Another embodiment of this application proposes an electronic device, and its specific structure is asFigure 4 As shown in the figure, it includes: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein, the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to execute an optimized method for drone-assisted edge storage service based on erasure code as described in the above method embodiments.
[0087] Among them, the memory and the processor can be connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described herein. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0088] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0089] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement an optimized method for drone-assisted edge storage service based on erasure code as described in the above method embodiments.
[0090] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0091] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.
Claims
1. A method for optimizing drone-assisted edge storage services based on erasure codes, characterized in that: include: Build a system architecture for drone-assisted edge storage based on erasure coding. In the system architecture, the original data is divided into blocks and encoded and stored on drones and edge servers to provide storage services for mobile users. The user mobility model is obtained by modeling the impact of random mobility of each mobile user on decision making, and a trajectory prediction model combining CNN and ConvLSTM is constructed to obtain the trajectory prediction results output by the trajectory prediction model; Based on the system architecture, under multiple constraints, an objective function is constructed to minimize the storage cost and the long-term average service delay of the system architecture; wherein the objective function represents the joint data placement and content delivery problem; The joint data placement and content delivery problem is decomposed into the erasure code-based data encoding and placement subproblem and the content delivery and resource allocation subproblem, and the two subproblems are modeled as Markov decision processes. Using the mobile-enhanced hierarchical deep reinforcement learning algorithm, the Markov decision process is solved based on the trajectory prediction results output by the trajectory prediction model, and the optimal solution for simultaneously optimizing the two sub-problems is obtained.
2. According to the method for optimizing drone-assisted edge storage services based on erasure codes in claim 1, it is characterized in that: The proposed erasure code-based drone-assisted edge storage system architecture consists of an edge server s and U drones hovering in the air, which work together to provide data storage services for D mobile users on the ground. The drone set and mobile user set are represented as and , , , dividing the time domain into time slots of equal duration ; In the system architecture, raw data is divided into blocks and encoded and stored on drones and edge servers, providing storage services for mobile users, including: Divide the original data f into k data blocks, and generate m parity check blocks according to the coding scheme. The total number of coding blocks is N. , the size of each coding block is , Indicates the size of the original data f; It is allowed to store at most one coded block on each UAV. At each time slot t, mobile user d has The probability of initiating a data request is , and the mobile user preferentially obtains the code block from the UAV covering its area; The coded blocks are transmitted between adjacent drones through the drone cluster network topology and delivered to the mobile users. If enough k coded blocks cannot be retrieved from the drones, the remaining coded blocks are obtained from the remote edge server through the drones; According to the drone cluster network topology, the neighbor nodes of drone u are defined as nodes that are directly connected to the current node and can communicate with it, which is represented as the neighbor node set , .
3. According to claim 2, a method for optimizing drone-assisted edge storage services based on erasure codes is characterized in that: The user mobility model is obtained by modeling the impact of random mobility of each mobile user on decision making, including: Consider a quasi-static scenario where the drone is hovering in the air and the position of the mobile user always remains unchanged during the data request, but movement is allowed at different periods; At any time slot t, the locations of edge server s, drone u, and mobile user d are expressed as: ; ; ; in, represents the location of edge server s at time slot t, , , They represent the horizontal coordinate, vertical coordinate and height of edge server s at time slot t, represents the position of UAV u at time slot t, , They represent the horizontal and vertical coordinates of UAV u at time slot t, respectively. UAV u is hovering at a constant height. , represents the location of mobile user d, , They represent the horizontal and vertical coordinates of mobile user d at time slot t respectively; It is stipulated that mobile user d is allowed to move in a fixed direction in time slot t Move preset distance , thus reaching the position of the next time slot , , , Indicates the maximum distance a user can move in a time slot. Calculated by the following formula: ; in, , They represent the mobile user d in time slot The horizontal and vertical coordinates of time; Defining coverage indicator variables To indicate whether the drone covers the mobile user, It is expressed as: in, The value of is 0 or 1. It means that mobile user d is within the coverage area of drone u at time slot t. It means that mobile user d is outside the coverage of UAV u at time slot t; The target area is divided into cells of equal size, and only two-dimensional coordinates are used to represent the location of mobile users. The historical trajectories of D mobile users from time slot 1 to time slot t are expressed as: : ; in, represents the historical trajectory of mobile user d, represents the two-dimensional coordinates of mobile user d in time slot t; At this point, the user mobility model is obtained.
4. According to the method for optimizing drone-assisted edge storage services based on erasure codes in claim 3, it is characterized in that: Construct a trajectory prediction model combining CNN and ConvLSTM to obtain the trajectory prediction results output by the trajectory prediction model, including: Construct a trajectory prediction model combining CNN and ConvLSTM to transform the historical trajectory of the past t time slots Input to the trajectory prediction model, the trajectory prediction model uses CNN from The spatial features are extracted from the ConvLSTM, and then the ConvLSTM further extracts the temporal features based on the spatial features to obtain the spatiotemporal features. Finally, the prediction is made based on the spatiotemporal features and the trajectory prediction results are output.
5. According to claim 4, a method for optimizing drone-assisted edge storage services based on erasure codes is characterized in that: Based on the system architecture, under multiple constraints, the objective function is constructed to minimize the storage cost and the long-term average service delay of the system architecture, including: Different data encoding and placement strategies will lead to different storage costs. First, consider the data placement strategy on the drone. For the k data blocks and m check blocks encoded from the original data, the data block placement decision vector and the check block placement decision vector are used to indicate whether the data block or check block is stored on each drone. The data block placement decision vector and the check block placement decision vector are expressed by the formula: ; ; in, and The value of is 0 or 1. Indicates that the data block is stored on the drone u. Indicates that there is no data block stored on drone u. Indicates that the check block is stored on drone u. Indicates that there is no checksum block stored on drone u; In order to improve the reliability and availability of storage, it is stipulated that each storage node is allowed to store at most one coding block. , the total number of coding blocks is , ; Therefore, the storage cost is defined as Cost, which is expressed by the formula: ; In the system architecture, the communications involved include A2G communication between drones and mobile users, A2A communication between drones, and G2A communication between edge servers and drones. All devices use FDMA technology to transmit coding blocks. The total downlink bandwidth of the edge server and drone is and ; Considering the occlusion in the 3D environment, A2G communication and G2A communication are more likely to encounter NLoS paths. Therefore, the communication channel models between drones and mobile users and between edge servers and drones are designed to include both LoS and NLoS path loss probabilities, and the communication channel model between drones is designed to include only the LoS path loss probability. For G2A communication, the composite channel gain between the edge server s and the UAV u combines the LoS component and the NLoS component and is expressed as: ; ; ; ; in, represents the composite channel gain between edge server s and UAV u, represents the distance between the edge server s and the drone u, Indicates the reference distance The channel gain at represents the LoS probability after adjusting for NLoS channel signal attenuation, represents the LoS probability, represents the path loss exponent, represents the NLoS attenuation factor, and is the environmental constant under a specific environment, represents the elevation angle from the edge server s to the drone u; In the same time slot, the edge server is allowed to communicate with multiple drones simultaneously, and the bandwidth resource allocation vector of the edge server to all drones is defined as , , , represents the percentage of spectrum allocated to UAV u in time slot t; based on and , the data transmission speed between the edge server s and the drone u is defined as: ; in, represents the transmission power of edge server s, is the noise spectral density, represents the data transmission speed between the edge server s and the drone u; For A2G communication, the composite channel gain between the drone and the mobile user also combines the LoS component and the NLoS component, which is expressed by the formula: ; ; ; ; in, represents the composite channel gain between UAV u and mobile user d, represents the distance between UAV u and mobile user d, represents the LoS probability after adjusting for NLoS channel signal attenuation, represents the LoS probability, represents the elevation angle from UAV u to mobile user d; In the same time slot, drone u is allowed to communicate with multiple mobile users simultaneously, and the bandwidth resource allocation vector of drone u to all mobile users is defined as , , , represents the percentage of spectrum allocated to mobile user d by drone u in time slot t; based on and , the data transmission speed between UAV u and mobile user d is defined as: ; in, represents the transmission power of UAV u, represents the data transmission speed between UAV u and mobile user d; For A2A communication, drone u and drone The channel gain between contains only the LoS component and is expressed as , allowing UAV u to communicate with multiple UAVs simultaneously in the same time slot, and defining the bandwidth resource allocation vector of UAV u to all UAVs as , , , indicating that UAV u is assigned to UAV in time slot t The percentage of the spectrum; based on and , define the data transmission speed between UAV v and UAV u as: ; in, represents the transmission power of the UAV v, It represents the data transmission speed between UAV v and UAV u; In the system architecture, a data request from mobile user d requires k coding blocks to decode the original data. There are three positions of these coding blocks, corresponding to different transmission delays. Using Access Indicators to indicate whether the mobile user d obtains the coding block from the drone u that directly covers it. , It means that mobile user d does not obtain the coding block from the drone u that directly covers itself. It means that mobile user d obtains the coding block from the drone u that directly covers itself. Indicates that the coding block is stored on the drone u. It means that mobile user d is within the coverage of drone u; Using Access Indicators to indicate whether the mobile user d obtains the coding block from the neighbor node v of the drone u that directly covers itself, , Indicates that v is a neighbor node of drone u; Using Access Indicators To indicate that mobile user d needs to obtain the coding block from the remote edge server s through the drone u that directly covers itself, It is expressed as: ; when Designated as When , it means that mobile user d obtains the coding block from the drone u that directly covers itself, then the direct access delay is expressed by the formula: ; in, Indicates direct access delay; when When is designated as 1, it indicates that the mobile user d obtains the coding block from the neighbor node v of the UAV u that directly covers itself, then the indirect access delay is composed of the communication delay between UAV v and UAV u and the communication delay between UAV u and the mobile user d, which is expressed by the formula: ; in, Indicates indirect access delay; when When , it means that the mobile user d directly covers its own drone u from the remote edge server The edge access delay is composed of the communication delay between the edge server s and the UAV u and the communication delay between the UAV u and the mobile user d, which can be expressed as follows: ; in, represents edge access latency; The total transmission delay of mobile user d accessing the coding block from the above three locations is expressed by the formula: ; in, represents the total transmission delay; Under multiple constraints, to minimize Cost and For this purpose, the objective function is constructed.
6. The method for optimizing drone-assisted edge storage services based on erasure codes according to claim 5 is characterized in that: The constructed objective function is expressed by the formula: ; ; ; ; ; ; ; ; ; ; ; ; 。 7. The method for optimizing drone-assisted edge storage services based on erasure codes according to claim 6 is characterized in that: The joint data placement and content delivery problem is decomposed into the erasure code-based data encoding and placement sub-problem and the content delivery and resource allocation sub-problem. The two sub-problems are modeled as Markov decision processes, including: If no coding blocks are stored on any drone, it will be impossible to determine which drones the mobile user should establish D2D communication with. Therefore, the data placement decision should be determined before the content delivery decision. To this end, the joint data placement and content delivery problem is decomposed into a data encoding and placement sub-problem based on erasure coding, and a content delivery and resource allocation sub-problem. First, the coding block access indicator and bandwidth resource allocation variables are fixed to solve the data encoding and placement sub-problem based on erasure codes. The data encoding and placement sub-problem based on erasure codes is expressed by the formula: ; ; ; In determining and Afterwards, the coding block access indicator and bandwidth resource allocation variables are further optimized through the content delivery and resource allocation sub-problem, which is expressed by the formula: ; ; ; ; ; ; ; ; ; ; ; The original approximate optimal solution is obtained by sequentially solving the data encoding and placement sub-problems based on erasure codes, as well as the content delivery and resource allocation sub-problems. To this end, a hierarchical deep reinforcement learning framework consisting of multiple UAV agents and an edge agent is proposed using the divide-and-conquer idea. The UAV agent is responsible for determining data placement decisions, the edge agent is responsible for determining content delivery decisions, and all agents use the same reward function. Cooperate with each other to minimize storage costs and long-term service delays of the system. The reward function Defined as: ; Determine the placement status of the drone. The mobility of the user will cause the coverage vector of the drone to change dynamically. Therefore, the drone agent needs to know the mobile user's position from time t to the future. The position at the moment is used to calculate the coverage vector. In addition, the set of neighbor nodes needs to be known. , the drone placement state is represented as , ; Determine the drone placement action, and the drone agent takes action on whether to store the data block or the check block. The drone placement action is expressed as , ,when When , it means that drone u does not store any coding blocks. Indicates that drone u stores data blocks, Indicates that the drone u stores the check block; Determine the edge state. The edge agent is responsible for specifying the coding block access location for each mobile user and allocating bandwidth resources to each mobile user. Therefore, the edge agent needs to know the bandwidth resources of the edge server, the bandwidth resource information of the drone, the drone cluster network topology, the user's movement trajectory, and the placement information of the coding block. The edge state is expressed as , ; Determine the edge action. The action generated by the edge agent is the location node and the allocated bandwidth resources for each mobile user to access the coding block. The edge action is expressed as ; 。 8. The method for optimizing drone-assisted edge storage services based on erasure codes according to claim 7 is characterized in that: The hierarchical deep reinforcement learning algorithm using mobility enhancement solves the Markov decision process based on the trajectory prediction results output by the trajectory prediction model, and obtains the optimal solution for simultaneously optimizing two sub-problems, including: The DDQN algorithm is used to solve the data encoding and placement sub-problem based on erasure codes. The DDQN algorithm contains two DNNs, namely the Q current network for parameter training and the Q target network for forward propagation to generate the target Q value. The Q value update function is: ; in, is the discount factor; The loss function of the DDQN algorithm is: ; In order to improve the exploration efficiency of DDQN, the UAV agent has The probability of choosing a random action; Since edge agents need to collect complex global information to make decisions, the PPO algorithm based on the actor-critic framework is used to solve the content delivery and resource allocation sub-problems; For actor networks, the probability ratio is To quantify the changes before and after the strategy update under the same state and action, It is expressed by the formula: ; On the other hand, temporal difference residual is used to calculate the advantage function , evaluate the actual return and expected return of the action selected under the current strategy, It is expressed by the formula: ; ; in, is the state value function; At the same time, in order to enhance the advantage function To improve the stability and efficiency of calculation, a generalized advantage estimate is introduced in the network, and the weight parameter is introduced To calculate the advantage function more smoothly, the calculated smooth advantage function It is expressed by the formula: ; In order to prevent excessive policy updates from causing unstable training, PPO uses a clipping function to ensure the range of updates. The clipping function is expressed by the formula: ; in, represents the clipping function; Based on this, the goal of the actor network is expressed as: ; in, represents the policy entropy that encourages exploration under the current policy, represents the weight of the policy entropy, Represents the goal of the actor network; The critic network is used to evaluate the expected reward in a certain state. As input, and output the relevant value function , updates itself by reducing the difference between the predicted value function and the calculated target value. The target of the critic network is expressed as: ; in, represents the target value of GAE, which is obtained by discounting the reward and The advantage estimate is calculated by combining Represents the target of the critic network.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a drone-assisted edge storage service optimization method based on erasure codes as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a method for optimizing drone-assisted edge storage services based on erasure codes as described in any one of claims 1 to 8.