Deep reinforcement learning-based common sensing calculation integration method and system in unmanned aerial vehicle edge calculation

By introducing a synesthesia computing integrated method with deep reinforcement learning in the UAV edge computing system, problems such as unreasonable resource allocation and excessive energy consumption under the multi-UAV collaborative architecture are solved, efficient system resource scheduling and energy efficiency improvement are achieved, and are suitable for complex and changeable aerial edge computing scenarios.

CN120547631APending Publication Date: 2025-08-26FUDAN UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510678986.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing UAV edge computing systems lack inter-system coordination mechanisms under the multi-UAV collaboration architecture, and there are bottlenecks in computing power, energy supply and service radius. Traditional optimization methods are not responding in a timely manner, have high computational complexity, and have poor adaptability in dynamic environments, making it difficult to achieve real-time and efficient system scheduling, and lack multi-dimensional resource optimization solutions.

Method used

The synesthesia computing integrated method based on deep reinforcement learning is adopted, and a multi-UAV collaborative system model is built and distributed DRL agent design is combined with A3C algorithm to realize the joint optimization of bandwidth allocation, power control and computing capabilities. A two-level computing architecture and asynchronous update mechanism are adopted to reduce communication overhead, improve system resource utilization and dynamic environment adaptability.

Benefits of technology

It significantly improves the service success rate and energy efficiency ratio of multi-UAV collaborative systems, realizes efficient resource scheduling in dynamic environments, and has good scalability and engineering application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547631A_ABST
    Figure CN120547631A_ABST
Patent Text Reader

Abstract

The invention provides a deep reinforcement learning-based method and a deep reinforcement learning-based system for integrating communication, sensing and calculation in unmanned aerial vehicle edge calculation, and the method comprises the steps: S1, constructing a system model, including constructing a network model which comprises a central unmanned aerial vehicle and a plurality of auxiliary unmanned aerial vehicles; s2, constructing a service process model; s3, constructing a sensing process model; s4, constructing a communication process model; s5, constructing a calculation process model; s6, constructing an energy consumption model; s7, constructing a joint optimization problem based on the network model, the service process model, the sensing process model, the communication process model, the calculation process model and the energy consumption model; s8, based on the joint optimization problem, constructing a Markov decision process conversion model; and S9, according to the Markov decision process conversion model, constructing a network architecture based on an A3C algorithm, and carrying out distributed DRL agent design and training, so that the affiliated unmanned aerial vehicle and the central unmanned aerial vehicle directly generate a resource allocation strategy based on a local state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drone edge computing technology, and in particular to a method and system for integrating synaesthesia and computing based on deep reinforcement learning in drone edge computing. Background Art

[0002] Against the backdrop of the continuous evolution of next-generation 6G network technology, Integrated Sensing, Communication and Computation (ISCC), which integrates perception, communication and computing, is becoming a key supporting technology. By integrating environmental perception, wireless communication and edge computing capabilities, this technology can provide low-latency, high-precision, and high-reliability service support for intelligent applications in complex mobile scenarios. At the same time, Mobile Edge Computing (MEC) has gradually expanded from ground base stations to aerial platforms, forming a new architecture for aerial edge computing. By deploying sensors and computing equipment on unmanned aerial vehicles (UAVs), the system can achieve rapid deployment and flexible response, filling the coverage gaps caused by obstacles, equipment failures, or resource constraints in the ground network.

[0003] Research has made some progress in the areas of Integrated Sensing and Computation (ISAC) and MEC, but this research primarily focuses on single UAV architectures. In these studies, UAVs typically serve as independent nodes providing services to end users, lacking inter-system coordination mechanisms. In practical applications, however, single UAVs face significant bottlenecks in computing power, energy supply, and service radius, making it difficult to meet the service demands of large-scale, high-density, and multi-task scenarios. Furthermore, traditional optimization methods based on mathematical modeling, such as convex optimization or heuristic algorithms, suffer from slow response times, high computational complexity, and poor adaptability in dynamic environments, making it difficult to achieve real-time and efficient system scheduling. To address these issues, recent attempts have combined deep reinforcement learning (DRL) with resource management. However, most focus on optimizing static tasks or single-dimensional resources (such as power or bandwidth). A comprehensive solution for jointly optimizing perception, communication, and computing resources in a multi-UAV collaborative architecture has yet to be established. Moreover, in a distributed environment, how to design an efficient architecture with decoupled training and execution, and how to reduce communication overhead and improve convergence speed during training are still core issues that need to be addressed urgently. Summary of the Invention

[0004] The present invention is made to solve the above-mentioned problems, and its purpose is to provide an integrated method and system for synaesthesia and computing based on deep reinforcement learning in drone edge computing.

[0005] The present invention provides an integrated method of synaesthesia and computing based on deep reinforcement learning in UAV edge computing, which has the following characteristics and specifically includes the following steps: S1, constructing a system model, including constructing a network model, the network includes a central UAV and multiple auxiliary UAVs; S2, constructing a service process model, the entire service cycle is divided into time slices of fixed length Δ, N = {1, 2, ..., N}, the state in each time slice is approximately stable, in each time slice, the auxiliary UAV performs perception, and uploads the pre-processed data to the central UAV through a communication link for fusion analysis and task decision-making; S3, constructing a perception process model, the auxiliary UAV performs downlink perception of targets in the area, collects echo signals, converts them into original radar data, and the echo signals meet the minimum mutual information threshold S4: Build a communication process model. The attached UAVs transmit the pre-processed raw radar data to the central UAV via an air-to-air communication link, and calculate the data transmission delay. S5: Build a computation process model. A two-level computation architecture is used, including local computation of attached UAVs and centralized computation of the central UAV. S6: Build an energy consumption model. The total energy consumption of each attached UAV is the sum of the energy consumption of perception, communication, and local computation:

[0006]

[0007] The total energy consumption of the system is:

[0008]

[0009] S7, based on the network model, service process model, perception process model, communication process model, computing process model and energy consumption model, optimizes the overall system performance within multiple time slices and constructs a joint optimization problem; S8, based on the joint optimization problem, constructs a Markov decision process transformation model to simulate the state, action and reward of the attached drones; S9, based on the Markov decision process transformation model, builds a network architecture based on the A3C algorithm, and designs and trains distributed DRL intelligent agents, so that the attached drones and central drones can directly generate resource allocation strategies based on local states.

[0010] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of the drone provided by the present invention, it can also have the following characteristics: wherein, step S1 specifically includes the following sub-steps: S1-1, network modeling, the functions of the central drone include: deploying edge servers, performing local perception, data preprocessing and communication tasks, the functions of the auxiliary drone include: serving as a central node, receiving the pre-processed data of the central drone and performing global data fusion and analysis, expressed as M = {1, 2, ..., M}, the auxiliary drone and the ground terminal K = {1, 2, ..., K} use perception beams to perform environmental sampling, and the central drone and the auxiliary drone use communication beams to transmit perception data, and the two use orthogonal frequency division multiplexing to avoid interference; S1-2, space constraints and flight safety, the distance between the central drone and the auxiliary drone meets Ensure that the communication link is stable, represents the three-dimensional position coordinates of the central drone o, The three-dimensional position coordinates of the attached drone m, R s and R e Indicates the minimum and maximum distances that the distance between attached drones must meet To avoid collisions, the auxiliary drones and the central drone are deployed at fixed altitudes, and their positions remain unchanged during the service time.

[0011] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones provided by the present invention, it can also have the following characteristics: wherein, step S3 specifically includes the following sub-steps: S3-1, the attached drone performs downlink perception of the target in the area through a directional radar beam, collects the echo signal, and converts it into raw radar data; S3-2, the amount of radar data is determined by multiple parameters, including data redundancy Radar beam switching frequency ν m , quantized angle number N θ , sampling frequency f s Number of bits per sample etc., can be expressed as: S3-3, in each time slice, the attached UAVs are connected according to the dynamic correlation variable χ mk [n]∈{0,1} determines whether the target vehicle k is perceived; S3-4, the effectiveness of the perceived echo is evaluated by the signal-to-noise ratio and mutual information, which must meet the minimum mutual information threshold

[0012]

[0013] in, represents the bandwidth, satisfying the following constraints, where represents the communication bandwidth, α m [n] represents the bandwidth ratio,

[0014]

[0015] Γ mk [n] represents the SNR of the sensing link from the attached UAV m to the terminal k, and the sensing interference between multiple UAVs needs to be considered. N0 represents the noise power. represents the perceived power and satisfies the following constraints, where represents the communication power, β m [n] is the power ratio,

[0016]

[0017] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of the UAV provided by the present invention, it can also have the following characteristics: wherein, step S4 specifically includes the following sub-steps: S4-1, the auxiliary UAV transmits the pre-processed radar data to the central UAV through the air-to-air communication link uplink, and the link model adopts the logarithmic distance path loss model; S4-2, calculates the data transmission delay as

[0018]

[0019] Among them, PL mo [n] represents the logarithmic distance path loss, N0 represents the signal-to-noise power, Indicates the communication signal transmission power, represents the communication bandwidth, Describes the output ratio of the attached UAV’s pre-processed radar perception data.

[0020] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of the UAV provided by the present invention, it can also have the following characteristics: wherein, step S5 specifically includes the following sub-steps: S5-1, local computing of the attached UAV: ​​mainly completing pre-processing operations such as clutter removal and target feature extraction; S5-2, local computing delay and energy consumption are respectively

[0021]

[0022]

[0023] in, represents the local computing power of the attached drone, ε m [n] is the processing complexity per bit (cycle / bit), κ m is the energy consumption coefficient; S5-3, central UAV center calculation: complete multi-UAV data fusion, environment modeling, and decision strategy generation; S5-4, central UAV calculation frequency is The delay and energy consumption are:

[0024]

[0025] S5-5, the total computation delay should meet the constraint, that is, to ensure that a complete perception-transmission-fusion service process is completed within each time slice.

[0026]

[0027] The integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones provided by the present invention may also have the following features: wherein step S7 specifically includes the following sub-steps: Joint optimization problem definition: Based on the aforementioned service process modeling, the present invention proposes to optimize the overall system performance within multiple time slices, and achieve dual guarantees of ISCC service quality and resource efficiency by maximizing the service success rate and minimizing the system energy consumption, and construct the following multi-objective joint optimization problem:

[0028]

[0029] Where B, P, and F represent the decision sets for bandwidth allocation, power control, and computing power control, respectively. Φ[n] represents the ISCC service success rate of the terminal. If the service process meets all constraints, it succeeds, otherwise it fails. The optimization problem is subject to the UAV's safe flight distance, variable values, upper and lower bounds of perception-communication-computing capabilities, and service delay constraints. The joint optimization problem is a time-dependent non-convex optimization problem with strong coupling between variables and high dimensionality. Traditional methods are difficult to solve efficiently. Therefore, reinforcement learning methods are used to achieve online policy learning.

[0030] The integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones provided by the present invention may also have the following features: wherein step S8 specifically includes the following sub-steps: S8-1, state: including the spatial relationship between each attached drone and the target terminal, the amount of perception data, the computing density, etc.,

[0031]

[0032] S8-2, Action: Bandwidth allocation ratio α for each SU m [n], power ratio β m [n], Local computing frequency ratio CU computing frequency

[0033]

[0034] S8-3, Reward: Consider the weighted combination of service completion incentives, failure penalties and energy consumption costs, where I m [n]∈{0,1} indicates whether the service of the attached drone is successful,

[0035]

[0036] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones provided by the present invention, it can also have the following characteristics: wherein, step S9 specifically includes the following sub-steps: S9-1 constructs a network architecture based on the A3C algorithm: each affiliated drone is an independent intelligent agent interacting with the local environment; the central drone is the global control center, asynchronously receives gradients, and uniformly updates the policy network; S9-2, the Critic network uses the least squares target driven by time difference error to estimate the state value function; S9-3, the Actor network uses policy gradients and advantage functions to output strategies, and in order to enhance exploratory power, an entropy term is introduced to calculate the policy gradient; S9-4, the working node uploads the local gradient every fixed number of steps, and the central drone performs asynchronous averaging to update the global network parameters until the algorithm converges or reaches the specified number of iterations to obtain the optimal policy network parameters; S9-5, after training is completed, only the Actor policy network is deployed on each drone, without the need to share parameters, reducing communication overhead; S9-6, all drones directly generate resource allocation strategies based on local states, with strong real-time and scalability.

[0037] In the integrated method of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones provided by the present invention, it can also have such features and also include: S10, experimental verification, specifically including the following sub-steps: S10-1, constructing a simulation scene based on an urban block model in a MATLAB environment, simulating K=15 terminals in an area of ​​1000m×1000m moving according to Poisson distribution and uniform motion mode; deep reinforcement learning adopts the A3C framework, the network structure is a multi-layer perceptron, parallel computing and GPU acceleration are enabled during training, and the comparison algorithms include proximal policy optimization, REINFORCE, deep deterministic policy gradient, particle swarm optimization, Greedy and Random strategies ; S10-2, the algorithm can converge stably during the training process, and the convergence speed is better than the traditional DRL method. Compared with REINFORCE, which has high variance, proximal policy optimization easily falls into local optimality, and deep deterministic policy gradient instability, MCAI significantly accelerates the policy learning process through the asynchronous update mechanism, and reaches the stable optimal value within 600 rounds on average during the training phase; S10-3, the proposed algorithm can achieve an average task completion rate of 99.20%, which is better than Greedy, REINFORCE and other strategies, and has stronger stability. While ensuring a high success rate, the algorithm achieves the lowest system energy consumption, and the energy efficiency is improved by 80.12% compared with Greedy, with the best energy efficiency.

[0038] The present invention also provides an integrated system of synaesthesia and computing based on deep reinforcement learning in UAV edge computing, including: a system modeling module, including building a network model, the network includes a central UAV and multiple auxiliary UAVs; building a service process model, the entire service cycle is divided into time slices of fixed length Δ, N = {1, 2, ..., N}, the state in each time slice is approximately stable, in each time slice, the auxiliary UAV performs perception, and uploads the pre-processed data to the central UAV through a communication link for fusion analysis and task decision-making; a perception process modeling module, the auxiliary UAV performs downlink perception of targets in the area, collects echo signals, converts them into original radar data, and the echo signals meet the minimum mutual information threshold In the communication process modeling module, the attached UAVs transmit the pre-processed raw radar data to the central UAV via the air-to-air communication link and calculate the data transmission delay. In the computation process modeling module, a two-level computation architecture is adopted, including local computation of the attached UAVs and central computation of the central UAV. In the energy consumption modeling module, the total energy consumption of each attached UAV is the sum of the perception, communication and local computation energy consumption:

[0039]

[0040] The total energy consumption of the system is:

[0041]

[0042] The joint optimization problem construction module optimizes the overall system performance within multiple time slices based on the network model, service process model, perception process model, communication process model, computing process model and energy consumption model to construct a joint optimization problem; the Markov decision process transformation module constructs a Markov decision process transformation model; the distributed DRL agent design and training module constructs a network architecture based on the A3C algorithm to design and train distributed DRL agents.

[0043] Functions and effects of the invention

[0044] According to the integrated method and system of synaesthesia and computing based on deep reinforcement learning in UAV edge computing involved in the present invention, in the collaborative multi-UAV scenario, the present invention addresses the problems of unreasonable resource allocation, low service efficiency, and excessive energy consumption in the multi-UAV collaborative ISCC system. A distributed deep reinforcement learning method based on the Asynchronous Advantage Actor-Critic (A3C) algorithm is proposed to construct a collaborative intelligent control mechanism of "central UAV + multiple auxiliary UAVs", jointly optimize the system's bandwidth allocation, power control and computing power scheduling strategy, and significantly reduce system energy consumption, significantly improve the system's responsiveness, adaptability and energy efficiency in dynamic environments while ensuring the service success rate. The method has excellent adaptability to dynamic environments and is suitable for complex and changeable aerial edge computing scenarios.

[0045] Experimental results show that compared with existing strategies, the method of the present invention can achieve an average task completion rate of 99.20% and an energy efficiency improvement of up to 80.12%. It has good scalability and deployability, is suitable for a variety of complex application environments, and has significant engineering application value and innovation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Schematic diagram of the composition of the network model in an embodiment of the present invention;

[0047] Figure 2 2 is a schematic diagram comparing the convergence performance of algorithms in an embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of the task success rate results of the algorithm in an embodiment of the present invention; and

[0049] Figure 4 Schematic diagram of energy saving efficiency of the algorithm in the embodiment of the present invention. DETAILED DESCRIPTION

[0050] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the integrated method and system of synaesthesia and computing based on deep reinforcement learning in the edge computing of drones of the present invention.

[0051] To overcome the shortcomings of traditional single-UAV systems in resource scheduling efficiency, energy consumption control, and dynamic environmental adaptability, this paper proposes a multi-UAV collaborative integrated sensing and computing service optimization method. This method combines the collaborative sensing capabilities of multiple UAVs with the advantages of a centralized computing architecture. By introducing distributed deep reinforcement learning, it achieves joint intelligent scheduling of communication bandwidth, power allocation, and computing power. The multi-UAV collaborative strategy effectively improves system resource utilization, reduces service energy consumption, and has excellent dynamic environmental adaptability, making it suitable for complex and changing aerial edge computing scenarios.

[0052] In the present invention, the integrated method of synaesthesia and computing based on deep reinforcement learning in drone edge computing specifically includes the following steps:

[0053] S1, builds the system model, including building a network model. The network includes a central UAV and multiple subsidiary UAVs.

[0054] Step S1 specifically includes the following sub-steps:

[0055] S1-1, Network modeling.

[0056] Figure 1 It is a schematic diagram of the composition of the network model in an embodiment of the present invention.

[0057] like Figure 1 As shown, the functions of the central drone include: deploying edge servers, performing local perception (radar detection), data preprocessing (such as clutter elimination) and communication tasks.

[0058] The functions of the subsidiary UAVs include: acting as a central node, receiving pre-processed data from the central UAV and performing global data fusion and analysis, which is represented as M = {1, 2, ..., M}.

[0059] Perception beams are used between the auxiliary UAVs and the ground terminals K = {1, 2, ..., K} to sample the environment, and communication beams are used between the central UAV and the auxiliary UAVs to transmit perception data. Orthogonal frequency division multiplexing is used between the two to avoid interference.

[0060] S1-2, Space Constraints and Flight Safety.

[0061] The distance between the central drone and the auxiliary drones meets Ensure that the communication link is stable, represents the three-dimensional position coordinates of the central drone o, The three-dimensional position coordinates of the attached drone m, R s and R e Indicates the minimum and maximum distance.

[0062] The distance between attached drones must meet Avoid collisions.

[0063] The auxiliary drones and central drone are deployed at fixed altitudes, and their positions remain unchanged during the service time.

[0064] S2, builds a service process model. The entire service cycle is divided into fixed-length time slices Δ, N = {1, 2, ..., N}. The state in each time slice is approximately stable. In each time slice, the auxiliary UAV performs perception and uploads the pre-processed data to the central UAV through the communication link for fusion analysis and task decision-making.

[0065] The entire service cycle is divided into time slices of fixed length Δ, N = {1, 2, ..., N}, and the state is approximately stable within each time slice.

[0066] During each time slice, the attached drones perform perception (radar data collection) and upload the pre-processed data via a communication link to the central drone for fusion analysis and mission decision-making. The entire mission service process includes: radar perception → local pre-processing → data uplink → central fusion processing, and the entire process is subject to service latency and energy consumption constraints.

[0067] S3, builds a perception process model, and the attached UAV performs downlink perception of the target in the area, collects the echo signal, and converts it into raw radar data. The echo signal meets the minimum mutual information threshold.

[0068] Step S3 specifically includes the following sub-steps:

[0069] S3-1, the attached UAV performs downlink perception of targets in the area through directional radar beams, collects echo signals, and converts them into raw radar data.

[0070] S3-2, the amount of radar data is determined by multiple parameters, including data redundancy Radar beam switching frequency ν m , quantized angle number N θ , sampling frequency f s Number of bits per sample etc., can be expressed as:

[0071]

[0072] S3-3, in each time slice, the attached UAVs are connected according to the dynamic correlation variable χ mk [n]∈{0,1} determines whether the target vehicle k is perceived.

[0073] S3-4, the effectiveness of the perceived echo is evaluated using signal-to-noise ratio and mutual information, which must meet the minimum mutual information threshold.

[0074]

[0075] in, represents the bandwidth, satisfying the following constraints, where represents the communication bandwidth, α m [n] represents the bandwidth ratio,

[0076]

[0077] Γ mk [n] represents the SNR of the sensing link from the attached UAV m to the terminal k, and the sensing interference between multiple UAVs needs to be considered. N0 represents the noise power. represents the perceived power and satisfies the following constraints, where represents the communication power, β m [n] is the power ratio,

[0078]

[0079] S4, builds a communication process model, where the subsidiary UAVs uplink the pre-processed raw radar data to the central UAV through the air-to-air communication link, and calculates the data transmission delay.

[0080] Step S4 specifically includes the following sub-steps:

[0081] In S4-1, the auxiliary UAV transmits the pre-processed radar data to the central UAV via an air-to-air communication link. The link model adopts the logarithmic distance path loss model.

[0082] S4-2, calculate the data transmission delay as

[0083]

[0084] Among them, PL mo [n] represents the logarithmic distance path loss, N0 represents the signal-to-noise power, Indicates the communication signal transmission power, represents the communication bandwidth, Describes the output ratio of the attached UAV’s pre-processed radar perception data.

[0085] S5,constructs the computing process model, which adopts a two-level computing architecture,,including local computing of the attached UAVs and central computing of the,central UAV.

[0086] Step S5 specifically includes the following sub-steps:

[0087] S5-1, local computing of the attached UAV: ​​mainly completes pre-processing operations such as clutter removal and target feature extraction.

[0088] S5-2, local computing delay and energy consumption are

[0089]

[0090] in, represents the local computing power of the attached drone, ε m [n] is the processing complexity per bit (cycle / bit), κ m is the energy consumption coefficient.

[0091] S5-3, Central UAV Center Computing: Complete multi-UAV data fusion, environment modeling, and decision strategy generation.

[0092] S5-4, the central drone calculates the frequency The delay and energy consumption are:

[0093]

[0094] S5-5, the total computation delay should meet the constraint, that is, to ensure that a complete perception-transmission-fusion service process is completed within each time slice.

[0095]

[0096] S6, build an energy consumption model. The total energy consumption of each attached UAV is the sum of the perception, communication and local computing energy consumption:

[0097]

[0098] The total energy consumption of the system is:

[0099]

[0100] S7, based on the network model, service process model, perception process model, communication process model, computing process model and energy consumption model, optimizes the overall system performance in multiple time slices and constructs a joint optimization problem.

[0101] Step S7 specifically includes the following sub-steps:

[0102] Joint optimization problem definition: Based on the aforementioned service process modeling, this paper proposes to optimize the overall system performance within multiple time slices, achieving dual guarantees of ISCC service quality and resource efficiency by maximizing service success rate and minimizing system energy consumption. The following multi-objective joint optimization problem is constructed:

[0103]

[0104] Where B, P, and F represent the decision sets for bandwidth allocation, power control, and computing power control, respectively. Φ[n] represents the ISCC service success rate of the terminal. If the service process meets all constraints, it succeeds, otherwise it fails. The optimization problem is subject to the UAV's safe flight distance, variable values, upper and lower bounds of perception-communication-computing capabilities, and service delay constraints.

[0105] The joint optimization problem is a time-dependent non-convex optimization problem with strong coupling between variables and high dimensionality. Traditional methods are difficult to solve efficiently, so reinforcement learning methods are used to achieve online strategy learning.

[0106] S8, based on the joint optimization problem, builds a Markov decision process transformation model to simulate the state, action and reward of the attached UAV.

[0107] Step S8 specifically includes the following sub-steps:

[0108] S8-1, Status: including the spatial relationship between each attached UAV and the target terminal, the amount of perception data, the calculation density, etc.

[0109]

[0110] S8-2, Action: Bandwidth allocation ratio α for each SU m [n], power ratio β m [n], Local computing frequency ratio CU computing frequency

[0111]

[0112] S8-3, Reward: Consider the weighted combination of service completion incentives, failure penalties and energy consumption costs, where I m [n]∈{0,1} indicates whether the service of the attached drone is successful,

[0113]

[0114] S9, based on the Markov decision process transformation model and the A3C algorithm, builds a network architecture and conducts distributed DRL agent design and training, so that the subsidiary UAVs and central UAVs can directly generate resource allocation strategies based on local states.

[0115] Step S9 specifically includes the following sub-steps:

[0116] S9-1 builds a network architecture based on the A3C algorithm: each subsidiary drone is an independent intelligent agent interacting with the local environment; the central drone is the global control center, asynchronously receiving gradients and uniformly updating the strategy network.

[0117] S9-2, the critic network uses a least squares objective driven by the time difference error to estimate the state value function.

[0118] S9-3, the Actor network uses policy gradient and advantage function to output the policy. To enhance exploration, the entropy term is introduced to calculate the policy gradient.

[0119] In S9-4, the working nodes upload local gradients every fixed number of steps, and the central UAV performs asynchronous averaging (FedAvg) to update the global network parameters until the algorithm converges or reaches the specified number of iterations to obtain the optimal policy network parameters.

[0120] S9-5, after training is completed, only the Actor strategy network is deployed on each drone, without the need to share parameters, reducing communication overhead.

[0121] In S9-6, all drones directly generate resource allocation strategies based on local status, which has strong real-time and scalability.

[0122] S10, experimental verification.

[0123] Step S10 specifically includes the following sub-steps:

[0124] S10-1: A simulation scenario based on a city block model was constructed in the MATLAB environment, simulating K = 15 terminals moving in a 1000m × 1000m area according to a Poisson distribution and uniform velocity. Deep reinforcement learning was performed using the A3C framework, a multilayer perceptron network, and parallel computing and GPU acceleration were enabled during training. Comparison algorithms included Proximal Policy Optimization (PPO), REINFORCE, Deep Deterministic Policy Gradient (DDPG), Particle Swarm Optimization (PSO), Greedy, and Random strategies.

[0125] Figure 2 3 is a schematic diagram comparing the convergence performance of the algorithms in the embodiments of the present invention.

[0126] S10-2, such as Figure 2 As shown in the figure, the algorithm can converge stably during training, and the convergence speed is better than the traditional DRL method. Compared with REINFORCE, which has high variance, proximal policy optimization easily falls into local optimality, and deep deterministic policy gradient instability, MCAI significantly accelerates the policy learning process through the asynchronous update mechanism, and reaches the stable optimal value within 600 rounds on average during the training phase.

[0127] Figure 32 is a schematic diagram of the task success rate results of the algorithm in the embodiment of the present invention. Figure 4 Schematic diagram of energy saving efficiency of the algorithm in the embodiment of the present invention.

[0128] S10-3, such as Figure 3-4 As shown in the figure, the proposed algorithm can achieve an average task completion rate of 99.20%, which is better than strategies such as Greedy and REINFORCE, and has stronger stability. While ensuring a high success rate, the algorithm achieves the lowest system energy consumption, which is 80.12% higher than Greedy, and has the best energy efficiency (energy efficiency is defined as (success rate*100) / energy consumption).

[0129] The present invention also provides an integrated system of synaesthesia and computing based on deep reinforcement learning in edge computing of drones, comprising:

[0130] The system modeling module is performed according to the method of the above step S1, including building a network model, where the network includes a central UAV and multiple auxiliary UAVs.

[0131] The service process model is constructed according to the method of step S2 above. The entire service cycle is divided into fixed-length time slices Δ, N = {1, 2, ..., N}. The state in each time slice is approximately stable. In each time slice, the auxiliary UAV performs perception and uploads the pre-processed data to the central UAV through the communication link for fusion analysis and task decision-making.

[0132] The perception process modeling module is carried out according to the method of step S3 above. The attached UAV performs downlink perception of the target in the area, collects the echo signal, and converts it into raw radar data. The echo signal meets the minimum mutual information threshold.

[0133] The communication process modeling module follows the method of step S4 above. The auxiliary UAV transmits the pre-processed raw radar data to the central UAV via the air-to-air communication link and calculates the data transmission delay.

[0134] The computational process modeling module is performed according to the method of step S5 above, and adopts a two-level computational architecture, including local computation of the subsidiary drones and central computation of the central drone.

[0135] The energy consumption modeling module is performed according to the method of step S6 above. The total energy consumption of each attached UAV is the sum of the energy consumption of perception, communication and local computing:

[0136]

[0137] The total energy consumption of the system is:

[0138]

[0139] The joint optimization problem construction module is carried out according to the method of step S7 above, and based on the network model, service process model, perception process model, communication process model, calculation process model and energy consumption model, the overall performance of the system is optimized within multiple time slices to construct a joint optimization problem.

[0140] The Markov decision process transformation module is implemented according to the method of step S8 above to construct a Markov decision process transformation model.

[0141] The distributed DRL agent design and training module is performed according to the method of step S9 above, and a network architecture is constructed based on the A3C algorithm to perform distributed DRL agent design and training.

[0142] Those skilled in the art will appreciate that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for integrating synaesthesia and computing based on deep reinforcement learning in UAV edge computing, characterized by: The specific steps include: S1, building a system model, including building a network model, wherein the network includes a central UAV and multiple subordinate UAVs; S2, constructing a service process model. The entire service cycle is divided into fixed-length time slices Δ, N = {1, 2, ..., N}. The state in each time slice is approximately stable. In each time slice, the auxiliary UAV performs perception and uploads the pre-processed data to the central UAV via a communication link for fusion analysis and task decision-making. S3, constructing a perception process model, the auxiliary UAV performs downlink perception of the target in the area, collects the echo signal, and converts it into raw radar data. The echo signal meets the minimum mutual information threshold. S4, constructing a communication process model, wherein the auxiliary UAV transmits the pre-processed raw radar data to the central UAV via an air-to-air communication link, and calculates the data transmission delay; S5, constructing a computing process model, adopting a two-level computing architecture, including local computing of the auxiliary drones and central computing of the central drone; S6, build an energy consumption model. The total energy consumption of each attached UAV is the sum of the perception, communication and local computing energy consumption: The total energy consumption of the system is: S7, based on the network model, the service process model, the perception process model, the communication process model, the computing process model, and the energy consumption model, optimizing the overall system performance within multiple time slices to construct a joint optimization problem; S8, constructing a Markov decision process transformation model based on the joint optimization problem to simulate the state, action, and reward of the attached UAV; S9, according to the Markov decision process transformation model, based on the A3C algorithm, a network architecture is constructed, and distributed DRL intelligent agent design and training are performed, so that the subsidiary UAV and the central UAV directly generate resource allocation strategies based on local states.

2. The method for integrating synaesthesia and computing based on deep reinforcement learning in UAV edge computing according to claim 1, characterized in that: in, The step S1 specifically includes the following sub-steps: S1-1, Network Modeling, The functions of the central drone include: deploying edge servers, performing local perception, data preprocessing and communication tasks, The functions of the auxiliary UAV include: serving as a central node, receiving pre-processed data from the central UAV and performing global data fusion and analysis, denoted as M = {1, 2, ..., M}, The auxiliary UAV and the ground terminal K = {1, 2, ..., K} use a sensing beam to perform environmental sampling, and the central UAV and the auxiliary UAV use a communication beam to transmit sensing data, and the two use orthogonal frequency division multiplexing to avoid interference; S1-2, space constraints and flight safety, the distance between the central UAV and the auxiliary UAV meets Ensure that the communication link is stable, represents the three-dimensional position coordinates of the central drone o, The three-dimensional position coordinates of the attached drone m, R s and R e Indicates the minimum and maximum distances, The distance between the attached drones must meet Avoid collisions, The auxiliary drones and the central drone are deployed at fixed altitudes, and their positions remain unchanged during the service time.

3. The method for integrating synaesthesia and computation based on deep reinforcement learning in UAV edge computing according to claim 1, characterized in that: in, The step S3 specifically includes the following sub-steps: S3-1, the auxiliary UAV performs downlink sensing of targets in the area through a directional radar beam, collects echo signals, and converts them into raw radar data; S3-2, the amount of radar data is determined by multiple parameters, including data redundancy Radar beam switching frequency ν m , quantized angle number N θ , sampling frequency f s and the number of bits per sample θ m etc., can be expressed as: S3-3, in each time slice, the attached UAVs are connected according to the dynamic correlation variable χ mk [n]∈{0,1} determines whether the target vehicle k is perceived; S3-4, the effectiveness of the perceived echo is evaluated using signal-to-noise ratio and mutual information, which must meet the minimum mutual information threshold. in, represents the bandwidth, satisfying the following constraints, where represents the communication bandwidth, α m [n] represents the bandwidth ratio, Γ mk [n] represents the SNR of the sensing link from the attached UAV m to the terminal k, and the sensing interference between multiple UAVs needs to be considered. N0 represents the noise power. represents the perceived power and satisfies the following constraints, where represents the communication power, β m [n] is the power ratio, 4. The integrated method of synaesthesia and computing based on deep reinforcement learning in drone edge computing according to claim 1, Its characteristics are: Wherein, the step S4 specifically includes the following sub-steps: S4-1, the auxiliary UAV transmits the pre-processed radar data to the central UAV via an air-to-air communication link. The link model adopts the logarithmic distance path loss model; S4-2, calculate the data transmission delay as Among them, PL mo [n] represents the logarithmic distance path loss, N0 represents the signal-to-noise power, Indicates the communication signal transmission power, represents the communication bandwidth, Describes the output ratio of the attached UAV’s pre-processed radar perception data.

5. The method for integrating synaesthesia and computing based on deep reinforcement learning in edge computing of drones according to claim 1, characterized in that: in, The step S5 specifically includes the following sub-steps: S5-1, local computing of the attached UAV: ​​mainly performs pre-processing operations such as clutter removal and target feature extraction; S5-2, local computing delay and energy consumption are in, represents the local computing power of the attached drone, ε m [n] is the processing complexity per bit (cycle / bit), κ m is the energy consumption coefficient; S5-3, Central UAV Center Computing: Complete multi-UAV data fusion, environment modeling, and decision strategy generation; S5-4, the central drone calculates the frequency The delay and energy consumption are: S5-5, the total computation delay should meet the constraint, that is, to ensure that a complete perception-transmission-fusion service process is completed within each time slice.

6. The method for integrating synaesthesia and computing based on deep reinforcement learning in UAV edge computing according to claim 1, characterized in that: in, The step S7 specifically includes the following sub-steps: Joint optimization problem definition: Based on the aforementioned service process modeling, this paper proposes to optimize the overall system performance within multiple time slices, achieving dual guarantees of ISCC service quality and resource efficiency by maximizing service success rate and minimizing system energy consumption. The following multi-objective joint optimization problem is constructed: Where B, P, and F represent the decision sets for bandwidth allocation, power control, and computing power control, respectively. Φ[n] represents the ISCC service success rate of the terminal. If the service process meets all constraints, it succeeds, otherwise it fails. The optimization problem is subject to the UAV's safe flight distance, variable values, upper and lower bounds of perception-communication-computing capabilities, and service delay constraints. The joint optimization problem is a time-dependent non-convex optimization problem with strong coupling between variables and high dimensionality. Traditional methods are difficult to solve efficiently, so reinforcement learning methods are used to achieve online strategy learning.

7. The method for integrating synaesthesia and computing based on deep reinforcement learning in UAV edge computing according to claim 1, characterized in that: in, The step S8 specifically includes the following sub-steps: S8-1, Status: including the spatial relationship between each attached UAV and the target terminal, the amount of perception data, the calculation density, etc. S8-2, Action: Bandwidth allocation ratio α for each SU m [n], power ratio β m [n], Local computing frequency ratio CU computing frequency S8-3, Reward: Consider the weighted combination of service completion incentives, failure penalties and energy consumption costs, where I m [n]∈{0,1} indicates whether the service of the attached drone is successful, 8. The method for integrating synaesthesia and computing based on deep reinforcement learning in UAV edge computing according to claim 1, characterized in that: in, The step S9 specifically includes the following sub-steps: S9-1 builds a network architecture based on the A3C algorithm: each subordinate drone is an independent intelligent agent interacting with the local environment; the central drone is the global control center, asynchronously receiving gradients and uniformly updating the policy network; S9-2, the critic network uses the least squares objective driven by the time difference error to estimate the state value function; S9-3, the Actor network uses policy gradient and advantage function to output the policy. To enhance exploration, the entropy term is introduced to calculate the policy gradient. S9-4, the working node uploads the local gradient every fixed number of steps, and the central drone performs asynchronous averaging to update the global network parameters until the algorithm converges or reaches the specified number of iterations to obtain the optimal strategy network parameters; S9-5, after training is completed, only the Actor strategy network is deployed on each drone, without the need to share parameters, thus reducing communication overhead; In S9-6, all drones directly generate resource allocation strategies based on local status, which has strong real-time and scalability.

9. The method for integrating synaesthesia and computing based on deep reinforcement learning in edge computing of drones according to claim 1, characterized in that: Also includes: S10, experimental verification, specifically includes the following sub-steps: S10-1: A simulation scenario based on a city block model was constructed in the MATLAB environment, simulating K = 15 terminals moving in a 1000m × 1000m area according to a Poisson distribution and uniform motion pattern. Deep reinforcement learning was conducted using the A3C framework, with a multilayer perceptron network structure. Parallel computing and GPU acceleration were enabled during training. Comparison algorithms included proximal policy optimization, REINFORCE, deep deterministic policy gradient, particle swarm optimization, Greedy, and random strategies. S10-2: The algorithm converges stably during training, and its convergence speed is superior to traditional DRL methods. Compared with REINFORCE, which suffers from high variance, the proximal policy optimization easily falling into local optimality, and unstable deep deterministic policy gradients, MCAI significantly accelerates the policy learning process through an asynchronous update mechanism, reaching a stable optimal value within an average of 600 rounds during the training phase. For S10-3, the proposed algorithm can achieve an average task completion rate of 99.20%, which is better than strategies such as Greedy and REINFORCE, and has stronger stability. While ensuring a high success rate, the algorithm achieves the lowest system energy consumption, which is 80.12% higher than Greedy, and has the best energy efficiency.

10. A synaesthesia-computing integrated system based on deep reinforcement learning in drone edge computing, characterized by: include: A system modeling module includes building a network model, wherein the network includes a central UAV and multiple subordinate UAVs; Construct a service process model. The entire service cycle is divided into fixed-length time slices Δ, N = {1, 2, ..., N}. The state in each time slice is approximately stable. In each time slice, the auxiliary UAV performs perception and uploads the pre-processed data to the central UAV via a communication link for fusion analysis and task decision-making. Perception process modeling module, the auxiliary UAV performs downlink perception of targets in the area, collects echo signals, and converts them into raw radar data. The echo signals meet the minimum mutual information threshold. a communication process modeling module, wherein the auxiliary UAV transmits the pre-processed raw radar data uplink to the central UAV via an air-to-air communication link and calculates the data transmission delay; A computational process modeling module adopts a two-level computational architecture, including local computations of the auxiliary drones and central computations of the central drone; Energy consumption modeling module, the total energy consumption of each attached UAV is the sum of perception, communication and local computing energy consumption: The total energy consumption of the system is: A joint optimization problem construction module, which optimizes the overall system performance in multiple time slices based on the network model, service process model, perception process model, communication process model, computing process model and energy consumption model to construct a joint optimization problem; Markov decision process transformation module, building a Markov decision process transformation model; The distributed DRL agent design and training module builds a network architecture based on the A3C algorithm to perform distributed DRL agent design and training.

Citation Information

Cited By

  • 6G-based power grid communication method, equipment and system

    CN121442417A

  • Unmanned aerial vehicle ISCC joint resource scheduling method and system based on lightweight DRL

    CN121487007A

  • Heterogeneous unmanned aerial vehicle task flow edge calculation unloading method based on multi-agent deep reinforcement learning

    CN121636167A