Unmanned aerial vehicle auxiliary sensing calculation integration method and system based on deep reinforcement learning
Through deep reinforcement learning, optimize the integrated collaborative processing of drone clusters of perception, communication and computing, the inefficiency and delay problems caused by resource separation in drone systems are solved, and efficient and low-energy consumption intelligent decision-making and dynamic environmental adaptation are achieved, which are suitable for future 6G intelligent transportation and emergency rescue scenarios.
Patent Information
- Application Number
- CN202510678982.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-29
AI Technical Summary
In the existing drone-assisted mobile edge computing system, the separation of perception, communication and computing functions leads to low resource utilization, high system delay, heavy communication load, and lacks an efficient coordination mechanism for multiple drones in dynamic environments, making it difficult to meet the real-time and global optimal requirements in high-dynamic and high-density scenarios.
Adopting the integrated method of drone-assisted synesthesia computing based on deep reinforcement learning, we design drone group deployment, service cycle phase division, energy consumption calculation and optimization models, build Actor and Critic networks, optimize resource allocation and task scheduling, and realize air-interface fusion and intelligent scheduling of perception, communication and computing.
Significantly improve system resource utilization and service efficiency, improve dynamic environment adaptability, support multi-UAV collaboration and real-time task processing, and is suitable for ubiquitous intelligent services in complex environments.
Smart Images

Figure CN120567271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drone technology, and in particular to a drone-assisted synaesthesia computing integrated method and system based on deep reinforcement learning. Background Art
[0002] With the rapid development of the Internet of Things (IoT), the dense access and computing demands of IoT terminals continue to grow, posing significant challenges to traditional terrestrial communication and computing architectures. On the one hand, IoT terminals generally have limited computing power, requiring reliance on edge computing for offloading. On the other hand, ground base station deployment is limited, resulting in blind spots. This makes it difficult for existing communication infrastructure to guarantee service coverage and timeliness, particularly in mountainous areas, emergency rescue operations, and extreme environments. Furthermore, sensing, communication, and computing functions are typically designed and deployed separately in traditional networks, resulting in fragmented resource scheduling, inefficient spectrum utilization, and significant processing latency, making it difficult to meet the demands of high-density, low-latency, multi-task converged services.
[0003] In recent years, ISAC (Integrated Sensing and Communication) architectures, which integrate communication and perception, have been widely researched. By sharing spectrum and hardware resources, they improve spectrum utilization and simplify system design. However, existing ISAC solutions still fail to address the collaborative processing of distributed computing tasks. Especially in dynamic scenarios, the integration of communication and perception alone remains insufficient to meet the real-time and efficient computing needs of terminals. Therefore, further integrating computing functions into communication and perception processes, achieving air interface convergence of the three, has become a new research hotspot.
[0004] Furthermore, aerial platforms can serve as a powerful complement to ground-based infrastructure, particularly in complex environments. Unmanned aerial vehicles (UAVs) are ideal aerial edge nodes due to their maneuverability, rapid deployment, and strong line-of-sight communication capabilities. However, individual UAVs face physical bottlenecks in terms of sensing range, computing power, and energy capacity, making them unable to independently perform complex system tasks. Therefore, the coordinated scheduling of UAV swarms is crucial. By integrating sensing information, allocating computing tasks, and complementing energy between UAV groups, the overall service capabilities of the system can be significantly improved.
[0005] Current research on UAV-assisted Mobile Edge Computing (MEC) primarily focuses on UAV flight trajectory optimization, task scheduling, computation offloading, or communication relay functions. There is little systematic consideration of the collaborative optimization of multiple UAV resources under the fusion of perception, communication, and computation. Furthermore, there is a lack of joint scheduling mechanisms for multiple UAVs using dynamic time, computation, and spectrum resources, and insufficient consideration of the balance between energy constraints and Quality of Service (QoS) constraints. Furthermore, the highly dynamic nature of system states, the complexity of multi-objective optimization problems, and the high-dimensional, non-convex solution space make it difficult for traditional analytical modeling and static optimization algorithms to respond to large-scale deployments and sudden task requests in real time. Summary of the Invention
[0006] The present invention is designed to solve the problems of low resource utilization, high system latency, heavy communication load, etc. caused by the separate deployment of perception, communication and computing functions in existing networks. The purpose is to provide a drone-assisted telepathy and computing integration method and system based on deep reinforcement learning, which solves the problem that existing solutions are difficult to simultaneously take into account the real-time computing requests of terminals, the energy consumption constraints of aerial platforms and the adaptability to dynamic mission environments, and lack an efficient collaborative mechanism for multiple drones in resource allocation, task scheduling, perception fusion, etc. Especially in highly dynamic and high-density scenarios, traditional optimization methods are difficult to meet the real-time and global optimal requirements.
[0007] The present invention provides an integrated method for unmanned aerial vehicle (UAV)-assisted synaesthesia and computing based on deep reinforcement learning, which has the following characteristics and comprises the following steps: S1, designing a system architecture, including deploying UAVs and ground equipment; S2, designing a service cycle phase division and a frame structure of the UAV; S3, calculating a radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the UAV based on the phase division and the frame structure of the UAV's service cycle; S4, constructing an optimization model: constructing a joint optimization target based on the perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the UAV, minimizing each energy consumption and maximizing the service success rate; S5, constructing a reinforcement learning model based on the optimization model; S6, optimizing the reinforcement learning model: constructing an actor network and a critic network and optimizing the algorithm strategies therein to obtain a strategy model; S7, training the strategy model: training the strategy model by adopting a single agent, multi-agent or parallel computing method to obtain an optimal actor network; S8, controlling the decision-making of the UAV according to the optimal actor network.
[0008] In the UAV-assisted synaesthesia and computing integration method based on deep reinforcement learning provided by the present invention, it can also have the following features: wherein, step S1 includes the following sub-steps: S1-1, deployment of a UAV group, deploying a UAV group consisting of multiple UAVs U = {1, 2, ..., m}, |U| = m, the UAVs are equipped with radar sensing units, communication antenna arrays and edge computing servers, the UAVs are fixed at a preset height and hover, orthogonal frequency division multiple access and self-interference elimination technology are used to avoid communication interference within the group, and a safety protection radius is set to prevent collisions; S1-2, ground equipment deployment, the distribution of ground Internet of Things devices I = {1, 2, ..., n}, |I| = n changes dynamically, and tasks that need to be processed in real time are generated. in Indicates the data size of the computing task, Indicates the number of CPU cycles required to complete the calculation of each bit task.
[0009] In the UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention, it can also have the following characteristics: wherein, step S2 includes the following sub-steps: S2-1, each service period T is evenly divided into N time slots of the same length, each time slot length is Δ=T / N, and each time slot is further subdivided into two sub-stages; S2-2, stage 1 is the perception stage: the duration is α[t]Δ, and the UAV uses the radar system to complete environmental perception; S2-3, stage 2 is the communication-computing stage: the duration is (1-α[t])Δ, which is used for data unloading and computing processing, that is, the Internet of Things terminal completes the task upload, and the UAV completes the calculation.
[0010] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: wherein step S3 includes the following sub-steps: S3-1, UAV perception: each UAVj transmits radar waves to the ground during the perception sub-period and receives reflected signals, and embeds information bits in the radar beam sidelobes; S3-2, converts the radar echoes into radar data, including information such as the location of the IoT terminal:
[0011]
[0012] in Including data such as location, obstacle identification, signal measurement and target identification, j ≥1 is a constant related to the introduced data redundancy, ν j is the switching speed of the radar beam, N θ is the number of quantized angles, f s [t] is the sampling frequency, is the number of quantization bits per sampling; S3-3, radar estimated information rate calculation
[0013]
[0014] Where B represents bandwidth, p j [t] represents the perceived power of the UAV, represents the channel gain of the radar echo signal, d i,j [t] represents the distance between UAVj and IoT terminal i, Represents the reference channel gain at a distance of 1 meter. The average information estimated by the radar should not be less than the required radar detection data. Perception needs to meet the constraints
[0015]
[0016] where c i,j represents the association decision between UAVj and IoT terminal i;
[0017] S3-4, calculation of perceived energy consumption:
[0018]
[0019] S3-5, data communication: The IoT terminal task data size is The task upload rate is Then the task upload delay is:
[0020]
[0021] S3-6, calculate communication energy consumption:
[0022]
[0023] Among them, q i [t] represents the signal transmission power of the IoT terminal. In S3-7, the UAV receives the offload data and uses dynamic voltage and frequency adjustment technology to control resource allocation. The CPU calculation time is:
[0024]
[0025] Among them, ε i,j [t] represents the proportion of computing resources allocated by UAVj to IoT terminal i, Indicates the maximum computing capability of UAVj; S3-8, calculates the computing energy consumption:
[0026]
[0027] Among them, κ j Represents hardware-related constants; S3-9, calculates the total energy consumption of the drone:
[0028]
[0029] S3-10, calculate the total delay of the task service, which must meet the delay constraint
[0030] T i ser [t] = T i off [t]+T i comp [t]
[0031] T i ser [t]≤(1-α j [t]Δ).
[0032] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: wherein step S4 includes the following sub-steps: constructing a joint optimization target with the goal of maximizing the service success rate and minimizing energy consumption:
[0033]
[0034] Among them, α and p represent the time slot ratio and signal transmission power decision set of UAV respectively, ε represents the computing resource ratio allocated by UAVj to IoT terminal i, C represents the association decision between UAVj and IoT terminal i and the association is unique, and E[t] represents the weighted sum of energy consumption of UAV and IoT terminal, that is,
[0035]
[0036] O[t]={o1[t],o2[t],...,o n [t]} represents the success sign of task i. If the UAV’s perception of IoT terminal i and the completion delay of task i meet the perception and delay constraints, then o i [t]=1, otherwise o i [t]=0.
[0037] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: wherein step S5 includes the following sub-steps:
[0038] S5-1, status
[0039]
[0040] in, Indicates the three-dimensional space coordinates of UAVj, x i [t] and y i [t] represents the two-dimensional horizontal and vertical coordinates of IoT terminal i; S5-2, action, i.e., optimization variables
[0041] a t =[α1[t],α2[t],...,α m [t],
[0042] p1[t],p2[t],...,p m [t],
[0043] ε 1,1 [t],ε 2,1 [t],...,ε n,m [t],
[0044] c 1,1 [t],c 2,1 [t],...,c n,m ] 1×(2m+2mn) ;
[0045] S5-3, Reward
[0046]
[0047] in, Represents the weight coefficient.
[0048] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: wherein step S6 includes the following sub-steps: S6-1, constructing an Actor network: outputting continuous actions (α, p, ε) and discrete service association decisions c i,j ; S6-2, build the Critic network: evaluate the state value V(s,w), and use time difference update; S6-3, algorithm strategy optimization, S6-3-1, use generalized advantage estimation and experience replay pool; S6-3-2, introduce the clipping function to limit the strategy update amplitude to prevent oscillation; S6-3-3, add entropy regularization term to promote exploration.
[0049] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: wherein step S8 includes the following sub-steps: S8-1, controlling the UAV decision: generating the optimal action based on the trained Actor network Guide the execution of service drones; S8-2, monitor the movement of IoT terminals and task changes in real time, and maintain service stability through online strategy updates.
[0050] The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention may also have the following features: S9, performance verification, building a network environment based on MATLAB, and comparing performance with deep deterministic policy gradient, REINFORCE, static strategy, and random strategy.
[0051] The present invention also provides an unmanned aerial vehicle (UAV)-assisted telepathy and computing integrated system based on deep reinforcement learning, which may also have the following features: a system architecture design model, which designs the system architecture, including deploying UAVs and ground equipment; a time slot division and frame structure design model, which designs the service cycle phase division and frame structure of the UAV; a telepathy and computing service process module, which calculates the radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the UAV based on the phase division and frame structure of the UAV's service cycle; an optimization modeling module, which constructs an optimization model: based on the perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the UAV, a joint optimization target is constructed to minimize each energy consumption and maximize the service success rate; a reinforcement learning modeling module, which constructs a reinforcement learning model based on the optimization model; a reinforcement learning optimization module, which optimizes the reinforcement learning model: constructs an actor network and a critic network and optimizes the algorithm strategy therein to obtain a strategy model; a model training module, which trains the strategy model: trains the strategy model using a single agent, multiple agents or parallel computing method to obtain an optimal actor network; and an online reasoning module, which controls the decision-making of the UAV based on the optimal actor network.
[0052] Functions and effects of the invention
[0053] According to the UAV-assisted telepathy and computing integrated method and system based on deep reinforcement learning involved in the present invention, the aerial perception-communication-computing integrated collaborative processing mechanism based on the UAV swarm can significantly improve the system's service efficiency and resource utilization for ground terminals. Compared with the traditional solution that separates perception, communication and computing functions, the present invention realizes the optimization of the entire process of task uploading, processing and result return through air interface fusion and intelligent scheduling strategy. At the same time, the use of deep reinforcement learning algorithm improves the system's adaptability and scalability to dynamic environments, supports complex scenarios such as multi-UAV collaboration, mobile terminal access, and real-time task processing, and has significant energy-saving, high-efficiency and intelligent decision-making advantages. It is suitable for ubiquitous intelligent services in complex environments such as future 6G intelligent transportation and emergency rescue. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flow chart in an embodiment of the present invention;
[0055] Figure 2 is a schematic diagram of the time slot division and frame structure of each service cycle in an embodiment of the present invention;
[0056] Figure 3 is the algorithm convergence result of PPO, DDPG and REINFORCE in the embodiment of the present invention; and
[0057] Figure 4This is the performance comparison result of PPO with DDPG, REINFORCE, static strategy, and random strategy in the embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate the present invention's drone-assisted synaesthesia computing integrated method and system based on deep reinforcement learning.
[0059] The present invention aims to realize efficient, low-energy, and low-latency perception, communication, and computing collaborative services of drone clusters in complex MEC environments through air interface fusion architecture and deep reinforcement learning optimization algorithm.
[0060] Aiming at the problem of integrated perception and computing in UAV-assisted MEC, this paper proposes an integrated aerial perception-communication-computing collaborative processing architecture based on drone clusters. Through an intelligent resource joint optimization algorithm based on deep reinforcement learning, it coordinates the perception, communication and computing functions of multiple drones, improves the system service quality, reduces energy consumption and communication burden, and ensures the timeliness and reliability of services in complex dynamic scenarios.
[0061] Figure 1 is a flow chart in an embodiment of the present invention.
[0062] like Figure 1 As shown, the UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning provided by the present invention specifically includes the following steps:
[0063] S1, design the system architecture, including deployment of drones and ground equipment.
[0064] Step S1 includes the following sub-steps:
[0065] S1-1, Deployment of drone swarms.
[0066] Deploy a drone swarm U = {1, 2, ..., m}, |U| = m, consisting of multiple drones equipped with radar sensing units, communication antenna arrays, and edge computing servers.
[0067] The drones hover at a fixed height, using orthogonal frequency division multiple access and self-interference cancellation technology to avoid communication interference within the group, and set a safety protection radius to prevent collisions.
[0068] S1-2, ground equipment deployment.
[0069] The distribution of ground IoT devices I={1,2,...,n},|I|=n changes dynamically, generating tasks that need to be processed in real time. in Indicates the data size of the computing task, Indicates the number of CPU cycles required to complete the calculation of each bit task.
[0070] S2, design the service cycle phase division and frame structure of the UAV.
[0071] Step S2 includes the following sub-steps:
[0072] Figure 2 It is a schematic diagram of the time slot division and frame structure of each service cycle in an embodiment of the present invention.
[0073] S2-1, such as Figure 2 As shown, each service period T is evenly divided into N time slots of equal length, each time slot length is Δ=T / N, and each time slot is further divided into two sub-phases.
[0074] S2-2, stage 1 is the perception stage: the duration is α[t]Δ, and the UAV uses the radar system to complete environmental perception.
[0075] S2-3, Phase 2 is the communication-computation phase: the duration is (1-α[t])Δ, which is used for data offloading and computation processing, that is, the IoT terminal completes the task upload and the UAV completes the computation.
[0076] S3, based on the phase division and frame structure of the UAV’s service cycle, calculates the radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the UAV.
[0077] Step S3 includes the following sub-steps:
[0078] S3-1, UAV perception: Each UAVj transmits radar waves to the ground during the perception sub-period and receives reflected signals, embedding information bits in the radar beam sidelobes.
[0079] S3-2, radar echo is converted into radar data, including information such as the location of the IoT terminal:
[0080]
[0081] in Including data such as location, obstacle identification, signal measurement and target identification, j ≥1 is a constant related to the introduced data redundancy, ν j is the switching speed of the radar beam, N θ is the number of quantized angles, f s [t] is the sampling frequency, is the number of quantization bits per sample.
[0082] S3-3, Radar Estimated Information Rate Calculation
[0083]
[0084] Where B represents bandwidth, p j [t] represents the perceived power of the UAV, represents the channel gain of the radar echo signal, d i,j [t] represents the distance between UAVj and IoT terminal i, Represents the reference channel gain at a distance of 1 meter. The average information estimated by the radar should not be less than the required radar detection data. Perception needs to meet the constraints
[0085]
[0086] where c i,j represents the association decision between UAVj and IoT terminal i.
[0087] S3-4, calculation of perceived energy consumption:
[0088]
[0089] S3-5, data communication: The IoT terminal task data size is The task upload rate is Then the task upload delay is:
[0090]
[0091] S3-6, calculate communication energy consumption:
[0092]
[0093] Among them, q i [t] represents the signal transmission power of the IoT terminal.
[0094] S3-7, UAV receives the offload data and uses dynamic voltage and frequency adjustment technology to control resource allocation. The CPU calculation time is:
[0095]
[0096] Among them, ε i,j [t] represents the proportion of computing resources allocated by UAVj to IoT terminal i, Indicates the maximum computing capability of UAVj.
[0097] S3-8, calculate the energy consumption:
[0098]
[0099] Among them, κ j Represents hardware-related constants.
[0100] S3-9, calculate the total energy consumption of the drone:
[0101]
[0102] S3-10, calculate the total delay of the task service, which must meet the delay constraint
[0103] T i ser [t] = T i off [t]+T i comp [t]
[0104] T i ser [t]≤(1-α j [t]Δ).
[0105] S4, build an optimization model: Based on perception energy consumption, communication energy consumption, computing energy consumption and total energy consumption of drones, build a joint optimization goal to minimize each energy consumption and maximize the service success rate.
[0106] Step S4 includes the following sub-steps:
[0107] With the goal of maximizing service success rate and minimizing energy consumption, we build a joint optimization goal:
[0108]
[0109] Among them, α and p represent the time slot ratio and signal transmission power decision set of UAV respectively, ε represents the computing resource ratio allocated by UAVj to IoT terminal i, C represents the association decision between UAVj and IoT terminal i and the association is unique, and E[t] represents the weighted sum of energy consumption of UAV and IoT terminal, that is,
[0110]
[0111] O[t]={o1[t],o2[t],...,o n [t]} represents the success sign of task i. If the UAV’s perception of IoT terminal i and the completion delay of task i meet the perception and delay constraints, then o i [t]=1, otherwise o i [t]=0.
[0112] S5, builds a reinforcement learning model based on the optimization model.
[0113] Step S5 includes the following sub-steps:
[0114] S5-1, status
[0115]
[0116] in, Indicates the three-dimensional space coordinates of UAVj, x i [t] and y i [t] represents the two-dimensional horizontal and vertical coordinates of IoT terminal i.
[0117] S5-2, action, i.e. optimization variables
[0118] a t =[α1[t],α2[t],...,α m [t],
[0119] p1[t],p2[t],...,p m [t],
[0120] ε 1,1 [t],ε 2,1 [t],...,ε n,m [t],
[0121] c 1,1 [t],c 2,1 [t],...,c n,m ] 1×(2m+2mn) .
[0122] S5-3, Reward
[0123]
[0124] in, Represents the weight coefficient.
[0125] S6, optimize the reinforcement learning model: build the Actor network and Critic network and optimize the algorithm strategy to obtain the strategy model.
[0126] Optimization based on deep reinforcement learning: The Proximal Policy Optimization (PPO) algorithm is used for optimization and solution.
[0127] Step S6 includes the following sub-steps:
[0128] S6-1, build an Actor network: output continuous actions (α, p, ε) and discrete service association decisions c i,j .
[0129] S6-2, build the critic network: evaluate the state value V(s,w) and use temporal difference (TD) update.
[0130] S6-3, algorithm strategy optimization,
[0131] S6-3-1, using generalized advantage estimation with an experience replay pool.
[0132] S6-3-2, introduce the clipping function to limit the strategy update amplitude to prevent oscillation.
[0133] S6-3-3, add entropy regularization term to promote exploration.
[0134] S7, training strategy model: Use single agent, multi-agent or parallel computing to train the strategy model to obtain the optimal Actor network.
[0135] S8, controls the decision-making of the drone based on the optimal Actor network.
[0136] Step S8 includes the following sub-steps:
[0137] S8-1, Drone control decision-making: Generate optimal action a based on the trained Actor network t * , guiding the execution of service drones.
[0138] S8-2 monitors IoT terminal movement and task changes in real time, and maintains service stability through online policy updates.
[0139] S9, performance verification, builds a network environment based on MATLAB, and compares the performance of the PPO-based method proposed in this invention with Deep Deterministic Policy Gradient (DDPG), REINFORCE, static strategy, and random strategy.
[0140] Result analysis:
[0141] Figure 3 These are the convergence results of the PPO, DDPG, and REINFORCE algorithms in the embodiments of the present invention. Figure 4 This is the performance comparison result of PPO with DDPG, REINFORCE, static strategy, and random strategy in the embodiment of the present invention.
[0142] like Figure 3-4 As shown in the figure, the algorithm is able to converge, indicating that the agent based on deep reinforcement learning has learned an effective strategy and has high learning efficiency and stability; compared with the baseline algorithm, the service success rate is significantly improved by 61.44%.
[0143] The present invention also provides a UAV-assisted synaesthesia-computing integrated system based on deep reinforcement learning, comprising:
[0144] The system architecture design model is designed according to the method of step S1 above, including the deployment of drones and ground equipment.
[0145] The time slot division and frame structure design model, according to the method of step S2 above, designs the service cycle phase division and frame structure of the UAV.
[0146] The synaptic computing service process module calculates the radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the drone based on the phase division and frame structure of the drone's service cycle according to the method of step S3 above.
[0147] The optimization modeling module constructs an optimization model according to the method of step S4 above: based on the perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the drone, a joint optimization goal is constructed to minimize the energy consumption of each and maximize the service success rate.
[0148] The reinforcement learning modeling module constructs a reinforcement learning model based on the optimization model according to the method of step S5 above.
[0149] The reinforcement learning optimization module optimizes the reinforcement learning model according to the method of step S6 above: constructs the actor network and the critic network and optimizes the algorithm strategy therein to obtain the strategy model.
[0150] The model training module trains the strategy model according to the method of step S7 above: the strategy model is trained using a single agent, multiple agents or parallel computing method to obtain the optimal Actor network.
[0151] The online reasoning module controls the drone's decision-making according to the method of step S8 above and the optimal Actor network.
[0152] Those skilled in the art will appreciate that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning, characterized in that: The steps include: S1, designing the system architecture, including deploying drones and ground equipment; S2, designing the service cycle phase division and frame structure of the UAV; S3, calculating the radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption, and total energy consumption of the UAV based on the phase division and frame structure of the UAV's service cycle; S4, constructing an optimization model: based on the sensing energy consumption, the communication energy consumption, the computing energy consumption and the total energy consumption of the UAV, constructing a joint optimization goal to minimize each energy consumption and maximize the service success rate; S5, constructing a reinforcement learning model based on the optimization model; S6, optimizing the reinforcement learning model: constructing an Actor network and a Critic network and optimizing the algorithm strategies therein to obtain a strategy model; S7, training the policy model: training the policy model using a single agent, multiple agents, or parallel computing to obtain an optimal Actor network; S8, controlling the decision of the UAV according to the optimal Actor network.
2. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 1 is characterized by: in, The step S1 includes the following sub-steps: S1-1, deployment of the drone swarm, Deploy a drone swarm U = {1, 2, ..., m}, |U| = m, consisting of multiple drones equipped with radar sensing units, communication antenna arrays, and edge computing servers. The drones are fixed and hovered at a preset height, using orthogonal frequency division multiple access and self-interference cancellation technology to avoid communication interference within the group, and setting a safety protection radius to prevent collisions; S1-2, Ground Equipment Deployment, The distribution of ground IoT devices I={1,2,...,n},|I|=n changes dynamically, generating tasks that need to be processed in real time. in Indicates the data size of the computing task, Indicates the number of CPU cycles required to complete the calculation of each bit task.
3. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 2 is characterized by: in, The step S2 includes the following sub-steps: S2-1, each service period T is evenly divided into N time slots of equal length, each time slot length is Δ = T / N, and each time slot is further divided into two sub-phases; S2-2, stage 1 is the perception stage: the duration is α[t]Δ, the UAV uses the radar system to complete the environment perception; S2-3, Phase 2 is the communication-computation phase: the duration is (1-α[t])Δ, which is used for data offloading and computation processing, that is, the IoT terminal completes the task upload and the UAV completes the computation.
4. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 3 is characterized by: in, The step S3 includes the following sub-steps: S3-1, UAV perception: Each UAVj transmits radar waves to the ground during the perception sub-period and receives reflected signals, embedding information bits in the radar beam sidelobes; S3-2, radar echo is converted into radar data, including information such as the location of the IoT terminal: in Including data such as location, obstacle identification, signal measurement and target identification, j ≥1 is a constant related to the introduced data redundancy, ν j is the switching speed of the radar beam, N θ is the number of quantized angles, f s [t] is the sampling frequency, is the number of quantization bits per sample; S3-3, Radar Estimated Information Rate Calculation Where B represents bandwidth, p j [t] represents the perceived power of the UAV, represents the channel gain of the radar echo signal, d i,j [t] represents the distance between UAVj and IoT terminal i, Represents the reference channel gain at a distance of 1 meter. The average information estimated by the radar should not be less than the required radar detection data. Perception needs to meet the constraints where c i,j represents the association decision between UAVj and IoT terminal i; S3-4, calculation of perceived energy consumption: S3-5, data communication: The IoT terminal task data size is The task upload rate is Then the task upload delay is: S3-6, calculate communication energy consumption: Among them, q i [t] represents the signal transmission power of the IoT terminal; S3-7, UAV receives the offload data and uses dynamic voltage and frequency adjustment technology to control resource allocation. The CPU calculation time is: Among them, ε i,j [t] represents the proportion of computing resources allocated by UAVj to IoT terminal i, Indicates the maximum computing capability of UAVj; S3-8, calculate the energy consumption: Among them, κ j Represents hardware-related constants; S3-9, calculate the total energy consumption of the drone: S3-10, calculate the total delay of the task service, which must meet the delay constraint T i ser [t]=T i off [t]+T i comp [t] T i ser [t]≤(1-α j [t]D).
5. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 4 is characterized in that: in, The step S4 includes the following sub-steps: With the goal of maximizing service success rate and minimizing energy consumption, we build a joint optimization goal: Among them, α and p represent the time slot ratio and signal transmission power decision set of UAV respectively, ε represents the computing resource ratio allocated by UAVj to IoT terminal i, C represents the association decision between UAVj and IoT terminal i and the association is unique, and E[t] represents the weighted sum of energy consumption of UAV and IoT terminal, that is, O[t]={o1[t],o2[t],...,o n [t]} represents the success sign of task i. If the UAV’s perception of IoT terminal i and the completion delay of task i meet the perception and delay constraints, then o i [t]=1, otherwise o i [t]=0.
6. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 5 is characterized by: in, The step S5 includes the following sub-steps: S5-1, status in, Indicates the three-dimensional space coordinates of UAVj, x i [t] and y i [t] represents the two-dimensional horizontal and vertical coordinates of IoT terminal i; S5-2, action, i.e. optimization variables a t =[α1[t],α2[t],...,α m [t], p1[t],p2[t],...,p m [t], e 1,1 [t],e 2,1 [t],...,e n,m [t], c 1,1 [t],c 2,1 [t],...,c n,m ] 1×(2m+2mn) ; S5-3, Reward in, Represents the weight coefficient.
7. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 6 is characterized by: in, The step S6 includes the following sub-steps: S6-1, build an Actor network: output continuous actions (α, p, ε) and discrete service association decisions c i,j ; S6-2, build the critic network: evaluate the state value V(s,w) and use temporal difference update; S6-3, algorithm strategy optimization, S6-3-1, using generalized advantage estimation and experience replay pool; S6-3-2, introduce a clipping function to limit the strategy update amplitude to prevent oscillation; S6-3-3, add entropy regularization term to promote exploration.
8. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 7 is characterized by: in, The step S8 includes the following sub-steps: S8-1, drone control decision-making: generating optimal actions based on the trained Actor network Direct service drone execution; S8-2 monitors IoT terminal movement and task changes in real time, and maintains service stability through online policy updates.
9. The UAV-assisted synaesthesia-computing integrated method based on deep reinforcement learning according to claim 8 is characterized in that: Also includes: S9, performance verification, builds a network environment based on MATLAB, and compares performance with deep deterministic policy gradient, REINFORCE, static strategy, and random strategy.
10. The UAV-assisted synaesthesia-computing integrated system based on deep reinforcement learning according to claim 1, characterized in that: include: System architecture design model, designing the system architecture, including the deployment of UAVs and ground equipment; Time slot division and frame structure design model, design the service cycle phase division and frame structure of the UAV; The synergistic computing service process module calculates the radar estimated information rate, perception energy consumption, communication energy consumption, computing energy consumption and the total energy consumption of the drone based on the phase division and frame structure of the drone's service cycle; An optimization modeling module constructs an optimization model: based on the perception energy consumption, the communication energy consumption, the computing energy consumption and the total energy consumption of the UAV, a joint optimization goal is established to minimize each energy consumption and maximize the service success rate; A reinforcement learning modeling module, which constructs a reinforcement learning model based on the optimization model; A reinforcement learning optimization module optimizes the reinforcement learning model: constructs the actor network and the critic network and optimizes the algorithm strategy therein to obtain a strategy model; Model training module, training the policy model: using single agent, multi-agent or parallel computing to train the policy model to obtain the optimal Actor network; An online reasoning module controls the decision-making of the UAV based on the optimal Actor network.