An unmanned aerial vehicle assisted offshore dynamic anti-interference communication method
Patent Information
- Application Number
- CN202611230778.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]随着海上活动与海洋经济的持续发展,建立高速且可靠的海上通信系统日益受到关注,然而,复杂的海洋环境使得传统地面通信基础设施的部署面临巨大挑战
[0056]本发明设计开发的一种基于无人机辅助的海上动态抗干扰通信方法,考虑干扰者的动态跟踪特点,以及低轨卫星频繁切换可能带来的乒乓效应,通过联合优化无人机的最佳飞行路径、最佳服务浮标选择以及最佳上传卫星选择,在最小化干扰影响的同时,确保高效、可靠的数据传输,通过LDSAC-AD算法,采用动作空间解耦策略,将离散决策变量(浮标选择和卫星选择)与连续决策变量(无人机轨迹)分离处理,从而有效解决了混合动作空间带来的维度爆炸和探索困难问题。对于离散变量,采用-贪心算法进一步探索和利用来优化变量的选择,对于连续变量,引入了基于大语言模型引导的超参数优化方案,该方案能够根据训练过程中的状态信息动态调节超参数(熵温度系数、学习率等),提升了算法对不同干扰场景的适应能力,还显著增强了训练过程的稳定性,从而提升整体策略的求解质量与收敛效率,在抗干扰性能、数据收集效率、能耗控制以及收敛速度以及鲁棒性方面都具有高效性和优越性。
Smart Images

Figure CN122802019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of maritime safety communication technology, and more specifically, to a dynamic anti-jamming communication method for the sea based on unmanned aerial vehicles (UAVs). Background Technology
[0002] With the continuous development of maritime activities and the marine economy, the establishment of high-speed and reliable maritime communication systems has attracted increasing attention. However, the complex marine environment poses significant challenges to the deployment of traditional terrestrial communication infrastructure. Against this backdrop, satellites with wide-area coverage capabilities are increasingly being used for maritime data transmission, effectively supporting information exchange between ships and between ships and shore. Low Earth orbit satellites, in particular, can significantly improve communication performance due to their lower operating altitude and shorter transmission latency. However, low-power devices (such as buoys) widely deployed at sea are often limited by antenna gain and transmission power, making it difficult to directly transmit data to satellites at greater distances.
[0003] To assist low-power maritime buoys in uploading data to satellites, drones are deployed as flexible relay platforms between the buoys and satellites, responsible for collecting buoy data and forwarding it to the satellites. However, due to the openness of maritime wireless links, unauthorized users (such as malicious unmanned surface vessels) may dynamically interfere with the drone's received signals, severely affecting the reliability of data transmission. To ensure efficient data transmission, the drone needs to dynamically optimize its flight path to minimize interference. Simultaneously, to improve data throughput, the drone also needs to select suitable satellites from the low-Earth orbit satellite constellation for data uploading. However, when multiple satellites have similar orbits, the drone may frequently switch between satellites, creating a ping-pong effect, thereby increasing communication overhead and transmission latency. Moreover, in the considered system, unauthorized users with detection and tracking capabilities can move relative to the drone, and their dynamic positional changes will significantly affect system performance.
[0004] Therefore, there is an urgent need to develop a dynamic anti-jamming communication method for maritime applications that can ensure the security, effectiveness, and stability of data transmission. Summary of the Invention
[0005] The purpose of this invention is to design and develop a UAV-assisted dynamic anti-jamming communication method for the sea, using the LDSAC-AD algorithm and Greedy algorithms handle continuous and discrete variables after decoupling action space strategies, improving the safety, stability, and adaptability of maritime communication.
[0006] The technical solution provided by this invention is as follows:
[0007] A method for dynamic anti-jamming communication at sea based on unmanned aerial vehicle (UAV) assistance includes the following steps:
[0008] Step 1: Collect information on the location of marine buoys, the orbital information of low-Earth orbit satellite constellations, the initial position of UAVs, and the initial position of unmanned surface vessels;
[0009] Step 2: Discretize the entire task cycle into T equal-length time slots and construct the objective function for the optimization problem;
[0010] The objective function of the optimization problem is to maximize the sum of the total reception rates of the maritime-UAV link and the UAV-satellite link.
[0011] Step 3: Model the objective function as a Markov decision process, and use a decoupled action space strategy to process discrete and continuous variables respectively. Discrete variables are decided through an ε-greedy architecture, while continuous variables are decided through a distributed soft actor-commentator guided by a large language model, to obtain the optimal flight position decision, optimal buoy selection strategy and optimal satellite switching strategy for the UAV.
[0012] Step 4: Based on the optimal flight position decision, optimal buoy selection strategy, and optimal satellite switching strategy of the UAV, complete the dynamic anti-jamming communication at sea.
[0013] Preferably, the receiving rate of the maritime-UAV link satisfies:
[0014] ;
[0015] In the formula, For the current time slot, the first The drone received the first The data reception rate of each buoy. For channel bandwidth, Indicates the buoy's transmission power. Indicates the first The transmission power of an unmanned surface vessel Indicates noise power. For the current time slot, the first The buoy and the first Path loss between drones For the current time slot, the first The unmanned surface vessel and the first Path loss between drones.
[0016] Preferably, the receiving rate of the UAV-satellite link satisfies:
[0017] ;
[0018] In the formula, For the current time slot, the first The satellite received the first The data reception rate of a drone. To indicate the first The transmission power of each drone, For the current time slot, the first The drone and the first Path loss between satellites.
[0019] Preferably, the objective function of the optimization problem needs to satisfy the following constraints:
[0020] (9)
[0021] In the formula, For time slices The amount of data received by the drone from all the buoys. For time slices The total amount of data received by the satellite For time slices, For the drone's cache capacity, This refers to the entire task cycle.
[0022] Preferably, the objective function of the optimization problem also needs to satisfy the following constraints:
[0023] ;
[0024] ;
[0025] ;
[0026] ;
[0027] ;
[0028] In the formula, For the receiving rate threshold, For data volume threshold, For single time slot duration, For the first Energy consumption during drone communication Energy consumption threshold For the number of satellite updates, This represents the maximum number of switching operations.
[0029] Preferably, the Markov decision process includes:
[0030] ;
[0031] ;
[0032] ;
[0033] In the formula, This is the state space under the current time slot. and These represent the positions of the drone and the unmanned surface vessel in the current time slot, respectively. It is the buoy index in the current time slot. = This represents the cumulative data received from each buoy. It is the first in the current time slot Energy consumption of a drone = It is the set of positions of all low-Earth orbit satellites. and These are the index of the selected satellite and the number of satellite switching attempts, respectively. = This represents the cumulative amount of data uploaded to each satellite. For the action space in the current time slot, The reward function for the current time slot. , and Scaling factor This is a penalty item.
[0034] Preferably, the ε-greedy architecture specifically includes the following steps:
[0035] Step 1: Obtain the state of the previous time slot Actions selected in the previous time slot This corresponds to the target buoy B[t-1] or target satellite S[t-1] in the previous time slot, and the reward for the previous time slot is calculated. ;
[0036] Step 2: Obtain the status of the current time slot At this point, directly change the state. Given a Q-table, iterate through all possible discrete sub-actions in the given state and find the action that maximizes the Q-value. ;
[0037] Step 3: Calculate the error and update the Q value with the future estimated value;
[0038] Step 4, through - Greedy algorithm updates discrete action space:
[0039] ;
[0040] In the formula, For parameter values, It is a random number. For the set of discrete action subspaces, As a heuristic exploration strategy, In the state Take action below The obtained Q value.
[0041] Preferably, the distribution of soft actors-critic guided by the large language model specifically includes the following steps:
[0042] Step I: Initialize the parameters of the distributed soft value network Policy network parameters Target value network parameters Target policy network parameters and experience playback buffer pool ;
[0043] Step II: Initialize the initial positions of the drone and malicious jammer, and perform motion space sampling according to the strategy. Update the current location of the drone and malicious jammer, calculate the current reward r[t] based on the buoy index and satellite index of the current service, obtain the state space s[t+1] of the next time slot, and update the experience replay buffer M;
[0044] Step III, if t mod And M Update the distributed soft value network;
[0045] Step IV: Update the policy network;
[0046] Step V: Maintain the target value network parameters using a soft update method. and target policy network parameters ;
[0047] Step VI, if mod By using a large language model to guide an adaptive hyperparameter tuning mechanism, dynamic hyperparameter adjustment can be achieved.
[0048] Preferably, the optimization objective of steps III and V is:
[0049] ;
[0050] ;
[0051] In the formula, The loss function of a value network, For quantile Huber loss, for The quantile and the th quantile Pairwise time-series difference error between quantiles , For the first quantiles, For the first quantiles, The loss function of a value network, For soft action-value functions, ( () is the expected value operator. This indicates that the state is sampled from the experience replay buffer M. It is a noise vector sampled from a fixed distribution. It is an adaptive temperature coefficient.
[0052] Preferably, step VI further includes:
[0053] ;
[0054] In the formula, For the first In the first iteration, the hyperparameters for actual application will be configured. For interpolation ratio, , This represents the feasible set of hyperparameters. ( This indicates projection onto a feasible set of hyperparameters. Configure candidate hyperparameters for the output of the large language model.
[0055] The beneficial effects of this invention are as follows:
[0056] This invention presents a UAV-assisted dynamic anti-jamming communication method for the sea. Considering the dynamic tracking characteristics of jammers and the ping-pong effect that may result from frequent switching of low-Earth orbit satellites, it jointly optimizes the UAV's optimal flight path, optimal service buoy selection, and optimal upload satellite selection. This minimizes interference impact while ensuring efficient and reliable data transmission. Using the LDSAC-AD algorithm and an action space decoupling strategy, discrete decision variables (buoy selection and satellite selection) are separated from continuous decision variables (UAV trajectory), effectively solving the problems of dimensional explosion and exploration difficulties caused by the mixed action space. For discrete variables, a... - The greedy algorithm is further explored and utilized to optimize variable selection. For continuous variables, a hyperparameter optimization scheme based on a large language model is introduced. This scheme can dynamically adjust hyperparameters (entropy temperature coefficient, learning rate, etc.) according to the state information during training, which improves the algorithm's adaptability to different interference scenarios and significantly enhances the stability of the training process. This improves the overall solution quality and convergence efficiency of the strategy. It has high efficiency and superiority in terms of anti-interference performance, data collection efficiency, energy consumption control, convergence speed, and robustness. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of the structure of the marine data transmission system described in this invention;
[0058] Figure 2 This is a flowchart illustrating the UAV-assisted dynamic anti-jamming communication method for the sea as described in this invention.
[0059] Figure 3 This is a schematic diagram of the convergence performance curves of the different algorithms described in this invention. Detailed Implementation
[0060] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0061] This invention provides a UAV-assisted dynamic anti-jamming communication method for the sea, which uses a maritime data transmission system for communication, such as... Figure 1 As shown, the maritime data transmission system consists of low Earth orbit satellites, buoys, a swarm of drones, and unmanned surface vessels. Among them, due to the limitations of the low power consumption of the buoy equipment at sea, it is usually difficult to directly transmit data to distant low Earth orbit satellites. The drone swarm, as a mobile and flexible communication relay platform between the buoys and the low Earth orbit satellites, is responsible for collecting buoy data and forwarding it to the satellites. However, there are also unmanned surface vessels on the sea surface that act as malicious jammers, affecting the reception of drone data.
[0062] The UAV-assisted dynamic anti-jamming communication method for the sea specifically includes the following steps:
[0063] Step 1: Collect information on the location of marine buoys, the orbital information of low-Earth orbit satellite constellations, the initial position of UAVs, and the initial position of malicious jammers (unmanned surface vessels);
[0064] The location of the maritime buoys is determined according to the maritime buoy observation mission schedule. The low-Earth orbit (LEO) satellite constellation orbit information includes the real-time location and orbit trajectory of all satellites in the LEO satellite constellation, which is determined based on the real historical data of the LEO satellite constellation. The UAV uses its onboard optical camera and detection equipment (radar) to detect the dynamic location of malicious jammers and uses GPS positioning equipment to determine the initial location, speed, and direction of the malicious jammers. The operating area of the UAV is determined based on the maritime buoys, and the initial location of the UAV is determined based on the location of the maritime buoys and the unmanned surface vessel.
[0065] In this embodiment, eight buoys constitute a service area. Within each service area, there is only one UAV and one unmanned surface vessel. In a set of time slots, one UAV serves different buoys in its service area, while the unmanned surface vessel interferes with the UAV's data reception.
[0066] Step 2: Discretize the entire mission cycle into T equal-length time slots. Within each time slot, the UAV simultaneously performs data acquisition and data upload in full-duplex mode. Design an optimization problem based on maritime communication requirements, with the objective of maximizing the sum of the UAV's receiving rate and the satellite's receiving rate.
[0067] (1)
[0068] In the formula, the entire task cycle is represented by a set. , The number of time slots. Indicates the current time slot. Let represent all decision variables, where ={ } represents the set of drone locations. This represents the three-dimensional position of the UAV in the current time slot. ={B[t] } represents the set of buoy indices selected by the drone. ={S[t] } Select the set of satellite indexes for the drone to transmit. For the currently serving float index, For the number of buoys, Select the satellite index for the current selection. For satellite collection, The number of satellites. For the current drone index, For the current time slot, the first The satellite received the first The data reception rate of the first drone, i.e., the data received by the first drone in the drone-satellite link. The reception rate of each satellite, For the current time slot, the first The drone received the first The data reception rate of the first buoy, i.e., the data from the first buoy in the maritime-UAV link. The drone reception rate of a single buoy;
[0069] The optimization problem means maximizing the sum of the total receiving rates of the two links (the maritime-UAV link and the UAV-satellite link). The UAV receiving rate is used to measure the anti-interference capability of the maritime-UAV link transmission, while also considering the transmission efficiency; the satellite receiving rate is used to measure the signal reception effectiveness in the actual transmission process.
[0070] The maritime-drone link originating from the first The effective signal for the drone's receiving rate from a single buoy is the legitimate signal power transmitted from the buoy to the drone. Invalid signals include interference signals transmitted to the drone by malicious jammers and noise power in the environment. Therefore, the calculation formula is:
[0071] (2)
[0072] In the formula, For channel bandwidth, Indicates the buoy's transmission power. Indicates the first The transmission power of an unmanned surface vessel Indicates noise power. For the current time slot, the first The buoy and the first Path loss between drones For the current time slot, the first The unmanned surface vessel and the first Path loss between drones;
[0073] The first time slot in the current time slot The buoy and the first The path loss between drones satisfies:
[0074] (3)
[0075] In the formula, For the current time slot, the first The buoy and the first Large-scale fading between drones, in units of , This refers to small-scale fading in the current time slot;
[0076] The first time slot in the current time slot The buoy and the first Large-scale fading between individual drones satisfies:
[0077] (4)
[0078] In the formula, For line-of-sight links, attenuation factor. For non-line-of-sight links, the attenuation factor is... This represents the first constant parameter of the sigmoid function. This represents the second constant parameter of the sigmoid function. As the first intermediate parameter, This is the second intermediate parameter;
[0079] In this embodiment, =2.3, =34, =5.0188, =0.3511.
[0080] The first intermediate parameter satisfies:
[0081] (5)
[0082] In the formula, For the current time slot, the first The buoy and the first The distance between the drones This represents the carrier frequency in MHz, and in this embodiment, =2.4MHz;
[0083] The second intermediate parameter satisfies:
[0084] (6)
[0085] In the formula, For the current time slot, the first The height of the drone (i.e.) );
[0086] The first time slot in the current time slot The satellite received the first The data reception rate of the drone transmission meets the following requirements:
[0087] (7)
[0088] In the formula, To indicate the first The transmission power of each drone, For the current time slot, the first The drone and the first Path loss between satellites;
[0089] The first The drone and the first The path loss between satellites satisfies:
[0090] (8)
[0091] In the formula, and These are the path loss coefficient and exponent for the UAV-satellite link, respectively. For the first The drone and the first The distance between satellites For small-scale fading in the current time slot, it passes through having The fading parameters are described by the Nakagami-m fading model;
[0092] Considering the potential data overflow risk from drones, the amount of data transmitted uplink by the drone cannot exceed the amount of data it receives, and must meet preset buffer constraints, i.e., for any given time slice... The difference between the amount of data received and the amount of data uploaded by the drone is less than the drone's cache capacity. :
[0093] (9)
[0094] In the formula, For time slices The amount of data received by the drone from all the buoys. For time slices The total amount of data received by the satellite;
[0095] The time slice The amount of data and time slices received by the drone from all buoys The total amount of data received by the next satellite satisfies:
[0096] (10)
[0097] (11)
[0098] In the formula, This refers to the duration of a single time slot.
[0099] Considering the limited marine resources, the constraints of the optimization problem include the limited energy consumption of UAVs, and to minimize the ping-pong effect, the number of satellite switching operations is also limited.
[0100] 1) The three-dimensional position of the UAV is crucial for secure data reception at sea and lays the foundation for subsequent UAV-satellite link connections. To cope with signal attenuation caused by interference, the UAV needs to continuously adjust its position. However, this process consumes energy, thus shortening mission duration. Therefore, to extend the operating time, the system considers minimizing the UAV's energy consumption as much as possible. Accordingly, a UAV energy consumption model is provided to mathematically model these dynamic costs.
[0101] (12)
[0102] In the formula, For the first Energy consumption during communication by a drone, of which Indicates the total flight time. This indicates the speed of the drone in the current time slot. Indicates the weight of the drone. Represents gravitational acceleration. Indicates the drone's altitude at the end of the flight. This indicates the initial altitude of the drone. The speed at the moment the flight terminates. The velocity at the initial moment of flight, This represents the propulsion power consumption of a drone flying in two-dimensional horizontal space, and it satisfies:
[0103] (13)
[0104] In the formula, and These are two constants, representing the blade profile and induced power in the hovering state, respectively. Indicates the speed of the drone. This represents the tip velocity of the rotor blades. This represents the average rotor blade induced velocity during hovering. Indicates the fuselage drag ratio. Indicates air density, Indicates the rigidity of the rotor blades. This indicates the area of the rotor disk.
[0105] 2) The UAV will dynamically select an optimal satellite from the low-Earth orbit satellite constellation to maintain connectivity between the UAV and the satellite. The selected satellite provides continuous tracking services during the time slice. Throughout the time slice, the system will reassess the feasibility of the connection every 10 seconds. If the satellite is not visible or has poor transmission performance, the UAV will select a more ideal satellite for transmission. This process generates a time-based switching sequence. ,in Indicates time slot Connecting to the low-Earth orbit satellite index for the drone; however, frequent satellite switching can lead to a ping-pong effect, causing communication delays and additional overhead. To quantify this overhead, the number of satellite updates is mathematically defined. At the same time, the maximum number of updates in the total time slots is limited:
[0106] (14)
[0107] In formula (14), if the satellite index selected in the current time slot is different from that selected in the previous time slot, the number of updates increases by one; otherwise, the number of updates remains unchanged.
[0108] Therefore, the objective function of the optimization problem also needs to satisfy the following constraints:
[0109] That is, through the maritime-drone link from the third The drone's reception rate for each buoy is higher than the reception rate threshold. To ensure anti-interference capability;
[0110] This ensures that enough data is collected from each buoy, so that the total amount of data received by the drone exceeds a data volume threshold. ;
[0111] Controlling the total energy consumption of drones within the energy consumption threshold the following;
[0112] To ensure the amount of buoy data received by the drone It can cover the amount of data it uploads. ;
[0113] Limit the number of satellite handovers to the maximum number of handovers. The following measures are taken to ensure transmission stability;
[0114] In this embodiment, the receiving rate threshold =1Mbps, data volume threshold =300Mbits, energy consumption threshold =60000J, maximum number of switching times .
[0115] Step 3: The constructed optimization problem is modeled as a Markov decision process, and a decoupled action space strategy is adopted to handle discrete and continuous variables separately. Discrete variables are decided using an ε-greedy architecture to balance exploration and exploitation. A large language model-guided distributed soft actor-critic with Decoupled Action Spaces (LDSAC-AD) algorithm is introduced to optimize continuous variables. An adaptive hyperparameter tuning mechanism guided by the large language model is used as a meta-controller to achieve dynamic hyperparameter adjustment, update the UAV's flight trajectory, and select the buoys to be served and the satellites to be transmitted.
[0116] The Markov decision process includes:
[0117] The state space includes information such as the drone's position, the location of the jammer, the location of the buoy, the satellite's position, the amount of data received by each satellite, the amount of data collected by the drone, the drone's energy consumption, and the number of satellite handovers.
[0118] (15)
[0119] In the formula, For the state space of the current time slot, and These represent the positions of the drone and the unmanned surface vessel in the current time slot, respectively. It is the buoy index in the current time slot. = This represents the cumulative data received by the drone from each buoy. It is the first Energy consumption during drone communication = It is the set of positions of all low-Earth orbit satellites. and These are the index of the selected satellite and the number of satellite switching attempts, respectively. = This indicates the cumulative amount of data uploaded by the drone to each satellite;
[0120] The action space includes the drone's three-dimensional position, the currently selected buoy index, and the selected satellite index:
[0121] (16)
[0122] The reward function is:
[0123] (17)
[0124] In the formula, , and Scaling factor This is a penalty term to ensure that the variable meets the constraint conditions.
[0125] Action space includes continuous and discrete variable values. A single deep reinforcement learning algorithm struggles to efficiently solve for this mixture of variables. While discretizing the continuous space can address the mixed action space, the dimensionality of the discrete solution space can grow exponentially with the number of floats, making deep reinforcement learning training difficult to converge. Furthermore, coupling continuous and discrete variables within the same action space exacerbates the exploration difficulty, making it hard for the agent to sample effective policies from the vast and mixed action space. Therefore, this invention employs a method of separating discrete and continuous action spaces, and processes them separately through independent policy modules:
[0126] The decoupled discrete action subspace 1 is the buoy index. The decoupled discrete action subspace 2 is the satellite index. , ( , The discrete action subspace is the set of actions; the decoupled continuous action subspace is the three-dimensional position of the UAV. .
[0127] Different solution schemes are adopted for discrete action spaces and continuous action spaces:
[0128] a. Solving for the discrete action subspace:
[0129] Step 1: Obtain the state of the previous time slot Actions selected in the previous time slot This corresponds to the target buoy B[t-1] or target satellite S[t-1] in the previous time slot, and the reward for the previous time slot is calculated. ;
[0130] Step 2: Obtain the status of the current time slot At this point, directly change the state. Input the Q-table (initially 0, iteratively updated), then iterate through all possible discrete sub-actions in that state and find the action that maximizes the Q-value. , represented as ;
[0131] Step 3: Calculate the error according to formula (18) and update the Q value with the future estimated value:
[0132] (18)
[0133] In the formula, The Q value is the value updated in the previous time slot. The Q value assumed in the previous time slot, The maximum Q value calculated for the current time slot. [t] represents the action that maximizes Q among all possible actions in the current time slot. This is a discount factor used to balance short-term and long-term rewards. The learning rate;
[0134] In this embodiment, =0.9, =0.15.
[0135] Step 4, through - Greedy algorithms update the discrete action space, balancing exploration and exploitation, thereby reducing the risk of getting trapped in suboptimal solutions:
[0136] (19)
[0137] In the formula, It is a random number. For parameter values, The heuristic exploration strategy is to select the nearest buoy / visible satellite to the drone as the discrete variable solution. In the state Take action below The obtained Q value;
[0138] In this embodiment, =0.3;
[0139] That is, when the random number rand≤ If the target is not immediately visible, the system enters the heuristic exploration phase, selecting the nearest candidate target buoy or visible satellite; otherwise, it enters the utilization phase, selecting the candidate target with the highest Q-value; thus obtaining discrete action solutions. That is, the target buoy B[t] or the target satellite S[t];
[0140] b. For continuous action subspaces ( The LDSAC-AD algorithm is used to optimize the 3D position of the UAV. This algorithm extends the traditional scalar action-value function to a complete distribution of soft rewards. It estimates the soft-discounted reward distribution through distributed soft Bellman operators and quantile regression to capture the uncertainty and potential risk of rewards. Specifically, it includes:
[0141] Step I: Initialize the parameters of the distributed soft value network Policy network parameters Target value network parameters Target policy network parameters and experience playback buffer pool ;
[0142] The parameters of the four networks are initialized randomly, and their linear layers are initialized with kaiming_uniform to ensure stable gradient propagation. The experience replay buffer is initially empty and has a size of 1e5.
[0143] Step II: Initialize the initial positions of the drone and malicious jammer, and perform motion space sampling according to the strategy. Update the current positions of drones and malicious jammers, obtain the buoy index and satellite index of the current service based on the solution of the discrete action subspace, calculate the current reward r[t] and obtain the state space s[t+1] of the next time slot, and update the experience replay buffer pool. ;
[0144] Step III, if t mod and That is, the current time slot t can be divided by the algorithm update interval number. Furthermore, the size of the experience replay buffer pool is not less than the size of the minimum replay buffer. hour:
[0145] The distributed soft value network update is based on the distributed soft Bellman operator. It calculates the pairwise temporal difference error between the target quantile and the predicted quantile, and minimizes this error using the quantile Huber loss function to update the value network parameters. Specifically:
[0146] First, strategy Corresponding soft action-value distribution satisfy:
[0147] (20)
[0148] In the formula, This represents the current continuous sub-action space value. For the entropy strategy of the algorithm, Indicates action By strategy Perform sampling. [0,1] is the adaptive discount factor. For adaptive temperature parameters;
[0149] Secondly, define the distributed soft Bellman operator. The input at this point is the reward distribution under the current state-action pair. The output is the new target distribution obtained after Bellman recursion:
[0150]
[0151] (twenty one)
[0152] In the formula, This indicates that the state of the next time slot is sampled from the probability distribution of the current state and the action. Indicates the next time slot from the strategy The sampling action is performed. This represents the reward distribution for the next time slot state-action pair. Indicates the next time slot strategy Obtain the action based on the state;
[0153] Since the distribution of returns is usually not explicitly represented, the algorithm approximates it using a finite number of quantiles. Specifically, it generates W+1 ordered quantile scores by dividing the distribution into equally spaced segments. ,in = < < < =1, the midpoint of each pair of consecutive quantiles is defined as calculated by formula (22):
[0154] (twenty two)
[0155] Next, calculate the first... The quantile and the th quantile quantiles ( Pairwise timing difference errors between :
[0156]
[0157] (twenty three)
[0158] In the formula, For the next time slot The target value network (whose parameters are updated more slowly and are used for stable training) outputs the first... The values of the location distribution, For the next time slot The target policy network is based on the policy The action of acquisition For the current time slot The value output of the network below is the first The values are distributed across the locations;
[0159] Calculate Huber quantile regression loss:
[0160] (twenty four)
[0161] (25)
[0162] In the formula, To calculate the threshold, Huber loss (when error) When the error is small, use the mean squared error (smooth and with a large gradient). When the value is large, use the mean absolute error (linear, to prevent gradient explosion). For quantile Huber loss, These are symmetric weighting coefficients, and their core function is to adjust the weighting coefficients based on the error. The positive or negative sign determines the severity of the punishment. It is an indicator function, if the error If the value is less than 0, its value is 1; otherwise, it is 0.
[0163] Therefore, the optimization objective of quantile numerical networks satisfies:
[0164] (26)
[0165] In the formula, The loss function of a value network;
[0166] Value network parameters By minimizing Perform gradient updates;
[0167] Step IV: The policy network generates actions using a reparameterization technique, updates the policy network parameters by maximizing entropy regularization to accumulate rewards, and maintains the target network parameters using a soft update method. This policy update uses the reparameterized policy network of formula (27). To achieve:
[0168] (27)
[0169] In the formula, Let the loss function be the policy network. For soft action-value functions, ( () is the expected value operator. This indicates that the state is sampled from the experience replay buffer M. It is a noise vector sampled from a fixed distribution;
[0170] The soft action-value function satisfies:
[0171] (28)
[0172] By minimizing Gradient descent to update parameters .
[0173] Step V: Maintain target network parameters using a soft update method:
[0174] (29)
[0175] (30)
[0176] in, To update parameters, , The old target network parameters before the update. , This is done to update the target network parameters, avoiding directly updating the target network to the current network, thus making the results more stable.
[0177] Step VI, if mod That is, the current time slot is divisible by the number of large language model update intervals. This framework employs a large language model-guided adaptive hyperparameter tuning mechanism, using the large language model as a meta-controller to achieve dynamic hyperparameter adjustment. It treats hyperparameter optimization as a sequential decision-making process, updating parameters based on recent learning performance to accelerate policy convergence and improve policy quality. Specifically, it includes:
[0178] Step 1-1, in the... During the second tuning event, the large language model receives a description of its current learning state, including the current configuration, the most recent reward sequence, training progress, and optional additional diagnostic statistics:
[0179] (31)
[0180] In the formula, for, For hyperparameter configuration, It represents the number of adjustable parameters. This represents the index of the hyperparameter update event, which varies depending on the specific implementation. This includes the learning rate for both critics and implementers, the entropy-temperature coefficient, the discount factor, the target network smoothness coefficient, the batch size, the number of quantile scores, and the exploration decay parameter. , and This represents the lower and upper bounds of different hyperparameters, thus limiting the range of different hyperparameters. It is a collection of recent A window of rewards for each iteration level (overall average). Indicates the normalized training progress;
[0181] Step 1-2, The information is converted into structured prompts containing the current configuration, recent performance trends, diagnostic statistics, allowed parameter ranges, and allowable adjustment ranges. These prompts are then used as parameter optimization cues within the large language model (LLM).
[0182] (32)
[0183] In the formula, A mapping function for generating parameters based on the current training state. Configure candidate hyperparameters for the output of the large language model;
[0184] In addition, to prevent invalid or overly aggressive adjustments, the proposed configuration is first processed by the constraint operator of formula (33) before being applied to the algorithm:
[0185] (33)
[0186] In the formula, For the first In the first iteration, the hyperparameters for actual application will be configured. Control the magnitude of each adjustment. This represents the feasible set of hyperparameters, while ( ) indicates projection onto the set;
[0187] For continuous hyperparameters, projection ( The following is given by formula (34):
[0188] (34)
[0189] In the formula, The independent variable for projection corresponds to specific parameters of the large language model. This represents the projection calculation over x;
[0190] Step VII: In each time slot, based on the UAV's three-dimensional position determined by the continuous variable optimization algorithm, and the actions to be performed by the UAV and the buoy and satellite determined by the discrete variable decision, the buoy data acquisition and satellite data upload tasks are completed simultaneously in full-duplex mode. Then, the current reward is obtained and the process is transferred to the next state. At the same time, the experience sample is stored in the experience replay pool.
[0191] Step VIII: Repeat the above steps until the maximum number of training rounds is reached to obtain the optimal flight position decision, optimal buoy selection strategy, and optimal satellite switching strategy for the UAV.
[0192] Step 4: Based on the optimal flight position decision, optimal buoy selection strategy, and optimal satellite switching strategy of the UAV, complete the dynamic anti-jamming communication at sea.
[0193] Therefore, the specific process of the LDSAC-AD algorithm described in this invention is as follows:
[0194]
[0195]
[0196] The LDSAC-AD algorithm described in this invention is compared with the original DSAC algorithm and benchmark deep reinforcement learning algorithms. Experimental results show that the LDSAC-AD algorithm has a faster convergence speed and better anti-interference security. Furthermore, this invention also comprehensively compares the reward function results, such as... Figure 3 As shown, the convergence performance curves of different algorithms are displayed. It can be seen that the convergence value of the LDSAC-AD algorithm is significantly better than that of the SAC, DDPG, TD3 and PPO algorithms, and the convergence speed is better than that of the standard DSAC algorithm.
[0197] Furthermore, Table 1 lists the comparison results of different algorithms in terms of UAV data reception volume and UAV flight energy consumption within a set of time slots. Among them, the LDSAC-AD algorithm performed best in maximizing the UAV data reception volume, and also achieved performance second only to DSAC and PPO algorithms in minimizing the total UAV flight energy consumption. This is because the LDSAC-AD algorithm sacrifices energy consumption to maximize transmission performance. Figure 3 It is evident that the LDSAC-AD algorithm has achieved a significant improvement in convergence speed, and this efficiency improvement further verifies the comprehensive superiority of the improved algorithm proposed in this invention. In summary, the UAV-assisted maritime dynamic anti-jamming communication method proposed in this invention, when using the LDSAC-AD algorithm, can effectively ensure the security, reliability, and efficiency of communication, and has significant theoretical value and engineering application prospects.
[0198] Table 1. Comparison results of different algorithms on two metrics.
[0199]
[0200] This invention presents a UAV-assisted dynamic anti-jamming communication method for the sea. Considering the dynamic tracking characteristics of jammers and the ping-pong effect that may result from frequent switching of low-Earth orbit satellites, it jointly optimizes the UAV's optimal flight path, optimal service buoy selection, and optimal upload satellite selection. This minimizes interference impact while ensuring efficient and reliable data transmission. Using the LDSAC-AD algorithm and an action space decoupling strategy, discrete decision variables (buoy selection and satellite selection) are separated from continuous decision variables (UAV trajectory), effectively solving the problems of dimensional explosion and exploration difficulties caused by the mixed action space. For discrete variables, a... - The greedy algorithm is further explored and utilized to optimize variable selection. For continuous variables, a hyperparameter optimization scheme based on a large language model is introduced. This scheme can dynamically adjust hyperparameters (entropy temperature coefficient, learning rate, etc.) according to the state information during training, which improves the algorithm's adaptability to different interference scenarios and significantly enhances the stability of the training process. This improves the overall solution quality and convergence efficiency of the strategy. It has high efficiency and superiority in terms of anti-interference performance, data collection efficiency, energy consumption control, convergence speed, and robustness.
[0201] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A method for dynamic anti-jamming communication at sea based on unmanned aerial vehicle (UAV) assistance, characterized in that, The steps include the following: Step 1: Collect information on the location of marine buoys, the orbital information of low-Earth orbit satellite constellations, the initial position of UAVs, and the initial position of unmanned surface vessels; Step 2: Discretize the entire task cycle into T equal-length time slots and construct the objective function for the optimization problem; The objective function of the optimization problem is to maximize the sum of the total reception rates of the maritime-UAV link and the UAV-satellite link. Step 3: Model the objective function as a Markov decision process, and use a decoupled action space strategy to process discrete and continuous variables respectively. Discrete variables are decided through an ε-greedy architecture, while continuous variables are decided through a distributed soft actor-commentator guided by a large language model, to obtain the optimal flight position decision, optimal buoy selection strategy and optimal satellite switching strategy for the UAV. Step 4: Based on the optimal flight position decision, optimal buoy selection strategy, and optimal satellite switching strategy of the UAV, complete the dynamic anti-jamming communication at sea.
2. The method for dynamic anti-jamming communication at sea based on unmanned aerial vehicle (UAV) assistance according to claim 1, characterized in that, The receiving rate of the maritime-UAV link satisfies: ; In the formula, For the current time slot, the first The drone received the first The data reception rate of each buoy. For channel bandwidth, Indicates the buoy's transmission power. Indicates the first The transmission power of an unmanned surface vessel Indicates noise power. For the current time slot, the first The buoy and the first Path loss between drones For the current time slot, the first The unmanned surface vessel and the first Path loss between drones.
3. The UAV-assisted dynamic anti-jamming communication method for maritime communication according to claim 2, characterized in that, The receiving rate of the UAV-satellite link satisfies: ; In the formula, For the current time slot, the first The satellite received the first The data reception rate of a drone. To indicate the first The transmission power of each drone, For the current time slot, the first The drone and the first Path loss between satellites.
4. The UAV-assisted dynamic anti-jamming communication method for maritime communication according to claim 3, characterized in that, The objective function of the optimization problem must satisfy the following constraints: (9) In the formula, For time slices The amount of data received by the drone from all the buoys. For time slices The total amount of data received by the satellite For time slices, For the drone's cache capacity, This refers to the entire task cycle.
5. The UAV-assisted dynamic anti-jamming communication method for maritime communication according to claim 4, characterized in that, The objective function of the optimization problem must also satisfy the following constraints: ; ; ; ; ; In the formula, For the receiving rate threshold, For data volume threshold, For single time slot duration, For the first Energy consumption during drone communication Energy consumption threshold For the number of satellite updates, This represents the maximum number of switching operations.
6. The UAV-assisted dynamic anti-jamming communication method for maritime operations according to claim 5, characterized in that, The Markov decision process includes: ; ; ; In the formula, This is the state space under the current time slot. and These represent the positions of the drone and the unmanned surface vessel in the current time slot, respectively. It is the buoy index in the current time slot. = This represents the cumulative data received from each buoy. It is the first in the current time slot Energy consumption of a drone = It is the set of positions of all low-Earth orbit satellites. and These are the index of the selected satellite and the number of satellite switching attempts, respectively. = This represents the cumulative amount of data uploaded to each satellite. For the action space in the current time slot, The reward function for the current time slot. , and Scaling factor This is a penalty item.
7. The UAV-assisted dynamic anti-jamming communication method for maritime operations according to claim 6, characterized in that, The ε-greedy architecture specifically includes the following steps: Step 1: Obtain the state of the previous time slot Actions selected in the previous time slot This corresponds to the target buoy B[t-1] or target satellite S[t-1] in the previous time slot, and the reward for the previous time slot is calculated. ; Step 2: Obtain the status of the current time slot At this point, directly change the state. Given a Q-table, iterate through all possible discrete sub-actions in the given state and find the action that maximizes the Q-value. ; Step 3: Calculate the error and update the Q value with the future estimated value; Step 4, through - Greedy algorithm updates discrete action space: ; In the formula, For parameter values, It is a random number. For the set of discrete action subspaces, As a heuristic exploration strategy, In the state Take action below The obtained Q value.
8. The UAV-assisted dynamic anti-jamming communication method for maritime operations according to claim 7, characterized in that, The introduction of a large language model-guided distributed soft actor-critic model specifically includes the following steps: Step I: Initialize the parameters of the distributed soft value network Policy network parameters Target value network parameters Target policy network parameters and experience playback buffer pool ; Step II: Initialize the initial positions of the drone and malicious jammer, and perform motion space sampling according to the strategy. Update the current location of the drone and malicious jammer, calculate the current reward r[t] based on the buoy index and satellite index of the current service, obtain the state space s[t+1] of the next time slot, and update the experience replay buffer M; Step III, if t mod And M Update the distributed soft value network; Step IV: Update the policy network; Step V: Maintain the target value network parameters using a soft update method. and target policy network parameters ; Step VI, if mod By using a large language model to guide an adaptive hyperparameter tuning mechanism, dynamic hyperparameter adjustment can be achieved.
9. The unmanned aerial vehicle-assisted dynamic anti-jamming communication method for the sea according to claim 8, characterized in that, The optimization objectives for steps III and V are: ; ; In the formula, The loss function of a value network, For quantile Huber loss, for The quantile and the th quantile Pairwise time-series difference error between quantiles , For the first quantiles, For the first quantiles, The loss function of a value network, For soft action-value functions, ( () is the expected value operator. This indicates that the state is sampled from the experience replay buffer M. It is a noise vector sampled from a fixed distribution. It is an adaptive temperature coefficient.
10. The UAV-assisted dynamic anti-jamming communication method for maritime communication according to claim 9, characterized in that, Step VI further includes: ; In the formula, For the first In the first iteration, the hyperparameters for actual application will be configured. For interpolation ratio, , This represents the feasible set of hyperparameters. ( This indicates projection onto a feasible set of hyperparameters. Configure candidate hyperparameters for the output of the large language model.