Unmanned aerial vehicle non-stop filling docking method based on depth enhanced predictive control algorithm
By employing a deep reinforcement predictive control algorithm, utilizing a probabilistic state-space model and a deep reinforcement learning strategy, the problem of docking window variation caused by the oscillation uncertainty of flexible hose joints under rotor was solved. This improved the success rate and robustness of uninterrupted refueling and docking for UAVs, and enabled adaptive parameter adjustment and safety control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-31
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN121763767A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) automatic control, and in particular to a method for unmanned aerial vehicle (UAV) refueling and docking without stopping, based on a deep reinforcement predictive control algorithm. Background Technology
[0002] With the increasing demands on endurance for drones in emergency rescue, border patrol, and long-distance transportation missions, non-stop refueling, a method that allows for rapid power replenishment without shutting down the power system, has attracted widespread attention due to its ability to significantly shorten support time and improve mission continuity. Existing non-stop refueling docking technologies typically employ airborne vision, laser ranging, or inertial navigation fusion to obtain the relative pose of the refueling connector and the drone's refueling interface, and achieve docking based on trajectory tracking control, servo control, or impedance control. Simultaneously, some research has introduced optimization control methods such as model predictive control to improve tracking performance and docking stability under input and state constraints. For complex wind fields and external disturbances, some solutions also attempt to improve the matching of control parameters through adaptive control or learning methods.
[0003] Existing technologies still have the following shortcomings in flexible hose connector docking scenarios:
[0004] 1. Most methods rely on rigid bodies or deterministic models for docking planning and control, which makes it difficult to characterize the random oscillation and uncertainty of flexible hose joints under rotor downwash excitation, resulting in unpredictable changes in the docking window and a high degree of influence on the docking success rate from the environment.
[0005] 2. Existing model-based predictive control schemes typically focus on instantaneous error constraints or deterministic constraints, lacking probabilistic constraints that address the event-driven objective of "the docking window persisting within the prediction time domain." This makes it difficult to make quantifiable risk trade-offs between safety and success rate.
[0006] 3. Control parameters are often set in a fixed manner or adjusted according to simple rules, making it difficult to adapt online to changes in swing intensity and window trend. When a learning strategy is introduced, there is often a lack of a feasible and safe gating mechanism, which may lead to the optimization problem becoming infeasible or generating unsafe control inputs.
[0007] Therefore, a method for unmanned aerial vehicle (UAV) refueling and docking without stopping is needed to address the shortcomings of the existing technology. This is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose a non-stop refueling and docking method for unmanned aerial vehicles (UAVs) based on a deep reinforcement predictive control algorithm. Addressing the challenges of accurately characterizing the swaying uncertainty of flexible hose joints under rotor downwash excitation, the low docking success rate due to random changes in the docking window, and the difficulty in balancing risk constraints in existing technologies, this invention proposes a technical solution that uses a probabilistic state-space model to estimate the relative docking state and predict the probability distribution of future relative poses, and calculates the docking window probability sequence based on docking criteria. Furthermore, the estimated relative docking state and the docking window probability sequence are input into a deep reinforcement learning strategy, which outputs online the risk budget allocation coefficients and the switching thresholds between the waiting and insertion phases of the chance-constrained model predictive controller. Boundary constraints and feasibility judgment backoffs are applied to these parameters through safety gating. A chance-constrained model predictive control problem incorporating the probability constraints of docking window events is constructed in rolling optimization to solve for the current control input. This invention achieves the technical effects of predicting the docking window in advance and quantifying risks under sway uncertainty conditions, improving docking success rate, enhancing control safety and robustness, and enabling adaptive parameter adjustment.
[0009] This invention provides a method for non-stop refueling and docking of unmanned aerial vehicles (UAVs) based on a deep reinforcement predictive control algorithm, comprising:
[0010] S1. Acquire the current observation data of the airborne sensors in the current control cycle, read the historical observation data and historical control inputs, and construct the time series input data in chronological order;
[0011] S2. Input the time series input data into the probabilistic state-space model to obtain the estimated docking relative state at the current moment;
[0012] S3. Based on the relative state estimation of docking, the probability distribution of the future relative pose prediction of the connector relative to the UAV docking interface is calculated in the preset prediction time domain using the probability state space model, and the docking window probability sequence is calculated from the future relative pose prediction probability distribution based on the preset docking criteria.
[0013] S4. Input the relative state estimate of docking and the probability sequence of docking window into the deep reinforcement learning strategy, and output the set of parameters of the opportunity constraint model predictive controller. The set of parameters includes risk budget allocation coefficient, waiting phase switching threshold and insertion phase switching threshold.
[0014] S5. Perform safety gating processing on the parameter set to obtain an executable parameter set. The safety gating processing includes: restricting the parameter set to a preset value range, and constructing an opportunity constraint model to predict the control problem based on the restricted parameter set for feasibility determination. If the feasibility determination result is infeasible, replace the restricted parameter set with a preset safety parameter set.
[0015] S6. Taking the relative state estimation of docking, the probability distribution of future relative pose prediction, the probability sequence of docking window, and the set of executable parameters as inputs, perform rolling optimization of the chance constraint model predictive control to obtain the control input for the current control cycle.
[0016] S7. Send the control input to the UAV flight control system for execution and enter the next control cycle.
[0017] Optionally, S1 includes:
[0018] At the sampling moment of the current control cycle, the current observation data of the airborne sensor is acquired. The current observation data includes the relative pose observation of the connector relative to the UAV docking interface and the UAV's own motion state observation.
[0019] Read historical observation data of a preset length corresponding to the sampling time and historical control inputs of a preset length from the buffer.
[0020] Time alignment is performed on the current observation data, the historical observation data, and the historical control input;
[0021] The current observation data, after time alignment, is concatenated with the historical observation data in chronological order to form an observation time series, and the historical control inputs are concatenated in chronological order to form a control input time series;
[0022] The observed time series and the control input time series are used together as the time series input data.
[0023] Optionally, S2 includes:
[0024] Input the time series data into the probabilistic state-space model;
[0025] The state prediction is performed based on historical control inputs by the probabilistic state-space model, and the state is updated based on the current observation data to obtain the estimated docking relative state at the current moment.
[0026] The docking relative state estimation includes the relative position, relative attitude, and swing state parameters of the connector relative to the UAV docking interface, and the docking relative state estimation includes estimated values of the relative position, the relative attitude, and the swing state parameters, as well as uncertainty representations corresponding to the estimated values.
[0027] Optionally, S3 includes:
[0028] Based on the relative state estimation of docking, the probability state space model is used to calculate the future relative pose prediction probability distribution of the joint relative to the UAV docking interface by time step recursively in the preset prediction time domain.
[0029] For each time step within the preset prediction time domain, the docking feasible region is determined according to the docking criterion. The docking feasible region is a set of future relative poses that simultaneously satisfy the following conditions: relative position error is less than the relative position error threshold, relative attitude error is less than the relative attitude error threshold, and relative velocity is less than the relative velocity threshold.
[0030] At each time step, the probability that the future relative pose falls into the docking feasible region is calculated based on the corresponding future relative pose prediction probability distribution, and the probability corresponding to each time step constitutes the docking window probability sequence.
[0031] Optionally, S4 includes:
[0032] The relative state estimation of docking and the probability sequence of docking window are input into the deep reinforcement learning strategy;
[0033] The deep reinforcement learning strategy outputs a set of parameters for the chance-constrained model predictive controller based on the joint swing state represented by the relative docking state estimate and the docking window change trend represented by the docking window probability sequence.
[0034] The parameter set includes a risk budget allocation coefficient for determining the event probability constraint of the docking window, a waiting phase switching threshold for transitioning from the waiting phase to the insertion phase, and an insertion phase switching threshold for transitioning from the insertion phase to the insertion phase.
[0035] Optionally, S5 includes:
[0036] Boundary constraint processing is performed on the parameter set to ensure that the risk budget allocation coefficient, the waiting phase switching threshold and the insertion phase switching threshold in the parameter set fall into their respective preset value ranges, thus obtaining the parameter set after boundary constraint.
[0037] Based on the parameter set after the boundary constraints, a chance-constrained model predictive control problem is constructed and a feasibility determination is performed. The feasibility determination includes determining whether there is a solution to the chance-constrained model predictive control problem that satisfies the input constraints and the state constraints.
[0038] If the feasibility assessment result is feasible, the parameter set after the boundary constraints will be used as the executable parameter set.
[0039] When the feasibility assessment result is infeasible, the preset safety parameter set is determined as the executable parameter set, wherein the preset safety parameter set is a set of parameters that are pre-set and meet the corresponding preset value range.
[0040] Optionally, S6 includes:
[0041] Using the relative state estimation of docking as the initial state, the set of executable parameters as the control parameters of the chance-constrained model predictive controller, and combining the probability distribution of future relative pose prediction and the probability sequence of docking window, a chance-constrained model predictive control problem is constructed, which includes cost function, input constraints, state constraints and docking window event probability constraints.
[0042] The probability constraint of the docking window event includes: within the preset prediction time domain, the probability that there exists a period of time that satisfies the docking criterion and has a duration of not less than the preset duration is not lower than the probability lower limit determined by the risk budget allocation coefficient.
[0043] At each rolling optimization time, the chance-constrained model predictive control problem is solved to obtain the optimal control input sequence, and the first item in the optimal control input sequence is selected as the control input for the current control cycle.
[0044] Optionally, the S7 includes:
[0045] The control input is converted into execution commands for the UAV flight control system and sent to the UAV flight control system, so that the UAV flight control system controls the attitude and position of the UAV according to the execution commands within the current control cycle;
[0046] At the start of the next control cycle, the current observation data of the airborne sensors is acquired again and the historical control input is updated, thereby returning to step S1 to execute the next rolling optimization control, so as to achieve non-stop refueling docking between the UAV and the refueling device.
[0047] Optionally, the probabilistic state-space model is a state-space model that includes external excitation inputs. The external excitation inputs include one or more of the rotor speed, pitch, or thrust command that characterize the rotor downwash intensity. The external excitation inputs are used as known inputs to update the oscillation state parameters during state prediction.
[0048] Optionally, the oscillation state parameters include one or more of the following: oscillation angle, oscillation velocity, equivalent pendulum length, and equivalent damping coefficient, and the probabilistic state space model updates the equivalent pendulum length and / or equivalent damping coefficient through online parameter identification.
[0049] Optionally, the probability constraint of the docking window event "the duration is not less than the preset duration" is implemented by discretization. Discretization includes: converting the preset duration into a continuous number of prediction steps K, and defining the docking window event as the event in which the future relative poses of K consecutive prediction steps fall into the docking feasible region within the preset prediction time domain.
[0050] Optionally, the risk budget allocation coefficients output by the deep reinforcement learning strategy are risk budget sequences allocated along the prediction time domain. The risk budget sequences satisfy that the sum of the risk budgets at each time step is equal to the preset total risk budget, and are used to determine the probability lower bound of each time step or each event.
[0051] The beneficial effects of this invention are:
[0052] 1. Improve docking success rate and window capture capability: The state estimation and future relative pose probability prediction of the flexible hose joint swing are performed by using a probabilistic state space model, and the docking window probability sequence is calculated from this. This allows the controller to make forward planning before the window appears, thereby improving the capture efficiency of the docking window and the final docking success rate.
[0053] 2. Achieve quantifiable risk constraints and stronger robustness: In the rolling optimization of model predictive control, the probability constraint of docking window event is introduced, and the probability lower limit is determined by the risk budget allocation coefficient. This enables the system to make a quantifiable trade-off between safety and success rate under input constraints and state constraints, thereby improving robustness and stability under rotor downwash excitation and uncertain oscillation conditions.
[0054] 3. Online Adaptation and Safety Feasibility Assurance: A deep reinforcement learning strategy is adopted to adjust the chance constraint model to predict controller parameters online, and boundary restrictions and feasibility judgment backoff are performed through safety gating to avoid the optimization problem becoming infeasible or the control input becoming unsafe due to improper learning parameters. This achieves parameter adaptation while ensuring control execution and system safety. Attached Figure Description
[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0056] Figure 1 This is a flowchart of a non-stop refueling and docking method for unmanned aerial vehicles (UAVs) based on a deep enhancement predictive control algorithm, as proposed in this invention. Detailed Implementation
[0057] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0058] refer to Figure 1 A method for unmanned aerial vehicle (UAV) refueling and docking without stopping, based on a deep reinforcement predictive control algorithm, includes:
[0059] S1. Acquire the current observation data of the airborne sensors in the current control cycle, read the historical observation data and historical control inputs, and construct the time series input data in chronological order;
[0060] S2. Input the time series input data into the probabilistic state-space model to obtain the estimated docking relative state at the current moment;
[0061] S3. Based on the relative state estimation of docking, the probability distribution of the future relative pose prediction of the connector relative to the UAV docking interface is calculated in the preset prediction time domain using the probability state space model, and the docking window probability sequence is calculated from the future relative pose prediction probability distribution based on the preset docking criteria.
[0062] S4. Input the relative state estimate of docking and the probability sequence of docking window into the deep reinforcement learning strategy, and output the set of parameters of the opportunity constraint model predictive controller. The set of parameters includes risk budget allocation coefficient, waiting phase switching threshold and insertion phase switching threshold.
[0063] S5. Perform safety gating processing on the parameter set to obtain an executable parameter set. The safety gating processing includes: restricting the parameter set to a preset value range, and constructing an opportunity constraint model to predict the control problem based on the restricted parameter set for feasibility determination. If the feasibility determination result is infeasible, replace the restricted parameter set with a preset safety parameter set.
[0064] S6. Taking the relative state estimation of docking, the probability distribution of future relative pose prediction, the probability sequence of docking window, and the set of executable parameters as inputs, perform rolling optimization of the chance constraint model predictive control to obtain the control input for the current control cycle.
[0065] S7. Send the control input to the UAV flight control system for execution and enter the next control cycle.
[0066] In this specific embodiment, S1 includes:
[0067] The controller uses the flight control clock or an equivalent monotonic clock as the unified time reference for the entire system, and records the sampling time of the current control cycle as... ,in Indicates the first Sampling timestamps for each control cycle This indicates the control cycle number, and the control cycle duration is denoted as... ,in Indicates the time interval between adjacent sampling times;
[0068] At sampling time Read the current observation data from the airborne sensors and form the current observation vector. The current observation vector includes at least the relative pose observation of the connector with respect to the UAV docking interface and the UAV's own motion state observation, organized according to a unified dimension and coordinate system as follows:
[0069] ;
[0070] in, Indicates the sampling time The current observation vector, This represents the relative positional observation of the connector with respect to the UAV docking interface, and its components are expressed in a pre-agreed docking interface coordinate system. The relative attitude observations of the connector with respect to the UAV docking interface are represented using unit quaternions to avoid Euler angle singularities. This represents the observed linear velocity of the UAV itself, and its components are expressed in a coordinate system consistent with the docking interface coordinate system or a coordinate system obtained through transformation of known external parameters. This represents the observation of the UAV's own angular velocity, and its components are expressed in a coordinate system consistent with the above.
[0071] Among them, the relative pose observation can be obtained by airborne visual measurement, laser ranging and attitude calculation, or multi-source fusion positioning output. The UAV's own motion state observation can be obtained by the output of the inertial measurement unit and the flight control state estimator. In order to avoid the mixing of coordinates in subsequent time series modeling, all observations are made consistent with the coordinate system and the unit before entering the buffer area. Threshold removal and saturation pruning are performed on obvious outliers to ensure the numerical stability of the observation sequence.
[0072] To ensure the traceability of "historical observation data" and "historical control inputs," the controller is equipped with a circular buffer to store the most recent data. The observation data and control input for each control cycle are recorded, and a timestamp is appended to each record. The number of sampling points, representing a preset length and pre-set to cover at least one swing characteristic time scale based on the swing frequency and control response requirements, is defined as follows: Historical control input is the control input record sent to the UAV flight control system in the previous control cycle. The control input can be any one of position and velocity commands, acceleration commands, or attitude and thrust commands, and in this embodiment, it is uniformly referred to as the control input vector. So that it can be used in conjunction with observation sequences;
[0073] Reading from the buffer and sampling time Corresponding historical observation data and historical control input Subsequently, time alignment is performed to eliminate time misalignments introduced by multi-sensor asynchrony, communication jitter, and different update rates, said time alignment to control the periodic time grid. As the target time axis, the original timestamped data is resampled onto the target time axis to form an aligned sequence. The alignment rules are as follows: for Euclidean space quantities such as position, linear velocity, and angular velocity, time-stamp-based linear interpolation or zero-order hold is used to generate aligned values; for quaternion attitude quantities, spherical linear interpolation is used to maintain the unit norm and avoid interpolation drift; for control inputs, zero-order hold consistent with flight control execution is used to maintain the physical meaning of "commands are constant within the period"; when there is a missing observation at a certain historical moment, the zero-order hold of the most recent valid observation is used and the missing flag is written synchronously so that subsequent models can selectively ignore or reduce the weight of that time step.
[0074] After time alignment is completed, the aligned current observation data and the aligned historical observation data are concatenated in chronological order to form an observation time series. The aligned historical control inputs are then concatenated in chronological order to form a control input time series. Then, both are used together as time series input data. The output is represented as:
[0075] ;
[0076] in, Represents the observed time series, This indicates the control input time series. This represents time series input data. This represents the time-aligned observation vector. This represents the time-aligned control input vector, and... The current observation data is explicitly included to ensure that the end of the sequence corresponds to the current control period, so that the probabilistic state space model in step S2 can simultaneously utilize the "correction information of the current observation" and the "dynamic context information of historical observation and historical control" to complete the docking relative state estimation under a unified time grid.
[0077] In this specific embodiment, S2 includes:
[0078] The controller receives time-series input data. ,in Indicates the first One control cycle is used to estimate the timing input data. This represents the time series of observations after time alignment. This represents the time-aligned control input time series. This represents the aligned observation vector, which contains at least the relative pose observation and the UAV's own motion state observation. This represents the aligned control input vector. The historical length is represented and preset by the system. The controller constructs a discrete-time probabilistic state-space model and uses the "dock-up relative state" as the estimated state. In this embodiment, the dock-up relative state is denoted as... ,in Indicates the sampling time The relative docking status, This indicates the relative position of the connector with respect to the UAV docking interface. The relative attitude of the connector with respect to the UAV docking interface is represented by a unit quaternion. This represents a oscillation state parameter vector, including at least the oscillation angle and angular velocity, and optionally further including the equivalent oscillation length and equivalent damping coefficient. In implementation... ,in Indicates the swing angle. Indicates the angular velocity of the pendulum. Indicates the equivalent pendulum length. Indicates the equivalent damping coefficient;
[0079] The probabilistic state-space model consists of a state transition model and an observation model, and is used to provide probabilistic calculations for "state prediction" and "state update," specifically expressed as follows:
[0080] ;
[0081] in, Indicates a set of parameters Parameterized state transition function At least includes sampling period With gravitational acceleration And in this embodiment, It can also include the discretization coefficients of the oscillation model and the excitation coupling gain. This indicates the alignment control input from the previous control cycle. This represents the external excitation input and is used to characterize the rotor downwash intensity. It can be constructed from rotor speed or thrust commands. This represents process noise and is used to characterize unmodeled flexibility effects and random disturbances, assumed to be zero-mean Gaussian noise with a covariance matrix of... In the engineering implementation, a diagonal matrix is selected, and its diagonal elements are set according to the random walk strength of the pendulum angle and the position drift tolerance. Indicates a set of parameters Parameterized observation function It includes the sensor extrinsic calibration parameters corresponding to relative pose observation and can be represented as a fixed combination of rotation matrix and translation vector to ensure that the observed coordinate system is consistent with the docking interface coordinate system. This represents observation noise and is used to characterize visual measurement or ranging error. It is assumed to be zero-mean Gaussian noise with a covariance matrix of... , The sensor accuracy index is given offline or calibrated online, and a diagonal matrix is selected.
[0082] Among them, the state transition function The structure adopts a combination of "relative rigid body kinematics + oscillation evolution", specifically implemented by... Based on the previous cycle's control input and the UAV's own motion state observations, discrete integration is performed to obtain the prior relative position. Quaternion recursion is performed based on small angle increments or exponential mappings, and normalization is performed after each recursion to preserve... , swing parameters Based on the damped pendulum model and superimposed with external excitation The influence is updated discretely, and the effect is... and Apply a projection of the physically feasible range to avoid non-physical values, where the equivalent pendulum length... Choice restrictions and Equivalent damping coefficient Choice restrictions and The aforementioned boundaries are determined prior to the hose structure dimensions and material damping range before installation.
[0083] Observation function The construction is from Extract and output the relative position and relative attitude components that are directly observable by the sensor to correlate with... The relative pose observations are compared for consistency, while the UAV's own motion state observations are used as known inputs in this embodiment. The recursion is used to improve prediction accuracy;
[0084] In the specific filtering solution, the controller selects unscented Kalman filtering or extended Kalman filtering to implement the recursive calculation of "prior prediction - observation update", and utilizes... The historical control inputs are used to progressively predict the state to obtain the prior distribution at the current moment, and then the current observations are used. The state update is completed to obtain the posterior distribution, thereby outputting the docking relative state estimate containing the estimated value and the uncertainty representation. The posterior distribution is expressed under the Gaussian approximation as follows:
[0085] ;
[0086] in, This indicates input data in a given time series. Relative state of docking under certain conditions The posterior probability distribution, Indicates the family of Gaussian distributions. Represents the mean of the relative state estimates of the docking and includes at least the following: , and Indicates and The corresponding covariance matrix is output as an uncertainty representation for use in step S3 for probability prediction and window probability calculation. Furthermore, in numerical implementation, this is achieved by... Symmetry and minimum eigenvalue truncation are performed to ensure its positive semidefiniteness and avoid filter divergence.
[0087] In this specific embodiment, S3 includes:
[0088] The controller is based on docking relative state estimation. and the state transition function of the probabilistic state-space model The preset prediction time domain is recursively predicted step by step, where Indicates the sampling time The mean of the relative state estimate of the docking, Indicates and The corresponding uncertainty covariance matrix, Indicates the first The sampling time of each control cycle, the prediction step size is the sampling cycle used in step S1. The prediction time domain length is set to Each prediction step is set offline by the system to cover the look-ahead time required for docking operations;
[0089] To explicitly propagate uncertainty into a "probability distribution for predicting future relative pose," this implementation method employs Monte Carlo particle recursion, i.e., in... Time from Gaussian posterior approximation Extraction The initial value is _ particles, denoted as _. ,in Indicates a given time series input data Posterior distribution of the relative docking state under the given conditions This represents time series input data. Indicates a Gaussian distribution. Represents the number of particles and selects... To strike a balance between computational load and accuracy, Indicates the particle number;
[0090] For each particle, future state particles are generated step-by-step in the prediction time domain, and an empirical probability distribution is formed from the particle set. The recursion adopts the same state transition model as step S2, and process noise and known input are explicitly injected, specifically:
[0091] ;
[0092] in, Indicates the first The prediction step corresponding to the th prediction step A future docking relative state particle Indicates the prediction step number. Indicates a set of parameters The parameterized state transition function is constructed in the same way as step S2 and includes the discrete evolution relationship between relative pose and oscillation state. Indicates the first one used for prediction The control input is controlled step by step, and in the implementation, either the zero-order hold of the control input executed in the previous control cycle or the nominal control sequence of the control output predicted by the model in the previous round is used to ensure that the predicted input is available. This represents the external excitation input term used for prediction, consistent with step S2, used to characterize the rotor downwash intensity, and constructed from rotor speed or thrust commands. It is also predicted using zero-order hold or nominal sequence. Indicates the first The particle in the first Step-by-step injection process noise and according to sampling, Represents the zero vector. This represents the process noise covariance matrix set in step S2 and is used to characterize the unmodeled flexibility effect and the intensity of random disturbances;
[0093] After the recursion is completed, the controller performs each prediction step. From the set of particles Extract random variables of future relative pose and form a "probability distribution for future relative pose prediction", where the future relative pose includes random variables of the future relative position of the connector relative to the UAV docking interface. relative to future attitude random variables The relative position components in the state particles are read directly and expressed in the docking interface coordinate system. The relative attitude quaternion components in the state particles are read directly and normalized after each recursion to maintain [the state]. To achieve the relative velocity criterion, the controller uses a finite difference between the relative positions of two adjacent steps for each particle and divides it by the sampling period. Obtain the future relative velocity vector Furthermore, its modulus length is taken as the relative velocity magnitude;
[0094] In terms of constructing docking criteria, the relative positions of the docking targets are pre-defined. attitude relative to the docking target ,in This indicates the target relative position when the connector and the mating interface are perfectly aligned, and is selected in the implementation selection. The target relative attitude quaternion represents the target relative attitude when fully aligned and is determined by assembly calibration;
[0095] And preset the relative position error threshold Relative attitude error threshold With relative velocity threshold ,in Based on the allowable lateral deviation of the interface-guided structure and the visual measurement error margin, values were selected within the centimeter range. The angle tolerance of the inserted structure is set and selected within a certain range of degrees. The upper limit of the safe relative speed during the insertion phase is set and a value is selected within the low-speed range.
[0096] Based on this, a docking feasible region is defined for each prediction step. The set of future relative poses that simultaneously satisfy the following three conditions, i.e., relative position error Not greater than The relative attitude error is no greater than Furthermore, the relative attitude error is determined by the relative attitude quaternion. and The quaternion difference is calculated and its equivalent minimum rotation angle is taken as the error measure, relative velocity magnitude. Not greater than ;
[0097] Then at each prediction step Based on the corresponding probability distribution of future relative pose prediction, the probability of "future relative pose falling into the docking feasible region" is calculated. The probability of "" is used to construct a docking window probability sequence, and the probability estimation is performed using particle ratio to avoid high-dimensional integration. Specifically:
[0098] ;
[0099] in, Indicates the first The docking window probability corresponding to each prediction step This represents the summation operation over particle numbers. This indicates an indicator function that takes the value 1 when the condition within the parentheses is true, and otherwise takes the value _____. Indicates the first The particle in the first The future relative position, future relative attitude, and future relative velocity corresponding to each prediction step simultaneously satisfy the docking criteria and thus fall into the docking feasible region. ;
[0100] Each prediction step Arranged chronologically to form a probability sequence of docking windows. The predicted probability distribution of the future relative pose formed step by step is output to steps S4 and S6 for subsequent policy adjustment and opportunity constraint model predictive control rolling optimization.
[0101] In this specific embodiment, S4 includes:
[0102] The relative state estimate of docking and the probability sequence of docking window are used together to construct the input of the deep reinforcement learning policy, and the deep reinforcement learning policy outputs the set of parameters of the opportunity constraint model predictive controller online.
[0103] The mean and uncertainty characterization of the docking relative state estimation are performed, where the mean of the docking relative state estimation is denoted as... It includes at least a relative position estimate. Relative attitude estimation And oscillation state parameter estimation Uncertainty is characterized using the covariance matrix. diagonal elements Or its equivalent scale quantity is used as input feature to reflect the estimation reliability of position, attitude, and sway parameters, where This represents the operator that retrieves the elements of the main diagonal of a matrix. Represent the posterior covariance matrix;
[0104] Let the docking window probability sequence be denoted as... ,in This indicates that the prediction time domain length is The column vector formed by the probabilities of each prediction step's docking window. Indicates the first The probability that the future relative pose falls into the docking feasible region for each prediction step. Indicates the preset prediction steps;
[0105] To ensure the numerical stability and cross-scenario generalization capability of online inference, the controller... and A unified normalization process is performed, and the normalization scale is selected by choosing the docking criterion threshold as the physical dimension benchmark, such as using the relative position error threshold. right Perform component normalization using a relative attitude error threshold. Scale constraints are applied to attitude-related features using a relative velocity threshold. Scale constraints are applied to the velocity quantities or their statistics derived from the oscillating state, and the diagonal elements of the covariance are truncated and logarithmically compressed to avoid input explosion caused by extreme uncertainty.
[0106] Based on this, construct the policy input vector. And input it into a deep reinforcement learning policy network. Get the parameter set Its form is:
[0107] ;
[0108] in, Indicates the first The policy input vector for each control cycle This represents the estimated relative position of the connector with respect to the UAV docking interface. This represents the relative attitude estimate of the connector with respect to the UAV docking interface, and is a unit quaternion. This represents the estimated values of the oscillation state parameters, including at least the oscillation angle and angular velocity, and optionally the equivalent oscillation length and equivalent damping coefficient. The diagonal characterization represents the uncertainty in the relative state estimation of docking. Represents the probability sequence of the docking window. This represents the set of parameters for the chance-constrained model predictor of the deep reinforcement learning policy output. Indicated by network parameters The parameterized policy function operates in a deterministic forward inference manner during the online phase;
[0109] The parameter set Including risk budget allocation coefficient Waiting phase switching threshold and the insertion phase switching threshold ,in This represents the risk budget sequence allocated along the prediction time domain and is used in step S6 to construct the probability lower bound or equivalent opportunity constraint tightness allocation for the docking window event probability constraints. This represents the threshold for transitioning from the waiting phase to the insertion phase, and is used to compare the "comprehensive dockability of the window probability sequence" with the threshold to trigger entry into the insertion phase. This represents the switching threshold for exiting the insertion phase and is used to trigger exiting the insertion phase and returning to the waiting phase or safety maneuver when the window probability drops rapidly or uncertainty increases;
[0110] In implementing the risk budget allocation coefficients, to ensure that the risk budget at each time step is non-negative and that the total risk budget is conserved, the policy network chooses to first output the unnormalized risk score vector. and through Normalization yields The calculation is as follows:
[0111] ;
[0112] in, This represents the sequence of risk budget allocation coefficients. This indicates that the preset total risk budget is set offline by the system and meets the requirements. , This represents a function that exponentially normalizes a vector so that the output components are non-negative and their sum is 0. This represents the unnormalized risk score vector output by the policy network;
[0113] The policy network structure includes a front-end fully connected feature extraction module for extracting swing intensity and uncertainty level, and a sequence feature extraction module for extracting window change trends. The sequence feature extraction module can employ one-dimensional convolutions or gated recurrent units. Encode the data to obtain trend characteristics that reflect "window rising, plateauing, and falling", and then combine them with... and The concatenated feature vectors are input into a multilayer perceptron and output. and and to and Use a bounded activation function to map to the expected dimensional range so that subsequent step S5 can perform safe gating boundary constraints.
[0114] The deep reinforcement learning strategy obtains network parameters through offline training before deployment. The offline training environment selection includes randomization factors such as rotor downwash intensity, hose equivalent swing length, equivalent damping coefficient, and measurement noise level to cover actual working conditions. The reward function selection simultaneously includes terms that increase the docking window capture probability and reduce drastic changes in control input, and imposes penalties for violations of safety boundaries or triggering infeasible solutions, thereby improving the online output. and It can adaptively adjust to the joint's swing state and the changing trend of the docking window.
[0115] In this specific embodiment, S5 includes:
[0116] For parameter set Perform security gating to obtain the set of executable parameters. ,in Indicates the first The set of controller parameters, including risk budget allocation coefficients, is predicted by the opportunity-constrained model of the deep reinforcement learning policy output for each control cycle. Waiting phase switching threshold and the insertion phase switching threshold Indicates the length along the prediction time domain The allocated risk budget sequence, This represents the switching threshold for transitioning from the waiting phase to the insertion phase. This indicates the switching threshold for exiting the insertion phase;
[0117] The safety gating system first performs boundary restriction processing to ensure that each parameter falls within its corresponding preset value range. This preset value range is given by offline engineering calibration and is fixed in the controller parameter table, where the total risk budget is denoted as... And satisfy and take The lower and upper bounds of a single-step risk budget are denoted as follows: and And satisfy and take and The lower and upper bounds of the waiting phase switching threshold are denoted as follows: and Furthermore, the lower and upper bounds of the insertion phase switching threshold are denoted as follows: [Insert lower bound here]. and It is also used to constrain the sensitivity of the policy exiting the insertion phase;
[0118] In implementation, the boundary constraints are obtained by using a projection method of "pruning + risk budget renormalization" to obtain the parameter set after boundary constraints. The calculation is as follows:
[0119] ;
[0120] in, This indicates the threshold for switching to the waiting phase after boundary restrictions. This indicates the threshold for switching during the insertion phase after boundary constraints. This represents the sequence of risk budget allocation coefficients after boundary constraints and renormalization. Indicates scalar Perform range clipping and output the range that falls into the range. As a result, Represents a vector Perform interval clipping based on elements and output the results of each element falling into its corresponding interval. The dimension is A single column vector, This indicates the transpose operation. This represents summing the elements of the clipped risk budget vector to achieve renormalization. express The One component;
[0121] After boundary constraints are met, the controller is based on Constructing a chance-constrained predictive control model for feasibility assessment ,in The same prediction time domain length is used as that of the chance-constrained predictive control problem actually solved in step S6. They share the same input and state constraints, the same docking window event probability constraints, and their initial conditions are taken as the mean of the relative docking state estimates. The docking window probability sequence is used during constraint evaluation. Scene samples with a probability distribution of future relative pose prediction, wherein the scene samples use Monte Carlo particles consistent with step S3 and select from them Representative particles are used to transform probabilistic constraints into finite scene constraints for rapid determination, where This represents the number of scenarios for feasibility determination and the selection criteria. To meet real-time requirements;
[0122] The specific implementation of feasibility assessment is in time budgeting. Internal call to the optimization solver consistent with step S6 Perform a solution attempt and read the solver's return status. When the solver returns that there exists a solution that satisfies all input constraints and state constraints, and also satisfies the conditions set by the solver... A solution is considered feasible if it meets the lower bound of the probability of the defined docking window event; otherwise, it is considered feasible if the solver returns infeasible, the numerical value is abnormal, or a timeout occurs. All scenarios are deemed infeasible, and to reduce the false positive rate, the optimal solution from the previous control cycle is used as the initial value for hot start. An additional conservative retry is performed when an infeasible scenario is identified. This conservative retry reduces the total risk budget. Alternatively, this can be achieved by increasing the constraint tightening margin;
[0123] When the feasibility assessment result is feasible, Directly identified as the set of executable parameters When the feasibility assessment result is infeasible, the preset safety parameter set will be used. Determined as the set of executable parameters ,in A set of parameters pre-set offline that meets the above-mentioned preset value range is selected, and a uniform risk allocation is chosen. and conservative stage threshold combinations and ,in This represents a conservative margin for the waiting phase threshold and is given by engineering calibration. This represents a conservative margin for the threshold in the insertion phase and is given by engineering calibration, thereby ensuring that the controller can still enter step S6 to perform rolling optimization with feasible and safe parameters when the learning output is unreliable or makes the optimization problem infeasible.
[0124] In this specific embodiment, S6 includes:
[0125] At sampling time Relative state estimation of receiver docking Future relative pose prediction probability distribution and docking window probability sequence and the set of executable parameters ,in This represents the mean of the relative state estimates for docking, and includes at least the relative position estimates. With relative attitude estimation The posterior covariance matrix is used to characterize uncertainty. Indicates the first The probability that the "future relative pose falls into the docking feasible region" in each prediction step. Indicates the preset prediction steps. This represents the sequence of risk budget allocation coefficients distributed along the forecast time domain. and These represent the switching thresholds for the waiting phase and the insertion phase, respectively.
[0126] The controller first bases its decisions on the docking window probability sequence. Forming "window trend volume" and , The comparison determines the stage mode adopted for the current rolling optimization, and the window trend quantity is selected as follows: And its position within the prediction domain is used to characterize "whether the window is about to appear and when it will appear," when the window trend exceeds Enter the insertion phase and set the control target to docking insertion. When the window trend value is lower than Exit the insertion phase and return to the waiting phase to avoid forced insertion when the window decays or uncertainty increases;
[0127] After determining the phase mode, the controller constructs a nominal predictive model for chance-constrained predictive control based on relative position and relative velocity, and solves for the control input of the current control cycle in rolling optimization, where the nominal state vector is denoted as... , This indicates the nominal state of relative translation. This indicates the relative position of the connector with respect to the UAV docking interface. The relative velocity is the derivative of relative position with respect to time, and the sampling period is denoted as . nominal initial value ,in Indicates at time The initial state of scroll optimization Represents the estimated relative velocity and in When not given directly Compared with the previous cycle Finite difference is used and first-order low-pass filtering is applied to suppress measurement noise;
[0128] Control input is denoted as It is defined as a three-axis acceleration command or equivalent control variable that the flight control system can track, and the decision variable within the prediction domain is denoted as... ,in Indicates at time The first planning For each prediction step of the control input, the controller constructs and solves the following rolling optimization problem:
[0129] ;
[0130] in The nominal forecast of the first The relative positions of the prediction steps are as follows: Position components, The nominal forecast of the first Each predicted step state This indicates the relative reference position related to the stage and takes the preset waiting position during the waiting stage. During the insertion phase, the docking target position is obtained. Representing a quadratic form Represents the position error weight matrix and takes And order To make the weights relative to the relative position error threshold Correlation of the same dimensions Indicates that the control energy weight matrix is taken And order With the maximum permissible control amplitude Related, This represents the terminal position weight matrix and is used to enhance terminal convergence. and These represent the sampling periods respectively. Construct the discrete double integral nominal model matrix and choose to take express identity matrix Represents the zero matrix. This represents the set of input constraints, given by the flight control executable boundaries, and selected as component-wise interval constraints to limit the magnitude and rate of change of acceleration or equivalent control quantities. This represents a set of state constraints constructed from safety distance, speed limits, and workspace boundaries, and includes a more stringent relative speed limit for the insertion phase to reduce collision risk. Indicates the probability of events in the docking window. This represents the lower probability limit determined by the risk budget allocation coefficient;
[0131] The event probability constraint of the docking window is modeled for "the existence of dockable time periods with a duration not less than a preset duration within the prediction time domain", where the preset duration is denoted as . And discretize it into continuous prediction steps ,in This represents the minimum number of consecutive dockable steps. This represents the floor operator, and the controller is based on the docking window probability sequence. Constructing event probabilities and the lower limit of risk The comparison is made, where the lower limit of risk is taken as... Make the risk budget allocation coefficient Corresponding to the "total permissible default risk", the event probability is calculated using the following formula under the conditional independence approximation:
[0132] ;
[0133] in This represents the chain multiplication operator. Indicates the starting prediction step index of the continuous window. Indicates the step index within the window. Indicates the first in the window The probability that a prediction step falls into the docking feasible region. Indicates "from the first Step start continuous The probability estimate of "all steps can be connected" Represents "all possible continuums" The probability estimate of "none of the step windows satisfy" is obtained by subtracting 1 from the quantity to obtain "at least one condition satisfies the continuous condition". Estimating the event probability of "step-by-step dockable window";
[0134] In numerical implementation, the controller... Perform lower bound truncation to avoid occurrence This causes the multiplicative value to underflow, and only occurs during the insertion phase if... Solving the optimization problem in the insertion phase is allowed if the condition is met; otherwise, a forced return to the reference in the waiting phase is imposed. With a more conservative set of constraints to ensure safety;
[0135] The optimal solver is selected using a quadratic programming solver or a sequential quadratic programming solver, with the optimal solution from the previous control cycle used as a warm start to meet real-time requirements, and the optimal control input sequence is obtained. Then select the first option As the control input for the current control cycle, the output is sent to step S7. At the same time, in the next control cycle, the observations and estimates are updated and the above rolling optimization process is repeated to form a closed-loop control of chance-constrained model predictive control.
[0136] In this specific embodiment, S7 includes:
[0137] Receive control input for the current control cycle ,in Indicates at the sampling time The control input to be issued needs to be selected as a three-axis acceleration command. Indicates the first Sampling time of each control cycle;
[0138] To ensure consistency with the command interface of the UAV flight control system and to avoid the accumulation of deviations caused by execution layer saturation, the controller first... The actual control inputs that can be sent out are obtained by applying amplitude and rate of change limits according to the flight control's permissible boundaries. ,in This indicates the execution control input after amplitude and speed limiting. The amplitude and speed limiting thresholds are calibrated offline by the UAV's dynamic capabilities and safety policies and are fixed in the parameter table. When amplitude or speed limiting is triggered, the trigger flag is also recorded for subsequent diagnosis and fallback strategies.
[0139] when When expressed in the docking interface coordinate system, the controller converts the known attitude transformation from the docking interface coordinate system to the navigation coordinate system into the navigation coordinate system expression desired by the flight controller. Furthermore, it converts acceleration commands into attitude and thrust commands or equivalent position and velocity commands executable by the flight controller. In the implementation using the "attitude + thrust" interface, the key quantities from acceleration to thrust can be calculated using the following formula:
[0140] ;
[0141] in, Indicates the coordinate system of the docking interface The following expresses the execution control input, Indicates transformation to navigation coordinate system The following expresses the execution control input, Indicates from coordinate system To coordinate system The rotation matrix is determined by the external parameter calibration of the docking interface and the real-time attitude calculation of the UAV, and is updated in each control cycle. Represents the expected force vector. Represents the gravitational acceleration vector and its magnitude is taken In coordinate system The middle direction is downward. Indicates the expected total thrust. This indicates the mass of the drone, which is a known parameter that can be obtained from pre-flight weighing or flight control parameter configuration. Represents the L2 norm;
[0142] In the When calculating attitude commands, the controller instructs the thrust axis of the machine system to... Orientation alignment is achieved to realize the desired acceleration, and the yaw angle is set to a preset yaw angle or maintained at the yaw angle of the previous cycle to avoid unnecessary yaw disturbances during docking. Subsequently, the attitude command and thrust command are packaged into execution commands supported by the flight control protocol and sent to the UAV flight control system through the airborne communication link. The execution commands include timestamps. The attitude and thrust targets, as well as the command validity period, are maintained in a zero-order hold mode within the current control cycle. The corresponding execution instructions continue until the next control cycle arrives to satisfy the discrete execution assumption of rolling optimization;
[0143] To enable step S1 to read the "historical control input", the controller will, after the command is successfully issued, Along with timestamps Write to the history control input buffer and maintain the most recent data in a circular queue. Records of each control cycle, among which This indicates a preset historical length that is consistent with step S1. At the same time, the execution confirmation information transmitted back by the flight control and the actual thrust or attitude tracking status are written into the diagnostic cache to support subsequent abnormal rollback.
[0144] At the start of the next control cycle The sampling is triggered again to obtain the current observation data of the airborne sensor, and the historical observation data and historical control input in the buffer are updated at the same time. Then, the process returns to step S1 to execute the next rolling optimization control and continue to complete the docking process between the UAV and the refueling device without stopping the machine.
[0145] In this specific embodiment, the probabilistic state-space model used in step S2 is constructed to include external stimulus inputs. The discrete-time state-space model is used, where the external excitation input term is used to characterize the rotor downwash intensity and participates in the update of the swing state parameters as a known input during the state prediction stage.
[0146] Specifically, the controller at each sampling time Read rotor-related quantities from the flight control system or electronic control telemetry and construct... ,in Indicates the first The external excitation input vector for each control cycle. This represents the rotor speed and is selected as either the arithmetic mean of the individual rotor speeds or a weighted average based on thrust distribution. This indicates the rotor pitch, and in fixed-pitch multirotors, it can be omitted and replaced with a preset constant value. This indicates the thrust command and allows selection of either the total thrust command after flight control and mixed control or the equivalent collective throttle conversion value.
[0147] To unify different types of rotor-related quantities into a "downwash intensity scalar" that can be used for oscillation updates, the controller calculates the induced velocity as the downwash intensity using the thrust command and defines it as follows:
[0148] ;
[0149] in Indicates the first The scalar value of the rinsing induction rate for each control cycle. Indicates thrust command. Indicates air density and selects... Alternatively, it can be obtained through online correction of airborne pressure, altitude, and temperature. It represents the equivalent disk area of the rotor, which is obtained by calibration of the rotor radius, and in the case of multiple rotors, it is the sum of the disk areas of each rotor.
[0150] When thrust command When the rotor speed cannot be directly obtained, the controller will set the rotor speed. With pitch Thrust mapping converted from offline calibration to The thrust mapping is obtained by fitting from ground bench tests or flight tests and stored in lookup table or polynomial form. Furthermore, to suppress excitation jitter caused by telemetry noise, the controller... Perform a first-order low-pass filter and apply upper and lower bound saturation to ensure its physical feasibility;
[0151] During the state prediction phase, the controller will The oscillating submodel is injected as known input to update the oscillating state parameters. The oscillation state parameters include at least the oscillation angle. With angular velocity It can further include the equivalent pendulum length. With equivalent damping coefficient In the implementation, the oscillating sub-model adopts the discretized form of a damped pendulum and superimposes a shuffling excitation term, specifically:
[0152] and ;
[0153] in This indicates the predicted swing angle value for the next control cycle. This indicates the estimated swing angle or particle value for the current control cycle. This indicates the predicted angular velocity of the pendulum in the next control cycle. This indicates the estimated angular velocity or particle value for the current control cycle. Indicates the sampling period. Represents the acceleration due to gravity and takes Indicates the equivalent pendulum length and limits it to a preset range. Inner and Represents the equivalent damping coefficient and is limited to a preset range. Inner and This represents the coupling gain from the induced velocity to the equivalent angular acceleration, determined offline and fixed as model parameters. This indicates the scalar value representing the intensity of the wash. Represents the sine function;
[0154] In practical implementation, to ensure that the excitation direction is consistent with the actual swing plane, the controller determines the swing plane based on the instantaneous offset direction of the joint in the docking interface coordinate system and... Perform directional projection to obtain Substituting the excitation components in the same plane into the above update, the predictability of the swing evolution is improved by using the external excitation input of the rotor downwash correlation in the state prediction of the probabilistic state space model, and the prior information that is more in line with the physical mechanism is provided for the subsequent prediction of the probability distribution of the relative pose.
[0155] In this specific embodiment, the probabilistic state-space model used in step S2 explicitly constructs the oscillation state parameter vector as follows: ,in Indicates the sampling time The oscillation state parameter vector, This indicates the angle of the joint relative to the vertical direction or the preset equilibrium direction. Indicates the angular velocity of the pendulum. Indicates the equivalent pendulum length. This represents the equivalent damping coefficient. Indicates the transpose operation;
[0156] in and The state estimates are given directly and denoted as follows: and and As a slow time-varying parameter, it is updated by online parameter identification and written back to the parameter table of the probabilistic state-space model for subsequent state prediction.
[0157] To ensure the feasibility and real-time performance of online identification, the controller employs online parameter identification based on recursive least squares and uses small-angle linearized oscillation dynamics as the identification model, when the following conditions are met... When linearization is enabled, This indicates the linearization enable threshold, which is given by engineering calibration and selected accordingly. to The sampling period is The acceleration due to gravity is And take The external excitation input is constructed using the thrust or downwash intensity obtained in step S7, and in this embodiment, the downwash induced velocity scalar is used. The downwash coupling gain is Furthermore, the results were obtained from offline experiments and solidified as constants;
[0158] The controller in After each control cycle completes the state update, it simultaneously reads from the cache. and Constructing identification samples and forming a linear regression form:
[0159] ;
[0160] in Indicates the first The regression output of each identified sample, Indicates the sampling time The estimated value of the pendulum angular velocity, Indicates the sampling time The estimated value of the pendulum angular velocity, Indicates the sampling period. Indicates the downwash coupling gain. Indicates the first The scalar value of the washing intensity per control cycle Represents the vector of independent variables in the regression. Indicates the sampling time The estimated value of the swing angle, This represents the regression residuals after combining unmodeled higher-order effects with noise. This represents the parameter vector to be identified, and its first component is... This is used to linearize the equivalent pendulum length in reciprocal form, and the second component is... Used to characterize the equivalent damping coefficient;
[0161] The controller then recursively calculates least squares pairs using the forgetting factor. Update online and convert the update results to obtain and Its recursive calculation is as follows:
[0162] ;
[0163] in Indicates the first The gain vector corresponding to each identified sample and Let represent the values of the parameter covariance matrix of the recursive least squares before and after the update, respectively. This represents the forgetting factor and is used to improve the tracking ability of slowly time-varying parameters and to select the appropriate value. and These represent the estimated values of the parameter vector before and after the update, respectively;
[0164] The controller will The first component is denoted as Based on this, the updated value of the equivalent pendulum length is calculated. ,Will The second component is directly used as the updated value of the equivalent damping coefficient. and to and Perform a projection of the physically feasible range to avoid non-physical values, where Restricted to and Restricted to and The aforementioned boundaries are obtained through prior offline calibration of the hose geometry and material damping.
[0165] To avoid recognition divergence due to insufficient stimulation, the controller must satisfy the following conditions: And regression residuals The above updates will only be performed at that time, among which Represents the L2 norm, Indicates the minimum excitation threshold. Represents the residual threshold and is used to remove outliers;
[0166] Updated and The state prediction module, written into the probabilistic state-space model, is used for state prediction in the next control cycle, enabling the swing evolution model to adapt online to changes in hose operating conditions and washing intensity, thereby improving the accuracy of docking relative state estimation and the consistency of its uncertainty characterization.
[0167] In this specific implementation, to transform the "duration duration is not less than the preset duration" event probability constraint in step 6 into a form that can be directly calculated and determined in the discrete prediction time domain, the controller records the preset duration as... And the sampling period is denoted as ,in This indicates the minimum duration for which the docking window needs to be maintained continuously. This indicates the time interval between sampling times of adjacent control cycles, and... Discretize into continuous prediction steps Used as a discrete criterion for the persistence of an event;
[0168] The discretization employs floor function rounding to ensure that the duration of discrete events is not underestimated. Specifically:
[0169] ;
[0170] in, This indicates the minimum number of consecutive prediction steps corresponding to the preset duration. This represents the floor operator. Indicates the preset duration. Indicates the sampling period;
[0171] Based on this, the controller defines the "docking window event" as: when the prediction time domain length is There exists a certain starting step in the discrete time series. So that from the first Step start continuous The future relative pose of each prediction step falls within the docking feasible region defined in step S3. ,in Indicates the number of prediction steps. This represents the set of future relative poses that simultaneously satisfy the following conditions: relative position error is less than a relative position error threshold, relative attitude error is less than a relative attitude error threshold, and relative velocity is less than a relative velocity threshold. The range of values is limited to To ensure continuity Do not cross the line;
[0172] To make the probability of this event calculable, the Monte Carlo particle prediction result generated in step S3 is reused in the first step. Each control cycle for each particle Step by step, check whether its future relative pose falls into And calculate the longest consecutive satisfying length. When the longest consecutive satisfying length is not less than When the particle satisfies the docking window event, its event indicator is recorded as follows. ,in Indicates the first The particles exist in a continuous time domain during prediction. During the time period when all steps can be connected, The statement indicates that the event does not exist, and then the probability of the docking window event is estimated based on the particle ratio. It is used for determining and solving the probabilistic constraints of predictive control in chance-constrained models, specifically:
[0173] ;
[0174] in, Indicates the first Estimated event probability for each control cycle's docking window. This represents the number of particles used for probability estimation. This represents the summation operation. Indicates the particle number, Indicates the first The docking window event indicator for each particle;
[0175] In step S6, when constructing the chance-constrained predictive control problem, the probability constraints are written as follows: The probability lower limit is not less than that determined by the risk budget allocation coefficient, thus equivalently realizing the requirement that "the duration is not less than the preset duration" as "there exists a continuous duration". Each prediction step falls within the feasible docking region. "This discrete event ensures that the constraints are computable, verifiable, and easy to update and solve in real time during rolling optimization."
[0176] In this specific implementation, the deep reinforcement learning strategy uses a risk budget allocation coefficient sequence distributed along the prediction time domain. ,in Indicates the first Risk budget allocation coefficient sequence for each control cycle Indicates the first The permissible default risk share corresponding to each prediction step Indicates the prediction step index. Indicates the preset prediction steps. Indicates the control cycle number. Indicates the transpose operation;
[0177] To ensure that the risk budget allocation coefficients are non-negative and that the total risk budget is conserved, the policy network first outputs the unnormalized risk score vector. ,in Represents the risk score vector. Indicates the first The risk score for each prediction step is obtained through normalized mapping. ,use:
[0178] ;
[0179] in, This indicates a preset total risk budget and meets the following conditions. And it is set offline by the system. This represents a function that exponentially normalizes a vector so that all components of the output vector are non-negative and the sum of the components is 1.
[0180] After the safety gating in step S5 passes boundary constraints and renormalization, the controller still records the final executable risk budget allocation coefficient sequence as... This is used in step S6 to determine the probability lower bound for each time step or each event, where "risk budget" means the default probability quota allowed by the opportunity constraint. The controller will allocate the default risk quota for each prediction step. Mapped to the lower probability bound of this step The total default risk over the entire prediction time domain is mapped to the lower bound of event probability. The mapping relationship is as follows:
[0181] ;
[0182] in, This represents the summation operation. Indicates the first The lower bound of the probability corresponding to each prediction step is used to construct the "probability constraint for each time step". It represents the lower bound of the probability of events in the docking window and is used to construct "event-type probability constraints";
[0183] During the insertion phase, the docking window probability sequence output in step S3 is used. Gradually aligning with the aforementioned lower probability bound, where This represents the probability sequence of the future relative pose falling into the docking feasible region at each prediction step within the prediction time domain. Indicates the first The docking window probability for each prediction step is given, and it is required that the set of consecutive prediction steps used for inserting action coverage satisfies the following conditions. To achieve "opportunity constraint tightness allocated by time step", and to adopt a method for handling window events with "a duration not less than a preset duration of dockable time periods". The total event probability is lowered in the form of "opportunity constraint imposed on events", so that the risk budget allocation coefficient sequence can reflect the risk forward or backward allocation under the changing trend of the docking window, and can also ensure the conservation of the total risk budget and provide quantifiable and consistent constraints on success rate and safety margin in rolling optimization.
[0184] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0185] This invention addresses the problem of random oscillations in flexible hose joints under rotor downwash excitation, leading to uncertain changes in the docking window over time. It combines a closed-loop approach using "probabilistic state-space model prediction," "rolling optimization of chance-constrained model predictive control," and "deep reinforcement learning self-tuning." The probabilistic state-space model outputs a docking relative state with uncertainty representation based on the fusion of historical control inputs and current observations, and provides a probability distribution of future relative poses in the prediction time domain. Furthermore, the "probability of the future pose falling into the docking feasible region" is explicitly transformed into a docking window probability sequence, enabling the controller to perform forward planning based on the probability trend of window occurrences, rather than relying solely on instantaneous errors. Chance-constrained model predictive control introduces docking window event probability constraints in the rolling optimization, constraining the existence of "dockable periods that satisfy docking criteria and last for a certain duration" in the form of a probability lower bound. This allows for a quantifiable trade-off between success rate and safety margin under input and state constraints, improving the docking success rate and robustness in uncertain oscillation scenarios.
[0186] This invention improves the algorithm structure to address the "randomness of the docking window": First, it introduces rotor downwash-related quantities as external excitation inputs into the state-space model, and updates the oscillation equivalent parameters through online parameter identification, thereby improving the predictability and model adaptability of oscillation evolution. Second, it extends the output of the state-space model from traditional point estimation to the construction basis of docking window probability and window event probability constraints, making the optimization objective directly correspond to the key technical bottleneck of "window capture". Third, it uses a deep reinforcement learning strategy to output risk budget allocation coefficients and switching thresholds between waiting and insertion phases online, enabling risk allocation and phase decisions to adaptively adjust with oscillation intensity and window change trends. Furthermore, it uses safety gating to constrain and backtrack the parameter value range and the feasibility of the optimization problem, avoiding infeasible solutions or unsafe control inputs caused by learned parameters. This improves the docking success rate while ensuring control execution and system safety.
Claims
1. A method for unmanned aerial vehicle (UAV) no-stop refueling docking based on a deep reinforcement predictive control algorithm, comprising: S1. obtaining current observation data of an on-board sensor in a current control period, reading historical observation data and historical control input, and constructing time series input data in time sequence; S2. inputting the time series input data into a probabilistic state space model to obtain docking relative state estimation at the current time; S3. based on the docking relative state estimation, using the probabilistic state space model to calculate a future relative pose prediction probability distribution of a docking adapter relative to a docking interface of the UAV within a preset prediction time domain, and based on a preset docking criterion, calculating a docking window probability sequence from the future relative pose prediction probability distribution; S4. inputting the docking relative state estimation and the docking window probability sequence into a deep reinforcement learning strategy to output a parameter set of an opportunity constraint model predictive controller, the parameter set including a risk budget allocation coefficient, a waiting phase switching threshold, and an insertion phase switching threshold; S5. performing safety gate processing on the parameter set to obtain an executable parameter set, the safety gate processing including limiting the parameter set within a preset value range, and based on the limited parameter set, constructing an opportunity constraint model predictive control problem for feasibility determination, and when the feasibility determination result is infeasible, replacing the limited parameter set with a preset safety parameter set; S6. inputting the docking relative state estimation, the future relative pose prediction probability distribution, the docking window probability sequence, and the executable parameter set as inputs to perform rolling optimization of the opportunity constraint model predictive control to obtain a control input for the current control period; S7. sending the control input to a UAV flight control system for execution and entering the next control period.
2. The method of claim 1, wherein S2 comprises: inputting the time series input data into the probabilistic state space model; performing state prediction based on the historical control input and state update based on the current observation data by the probabilistic state space model to obtain the docking relative state estimation at the current time; wherein the docking relative state estimation includes relative position, relative attitude, and swing state parameters of the docking adapter relative to the docking interface of the UAV, and the docking relative state estimation includes estimated values of the relative position, the relative attitude, and the swing state parameters and uncertainty representations corresponding to the estimated values.
3. The method of claim 1, wherein S3 comprises: based on the docking relative state estimation, recursively calculating the future relative pose prediction probability distribution of the docking adapter relative to the docking interface of the UAV within the preset prediction time domain by the probabilistic state space model in time steps; for each time step within the preset prediction time domain, determining a docking feasible region according to the docking criterion, the docking feasible region being a future relative pose set that simultaneously satisfies a relative position error being less than a relative position error threshold, a relative attitude error being less than a relative attitude error threshold, and a relative velocity being less than a relative velocity threshold. and a probability of the future relative pose falling into the docking feasible region is calculated based on the corresponding future relative pose prediction probability distribution at each time step, and the probability at each time step corresponds to the docking window probability sequence.
4. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 1, S4 comprising: inputting the docking relative state estimation and the docking window probability sequence into a deep reinforcement learning strategy; outputting, by the deep reinforcement learning strategy, a parameter set of an opportunity constraint model predictive controller according to a joint swing state represented by the docking relative state estimation and a docking window change trend represented by the docking window probability sequence; wherein the parameter set comprises a risk budget allocation coefficient for determining the docking window event probability constraint, a waiting phase switching threshold for representing the transition from the waiting phase to the insertion phase, and an insertion phase switching threshold for representing the exit from the insertion phase.
5. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 1, S5 comprising: performing boundary restriction processing on the parameter set, so that the risk budget allocation coefficient, the waiting phase switching threshold and the insertion phase switching threshold in the parameter set fall into corresponding preset value ranges, to obtain a boundary-restricted parameter set; constructing an opportunity constraint model predictive control problem based on the boundary-restricted parameter set and performing a feasibility determination, the feasibility determination comprising determining whether the opportunity constraint model predictive control problem has a solution that satisfies the input constraint and the state constraint; when the feasibility determination result is feasible, taking the boundary-restricted parameter set as an executable parameter set; when the feasibility determination result is infeasible, determining a preset safety parameter set as the executable parameter set, wherein the preset safety parameter set is a parameter set that is preset and satisfies the corresponding preset value range.
6. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 1, S6 comprising: taking the docking relative state estimation as the initial state and taking the executable parameter set as the control parameter of the opportunity constraint model predictive controller, and combining the future relative pose prediction probability distribution and the docking window probability sequence to construct an opportunity constraint model predictive control problem comprising a cost function, an input constraint, a state constraint and a docking window event probability constraint; wherein the docking window event probability constraint comprises: within a preset prediction time domain, the probability that there is a period that satisfies the docking criterion and has a duration not less than a preset duration is not less than a probability lower limit determined by the risk budget allocation coefficient; at each rolling optimization time, solving the opportunity constraint model predictive control problem to obtain an optimal control input sequence, and selecting the first item in the optimal control input sequence as the control input of the current control period.
7. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 1, characterized in that, The probabilistic state space model is a state space model including an external excitation input term, the external excitation input term including one or more of a rotor speed, a pitch or a thrust command representing a rotor downwash intensity, and the external excitation input term being taken as a known input to update a swing state parameter in state prediction.
8. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 2, characterized in that, The swing state parameter includes one or more of a swing angle, a swing angular velocity, an equivalent swing length and an equivalent damping coefficient, and the probabilistic state space model updates the equivalent swing length and / or the equivalent damping coefficient through online parameter identification.
9. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 6, characterized in that, The docking window event probability constraint that the duration is not less than a preset duration is achieved through discretization, which includes converting the preset duration into a continuous prediction step number K and defining the docking window event as a future relative pose falling into the docking feasible region existing for K continuous prediction steps within the preset prediction time domain.
10. The unmanned aerial vehicle non-stop refueling docking method based on the deep reinforcement predictive control algorithm according to claim 1, characterized in that, The risk budget allocation coefficient output by the deep reinforcement learning strategy is a risk budget sequence allocated along a prediction time domain, the risk budget sequence satisfying a sum of risk budgets at each time step being equal to a preset total risk budget, and being used to respectively determine a probability lower limit at each time step or each event.