Unmanned aerial vehicle autonomous flight control method combined with overturning angle adjustment of photovoltaic cell panel
By employing a two-tiered collaborative mechanism of reinforcement learning and model predictive control, the unified optimization problem of flight control and energy management for UAVs in complex environments is solved, achieving an adaptive dynamic balance between energy and flight performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JETLINE AVIATION (SHANGHAI) CO LTD
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-21
AI Technical Summary
Existing UAV control methods cannot achieve deep and coordinated optimization of flight control and energy management within a unified framework, and the control system cannot adaptively adjust priorities according to long-term mission requirements and dynamic environments.
A two-layer collaborative mechanism is constructed by combining a deep neural network decision model based on reinforcement learning with model predictive control. The high-level decision layer generates importance indicators through overall task and environmental information, while the low-level execution layer performs multi-objective optimization to achieve collaborative optimization of the photovoltaic panel flip angle and the conventional control surface angle.
It achieves the integration and dynamic balance of long-term and short-term optimization goals. The system can adaptively adjust control objectives according to mission requirements and environmental changes, thereby improving the intelligent trade-off between energy and flight performance.
Smart Images

Figure CN121900440A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control methods, specifically to an autonomous flight control method for UAVs that incorporates the adjustment of the flip angle of photovoltaic panels. Background Technology
[0002] Unmanned aerial vehicles (UAVs) require extremely high energy self-sufficiency for long-duration, long-endurance missions. Solar-powered UAVs, due to their clean and renewable energy source, have become a key technological direction for extending endurance. Traditional solar-powered UAVs primarily use fixed photovoltaic panels, whose energy harvesting efficiency is limited by the aircraft's attitude, making it difficult to achieve maximum power point tracking (MPPT). This leads to a sharp decline in power generation under complex maneuvers or non-ideal flight attitudes. In recent years, there has been a trend of integrating photovoltaic panels into flight control surfaces, forming solar control wings. These wings can both harvest energy and serve as auxiliary control surfaces for attitude control. This design presents a core technological challenge: there is a natural and strong coupling and conflict between maximizing the perpendicularity of the solar panel to sunlight and minimizing drag to achieve rapid attitude response. For example, when the photovoltaic wing is adjusted to its maximum power generation angle, the additional aerodynamic forces and torques generated can disrupt the UAV's attitude, requiring additional conventional control surface energy to offset, potentially leading to a decrease in net energy gain and even affecting flight safety and trajectory accuracy.
[0003] The aforementioned patent documents and prior art have the following technical problems when used:
[0004] Problem 1: Existing methods typically treat flight control and energy management as independent subsystems. Energy management focuses only on instantaneous power generation, while the control system focuses only on flight stability; neither is optimized in a deep, forward-looking manner within a unified framework.
[0005] The second problem is that traditional control systems cannot adaptively adjust the priority of energy and flight accuracy based on complex long-term mission objectives and dynamic environmental information. When battery power is sufficient, trajectory accuracy and efficiency should be prioritized; when battery power is insufficient, energy priority should be switched, sacrificing some flight accuracy. This macro-level strategic trade-off and dynamic adjustment of priorities is a gap in existing methods. Summary of the Invention
[0006] Technical problems to be solved
[0007] To address the shortcomings of existing technologies, this invention provides an autonomous flight control method for unmanned aerial vehicles (UAVs) that incorporates photovoltaic panel flip angle adjustment, solving the following problems:
[0008] 1. Addressing the problem that control methods cannot perform in-depth collaborative optimization of multiple conflicting objectives within a single, unified framework;
[0009] 2. This addresses the problem that the control system is static and cannot adaptively adjust the priority of different control objectives according to long-term task requirements and dynamically changing environmental conditions.
[0010] Technical solution
[0011] To achieve the above objectives, the present invention provides the following technical solution: an autonomous flight control method for unmanned aerial vehicles (UAVs) that combines photovoltaic panel flip angle adjustment, the control method comprising the following steps:
[0012] Sp1: Data Acquisition and Classification: Acquire the overall mission and environmental information and current flight data of the UAV. The overall mission and environmental information are used for the high-level decision-making layer, and the current flight data is used for the low-level execution layer.
[0013] Sp2: Macro Strategy Generation: The advanced decision-making layer adopts a deep neural network decision-making model based on reinforcement learning. Based on the overall task and environmental information, it outputs a set of dynamically changing importance indicators in real time. The importance indicators are used to reflect the current priority of multiple control objectives such as flight tracking and energy gain.
[0014] Sp3: Low-level predictive control: The low-level execution layer adopts a model-based predictive control framework and incorporates a comprehensive predictive model of energy, aerodynamics and control. The comprehensive predictive model includes the coupling relationship between the deflection angle of the solar control wing, the angle of the conventional flight control surface and the aerodynamic performance and energy gain of the UAV.
[0015] Sp4: Multi-objective optimization solution: The underlying execution layer calculates the future state transition of the UAV in the prediction time domain based on the comprehensive prediction model, and substitutes the prediction results into the multi-objective optimization evaluation formula for solution, so as to obtain a set of optimal control commands that simultaneously satisfy multiple objectives in each control cycle;
[0016] Sp5: Cooperative Command Execution: The optimal control command includes the deflection angles of all solar control wings and the angles of all conventional flight control surfaces. The UAV executes the optimal control command to achieve autonomous flight.
[0017] Preferably, the control method is applicable to UAVs equipped with solar control wings. The solar control wings serve as both an energy harvesting component and an auxiliary flight attitude control component for the UAV. The solar control wings include a photovoltaic wing plate that can deflect around a rotation axis and a drive actuator connected to the photovoltaic wing plate. The surface of the photovoltaic wing plate is provided with a photovoltaic cell array. Under the drive actuator, the photovoltaic wing plate changes its own deflection angle, thereby adjusting the incident angle of the photovoltaic array and generating additional aerodynamic forces and torques during the deflection process, thereby achieving synergistic optimization of energy harvesting and auxiliary attitude control.
[0018] Preferably, the integrated prediction model consists of a flight dynamics model and an energy model. The flight dynamics model is used to calculate the additional lift, drag, and torque changes generated by the flipping attitude of the solar-powered control wing. The energy model is used to calculate the total instantaneous power generation based on the direction of sunlight, the attitude of the UAV, and the flipping angle of the solar-powered control wing.
[0019] Preferably, the overall mission and environmental information acquired by the advanced decision-making layer includes at least one of the following: the current battery level of the UAV, the real-time position of the direction of sunlight, the estimated wind speed, the intensity of airflow, the time until sunset, and the distance to the mission target point.
[0020] Preferably, the multi-objective optimization evaluation formula is a function that sums multiple costs over a prediction time range. The multiple costs include at least: a trajectory tracking error term describing the accuracy of the flight path, a control component wear term describing the wear of the control surfaces and motors, and an energy gain term describing the power generation revenue. The optimization variables of the multi-objective optimization evaluation formula include the deflection angle of the solar control wing and the deflection angle of the conventional control surfaces. The constraints include UAV attitude stability constraints, flight mechanics constraints, and power generation output constraints.
[0021] Preferably, the higher-level decision-making layer operates at a longer interval to update the importance index, while the lower-level execution layer operates at a shorter interval to solve for the optimal control command, thereby forming a hierarchical intelligent control system.
[0022] Preferably, when the advanced decision-making layer determines that the UAV is in an energy-priority state based on overall task and environmental information, it outputs a higher energy gain importance index. Accordingly, the underlying execution layer prioritizes adjusting the solar control wings to maximize alignment with the sun, and at the same time collaboratively calculates compensatory control commands to counteract the aerodynamic disturbances generated.
[0023] Preferably, the advanced decision-making layer determines that the UAV is in a control priority state based on overall mission and environmental information. When performing complex maneuvers or encountering strong winds, it outputs a high trajectory tracking importance index. The underlying execution layer then coordinates the use of solar-powered control wings and conventional flight control surfaces to achieve flight control with minimal drag and the fastest response.
[0024] Preferably, the intelligent learning decision-making model adopted by the advanced decision-making layer is obtained by pre-training in a simulation environment based on historical flight data and mission environment. The simulation environment has the integrated prediction model of energy, aerodynamics and control built in. The intelligent learning decision-making model learns the optimal importance index output strategy by maximizing the total reward over a long period.
[0025] Preferably, the control method is executed by the central processing unit in the UAV flight control system, and forms a closed-loop control structure by combining the sensor acquisition module and the actuator control module. The sensor acquisition module includes an attitude sensor, a light intensity sensor, a wind speed sensor and a battery power detection module, which collects the UAV flight status and environmental parameters in real time. The actuator control module includes a control surface driver, a photovoltaic wing deflection driver and a motor controller, and receives the optimal control command output by the central processing unit and drives the corresponding actuators to move.
[0026] Beneficial effects
[0027] This invention provides an autonomous flight control method for unmanned aerial vehicles (UAVs) that incorporates photovoltaic panel tilt angle adjustment. It offers the following advantages:
[0028] 1. This invention introduces a two-layer collaborative mechanism of reinforcement learning and model predictive control in its overall architecture, achieving a unification of long-term strategy optimization and short-term dynamic control. The high-level decision layer, based on a deep reinforcement learning model, learns a "multi-objective weight allocation strategy" under different states, considering both the global mission objective and long-term energy gains. This strategy dynamically determines the relative importance of indicators such as trajectory accuracy, energy gains, and system stability under different flight phases and energy states. The lower-level execution layer, centered on model predictive control, reads the aforementioned weights and current state estimates in real time. Within the prediction time domain, it integrates the flight dynamics model and the photovoltaic energy model to construct a multi-objective weighted cost function and performs online optimization to obtain the optimal control command for the current moment. This two-layer architecture achieves information transmission and collaboration through shared memory, enabling reinforcement learning to focus on long-term global optima and model predictive control to focus on instantaneous local optima. The two form a closed-loop iteration, fundamentally achieving the fusion and dynamic balance of long-term and short-term optimization objectives.
[0029] 2. The advanced decision-making layer of this invention, through mission planning and environmental perception, calculates weights in real time for factors such as remaining range, solar altitude angle, battery power, and airflow disturbances. When battery power is sufficient, the system automatically increases the weights of trajectory accuracy and mission completion; when battery power is insufficient or near sunset, it automatically switches to an energy-priority mode, increasing the weights of photovoltaic panel angle optimization and energy harvesting efficiency. In this way, the system can adaptively adjust control objectives according to mission requirements and environmental changes, achieving an intelligent trade-off between energy and flight performance. Attached Figure Description
[0030] Figure 1 This is a flowchart of the present invention;
[0031] Figure 2 This is a system architecture diagram of the present invention;
[0032] Figure 3 This is a diagram of the method architecture of the present invention. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Specific Implementation Example 1:
[0035] like Figures 1 to 2 As shown, an autonomous flight control method for a drone that combines photovoltaic panel flip angle adjustment is described. The control method includes the following steps:
[0036] Sp1: Current flight data is provided in real-time by an onboard multi-source sensor system, covering key physical quantities such as attitude, position, environment, energy, and actuator status. This system constitutes the foundational information layer for the UAV's autonomous control and energy management. Specifically, it includes: Attitude Sensor (IMU): Employing a six- or nine-axis inertial measurement unit (accelerometer, gyroscope, and optional magnetometer), it outputs three-axis angular velocity and linear acceleration for attitude estimation and flight dynamics calculation, serving as the primary feedback source for lower-level control. Global Navigation Satellite System (GNSS) Receiver: Providing three-dimensional position, ground velocity, and track angle information for positioning and velocity constraint correction. Illumination and Solar Vector Sensor: Acquiring incident light direction, illumination intensity, and solar azimuth angle information, providing real-time data for optimizing photovoltaic panel deflection angles and predicting power generation. Wind Speed and Turbulence Sensor: Including pitot tubes, five-hole probes, or onboard micro-wind measurement modules, serving as auxiliary calibration or redundancy. The system primarily employs a synthetic wind field estimation method based on a dynamic model, combining airframe attitude, airspeed sensor readings, and ground velocity information provided by GNSS. It utilizes extended Kalman filtering or a disturbance observer to identify and estimate real-time wind speed, direction, and turbulence online, meeting the accuracy requirements for lightweight and low-speed flight. The battery monitoring module detects voltage, current, temperature, and remaining capacity, and estimates remaining flight time and state of energy using power integration. Actuator feedback includes motor speed and current feedback, as well as servo angle and position encoded signals for precise closed-loop control and actuator health monitoring. Photovoltaic wingplate feedback uses deflection angle encoders and driver current feedback to ensure accurate photovoltaic wingplate angle adjustment and actuator safety. A temperature sensor array monitors the temperature of photovoltaic cells and critical airframe components for thermal protection and efficiency correction calculations.
[0037] To achieve temporal consistency and real-time computation of multi-source data, each sensor is configured with a different sampling frequency based on its physical characteristics: the inertial measurement unit (IMU) operates at a high-frequency sampling rate of 100–1000 Hz; the GNSS updates at 1–10 Hz; the illumination and solar vector sensor at 10 Hz; the battery and motor monitoring module at 1–10 Hz; and the servo motor and yaw angle feedback maintain an update rate of 10–100 Hz. All sampled signals are first timestamped and synchronized via the sensor acquisition module before entering the state estimation and preprocessing module. This module performs the following multi-level filtering and data fusion: Attitude estimation based on Extended Kalman Filter (EKF): The acceleration and angular velocity data provided by the IMU have high-frequency characteristics, but they accumulate drift errors over time; while absolute sensors such as GNSS and magnetometers update slowly but are stable. The system uses the Extended Kalman Filter algorithm to establish the nonlinear state equations and observation equations of the aircraft, and achieves optimal fusion of multi-source information by recursively estimating the state variables, namely attitude angle, angular velocity, position, and velocity. Specifically, the EKF uses IMU integration to obtain prior estimates of attitude and velocity during the prediction phase, and performs corrections based on GNSS position and magnetometer azimuth measurements during the update phase, thus balancing dynamic response and long-term stability. The filter can also be extended to incorporate solar vector direction and aerodynamic disturbance estimates, enabling the system to maintain high-precision attitude calculations even under illumination disturbances or wind field changes. The attitude and velocity states processed by the EKF are used for control law solving and trajectory constraint correction; synchronization compensation is based on weighted averaging or time-series interpolation: due to significant differences in sampling frequencies among various sensors, directly using raw data would cause timing misalignment and numerical abrupt changes, affecting controller stability. Therefore, the system performs time alignment and synchronization compensation processing before data fusion.
[0038] For low-frequency signals, the system resamples them to the same time base as the IMU using linear or spline interpolation based on adjacent sampling points. For quantities with multi-sensor redundancy, such as attitude angle or wind speed estimation, the system performs a weighted average based on the real-time confidence of each sensor (determined by indicators such as noise variance and signal integrity) to obtain a smoother and more reliable fusion result. This synchronization and weighted compensation mechanism ensures the phase consistency of each data channel, enabling the controller to respond accurately on a millisecond timescale. Sensor fault detection and redundancy switching based on residuals and statistical characteristics: In long-duration flights or complex environments, individual sensors may produce abnormal readings due to interference, obstruction, or aging. The system monitors data consistency in real time by constructing observation residuals, i.e., the difference between measured values and filtered predicted values, and uses sliding window statistical analysis (such as residual mean, variance, and rate of change) to determine abnormal patterns. When an anomaly continuously deviating from a set threshold is detected, the system automatically triggers redundancy switching: if the attitude sensor fails, the backup IMU or estimation model takes over the attitude information; if the GNSS fails, it temporarily switches to relative navigation based on inertial and barometer measurements; if there is an anomaly in illumination or solar vector, historical trends or geographic models are used for prediction and compensation. After the fault is recovered, the system automatically performs sensor recalibration and weight reallocation to ensure continuous and stable operation of flight control and energy optimization. Through the above three mechanisms, the state estimation and preprocessing module achieves high-precision fusion of attitude, position, and environmental states, and provides continuous, reliable, and low-latency data support for the control and decision-making layers.
[0039] After fusion and correction, the data is used bidirectionally within the system: on one hand, it serves as direct input to the lower-level execution layer, Model Predictive Control (MPC), for real-time solution of optimal attitude, thrust, and photovoltaic deflection control commands; on the other hand, after feature extraction and aggregation, it generates higher-level feature quantities, such as instantaneous power generation, short-term energy prediction, and wind speed trend estimation, which serve as input to the higher-level decision layer, namely reinforcement learning-based policy decisions, providing support for macro-level flight strategies, energy allocation, and mission planning. Furthermore, long-term system operation data is stored and managed through a log and learning module. This module is responsible for flight log compression, key event labeling, data cleaning, and uploading. Data sources include airborne real-time recordings, historical mission data aggregated by ground stations, and expanded samples generated from the simulation environment, used for periodic offline training and iterative updates of the policy model, thereby enabling the self-evolution and long-term performance optimization of the intelligent control algorithm.
[0040] Sp2: The advanced decision-making layer is the core intelligent module of the entire UAV autonomous control system. Its main function is to generate a time series of importance indicators that reflect the priority of each control objective based on long-term mission objectives and real-time system status under complex mission and dynamic environmental conditions. This provides a weighting basis for the lower control layer and achieves global optimal coordination among mission execution, energy management and system health.
[0041] The input information for the high-level decision layer integrates multi-source data from the task layer, environment layer, and system layer, covering the following: Task layer information: task type and priority, remaining range, distance to the target point, estimated task completion time, sunshine duration, and sunset time, used to characterize task constraints and time pressure; Environment layer information: including short-term weather forecasts, i.e., wind speed, wind direction, temperature, solar altitude and azimuth angles, current illumination intensity and incident angle estimates, and historical wind field statistical distribution, used to assess external disturbances and energy availability; System layer information: including current battery state of charge, energy consumption rate, estimated maximum available thrust, actuator temperature status, real-time disturbance levels, such as airflow disturbances and attitude jitter indicators, and short-term energy prediction results fed back from the lower layers. After feature normalization and time weighting, these inputs are assembled into high-dimensional feature vectors and input into the reinforcement learning decision network. The model outputs several important metrics to guide the optimization of the underlying control system, specifically: a trajectory tracking importance metric, used to balance path accuracy and energy expenditure; an energy gain importance metric, reflecting the optimization priority of photovoltaic power generation efficiency; a device health importance metric, corresponding to actuator wear and temperature rise constraints; and a response speed weighting metric, used to adjust the dynamic sensitivity of attitude and thrust control. Each metric is constrained within a preset numerical range so that the underlying execution layer can directly read and apply it to the weighting function of model predictive control, achieving adaptive adjustment of multi-objective trade-offs.
[0042] The advanced decision-making layer employs a deep neural network model based on reinforcement learning as its core policy engine. This model achieves optimal decision-making over long time scales by learning the nonlinear mapping relationship between task objectives, environmental states, and system feedback. Its structure typically includes: an input layer that receives fused task, environment, and system feature vectors; an intermediate layer composed of multi-layer fully connected structures and time-series feature extraction units such as Long Short-Term Memory networks or gated recurrent units, used to capture the temporal dependencies and dynamic trends between states; and an output layer that generates weighted importance indices for each control objective. Model training is based on a reinforcement learning framework, aiming to maximize long-term cumulative rewards. This reward function is composed of multiple evaluation dimensions, including: task completion and time efficiency; photovoltaic energy self-sufficiency rate and remaining power level; flight stability and attitude safety margin; device lifetime consumption and temperature upper limit constraints; overall mission success rate and emergency obstacle avoidance results, etc. Training is conducted in a high-fidelity simulation environment, which incorporates aerodynamic models, energy balance models, disturbance generation modules, and control execution sub-modules consistent with the actual flight control system, thereby ensuring the transferability and consistency of the policy. Training samples are sourced from historical flight logs, randomly generated wind and lighting conditions, and manually set extreme conditions. A domain randomization mechanism is introduced during training, continuously varying wind speed, lighting, load, and actuator parameters to allow the model to learn universally applicable policies in diverse environments. Adversarial training or policy entropy regularization enhances the model's robustness. This ensures that when sensor input data is noisy or deviates significantly from the training distribution, the model tends to output conservative policies or maintain previous weights, thus avoiding high-frequency oscillations and unstable control decisions. The algorithm can employ proximal policy optimization, deep deterministic policy gradient, or other reinforcement learning methods suitable for continuous action spaces. After training, model parameters are compressed and quantized and stored in an onboard model repository for real-time access during flight missions. To improve long-term adaptability, the system allows for limited online fine-tuning. When the UAV's long-term operating environment deviates from the training distribution, lightweight policy updates can be performed based on data from several recent missions. This process is only performed under conditions of stable flight, sufficient computing resources, and ground authorization, ensuring safety and consistency. The reasoning cycle of the high-level decision-making layer is relatively long, typically performing a complete calculation every ten seconds to one minute, outputting a new set of importance indicators and writing it to the system's shared memory. The lower-level execution layer operates at high frequency, reading the latest indicators in each short cycle and using them as optimization weights input to the model predictive control solver, ensuring that flight control maintains a continuous response to the intentions of the higher levels over short timescales. When the high-level decision-making layer becomes temporarily unavailable due to computational delays, communication interruptions, or policy failures, the system automatically enters a redundancy protection mode, using ground-preset or the default importance indicators from the last normal output to maintain operation, thereby ensuring flight safety and control continuity.
[0043] Sp3: The integrated prediction model is used to characterize the dynamic coupling relationship between the solar-powered control wing deflection angle, conventional control surface angle, and the aerodynamic performance and energy gains of the UAV, providing a basis for the underlying predictive control module to perform state evolution and cost assessment in the prediction time domain. The model consists of two main parts: a flight dynamics model and an energy model, and comprehensively considers the dynamic characteristics of actuators, sensor delays, and environmental disturbances to ensure the physical consistency and real-time feasibility of the prediction results.
[0044] The aerodynamics and control model is a flight dynamics model based on six-degree-of-freedom rigid body dynamics equations. It describes the forces and motion of the UAV in three-dimensional space and is used to calculate the impact of photovoltaic wing deflection and conventional control surface deflection on the airframe attitude and aerodynamic performance. The model mainly includes the following parts: Aerodynamics module: This module calculates lift, drag, and aerodynamic components such as pitch, yaw, and roll moments. In real-time calculations, lookup table interpolation or low-rank approximation models are used to quickly solve for the aerodynamic coefficients. The aerodynamic correction terms caused by the photovoltaic wing deflection angle, especially the pitch and roll moments, are crucial. To address the unsteady aerodynamic effects and model uncertainties caused by variable configurations, a generalized disturbance observer or incremental adaptive law is embedded in the flight dynamics model. This is used to compensate for the deviation between the model-predicted aerodynamic forces and actual aerodynamic characteristics in real time, enhancing robustness to variable configurations and wind field disturbances. Aerodynamic characteristic data can be obtained from wind tunnel experiments or high-fidelity computational fluid dynamics simulations and embedded into the system through fitting or interpolation tables. The mass and inertia module defines the UAV's mass distribution and moment of inertia matrix, and updates mass parameters in real time based on power consumption or load changes to ensure the accuracy of dynamic calculations. The propulsion and gravity module determines the motor thrust provided by the propulsion system based on the current speed and voltage and applies it along the longitudinal axis of the fuselage; gravity always acts on the center of mass along the geocentric direction. The actuator dynamics module explicitly models the actuator response as a first- or second-order hysteresis system to reflect the actual dynamic characteristics of the servo and photovoltaic wing drive mechanisms. That is, the rate of change of the actual control surface deflection angle is equal to the difference between the target command angle and the current angle multiplied by the servo response coefficient, thereby limiting the controller from generating unrealizable transient commands and improving prediction accuracy. The energy model is used to predict the instantaneous power generation and cumulative energy output of the photovoltaic system in the future time domain under given solar illumination direction, fuselage attitude, and photovoltaic wing deflection angle. The instantaneous power generation of a photovoltaic (PV) panel is closely related to the incident irradiance, the cosine factor of the incident angle, the PV panel temperature, and the cell conversion efficiency. The relationship is expressed as: instantaneous power generation of the PV wing equals the incident irradiance multiplied by the effective projected area of the PV wing multiplied by the PV conversion efficiency. The incident angle is determined by the angle between the solar vector and the PV wing normal vector. The solar vector is calculated using a solar position algorithm combined with time, latitude, longitude, and date; the PV wing normal vector is determined by the aircraft attitude and wing deflection angle. The cosine factor of the incident angle varies with the cosine value of the angle, reaching its maximum value when the light is perpendicularly incident. The PV conversion efficiency is a function of temperature and operating conditions. Efficiency decreases as the panel temperature increases. The model uses an empirical linear correction relationship to dynamically compensate for the PV efficiency. To reflect the control characteristics of the power electronic converter, the model incorporates a simplified MPPT module to estimate the optimal operating point of the current voltage-current curve and calculate the maximum extractable power. This process avoids overestimation of energy prediction and improves the overall physical realism of the model.In the integrated prediction model, the flight dynamics model and the energy model are bidirectionally coupled through the deflection angle of the solar control wing: on the one hand, the deflection angle of the photovoltaic wing changes the local airflow direction of the wing, thus affecting the distribution of lift, drag, and torque; on the other hand, this angle directly determines the incident angle of sunlight on the photovoltaic array, thus affecting the light energy conversion efficiency and power generation. Therefore, the deflection angle simultaneously acts on both the flight control performance and energy gain performance system objectives, making them naturally correlated during the optimization process. However, the flight environment itself has strong randomness and uncertainty, such as fluctuations in solar irradiance due to cloud cover, changes in wind speed and direction with altitude and time, aerodynamic fluctuations caused by turbulent disturbances, and changes in battery performance due to temperature. If the model makes predictions based on only a single deterministic value, the resulting control commands may fail in the actual environment. To ensure the stability and feasibility of the control strategy in the real environment, this system introduces a probabilistic scenario description mechanism and an uncertainty propagation analysis method into the integrated prediction model.
[0045] Sp4: Probabilistic Modeling of Environmental Variables: The model first establishes a statistical distribution model for external disturbances, such as wind speed, solar irradiance, airflow turbulence energy, and air density. For wind speed and direction, a Gaussian mixture distribution or a probability density function fitted based on measured data is used; for solar irradiance, a confidence interval varying over time is generated by combining real-time illumination sensor readings with a weather forecast model; for turbulence disturbances, a time-domain noise model with random phase and power spectral density is introduced; and for sensor delay and measurement error, an additive noise model is used to characterize them. These variables are initialized with statistical parameters generated jointly by the sensor acquisition module and the prediction model at the beginning of the control cycle and input into the predictive control module. Within each prediction time domain, the system randomly samples according to the above probability distribution to generate several representative environmental scenarios, such as several different combinations of wind speed, solar angle, and irradiance. Each scenario propagates independently under the integrated prediction model, that is, the attitude evolution, trajectory deviation, energy gain, and control input response of the UAV in that scenario are calculated using flight dynamics and energy models. This yields a set of multi-scenario predicted trajectories. If wind speed is sampled at three confidence levels (low, medium, and high), and irradiance is sampled under two weather conditions (sunny and partly cloudy), the system calculates the future trajectory and energy output for six environmental scenarios in each control cycle. After multi-scenario prediction is completed, the system substitutes the trajectory deviation, energy output, attitude stability, and other indicators obtained for each scenario into the multi-objective evaluation function, and performs a weighted summary of the results for each scenario or uses the worst-case evaluation. In practical implementation, pipe model predictive control or constraint tightening strategies can be used to pre-shrink the safety constraint range by evaluating the impact of the worst-case environmental scenario on the system state boundary. This ensures that even under the worst disturbances in the real environment, the actual trajectory of the solved nominal control command is still confined within a safe channel, thus achieving optimal performance while ensuring absolute safety.
[0046] In the underlying execution layer, the system employs rolling predictive control for multi-objective optimization. Within each control cycle, the controller uses a comprehensive predictive model to make rolling predictions of the UAV's state at several future moments. The prediction results are then substituted into the multi-objective optimization evaluation formula to obtain the optimal control sequence within the prediction time domain. Only the first control command is actually executed in the current cycle; subsequent cycles involve re-predicting and re-solving, thus forming a continuously rolling closed-loop control process. Optimization variables include the deflection angle sequences of the solar-powered control wing and the conventional control surfaces, used to achieve a balance between energy harvesting and flight attitude control. Since the prediction time domain includes multiple future moments, to reduce the dimensionality and computational complexity, variable-step sampling or piecewise parameterization methods can be used to represent the continuous control sequence in the form of finite basis function coefficients. Alternatively, a future constant assumption can be adopted to keep the control quantity constant within several prediction steps, thereby improving real-time performance and computational efficiency without significantly affecting performance.
[0047] Within the predicted timeframe, the system employs a comprehensive cost function to weightedly sum different control objectives, aiming to achieve an optimal balance between flight accuracy, energy gain, and structural safety. The total cost comprises multiple weighted cost terms:
[0048] in The cost of trajectory tracking error; Tracking weights; Wear and tear on control components; Energy gain at a cost; Energy weighting; Energy consumption costs of the implementing agency; Energy consumption weight; Safety and constraint penalties; : Corresponding weights; In order to eliminate the huge differences in numerical magnitude between position (meter), angle (radian) and energy (joule), each cost term is normalized or scaled before weighting to ensure that the optimizer can consider each control objective in a balanced way. The trajectory tracking error cost describes the deviation between the UAV's desired trajectory and the predicted trajectory. It can be represented by a weighted integral of the sum of squares of position error and attitude error to reflect the flight accuracy requirements. The actuator energy consumption cost penalizes the electrical energy consumed by the control surfaces and photovoltaic wing drive motors in the prediction time domain, thereby ensuring that the net energy gain of the strategy selection is positive. The controller wear cost measures the combined consumption of control surface action and servo start-stop frequency, usually represented as a linear combination of the sum of squares of control surface deflection angle changes and the number of start-stop frequencies. The energy gain cost reflects the gain of photovoltaic power generation in the prediction time domain, usually taken as a negative value of the generated energy, i.e., the higher the energy gain, the lower the total cost. It can also be equivalently achieved by using a reward. The constraint penalty term penalizes situations such as attitude over-limit, acceleration over-limit, or energy system out-of-bounds, so that the controller can still output a feasible and safe control solution under disturbance or uncertainty conditions.
[0049] This optimization problem requires simultaneous satisfaction of multiple constraints, including attitude stability, flight mechanics, energy system constraints, and actuator physics constraints. Specifically, these include: attitude angles and angular velocities must not exceed stability margins; angle of attack must not exceed stall thresholds; overload and linear acceleration must not exceed structural limits; photovoltaic power generation is limited by converter rated current and panel temperature; control surface deflection angles, angular rates, and drive currents must all meet hardware constraints; and altitude, speed, and time window constraints are set according to mission requirements. To ensure smooth control commands and prevent command abrupt changes caused by sudden shifts in higher-level weights, the system must limit the rate of change of the optimization variables within adjacent control cycles. This rate of change must be less than or equal to a preset maximum rate of change limit. Furthermore, to improve robustness, the system introduces a soft penalty mechanism into the constraints. When a disturbance causes temporary infeasibility, a cost-smoothing correction is used to avoid solution failure.
[0050] Due to the nonlinear coupling characteristics of this optimization problem, the system employs multiple engineering solution strategies to meet the requirements of embedded real-time control. In the linearized approximation scenario, the UAV dynamics model is linearized at the current operating point, making the cost function quadratic and the constraints approximately linear. A high-performance quadratic programming solver can quickly obtain an approximate optimal solution. To maintain nonlinear accuracy, sequential quadratic programming or interior-point methods are used for iterative solutions, with a balance between accuracy and real-time performance achieved by limiting the number of iterations and setting an upper limit on the solution time. A warm-start mechanism is also employed, using the optimal solution from the previous control cycle as the initial point to improve convergence speed. If the solution fails to complete within the timeout period, the solution from the previous cycle or a fast approximation based on linearization is automatically adopted to ensure continuous and stable control output. When computational resources are limited, real-time performance can be maintained by reducing the order of states or shortening the prediction step size. The system can utilize high-performance optimizers suitable for embedded real-time control, such as fast quadratic programming or real-time nonlinear programming solvers designed for MPC, with a preset maximum computation time for each solution, typically not exceeding half of the current control cycle. The entire rolling optimization control execution process includes the following steps: In each control cycle, the underlying execution layer reads the current state estimate and environmental data, updates the importance weights of each objective, and calls the comprehensive prediction model to perform rolling predictions of the state at several future moments; subsequently, in the prediction time domain, it calculates cost terms such as trajectory, energy, wear, and constraint penalties to form a comprehensive cost function and performs constrained optimization; the system outputs the optimal control command for the current cycle, i.e., the first step of the optimal sequence, driving the control surfaces and solar control wings to perform corresponding actions; after the next control cycle arrives, the system reacquires the state and repeats the above process. Through this rolling closed-loop mechanism, the controller can achieve a dynamic balance and optimal coordination of flight performance, energy gain, and structural safety in complex tasks and multi-disturbance environments, ensuring that the UAV performs flight and energy management tasks in an adaptive, stable, and efficient manner throughout the entire process.
[0051] Sp5: After the central processing unit obtains the optimal control command through optimization, the output includes the deflection angles of each solar control wing and the conventional control surface. These commands represent the optimal attitude adjustment and energy harvesting strategy that the UAV should adopt in the current flight state. The role of the actuator control module is to map these high-level control quantities into low-level electrical signal commands that can be directly executed by the drive system, realizing a closed-loop conversion from control commands to physical actions. In the specific implementation process, the module first filters and smooths the input angle commands to eliminate high-frequency jitter caused by the optimization algorithm and prevent structural oscillations and mechanical wear caused by frequent reverse movements of the servos. The filtered commands are then corrected through anti-saturation and safety limiting mechanisms. The system dynamically adjusts the upper and lower limits of the angles according to the mechanical limits, attitude constraints, and real-time health status of each servo and solar control wing, so that the control signals are always kept within the safe and feasible domain, avoiding overload or abnormal deflection. Subsequently, the execution module performs time planning and trajectory generation for angle commands based on the actual position and velocity information of the control surfaces. It generates a smooth transition curve using continuous acceleration or polynomial interpolation, ensuring the angle, angular velocity, and angular acceleration of the control surface motion remain continuous over time, reducing mechanical impact and energy loss. Based on trajectory planning, the system achieves precise angle tracking through closed-loop servo control. The actual deflection angle of the servo motor is detected in real-time by a high-precision encoder. After comparing the error with the desired angle, the controller calculates the corresponding drive current or voltage command based on proportional, integral, and derivative laws, combined with model feedforward compensation, thereby achieving fast, stable, and high-precision position tracking. This process constitutes a key feedback closed loop from the central control layer to the actuator layer. During execution, the drive units of each control surface and solar control wing continuously output real-time status feedback, including information such as actual deflection angle, drive current, temperature, and dynamic response characteristics. This feedback data is synchronously transmitted to the state estimation module, which is used not only to correct the actuator model and improve the control accuracy of the next cycle but also to assess servo motor wear and actuator health, supporting the system's self-learning and predictive maintenance. The central control layer then dynamically adjusts energy allocation and attitude strategies based on this data, achieving deep synergy between flight control and energy utilization.
[0052] Before mission initiation or during flight, the system first loads overall mission and environmental information, including mission waypoint sequence, mission priority, estimated flight time, current location, initial battery level, and geographic location and weather forecast information to calculate solar trajectory and illumination conditions. Simultaneously, the sensor acquisition module starts in real-time, continuously sampling the UAV's attitude, illumination intensity, wind speed, and battery level. The state estimation module fuses this data and outputs initial state estimates, providing accurate flight state information for subsequent control. The advanced decision layer, based on the overall mission and environmental information, invokes a pre-trained deep reinforcement learning model to generate a set of importance indicators in real-time. These indicators reflect the current priority of multiple objectives, such as trajectory tracking, energy gain, and control component wear. These indicators are written to shared memory for the underlying execution layer to read. If an anomaly or model failure is detected, the advanced decision layer will activate preset default indicators and generate a fault report to ensure safe system operation.
[0053] The underlying execution layer reads the latest state estimate and importance indicators within each control cycle, and invokes the integrated prediction model to simulate the state transitions for several future steps within the prediction time domain. The integrated prediction model consists of a flight dynamics model and an energy model. The flight dynamics model calculates additional lift, drag, and torque changes based on the UAV's attitude, speed, photovoltaic wing deflection angle, and control surface angle. The energy model calculates instantaneous power generation based on the solar azimuth, UAV attitude, and photovoltaic wing deflection angle. During the prediction process, the underlying execution layer constructs a multi-objective weighted cost function based on the predicted state. The cost components include trajectory tracking error, control component wear, and negative energy gains. The optimization variables are the photovoltaic wing deflection angle and the conventional control surface angle. Constraints include attitude stability, flight dynamics limitations, and power generation output range. The optimization problem is processed by a real-time solver to obtain the optimal control sequence for each control cycle. The system then sends the optimal control commands, including the photovoltaic wing deflection angle and the conventional control surface angle, to the actuator control module. The module smooths and limits the commands before execution, while simultaneously collecting feedback data such as actual control surface angles, wing deflection angles, and drive currents, which are then transmitted back to the state estimation module for state updates in the next cycle, forming a closed-loop control system. Throughout the flight, the system continuously records raw sensor data, state estimation results, optimization results, control commands and execution feedback, and alarm information. When a safety threshold is triggered, such as low battery power, sensor failure, or actuator malfunction, the system automatically enters a degraded or safety mode, such as a forced energy priority mode, a backup state estimation mode, or a backup control surface or inertial stabilization strategy, to ensure flight safety and mission continuation. Simultaneously, the ground station receives summary information for manual intervention and mission adjustments. Specific Implementation Example 2:
[0055] like Figures 1 to 2As shown, based on the content of the above specific embodiments, the following content is further disclosed:
[0056] In this embodiment, the UAV's autonomous flight control system first acquires multi-source raw data through the sensor acquisition module, including accelerometer and gyroscope data from the inertial measurement unit, GPS positioning data, illumination and solar position sensor information, wind speed sensor readings, and parameters such as battery voltage and current. All data undergoes time synchronization and noise reduction processing before entering the control system to ensure accuracy and consistency. The inertial data is fused with GPS data through complementary filtering or expansion and an unscented Kalman filter to obtain high-precision estimates of the UAV's position, velocity, and attitude. Illumination and solar sensor data are low-pass filtered and compared with the solar position algorithm output for correction to obtain accurate sunlight incidence direction. Battery voltage and current data are averaged through a sliding window and Kalman filtered to estimate the battery's internal impedance and remaining energy, achieving accurate prediction of charge state. Wind field estimation combines airborne speed, airspeed sensor readings, and aircraft attitude information. Real-time wind speed and direction estimation is performed using differential calculations and inverse kinematics of the observation model. Extended Kalman filtering or recursive least squares methods are employed for online identification, and uncertainties are described in probability distribution form when necessary for robust control in subsequent optimization. The system identifies abnormal sensor data using thresholding and multi-sensor consistency detection, isolating data exceeding physical range or exhibiting abrupt changes, and replacing them with redundant sensors or estimates to ensure continuous and reliable state estimation. All intermediate estimation results, including instantaneous power generation and cumulative available energy in the near future, are written to shared memory for use in high-level decision-making strategy calculations and log recording. The system adopts a hierarchical intelligent control architecture. The high-level decision-making layer has a longer cycle and is responsible for updating the importance indicators of each control objective based on overall task and environmental information; the lower execution layer has a shorter cycle and solves for the optimal control command in real time for rapid response. The system's priority determination logic is driven by the output indicators of the high-level decision-making layer. When the system is in an energy-priority state, such as when battery power is low or sunset is approaching, the higher-level decision layer outputs high energy gain importance indicators. The lower-level controller increases the weight of the energy term in multi-objective optimization, adjusting the solar control wing to the maximum angle of alignment with the sun to obtain maximum power generation. Simultaneously, it calculates compensating control commands, such as control surface adjustments, to counteract the additional torque caused by wing deflection, thereby ensuring flight stability. When the system is in a control-priority state, such as when performing complex maneuvers or encountering strong winds, the higher-level decision layer outputs high trajectory tracking importance indicators. The lower-level optimization focuses more on reducing trajectory deviation and response time. Solar wing deflection is restricted or used in conjunction with control surfaces to achieve minimum drag and fastest response. During priority switching, the output indicators of the higher-level decision layer transition through a time-series smoothing strategy based on polynomial interpolation, ensuring that the changes in each weight indicator are continuous and bounded within an inference cycle of 10 seconds to 1 minute, avoiding the introduction of step disturbances to the lower-level control system.If a conflict arises between high-level policies and low-level security constraints, the low-level policy will reject dangerous actions by setting constraints and prioritizing them, and will then report back to the high-level policy, triggering policy adjustments or a switch to a new security mode.
[0057] The advanced decision-making layer employs a reinforcement learning deep neural network model pre-trained in a simulation environment. This environment incorporates a comprehensive prediction model, including flight dynamics and energy models, to predict the impact of solar array deflection on lift, drag, torque, and power generation. It supports randomization across multiple scenarios, including solar trajectory variations, cloud cover, wind fluctuations, mission configuration differences, and equipment aging models. Training aims to maximize long-term cumulative rewards, with the reward function considering energy utilization, mission completion efficiency, flight safety indicators, and device lifespan. After training, the model is validated on a hardware-in-the-loop platform and then subjected to small-scale flight tests in controlled environments, gradually expanding to complex operating conditions. Data generated from simulations and experiments is stored in logs and the learning module for subsequent model iterations and optimizations.
[0058] The central processing unit (CPU) consists of a high-performance real-time embedded processor. Its software architecture includes operating system scheduling, sensor interfaces, state estimation services, advanced decision inference, low-level optimization solutions, actuator command issuance, and log recording services. Modules transmit data via shared memory and message queues, and are precisely clock-synchronized. The sensor acquisition module collects, filters, and calibrates data in real time. The actuator control module receives optimal control commands to drive the control surfaces and solar panels, and feeds back the execution status to the state estimation module. Critical events such as sensor failure, actuator saturation, insufficient energy, or critical attitude trigger predefined safety policies, record logs, and send alarms to the ground station via the communication link. In operation, the sensor acquisition module collects raw signals and performs analog-to-digital conversion and timestamping. The state estimation and preprocessing module filters and fuses the data to generate a stable state estimate, including position, velocity, attitude, wind speed, instantaneous power generation, and battery status. Part of the state estimation results is used as initial state and constraint inputs for the low-level predictive controller, while another part is aggregated with short-term energy predictions and disturbance indicators and sent to the high-level decision layer. Importance indicators output by the high-level layer are written to shared memory. The underlying predictive controller reads the latest indicators, calls the integrated prediction model to simulate the future state within the prediction time range, and constructs a multi-objective cost function. Under attitude stability constraints, flight mechanics constraints, and power generation constraints, it solves for the optimal control sequence and issues the initial control command to the actuator control module. After the actuator executes its action, it reports back the status, forming a closed-loop control. All raw data and key intermediate quantities are periodically written to the log. Key events and statistical data are uploaded to the ground station via the communication link. When offline, the log and learning modules upload flight data to the training server for model iteration and optimization. The multi-objective optimization formula can be described in Chinese as follows: Within the prediction time range, multiple costs such as trajectory tracking error, control surface and motor wear, and energy gain are weighted and summed to obtain the total cost value. The optimization variables are the solar panel deflection angle and control surface angle, and the constraints are attitude stability, flight mechanics, and power generation output constraints. Through the above process, this embodiment realizes a complete autonomous flight closed loop from multi-source sensor data acquisition, state estimation, wind field prediction, high-level decision output, low-level predictive control solution, actuator command execution, to closed-loop feedback and log recording, ensuring the coordinated optimization of energy harvesting and flight control, and providing data support for the offline training and iteration of intelligent control models. Specific Implementation Example 3:
[0060] like Figures 1 to 2 As shown, based on the content of the above specific embodiments, the following content is further disclosed:
[0061] The following implementation examples will further illustrate this.
[0062] Case background: Long-range reconnaissance mission.
[0063] The drone is conducting a long-range, long-endurance reconnaissance mission. The current time is 3:30 PM, with 2.5 hours remaining until the scheduled sunset. The battery charge is 40%, below the safe threshold of 50%. The ambient wind speed is stable at 3 meters per second, and the target's flight path is a straight cruise.
[0064] At the start of the mission, the system first acquires real-time state parameters of the UAV, including attitude, position, speed, light intensity, wind speed, and battery level, through a sensor acquisition module. These parameters are then fused and calculated by a state estimation module to obtain a precise current flight state. After analyzing this data, the advanced decision-making layer invokes a pre-trained deep reinforcement learning model. Based on the low battery status and the approaching sunset, it dynamically outputs a set of importance indicators for the current moment: trajectory tracking has a weight of 0.4, energy gain has the highest weight of 0.9, execution power has a weight of 0.8 to suppress ineffective adjustments, and component wear has a weight of 0.5. This set of weights is smoothed and written to shared memory for the underlying execution layer to read and use. Based on this, the system determines that an energy-priority strategy should be adopted, with the core objective of maximizing photovoltaic power generation while ensuring flight attitude stability and rational energy regulation.
[0065] The underlying execution layer then invokes the integrated prediction model to simulate and extrapolate the flight state for several future steps within the prediction time domain. The integrated prediction model consists of a flight dynamics model and an energy model. The energy model calculates the power generation of the photovoltaic panel if it deflects to its maximum solar angle under the current attitude; the results show that the power can be increased from 30... Increased to 45 The flight dynamics model predicts that this deflection will introduce a leftward roll moment and additional drag. Subsequently, a multi-objective weighted cost function is constructed based on model predictive control methods. Among the cost terms, the energy gain term is optimized with the highest weight, while also considering execution energy consumption, flight stability, and constraints. The real-time optimizer then solves for the optimal control sequence for the current cycle. In this sequence, the photovoltaic wing deflection angle is commanded to 15° to obtain the maximum solar angle, while the conventional control surface angles are compensated through co-optimization calculations, such as a 3° downward deflection of the right aileron and a 0° deflection of the left aileron, to counteract the roll moment generated by the photovoltaic wing deflection and ensure flight attitude stability. This optimal control sequence, after constraint verification, meets the attitude angle safety range and power output constraints, with energy gain significantly exceeding execution energy consumption, achieving a net energy gain.
[0066] After receiving the optimal control commands, the actuator control module performs smoothing interpolation and amplitude limiting on the photovoltaic wing deflection angle and control surface angle, and then drives the actuator to move. The UAV maintains stable attitude during the deflection process, and the power generation immediately increases to 45. The state estimation module transmits newly acquired power and attitude information back to the system in real time for use in the next cycle, forming a dynamic closed-loop optimization. Through this hierarchical control strategy based on importance index scheduling, the UAV achieves coordinated control of energy harvesting and attitude stabilization when battery power is insufficient, significantly extending endurance and maintaining flight safety.
[0067] Case background: High-precision path tracking task.
[0068] The UAV is performing a high-precision path tracking mission, requiring it to maintain a precise "S"-shaped trajectory while traversing a canyon area. During flight, the synthetic wind field estimation module detects a one-second transient cross gust with a wind speed reaching ten meters per second, while the battery's state of charge is high at 85%, indicating ample energy reserves. At this moment, the sensor acquisition module detects a 5° roll deviation in attitude, and the state estimation module reports this anomaly to the higher-level decision-making layer. The reinforcement learning decision model, combining the high-precision trajectory requirements of the task layer with the sudden disturbance information detected by the system layer, immediately adjusts the importance indicators, outputting a trajectory tracking weight of 0.95, a response speed weight of 0.9, an energy gain weight of 0.1, and an execution energy consumption weight of 0.1. The system smoothly switches weights within one second using polynomial interpolation to avoid abrupt changes in the cost function that could lead to control instability. Based on this, the decision-making layer determines the current state to be a control-priority mode, instructing the lower-level execution layer to prioritize rapid response and suppress attitude errors, temporarily reducing the importance of energy harvesting.
[0069] After reading the new importance indicators, the underlying execution layer calls the integrated prediction model and the generalized disturbance observer to estimate the disturbance torque generated by the gusts in real time. A multi-objective optimization problem is then constructed using a model predictive control framework, with trajectory error and attitude error as the main cost terms. The optimization solution indicates that the photovoltaic wing should remain nearly flat, allowing only a small deflection of 0° to ±2°, and, if necessary, a small coordinated movement of about 3° to assist the conventional control surfaces in counteracting the disturbance torque. Simultaneously, the conventional control surfaces are instructed to deflect at a large angle in the opposite direction, for example, the aileron on the windward side deflects upward by 10°, to maximize the counteraction of gust disturbances. Control increment constraints ensure that the control surface movements are rapid but do not exceed the safe angular rate range. Although the energy gain term has a very low weight at this time, the system allows the motors to output a higher current to provide instantaneous thrust, achieving rapid attitude recovery. After receiving the command, the actuator control module drives the servos at the maximum safe rate, and the UAV's attitude stabilizes within 0.5 seconds, with trajectory deviation suppressed to an acceptable range. The photovoltaic power generation capacity decreased slightly temporarily, but after the disturbance disappeared, the senior decision-makers readjusted the importance indicators, and the system smoothly returned to normal cruise status and continued to perform the original flight path mission.
[0070] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising a reference structure" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0071] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for autonomous flight control of a drone combining photovoltaic panel flip angle adjustment, characterized in that: The control method includes the following steps: Sp1: Data Acquisition and Classification: Acquire the overall mission, environmental information, and current flight data of the UAV; the overall mission and environmental information are used for the high-level decision-making layer, and the current flight data are used for the low-level execution layer. Sp2: Macro Strategy Generation: The advanced decision-making layer adopts a deep neural network decision-making model based on reinforcement learning. Based on the overall task and environmental information, it outputs a set of dynamically changing importance indicators in real time. The importance indicators are used to reflect the current priority of multiple control objectives such as flight tracking and energy gain. Sp3: Low-level predictive control: The low-level execution layer adopts a model-based predictive control framework and incorporates a comprehensive predictive model of energy, aerodynamics, and control; the comprehensive predictive model includes the coupling relationship between the deflection angle of the solar control wing, the angle of the conventional flight control surfaces, and the aerodynamic performance and energy gain of the UAV; Sp4: Multi-objective optimization solution: The underlying execution layer calculates the future state transition of the UAV in the prediction time domain based on the comprehensive prediction model, and substitutes the prediction results into the multi-objective optimization evaluation formula for solution, so as to obtain a set of optimal control commands that simultaneously satisfy multiple objectives in each control cycle; Sp5: Cooperative command execution: The optimal control commands include the deflection angles of all solar control wings and the angles of all conventional flight control surfaces; the UAV executes the optimal control commands to achieve autonomous flight.
2. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The control method described herein is applicable to unmanned aerial vehicles (UAVs) equipped with solar-powered control wings. The solar-powered control wings serve as both an energy harvesting component and an auxiliary flight attitude control component for the UAV. The solar-powered control wings include a photovoltaic wing plate that can deflect around a rotation axis and a drive actuator connected to the photovoltaic wing plate. The surface of the photovoltaic wing plate is provided with a photovoltaic cell array. Under the drive actuator, the photovoltaic wing plate changes its own deflection angle, thereby adjusting the incident angle of the photovoltaic array and generating additional aerodynamic forces and torques during the deflection process, thereby achieving synergistic optimization of energy harvesting and auxiliary attitude control.
3. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The comprehensive prediction model consists of a flight dynamics model and an energy model. The flight dynamics model is used to calculate the additional lift, drag, and torque changes generated by the flipping attitude of the solar-powered control wing. The energy model is used to calculate the total instantaneous power generation based on the direction of sunlight, the attitude of the UAV, and the flipping angle of the solar-powered control wing.
4. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The overall mission and environmental information acquired by the advanced decision-making layer includes at least one of the following: the current battery level of the UAV, the real-time position of the direction of sunlight, the estimated wind speed, the intensity of airflow, the time until sunset, and the distance to the mission target point.
5. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The multi-objective optimization evaluation formula is a function that sums multiple costs over a prediction time range. These costs include at least: a trajectory tracking error term describing the accuracy of the flight path, a control component wear term describing the wear of the control surfaces and motors, and an energy gain term describing the power generation revenue. The optimization variables of the multi-objective optimization evaluation formula include the deflection angle of the solar control wing and the deflection angle of the conventional control surfaces. The constraints include UAV attitude stability constraints, flight mechanics constraints, and power generation output constraints.
6. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The higher-level decision-making layer operates at a longer interval to update the importance index, while the lower-level execution layer operates at a shorter interval to solve for the optimal control command, thus forming a hierarchical intelligent control system.
7. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: When the advanced decision-making layer determines that the UAV is in an energy-priority state based on overall mission and environmental information, it outputs a higher energy gain importance index; accordingly, the underlying execution layer prioritizes adjusting the solar control wings to maximize alignment with the sun, while simultaneously calculating compensatory control commands to counteract the aerodynamic disturbances generated.
8. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The advanced decision-making layer determines that the UAV is in a control-priority state based on overall mission and environmental information. When performing complex maneuvers or encountering strong winds, it outputs a high trajectory tracking importance index. The underlying execution layer then coordinates the use of solar-powered control wings and conventional flight control surfaces to achieve flight control with minimal drag and the fastest response.
9. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The intelligent learning decision-making model adopted by the advanced decision-making layer is obtained by pre-training in a simulation environment based on historical flight data and mission environment; the simulation environment has the built-in energy, aerodynamic and control integrated prediction model, and the intelligent learning decision-making model learns the optimal importance index output strategy by maximizing the total reward over a long period.
10. The autonomous flight control method for a drone combined with photovoltaic panel flip angle adjustment according to claim 1, characterized in that: The control method is executed by the central processing unit in the UAV flight control system, and forms a closed-loop control structure by combining the sensor acquisition module and the actuator control module. The sensor acquisition module includes an attitude sensor, a light intensity sensor, a wind speed sensor, and a battery power detection module, which collects the UAV's flight status and environmental parameters in real time. The actuator control module includes a control surface driver, a photovoltaic wing deflection driver, and a motor controller, and receives the optimal control commands output by the central processing unit and drives the corresponding actuators to move.