Intelligent control method for airborne optoelectronic stabilized platform

By constructing an integrated electromechanical environment simulation platform and a teacher-student controller architecture, the problem of insufficient control performance of traditional optoelectronic stabilization platforms in complex environments was solved, enabling rapid migration and efficient control strategy optimization, and improving the adaptability and robustness of airborne optoelectronic stabilization platforms.

CN122172658APending Publication Date: 2026-06-09CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
Filing Date
2026-02-06
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional optoelectronic stabilization platform control algorithms struggle to cope with model uncertainties and parameter disturbances in complex airborne environments. Existing simulation platforms lack integrated fusion of electromechanical systems, external environment, and mission requirements, making it difficult to achieve comprehensive verification and optimization.

Method used

An integrated electromechanical environment simulation platform is constructed, adopting a teacher-student controller architecture. Through privileged learning and meta-reinforcement learning, the teacher controller learns the optimal control strategy in the high-fidelity simulation platform, while the student controller transfers the learning strategy to the actual environment and adjusts the control strategy using airborne sensor information.

Benefits of technology

It achieves rapid and stable control performance improvement in complex airborne environments, possesses high adaptability and robustness, reduces dependence on models, and improves learning efficiency and control performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122172658A_ABST
    Figure CN122172658A_ABST
Patent Text Reader

Abstract

This invention relates to the field of optoelectronic imaging technology, specifically providing an intelligent control method for an airborne optoelectronic stabilization platform. The method involves constructing a simulation platform for the airborne optoelectronic stabilization platform that integrates electromechanical environment and task functions. The simulation platform outputs privileged information and observable information from airborne sensors. A teacher-student privileged learning architecture is constructed, where the teacher controller's input is privileged information and its output is the teacher's control strategy. The teacher controller is trained based on the privileged information, and the student controller learns and imitates the teacher's control strategy based on the observable information from the airborne sensors. After training on the simulation platform, the student controller is further trained through meta-reinforcement learning. The student controller's input is the actual observation information from the airborne sensors, and its output is the actual control strategy for the airborne optoelectronic stabilization platform. The trained student controller is then deployed on the airborne optoelectronic stabilization platform, generating the actual control strategy based on the real-time acquired actual observation information from the airborne sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of optoelectronic imaging technology, and in particular relates to an intelligent control method for an airborne optoelectronic stabilization platform. Background Technology

[0002] Electro-Optical Stabilized Platforms (EOSPs) are widely used in aviation, UAVs, and shipboard applications for Earth observation, tracking, and imaging. Their primary task is to maintain line-of-sight stability and accurately point to targets under various flight attitudes and external disturbances.

[0003] In complex airborne operating environments, optoelectronic stabilization platforms face a variety of disturbances, such as wind field variations, load fluctuations, flight attitude coupling, temperature drift, and cable elasticity changes. These factors collectively lead to significant uncertainties in the system model, with frequent external disturbances that are difficult to model accurately.

[0004] Traditional control algorithms for optoelectronic stabilization platform controllers (PID, LQG, DOB) typically rely on system models to design control parameters. While some adaptive and robust control methods can address model uncertainties and parameter disturbances to a certain extent, these methods remain limited by the preset model structure and parameter tuning range when the load conditions, disturbance environment, flight attitude, or mission requirements of the optoelectronic stabilization platform change significantly. This makes it difficult to maintain stable control performance. Currently, the industry lacks an integrated simulation platform that combines electromechanical systems, external environment, and mission requirements. Existing simulation platforms are mostly isolated simulations from different disciplines, failing to establish effective linkage between algorithmic control, mechanical dynamics, and external disturbances. This makes it difficult to comprehensively verify and optimize control strategies for specific missions. Summary of the Invention

[0005] In view of this, the present invention aims to provide an intelligent control method for an airborne optoelectronic stabilization platform. By constructing an integrated electromechanical environment and mission simulation platform for the airborne optoelectronic stabilization platform, the performance, robustness, and reliability of the control strategy under real-world conditions can be systematically and comprehensively evaluated during the design phase for different mission scenarios. This integrated simulation platform simultaneously covers electromechanical dynamics models, typical airborne disturbance environments, and various mission requirements (such as attitude steady-state control, target tracking, and disturbance rejection control), providing a highly realistic and scalable simulation foundation for the training, testing, and optimization of control strategies. Based on this, the present invention employs a "teacher-student" privileged learning architecture for controller model training. The teacher controller learns the optimal control strategy offline in a simulation platform with complete privileged information (including the internal state of the electromechanical system, environmental disturbances, and mission parameters), avoiding the limitation of difficulty in obtaining global information under real-world conditions. Subsequently, the student controller is trained only based on actual observation information from a limited number of practically equippable airborne sensors. By mimicking the teacher's control strategy and combining it with a meta-reinforcement learning mechanism, a rapid and stable migration from the simulation platform to the real flight environment is achieved.

[0006] To achieve the above objectives, the technical solution created by this invention is implemented as follows: This invention provides an intelligent control method for an airborne optoelectronic stabilization platform, comprising: S1: A simulation platform for building an airborne optoelectronic stabilization platform that integrates electromechanical environment and mission. The simulation platform outputs two types of information, including privileged information and observable information from airborne sensors. S2: Construct a teacher-student controller based on a privileged learning architecture. The teacher-student controller includes a teacher controller and a student controller. The input of the teacher controller is privileged information, and the output of the teacher controller is the teacher control policy. The teacher controller is trained based on the privileged information, and the student controller learns the teacher control policy by imitating the observable information from the airborne sensors. S3: The student controller is trained through meta-reinforcement learning. During the meta-reinforcement learning process, the input of the student controller is the actual observation information of the airborne sensor, and the output of the student controller is the actual control strategy of the airborne optoelectronic stabilization platform. S4: The trained student controller is deployed on the airborne opto-stabilized platform. The student controller generates the actual control strategy based on the real-time observation information collected by the airborne sensors.

[0007] Preferably, the privileged information of the simulation platform includes: the internal state of the electromechanical system, environmental disturbances, and task parameters. Represented as: ; in: This refers to the internal state of the electromechanical system. For environmental disturbance, These are task parameters.

[0008] Preferably, the teacher controller is trained using a proximal policy optimization algorithm, and the objective function of the pruning policy in the proximal policy optimization algorithm is... Represented as: ; in: ; Prior teacher control strategies, For current teacher control strategies, These are the trainable parameters for the teacher controller. It is the ratio of the probability of a certain action to the probability of the prior teacher control strategy and the current teacher control strategy. For the dominant function, This is a truncation term used to control the training-to-update ratio within a certain range. , This is a hyperparameter for the cropping range. This refers to the expected trajectory of interaction between time and environment.

[0009] Preferably, the composite reward function for teacher controller training for: ; in: To control the weighting coefficients of reward items, As a reward for stable posture, To reward tracking accuracy, For energy consumption constraints, For perturbation robustness rewards; Among them: attitude stability reward Represented as:

[0010] in: This is the attitude angle error vector. This is the platform's angular velocity vector; Tracking accuracy bonus Represented as:

[0011] in: To track errors, Represented as: ; in: For the output of the simulation platform, For task commands; Energy consumption constraints Represented as:

[0012] in: This refers to the motor drive torque or control input vector of the airborne optoelectronic stabilization platform. Perturbation Robustness Reward Represented as:

[0013] in: This represents the change in environmental disturbance.

[0014] Preferably, the loss function of the student controller mimics the teacher's control strategy. for: ; in: The parameters of the student controller are used to mimic and learn the teacher's control strategy. These are trainable parameters for the teacher's control strategies. The strategy function is used by the student controller to mimic the teacher's control strategy. For the teacher's control strategy, the strategy function is... For complete privileged information observed by the teacher controller, The observable information acquired by the student controller through onboard sensors. For expectation operators.

[0015] Preferably, meta-reinforcement learning includes an inner loop and an outer loop, with the inner loop iteration formula being: ; in: For the i-th task, each task represents a perturbation configuration. These are the initial parameters for the objective of meta-reinforcement learning training. These are the parameters of the student controller after the last gradient update. The learning rate is used to control the step size of policy updates in each task. To calculate policy parameters in the task Directional improvements; The outer loop iteration formula is: ; in: For the outer loop learning rate, The accumulated gradients across all tasks are used to update the meta-parameters. , To calculate policy parameters in the task Directional improvements.

[0016] Preferably, an online fine-tuning mechanism is introduced into the airborne optoelectronic platform, which adjusts the parameters of the student controller based on the actual observation information collected in real time from the airborne sensors.

[0017] Compared with the prior art, the present invention can achieve the following beneficial effects: Traditional control algorithms for optoelectronic stabilization platforms (PID, LQG, DOB) typically rely on system models to design control parameters. While some adaptive and robust control methods can address model uncertainties and parameter disturbances to some extent, they remain limited by the preset model structure and parameter tuning range when the load conditions, disturbance environment, flight attitude, or mission requirements of the optoelectronic stabilization platform change significantly. This makes it difficult to maintain stable control performance. Currently, the industry lacks an integrated simulation platform that combines electromechanical systems, external environment, and mission requirements. Existing simulation platforms are mostly isolated simulations from different disciplines, failing to establish effective linkage between algorithmic control, mechanical dynamics, and external disturbances. This makes it difficult to comprehensively verify and optimize control strategies for specific missions. Compared with existing airborne optoelectronic stabilization platform controller design methods, this invention has significant technical advantages. It resolves the contradiction between the need for rapid iteration in algorithm development and the high cost and high risk of real-world environments by building a simulation platform that integrates electromechanical environment and mission requirements for an airborne optoelectronic stabilization platform. This invention eliminates the need for precise mathematical modeling, allowing the teacher control strategy to learn optimal control laws directly in a high-fidelity simulation platform through a privileged learning phase, effectively reducing reliance on models. Utilizing a meta-reinforcement learning mechanism, the airborne sensors can rapidly adjust parameters and transfer strategies for different mission objectives when facing various disturbances, loads, and flight attitude changes, significantly improving the system's adaptability and environmental robustness. Compared to traditional reinforcement learning, which requires a large amount of interactive data and training, and addresses the domain gap problem between simulation and real-world deployment, this invention introduces a teacher controller supervision through a teacher-student two-stage architecture, greatly accelerating the convergence speed of the student controller and improving learning efficiency. Through simulation-training-verification-deployment, it resolves the contradiction between the safety and reliability requirements of optoelectronic stabilization platform development and technological uncertainties. Furthermore, this invention fully leverages the complementarity between privileged information and observable information from airborne sensors. During the training phase, the teacher controller acquires full-state information of the system for optimal action learning, while the student controller relies solely on airborne sensor input for generalization and transfer, thereby improving control performance without increasing sensor load. Finally, the airborne sensors of this invention maintain stability and real-time performance under conditions of multiple disturbances and tasks, exhibiting stronger generalization and online adaptive capabilities, enabling high-precision, low-energy attitude stabilization control in complex airborne environments. Attached Figure Description

[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a technical flowchart of the intelligent control method for an airborne optoelectronic stabilization platform provided according to an embodiment of the present invention; Detailed Implementation To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and do not constitute a limitation thereof. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of the invention. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, some operations related to the invention are not shown or described in the specification. This is to avoid obscuring the core parts of the invention with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; the relevant operations can be fully understood based on the description in the specification and general technical knowledge in the art.

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined to form various implementations. Furthermore, the order of the steps or actions in the method description can be changed or adjusted in a manner readily apparent to those skilled in the art. Therefore, the various orders in the specification and drawings are merely for the clear description of a particular embodiment and do not imply a mandatory order, unless otherwise stated that a particular order must be followed.

[0020] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0021] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0022] The invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] Please see Figure 1 In one embodiment of the present invention, an intelligent control method for an airborne optoelectronic stabilization platform is provided, comprising: S1: A simulation platform for building an airborne optoelectronic stabilization platform that integrates electromechanical environment and mission. The simulation platform outputs two types of information, including privileged information and observable information from airborne sensors. S2: Construct a teacher-student controller based on a privileged learning architecture. The teacher-student controller includes a teacher controller and a student controller. The input of the teacher controller is privileged information, and the output of the teacher controller is the teacher control policy. The teacher controller is trained based on the privileged information, and the student controller learns the teacher control policy by imitating the observable information from the airborne sensors. S3: The student controller is trained through meta-reinforcement learning. During the meta-reinforcement learning process, the input of the student controller is the actual observation information of the airborne sensor, and the output of the student controller is the actual control strategy of the airborne optoelectronic stabilization platform. S4: The trained student controller is deployed on the airborne opto-stabilized platform. The student controller generates the actual control strategy based on the real-time observation information collected by the airborne sensors.

[0024] In step S1, the simulation platform of this embodiment adopts a scalable modeling framework for an airborne optoelectronic stabilization platform that integrates electromechanical environment and task requirements. This framework can integrate various potential influencing factors according to the actual application needs of the optoelectronic stabilization platform. The simulation platform includes not only an electromechanical dynamics model but also an environmental disturbance model and a task parameter randomization model. The electromechanical dynamics model determines the internal state of the electromechanical system, such as the simulation platform's attitude angle, angular velocity, and driving torque. The environmental disturbance model determines typical airborne environmental disturbances, such as wind field changes, load fluctuations, and structural nonlinearities, introducing more complex environmental and electromechanical factors such as frictional drift, temperature effects, material aging, uncertainties in airborne sensors, and multi-source coupling interference. The task parameter randomization model is used to set task parameters (target pointing, tracking angle, task commands, etc.) according to task requirements for simulation training. Through this scalable modeling mechanism, the simulation platform can support the construction of multi-level scenarios from ideal to extreme conditions, providing comprehensive, safe, and systematic evaluation and training conditions for control strategies.

[0025] The complete state (privileged information) of the simulation platform is defined as follows: ; in: This refers to the internal state of the electromechanical system, such as the attitude angle, angular velocity, joint angle, joint velocity, and driving torque of the simulation platform. For environmental disturbances (wind-driven models, friction parameters, load coefficients, temperature drift, etc.). These are the mission parameters (target pointing, tracking angle, mission commands, etc.).

[0026] The simulation platform modules and implementation methods of the airborne optoelectronic stabilization platform integrating electromechanical environment and mission are shown in Table 1:

[0027] In step S2, this embodiment of the invention constructs a teacher-student controller based on a privileged learning architecture. The teacher-student controller includes: a teacher controller as an upper-level policy generation module, whose function is to learn the optimal reference control policy in a simulation platform with complete state information; and a student controller as an execution module ultimately deployed on an airborne opto-stabilized platform, making control decisions based solely on observable information from limited airborne sensors and achieving efficient learning by mimicking the teacher's control policy. The relationship between the two can be understood as follows: the teacher controller provides optimal behavior demonstration and supervision signals, while the student controller learns transferable teacher control policies under observable conditions of a practically deployable airborne controller.

[0028] During the training phase, the teacher controller has direct access to all privileged information within the simulation platform, including but not limited to the platform's attitude angles and angular velocities, joint positions and velocities, external disturbance torques, friction coefficients, load parameters, environmental disturbance states, and other hidden physical quantities. The teacher controller employs the Proximal Policy Optimization (PPO) algorithm for training with this privileged information. The input is the teacher control strategy, and the output is the teacher control policy. The teacher control strategy includes specific parameters such as control variables used to drive the motors or servos of the airborne opto-stabilized platform. Privileged information is also input to the teacher controller. Represented as: ; in: Complete privileged information for the teacher controller observations (the teacher's observations at time t). All privileged information provided to the simulation platform at time t, including but not limited to the simulation platform's attitude angles and angular velocities, joint positions and velocities, external disturbance torques, friction coefficients, load parameters, wind field disturbances, temperature parameters, airborne sensor noise model status, and parameters required by the actual mission.

[0029] The objective function of the pruning strategy in the proximal policy optimization algorithm during teacher controller training. Represented as: ; in:

[0030] in: Prior teacher control strategies, For current teacher control strategies, These are the trainable parameters for the teacher controller. It is the ratio of the probability of a certain action to the probability of the prior teacher control strategy and the current teacher control strategy. As the advantage function, it guides the direction of whether an action is good or bad. This is a truncation term used to control the training-to-update ratio within a certain range. , This is the clipping range hyperparameter, with a value between 0.1 and 0.2. This refers to the expected trajectory of interaction between time and environment.

[0031] To ensure the stability and effectiveness of the teacher control strategy under various working conditions, this invention designs a scalable composite reward function. Represented as: ; in: To control the weighting coefficients of reward items, As a reward for stable posture, To reward tracking accuracy, For energy consumption constraints, For perturbation robustness rewards.

[0032] An attitude stability reward is constructed based on the pitch, yaw, roll angle errors, and angular velocity deviations of the simulation platform. Represented as:

[0033] in: This is the attitude angle error vector. This refers to the angular velocity vector of the simulation platform (an IMU measurement).

[0034] A tracking accuracy reward is constructed based on the errors in the target pointing direction, target angular velocity, or mission trajectory of the simulation platform. Represented as:

[0035] in: To track errors, Represented as: ; in: The output of the simulation platform (such as pointing angle, angular velocity, target trajectory). For task commands (external input or upper-level planning).

[0036] Energy consumption constraints are constructed based on the servo motor current, motor torque, and control variable change rate of the simulation platform. Represented as:

[0037] in: This refers to the motor drive torque or control input vector of the airborne optoelectronic stabilization platform.

[0038] Disturbance robustness incentives encourage simulation platforms to maintain recoverability and stability under conditions such as wind disturbances and sudden load changes. Represented as:

[0039] in: This refers to environmental disturbances, including but not limited to changes in wind field, load, friction coefficient drift, material parameter drift caused by temperature changes, noise, or external impacts.

[0040] Because the student controller exhibits imitation error when learning the teacher's control strategy, a loss function is introduced to describe the student controller's imitation learning of the teacher's control strategy. Represented as: ; in: The parameters of the student controller are used to mimic and learn the teacher's control strategy. These are trainable parameters for the teacher's control strategies. The student controller learns the teacher's control strategy through a strategy function. It inputs actual observation information from onboard sensors and outputs the actual control strategy of the onboard photoelectric stabilization platform. This is the policy function for the teacher control strategy. The teacher controller takes privileged information as input and outputs the teacher control strategy. For complete privileged information observed by the teacher controller, The observable information acquired by the student controller through onboard sensors, such as gyroscope angular velocity and yaw angle, is used. Let be the expectation operator, representing the average of all observable samples accessed by the student controller through the onboard sensors.

[0041] During the training of the teacher controller, a randomization method was employed on the simulation platform to construct multi-perturbation, multi-task scenarios, including random wind fields, random loads, friction coefficient variations, attitude changes, and airborne sensor noise, to enhance the teacher controller's generalization ability to complex environments. Through extensive simulation interactions, the teacher controller gradually learns the optimal strategy for maintaining the stability and high-precision pointing of the photoelectric stabilized platform under different perturbation modes and task objectives. The teacher controller takes privileged information from the system as input and outputs actuator control signals to achieve precise tracking of attitude angles and angular velocities. The teacher control strategy employs a probabilistic policy network with a Gaussian distribution, its mean function implemented by a multilayer perceptron, and its variance set as a learnable parameter. After training, the teacher controller acts as a "reference expert strategy," providing behavioral labels and supervision signals to the student controller. The student controller then performs imitation learning based on its limited available airborne sensor data, thereby achieving effective transfer from a fully information-rich environment to real airborne conditions.

[0042] Since the student controller is trained based on a teacher-student architecture in step S2, and its training samples are virtual samples generated by the simulation platform, this training process can be carried out in a multi-task environment (including different wind speeds, loads, friction coefficients, and noise levels). The student controller can learn a set of initial parameters that can quickly adapt to different disturbance tasks. However, the actual environment is not exactly the same as the virtual environment of the simulation platform. Therefore, in step S3, model-independent meta-reinforcement learning (MAML) is introduced to further enhance the student controller. In this training process, the student controller only uses the actual observation information from the airborne sensors (such as angular velocity, attitude angle, historical action sequences, etc.) as input and does not access privileged information. Through meta-reinforcement learning, optimization can be performed based on the initial parameters, allowing the student controller to be transferred from virtual environment applications to complex and uncertain environments. When facing new environmental or disturbance conditions, the student controller only needs a few iterations to complete parameter fine-tuning, thus possessing rapid adaptive capabilities. This design significantly improves the generalization and transfer performance of the control system in complex and uncertain environments. It avoids the problems of difficulty in obtaining real samples and high acquisition costs associated with traditional model training using real data.

[0043] Meta-reinforcement learning consists of two loops: an inner loop and an outer loop. The inner loop performs reinforcement learning for a specific task (under a specific load disturbance environment) to learn the optimal policy for that task. The inner loop is responsible for rapid in-task adaptation during student controller training. The training iteration formula is expressed as: ; in: For the i-th task, each task represents a perturbation configuration, such as different wind fields, different loads, different friction parameters, etc. The initial parameters for the goal of meta-reinforcement learning training are such that optimal results are achieved with the fewest training iterations across multiple tasks. These are the parameters of the student controller after the last gradient update. The learning rate is used to control the step size of policy updates in each task. To calculate policy parameters in the task Directional improvements; The outer loop observes the training process of multiple similar tasks and, during the learning process across different tasks, identifies the optimal model initialization parameters. These parameters are tailored to a specific set of tasks and are quickly optimized through a small amount of training. The iterative formula for the outer loop is expressed as: ; in: For the outer loop learning rate, The accumulated gradients across all tasks are used to update the meta-parameters. , To calculate policy parameters in the task Directional improvements.

[0044] In step S4, after the student controller's reinforcement training is completed, the student controller is exported as an embeddable model and deployed in the airborne sensor hardware of the airborne optoelectronic stabilization platform. The deployment platform can be a high-performance embedded computing unit such as a DSP or FPGA. The student controller makes real-time decisions by acquiring angular velocity, attitude angle, and actuator feedback signals from observable information from the airborne sensors. To ensure real-time performance and stability, this invention introduces an online strategy fine-tuning mechanism during the deployment phase. This mechanism adjusts the student controller's basic parameters based on real-time observation information from the airborne sensors. When external disturbances exceed the training range, strategy correction can be achieved through minor online updates. This enables the UAV controller to maintain high-precision and stable control under various flight attitudes, wind disturbances, and load variations, exhibiting excellent robustness, adaptability, and energy efficiency.

[0045] In summary, the above are merely preferred embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

[0046] The systems, apparatuses, modules, or units described in one or more of the above embodiments may be implemented by a computer chip or entity, or by a product having a certain function. A typical implementation device is a computer. Specifically, a computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0047] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0048] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

Claims

1. An intelligent control method for an airborne opto-electronic stabilized platform, characterized in that, include: S1: A simulation platform for building an airborne optoelectronic stabilization platform that integrates electromechanical environment and mission. The simulation platform outputs two types of information, including privileged information and observable information from airborne sensors. S2: Construct a teacher-student controller based on a privileged learning architecture. The teacher-student controller includes a teacher controller and a student controller. The input of the teacher controller is the privileged information, and the output of the teacher controller is the teacher control policy. The teacher controller is trained based on the privileged information, and the student controller learns the teacher control policy by imitation based on observable information from airborne sensors. S3: The student controller is trained through meta-reinforcement learning. During the meta-reinforcement learning process, the input of the student controller is the actual observation information of the airborne sensor, and the output of the student controller is the actual control strategy of the airborne optoelectronic stabilization platform. S4: The trained student controller is deployed on the airborne optoelectronic stabilization platform. The student controller generates an actual control strategy based on the actual observation information collected in real time from the airborne sensors.

2. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, The privileged information of the simulation platform includes: the internal state of the electromechanical system, environmental disturbances, and task parameters. Represented as: ; in: This refers to the internal state of the electromechanical system. For environmental disturbance, These are task parameters.

3. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, The teacher controller is trained using a proximal policy optimization algorithm, and the objective function of the pruning policy in the proximal policy optimization algorithm is... Represented as: ; in: ; Prior teacher control strategies, For current teacher control strategies, These are the trainable parameters of the teacher controller. It is the ratio of the probability of a certain action to the probability of the prior teacher control strategy and the current teacher control strategy. For the dominant function, This is a truncation term used to control the training-to-update ratio within a certain range. , This is a hyperparameter for the cropping range. This refers to the expected trajectory of interaction between time and environment.

4. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, The composite reward function trained by the teacher controller for: ; in: To control the weighting coefficients of reward items, As a reward for stable posture, To reward tracking accuracy, For energy consumption constraints, For perturbation robustness rewards; Wherein: the attitude stability reward Represented as: in: This is the attitude angle error vector. This is the platform's angular velocity vector; The tracking accuracy bonus Represented as: in: To track errors, Represented as: ; in: The output of the simulation platform, For task commands; The energy consumption constraint Represented as: in: This refers to the motor drive torque or control input vector of the airborne optoelectronic stabilization platform. The perturbation robustness reward Represented as: in: This represents the change in environmental disturbance.

5. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, The student controller mimics the loss function of the teacher's control strategy. Represented as: ; in: The parameters for the student controller to mimic and learn the teacher's control strategy. These are the trainable parameters of the teacher control strategy. The policy function is the one by which the student controller learns and mimics the teacher's control strategy. Let be the policy function of the teacher control strategy. For complete privileged information observed by the teacher controller, The observable information acquired by the student controller through onboard sensors. For expectation operators.

6. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, The meta-reinforcement learning includes an inner loop and an outer loop, and the inner loop iteration formula is expressed as follows: ; in: For the i-th task, each task represents a perturbation configuration. These are the target initial parameters for the meta-reinforcement learning training. The parameters of the student controller after the last gradient update. The learning rate is used to control the step size of policy updates in each task. To calculate policy parameters in the task Directional improvements; The outer loop iteration formula is expressed as: ; in: The learning rate for the outer loop. The accumulated gradients across all tasks are used to update the meta-parameters. , To calculate policy parameters in the task Directional improvements.

7. The intelligent control method for the airborne optoelectronic stabilization platform according to claim 1, characterized in that, An online fine-tuning mechanism is introduced into the airborne optoelectronic platform, which adjusts the parameters of the student controller based on the actual observation information collected in real time by the airborne sensors.