An attitude error compensation method for aerial imaging systems

The attitude error compensation method combining NAR-LSTM and DDPG algorithms solves the problem of attitude error affecting imaging accuracy in linear array airborne imaging systems, achieving real-time and high-precision attitude error compensation and improving the stability and adaptability of the imaging system.

CN121632208BActive Publication Date: 2026-04-28CHANGCHUN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGCHUN UNIV OF TECH
Filing Date
2026-02-04
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

During flight, the linear array airborne imaging system is affected by multiple sources of disturbance to the platform attitude, which causes pointing errors in the optical line of sight, affecting the imaging geometric accuracy and stability. Traditional control strategies are difficult to achieve real-time high-precision compensation.

Method used

Nonlinear attitude error prediction is performed using a nonlinear autoregressive long short-term memory network (NAR-LSTM), and the compensation parameters are optimized by combining deep deterministic policy gradient (DDPG) to generate attitude error feedforward compensation commands, thereby realizing multi-axis collaborative attitude error compensation.

Benefits of technology

It improves the imaging stability and geometric accuracy of the airborne imaging system, effectively offsets the dynamic response lag of the actuator, and enhances the system's adaptability to complex disturbance environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121632208B_ABST
    Figure CN121632208B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of aviation remote sensing and intelligent control technology, and particularly relates to a kind of aviation imaging system attitude error compensation method, comprising: S1: real-time acquisition of the attitude measurement data of aviation imaging system;S2: attitude error is decomposed into linear error component and nonlinear error component;S3: construct nonlinear error prediction model based on NAR-LSTM network;S4: construct compensation parameter optimization model based on DDPG strategy, attitude error and nonlinear error prediction value are used as state input, and Actor network is used to output feedforward compensation parameter;S5: according to the prediction advance, the compensation proportion coefficient and the nonlinear attitude error obtained by prediction, generate attitude error feedforward compensation instruction, and the attitude error feedforward compensation instruction is superimposed with linear error compensation result and then output to actuator, the advantage of the present application is: can effectively reduce aerial surveying and mapping optical axis deviation, significantly improves imaging stability and geometric accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of airborne remote sensing and intelligent control technology, and in particular relates to an attitude error compensation method for airborne imaging systems based on nonlinear autoregressive long short-term memory network (NAR-LSTM) and deep deterministic policy gradient (DDPG). Background Technology

[0002] Linear array airborne imaging systems, due to their high resolution and high frame rate, are widely used in surface mapping, target identification, and remote sensing monitoring tasks. However, during flight, the platform attitude is often affected by multiple sources of disturbance (such as airflow, structural coupling, actuator hysteresis, and sensor noise), leading to line-of-sight error (LOS error) and thus affecting the geometric accuracy and stability of imaging. Traditional PID or sliding mode control strategies exhibit sluggish response to strong nonlinear and uncertain disturbances, making it difficult to achieve real-time high-precision compensation. Therefore, it is necessary to introduce intelligent control algorithms with predictive and adaptive capabilities to achieve dynamic compensation for attitude errors. Summary of the Invention

[0003] In view of the above problems, the purpose of this invention is to provide an attitude error compensation method for an airborne imaging system, which introduces a combination of nonlinear error prediction and reinforcement learning parameter optimization to achieve predictive compensation and adaptive adjustment of attitude errors, thereby effectively improving the imaging stability and geometric accuracy of the airborne imaging system and overcoming the shortcomings of the prior art.

[0004] The present invention provides a method for attitude error compensation in an airborne imaging system, comprising the following steps:

[0005] S1: Real-time acquisition of attitude measurement data from the airborne imaging system to obtain the actual attitude information of the airborne imaging system, and calculation of the attitude error based on the desired attitude command;

[0006] S2: Decompose the attitude error into linear error components and nonlinear error components;

[0007] S3: Construct a nonlinear error prediction model based on NAR-LSTM network (nonlinear autoregressive long short-term memory network), using historical nonlinear attitude error and airborne imaging system control input as time series inputs to predict nonlinear attitude error at future moments;

[0008] S4: Construct a compensation parameter optimization model based on the DDPG (Deep Deterministic Policy Gradient) strategy. The predicted values ​​of attitude error and nonlinear error (obtained from NAR-LSTM) are used as state inputs. An Actor network is used to output feedforward compensation parameters, including the prediction lead. and compensation ratio coefficient ;

[0009] S5: Based on the predicted lead time Compensation ratio coefficient The system generates attitude error feedforward compensation commands based on the predicted nonlinear attitude error, and then outputs these commands, along with the linear error compensation results, to the actuator to achieve real-time compensation of attitude error in the airborne imaging system.

[0010] As a preferred embodiment of the present invention, the nonlinear error prediction model adopts a network structure including an input layer, at least one long short-term memory hidden layer, and an output layer, wherein the hidden layer selectively memorizes historical nonlinear attitude error information through forget gates, input gates, and output gates.

[0011] As a preferred embodiment of the present invention, the input of the NAR-LSTM network is normalized nonlinear attitude error time series data, which includes at least historical nonlinear attitude error values ​​and corresponding control inputs for multiple consecutive sampling periods.

[0012] As a preferred embodiment of the present invention, the compensation parameter optimization model includes an Actor network and a Critic network, wherein: the Actor network is used to output the predicted lead amount. and compensation ratio coefficient The Critic network is used to evaluate the action value function corresponding to the current combination of compensation parameters.

[0013] As a preferred embodiment of the present invention, the compensation parameter optimization model stores the next state samples of state, action, and reward through an experience replay mechanism, and updates the network parameters using a small-batch random sampling method.

[0014] As a preferred embodiment of the present invention, the DDPG strategy adopts a target network soft update strategy, which uses a preset soft update coefficient to synchronously update the online network parameters and the target network parameters.

[0015] As a preferred embodiment of the present invention, the compensation parameter optimization model constructs an objective function or reward function for evaluating the compensation effect. The objective function or reward function affects the magnitude of the attitude error and a preset error threshold, and is used to guide the compensation parameters to be optimized in the direction of reducing the attitude error.

[0016] As a preferred embodiment of the present invention, the attitude error feedforward compensation command calculates the error value of the next moment based on the predicted nonlinear attitude error of the future moment, in order to offset the dynamic response lag of the actuator in advance.

[0017] As a preferred embodiment of the present invention, the attitude error feedforward compensation command is generated in the pitch axis and roll axis directions of the airborne imaging system, respectively, to achieve multi-axis cooperative attitude error compensation.

[0018] As a preferred embodiment of the present invention, the linear error component is estimated and compensated by a linear model or a filtering method, and the nonlinear error component is predicted and compensated by the nonlinear error prediction model.

[0019] The beneficial effects of this invention are as follows:

[0020] 1. This invention introduces a NAR-LSTM network to predict nonlinear attitude errors, which can effectively characterize the nonlinear dynamic characteristics of the system that are difficult to describe by traditional models, and improve the accuracy of error prediction.

[0021] 2. This invention utilizes a deep deterministic strategy gradient algorithm to optimize the prediction lead and compensation ratio coefficient online, avoiding manual parameter tuning and improving the system's adaptability to complex disturbance environments.

[0022] 3. This invention introduces a predictive lead to feedforward compensation for future errors, effectively offsetting the dynamic response lag of the actuator and improving the real-time performance and stability of attitude compensation. Attached Figure Description

[0023] Other objects and results of the invention will become more apparent and readily understood with reference to the following description taken in conjunction with the accompanying drawings. In the drawings:

[0024] Figure 1 Flowchart for attitude error modeling and decomposition;

[0025] Figure 2 This is a schematic diagram of the NAR-LSTM network structure;

[0026] Figure 3 Flowchart for optimizing the DDPG strategy. Detailed Implementation

[0027] Example 1

[0028] See Figure 1-3 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] This embodiment provides a method for attitude error compensation in an airborne imaging system, including the following steps:

[0030] S1: Real-time acquisition of attitude measurement data from the airborne imaging system to obtain the actual attitude information of the airborne imaging system, and calculation of the attitude error based on the desired attitude command;

[0031] S2: Decompose the attitude error into linear error components and nonlinear error components;

[0032] S3: Construct a nonlinear error prediction model based on NAR-LSTM network (nonlinear autoregressive long short-term memory network), using historical nonlinear attitude error and airborne imaging system control input as time series inputs to predict nonlinear attitude error at future moments;

[0033] S4: Construct a compensation parameter optimization model based on the DDPG (Deep Deterministic Policy Gradient) strategy. The predicted values ​​of attitude error and nonlinear error (obtained from NAR-LSTM) are used as state inputs. An Actor network is used to output feedforward compensation parameters, including the prediction lead. and compensation ratio coefficient ;

[0034] S5: Based on the predicted lead time Compensation ratio coefficient The system generates attitude error feedforward compensation commands based on the predicted nonlinear attitude error, and then outputs these commands, along with the linear error compensation results, to the actuator to achieve real-time compensation of attitude error in the airborne imaging system.

[0035] Furthermore, the nonlinear error prediction model in this embodiment adopts a network structure that includes an input layer, at least one long short-term memory hidden layer, and an output layer. The hidden layer selectively memorizes historical nonlinear attitude error information through forget gates, input gates, and output gates.

[0036] Furthermore, in this embodiment, the input to the NAR-LSTM network is normalized nonlinear attitude error time series data, which includes at least historical nonlinear attitude error values ​​and corresponding control inputs for multiple consecutive sampling periods.

[0037] Furthermore, the compensation parameter optimization model in this embodiment includes an Actor network and a Critic network, wherein the Actor network is used to output the prediction lead. and compensation ratio coefficient The Critic network is used to evaluate the action value function corresponding to the current combination of compensation parameters.

[0038] Furthermore, the compensation parameter optimization model in this embodiment stores the next state samples of state, action, and reward through an experience replay mechanism, and updates the network parameters using a small-batch random sampling method.

[0039] Furthermore, the DDPG strategy in this embodiment adopts a target network soft update strategy, which uses a preset soft update coefficient to synchronously update the online network parameters and the target network parameters.

[0040] Furthermore, in the compensation parameter optimization model of this embodiment, an objective function or reward function is constructed to evaluate the compensation effect. The objective function or reward function affects the magnitude of the attitude error and the preset error threshold, and is used to guide the compensation parameters to be optimized in the direction of reducing the attitude error.

[0041] Furthermore, in this embodiment, the attitude error feedforward compensation command calculates the error value at the next moment based on the predicted nonlinear attitude error at the future moment, in order to offset the dynamic response lag of the actuator in advance.

[0042] Furthermore, in this embodiment, the attitude error feedforward compensation commands are generated in the pitch and roll axes of the airborne imaging system, respectively, to achieve multi-axis cooperative attitude error compensation.

[0043] Furthermore, in this embodiment, the linear error component is estimated and compensated using a linear model or filtering method, and the nonlinear error component is predicted and compensated using the nonlinear error prediction model.

[0044] Furthermore, the aerial imaging system in this embodiment includes a camera imaging module, a dual-axis stabilized gimbal, an inertial measurement unit (IMU), a flight control unit, and an intelligent control module. The intelligent control module integrates a NAR-LSTM network and a DDPG strategy to perform attitude error modeling and compensation.

[0045] Example 2

[0046] This embodiment provides an attitude error modeling method for attitude error compensation in an airborne imaging system, including:

[0047] Attitude error:

[0048] Airborne imaging system at each sampling time Generate desired attitude command The actual angular velocity feedback is obtained by measuring with an inertial measurement unit (IMU). and proportional gain Define linearity error:

[0049]

[0050] Nonlinear error:

[0051]

[0052] in: For attitude error, For linear error components, This is a nonlinear error component;

[0053] NAR-LSTM networks use historical state vectors As inputs to angular velocity, angular acceleration, and jerk, the prediction is... The output is the total error prediction value. It should be noted that the above input features are only an exemplary selection method. In practical applications, other equivalent state variables or their combinations can be selected as network inputs based on the system sensor configuration and error characteristics. The final predicted total error... for:

[0054]

[0055] in: Predicting linear error components, Predicting nonlinear error components;

[0056] NAR-LSTM networks:

[0057] Open-loop structure: Historical real nonlinear pose errors are used as input during the training phase; the corresponding mathematical expression is as follows:

[0058]

[0059] in: This is the historical nonlinear attitude error sequence of the open-loop structure;

[0060] Closed-loop structure: In the prediction phase, the predicted value from the previous time step is used to replace the actual error for iterative prediction; the corresponding mathematical expression is as follows:

[0061]

[0062] in, For deep learning network functions, This is the historical nonlinear attitude error sequence of the closed-loop structure. It is the first The input feature vector at time step 1 includes:

[0063]

[0064] For aerial imaging systems, every moment The dual-axis stabilized gimbal generates reference attitude angles based on the output of the flight control unit. , representing the desired attitude or optical axis projection direction in both the pitch and roll axes; the resulting nonlinear error prediction model predicts the actual attitude. for:

[0065]

[0066] By integrating the predicted position points at all times, the projected trajectory of the optical axis of the dual-axis stabilized gimbal can be constructed over the entire flight path:

[0067]

[0068] in, To predict the actual attitude of the roll axis, Predict the actual attitude for the pitch axis;

[0069] To quantify the deviation between the actual optical axis and the desired path, the geometric minimum distance method is used to define the line-of-sight attitude error (Pointing Error) of the dual-axis stabilized gimbal at each time step. , will predict the current position Projected onto the reference trajectory segment composed of ideal path points:

[0070]

[0071] in: Indicates the time before the current moment Control input within each sampling period;

[0072] Error is defined as the prediction point Shortest distance to the nearest point in this segment:

[0073]

[0074] This shortest distance can be considered as the first... The instantaneous offset in the optical axis plane of the aerial image at a given moment reflects the magnitude of the current line-of-sight attitude error;

[0075] Construct a compensation parameter optimization model based on DDPG (Deep Deterministic Policy Gradient) strategy:

[0076] To effectively offset the actuator response lag, the compensation parameter optimization model retains two key parameters: feedforward compensation delay. Compensation ratio coefficient ; Desired attitude command at a given time Its compensation instructions It can be represented in the following form:

[0077]

[0078] The following feedforward compensation calculation method is a preferred implementation of this embodiment. Its compensation relationship can be equivalently transformed or extended according to the actuator type, coordinate system definition and control requirements. Indicates the current time Applications for the Future The line-of-sight error predicted at each moment is corrected to offset the hysteresis effect; further, the compensation amounts in the roll axis and pitch axis directions are expressed as follows:

[0079]

[0080]

[0081] in: The angle between the current attitude direction and the normal to the pitch axis; This is the compensation amount in the roll axis direction; This is the compensation amount in the pitch axis direction;

[0082] Based on the above compensation amounts, the compensation result for the reference attitude position is as follows:

[0083]

[0084]

[0085] in: For the roll axis compensation result, This is the result of pitch axis compensation;

[0086] The compensated command will be transmitted to the intelligent control module, making the attitude adjustment trajectory generated by the airborne imaging system closer to the target's line of sight, thereby significantly reducing imaging errors.

[0087] The following reward function is an example used to illustrate the construction of reinforcement learning policies. Its function form can be adjusted according to the performance indicators and convergence characteristics of the aerial imaging system without affecting the overall technical concept of this embodiment. Considering the high requirements of aerial imaging for line-of-sight stability, this embodiment designs the following reward function:

[0088]

[0089] in: For the first The time-prediction error index is defined as a weighted combination of mean square error and maximum absolute error. For the error threshold, This is the scaling factor;

[0090] When the prediction error is lower than the system tolerance threshold, the agent receives a positive reward, thereby guiding the policy to converge toward a high-precision compensation scheme.

[0091] Under the DDPG strategy, future... Constant line-of-sight angle deviation:

[0092]

[0093] in, For the predicted direction of the line of sight, The ideal direction vector;

[0094] Compensation parameters output by the policy network (Actor network and Critic network) of the DDPG policy and Calculate the correction compensation amount :

[0095]

[0096] in: To predict the direction in the first The components of the axis; These are the roll and pitch directions, respectively;

[0097] Finally, the corrected attitude commands are generated:

[0098]

[0099] This instruction is executed by the servo controller to achieve dual-axis collaborative compensation control of the dual-axis stabilized gimbal. It can be widely used in high-precision remote sensing and dynamic stabilization imaging tasks, effectively improving the error suppression capability and imaging geometric consistency of the airborne imaging system.

[0100] In each control cycle, the system state vector As input to the agent, it comprehensively describes the current flight and compensation state, including:

[0101]

[0102] in: For the roll axis at time Shaft angle error, pitch axis at time Shaft angle error, For prediction error, The rate of change of attitude command. This is an optional IMU jitter metric.

[0103] The DDPG strategy introduces a historical reward decay update mechanism: if consecutive Step by step, use the same action And currently, high rewards are available. The update can be traced back to the previous state using the following recursive method. Step experience:

[0104]

[0105] in: This is the time decay factor;

[0106] The following description of the policy network and value network update methods illustrates the basic implementation process of the reinforcement learning algorithm in this embodiment. The specific network structure and update strategy can be adjusted according to computing resources and real-time requirements. The DDPG policy is trained using a dual-network structure (main Actor / Critic network + corresponding Target network), and its goal is to minimize the mean squared error of the Critic network in the following way. :

[0107]

[0108] in, For the present Value estimation, For the goal value:

[0109]

[0110] in, For the target Critic network, The system state at the next moment. The target Actor network outputs an action. Discount factor;

[0111] Meanwhile, the Actor network is optimized through the following policy gradient:

[0112]

[0113] in, For performance indicators Regarding Actor network parameters gradient, for The gradient of the function with respect to the action. The gradient of the policy with respect to the parameters;

[0114] To effectively guide policy convergence, the following nonlinear reward function is designed:

[0115]

[0116] in: This represents the angular error between the current prediction and the actual attitude. This is the error tolerance threshold. For reward scaling factor;

[0117] This nonlinear reward function is highly sensitive to small error segments, which can drive the DDPG strategy to converge toward high-precision control and significantly penalize failed actions.

[0118] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for attitude error compensation in an airborne imaging system, characterized in that, Includes the following steps: S1: Real-time acquisition of attitude measurement data from the airborne imaging system to obtain the actual attitude information of the airborne imaging system, and calculation of the attitude error based on the desired attitude command; S2: Decompose the attitude error into linear error components and nonlinear error components; S3: Construct a nonlinear error prediction model based on NAR-LSTM network, using historical nonlinear attitude error and airborne imaging system control input as time series inputs, to predict the nonlinear attitude error at future moments; S4: Construct a compensation parameter optimization model based on the DDPG strategy, using the predicted values ​​of attitude error and nonlinear error as state inputs, and utilizing an Actor network to output feedforward compensation parameters, including the prediction lead. and compensation ratio coefficient ; S5: Based on the predicted lead time Compensation ratio coefficient The system generates attitude error feedforward compensation commands based on the predicted nonlinear attitude error, and then outputs these commands, along with the linear error compensation results, to the actuator to achieve real-time compensation of attitude error in the airborne imaging system.

2. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The nonlinear error prediction model adopts a network structure that includes an input layer, at least one long short-term memory hidden layer, and an output layer. The hidden layer selectively memorizes historical nonlinear attitude error information through forget gates, input gates, and output gates.

3. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The input to the NAR-LSTM network is normalized nonlinear attitude error time series data, which includes at least historical nonlinear attitude error values ​​and corresponding control inputs for multiple consecutive sampling periods.

4. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The compensation parameter optimization model includes an Actor network and a Critic network, wherein the Actor network is used to output the predicted lead. and compensation ratio coefficient The Critic network is used to evaluate the action value function corresponding to the current combination of compensation parameters.

5. The attitude error compensation method for an airborne imaging system according to claim 4, characterized in that, The compensation parameter optimization model stores the next state samples of state, action, and reward through an experience replay mechanism, and updates the network parameters using a small-batch random sampling method.

6. A method for attitude error compensation in an airborne imaging system according to claim 4 or 5, characterized in that, The DDPG strategy adopts a target network soft update strategy, which uses a preset soft update coefficient to synchronously update the online network parameters and the target network parameters.

7. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The compensation parameter optimization model constructs an objective function or reward function to evaluate the compensation effect. The objective function or reward function affects the magnitude of the attitude error and the preset error threshold, and is used to guide the compensation parameters to optimize in the direction of reducing the attitude error.

8. The method for attitude error compensation of an airborne imaging system according to claim 1, characterized in that, The attitude error feedforward compensation command calculates the error value for the next moment based on the predicted nonlinear attitude error at the future moment, in order to offset the dynamic response lag of the actuator in advance.

9. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The attitude error feedforward compensation commands are generated in the pitch and roll axes of the airborne imaging system, respectively, to achieve multi-axis collaborative attitude error compensation.

10. The attitude error compensation method for an airborne imaging system according to claim 1, characterized in that, The linear error component is estimated and compensated using a linear model or filtering method, and the nonlinear error component is predicted and compensated using the nonlinear error prediction model.

Citation Information

Patent Citations

  • Self-adaptive track control method of unmanned surface vessel

    CN117111594A

  • Self-adaptive optimization regulation and control method for performance of aero-engine combustion chamber

    CN118395884A