Control surface oscillation fault feature extraction method based on physical model and reinforcement learning fusion
By combining physical models and reinforcement learning algorithms in the oscillation fault monitoring of aircraft rudder surface oscillation fault signal, high-precision mapping between flight control available signals and rudder surface oscillation fault signal characteristics is achieved, solving the problem of difficulty in extracting fault features in the prior art, and improving the accuracy and real-timeness of fault monitoring.
Patent Information
- Application Number
- CN202510269494.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-10
AI Technical Summary
When handling aircraft rudder oscillation faults, it is difficult to effectively use flight control available signals to extract fault characteristics, and there are local optimization problems.
A method based on the integration of physical model and reinforcement learning is adopted to establish a high-precision nonlinear servo actuator physical model, and iterative training is used to achieve high-precision mapping between flight control available signals and rudder surface oscillation fault signal characteristics.
It realizes the frequency and amplitude of the rudder surface oscillation fault signal in real time when the fault model is unknown, avoiding local optimization problems and improving the accuracy of fault feature extraction.
Smart Images

Figure CN120123741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of aircraft manufacturing, specifically a method for extracting the fault characteristics of a control surface oscillation based on the fusion of a physical model and reinforcement learning. Background Art
[0002] Control surface oscillation refers to the periodic oscillation of a certain control surface of an aircraft, resulting in structural fatigue damage to the control surface. The main reasons come from the signal noise and clutter in the high-gain fly-by-wire system, which are coupled to the actuator controller, causing the control surface to vibrate repeatedly, and then triggering the overall oscillation phenomenon of the flight control system, reducing the flight control quality and flight safety performance. Since the data collected under actual flight conditions with noise and interference will have a situation where the local data error becomes larger, the existing fault prediction technologies have local optimization problems, and the existing control surface oscillation fault monitoring methods cannot extract the fault characteristics of the control surface oscillation using the available signals of the flight control. Summary of the Invention
[0003] In view of the above deficiencies of the existing technology, the present invention proposes a method for extracting the fault characteristics of a control surface oscillation based on the fusion of a physical model and reinforcement learning. First, a high-precision nonlinear servo actuator physical model is established based on the working principle and parameter information of the electro-hydraulic servo actuator. Then, the available signals of the flight control output by the servo actuator and physical models such as aircraft dynamics are used as the state observation inputs of the control surface oscillation fault characteristic extraction system, and a reward function is set. Then, an iterative training is carried out using the training method based on DDPG to obtain a mapping relationship between the available signals of the flight control and the fault signal characteristics of the control surface oscillation. Finally, the amplitude and frequency of the control surface oscillation fault are extracted in real time through the physical model fusion reinforcement learning algorithm, realizing a high-precision mapping between the available signals of the flight control and the fault signals in the case of an unknown fault model, so as to accurately extract the fault signal characteristics.
[0004] The present invention is realized through the following technical solutions:
[0005] The present invention relates to a method for extracting the fault characteristics of a control surface oscillation based on the fusion of a physical model and reinforcement learning, including:
[0006] Step 1: Generate a reference control surface control command through an autopilot according to the error between the flight reference command and the actual flight parameter signal.
[0007] The autopilot calculates the final reference control surface control command δ according to the error e(t) between the flight reference command and the actual flight parameter signal through a PID control law d .
[0008] Step 2: According to the servo actuator model, the aircraft dynamics model, and the reference control surface control command δd Thus, a rudder surface servo control system is established, and then the available flight control parameter signals of the aircraft during flight are generated.
[0009] The described rudder surface servo control system includes an autopilot unit, a servo actuator model unit, and an aircraft dynamics model unit, where: The autopilot unit calculates the final reference rudder surface control command δ according to the error e(t) between the flight reference command and the actual flight parameter signal d ; the servo actuator model unit obtains the actual aircraft servo actuator deflection signal δ according to the reference rudder surface control command δ d and the oscillation fault injection; and the aircraft dynamics model unit generates the available flight control parameter signals of the aircraft during flight according to the actual aircraft servo actuator deflection signal δ d true true
[0010] The described servo actuator model unit adopts a non-linear electro-hydraulic servo actuator system model, and the aircraft dynamics model unit adopts an aircraft longitudinal flight control system model.
[0011] Step 3: Take the error between the expected frequency and amplitude of the rudder surface oscillation fault and the frequency and amplitude output by the rudder surface oscillation fault feature extraction system as the input of the reward function. After taking the available flight control signal generated by the rudder surface servo control system as the input of the state observation of the rudder surface oscillation fault feature extraction system, train according to the reinforcement learning algorithm.
[0012] The described rudder surface oscillation fault feature extraction system includes: a state observation unit, a reward function unit, and a reinforcement learning training unit, where: The state observation unit generates an available flight control error signal by taking the difference between the available flight control signal generated by the rudder surface servo control system and the flight reference command signal, and performs state observation by using the available flight control error signal; the reward function unit sets a reward function according to the error between the expected frequency and amplitude of the rudder surface oscillation fault and the frequency and amplitude output by the rudder surface oscillation fault feature extraction system; the reinforcement learning training unit performs iterative training by taking the output results of the state observation and the output of the reward function as the input of the algorithm according to the training method based on Deep Deterministic Policy Gradient (DDPG), and finally extracts the frequency and amplitude of the rudder surface oscillation fault.
[0013] The described reward function is where: the expected frequency f true and amplitude A true of the rudder surface oscillation fault and the frequency f α and amplitude A α output by the rudder surface oscillation fault feature extraction system, and the error signal (f α - f true ) and (Aα -A true ), -w 1 and -w 2 Both are reward function gains. The difference is taken separately between the frequencies and amplitudes of the outputs of the previous moment and the current moment of the control surface oscillation fault feature extraction system to generate an error signal Δf α and ΔA α .
[0014] The described training method based on DDPG is characterized in that: there are an action network (Actor) unit and an evaluation network (Critic), where: the Critic network takes the action generated by the Actor network as input and then outputs a Q value; the update of the Actor network is to optimize the policy by maximizing the Q value evaluated by the Critic, and the network parameters are updated by gradient ascent.
[0015] Step 4: Extract the amplitude and frequency of the control surface oscillation fault signal corresponding to the current flight control available signal through the trained control surface oscillation fault feature extraction system, and verify the performance of the trained control surface oscillation fault feature extraction system through the error constraint range.
[0016] The present invention relates to a control surface oscillation fault feature extraction system for implementing the above method, including: a state observation module, a reward function module, and a reinforcement learning training module, where: the state observation module performs state observation according to the flight control available signal generated by the control surface servo control system; the reward function module sets the reward function according to the error signals (f α -f true ) and (A α -A true ) and the reward function gains -w 1 and -w 2 as well as the difference signals Δf α and ΔA α , and the reinforcement learning training module processes the output results of the state observation and the output results of the reward function according to the training method based on DDPG, and performs iterative training to finally extract the features of the control surface oscillation fault. Technical effects
[0017] In view of the local optimization problem existing in model-free reinforcement learning and the problem that it is difficult to extract the characteristics of aileron oscillation faults from the available signals of the flight control in the prior art, the present invention proposes a method for extracting the characteristics of aileron oscillation faults based on the fusion of physical models and reinforcement learning. First, a high-precision nonlinear servo actuator physical model is established based on the working principle and parameter information of the electro-hydraulic servo actuator. Then, the available signals of the flight control output by physical models such as the servo actuator and aircraft dynamics are used as the state observation inputs of the aileron oscillation fault feature extraction system, and a reward function is set. Then, an iterative training is carried out using the training method based on DDPG to obtain a mapping relationship between the available signals of the flight control and the characteristics of the aileron oscillation fault signals. Finally, the characteristics of the aileron oscillation fault signals are extracted using the available signals of the flight control. Compared with the prior art, the present invention can extract the frequency and amplitude of the aileron oscillation fault signals in real time. The performance of the improved method for extracting the characteristics of aileron oscillation faults is verified by the data of the validation set, and the frequency and amplitude of the extracted aileron oscillation fault signals are within the error constraint range. Description of the Drawings
[0018] Figure 1 Schematic diagram of a fly-by-wire flight control system and possible oscillation fault schematic diagrams;
[0019] Figure 2 Structural block diagram of the extraction of aileron oscillation fault characteristics based on the fusion of physical models and reinforcement learning according to the present invention;
[0020] Figure 3 Structural diagram of the DDPG network according to the present invention;
[0021] Figure 4 Available signal diagram of the flight control when an aileron oscillation fault occurs according to the present invention;
[0022] In the figure: (a) flight path angle curve, (b) load curve graph, (c) other available signals of the flight control;
[0023] Figure 5 Convergence process diagram of the reward function during the DDPG training when an aileron oscillation fault occurs according to an embodiment of the present invention; Frequency and amplitude training process diagrams in a single episode;
[0024] In the figure: (a) Convergence process of the reward function during the DDPG training, (b) Frequency training process in a single episode, (c) Amplitude training process in a single episode;
[0025] Figure 6 Frequency and amplitude diagrams extracted during the training of the designed aileron oscillation fault feature extraction;
[0026] In the figure: (a) Frequency extracted during the training of the aileron oscillation fault feature extraction, (b) Amplitude extracted during the training of the aileron oscillation fault feature extraction;
[0027] Figure 7 Frequency and amplitude diagrams extracted during the verification process for the designed rudder surface oscillation fault feature extraction;
[0028] In the figure: (a) The frequency extracted during the verification process for the rudder surface oscillation fault feature extraction, (b) The amplitude extracted during the verification process for the rudder surface oscillation fault feature extraction. Specific implementation manner
[0029] As Figure 1 shown, it is the structure of the fly-by-wire flight control system and the rudder surface oscillation fault in the action scenario of this embodiment, including: a liquid form fault where the oscillation signal is superimposed on the original signal and a solid fault form where the oscillation signal replaces the original signal. This oscillation signal is transmitted to the rudder surface through the servo control loop, resulting in rudder surface oscillation.
[0030] As Figure 2 shown, it is a method for extracting rudder surface oscillation fault features based on the fusion of physical models and reinforcement learning involved in this embodiment, including:
[0031] Step 1. According to the error between the flight reference command and the actual flight parameter signal, generate a reference rudder surface control command through the autopilot, specifically including:
[0032] 1.1 Calculate the error e(t) between the flight reference command and the actual flight parameter signal;
[0033] 1.2 Generate a reference rudder surface control command δ d , specifically: Where: control law parameters, k 1 = 0.8, k 2 = 3, k 3 = -3.
[0034] Step 2. According to the servo actuator model, the aircraft dynamics model, and the reference rudder surface control command δ d , generate the fly-by-wire available parameter signals during the flight of the aircraft, specifically including:
[0035] 2.1 Establish a non-linear servo actuator dynamics model, specifically: Where: v(t) is the actuator rod displacement rate, ΔP refers to the hydraulic pressure difference, ΔP ref is the pressure difference reference value when the rod speed reaches the maximum value, S is the actuator piston surface area, the external force F aero = -sgn(δ)K aero v c and F damping = K dv 2 ,K d = 8.45 and K aero = 647.7 are the actuator damping coefficient and the aerodynamic load coefficient respectively,
[0036] In this embodiment, ΔP ref = 33.5 Mpa, S = 5800 cm2, and the transformation relationship between the rod displacement p(t) and the rudder surface deflection δ can be obtained by interpolation and look-up table.
[0037] The control law of the rod displacement includes: i c (t) = K p [p REF (t) - p(t)], v c (t) = K c i(t), K p = 0.6 is the gain of the rod displacement control law, and K c = [-250, -250, -32, 0, 32, 250, 250] is the conversion coefficient from the servo actuator input current to the actuator rod speed.
[0038] 2.2 Establishing the aircraft dynamics model specifically includes: Where: x is the aircraft state vector, u is the control input vector, y is the output vector, the state matrix The control input matrix The output matrix The feedforward matrix
[0039] 2.3 Generating the fly-by-wire available parameter signals during flight, including: flight path angle γ, overload N z , angle of attack α, pitch rate p, pitch angle θ, speed V.
[0040] Step 3: Taking the fly-by-wire available signals generated by the rudder surface servo control system as the input of the state observation of the rudder surface oscillation fault feature extraction system, and taking the error between the expected frequency and amplitude of the rudder surface oscillation fault and the frequency and amplitude output by the rudder surface oscillation fault feature extraction system as the input of the reward function, and then using the DDPG-based method for iterative training, specifically including:
[0041] 3.1 Generating the fly-by-wire available error signals by taking the difference between the fly-by-wire available signals generated by the rudder surface servo control system and the flight reference command signals.
[0042] 3.2 The reward function is set according to the error between the expected frequency and amplitude of the rudder surface oscillation fault and the frequency and amplitude output by the rudder surface oscillation fault feature extraction system.
[0043] The said reward function Among them: the reward function gain is specifically set to -w 1 =-1.0 and -w 2 =-0.1.
[0044] 3.3 Iterative training is performed based on the DDPG-based reinforcement learning method, specifically: the total number of episodes is 1000, the duration of a single episode is T = 100s, the learning parameter of the main network is 0.001, the discount factor is 0.9, the target network replacement parameter τ is 0.02, and the maximum storage of the memory space is 10000.
[0045] Step 4: Extract the amplitude and frequency of the control surface oscillation fault signal corresponding to the currently available flight control signals through the trained control surface oscillation fault feature extraction system, and verify the performance of the trained control surface oscillation fault feature extraction system through the error constraint range.
[0046] The error constraint range mentioned above refers to:
[0047] As Figure 3 shown, the DDPG network structure diagram includes: an action network (Actor) unit and a critic network. Among them: the Actor unit includes a state input layer, a hidden layer, and an action output layer. The hidden layer has three layers, and the ReLU activation function is used as the transfer relationship between the hidden layer networks. The Critic unit includes a state input layer, an action input layer, a hidden layer, and an evaluation output layer. The hidden layer also has three layers, and the ReLU activation function is also used as the transfer relationship between the hidden layer networks.
[0048] After specific actual experiments, during the flight of a civil aircraft executing a tracking flight reference instruction, after injecting a control surface oscillation fault, the characteristics of the control surface oscillation fault signal are extracted in real time through the above method, and the simulation results are as Figure 6 、 Figure 7 and Table 1 show.
[0049] When a control surface oscillation fault occurs, set the frequency of the injected control surface oscillation fault to 10 (Hz) and the amplitude to 2 (mm). As Figure 4 (a) shows, the straight line is the flight path angle instruction, and the curve with a waveform shape is the flight path angle during the actual flight with the injected control surface oscillation fault. As Figure 4 (b) shows, the curve represents the load curve with the injected control surface oscillation fault. As Figure 4 (c) shows, the curve represents the curves of other available flight control signals with the injected control surface oscillation fault.
[0050] When a control surface oscillation fault occurs, based on DDPG training, achieve reward function convergence and training of frequency and amplitude in a single episode: as Figure 5(a) As shown, the reward function finally converges to the maximum value. As Figure 5 (b) As shown, the frequency finally converges to around 10 Hz in a single round. As Figure 5 (c) As shown, the amplitude finally converges to around 2 mm in a single round.
[0051] As Figure 6 (a) As shown, the expected frequency is 10 Hz and the output frequency is 10.1 Hz during the training process. As Figure 6 (b) As shown, the expected amplitude is 2 mm and the output amplitude is 2.08 mm during the training process.
[0052] Verification is carried out according to the frequency and amplitude diagrams that can be extracted during the verification process of the present invention: As Figure 7 (a) As shown, the expected frequency is 8 Hz and the output frequency is 8.36 Hz during the verification process. As Figure 7 (b) As shown, the expected amplitude is 1.5 mm and the output amplitude is 1.47 mm during the verification process. As shown in Table 1, the performance indicators of the designed rudder surface oscillation fault feature extraction method are shown, and the constraint error meets the design requirements.
[0053] Table 1
[0054] During the training process, the frequency constraint error is The amplitude constraint error is During the verification process, the frequency constraint error is The amplitude constraint error is It can be seen that it has a good fault extraction effect.
[0055] Compared with the prior art, the present invention realizes the extraction of the characteristics of the rudder surface oscillation fault signal by establishing a high-precision physical model of the nonlinear servo actuator and fusing the servo actuator with physical models such as aircraft dynamics and the reinforcement learning intelligent algorithm. Compared with the existing methods, it realizes the mapping between the available signals of the flight control and the characteristics of the rudder surface oscillation fault signal, and also avoids the local optimum problem of the model-free reinforcement learning algorithm.
[0056] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.
Claims
1. A method for extracting rudder oscillation fault features based on the fusion of physical model and reinforcement learning, characterized in that: include: Step 1: Generate reference control commands through the autopilot according to the error between the flight reference command and the actual flight parameter signal; Step 2: According to the servo actuator model, the aircraft dynamics model and the reference control surface control instruction δ d Thus, a control surface servo control system is established, and then a flight control parameter signal is generated for the aircraft during flight; Step 3: Using the error between the expected frequency and amplitude of the rudder oscillation fault and the frequency and amplitude output by the rudder oscillation fault feature extraction system as the input of the reward function, using the flight control available signal generated by the rudder servo control system as the state observation input of the rudder oscillation fault feature extraction system, and then training is performed using a reinforcement learning algorithm; Step 4: The amplitude and frequency of the rudder oscillation fault signal corresponding to the current flight control available signal are extracted through the trained rudder oscillation fault feature extraction system, and the performance of the rudder oscillation fault feature extraction system after training is verified through the error constraint range.
2. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1 is characterized in that: The autopilot uses the PID control law to calculate the error e(t) between the flight reference command and the actual flight parameter signal to obtain the final reference control command δ d .
3. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1 is characterized in that: The control surface servo control system comprises an autopilot unit, a servo actuator model unit and an aircraft dynamics model unit, wherein: the autopilot unit is based on the delta between the flight reference command and the actual flight parameter signal d The error e(t) is solved to obtain the final reference control command δ d The servo actuator model unit responds to the reference control surface control command δ d and oscillation fault injection to obtain the actual aircraft servo actuator deflection signal δ true The aircraft dynamics model unit is based on the actual aircraft servo actuator deflection signal δ true Generates available flight control parameter signals during the flight of the aircraft.
4. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1, characterized in that: The servo actuator model unit adopts a nonlinear electro-hydraulic servo actuator system model, and the aircraft dynamics model unit adopts an aircraft longitudinal flight control system model.
5. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1 is characterized in that: The rudder oscillation fault feature extraction system comprises: a state observation unit, a reward function unit and a reinforcement learning training unit, wherein: the state observation unit generates a flight control available error signal by subtracting a flight control available signal generated by a rudder servo control system from a flight reference command signal, and uses the flight control available error signal to perform state observation; the reward function unit sets a reward function according to the error between the expected frequency and amplitude of the rudder oscillation fault and the frequency and amplitude output by the rudder oscillation fault feature extraction system; the reinforcement learning training unit uses the output result of the state observation and the output of the reward function as the input of the algorithm and performs iterative training according to a training method based on deep deterministic policy gradient (DDPG), and finally extracts the frequency and amplitude of the rudder oscillation fault.
6. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1, characterized in that: The reward function is Where: the expected frequency f of the rudder oscillation fault true and amplitude A true The frequency f output by the rudder oscillation fault feature extraction system α and amplitude A α The error signal f α -f true and A α -A true , -w1 and -w2 are both reward function gains, and the frequency and amplitude outputted by the rudder oscillation fault feature extraction system at the previous moment and the current moment are respectively subtracted to generate an error signal Δf α and ΔA α .
7. The method for extracting the characteristics of the rudder surface oscillation fault according to claim 1 is characterized in that: The DDPG-based training method is characterized by: an action network (Actor) unit and an evaluation network (Critic), wherein: the Critic network takes the action generated by the Actor network as input and then outputs a Q value; the Actor network is updated by optimizing the strategy by maximizing the Q value evaluated by the Critic, and updating the network parameters by gradient ascent.
8. A control surface oscillation fault feature extraction system for implementing the method described in any one of claims 1 to 7, characterized in that: include: State observation module, reward function module and reinforcement learning training module, wherein: the state observation module performs state observation according to the available flight control signal generated by the rudder servo control system; the reward function module performs state observation according to the error signal f α -f true and A α -A true and the reward function gains -w1 and -w2 and the difference signal Δf α and ΔA α , to set the reward function; the reinforcement learning training module processes the output results of the state observation and the output results of the reward function according to the DDPG-based training method, and performs iterative training to finally extract the characteristics of the rudder oscillation fault.