Self-repair learning control method and system for semi-healthy device with adaptive torque observer

By combining an adaptive torque observer with a neural network, self-repair control of the flow control valve is achieved, solving the problem in the existing technology that fault diagnosis relies on precise models and manual experience, and improving the stability and adaptability of the system.

CN119105291BActive Publication Date: 2025-09-16SHANGHAI JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411482992.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-09-16
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

In the existing technology, the fault diagnosis and self-repair control methods of flow control valves rely on precise mathematical models or manual experience and lack adaptive capabilities, resulting in inaccurate fault detection and repair, which may cause safety accidents.

Method used

An adaptive torque observer is combined with a neural network to establish a multi-degree-of-freedom system dynamics model. Fault diagnosis and self-repair control are achieved through fuzzy parameter estimation and reinforcement learning. The system response data is integrated for fault mode diagnosis and status monitoring, and a self-repair control strategy is established.

Benefits of technology

It improves the accuracy of fault diagnosis and the stability of the system, realizes self-repair control without the need for precise dynamic models and neural network fitting, and enhances the reliability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119105291B_ABST
    Figure CN119105291B_ABST
Patent Text Reader

Abstract

The present invention provides a self-repair learning control method and system for a semi-healthy device with an adaptive torque observer, comprising: establishing a system dynamics model for multi-degree-of-freedom control; establishing an adaptive torque observer based on generalized momentum based on the system dynamics model; collecting system response data and obtaining fuzzy parameters based on the adaptive torque observer to achieve adaptive estimation of external torque; fusing and classifying the system response data and the external torque estimate to perform fault mode diagnosis and state monitoring; and establishing a self-repair control strategy based on reinforcement learning based on the state monitoring results to achieve self-repair control of the multi-degree-of-freedom semi-healthy device. This invention has the advantages of being independent of a precise dynamics model for semi-healthy systems, highly adaptable external torque estimation, strong risk perception capabilities, and higher control system reliability and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-degree-of-freedom motion control, and in particular to a self-repairing learning control method and system for a semi-healthy device of an adaptive torque observer. Background Art

[0002] In modern industrial production, precise regulation and control of fluid flow has become an indispensable component in key areas such as robotics, oil and gas, chemicals, and aerospace. As a crucial component in pipeline systems, regulating valves play a vital role in regulating and controlling flow. Typically, a multi-degree-of-freedom robotic motion control device consists of multiple flow control valves. This device uses a uniform structure of regulating valves to control the valve opening, thereby precisely adjusting the fluid flow through the actuator and driving the movement of the entire structure.

[0003] In this system, each flow control valve shares a common input fluid chamber. The magnitude of its output force is closely related to the valve opening and the fluid pressure within the chamber. The chamber fluid pressure is, in turn, affected by the combined opening of all flow control valves. Notably, valve port blockage is a common fault in control valves. However, there is currently a lack of direct detection devices for valve blockage, posing significant challenges to control valve fault diagnosis and condition monitoring.

[0004] More seriously, if effective self-repair control measures are not implemented promptly when a valve malfunctions, it is very likely to lead to serious deviations in the control output and even serious safety accidents. Therefore, the research and application of fault diagnosis, condition monitoring, and self-repair control methods for multi-valve control systems are of great significance for improving the stability and reliability of the entire system and extending the service life of the equipment.

[0005] In the existing technology, Zhang Yudong conducted research on key technologies for fault diagnosis of flow control valves under multi-opening interference, and proposed using machine learning methods such as clustering to realize fault diagnosis of flow control valves. However, this method requires a lot of manual experience in multi-model modeling and data processing, and lacks adaptive capabilities.

[0006] Zhang Yingguang and Yan Xu conducted research on fault diagnosis of regulating valves in the temperature control water system of polyethylene plants. They diagnosed the faults of regulating valves based on actual cases of polyethylene plants and proposed corresponding solutions. However, these solutions relied on manual experience and lacked universality.

[0007] Chinese invention patent publication number CN117864087A provides a method for fault diagnosis and fault-tolerant control of a current sensor in an ESC-based electronically controlled power-assisted braking system. This solution relies on manual experience and lacks adaptability by determining faults using a threshold value.

[0008] A Chinese invention patent, publication number CN117806276A, provides a method for sensor fault diagnosis and fault-tolerant control of a variant aircraft. This solution achieves fault-tolerant control of the aircraft by building a sensor fault diagnosis model based on an expert system and reconstructing the observation matrix based on the fault state. However, the fault diagnosis method relies on critical value settings, which does not provide significant advantages.

[0009] A Chinese invention patent with publication number CN114115195A provides a method for fault diagnosis and fault-tolerant control of underwater robot thrusters. This scheme locates the fault location by observing the residual component and directly uses the Elman network to identify the nonlinear model. It relies on the fitting accuracy of the neural network and has weak interpretability. Summary of the Invention

[0010] In view of the defects in the prior art, the purpose of the present invention is to provide a self-repair learning control method and system for a semi-healthy device of an adaptive torque observer.

[0011] According to one aspect of the present invention, a method for self-repairing learning control of a semi-healthy device using an adaptive torque observer is provided, comprising:

[0012] Establish a system dynamics model for multi-degree-of-freedom control;

[0013] Based on the system dynamics model, an adaptive torque observer based on generalized momentum is established;

[0014] Collecting system response data and obtaining fuzzy parameters based on the adaptive torque observer to achieve adaptive estimation of external torque;

[0015] fusing and classifying the system response data and the external torque estimation, performing fault mode diagnosis, and obtaining a condition monitoring result based on the fault mode diagnosis;

[0016] Based on the state monitoring results, a self-repair control strategy based on reinforcement learning is established to realize self-repair control of multi-degree-of-freedom semi-healthy devices.

[0017] Preferably, the system dynamics model of the multi-degree-of-freedom control includes dynamic models of multiple regulating valves, wherein the dynamic model of a single regulating valve is:

[0018]

[0019] where τ m Represents the motor input torque, τ f Represents the friction torque of the motor rotor, F f1represents the friction between the cam pin and the cam groove, l represents the distance between the axis of the cam pin and the cam axis, θ represents the output angle of the motor reducer, R represents the radius of the cam pin, F x1 Represents the normal contact force between the cam pin and the cam groove, F f2 Represents the friction force of the gas valve body on the valve stem, F L represents the external load, J1 represents the sum of the motor rotor inertia and the reducer inertia, J2 represents the rotational inertia of the cam, and m represents the mass of the valve core.

[0020] Preferably, the generalized momentum-based adaptive torque observer includes the system dynamics model and a neural network, wherein the neural network adaptively estimates the fuzzy parameters of the system dynamics model based on system response data collected in a no-collision and no-fault mode.

[0021] Preferably, the acquisition system response data and the fuzzy parameters obtained based on the adaptive torque observer to achieve adaptive estimation of the external torque include:

[0022]

[0023] in, represents the estimate of the external torque, K is a positive real value, τ m is the actuator input torque, τ e is the external torque, p is the system generalized momentum, and α is a physical quantity containing fuzzy parameters;

[0024]

[0025] τ e =-F L lcosθ=-(F g -F c )lcosθ

[0026]

[0027] x=lsinθ

[0028] Where M is the generalized mass, M=J1+J2+ml 2 cos 2 θ, F g For normal load, F c is the interference load caused by factors such as collision, x is the displacement of the valve core, β i That is the fuzzy dynamic parameter;

[0029] The expression in α is:

[0030]

[0031] in

[0032]

[0033] w i,n are weight values, which are calculated by their respective expressions, where the subscript n is the serial number of the discrete sampling value, indicating that these are the four weight values ​​w and the constant value b calculated for the nth group of signals, and then the nth α is calculated.

[0034] Preferably, the neural network adopts a multi-layer perceptron MLP; the input of the neural network is the system response data, and the output is the fuzzy parameter β i , i=1~4 estimation, and finally through the formula Calculate α n ;

[0035] According to the system response data in the no-collision no-fault mode, the neural network in the adaptive external torque estimator is trained by the back-propagation formula, specifically:

[0036]

[0037] E represents the error (MSE), δw represents the calculation error of gradient descent, and η is the learning rate.

[0038] Preferably, the system response data includes: actuator current data, actuator displacement data, actuator speed data and environmental state data, and the environmental state data includes fluid pressure data, fluid density data and fluid temperature data;

[0039] The failure modes include: actuator blockage or actuator jam;

[0040] The fault mode determination includes: utilizing a classification network, inputting system response data and external torque estimation data, and outputting the fault mode;

[0041] The state monitoring includes: determining the degree of valve port blockage based on the determined failure mode.

[0042] Preferably, the classification network is trained by back propagation based on the response data of the faulty system, specifically:

[0043]

[0044] y c,n is the real classification label, is the output prediction of the network, M is the number of samples involved in the calculation, and τ is the learning rate.

[0045] Preferably, the state monitoring includes:

[0046] When the fault mode is diagnosed as actuator blockage, calculate the valve port blockage fault degree ρ;

[0047]

[0048] i c,n is the current value when the collision occurs, is the estimated current value without external load calculated by the adaptive dynamic model, and i0 is the threshold selected according to the motor performance parameters;

[0049] When the fault mode is diagnosed as actuator jam, the jam fault degree is not calculated.

[0050] Preferably, the self-repair control strategy based on reinforcement learning is established based on the state monitoring results to realize the self-repair control of the multi-degree-of-freedom system, including:

[0051] The goal of the self-repair control strategy is to track the given cavity pressure and output thrust vector;

[0052] The inputs of the self-repair control strategy are set as: a given cavity pressure target, a given output thrust vector target, the collected actual cavity pressure, the collected actual thrust vector, and the calculated fault degree of each actuator;

[0053] The output of the self-repair control strategy is set as: the servo displacement target of each actuator;

[0054] Based on the PPO algorithm, the self-repair control strategy training is carried out, and the training loss function is:

[0055]

[0056] π is the current control strategy, π old is the old policy corresponding to the data currently being trained, s is the system state, a is the action output by the policy based on the state s, A represents the advantage function, and ε is a value between (0, 1) that is used to limit the gap between the old and new policies to a set range;

[0057] The form of the training reward function is set as:

[0058] R=K1·|P target -P actor |+K2·|F a,target -F a,actor |+K3·|F l,target -F l,actor |

[0059] R is the reward value, K1, K2, K3 are weight values, P target is the target pressure of the cavity, P actoris the actual chamber pressure obtained by the control strategy, F a,target is the given output thrust angle, F a,actor is the actual output thrust angle, F l,target is the given output thrust, F l,actor is the actual output thrust.

[0060] According to a second aspect of the present invention, there is provided an adaptive torque observer semi-health device self-repair learning control device, comprising: a simulation loading module, a signal acquisition module, a fault diagnosis and status detection module, and a servo and self-repair control module; wherein,

[0061] The simulation loading module is used to simulate actuator failure;

[0062] The signal acquisition module is used to collect actual response data of the simulation loading module;

[0063] The fault diagnosis and state detection module uses an adaptive torque observer of generalized momentum based on the actual response data to obtain fuzzy parameters and achieve adaptive estimation of external torque; fuses and classifies the system response data and the external torque estimate to perform fault mode diagnosis, and obtains a state monitoring result based on the fault mode diagnosis;

[0064] The servo and self-repair control module establishes a self-repair control strategy based on reinforcement learning to achieve self-repair control of multi-degree-of-freedom semi-healthy devices.

[0065] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0066] (1) The self-repair learning control method and system for the semi-healthy device of the adaptive torque observer provided by the embodiment of the present invention, by collecting system response data as data drive, combines the dynamic model with the neural network. It does not rely on the precise dynamic model, nor does it rely solely on the neural network fitting modeling. It solves the problems of the existing flow control valve collision detection and fault diagnosis methods that rely on precise mathematical models, have poor fitting ability of simple neural networks, and are weak in interpretability.

[0067] (2) The self-repair learning control method and system of the adaptive torque observer semi-health device provided by the present invention integrates the neural network fitting technology into the dynamic model to achieve adaptive estimation of fuzzy parameters, solves the problem of difficulty in accurately calibrating nonlinear dynamic models, and improves the accuracy of dynamic model estimation and fault diagnosis;

[0068] (3) The self-repair learning control method and system of the adaptive torque observer semi-health device provided by the present invention designs a self-repair control strategy based on fault diagnosis results and reinforcement learning technology, and does not rely on manual control experience. The self-repair control strategy based on reinforcement learning improves the system reliability, stability and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0070] Figure 1 This is a flow chart of a self-repair learning control method for a semi-healthy device of an adaptive torque observer according to an embodiment of the present invention;

[0071] Figure 2 Schematic diagram of a self-repairing learning control system for a multi-degree-of-freedom semi-healthy device based on an adaptive torque observer in another embodiment of the present invention;

[0072] Figure 3 The comparison results between the adaptive torque observer and the initial dynamic model of a preferred embodiment of the present invention are shown;

[0073] Figure 4 Comparison results between a collision detection and fault diagnosis method based on an adaptive external torque observer according to a preferred embodiment of the present invention and the prior art;

[0074] Figure 5 This figure shows the comparison results between a self-repair control method based on reinforcement learning in a preferred embodiment of the present invention and a common controller without risk perception capability. DETAILED DESCRIPTION

[0075] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0076] In one embodiment of the present invention, a self-repair learning control method for a semi-healthy device of an adaptive torque observer is provided. Figure 1 As shown, the main steps are as follows:

[0077] S1, establish a system dynamics model for multi-degree-of-freedom control;

[0078] S2, based on the system dynamics model established in S1, establishes the adaptive torque observer DMAO based on generalized momentum;

[0079] S3, collects system response data and obtains fuzzy parameters based on the adaptive torque observer DMAO established in S2, thereby realizing adaptive estimation of external torque;

[0080] S4, fusing and classifying the system response data and the external torque estimation obtained in S3, performing fault mode diagnosis, and obtaining a condition monitoring result based on the fault mode diagnosis;

[0081] S5, based on the state monitoring results obtained in S4, establishes a self-repair control strategy based on reinforcement learning to realize the self-repair control of the multi-degree-of-freedom semi-healthy device.

[0082] The above embodiment has the advantages of being independent of an accurate dynamic model for a semi-healthy system, having high adaptability of external torque estimation, strong risk perception capability, and higher reliability and stability of the control system.

[0083] In a preferred embodiment of the present invention, S1 is implemented to establish a system dynamics model of multi-freedom control. The system dynamics model of multi-freedom control includes dynamic models of multiple control valves, and the dynamic model of a single control valve is:

[0084]

[0085] where τ m Represents the motor input torque, τ f Represents the friction torque of the motor rotor, F f1 represents the friction between the cam pin and the cam groove, l represents the distance between the axis of the cam pin and the cam axis, θ represents the output angle of the motor reducer, R represents the radius of the cam pin, F x1 Represents the normal contact force between the cam pin and the cam groove, F f2 Represents the friction force of the gas valve body on the valve stem, F L represents the external load, J1 represents the sum of the motor rotor inertia and the reducer inertia, J2 represents the rotational inertia of the cam, and m represents the mass of the valve core.

[0086] In a preferred embodiment of the present invention, S2 is implemented to establish a generalized momentum-based adaptive torque observer DMAO based on the system dynamics model established in S1.

[0087] The adaptive torque observer primarily consists of a system dynamics model and a neural network. The neural network is used to adaptively estimate the fuzzy parameters in the dynamics model. These parameters are difficult to precisely calibrate within the system dynamics model and primarily consist of various mechanical friction coefficients.

[0088] In a preferred embodiment, the formula of the adaptive torque observer is:

[0089]

[0090] in, represents the estimate of the external torque, K is a positive real value, τ m is the actuator input torque, τ e is the external torque, p is the system generalized momentum, and α is a physical quantity containing fuzzy parameters;

[0091]

[0092] τ e =-F L lcosθ=-(F g -F c )lcosθ

[0093]

[0094] x=lsinθ

[0095] Where M is the generalized mass, M=J1+J2+ml 2 cos 2 θ, F g For normal load, F c is the interference load caused by factors such as collision, x is the displacement of the valve core, β i That is the fuzzy dynamic parameter;

[0096] The expression of α can be sorted out as follows:

[0097]

[0098] in

[0099]

[0100] w i,n are weight values, which are calculated by their respective expressions, where the subscript n is the serial number of the discrete sampling value, indicating that these are the four weight values ​​w and the constant value b calculated for the nth group of signals, and then the nth α is calculated.

[0101] In the above embodiment, the dynamic model is combined with the neural network through data driving. It does not rely on the precise dynamic model, nor does it rely solely on the neural network fitting modeling. It solves the problems of the existing flow control valve collision detection and fault diagnosis methods relying on precise mathematical models, poor fitting ability of simple neural networks, and weak interpretability.

[0102] In a preferred embodiment, the neural network in the adaptive external torque estimator is trained based on the fault-free data, and the back propagation formula is:

[0103]

[0104] E represents the error (MSE), δw represents the calculation error of gradient descent, and η is the learning rate.

[0105] In a preferred embodiment of the present invention, S3 is implemented to collect system response data and obtain fuzzy parameters based on the adaptive torque observer DMAO established in S2, thereby realizing adaptive estimation of external torque. Specifically, it can be divided into two steps:

[0106] S31, based on the system response data actually collected in the no-collision and no-fault mode (In turn, they are actuator current data, actuator displacement data, actuator speed data, fluid pressure data, fluid temperature data, and fluid density data). Fuzzy parameter estimation is performed using a neural network. The formula is:

[0107]

[0108] S32, executes multi-degree-of-freedom system servo control and collects real-time system status signals

[0109] Based on the fuzzy parameters obtained in S31 and the adaptive torque observer established in S2, the adaptive estimation of the external torque is realized. The formula is:

[0110]

[0111] In the above embodiment, the neural network fitting technology is integrated into the dynamic model to realize the adaptive estimation of fuzzy parameters, which solves the problem that the nonlinear dynamic model is difficult to accurately calibrate and improves the estimation accuracy of the dynamic model.

[0112] Accuracy of measurement and fault diagnosis.

[0113] In a preferred embodiment of the present invention, S4 is implemented to fuse and classify the actual system response data collected by S3 and the external torque estimation obtained by S3 to perform fault mode diagnosis and status monitoring.

[0114] S41: Failures in gas regulating valve actuators are generally caused by sediment clogging the valve port (causing blockage / stuckness, leading to collisions). Therefore, to implement self-repair learning control of the device, it is necessary to detect the degree of sediment blockage. This method utilizes the actuator's collision characteristics to detect blockage.

[0115] Therefore, in this step, the collision detection classification network is constructed and trained, and its input is set as: the actual system response data and the external torque estimation obtained by S3, and the output is the fault mode and the degree of congestion.

[0116] During training, the input of the classification network is the response data with faults, the external torque estimation data, and the corresponding fault labels. The back-propagation formula is used for training:

[0117]

[0118] y c,n is the real classification label, is the output prediction of the network, M is the number of samples involved in the calculation, and τ is the learning rate.

[0119] S42, the actuator displacement signal x n , current signal i n With external torque estimation Pass it into the classification network to get the collision detection result cd n , the formula is:

[0120]

[0121] is the output of the classification network, which represents the classification prediction result. The output of the classification network has two dimensions, that is, c = 1 or 2; when y 1,n >y 2,n Time CD n is 1, otherwise it is 0. 1 represents congestion, that is, there is a collision.

[0122] When a cd is detected n When it is 1, the moment of collision is recorded and the actual system response data at that moment is extracted. c,n and i c,n is the current value when the collision occurs, The current estimate value without external load is calculated by the adaptive dynamic model, and i0 is the threshold value selected according to the motor performance parameters.

[0123] S43, performing status monitoring, that is, calculating the relevant fault degree.

[0124] Based on the collision detection results and the actual system current response, the valve port blockage fault degree ρ is calculated, which realizes the actuator fault diagnosis and early warning, laying the foundation for the subsequent self-repair control method based on reinforcement learning. In one embodiment, the formula is:

[0125]

[0126] i0 is a threshold selected based on the motor performance parameters.

[0127] In a robot multi-degree-of-freedom motion control device composed of multiple flow control valves, its control requirements are reflected in tracking the cavity pressure and the output thrust vector. According to the control requirements, the entire control device can move along the expected trajectory. On the basis of the state monitoring results calculated in the above embodiment, it is necessary to repair its entire control strategy so that it still moves along its expected trajectory. In a control device composed of multiple flow control valves, each actuator shares a pressure cavity, and its cavity pressure is related to the total opening of all actuators. The output thrust vector is the composite thrust of the output thrusts of each actuator, and the composite thrust determines the final motion trajectory. Therefore, in a preferred embodiment of the present invention, step 5 is implemented, and based on the state monitoring results obtained in S4, a self-repairing control strategy based on reinforcement learning is established to realize self-repairing control of the multi-degree-of-freedom system, so that in the presence of an unknown blockage fault, each actuator is controlled to cooperate with each other so that the cavity pressure and the total output thrust meet the given target requirements. Specifically, the following steps can be adopted:

[0128] S4.1, set the control target of the intelligent agent (self-repair control strategy) to track the given cavity pressure and output thrust vector, which is generally set manually.

[0129] S4.2, set the input of the self-repair control strategy as the target cavity pressure, the output thrust vector target, the actual cavity pressure, the actual thrust vector, the fault degree of each actuator (i.e., the ρ calculated in the above embodiment) n );

[0130] S4.3, set the servo displacement target of each actuator output by the self-repair control strategy.

[0131] S4.4, self-repair control strategy training is carried out based on the PPO algorithm. In one embodiment, the training loss function is:

[0132]

[0133] π is the current control strategy, π old is the old policy corresponding to the data used for the current training, s is the system state, a is the action output by the policy according to the state s, A represents the advantage function, and ε is a value between (0,1);

[0134] In order to drive the agent to accurately track the chamber pressure and output thrust vector, in a preferred embodiment, the training reward function is set as follows:

[0135] R=K1·|P target -P actor |+K2·|F a,target -F a,actor |+K3·|F l,target -F l,actor |

[0136] R is the reward value, K1, K2, K3 are weight values, P target is the target pressure of the cavity, P actor is the actual chamber pressure obtained by the control strategy, F a,target is the given output thrust angle, F a,actor is the actual output thrust angle, F l,target is the given output thrust, F l,actor is the actual output thrust.

[0137] Based on the same inventive concept, in a preferred embodiment of the present invention, a multi-degree-of-freedom semi-healthy device self-repair learning control system based on an adaptive torque observer is provided. Figure 2 As shown, it includes: simulation loading module, signal acquisition module, fault diagnosis and status detection module and servo control and self-repair control module. Specifically,

[0138] The simulation loading module is used to simulate actuator failure;

[0139] The signal acquisition module is used to collect the actual response data of the simulation loading module;

[0140] The fault diagnosis and state detection module uses an adaptive torque observer with generalized momentum based on actual response data to obtain fuzzy parameters and achieve adaptive estimation of external torque. It fuses and classifies system response data and external torque estimates to perform fault mode diagnosis and obtain state monitoring results based on the fault mode diagnosis.

[0141] The servo and self-repair control module establishes a self-repair control strategy based on reinforcement learning to realize the self-repair control of multi-degree-of-freedom semi-healthy devices.

[0142] Furthermore, in a preferred embodiment, the simulation loading module includes a linear drive submodule, a loading controller, a simulation collision rod, and a pressure sensor, and the simulation loading module is used to simulate actuator failure; wherein the simulation collision rod is arranged on the linear drive submodule, and the linear drive controls the collision rod to extend into the interior of the valve to simulate blockages at different positions, thereby obtaining response data under different degrees of blockage.

[0143] The signal acquisition module includes a data collector for regularly collecting multi-channel data. The multi-channel data includes actuator current data, actuator displacement data, actuator speed data and environmental status data. The environmental status data includes fluid pressure data, fluid density data and fluid temperature data.

[0144] The servo and self-repair control module adopts a self-repair control strategy based on the actual response signal and status monitoring results to obtain the displacement target of the linear drive sub-module in the simulated loading module, control the simulated collision rod to extend into the valve according to the displacement target, and track the given cavity pressure and output thrust vector, thereby realizing self-repair control of the multi-degree-of-freedom semi-healthy device.

[0145] The modules / units in the above example of the present invention may specifically refer to the implementation technology of the corresponding steps of the self-repair learning control method of an adaptive torque observer semi-healthy device in the above embodiment, which will not be repeated here.

[0146] In order to verify the feasibility and effectiveness of the self-repair learning control method and system of an adaptive torque observer semi-healthy device in the above embodiment, in a specific embodiment of the present invention, the adaptive dynamic model and the fixed parameter dynamic model are compared for the system current response estimation. The estimation accuracy of the former is improved by 25%, as shown in FIG. Figure 3 As shown in the figure, the collision detector based on the adaptive dynamic model is compared with other classification methods, including SVM, 1D-CNN, 1D-LeNet, 1D-ResNet, 1D-AlexNet, Th-2A (threshold setting method), and GMO (generalized momentum observer). The collision detector based on the adaptive dynamic model (DMAO) has the best classification effect, as shown in the figure. Figure 4 As shown in Figure 2, the adaptive torque observer semi-health device self-repair learning control method is compared with the nominal control method (without self-repair control capability) to track the given pressure and inference signal. The control effect of the former is significantly better than the latter, as shown in Figure 2. Figure 5 shown.

[0147] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various modifications or variations within the scope of the claims without affecting the essence of the present invention. The above preferred features may be used in any combination as long as they do not conflict with each other.

Claims

1. A self-repair learning control method for a semi-healthy device with an adaptive torque observer, characterized in that: include: Establish a system dynamics model for multi-degree-of-freedom control; Based on the system dynamics model, an adaptive torque observer based on generalized momentum is established; Collecting system response data and obtaining fuzzy parameters based on the adaptive torque observer to achieve adaptive estimation of external torque; fusing and classifying the system response data and the external torque estimation, performing fault mode diagnosis, and obtaining a condition monitoring result based on the fault mode diagnosis; Based on the state monitoring results, a self-repair control strategy based on reinforcement learning is established to realize self-repair control of the multi-degree-of-freedom semi-healthy device; The generalized momentum-based adaptive torque observer includes the system dynamics model and a neural network, wherein the neural network adaptively estimates fuzzy parameters of the system dynamics model based on system response data collected in a no-collision and no-fault mode; The acquisition system response data and the fuzzy parameters are obtained based on the adaptive torque observer to achieve adaptive estimation of the external torque, including: in, represents the estimate of the external torque, K is a positive real value, τ m is the actuator input torque, τ e is the external torque, p is the system generalized momentum, and α is a physical quantity containing fuzzy parameters; τ e =-F L lcosθ=-(F g -F c )lcosθ x=lsinθ Where M is the generalized mass, M=J1+J2+ml 2 cos 2 θ, F g For normal load, F c is the interference load caused by collision factors, x is the displacement of the valve core, β i That is the fuzzy dynamic parameter; The expression in α is: in w i,n are weight values, which are calculated by their respective expressions, where the subscript n is the serial number of the discrete sampling value, indicating that these are the four weight values ​​w and the constant value b calculated for the nth group of signals, and then the nth α is calculated.

2. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 1, characterized in that: The system dynamics model of the multi-degree-of-freedom control includes dynamic models of multiple regulating valves, wherein the dynamic model of a single regulating valve is: where τ m Represents the motor input torque, τ f Represents the friction torque of the motor rotor, F f1 represents the friction between the cam pin and the cam groove, l represents the distance between the axis of the cam pin and the cam axis, θ represents the output angle of the motor reducer, R represents the radius of the cam pin, F x1 Represents the normal contact force between the cam pin and the cam groove, F f2 Represents the friction force of the gas valve body on the valve stem, F L represents the external load, J1 represents the sum of the motor rotor inertia and the reducer inertia, J2 represents the rotational inertia of the cam, and m represents the mass of the valve core.

3. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 1, characterized in that: The neural network adopts a multi-layer perceptron MLP; the input of the neural network is the system response data, and the output is the fuzzy parameter β i , i=1~4 estimation, and finally through the formula Calculate α n ; According to the system response data in the no-collision no-fault mode, the neural network in the adaptive external torque estimator is trained by the back-propagation formula, specifically: E represents the error, δw represents the calculation error of gradient descent, and η is the learning rate.

4. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 1, characterized in that: The system response data includes: actuator current data, actuator displacement data, actuator speed data and environmental status data, and the environmental status data includes fluid pressure data, fluid density data and fluid temperature data; The failure modes include: actuator blockage or actuator jam; The fault mode determination includes: utilizing a classification network, inputting system response data and external torque estimation data, and outputting the fault mode; The state monitoring includes: determining the degree of valve port blockage based on the determined failure mode.

5. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 4, characterized in that: The classification network is trained by back propagation based on the response data of the faulty system, specifically: y c,n is the real classification label, is the output prediction of the network, M is the number of samples involved in the calculation, and τ is the learning rate.

6. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 4, characterized in that: The condition monitoring includes: When the fault mode is diagnosed as actuator blockage, calculate the valve port blockage fault degree ρ; i c,n is the current value when the collision occurs, is the estimated current value without external load calculated by the adaptive dynamic model, and i0 is the threshold selected according to the motor performance parameters; When the fault mode is diagnosed as actuator jam, the jam fault degree is not calculated.

7. The self-repair learning control method of a semi-healthy device with an adaptive torque observer according to claim 1, characterized in that: The self-repair control strategy based on reinforcement learning is established based on the state monitoring results to realize the self-repair control of the multi-degree-of-freedom system, including: The goal of the self-repair control strategy is to track the given cavity pressure and output thrust vector; The inputs of the self-repair control strategy are set as: a given cavity pressure target, a given output thrust vector target, the collected actual cavity pressure, the collected actual thrust vector, and the calculated fault degree of each actuator; The output of the self-repair control strategy is set as: the servo displacement target of each actuator; Based on the PPO algorithm, the self-repair control strategy training is carried out, and the training loss function is: π is the current control strategy, π old is the old policy corresponding to the data currently being trained, s is the system state, a is the action output by the policy based on the state s, A represents the advantage function, and ε is a value between (0, 1) that is used to limit the gap between the old and new policies to a set range; The form of the training reward function is set as: R=K1·|P target -P actor |+K2·|F a,target -F a,actor |+K3·|F l,target -F l,actor | R is the reward value, K1, K2, K3 are weight values, P target is the target pressure of the cavity, P actor is the actual chamber pressure obtained by the control strategy, F a,target is the given output thrust angle, F a,actor is the actual output thrust angle, F l,target is the given output thrust, F l,actor is the actual output thrust.

8. An adaptive torque observer semi-healthy device self-repair learning control system, used to implement the adaptive torque observer semi-healthy device self-repair learning control method according to any one of claims 1 to 7, characterized in that: include: Simulation loading module, signal acquisition module, fault diagnosis and status detection module and servo and self-repair control module; among them, The simulation loading module is used to simulate actuator failure; The signal acquisition module is used to collect actual response data of the simulation loading module; The fault diagnosis and state detection module uses a generalized momentum-based adaptive torque observer based on the actual response data to obtain fuzzy parameters and achieve adaptive estimation of external torque; fuses and classifies the system response data and the external torque estimate to perform fault mode diagnosis, and obtains a state monitoring result based on the fault mode diagnosis; The servo and self-repair control module establishes a self-repair control strategy based on reinforcement learning to achieve self-repair control of multi-degree-of-freedom semi-healthy devices.

Citation Information

Patent Citations

  • Fault diagnosis and fault-tolerant control method for underwater robot propeller

    CN114115195A

  • Variant aircraft sensor fault diagnosis and fault-tolerant control method

    CN117806276A

  • Electric control power-assisted braking system current sensor fault diagnosis and fault-tolerant control method

    CN117864087A

  • Self-adaptive compensation method for out-of-control fault of actuator of satellite attitude control system

    CN108563131A

  • Quadruped robot intelligent control method and system based on deterministic learning

    CN118131807A