Online self-identification predictive control method and system for underwater robot with unknown propulsion configuration

CN122431154BActive Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610902645.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-25
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0004]此外,现有自适应控制方法虽然能够补偿部分模型不确定性,但其构造方式多以残差项、误差项或扰动项补偿为主,通常不能直接学习与控制求解相关的有效作用映射关系,并且仍依赖参数近似正确的基础模型

Benefits of technology

指令映射与执行模块,用于将所述执行器实际输入目标转换为推进器转速指令、推力指令或电流指令,并发送至推进器驱动层完成闭环执行;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431154B_ABST
    Figure CN122431154B_ABST
Patent Text Reader

Abstract

The application discloses an online self-identification predictive control method for an underwater robot with unknown propulsion configuration, and comprises the following steps: establishing a normalized dynamic model base; performing online identification and iterative updating on parameters of the normalized dynamic model base based on actuator input and corresponding robot state and response, and outputting model identification accuracy index and identification sufficiency index; constructing a model predictive controller, and solving control components according to robot state estimation and control targets; injecting a zero-bias excitation signal as an excitation component; continuously generating and adjusting control budget and excitation budget according to the model identification accuracy index and the identification sufficiency index, and superimposing the control components and the excitation components under the constraints of the control budget and the excitation budget to obtain an actuator actual input target. The application further discloses an online self-identification predictive control system for an underwater robot with unknown propulsion configuration. The application can realize online self-identification and predictive control under the condition of unknown propulsion configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of adaptive control methods, and specifically relates to an online self-identification and predictive control method and system for underwater robots with unknown propulsion configurations. Background Technology

[0002] Existing underwater robot control technologies largely rely on strong prior conditions. For example, methods based on complete dynamic models or nominal thruster allocation matrices typically require that the body mass, damping, restoring force parameters, and thruster configuration relationships be known in advance or obtained through calibration. When the platform model, thruster layout, installation direction, or load conditions change, the original model and allocation relationships are prone to mismatch, leading to a decline in control performance. For engineering scenarios requiring cross-platform deployment or rapid on-site integration, the applicability of these methods is limited.

[0003] In recent years, some solutions have introduced neural networks, reinforcement learning, or other data-driven algorithms to improve control capabilities. However, these methods typically rely on offline training, policy search, or large-scale sample accumulation, and have high requirements for the consistency between the training environment and actual working conditions. For underwater robot systems with unknown thruster configurations, inaccurate dynamic parameters, or continuously changing field conditions, such solutions usually require additional training and transfer processes, resulting in high on-site deployment costs. For example, Chinese patent CN115167486A discloses an online hydrodynamic parameter identification method based on a dynamic model. This method utilizes the error between navigation sensor measurements and the velocity prior of the dynamic model, as well as the data correlation feedback between hydrodynamic parameters and velocity priors, to correct hydrodynamic parameters and estimate the accuracy of parameter identification. Chinese patent CN122072452A discloses a multi-task fast adaptive control method for underwater robots based on meta-reinforcement learning, including: constructing an underwater robot dynamics model and a thrust allocation strategy matrix; designing three reinforcement learning controllers for different task requirements: position without overshoot, position with allowable overshoot, and thruster flexible control; introducing a meta-learning mechanism to build a meta-training platform to train the three reinforcement learning controllers and obtain a set of optimal initialization parameters; deploying the obtained optimal initialization parameters and sub-task reinforcement learning controllers to the underwater robot and performing two-stage training according to different tasks; finally, when the controller is deployed in engineering practice, it can intelligently adjust the thruster output based on the real-time calculated relative distance and velocity information to the target position, ensuring accurate position control in various task scenarios.

[0004] Furthermore, while existing adaptive control methods can compensate for some model uncertainties, their construction primarily relies on residual terms, error terms, or disturbance terms for compensation. They typically cannot directly learn the effective action mapping relationship related to the control solution and still depend on a basic model with approximately correct parameters. When the thruster configuration uncertainty and hydrodynamic uncertainty are significant, residual compensation methods struggle to uniformly address the corresponding mismatches.

[0005] In summary, at least the following technical problems exist: First, when the thruster configuration is unknown or only partially known, or when a complete prior hydrodynamic model is missing or inaccurate, existing methods struggle to directly establish an effective model for control solutions. Second, when thruster configuration and hydrodynamic uncertainties are significant, conventional residual compensation adaptive methods struggle to uniformly handle and effectively absorb the characteristics of the physical model. Third, during zero-model start-up and abrupt changes in operating conditions (such as changes in platform load, thruster installation differences, thruster effectiveness changes, and sudden changes in environmental conditions), existing solutions are prone to problems such as abrupt control takeover switching and difficulty in timely recovery after parameter inaccuracies, thus affecting system safety, availability, and engineering deployment efficiency.

[0006] Therefore, it is still necessary to propose a technical solution that is more suitable for actual deployment conditions, so that underwater robot systems can still be directly put into the mission and complete model identification and control takeover online from a zero-model cold start state when the thruster configuration is unknown or only partially known, the complete prior hydrodynamic model is missing or the initial model is inaccurate. During the mission execution, the model estimation can be continuously optimized, and the control takeover capability can be quickly restored when the working conditions change significantly. Summary of the Invention

[0007] The purpose of this invention is to provide an online self-identification and predictive control method and system for underwater robots with unknown propulsion configurations, which can achieve online self-identification and predictive control in the case of unknown propulsion configurations.

[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution: An online self-identification and predictive control method for an underwater robot with an unknown propulsion configuration includes: 1) Establish a normalized dynamics model base to uniformly characterize the dynamics and propulsion configuration of underwater robots with unknown propulsion configurations; 2) Based on the actuator input and its corresponding robot state and response, the parameters of the normalized dynamic model base are identified and iteratively updated online, and the model identification accuracy index and identification sufficiency index are output; 3) Construct a model predictive controller based on the normalized dynamic model base, solve the control components of the actuator input according to the robot state estimation and control objectives, and update the model parameters in real time according to the online identification results of step 2); 4) Inject a zero-bias excitation signal into the actuator input terminal as the excitation component of the actuator input; 5) Based on the model identification accuracy index and identification sufficiency index output in step 2), continuously generate and adjust the control budget and incentive budget, and under the constraints of the control budget and incentive budget, superimpose the control component and incentive component to obtain the actual input target of the actuator.

[0009] In step 1), the normalized dynamic model base includes an effective control action mapping, a linear damping term, a quadratic damping term, and an attitude-related recovery term; the unknown propulsion configuration includes the number of propellers, the correspondence between propeller numbers and control channels, the propeller installation direction, the propeller installation position, the thrust distribution relationship, or any combination thereof, and allows online identification to be initiated when the nominal propeller allocation matrix is ​​missing (allows cold start with zero values, preset initial values, or engineering approximations when the accurate propeller allocation matrix and complete mass, damping, and recovery force parameters are missing).

[0010] In step 2), the actuator input includes one or more of the following: thruster thrust command, rotation speed command, current command, thruster thrust estimate, or thruster channel feedback; the robot state and response include one or more of the following: pose observation estimate, velocity estimate, acceleration estimate, body axis velocity change, or generalized motion response.

[0011] In step 2), the online identification and iterative update adopts a recursive least squares update mechanism, and is adjusted by combining an adaptive forgetting factor and covariance constraints, covariance contraction or covariance pullback mechanisms to improve numerical stability and parameter tracking ability under changing operating conditions; the model identification accuracy index includes one or more of the following: prediction error index or residual statistics index; the identification sufficiency index includes one or more of the following: parameter convergence index, covariance stability index, incentive coverage index or operating condition novelty index.

[0012] In step 2), online identification and iterative updates can make the model parameters approximate the dynamics of the real robot.

[0013] In step 3), the model predictive controller is a nonlinear model predictive controller, which directly uses the actuator input vector or the thrust scalar thrust combination as decision variables to generate the control components, and applies thrust amplitude constraints and thrust rate of change constraints to the control components.

[0014] In step 3), the model parameters of the model predictive controller are updated in real time with the online identification results of step 2) to form indirect adaptive control.

[0015] In step 4), the zero-bias excitation signal is one or more of a multi-channel quadrature excitation signal, a coded periodic excitation signal, or a band-limited perturbation signal, and satisfies the characteristics of zero mean bias and zero integral bias.

[0016] Specifically, the zero-bias excitation signal can be a Hadamard-coded multi-frequency sinusoidal excitation signal, or a multi-channel excitation signal constructed from multi-frequency sinusoidal basis functions through symbol encoding, phase encoding, or orthogonal encoding. It can also be a piecewise positive and negative symmetrical perturbation signal, a band-limited periodic perturbation signal, or a band-limited random signal.

[0017] In step 4), in order to meet the continuous excitation condition, a zero-bias excitation signal needs to be injected into the actuator input.

[0018] In step 5), the control budget and the excitation budget are generated respectively, satisfying that they are non-negative and their sum does not exceed the preset budget upper limit, and are used as the generation side amplitude constraints of the control component and the excitation component respectively, so that the actual input target of the actuator after the two are directly mixed satisfies the actuator saturation constraint.

[0019] Furthermore, when the prediction error decreases or the identification sufficiency improves, the control budget is increased and the incentive budget is decreased. When at least one of the following occurs: increased prediction error, abrupt change in residual statistics, worsening covariance, decreased thruster effectiveness, or increased operational novelty, the control budget is decreased and the incentive budget is increased to enter the rollback and re-identification phase. When the model quality recovers, the control budget is gradually released and the incentive budget is decreased, thereby restoring stable control takeover.

[0020] The method further includes: 6) converting the actual input target of the actuator into a thruster speed command, thrust command or current command, and sending it to the thruster drive layer to complete closed-loop execution.

[0021] The present invention achieves online identification, model predictive control, stimulus injection, budget adjustment, generation-side limiting and instruction mapping in sequence through the above steps 1)-6), and executes them continuously within the same real-time control loop.

[0022] The present invention also provides an online self-identification and predictive control system for an underwater robot with an unknown propulsion configuration employing the above-described method, characterized in that it comprises: The model base establishment module is used to establish a normalized dynamic model base to uniformly represent the dynamics and propulsion configuration of underwater robots with unknown propulsion configurations. The status and reference information acquisition module is used to acquire actuator inputs and their corresponding robot states and responses, as well as control objectives; The online identification module is used to identify and iteratively update the parameters of the normalized dynamic model base based on the actuator input and its corresponding robot state and response, and output the model identification accuracy index and identification sufficiency index. The predictive control module is used to construct a model predictive controller based on a normalized dynamics model base, and to solve for the control components of the actuator input based on the robot state estimation and control objectives. The excitation injection module is used to inject a zero-bias excitation signal into the actuator input terminal as the excitation component of the actuator input. The budget adjustment module is used to generate and continuously adjust the control budget and incentive budget based on the model identification accuracy index and identification sufficiency index, respectively. The budget limiting and input mixing module is used to limit the generation amplitude of the control component and the excitation component according to the control budget and the excitation budget respectively, and directly mix the limited control component and the excitation component to obtain the actual input target of the actuator that satisfies the actuator saturation constraint and the total input budget constraint.

[0023] Specifically, the state and reference information acquisition module acquires actuator input, robot state estimation, and actuator feedback state from the perception software and hardware, and acquires control targets from the planning and control upper layer; the robot state estimation includes robot pose observation estimation and velocity estimation, and the actuator feedback state includes thruster channel feedback and motion response; the predictive control module updates the model parameters of the model predictive controller in real time with the online identification results; The system includes: The instruction mapping and execution module is used to convert the actual input target of the actuator into a thruster speed command, thrust command or current command, and send it to the thruster drive layer to complete closed-loop execution; The online identification module, predictive control module, excitation injection module, budget adjustment module, budget limiting and input mixing module, and instruction mapping and execution module are deployed in the same real-time control loop and form a closed loop through feedback from the thruster channel.

[0024] The system provided by this invention continuously completes state acquisition, model update, control solution, excitation injection, budget adjustment, generation side limiting, input mixing and execution output within the same real-time control loop, and converts the actual input target of the actuator into thruster speed command, thrust command or current command and sends it to the thruster drive layer to complete closed-loop execution.

[0025] Compared with existing technologies, this invention has at least the following advantages: First, it can still achieve zero-model or weak-model startup and directly enter the closed-loop control and online identification process without relying on precise thruster allocation matrices and complete prior hydrodynamic parameters, reducing the threshold for cross-platform deployment and field access; Second, by online identification of effective control action mapping and related dynamic parameters, it can uniformly absorb thruster configuration deviations, hydrodynamic uncertainties, and thruster effectiveness changes, rather than only handling mismatches through residual compensation; Third, by limiting the generation amplitudes of control and excitation components through control and excitation budgets respectively, the two can be directly mixed to still satisfy actuator constraints, reducing the risk of input mutations and actuator saturation, and avoiding changes in control gain through the mixing process; Fourth, it can automatically enter backtracking and re-identification under conditions such as model quality degradation, load mutations, or environmental disturbances, and resume control takeover after model recovery, thereby improving the robustness, availability, and engineering deployment efficiency of the system. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the control system structure of the present invention; Figure 2 A comparison diagram of three-dimensional trajectory tracking of three control methods under mid-load sudden change conditions; Figure 3 The timing diagram shows the actual applied control budget coefficient and excitation budget coefficient of the method of the present invention under the condition of sudden load change during the middle of the operation. Figure 4 The diagram shows the time series of the online identification quality index of the method of the present invention under the condition of sudden load change, illustrating the time series changes of the one-step prediction error and the identification quality index before and after the sudden load change. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following embodiments are used to illustrate the technical solution of the present invention, and are not intended to limit the scope of protection of the present invention.

[0028] like Figure 1 As shown, the online self-identification and predictive control method for underwater robots with unknown propulsion configurations provided by this invention is executed through a control system, which mainly consists of an identification layer, an evaluation layer, a control layer, an interface layer, and a platform layer. The interface layer and platform layer can serve as peripheral adapters, used to acquire the underwater robot's body state and publish the underlying control objectives. The identification layer, evaluation layer, and control layer together form an indirect adaptive model predictive controller, with the adaptive capability derived from the real-time updating of the predictive model parameters of the model predictive controller by the identifier.

[0029] In one specific application, the system is deployed on a fully driven multi-thruster underwater robot platform, and the whole system consists of a host control computing unit, a state estimation unit, a bottom-level thruster execution unit, and a sensor synchronous acquisition unit.

[0030] The upper-level control computing unit is used to run the predictive controller, online identifier, excitation injection module, budget adjustment module, and budget limiting and input hybrid module. The state estimation unit can be composed of one or more of an inertial measurement unit, depth gauge, Doppler velocimeter, visual inertial odometry, or external motion capture system, used to output attitude, position, velocity, or acceleration estimates. The lower-level thruster execution unit can consist of an electronic speed controller, thruster driver, encoder, or speed sampling module, used to execute thrust control according to the final thruster commands. The sensor synchronization acquisition unit is used to synchronously acquire body state, thruster channel feedback, and environmental information, ensuring that identification and control operate on a unified time reference.

[0031] In terms of connectivity, the state estimation unit can be connected to the upper control computing unit via serial port, fieldbus, Ethernet or other data communication interfaces; the upper control computing unit sends the final thruster command to the lower thruster execution unit, and the lower execution unit returns the speed, thrust estimate or execution status; thruster channel feedback, attitude and velocity observation continue to be transmitted back to the online identification module and the budget adjustment module, thus forming a complete closed loop.

[0032] Specifically, the online self-identification and predictive control method for underwater robots with unknown propulsion configurations provided by this invention includes: S1. Establish a normalized dynamic model base to uniformly represent the dynamics and propulsion configuration of an underwater robot with an unknown propulsion configuration: The normalized dynamic model base includes an effective control action mapping, a linear damping term, a quadratic damping term, and an attitude-related recovery term; The unknown propulsion configuration includes the number of thrusters, the correspondence between thruster numbers and control channels, the thruster installation direction, the thruster installation position, the thrust distribution relationship, or any combination thereof, and allows online identification to be initiated when the nominal thruster allocation matrix is ​​missing.

[0033] Taking the parametric modeling of a six-DOF underwater robot as an example, in one implementation, the normalized dynamic model base has the following form:

[0034] in, For the six-degree-of-freedom body axis velocity vector, For multi-thrust scalar thrust combination, To effectively control the mapping of effects, and These represent the linear and quadratic damping terms, respectively. This indicates attitude-related recovery terms.

[0035] When the number of thruster control channels is At that time, the normalized dynamic model can have The model has 1,000 degrees of freedom, where the effective control action is mapped to the number of thruster channels. Linear damping, quadratic damping, and recovery terms are used to absorb the main dynamic effects in low-speed underwater operations. This model does not require the robot to have a specific symmetric structure, nor does it require a precise thruster allocation matrix before startup. The above model parameters can be combined into a parameter vector. The model prediction controller in S2 is updated by the online identifyer in each identification cycle.

[0036] S2. Based on the actuator inputs and their corresponding robot states and responses, the parameters of the normalized dynamic model base are identified and iteratively updated online, and the online identification results, model identification accuracy indicators, and identification sufficiency indicators are output. The actuator inputs include one or more of the following: thruster thrust command, rotation speed command, current command, thruster thrust estimate, or thruster channel feedback. The robot states and responses include one or more of the following: pose observation estimate, velocity estimate, acceleration estimate, body axis velocity change, or generalized motion response. The online identification and iterative update adopts a recursive least squares update mechanism, combined with an adaptive forgetting factor and covariance constraints, covariance contraction, or covariance pullback mechanisms for adjustment. The model identification accuracy indicators include one or more of the following: prediction error indicators or residual statistics indicators. The identification sufficiency indicators include one or more of the following: parameter convergence indicators, covariance stability indicators, excitation coverage indicators, or working condition novelty indicators.

[0037] The recursive least squares update mechanism is implemented through the following methods: (1) Run the discretized observation model on the six degrees of freedom of the organism. For the i-th degree of freedom, the following observation relationship can be constructed: ; in, For the first i Each degree of freedom at time... k The observed output, Let be the regression vector for the i-th degree of freedom, constructed from the system state, input, and historical data. Let be the parameter vector to be identified for the i-th degree of freedom. This is for observing the noise term.

[0038] (2) Update using a row-by-row recursive least squares method: ; ; ; ; ; ; The error (prior error) is the prediction error using the parameters from the previous step, where For the first i Each degree of freedom at time... k Actual observation output, For the first i Transpose of a regression vector with 1 degree of freedom For the first i Each degree of freedom at time... k -1 is the parameter estimate; λ[k] is the time interval. k The adaptive forgetting factor, where Let pe[k] be the minimum forgetting factor, and pe[k] be the combined one-step prediction error. As an error reference scale; For the first i A gain vector with n degrees of freedom, where For the first i Each degree of freedom at time... k -1 covariance matrix; For the first i Each degree of freedom at time... k The parameter estimates; The covariance matrix after standard RLS update. This is the covariance matrix after pullback, where σ is the covariance pullback coefficient. For the first i The initial covariance matrix of degrees of freedom.

[0039] Through the above updates, S2 obtained (That is, the online identification results) are not used for offline modeling and restoration, but are directly used as control-related model parameters in the next control cycle, thereby forming an indirect adaptive closed loop between the identification layer and the control layer.

[0040] S3. Construct a model predictive controller based on the normalized dynamic model base, solve the control components of the actuator input according to the robot state estimation and control objectives, and update the model parameters in real time with the online identification results of S2: The model predictive controller is a nonlinear model predictive controller, which directly uses the actuator input vector or the scalar thrust combination of the thruster as the decision variable to generate the control components, and applies thrust amplitude constraints and thrust change rate constraints to the control components.

[0041] Specifically, the nonlinear model predictive controller (NMPC) transforms motion control into the following optimization problem: ; ; ; ; ; in, It includes the stage cost and terminal cost at each time point in the prediction time domain. The stage cost includes state tracking terms. Control Quantity Items and control increments Terminal cost Used to constrain the terminal state. Indicates the current state The optimal control sequence obtained under the initial conditions; Input control components to the actuator and take the first term of the optimal control sequence. To control the budget coefficient, The nominal thrust limit for the actuator.

[0042] In S3, the model predictive controller directly uses the actuator input vector or the propeller scalar thrust combination as the decision variable, avoiding reliance on pre-existing, precisely known propeller decoupling relationships. Because the predictive model parameters are updated in real time with the online identification results, the controller can continuously correct the predictive model in response to propeller configuration deviations, changes in hydrodynamic parameters, or load variations.

[0043] S4. Inject a zero-bias excitation signal into the actuator input terminal as the excitation component of the actuator input: the zero-bias excitation signal is one or more of the following: multi-channel quadrature excitation signal, coded periodic excitation signal or band-limited disturbance signal, and satisfies the characteristics of zero mean bias and zero integral bias.

[0044] Since online model identification requires the robot to be under continuous excitation, but under cold start or weak model conditions, the model predictive controller itself may not be able to generate sufficiently rich excitation. Therefore, this invention generates excitation components at the actuator input through an excitation injection module. The excitation signal meets the requirements of consistent multi-channel spectral distribution, approximately orthogonality between channels, controllable amplitude, zero-mean bias, and zero-integral bias, so as to reduce additional drift while maintaining the excitation required for identification.

[0045] In one implementation, the excitation component can be in the form of a Hadamard-coded multi-frequency sine wave: ; in For the ()th of the Hadamard matrix 𝑯 i , j ) element, satisfying ; For the firsti Channel amplitude, satisfying ; For the excitation frequency, For phase parameters, in one implementation, the excitation signal can be made to satisfy integral zero bias by setting the phase.

[0046] S5. Based on the model identification accuracy index and identification sufficiency index output by S2, continuously generate and adjust the control budget and incentive budget, and superimpose the control component and incentive component under the constraints of the control budget and incentive budget to obtain the actual input target of the actuator; the control budget and incentive budget are generated separately, satisfying that they are non-negative and their sum does not exceed the preset budget upper limit, and are used as the generation side amplitude constraints of the control component and incentive component respectively, so that the actual input target of the actuator after the two are directly mixed satisfies the actuator saturation constraint.

[0047] Specifically, the control input of the actuator is the sum of the control component and the excitation component: ; To ensure that the total control signal still meets the control input limiting constraint, the control budget and the excitation budget must satisfy the following conditions: ; In one type of implementation, the above coupling conditions can be satisfied by scaling.

[0048] In S5, the present invention limits the range of each component from the generation side based on budget constraints, rather than changing the component gain by scaling.

[0049] S6. Convert the actual input target of the actuator into a thruster speed command, thrust command, or current command, and send it to the thruster drive layer to complete the closed-loop execution.

[0050] In a real-time control loop, the specific working method can be executed as follows: Receive the current machine state. Compared with reference trajectory Based on parameter estimation from the previous period Construct the current prediction model; solve the NMPC optimization problem to obtain the basic control vector with controlled budget limit. The incentive injection module generates incentive vectors that are subject to incentive budget constraints and satisfy zero bias constraints. The online identifier reads the latest observations and updates the parameter vector and covariance matrix; control budget coefficients are generated based on prediction error and identification quality indicators. and incentive budget coefficient The limited base control vector and excitation vector are directly mixed to output the final thruster command. The thruster execution layer maps the commands to speed, thrust, or current targets and completes closed-loop execution.

[0051] In this invention, to determine whether the online identification results meet the control usage conditions, the evaluation layer performs real-time evaluation of the identification process. This evaluation is based on the one-step prediction error of the current identification parameters and combines indicators such as coverage, stability, consistency, and novelty of new operating conditions related to the control mapping to generate a budget scheduling basis.

[0052] The one-step prediction error index can be defined as: ; Further smoothing of the prediction error yields: ; Then, control budget coefficients are generated based on this: ; in, The smaller the value, the higher the corresponding control budget coefficient.

[0053] The incentive budget coefficient can be generated based on the identification quality of the control-related mapping and the novelty of the new operating condition, and can be defined as: ; In one type of implementation, let Indicates time Control mapping Sub-block regression vector, window length is Then, we can first construct the local excitation Gram matrix: ; Therefore, the coverage index can be defined as: ; in, As a reference scale for coverage, This indicates that the more fully the control mapping subspace is stimulated within the current window.

[0054] The stability index can be defined based on whether the covariance matrix remains within an acceptable range: ; in, This represents the covariance submatrix associated with the control mapping subblock. It serves as a stability reference scale; when the covariance is too large, it indicates that the current estimate is unstable, and the corresponding stability index decreases.

[0055] The consistency index can be defined as the relative drift between the current control mapping estimate and the most recent historical mean: ; ; in, It serves as a consistency reference scale; the smaller the change in the estimation results within adjacent windows, the higher the corresponding consistency index.

[0056] The novelty index for new operating conditions can be defined based on the degree of sudden drop in readiness: ; This represents the normalized recognition readiness score. The reference scale for readiness is decreased; when the model quality deteriorates significantly in a short period of time, this index increases, which is used to characterize the system entering a new operating condition or deviating from the original operating range.

[0057] ; ; When the model is in a covered and relatively stable state, the incentive budget coefficient decreases; when the prediction error increases, the coverage decreases, or the novelty of new operating conditions increases, the incentive budget coefficient increases, and the control budget coefficient decreases accordingly, causing the closed loop to enter the back-off and re-identification stage. Through the above evaluation mechanism, the identifier, evaluation layer, and controller are linked in the same closed loop without the need for manual switching.

[0058] In this embodiment, the classic RexROV model is used as the controlled object for simulation testing. This model corresponds to a fully driven eight-thrust underwater robot platform, and the reference trajectory is a spatial figure-eight trajectory. For ease of explanation, discrete time intervals are used below. The notation, attitude and velocity observations are respectively denoted as and The thruster input vector is denoted as .

[0059] After the system is powered on, the dynamic base required for a zero-model cold start is first established, and the six-degree-of-freedom velocity dynamic expression is constructed: ; In this embodiment, , 、 and Precise calibration before startup is not required; instead, parameters are uniformly incorporated into the parameter vector for online estimation. Initial parameters can be set to zero or engineering approximations, allowing the control loop to start from a weak or zero model state.

[0060] After obtaining the look-ahead reference trajectory, the upper-level trajectory planning module outputs the reference state. The nonlinear model predictive controller is based on the current state estimate. Parameter estimation from the previous time step Generate the basic thrust vector of the thruster: ; The controller directly uses the scalar thrust combination of the thrusters as the decision variable, and simultaneously applies thrust amplitude constraints and thrust rate of change constraints to limit input mutations in the initial stage of the zero model.

[0061] To improve the identifiability of the zero-model cold start phase, an excitation input is superimposed on top of the basic control input. The excitation satisfies the mean-zero bias and integral-zero bias constraints, and can be expressed as: ; In the specific implementation, Segmented positive and negative symmetrical perturbations or band-limited periodic perturbations can be used to provide continuous excitation to the identifier and control the additional position drift and velocity drift within a small range.

[0062] Subsequently, a budget-coupled hybrid approach is applied to the control inputs and excitation inputs. Let the control budget coefficient be... The incentive budget coefficient is Then, an unconstrained combined input is first formed: ; Subsequently, the budget coupling mixer will Project or compress the data into the actuator feasible region to obtain the final thruster command. And satisfy: ; The above processing can avoid actuator saturation and discontinuous switching problems caused by the direct superposition of control and excitation quantities.

[0063] During system operation, real-time data collection , And the thruster channel feedback, and construct a discretized observation model:

[0064] For the For each degree of freedom, the parameters are updated using a row-by-row recursive least squares approach:

[0065]

[0066]

[0067] From this, we obtain These parameters are directly used as control-related model parameters for the next control cycle. The evaluation layer generates control budget coefficients and incentive budget coefficients in real time, based on the calculation methods of one-step prediction error, smoothing error, and identification quality indicators. When the model error is large or the indicators of the new operating condition increase, the control budget coefficient remains at a low level, while the incentive budget coefficient increases accordingly, corresponding to the rollback and re-identification stage. When the prediction error decreases and the identification sufficiency improves, the control budget coefficient gradually increases, while the incentive budget coefficient gradually decreases, corresponding to the control takeover recovery process.

[0068] In a simulation test of trajectory tracking under a sudden load disturbance, the total simulation time was 110 s. For the first 50 s, the classic RexROV nominal model was used to perform spatial figure-eight trajectory tracking. At 50 s, a 200 kg sudden eccentric load disturbance was introduced through model switching. Afterwards, the simulation continued under the same spatial figure-eight reference trajectory. The results of comparing PID, fixed model NMPC, and the method of this invention are as follows: 1. The root mean square error of the position trajectory using the PID method is approximately 1.3657 m, and the terminal position error is approximately 1.8105 m; 2. The root mean square error of the position trajectory of the fixed model NMPC method is approximately 0.3353 m, and the end position error is approximately 0.4154 m; 3. The root mean square error of the position trajectory of the method of the present invention is approximately 0.4915 m, and the terminal position error is approximately 0.2103 m.

[0069] Under this sudden load condition, the PID method exhibits a large positional error; the fixed model NMPC has a smaller overall root mean square error, but its terminal positional error is still affected by model mismatch; the method of this invention has a smaller terminal positional error. The corresponding trajectory results are shown below. Figure 2 (In the classic RexROV model spatial figure-eight trajectory tracking task, after introducing a 200 kg sudden load at 50 s, the spatial trajectories of PID, fixed model NMPC, and the method of this invention under the same reference trajectory), the timing of control budget coefficients and excitation budget coefficients is shown in [the table]. Figure 3 (Time series changes of control budget coefficients and excitation budget coefficients after introducing a 200 kg sudden load at 50 s), the time series changes of one-step prediction error and online identification quality index before and after the sudden load are shown in [reference needed]. Figure 4 .

[0070] Unlike comparative reference methods, the method of this invention does not input the hydrodynamic model and thrust distribution model of the actual machine into the controller, yet it still achieves the control performance of similar models and exhibits adaptive capabilities.

[0071] Based on the same inventive concept, this invention also provides an online self-identification and predictive control system for an underwater robot with an unknown propulsion configuration, comprising: a model base establishment module for establishing a normalized dynamic model base to uniformly represent the dynamics and propulsion configuration of the underwater robot with an unknown propulsion configuration; a state and reference information acquisition module for acquiring actuator inputs and their corresponding robot states and responses, as well as control objectives; an online identification module for online identification and iterative updating of the parameters of the normalized dynamic model base based on actuator inputs and their corresponding robot states and responses, and outputting model identification accuracy and identification sufficiency indices; and a predictive control module for... A normalized dynamics model base is used to construct a model predictive controller, and the control components of the actuator input are solved based on robot state estimation and control objectives. An excitation injection module injects a zero-bias excitation signal into the actuator input as the excitation component. A budget adjustment module generates and continuously adjusts the control budget and excitation budget based on model identification accuracy and sufficiency indices, respectively. A budget limiting and input mixing module limits the generation amplitudes of the control and excitation components based on the control and excitation budgets, and directly mixes the limited control and excitation components to obtain the actual actuator input target that satisfies the actuator saturation constraint and the total input budget constraint. For detailed implementation process, please refer to the method provided above.

Claims

1. A method for online self-identification and predictive control of an underwater robot with an unknown propulsion configuration, characterized in that, include: 1) Establish a normalized dynamic model base to uniformly represent the dynamics and unknown propulsion configuration of underwater robots with unknown propulsion configurations; 2) Based on the actuator input and its corresponding robot state and response, the parameters of the normalized dynamic model base are identified and iteratively updated online, and the model identification accuracy index and identification sufficiency index are output; 3) Construct a model predictive controller based on the normalized dynamic model base, solve the control components of the actuator input according to the robot state estimation and control objectives, and update the model parameters in real time according to the online identification results of step 2); 4) Inject a zero-bias excitation signal into the actuator input terminal as the excitation component of the actuator input; 5) Based on the model identification accuracy index and identification sufficiency index output in step 2), continuously generate and adjust the control budget and incentive budget, and under the constraints of the control budget and incentive budget, superimpose the control component and incentive component to obtain the actual input target of the actuator; In step 1), the normalized dynamic model base includes an effective control action mapping, a linear damping term, a quadratic damping term, and an attitude-related recovery term; the unknown propulsion configuration includes the number of propellers, the correspondence between propeller numbers and control channels, the propeller installation direction, the propeller installation position, the thrust distribution relationship, or any combination thereof, and allows online identification to be initiated when the nominal propeller allocation matrix is ​​missing; In step 5), the control budget and the excitation budget are generated respectively, satisfying that they are non-negative and their sum does not exceed the preset budget upper limit, and are used as the generation side amplitude constraints of the control component and the excitation component respectively, so that the actual input target of the actuator after the two are directly mixed satisfies the actuator saturation constraint.

2. The method according to claim 1, characterized in that, In step 2), the actuator input includes one or more of the following: thruster thrust command, rotation speed command, current command, thruster thrust estimate, or thruster channel feedback; the robot state and response include one or more of the following: pose observation estimate, velocity estimate, acceleration estimate, body axis velocity change, or generalized motion response.

3. The method according to claim 1, characterized in that, In step 2), the online identification and iterative update adopt a recursive least squares update mechanism, and are adjusted by combining an adaptive forgetting factor and covariance constraints, covariance contraction or covariance pullback mechanisms; the model identification accuracy index includes one or more of the following: prediction error index or residual statistics index; the identification sufficiency index includes one or more of the following: parameter convergence index, covariance stability index, excitation coverage index or operating condition novelty index.

4. The method according to claim 1, characterized in that, In step 3), the model predictive controller is a nonlinear model predictive controller, which directly uses the actuator input vector or the thrust scalar thrust combination as decision variables to generate the control components, and applies thrust amplitude constraints and thrust rate of change constraints to the control components.

5. The method according to claim 1, characterized in that, In step 4), the zero-bias excitation signal is one or more of a multi-channel quadrature excitation signal, a coded periodic excitation signal, or a band-limited perturbation signal, and satisfies the characteristics of zero mean bias and zero integral bias.

6. The method according to claim 1, characterized in that, The method further includes: 6) converting the actual input target of the actuator into a thruster speed command, thrust command or current command, and sending it to the thruster drive layer to complete closed-loop execution.

7. An online self-identification and predictive control system for an underwater robot with an unknown propulsion configuration employing the method described in any one of claims 1-6, characterized in that, include: The model base establishment module is used to establish a normalized dynamic model base to uniformly represent the dynamics and propulsion configuration of underwater robots with unknown propulsion configurations. The status and reference information acquisition module is used to acquire actuator inputs and their corresponding robot states and responses, as well as control objectives; The online identification module is used to identify and iteratively update the parameters of the normalized dynamic model base based on the actuator input and its corresponding robot state and response, and output the model identification accuracy index and identification sufficiency index. The predictive control module is used to construct a model predictive controller based on a normalized dynamics model base, and to solve for the control components of the actuator input based on the robot state estimation and control objectives. The excitation injection module is used to inject a zero-bias excitation signal into the actuator input terminal as the excitation component of the actuator input. The budget adjustment module is used to generate and continuously adjust the control budget and incentive budget based on the model identification accuracy index and identification sufficiency index, respectively. The budget limiting and input mixing module is used to limit the generation amplitude of the control component and the excitation component according to the control budget and the excitation budget respectively, and directly mix the limited control component and the excitation component to obtain the actual input target of the actuator that satisfies the actuator saturation constraint and the total input budget constraint.

8. The system according to claim 7, characterized in that, The system includes: The instruction mapping and execution module is used to convert the actual input target of the actuator into a thruster speed command, thrust command, or current command, and send it to the thruster drive layer to complete closed-loop execution; The online identification module, predictive control module, excitation injection module, budget adjustment module, budget limiting and input mixing module, and instruction mapping and execution module are deployed in the same real-time control loop and form a closed loop through feedback from the thruster channel.

Citation Information

Patent Citations

  • Online hydrodynamic parameter identification method suitable for underwater robot

    CN115167486A

  • Multi-task rapid adaptive control method for underwater robot based on meta-reinforcement learning

    CN122072452A

  • STM32-based autonomous underwater robot intelligent attitude control system and method

    CN120779999A

  • Industrial mechanical arm joint rigidity online identification method based on optimal excitation trajectory

    CN121680284A