A Multimodal Adaptive Enhancement-Based Method for Stabilizing Vehicle Masts
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0010]本发明的一个目的在于提出一种基于多模态自适应强化的车载桅杆稳平方法,针对现有技术将长桅杆按刚体建模导致柔性振动难以抑制、主导频率与相位漂移下稳控性不足、难以兼顾角度行程与应变等安全约束的问题,提出了统一时钟与坐标的多源感知,刚体与柔性分层的库普曼映射并施加振型正交与能量一致约束,主导频率与相位估计及相位增广,基于强化协同的成本权重与短地平线和长地平线调度及热启动,包含反相位前馈的双地平线学习型模型预测控制,以及融入频带能量与结构应变边界的控制障碍函数约束,并以稳平误差与频带能量的闭环指标在线更新模型与策略的技术方案
Smart Images

Figure CN121523435B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of vehicle engineering and automation control, and in particular to a method for stabilizing and leveling a vehicle mast based on multimodal adaptive reinforcement. Background Technology
[0002] Vehicle-mounted telescopic masts are widely used in emergency communication, monitoring and surveying scenarios. As the height of the mast increases and the load on the top becomes more diverse, the system exhibits obvious strong-flexible coupling characteristics. Base disturbances, wind loads and start-stop accelerations can easily excite the first and higher-order vibration modes of the mast, resulting in whip-like swinging, which leads to long leveling time, tilt overshoot and increased risk of structural strain.
[0003] Existing technologies mainly include: rigid body leveling control based on inertial or tilt sensing, such as proportional-integral-derivative and linear quadratic control; model predictive control based on finite-degree-of-freedom flexible modeling and state estimation; vibration suppression based on frequency domain filtering and notch filtering; and passive or semi-active damping devices, etc.
[0004] These methods are effective for short and medium masts or under single operating conditions, but their stability and safety are limited under long masts, frequency drift, and load variation conditions.
[0005] The main shortcomings of existing technologies are:
[0006] 1. Insufficient model and representation: Approximating the mast as a rigid body or a low-order lumped parameter model makes it difficult to accurately capture the dominant frequency and phase drift that vary with height and load, as well as the energy distribution of the flexible modes, resulting in insufficient control over the suppression of flexible vibrations.
[0007] 2. Insufficient coupling between control and constraints: Conventional model predictive control does not take safety boundaries such as frequency domain energy and structural strain as unified constraints within the feasible domain. It lacks coordinated guarantees for the convergence of energy envelope and multiple constraints such as angle, angular velocity, and travel, which can easily lead to conservative or constraint violations.
[0008] 3. Insufficient Adaptability and Collaboration: Weights and horizons are usually fixed, lacking online scheduling and warm start based on operating conditions. Sensor information fusion is insufficient in terms of clock and coordinate consistency, making it difficult to track changes in the main frequency in a timely manner and maintain the reliability and closed-loop robustness of the optimization solution.
[0009] Therefore, a method for stabilizing vehicle-mounted masts that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0010] One objective of this invention is to propose a multimodal adaptive reinforcement-based method for leveling a vehicle-mounted mast. Addressing the problems of existing technologies that model long masts as rigid bodies, leading to difficulties in suppressing flexible vibrations, insufficient stability under dominant frequency and phase drift, and difficulty in balancing safety constraints such as angle travel and strain, this invention proposes a multi-source sensing method with unified clock and coordinates, Koopman mapping for rigid and flexible layers with applied mode shape orthogonality and energy consistency constraints, dominant frequency and phase estimation and phase augmentation, cost weighting based on reinforcement synergy, scheduling of short and long horizons and hot start, dual-horizontal learning model predictive control including anti-phase feedforward, and control barrier function constraints incorporating frequency band energy and structural strain boundaries. The invention also uses closed-loop indices of leveling error and frequency band energy to update the model and strategy online. This invention provides the technical benefits of rapid leveling and suppression of dominant frequency band energy under different mast heights and loads, reduced tilt error and overshoot, ensured structural and travel safety, and improved energy efficiency and robustness.
[0011] A method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to an embodiment of the present invention is characterized by comprising:
[0012] S1. Under a unified clock and coordinate system, collect and calibrate the data of the base inertial measurement unit, the data of the mast multi-point accelerometer, the data of the mast multi-point strain gauge, and the visual displacement data of the mast top. Read the mast extension height parameters and the top load parameters and set the leveling target parameters.
[0013] S2. Based on the above data, perform hierarchical Koopman mapping to enhance the observations in the rigid body subspace and flexible mode subspace. The flexible modes satisfy the mode orthogonality constraint and the energy consistency constraint, thus obtaining the rigid body state vector and the flexible state vector.
[0014] S3. Based on the flexible state vector, estimate the dominant frequency parameters and phase parameters;
[0015] S4. Based on the rigid and flexible state vectors and combined with the dominant frequency parameter and phase parameter, state augmentation is performed. The dominant frequency parameter and the sine and cosine characteristics based on the phase parameter are concatenated with the aforementioned state to form a dominant frequency phase augmented state vector.
[0016] S5. Based on the main frequency phase augmentation state vector, mast extension height parameters and top load parameters, the enhanced coordination module is called to perform weight scheduling and hot start, and the cost weight vector, short horizon step parameters, long horizon step parameters and optimization initial values of model predictive control are obtained.
[0017] S6. Based on the cost weight vector of model predictive control, short horizon step parameters, long horizon step parameters, initial optimization values, and main frequency phase augmentation state vector, construct and solve the dual-horizontal learning model predictive control optimization problem to obtain the control command sequence.
[0018] S7. Apply the first control command of the control command sequence to the vehicle-mounted mast actuator and collect the execution response data after the action;
[0019] S8. Based on the execution response data, the stabilization target parameters and the dominant frequency parameters, calculate the stabilization error feedback and the energy feedback in the frequency band centered on the dominant frequency parameters to form a closed-loop feedback index vector.
[0020] S9. Based on the closed-loop feedback index vector, update the hierarchical Koopman mapping parameters and the enhanced collaborative module parameters for the next control cycle.
[0021] Optionally, step S1 specifically includes:
[0022] Establish a unified system clock and add timestamps to the data from the base inertial measurement unit, the mast multi-point accelerometer, the mast multi-point strain gauge, and the visual displacement data at the top of the mast. Perform time synchronization correction to align all data under a unified time reference.
[0023] Zero-bias calibration, scaling factor calibration, and temperature compensation are performed on the data from the base inertial measurement unit, the mast multi-point accelerometer, and the mast multi-point strain gauge, and the above data are expressed in a unified coordinate system;
[0024] Lens distortion correction and intrinsic / extrinsic parameter calibration are performed on the visual displacement data at the top of the mast, and the displacement is expressed in a unified coordinate system;
[0025] Noise suppression and anti-interference filtering are applied to each data point without altering the peak value of the dominant frequency band.
[0026] Read and calibrate the mast extension height parameters from the travel sensor;
[0027] Determine and verify the top load parameters based on the equipment identification and calibration table;
[0028] Set the leveling target parameters according to the operation instructions. The leveling target parameters should include at least the target tilt angle threshold, the upper limit of the dominant frequency band energy, and the maximum allowable leveling time.
[0029] The data and parameters processed above are subjected to integrity and quality checks, and unqualified samples are removed and appropriate interpolation is performed.
[0030] Output base inertial measurement unit data, mast multi-point accelerometer data, mast multi-point strain gauge data, mast top visual displacement data, mast extension height parameters, top load parameters, and leveling target parameters.
[0031] Terminology definition:
[0032] The unified clock is a clock system that uses the same time base for sampling and timestamp marking of all sensor channels in the system, ensuring that different data can be aligned in the same time domain;
[0033] The unified coordinate system is a coordinate system that unifies the outputs of each sensor to the same reference coordinate (e.g., base-fixed coordinate or inertial reference coordinate) through calibration and coordinate transformation;
[0034] The data from the base inertial measurement unit are time-series data such as acceleration, angular velocity, and derived attitude collected by the inertial measurement unit installed on the mast base or its rigid connection.
[0035] The mast multi-point accelerometer data consists of one-dimensional or three-dimensional acceleration time-series data collected by one or more accelerometers deployed at different heights or positions along the mast.
[0036] The data from the multi-point strain gauges on the mast are time-series data of strain or strain rate collected by strain sensors deployed at different positions along the mast. The strain sensors include resistance strain gauges, fiber optic gratings, or equivalent devices.
[0037] The visual displacement data of the mast top is time-series data measured by an image / visual sensor of the displacement of the mast top relative to a unified coordinate system reference, including translation and / or displacement components corresponding to the tilt angle.
[0038] The mast extension height parameter is a geometric height parameter characterizing the current extension length of the mast, which can be obtained and calibrated by a stroke or length measuring device;
[0039] The top load parameters are a set of parameters such as mass, moment of inertia, and equivalent aerodynamic / wind load characteristics corresponding to the load installed at the top of the mast, which can be obtained from equipment identification and calibration tables or estimated online.
[0040] The stability target parameters are a set of targets used for closed-loop control performance and safety constraints, including at least the target tilt angle threshold, the upper limit of the dominant frequency band energy, and the maximum allowable stability time.
[0041] The dominant frequency band peak is the energy peak that appears near the dominant frequency in the power spectrum or energy spectrum of the flexible vibration, and is used to characterize the frequency domain intensity of the main mode.
[0042] The stroke sensor is a sensor used to measure the displacement or length of the mast telescopic mechanism, including but not limited to potentiometers, encoders, laser rangefinders, or magnetostrictive displacement sensors;
[0043] The noise suppression and anti-interference filtering that does not change the peak value of the dominant frequency band is a filtering process in which the amplitude-frequency response is approximately 1 and the phase distortion is controlled within the frequency band corresponding to the dominant frequency, so as to reduce noise while keeping the position and amplitude of the peak value of the dominant frequency band from being significantly changed.
[0044] The reasonable interpolation refers to the process of interpolating or reconstructing time series data for missing or unqualified samples while ensuring the physical consistency of the data and the absence of distortion in the dominant frequency band characteristics.
[0045] Optionally, step S2 specifically includes:
[0046] Within a preset time window and sampling step size, the data from the base inertial measurement unit, the mast multi-point accelerometer, the mast multi-point strain gauge, and the visual displacement data at the top of the mast are spliced together in chronological order to form a time-series observation vector.
[0047] The time-series observation vector is lifted by performing a lifting transformation based on the hierarchical Koopman mapping, generating corresponding lifting features in the rigid body subspace and the flexible mode subspace respectively, while keeping the parameters of the two types of features from being shared.
[0048] An optimization problem with modal orthogonal constraints and energy consistency constraints is constructed. The modal orthogonal constraints limit the pairwise inner product of the flexible state vector within the time window to zero. The energy consistency constraints limit the energy of the flexible state vector to be consistent with the energy calculated based on the mast multi-point accelerometer data and mast multi-point strain gauge data within the allowable error range. The rigid body state vector and the flexible state vector are obtained by solving the optimization problem.
[0049] Output the rigid body state vector and the flexible body state vector.
[0050] Terminology definition:
[0051] The hierarchical Koopman mapping is a data-driven mapping that lifts time-series observations to a high-dimensional feature space through basis functions / dictionaries, and approximates linear dynamics in both rigid body subspace and flexible modal subspace. It consists of two parts: lifting transformation and corresponding linear evolution operators.
[0052] The rigid body subspace is a feature subspace used to characterize the rigid body motion components of the mast-base as a whole. Its state does not contain flexible deformation energy components and is used for subsequent linear prediction and control.
[0053] The flexible modal subspace is a feature subspace used to characterize flexible modal responses such as mast bending and torsion. Its state corresponds to several modal generalized coordinates or equivalent representations, which are used for frequency / phase estimation and vibration suppression control.
[0054] The preset time window is a fixed-length time interval that traces back to the past from the current moment according to a unified time sequence, and is used to construct time series observation and statistical constraints;
[0055] The sampling step size is the time interval between adjacent sampling points within the time window;
[0056] The time-series observation vector is a vector or vector sequence formed by aligning and splicing multi-source sensor data in chronological order within the preset time window.
[0057] The lifting transformation is a transformation that maps the time-series observations to high-dimensional features. It can be a combination of preset or learned basis functions / dictionaries to enhance linear predictability.
[0058] The enhanced features are high-dimensional feature representations obtained through the enhanced transformation, and are divided into corresponding sub-vectors according to rigid body and flexible modes;
[0059] The parameter non-sharing means that the parameter sets of the rigid body subspace and the flexible modal subspace on the lifting transformation and linear evolution operators are set independently and without common weights;
[0060] The mode orthogonal constraint is an orthogonal condition applied to each component of the flexible mode within the preset time window, that is, the weighted discrete inner product of any two mode components is zero. The weighted discrete inner product is a quantity obtained by multiplying the corresponding samples point by point within the time window and summing them with non-negative weights.
[0061] The energy consistency constraint is a condition that the energy of the flexible state vector is consistent with the equivalent energy calculated by the accelerometer / strain gauge within a preset tolerance.
[0062] The flexible state energy is an energy metric obtained by summing the quadratic forms of the flexible state vectors formed by a positive semi-definite weighted matrix within the preset time window, and is used to characterize the magnitude of the flexible modal energy.
[0063] The sensing equivalent energy is an energy metric obtained by summing the quadratic forms formed by the mast multi-point accelerometer data and the mast multi-point strain gauge data respectively according to the semi-positive definite weighting matrix within the preset time window, and is used to constrain the consistency with the flexible state energy.
[0064] The optimization problem aims to minimize the reconstruction / prediction error and the regularization term, while satisfying the mode orthogonality constraint and the energy consistency constraint. It is used to solve the parameter estimation problem of the lifting transformation parameters and / or state vector.
[0065] The rigid body state vector is a state representation composed of the rigid body subspace lifting features and used for linear prediction and control solutions.
[0066] The flexible state vector is a state representation composed of the enhanced features of the flexible modal subspace, used for main frequency phase estimation and vibration suppression control.
[0067] Optionally, step S3 specifically includes:
[0068] Within a preset time window, the flexible state vector is de-trended and amplitude normalized to obtain the processed flexible state vector.
[0069] The power spectral density is calculated based on the processed flexible state vector. The energy peak is located in the preset search frequency band by the joint estimation of short-time Fourier transform and autoregressive spectrum. The peak position is refined by parabolic interpolation of nearby frequency points to determine the dominant frequency parameter.
[0070] A narrowband demodulator is constructed around the dominant frequency parameter to perform bandpass filtering and analytical signal construction on the processed flexible state vector. The instantaneous phase is obtained using Hilbert transform, phase expansion is performed, and the phase parameter is obtained using the most recent zero crossover as the phase reference.
[0071] When the flexible state vector contains multiple components, the power spectrum and phase are weighted and fused according to the energy weight of each component within the preset time window to suppress the influence of anomalous components on the dominant frequency parameters and phase parameters.
[0072] Output the dominant frequency and phase parameters.
[0073] Terminology definition:
[0074] The detrending process involves removing the mean, linear, or low-order polynomial drift from the flexible state vector within a preset time window, causing it to fluctuate stably around the zero mean.
[0075] The amplitude normalization is a process of scaling the flexible state vector according to a preset amplitude specification so that its amplitude meets a uniform dimension, such as normalizing it to a unit mean square value or a unit peak value.
[0076] The power spectral density is a spectral function that describes the energy distribution of the processed flexible state vector with frequency, and is used to characterize the energy intensity of each frequency component.
[0077] The short-time Fourier transform is a processing method that performs a discrete Fourier transform on the processed flexible state vector within a sliding time window to obtain the time-frequency local spectrum.
[0078] The autoregressive spectrum is a power spectrum obtained by parametrically estimating the processed flexible state vector based on an autoregressive model, which is used to improve peak positioning resolution.
[0079] The preset search frequency band is a frequency range pre-set for the dominant frequency positioning, and its upper and lower limits are determined based on prior information such as historical working conditions, mast height and top load.
[0080] The energy peak is a local maximum value presented by the power spectral density or autoregressive spectrum within the preset search frequency band, used to indicate the frequency domain intensity of the dominant mode;
[0081] The parabolic interpolation method is to fit a quadratic curve to the spectral values of the frequency point where the energy peak is located and the two adjacent frequency points, and then find its extreme points, thereby refining the interpolation method for the dominant frequency position.
[0082] The dominant frequency parameter is the center frequency parameter representing the flexible master mode obtained by joint spectrum estimation and parabolic refinement, which is used for subsequent demodulation and control augmentation.
[0083] The narrowband demodulator is a filtering and demodulation device or algorithm that extracts and tracks the corresponding vibration components within a preset narrow bandwidth, centered on the dominant frequency parameter.
[0084] The analytical signal is a complex signal composed of a bandpass signal and its Hilbert transform, with its real part being the original bandpass signal and its imaginary part being the result of its Hilbert transform.
[0085] The Hilbert transform is an integral transform of a real signal with a 90-degree phase shift, used to construct an analytic signal to extract the instantaneous phase;
[0086] The instantaneous phase is a function of the complex argument of the analytical signal over time, used to characterize the phase position of the vibration component at the current moment;
[0087] The phase expansion is to perform de-jump processing on the instantaneous phase, and continuously expand the phase jumps spanning ±π into a time-continuous phase sequence by 2π.
[0088] The zero crossing is an event in which a bandpass signal passes through a zero value on the time axis, including the zero crossing point of either the rising or falling edge;
[0089] The phase reference is a reference event or reference value used to align the instantaneous phase; in this case, the most recent zero crossover is used as the phase reference.
[0090] The phase parameter is the instantaneous phase value corresponding to the current sampling time under the selected phase reference, which is used for state augmentation and feedforward generation;
[0091] The energy weight is a non-negative weight obtained by normalizing the energy ratio of each flexible component within a preset time window, and is used to suppress the influence of abnormal components.
[0092] The weighted fusion is performed by weighting the power spectrum and instantaneous phase of each flexible component according to energy weights to obtain robust dominant frequency parameters and phase parameters.
[0093] Optionally, step S4 specifically includes:
[0094] Under a unified clock, sampling points corresponding to the current moment of a preset time window are selected, and the rigid body state vector and flexible state vector are aligned in dimension and consistent in index.
[0095] Based on the dominant frequency parameter and the phase parameter, a sine and cosine feature based on the phase parameter is generated. The sine and cosine feature includes at least the cosine value and the sine value of the phase parameter.
[0096] The rigid body state vector, flexible state vector, dominant frequency parameter, and sine and cosine features based on phase parameter are spliced together in a predetermined order to form a dominant frequency phase augmented state vector with fixed dimensions and consistent in each control cycle.
[0097] Output the phase augmented state vector of the main frequency.
[0098] Terminology definition:
[0099] The sampling point corresponding to the current moment is the observation sample at the time index selected under a unified clock and preset time window for reading the state of this control cycle;
[0100] The dimension alignment is a process of unifying the dimension and arrangement of the rigid body state vector and the flexible state vector at selected sampling points, so that they can be directly combined and calculated without changing the physical meaning.
[0101] The indexing uniformity is to uniformly map the time index and sensor / spatial index of each component in the state vector so that the same physical moment and location have a consistent index position in the vector representation.
[0102] The sine and cosine features based on the phase parameter are periodic basis function features calculated with the phase parameter φ, including at least cos(φ) and sin(φ), used to explicitly inject phase information;
[0103] The splicing is an operation that arranges several sub-vectors or scalars in a predetermined order to form a single high-dimensional vector.
[0104] The predetermined order is the order in which vector fields are arranged and stored during the system design phase, in order to ensure interpretability and consistency across cycles.
[0105] The dimension is fixed so that the length and field structure of the concatenated vector remain unchanged in each control cycle, ensuring that the input dimension of the prediction and optimization model is constant.
[0106] The control cycle is a basic time unit of closed-loop control, which includes three stages: sampling, calculation and execution, and serves as the scheduling benchmark for state updates and command output.
[0107] The dominant frequency phase augmented state vector is an extended state representation for prediction and control obtained by splicing together the rigid body state vector, the flexible state vector, the dominant frequency parameter, and the sine and cosine features based on the phase parameter in a predetermined order.
[0108] Optionally, step S5 specifically includes:
[0109] Based on the dominant frequency phase augmented state vector and with reference to the mast extension height parameter and the top load parameter, an input quantity for scheduling is constructed. The input quantity includes the dominant frequency parameter, the sine and cosine characteristics of the phase parameter, the amplitude characteristics of the flexible mode component, the mast extension height parameter, and the top load parameter.
[0110] Based on the input, the estimated value of the frequency band energy and the phase change rate corresponding to the dominant frequency parameter are calculated and used as scheduling indicators for weight and horizon step count;
[0111] The weight scheduling submodule of the enhanced coordination module is invoked to generate a cost weight vector for model predictive control based on the scheduling index and the input quantity. The cost weight vector includes at least the tilt error term weight, the dominant frequency band energy term weight, the control increment term weight, and the energy consumption term weight. The upper and lower bound constraints and normalization processing are performed on the cost weight vector to ensure numerical stability.
[0112] The hot start submodule of the enhanced coordination module is called. It prioritizes the candidate sequence after time shifting the control instruction sequence of the previous control cycle as the control instruction part of the initial value of optimization, and uses the rolling prediction of the main frequency phase augmented state vector as the state part of the initial value of optimization. When the data of the previous control cycle is missing, the combination of the static sequence and the current state is used as the initial value of optimization.
[0113] Based on the threshold rules of the dominant frequency band energy estimate and the phase change rate, the short horizon step parameter and the long horizon step parameter are determined. When the dominant frequency band energy estimate is large or the phase change rate is large, the short horizon step parameter is increased while the long horizon step parameter is kept not lower than the preset lower limit. When the dominant frequency band energy estimate is small and the phase change rate is small, the short horizon step parameter is decreased while the long horizon step parameter is moderately increased within the allowable range of the stabilization target parameter.
[0114] Output the cost weight vector, short horizon step number parameter, long horizon step number parameter, and initial optimization value.
[0115] Terminology definition:
[0116] The input for scheduling is a feature set consisting of the dominant frequency phase augmented state vector, mast extension height parameter, and top load parameter, including the dominant frequency parameter, the sine and cosine features of the phase parameter, the amplitude features of the flexible mode component, the mast extension height parameter, and the top load parameter, which is used to drive the scheduling of weights and horizon steps.
[0117] The amplitude characteristics of the flexible modal components are amplitude statistics calculated for each modal component of the flexible state vector within a preset time window, including at least the root mean square value, peak value or energy envelope peak value, used to characterize the intensity of each flexible component.
[0118] The frequency band energy estimate corresponding to the dominant frequency parameter is the energy metric calculated from the flexible modal subspace state or its bandpass signal mapped to the sensing signal after setting a narrow band range around the dominant frequency parameter. It can be in the form of quadratic or integral energy.
[0119] The phase change rate is the rate of change of the phase parameter per unit time, which is obtained by dividing the time derivative of the instantaneous phase or the discrete difference by the sampling interval, and is used to characterize phase dynamics.
[0120] The scheduling metrics are a set of scalar / vector metrics used to determine cost weights and horizon steps, including at least the dominant frequency band energy estimate and phase change rate;
[0121] The enhanced coordination module is a strategy module that generates a model to predict and control relevant scheduling results based on the input quantities and scheduling indicators used for scheduling, and outputs a cost weight vector, short horizon step parameters, long horizon step parameters, and initial optimization values.
[0122] The weighted scheduling submodule is a strategy unit in the enhanced coordination module that calculates and outputs the model prediction control cost weight vector based on scheduling indicators.
[0123] The cost weight vector of the model predictive control is a set of coefficients used for weighted optimization of each objective term, which includes at least the tilt error term weight, the dominant frequency band energy term weight, the control increment term weight, and the energy consumption term weight.
[0124] The tilt angle error term weight is a non-negative weighting coefficient applied to the tilt angle deviation error term in the optimization objective;
[0125] The weight of the dominant frequency band energy term is a non-negative weighting coefficient applied to the dominant frequency band energy suppression term in the optimization objective;
[0126] The control increment term weight is a non-negative weighting coefficient applied to the difference between adjacent control inputs (control increment) in the optimization objective, used to limit the rate of control change;
[0127] The weight of the energy consumption item is a non-negative weighting coefficient applied to the energy or power consumption metric of the actuator in the optimization objective;
[0128] The hot start submodule is a submodule in the enhanced collaboration module used to generate the initial optimization value. It combines the information from the previous control cycle with the current state rolling prediction to improve the solution convergence.
[0129] The time shift is an operation that moves the control command sequence of the previous control cycle forward by one or more sampling steps along the time axis and truncates / pastes it to align it with the current prediction time domain.
[0130] The candidate sequence is a control command sequence that has been time-shifted and meets the limits of amplitude and rate of change, and is used as a candidate for the control command part of the initial value optimization.
[0131] The initial optimization value is the initial solution used for model predictive control optimization, which consists of two parts: control command and state.
[0132] The control command sequence is a discrete control input sequence arranged in sampling steps within the prediction time domain, used to drive the actuator;
[0133] The rolling prediction is a state sequence obtained by progressive forward deduction through a hierarchical Koppman prediction model, using the current dominant frequency phase augmented state vector as the initial condition.
[0134] The static sequence is the control instruction sequence used when the initial value is defaulted during optimization, and the control quantity at each step is zero or keeps the current value unchanged.
[0135] The short horizon step number parameter is the prediction step number used in short time domain optimization, which is used to prioritize the suppression of dominant frequency band energy and fast phase response;
[0136] The long horizon step number parameter is the prediction step number used in long time domain optimization, which is used to balance the global objectives of tilt stability and energy consumption.
[0137] The threshold rule is a set of logical rules that determine whether to increase or decrease the short horizon step number parameter and constrain the long horizon step number parameter after comparing the dominant frequency band energy estimate and phase change rate with a preset threshold.
[0138] The preset lower limit is the minimum step limit allowed by the long horizon step parameter, which is used to ensure that long-time domain targets are not excessively weakened.
[0139] Optionally, step S6 specifically includes:
[0140] The current prediction initial state is determined based on the dominant frequency phase augmented state vector, and a discrete-time prediction model and prediction time domain are established using hierarchical Koopman mapping based on the cost weight vector, short horizon step parameters, long horizon step parameters and optimization initial values.
[0141] Within the short horizon, an anti-phase feedforward term is generated based on the dominant frequency parameter and phase parameter. The sum of the anti-phase feedforward term and the control command to be optimized is used as the equivalent input of the prediction model. Limits are applied to the amplitude and rate of change of the anti-phase feedforward term to avoid excitation overshoot.
[0142] The optimization objective of the learning model predictive control is constructed by superimposing the dominant frequency band energy suppression term and the control command change rate penalty term in the short horizon and the tilt error term and energy consumption term in the long horizon using a cost weight vector. The tilt error term is used to determine the target tilt angle with reference to the stable target parameter, and the dominant frequency band energy suppression term is calculated based on the energy of the flexible mode subspace state component in the frequency band corresponding to the dominant frequency parameter.
[0143] By controlling the obstacle function, angle limit constraints, angular velocity limit constraints, travel limit constraints, cable or structural strain safety boundary constraints, and vibration energy envelope upper bound and monotonically decreasing constraints are gradually applied within the prediction step, so that the corresponding safety function remains non-negative in the prediction time domain and does not increase at a predetermined convergence rate.
[0144] Under the above objectives and constraints, a hot start solution is performed using optimized initial values to obtain a control command sequence that meets the requirements of the control barrier function;
[0145] Output control command sequence.
[0146] Terminology definition:
[0147] The current prediction initial state is the current state extracted from the main frequency phase augmentation state vector before entering this optimization, and is used as the initial condition of the prediction model.
[0148] The discrete-time prediction model is a discrete-time state-space model based on hierarchical Koopman mapping, used to predict the evolution of the state with control input under sampling steps, and can be in linear or piecewise linear approximation form.
[0149] The prediction time domain is the time interval used for model prediction and optimization calculation, which includes two sub-intervals, the short horizon and the long horizon, and their corresponding prediction step sequences.
[0150] The anti-phase feedforward term is a feedforward control component constructed based on the dominant frequency parameter and phase parameter, which is opposite in phase (approximately π) to the main vibration component. Its amplitude and phase are updated over time to cancel the main modal response, and it is not limited to a single sinusoidal form.
[0151] The equivalent input is a synthetic input obtained by adding the anti-phase feedforward term and the control command to be optimized in the prediction model, which is used for state rolling prediction.
[0152] The amplitude limit is an upper and lower bound constraint imposed on the absolute amplitude of the inverse phase feedforward term or control input to avoid drive saturation or overshoot.
[0153] The rate of change limit is an upper limit constraint imposed on the magnitude of change of the anti-phase feedforward term or control input between adjacent prediction steps, in order to limit jerk and excitation abrupt changes.
[0154] The learning-based model predictive control is a control method that, within the framework of model predictive control, utilizes online learning or historical data to update the predictive model, weights, or constraint parameters to improve closed-loop performance.
[0155] The optimization objective of the learning model predictive control is an objective function that is a weighted sum of multiple performance items in the prediction time domain according to the cost weight vector, which combines the control objectives of the short horizon and the long horizon.
[0156] The dominant frequency band energy suppression term is a target term that penalizes the energy metric of the flexible modal subspace state in the frequency band corresponding to the dominant frequency within the short horizon, and is used to suppress the dominant mode vibration.
[0157] The control command change rate penalty term is a penalty term for the difference in control input between adjacent prediction steps, used to smooth control and reduce execution shock;
[0158] The tilt angle error term is a target term that penalizes the deviation of the predicted tilt angle from the target tilt angle set by the leveling target parameters;
[0159] The energy consumption term is a target term that penalizes the measurement of the energy or power consumption of the actuator, and is often represented by a quadratic form of the control input or its weighted form.
[0160] The control barrier function is a function that transforms the safety constraint into a non-negative safety function and maintains its convergence condition within the prediction step, thereby ensuring the forward invariance and safety of the feasible region.
[0161] The angle limit constraint is the upper and lower limit boundary constraint applied to the system attitude angle by controlling the obstacle function;
[0162] The angular velocity limit constraint is an upper and lower limit boundary constraint applied to the system's angular velocity by controlling the obstacle function;
[0163] The travel limit constraint is an upper and lower limit boundary constraint applied to the displacement (travel) of the actuator by controlling the obstacle function;
[0164] The cable or structural strain safety boundary constraint is a safety threshold constraint applied to cable tension, structural strain or their equivalent indices by controlling the obstacle function.
[0165] The vibration energy envelope is a time-varying envelope obtained by envelope detection and smoothing of the bandpass signal energy related to the dominant frequency, used to describe the slow variation trend of vibration intensity.
[0166] The upper bound of the vibration energy envelope and the monotonically non-increasing constraint are to force the vibration energy envelope not to exceed the preset or adaptively generated time-decreasing upper bound in the prediction time domain, and to satisfy the constraint of gradually not increasing or shrinking proportionally.
[0167] The safety function is a non-negative function used to describe the feasibility of various safety constraints (angle, angular velocity, travel, strain, energy envelope, etc.), and its non-negativity indicates that the safety boundary is satisfied.
[0168] The predetermined convergence rate is a parameter set within the prediction step for the safety function or the energy envelope decay rate, used to limit its convergence speed or non-increase interval over time.
[0169] The hot-start solution is a process of calling the solver to perform rapid iterations starting from the optimized initial value to obtain a feasible suboptimal or optimal solution, which is used to shorten the online solution time.
[0170] The prediction step is a time index unit obtained by discretizing the sampling interval within the prediction time domain, used for the time sequence of organization status and control.
[0171] Optionally, step S7 specifically includes:
[0172] Under a unified clock, the start time of the current control cycle is selected. Based on the first control instruction in the control instruction sequence, the execution control signal is generated through the actuator interface. The amplitude limit, rate of change limit, and action time limit of the execution control signal are implemented to meet the requirements of travel limit and stabilization target parameters.
[0173] The execution control signal is applied to the vehicle-mounted mast actuator, and the multimodal response after the action is collected. The multimodal response includes at least the base inertial measurement unit data, mast multi-point accelerometer data, mast multi-point strain gauge data, mast top visual displacement data, and actuator status data. The actuator status data includes at least the actuator position, actuator speed, drive current, and supply voltage.
[0174] The calibration and unified coordinate system established by S1 is reused to perform time synchronization, coordinate transformation and basic filtering on the multimodal response, and to perform integrity checks and outlier removal to generate execution response data.
[0175] Output execution response data.
[0176] Terminology definition:
[0177] The vehicle-mounted mast actuator is an actuation device and its drive unit used to drive the mast's posture or extension, including electric, hydraulic or electro-hydraulic servo types;
[0178] The actuator interface is converted into an interface mapping that converts numerical commands of control instructions into drive signals that the actuator can recognize, including the generation of voltage / current / PWM / valve opening or bus protocol commands;
[0179] The execution control signal is the drive input actually applied to the actuator after being converted by the interface, and serves as a control quantity with limited amplitude, direction and duration;
[0180] The action time limit is a setting that constrains the minimum or maximum duration of the execution control signal during a single application process, in order to meet travel and safety requirements;
[0181] The multimodal response is a set of system responses synchronously collected by multiple types of sensors after the control signal is applied, and includes at least inertial, acceleration, strain, vision, and actuator status information;
[0182] The actuator status data is a set of measurement or estimation data characterizing the working status of the actuator, including at least the actuator position, actuator speed, drive current and supply voltage;
[0183] The position of the actuator is a measured value of displacement or angle at the output end of the actuator or associated with the mast, used to reflect changes in stroke or attitude;
[0184] The actuator speed is a measured value of the speed or angular velocity at the output end of the actuator or associated with the mast, used to reflect the dynamic response;
[0185] The drive current is the operating current of the electric drive part of the actuator, which is used to reflect the load and energy consumption status and participate in safety monitoring.
[0186] The supply pressure is the supply pressure of a hydraulic or pneumatic actuator or the working pressure of an electro-hydraulic system, used to reflect the driving capability and safety boundary.
[0187] The time synchronization process involves aligning and interpolating the data from each channel of the multimodal response using a unified clock, so that the data at the same physical moment are consistent on the time axis.
[0188] The coordinate transformation is a process of mapping the original measurement values of each sensor to a unified coordinate system through calibration parameters, so that spatial measurements can be compared under a unified reference.
[0189] The basic filtering is a low-order or band-limited filtering and denoising process applied to the multimodal response, which reduces noise without changing the dominant frequency band characteristics and key phase relationships.
[0190] The integrity check is a process of checking the existence, temporal continuity, scope validity, and format correctness of the collected data.
[0191] The outlier removal process involves identifying, deleting, or replacing detected outliers, saturated segments, or sensor distortion samples.
[0192] The execution response data is a set of response data for feedback and evaluation, formed after time synchronization, coordinate transformation, basic filtering and quality control.
[0193] Optionally, step S8 specifically includes:
[0194] Under a unified clock and coordinate system, a preset time window corresponding to the current control cycle is selected to extract the execution response data;
[0195] The current tilt angle is estimated based on the base inertial measurement unit data and the visual displacement data of the mast top in the execution response data. The target tilt angle is determined with reference to the leveling target parameters. The instantaneous tilt angle error and the mean square tilt angle error within the time window are calculated to form leveling error feedback.
[0196] A bandpass filter is constructed around the dominant frequency parameter to set the frequency band range. The mast multi-point accelerometer data and mast multi-point strain gauge data in the execution response data are bandpass filtered to calculate the dominant frequency band energy estimate. The energy envelope of the bandpass signal is obtained by Hilbert transform. The upper bound deviation of the energy envelope is calculated based on the upper limit of the dominant frequency band energy in the steady-state target parameter. At the same time, the monotonic descent of the energy envelope within the time window is statistically analyzed to form frequency domain energy feedback.
[0197] The steady-state error feedback and the frequency domain energy feedback are combined in a predetermined order to form a closed-loop feedback index vector;
[0198] Output closed-loop feedback index vector.
[0199] Terminology definition:
[0200] The tilt angle is the attitude angle of the mast relative to the reference gravity direction or horizontal plane in a unified coordinate system, and is used to characterize the current attitude of the mast.
[0201] The tilt angle estimation is the current tilt angle value calculated by fusing the base inertial measurement unit data and the visual displacement data of the mast top in a unified coordinate system;
[0202] The target tilt angle is the attitude target or threshold specified by the leveling target parameters, and is used for leveling error calculation and control constraints.
[0203] The instantaneous tilt angle error is the difference between the current tilt angle and the target tilt angle, used to reflect the degree of instantaneous attitude deviation;
[0204] The mean square tilt angle error is an error metric that averages the squares of the instantaneous tilt angle error within a preset time window, and is used for leveling performance evaluation.
[0205] The frequency band range is a narrow bandwidth interval set with the dominant frequency parameter as the center, used to extract the main mode related frequency components;
[0206] The bandpass filter is a filter that approximates the passband within the specified frequency band and attenuates out-of-band frequency components, used to obtain a bandpass signal;
[0207] The bandpass signal is the execution response data that has been processed by a bandpass filter and retains only the components within the specified frequency band.
[0208] The dominant frequency band energy estimate is a measure of the energy of the bandpass signal within a preset time window, and can be expressed as the mean square value or quadratic integral of the bandpass signal.
[0209] The energy envelope is the envelope curve of the instantaneous amplitude sequence obtained by Hilbert transform or equivalent detection of the bandpass signal and then appropriately smoothed, used to describe the slow variation trend of vibration intensity.
[0210] The upper bound of the energy envelope is a time-decreasing upper boundary function constructed based on the upper limit of the dominant frequency band energy in the stable target parameters, which is used to constrain the energy envelope from not exceeding the set upper limit.
[0211] The upper bound deviation of the energy envelope is the amount or difference that the energy envelope exceeds the upper bound, and is used to measure the degree of exceeding the bound.
[0212] The monotonic decay rate is the percentage of off-steps or an equivalent normalized decay index in which the energy envelope does not increase within a preset time window, used to measure the degree of monotonic decay of the envelope.
[0213] The frequency domain energy feedback is a feedback vector composed of the dominant frequency band energy estimate, the upper bound deviation of the energy envelope, and the monotonic descent degree in a predetermined order.
[0214] The leveling error feedback is a feedback vector composed of instantaneous tilt angle error and mean square tilt angle error in a predetermined order;
[0215] The closed-loop feedback index vector is a comprehensive feedback vector formed by splicing the steady-state error feedback and the frequency domain energy feedback in a predetermined order, which is used to drive online updates and scheduling.
[0216] Optionally, step S9 specifically includes:
[0217] Based on the closed-loop feedback index vector extraction of steady-state error feedback and frequency domain energy feedback, an online adaptive loss function is constructed. The loss function is formed by weighting the mean square tilt angle error, the deviation of the dominant frequency band energy estimate relative to the steady-state target parameter, the upper bound deviation of the energy envelope, and the monotonic descent penalty with non-negative weights.
[0218] Online updates are performed on the parameters of the hierarchical Koopman map, which include the one-step prediction matrix of the rigid subspace, the one-step prediction matrix of the flexible modal subspace, and the lifting transformation parameters. The loss function is minimized by recursive least squares with a forgetting factor or small-step gradient descent. A larger forgetting factor is used for the one-step prediction matrix of the flexible modal subspace to track the drift of the dominant frequency parameter, and a smaller learning rate is used for the one-step prediction matrix of the rigid subspace to maintain steady state.
[0219] Projection-based constraints are applied to the updated flexible modal subspace related parameters to satisfy mode orthogonality constraints and energy consistency constraints, and the parameter change rate is limited by spectral radius constraints and regularization to ensure the stability of one-step prediction.
[0220] Online updates are performed on the parameters of the enhanced collaboration module, which include weight scheduling strategy parameters and hot-start generation strategy parameters. The negative of the aforementioned loss function is used as the reward signal, and the weight scheduling strategy parameters are updated and improved using policy gradient or approximate value, so that the generated cost weight vector, short horizon step number parameter and long horizon step number parameter improve the closed-loop performance within their respective preset ranges.
[0221] The parameters of the hot start generation strategy are updated in a supervised manner, so that the control instruction part of the optimized initial value gradually approaches the time-shifted version of the control instruction sequence of the previous control cycle in the internal cache of the module and the rate of change is bounded.
[0222] When the frequency domain energy feedback shows that the upper bound deviation of the energy envelope is positive or the monotonic descent is below the threshold, the weights of the dominant frequency band energy term and the tilt error term in the cost weight vector are increased and the reduction of the short horizon step parameter is prohibited. When both the steady-state error feedback and the frequency domain energy feedback are below the threshold, the learning rate is reduced to suppress overfitting.
[0223] The updated hierarchical Koopman mapping parameters and the updated enhanced co-module parameters are output for use in the next control cycle.
[0224] Terminology definition:
[0225] The online adaptive loss function is a target function dynamically constructed based on the closed-loop feedback index vector. It is at least a weighted sum of mean square tilt error, deviation of the dominant frequency band energy estimate from the stable target parameter, deviation of the upper bound of the energy envelope, and monotonic descent penalty, which is used to drive the online update of the model and policy.
[0226] The non-negative weights are a set of weighting coefficients for each item in the loss function or optimization objective, and each element is not less than zero, in order to ensure the physical interpretability and numerical stability of the metric.
[0227] The one-step prediction matrix is an operator matrix that describes the linear evolution of the state from the current step to the next step within a discrete sampling interval, and is used for single-step state updates in the Koopman lifting space.
[0228] The one-step prediction matrix of the rigid body subspace is a linear evolution operator of the rigid body state vector in discrete time, used to describe the single-step dynamics of attitude and rigid body motion components.
[0229] The one-step prediction matrix of the flexible modal subspace is a linear evolution operator of the flexible state vector in discrete time, used to describe the single-step dynamics of the main mode and higher-order flexible components.
[0230] The lifting transformation parameters are a set of parameters for a dictionary / basis function that maps time series observations to the lifting feature space, including basis function coefficients, selection structure, and normalization parameters, etc.
[0231] The recursive least squares with forgetting factor is a parameter estimation method that minimizes the weighted error online when data arrives. It applies a forgetting factor to historical samples to reduce the weight of old data and enhance the ability to track new data.
[0232] The forgetting factor is a coefficient that controls the decay rate of historical sample weights in recursive least squares. Its value is usually in (0,1], and the smaller the value, the faster the forgetting.
[0233] The small-step gradient descent is a gradient method that iteratively updates parameters online with a small learning rate, in order to improve the stability of the update process and avoid oscillations.
[0234] The drift of the dominant frequency parameter refers to the slow or medium-speed shift of the center frequency of the flexible master mode as it changes with time and operating conditions (height, load, wind load, etc.).
[0235] The learning rate is the step size coefficient for each parameter correction in gradient update, used to balance convergence speed and stability;
[0236] The projection-based constraint processing is an operation that projects the parameter vector onto a feasible set that satisfies the mode orthogonality constraint and the energy consistency constraint after the parameter update, so as to ensure that the physical and numerical constraints are not violated.
[0237] The spectral radius constraint is an upper limit imposed on the spectral radius (absolute value of the largest eigenvalue) of the one-step prediction matrix to ensure the stability of discrete-time evolution. It is usually set to no more than 1 or slightly less than 1.
[0238] The regularization involves adding a penalty term (such as an L2 or L1 term) to the loss function to penalize the magnitude or rate of change of the parameters, in order to suppress overfitting and smooth parameter updates.
[0239] The weighted scheduling strategy parameters are a set of strategy parameters generated from the scheduling input / scheduling index mapping to the cost weight vector and the horizon step number in the enhanced collaboration module;
[0240] The hot start generation strategy parameters are a set of parameters in the enhanced collaboration module that control the logic for generating the initial value, including the time shift ratio of the previous cycle sequence, the truncation length, and the amplitude and speed limiting rules, etc.
[0241] The reward signal is an evaluation signal used for reinforcement learning updates. In this case, the negative of the aforementioned loss function is taken so that a larger reward indicates better closed-loop performance.
[0242] The policy gradient is a method for updating policy parameters by gradient ascent / descent based on the reward signal, used to directly optimize the expected reward;
[0243] The approximate value update is an update method that improves the policy by estimating the value function or advantage function. It can use temporal difference or function approximation to obtain the policy update direction.
[0244] The supervised update is a learning method that uses the target output (such as a time-shifted version of the control instruction sequence from the previous control cycle) as the supervision signal to minimize the difference between the prediction and the target in order to adjust the parameters of the hot-start generation strategy.
[0245] The internal cache of the module is a storage structure within the enhanced collaborative module used to store the control instruction sequence and related states of the previous control cycle, in order to support hot start and continuous updates.
[0246] The online update is a method of updating model parameters and strategy parameters by adjusting them recursively over time without stopping the control loop during system operation, in order to continuously adapt to changes in operating conditions.
[0247] The beneficial effects of this invention are:
[0248] 1. Accurate characterization and rapid leveling: The Koopman mapping of rigid and flexible layers is adopted and mode orthogonality and energy consistency constraints are applied. Combined with the dominant frequency and phase augmentation, as well as the learning model predictive control of short and long horizons and anti-phase feedforward, the directional suppression of the flexible master mode is achieved, the leveling time is shortened, the tilt error and overshoot are reduced, and the dominant frequency band energy is significantly suppressed.
[0249] 2. Enhanced safety and controllable convergence: By controlling the barrier function and simultaneously applying constraints on angle, angular velocity, stroke, structural strain, and the upper bound and monotonic non-increasing energy envelope of the dominant frequency band, the safety boundary is ensured within the prediction domain and the vibration energy decays at a predetermined rate, reducing the risk of limit and structural problems.
[0250] 3. Enhanced Adaptability and Robustness: By leveraging the cost weights of enhanced collaboration, horizon step scheduling, and warm start, combined with hierarchical Koopman algorithm and online policy updates, the closed-loop performance can be adaptively maintained according to mast height, top load, and frequency drift, improving the convergence and computational efficiency of the optimization solution and reducing the workload of parameter tuning. Attached Figure Description
[0251] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0252] Figure 1 This is a flowchart of a vehicle-mounted mast stabilization method based on multimodal adaptive reinforcement proposed in this invention. Detailed Implementation
[0253] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0254] refer to Figure 1 A method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement, characterized by comprising:
[0255] S1. Under a unified clock and coordinate system, collect and calibrate the data of the base inertial measurement unit, the data of the mast multi-point accelerometer, the data of the mast multi-point strain gauge, and the visual displacement data of the mast top. Read the mast extension height parameters and the top load parameters and set the leveling target parameters.
[0256] S2. Based on the above data, perform hierarchical Koopman mapping to enhance the observations in the rigid body subspace and flexible mode subspace. The flexible modes satisfy the mode orthogonality constraint and the energy consistency constraint, thus obtaining the rigid body state vector and the flexible state vector.
[0257] S3. Based on the flexible state vector, estimate the dominant frequency parameters and phase parameters;
[0258] S4. Based on the rigid and flexible state vectors and combined with the dominant frequency parameter and phase parameter, state augmentation is performed. The dominant frequency parameter and the sine and cosine characteristics based on the phase parameter are concatenated with the aforementioned state to form a dominant frequency phase augmented state vector.
[0259] S5. Based on the main frequency phase augmentation state vector, mast extension height parameters and top load parameters, the enhanced coordination module is called to perform weight scheduling and hot start, and the cost weight vector, short horizon step parameters, long horizon step parameters and optimization initial values of model predictive control are obtained.
[0260] S6. Based on the cost weight vector of model predictive control, short horizon step parameters, long horizon step parameters, initial optimization values, and main frequency phase augmentation state vector, construct and solve the dual-horizontal learning model predictive control optimization problem to obtain the control command sequence.
[0261] S7. Apply the first control command of the control command sequence to the vehicle-mounted mast actuator and collect the execution response data after the action;
[0262] S8. Based on the execution response data, the stabilization target parameters and the dominant frequency parameters, calculate the stabilization error feedback and the energy feedback in the frequency band centered on the dominant frequency parameters to form a closed-loop feedback index vector.
[0263] S9. Based on the closed-loop feedback index vector, update the hierarchical Koopman mapping parameters and the enhanced collaborative module parameters for the next control cycle.
[0264] In this specific embodiment, S1 specifically refers to:
[0265] Establish a unified clock for each sensor channel and perform time synchronization correction for the sensor index. The original timestamp uses an identification relationship:
[0266] ;
[0267] in For sensors Synchronized timestamps under a unified clock For time scale correction coefficients, For time offset correction amount, For sensors The original timestamp, It is an index set that includes the base inertial measurement unit channel, the mast multi-point accelerometer, the mast multi-point strain gauge, and the vision channel;
[0268] After completing time-domain alignment, the inertial, acceleration, and strain data are calibrated and their coordinates are unified, using a simplified integrated calibration and transformation method:
[0269] ;
[0270] in For sensors Measurement vectors in a unified coordinate system For the rotation matrix from the sensor coordinate system to the unified reference coordinate system, For the scaling factor matrix, For sensors The original measurement vector, For zero bias vectors, For temperature compensation coefficient vector, For sensors Real-time temperature, For sensors Reference temperature The current time under a unified clock;
[0271] The visual displacement data at the top of the mast is first corrected for lens distortion, and then intrinsic and extrinsic parameter mapping is performed. The distortion correction is denoted as:
[0272] ;
[0273] in For a moment The distortion-free pixel coordinate vector For the inverse mapping operator of the lens distortion model, For a moment The original pixel coordinate vector This is the distortion coefficient vector;
[0274] Then, the camera's intrinsic and extrinsic parameters are used to map it into a top displacement in a unified coordinate system:
[0275] ;
[0276] in The vector of the mast tip displacement in a unified coordinate system. For the rotation matrix from the camera coordinate system to the unified coordinate system, Camera intrinsic parameter matrix The reverse It is a homogeneous vector of the distorted pixel coordinates;
[0277] In the noise suppression and anti-interference filtering stage, to ensure that the relationship between the peak value of the dominant frequency band and the critical phase is not changed, the filter is selected to meet the following requirements:
[0278] ;
[0279] in For the frequency response of the filter, The frequency point where the dominant frequency is located, The phase response at this frequency point, The maximum allowable phase distortion threshold;
[0280] The mast extension height is obtained by linearly calibrating the readings from the travel sensor.
[0281] ;
[0282] in For mast extension height parameters, For travel sensor readings, For height ratio coefficient, This is the height offset.
[0283] The top load parameters are determined by referring to a table based on the equipment identification and verified, and recorded as follows:
[0284] ;
[0285] in For load parameter vector, For the lookup function that maps device identifiers (IDs) to calibration tables, For top load quality, Equivalent moment of inertia of top load, This is the equivalent aerodynamic or wind load coefficient;
[0286] The target parameters for leveling are set according to the work instructions as follows:
[0287] ;
[0288] in To stabilize the target parameter set, For the target tilt angle threshold, As the upper limit of the dominant frequency band energy, The maximum permissible stabilization time;
[0289] After completing the above time synchronization, calibration, coordinate unification and filtering, the integrity and quality of each data are checked to remove unqualified samples. Under the premise of ensuring that the dominant frequency band characteristics are not distorted, reasonable interpolation is performed. Finally, the data of the base inertial measurement unit, mast multi-point accelerometer data, mast multi-point strain gauge data, mast top visual displacement data, mast extension height parameters, top load parameters and leveling target parameters are output under a unified clock and coordinate system.
[0290] In this specific embodiment, S2 specifically refers to:
[0291] Under a unified clock and coordinate system, a preset time window is selected and the data of each channel is aligned according to the sampling step size. The inertial measurement unit of the base, the multi-point acceleration of the mast, the multi-point strain of the mast, and the visual displacement of the mast top are concatenated in chronological order to form an observation vector, which is expressed in a compact form as follows:
[0292] ;
[0293] in Indicates time The time-series observation vector, This represents the observation vector of the base inertial measurement unit in a unified coordinate system. This represents a multi-point acceleration aggregation vector deployed along the mast. This represents a multi-point strain aggregation vector deployed along the mast. This represents the visual displacement vector at the top of the mast. Represents any sampling moment within the time window, symbol Indicates transpose;
[0294] Based on the hierarchical Koopman concept, observations are boosted by generating states in both rigid and flexible modal subspaces, ensuring that the two types of boosting parameters are not shared. The boosting relation is written as follows:
[0295] ;
[0296] in Indicates time The rigid body state vector, Indicates time The flexible state vector, This indicates the observed lifting transformation of the rigid body subspace. This indicates the observed lifting transformation of the flexible subspace;
[0297] To ensure the separability and physical consistency of the flexible modal characterization, two types of constraints—modal orthogonality and energy consistency—are applied to the flexible state within the time window. The modal orthogonality constraint is denoted as:
[0298] ;
[0299] in Indicates the discrete sampling index within the time window. This represents the number of samples within the time window. Indicates the first The inner product weights of each sample, Indicates the flexible state in the 1st... Sample vectors on each modal component Indicates the flexible state in the 1st... Sample vectors on each modal component Indicates the first time window Each sampling time, Represents different modal indices;
[0300] Energy consistency constraint is denoted as:
[0301] ;
[0302] in This represents the energy metric calculated based on flexible states. This represents the equivalent energy measure calculated based on sensor data. Indicates energy consistency tolerance. The semi-positive definite matrix representing the energy weighting of the flexible state. This represents a positive semidefinite matrix weighted by acceleration energy. This represents a positive semidefinite matrix weighted by strain energy;
[0303] Under the aforementioned enhancements and constraints, an optimization problem is constructed with the objective of minimizing the time-series observation reconstruction and short-step prediction errors. By solving this problem, the state sequence within the time window is obtained, and the current time step is extracted. The rigid body state vector and the flexible body state vector are used in subsequent steps, where This indicates the sampling time at the end of the time window corresponding to the current control cycle.
[0304] In this specific embodiment, S3 specifically refers to:
[0305] Under a unified clock and coordinate system, the flexible state is first detrended and its amplitude normalized within a preset time window. Then, the dominant frequency is jointly estimated and the instantaneous phase is calculated. If necessary, multiple components are fused according to energy weights to improve robustness. Among them, the first... The preprocessing of the flexible component can be written as:
[0306] ;
[0307] in Indicates at time Detrending and normalized flexible state vectors Indicates at time The original flexible state vector, The detrending operator within the time window, Indicates the first The amplitude gauge constant of each component, Represents flexible component index, This represents any sampling moment within the time window;
[0308] Subsequently, a weighted sum of the short-time Fourier and autoregressive spectra was used for joint estimation, and the peak value was located within the preset search frequency band. The joint spectrum is denoted as:
[0309] ;
[0310] The initial peak value is:
[0311] ;
[0312] The dominant frequency parameters are then obtained by refining the peak value using parabolic interpolation:
[0313] ;
[0314] in, Indicates the first Components at frequency The joint power spectral density at the location Indicates the first Short-time Fourier spectrum estimation of components Indicates the first Autoregressive spectral estimation of components Represents the weighting coefficients of the two types of spectra, Representing frequency variables, Peak frequency indicating preliminary positioning Indicates the preset search frequency band range, This represents the refined dominant frequency parameters. Indicates the frequency sampling interval;
[0315] around Construct a narrowband bandpass signal and calculate its instantaneous phase using an analytic signal. Let the analytic signal and phase be denoted as:
[0316] ;
[0317] And using the phase at the most recent zero-crossing point as a reference:
[0318] ;
[0319] in Indicates surrounding bandpass signal, Indicates the first Analytical signals of components Represents the imaginary unit, Represents the Hilbert transform operator, Indicates the first Component instantaneous phase, Represents the phase expansion operator, Indicates the argument of a complex signal, Indicates phase reference value, Indicates the first The time of the most recent zero-crossing of the component bandpass signal;
[0320] When multiple flexible components exist, weighted fusion is performed according to bandpass energy within the time window to suppress the influence of anomalous components. The energy, weights, and fusion output are denoted as:
[0321] ;
[0322] in Indicates the first The bandpass energy of the component within the time window, Indicates the first The energy weight of the component Indicates the dominant frequency parameter after fusion. Indicates the instantaneous phase parameters after fusion. Indicates the first time window Each discrete sampling time, Indicates the number of samples within the time window, Indicates the total number of flexible components, The index represents the sum of energy weights; when there is only a single component, it is taken directly. and The dominant frequency and phase parameters are used as the output.
[0323] In this specific embodiment, S4 specifically refers to:
[0324] Under a unified clock, sampling points corresponding to the current moment of a preset time window are selected, and the dimensions of the rigid body state and the flexible state are aligned and their indices are consistent. Then, based on the dominant frequency parameter and the phase parameter, sine and cosine features of the phase are generated and concatenated with the state variables in a predetermined field order to form a fixed-dimensional dominant frequency phase augmented state vector for subsequent prediction and control. The sine and cosine features are written as follows:
[0325] ;
[0326] in Cosine characteristic of phase parameter Sine characteristics representing phase parameters Represents the cosine function, Represents the sine function, Represents the instantaneous phase parameter at the current sampling moment under a unified clock. This indicates the sampling time corresponding to the current control cycle;
[0327] Augmented state vector writing:
[0328] ;
[0329] in Represents the phase augmentation state vector of the main frequency. Indicates at time rigid body state vector, Indicates at time Flexible state vector, Indicates the dominant frequency parameter, symbol [ [] indicates a vector concatenation operation in a predetermined field order, symbol The transpose operation represents a vector or matrix;
[0330] To ensure interpretability and consistency across control cycles, field indexes are kept fixed during splicing, and this order is recorded during the system design phase to ensure that the input dimension of the prediction model remains constant.
[0331] In this specific embodiment, S5 specifically includes:
[0332] Based on the main frequency phase augmentation state and operating parameters, the input quantities for scheduling are first constructed, and cost weights, horizon steps, and initial values for hot start are generated accordingly. The scheduling input vector is compactly represented as follows:
[0333] ;
[0334] in The scheduling input vector at the current sampling time, Indicates the dominant frequency parameter, Cosine characteristic of phase parameter Sine characteristics representing phase parameters Indicates at time instantaneous phase, Indicates at time Flexible state vector, Representing the 2-norm, Indicates the mast extension height parameter. Represents the top load parameter vector and its transpose Used for concatenation with scalars;
[0335] Then, the two core scheduling metrics, namely the dominant frequency band energy and the phase change rate, are calculated in a simplified form:
[0336] ;
[0337] in Indicates the dominant frequency band energy estimation, The positive semidefinite weighting matrix that matches the dominant frequency band. Indicates the rate of phase change, Indicates the instantaneous phase at the previous sampling time. Indicates the sampling step size;
[0338] The weight scheduling submodule of the enhanced collaboration module maps the above inputs and indicators to unnormalized cost-weight vectors and performs normalization to ensure numerical stability, which can be abbreviated as:
[0339] ;
[0340] in Represents the unnormalized cost weight vector, Represents the normalized cost weight vector. Represents the weighted scheduling strategy mapping function, Describe a norm and The components correspond to the tilt error term weight, the dominant frequency band energy term weight, the control increment term weight, and the energy consumption term weight, respectively.
[0341] In the hot-start generation, the time-shift sequence of the previous control cycle is preferentially used, and the current augmented state is used as the initial condition for prediction, expressed as:
[0342] ;
[0343] in Indicates the initial sequence of optimized control instructions, This represents the time shift operator of the control sequence from the previous period. This indicates the sequence of control commands from the previous control cycle. Indicates the prediction of the initial state, Indicates at time The dominant frequency phase augmented state vector;
[0344] In horizon setting, interval constraints are applied to both the short and long horizons to ensure target compatibility. And specify the threshold rule in the text when Larger or Increase when larger And maintain ,when Smaller and Reduce appropriately when smaller Increase within the allowable range ;
[0345] in Indicates the short horizon step number parameter, Indicates the long horizon step count parameter, and Represent the minimum and maximum number of steps for the short horizon, respectively. and Let these represent the minimum and maximum number of steps for the long horizon, respectively.
[0346] The final output cost weight vector Short horizon step count parameter Long horizon step count parameter With optimization initial values .
[0347] In this specific embodiment, S6 specifically refers to:
[0348] Using the dominant frequency phase augmentation state as the initial condition for the current prediction, a discrete-time prediction model and optimization objectives for short and long horizons are established. Simultaneously, an anti-phase feedforward is introduced into the short horizon, and control barrier function constraints are applied across the entire prediction domain to ensure multiple safety boundaries. Hot-start initial values are used to accelerate the online solution, where the initial state is denoted as:
[0349] ;
[0350] in Indicates the initial state of the prediction model, This represents the dominant frequency phase augmented state vector at the current sampling time. This indicates the unified clock downsampling time corresponding to the current control cycle;
[0351] The prediction model adopts a first-order discrete form in the Koopman lifting space:
[0352] ;
[0353] in Indicates the first Augmented state of each prediction step, Represents the one-step prediction matrix of hierarchical Koopman. Represents the input mapping matrix, Indicates equivalent input, Indicates the first The control instructions to be optimized for each prediction step Indicates the inverse phase feedforward term, Indicates the prediction step index, Indicates the short horizon step number parameter, This represents the number of steps for the long horizon.
[0354] Inverse phase feedforward is generated within the short horizon based on the dominant frequency and phase, written as:
[0355] And apply limits ;
[0356] in Indicates the feedforward amplitude coefficient, Represents the unit direction vector consistent with the actuator channel. Represents the cosine function, Indicates the current instantaneous phase, Indicates the dominant angular frequency, Indicates the dominant frequency parameter, Indicates sampling step size, Represents the phase reversal constant, Representing the 2-norm, Indicates the upper bound of the feedforward amplitude, This indicates the upper bound of the rate of change of adjacent feedforward steps;
[0357] The optimization objective combines short-domain vibration suppression and long-domain stability energy consumption, denoted as:
[0358] ;
[0359] in Represents the total cost function, Indicates the weight of the dominant frequency band energy term, Indicates the first Energy measurement of flexible modes in the dominant frequency band Indicates from augmented state The extracted flexible subspace state The positive semidefinite weighting matrix that matches the dominant frequency band. Indicates the weight of the control increment term, Indicates the difference between adjacent control commands and for The text refers to the end-of-cycle instruction from the previous cycle. Indicates the weight of the tilt error term, Indicates the state The tilt angle estimated by geometric mapping, Indicates the target tilt angle in the leveling target parameters, Indicates the weight of the energy consumption item;
[0360] Safety constraints are uniformly applied using a control barrier function, written as... With energy envelope convergence constraint ;
[0361] in Indicates the first Class security functions , rate, stroke, strain, envelope This represents a set of safety constraint indices, corresponding to angular limits, angular velocity limits, travel limits, structural strain boundaries, and energy envelope boundaries, respectively. Indicates the first The upper bound of the allowed energy envelope, This represents the predetermined convergence rate of the energy envelope;
[0362] The solver enters the iteration with a hot-start initial value, denoted as:
[0363] ;
[0364] in Indicates the initial sequence of control commands, Represents the time shift operator, This indicates the sequence of control commands from the previous control cycle;
[0365] Finally, the optimal control command sequence that satisfies the control barrier function and the amplitude and speed limiting requirements is obtained. ,in This represents the control sequence obtained through optimization. Indicates the first The optimal control command for each prediction step.
[0366] In this specific embodiment, S7 specifically refers to:
[0367] Under a unified clock, the start time of the current control cycle is selected, and the first control quantity in the control command sequence is converted into a drive signal recognizable by the actuator through the interface. The first optimized control quantity is initially recorded as... And obtain the applied command through interface mapping:
[0368] ;
[0369] in This represents the first optimal control vector in the current control cycle. Indicates the first The drive signal vector applied to the actuator in each control cycle, Operators that represent the mapping from numerical instructions to actuator interface signals. Indicates the control cycle index;
[0370] Then, the applied command's amplitude and rate of change are limited, and the action time window is constrained to meet travel and safety requirements, using compact constraints:
[0371] ;
[0372] in Representing the 2-norm, Indicates the upper limit of the instruction amplitude, This represents the drive signal vector applied in the previous control cycle. Indicates sampling step size, Indicates the maximum allowable rate of change, Representing continuous time variables, Indicates the start time of the current control cycle. This indicates the upper limit of the duration of the applied control signal;
[0373] After the driving signal is applied, the multimodal response is synchronously acquired and organized into a response vector:
[0374] ;
[0375] in This represents the vector of the base inertial measurement unit data in a unified coordinate system. Represents the aggregated acceleration vector of the mast at multiple points. Represents the multi-point strain aggregation vector of the mast. Represents the visual displacement vector at the top of the mast. Represents the scalar or vector quantity of the actuator position (stroke). Represents the speed scalar or vector of the actuator. Indicates drive current, Indicates pressure supply, symbol This indicates transposition for compact splicing;
[0376] To ensure the consistency and usability of the feedback data, the calibration and unified coordinate system established in step S1 is reused to perform time synchronization, coordinate transformation, basic filtering, and quality control on the above response, denoted in operator form as follows: ,in This indicates the processed execution response data. This indicates a processing operator that includes time alignment, coordinate unification, noise suppression, and anomaly removal;
[0377] The final output is the execution response data vector of the current control cycle. .
[0378] In this specific embodiment, S8 specifically refers to:
[0379] Under a unified clock and coordinate system, a preset time window corresponding to the current control cycle is selected to capture and evaluate the multimodal execution response. First, a set of time window indices is defined:
[0380] ;
[0381] in Represents the set of discrete times in the current period. Indicates the start time of the current control cycle. Indicates sampling step size, This represents the number of samples within the time window, and the starting point of the time window is denoted as . ;
[0382] At the attitude level, tilt angle estimation is obtained by fusing IMU and top visual displacement to form a leveling error metric, which is expressed in a compact form:
[0383] ;
[0384] in Indicates instantaneous tilt angle error, Indicates the current tilt angle estimate. Indicates the target inclination angle, Indicates the mean square tilt error within the time window, Indicates the first time window Each sampling time;
[0385] In the frequency domain, a passband is defined around the dominant frequency, and the energy and envelope are extracted. Let the passband be:
[0386] ;
[0387] in Indicates the dominant frequency band, Indicates the dominant frequency parameter, Indicates bandwidth;
[0388] Let the bandpass signals of acceleration and strain be respectively and And linearly fuse them to obtain a scalar bandpass signal:
[0389] ;
[0390] in This represents the multi-point acceleration vector after bandpass. Represents the multi-point strain vector after bandpass. and This represents the corresponding non-negative weighted vector;
[0391] The dominant frequency band energy is estimated as follows:
[0392] ;
[0393] in The dominant frequency band energy estimate within the time window is obtained by applying a Hilbert transform to the envelope. ,in Indicates energy envelope, Represents the Hilbert transform operator, Represents the imaginary unit;
[0394] To measure the degree of time out-of-bounds access, a time-decreasing upper bound function is constructed:
[0395] Deviation from upper bound ;
[0396] in Indicates the upper bound of the energy envelope, Indicates the upper limit of the dominant frequency band energy, Indicates the attenuation coefficient, Represents the positive part operator, Indicates the upper bound deviation;
[0397] To measure the monotonically convergent nature of the envelope, the degree of monotonically decreasing is defined as follows:
[0398] ;
[0399] in Indicates the proportion of time within which no step increment occurs. Indicates indicator functions, and Indicates adjacent sampling times;
[0400] Finally, the steady-state error feedback and the frequency domain energy feedback are concatenated in sequence to form a closed-loop feedback index vector. ,in This represents the feedback vector used to drive online updates and scheduling in the current control cycle.
[0401] In this specific embodiment, S9 specifically refers to:
[0402] The prediction model and scheduling strategy are updated online using execution instructions and closed-loop feedback. Safety and convergence are ensured through gating, trust regions, and small step size constraints. First, the gating variables are constructed based on the leveling error, band energy, and envelope index.
[0403] ;
[0404] in A binary gated variable indicating whether parameter updates are allowed in the current period. Indicates indicator functions, Indicates the upper bound deviation of the envelope, Indicates the envelope bias threshold, Indicates the monotonically decreasing degree of the envelope. Indicates single scheduling threshold, Indicates the mean square error of the inclination angle within the time window. This represents the mean square error of the tilt angle corresponding to the previous cycle;
[0405] Then, the first-order prediction error is used as the main loss construct for model updates, written as:
[0406] ;
[0407] in Represents the first-order state prediction error vector. The augmented state vector representing the current period, Represents the augmented state vector at the next sampling time. This represents the estimation of the predicted state transition matrix for the current period. This represents the estimation of the input mapping matrix for the current period. This represents the equivalent input vector after interface mapping and feedforward merging;
[0408] To balance performance and security, the prediction error and the feedback target deviation are combined into an online learning cost, which is simplified as follows:
[0409] ;
[0410] in The online learning cost function for the current period. Indicates the weight of the prediction error term, Representing the 2-norm, Indicates the weight of the feedback target item, This represents the feedback index vector obtained in step S8. This represents a reference index vector set based on the stability target;
[0411] Under the trust region constraint, the learnable parameters are updated with gated gradient steps and projected back to the feasible region, which can be compactly expressed as:
[0412] ;
[0413] in This represents the parameter vector for the current period and includes the prediction model parameters and the weight scheduling strategy parameters. Represents the updated parameter vector, Indicates learning rate, Represents the gradient with respect to the parameter vector. The projection operator to the feasible set of the trust region, This indicates that the radius is centered on the current parameter. Trust domain The radius of the trust domain is represented and adaptively scheduled based on historical convergence and security.
[0414] In implementation, gating variables control whether to perform synchronous updates of the model and scheduler, and an empirical buffer maintains samples from the most recent multiple periods for mini-batch gradient estimation. Simultaneously, numerical and physical consistency checks are performed on the update results; if these checks fail, a rollback is initiated. and reduce Through the above steps, the updated prediction model parameters and scheduling strategy parameters for the next control cycle are finally obtained and fed back into steps S5 and S6 for use.
[0415] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0416] This project integrates multimodal sensing, hierarchical Koopman representation, dominant frequency and phase estimation and augmentation, enhanced cooperative scheduling, dual-horizontal learning model predictive control, and control barrier function safety constraints into a closed-loop control chain. Hierarchical Koopman decouples rigid body motion from flexible modes and provides predictable states. Dominant frequency and phase augmentation explicitly injects vibration energy and phase information into the control variables. The short horizon, under anti-phase feedforward, directionally suppresses the dominant frequency band energy, while the long horizon balances tilt stabilization and energy dissipation targets. The control barrier function simultaneously constrains angle, angular velocity, stroke, structural strain, and energy envelope within the prediction domain, forcing the safety boundary to remain unbroken. Enhanced cooperative scheduling adaptively adjusts weights and steps based on energy and phase change rates and provides hot start. Online updates enable the model and strategy to continuously track frequency drift and load changes, thereby achieving rapid stabilization, low overshoot, energy decay at a predetermined rate, and controllable structural safety.
[0417] In terms of algorithm structure, this study makes targeted improvements to address the characteristics of whip-like vibrations of long masts: It employs a layered Koopman mapping with non-shared parameters between the rigid and flexible subspaces, and adds mode orthogonal constraints and energy consistency constraints to improve the identifiability and energy matching of flexible modes; it designs augmented states for dominant frequencies and phases, making the instantaneous phase a direct reference for control, and achieves phase cancellation through anti-phase feedforward; it proposes a collaborative optimization framework for short and long horizons, combined with adaptive scheduling of cost weights and hot-start optimization initial values, prioritizing the densification of short horizons to accelerate decay under high energy or rapid phase changes, while maintaining the global guidance of long horizons on tilt angle and energy consumption; it introduces an upper bound on the energy envelope and a monotonically non-increasing condition into the control barrier function, and implements projection, spectral radius, and amplitude and velocity limiting constraints on parameters to enhance feasibility and stability. These structural improvements enable the algorithm to suppress vibrations and stabilize the mast more quickly under multiple heights and loads, while reducing solution time and improving robustness.
Claims
1. A method for stabilizing and leveling a vehicle-mounted mast based on multimodal adaptive reinforcement, characterized in that, include: S1. Under a unified clock and coordinate system, collect and calibrate the data of the base inertial measurement unit, the data of the mast multi-point accelerometer, the data of the mast multi-point strain gauge, and the visual displacement data of the mast top. Read the mast extension height parameters and the top load parameters and set the leveling target parameters. S2. Based on the above data, perform hierarchical Koopman mapping to enhance the observations in the rigid body subspace and flexible mode subspace. The flexible modes satisfy the mode orthogonality constraint and the energy consistency constraint, thus obtaining the rigid body state vector and the flexible state vector. S3. Based on the flexible state vector, estimate the dominant frequency parameters and phase parameters; S4. Based on the rigid and flexible state vectors and combined with the dominant frequency parameter and phase parameter, state augmentation is performed. The dominant frequency parameter and the sine and cosine characteristics based on the phase parameter are concatenated with the aforementioned state to form a dominant frequency phase augmented state vector. S5. Based on the main frequency phase augmentation state vector, mast extension height parameters and top load parameters, the enhanced coordination module is called to perform weight scheduling and hot start, and the cost weight vector, short horizon step parameters, long horizon step parameters and optimization initial values of model predictive control are obtained. S6. Based on the cost weight vector of model predictive control, short horizon step parameters, long horizon step parameters, initial optimization values, and main frequency phase augmentation state vector, construct and solve the dual-horizontal learning model predictive control optimization problem to obtain the control command sequence. S7. Apply the first control command of the control command sequence to the vehicle-mounted mast actuator and collect the execution response data after the action; S8. Based on the execution response data, the stabilization target parameters and the dominant frequency parameters, calculate the stabilization error feedback and the energy feedback in the frequency band centered on the dominant frequency parameters to form a closed-loop feedback index vector. S9. Based on the closed-loop feedback index vector, update the hierarchical Koopman mapping parameters and the enhanced collaborative module parameters for the next control cycle.
2. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S1 specifically refers to: Establish a unified system clock and add timestamps to the data from the base inertial measurement unit, the mast multi-point accelerometer, the mast multi-point strain gauge, and the visual displacement data at the top of the mast. Perform time synchronization correction to align all data under a unified time reference. Zero-bias calibration, scaling factor calibration, and temperature compensation are performed on the data from the base inertial measurement unit, the mast multi-point accelerometer, and the mast multi-point strain gauge, and the above data are expressed in a unified coordinate system; Lens distortion correction and intrinsic / extrinsic parameter calibration are performed on the visual displacement data at the top of the mast, and the displacement is expressed in a unified coordinate system; Noise suppression and anti-interference filtering are applied to each data point without altering the peak value of the dominant frequency band. Read and calibrate the mast extension height parameters from the travel sensor; Determine and verify the top load parameters based on the equipment identification and calibration table; Set the leveling target parameters according to the operation instructions. The leveling target parameters include at least the target tilt angle threshold, the upper limit of the dominant frequency band energy, and the maximum allowable leveling time. The data and parameters processed above are subjected to integrity and quality checks, and unqualified samples are removed and appropriate interpolation is performed. Output base inertial measurement unit data, mast multi-point accelerometer data, mast multi-point strain gauge data, mast top visual displacement data, mast extension height parameters, top load parameters, and leveling target parameters.
3. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S2 specifically refers to: Within a preset time window and sampling step size, the data from the base inertial measurement unit, the mast multi-point accelerometer, the mast multi-point strain gauge, and the visual displacement data at the top of the mast are spliced together in chronological order to form a time-series observation vector. The time-series observation vector is lifted by performing a lifting transformation based on the hierarchical Koopman mapping, generating corresponding lifting features in the rigid body subspace and the flexible mode subspace respectively, while keeping the parameters of the two types of features from being shared. An optimization problem with modal orthogonal constraints and energy consistency constraints is constructed. The modal orthogonal constraints limit the pairwise inner product of the flexible state vector within the time window to zero. The energy consistency constraints limit the energy of the flexible state vector to be consistent with the energy calculated based on the mast multi-point accelerometer data and mast multi-point strain gauge data within the allowable error range. The rigid body state vector and the flexible state vector are obtained by solving the optimization problem. Output the rigid body state vector and the flexible body state vector.
4. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S3 specifically refers to: Within a preset time window, the flexible state vector is de-trended and amplitude normalized to obtain the processed flexible state vector. The power spectral density is calculated based on the processed flexible state vector. The energy peak is located in the preset search frequency band by the joint estimation of short-time Fourier transform and autoregressive spectrum. The peak position is refined by parabolic interpolation of nearby frequency points to determine the dominant frequency parameter. A narrowband demodulator is constructed around the dominant frequency parameter to perform bandpass filtering and analytical signal construction on the processed flexible state vector. The instantaneous phase is obtained using Hilbert transform, phase expansion is performed, and the phase parameter is obtained using the most recent zero crossover as the phase reference. When the flexible state vector contains multiple components, the power spectrum and phase are weighted and fused according to the energy weight of each component within the preset time window to suppress the influence of anomalous components on the dominant frequency parameters and phase parameters. Output the dominant frequency and phase parameters.
5. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S4 specifically refers to: Under a unified clock, sampling points corresponding to the current moment of a preset time window are selected, and the rigid body state vector and flexible state vector are aligned in dimension and consistent in index. Based on the dominant frequency parameter and the phase parameter, a sine and cosine feature based on the phase parameter is generated. The sine and cosine feature includes at least the cosine value and the sine value of the phase parameter. The rigid body state vector, flexible state vector, dominant frequency parameter, and sine and cosine features based on phase parameter are spliced together in a predetermined order to form a dominant frequency phase augmented state vector with fixed dimensions and consistent in each control cycle. Output the phase augmented state vector of the main frequency.
6. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S5 specifically refers to: Based on the dominant frequency phase augmented state vector and with reference to the mast extension height parameter and the top load parameter, an input quantity for scheduling is constructed. The input quantity includes the dominant frequency parameter, the sine and cosine characteristics of the phase parameter, the amplitude characteristics of the flexible mode component, the mast extension height parameter, and the top load parameter. Based on the input, the estimated value of the frequency band energy and the phase change rate corresponding to the dominant frequency parameter are calculated and used as scheduling indicators for weight and horizon step count; The weight scheduling submodule of the enhanced coordination module is invoked to generate a cost weight vector for model predictive control based on the scheduling index and the input quantity. The cost weight vector includes at least the tilt error term weight, the dominant frequency band energy term weight, the control increment term weight, and the energy consumption term weight. The upper and lower bound constraints and normalization processing are performed on the cost weight vector to ensure numerical stability. The hot start submodule of the enhanced coordination module is called. The candidate sequence after time shifting the control instruction sequence of the previous control cycle is used as the control instruction part of the initial value of optimization, and the rolling prediction of the main frequency phase augmented state vector is used as the state part of the initial value of optimization. When the data of the previous control cycle is missing, the combination of the static sequence and the current state is used as the initial value of optimization. Based on the threshold rules of the dominant frequency band energy estimate and the phase change rate, the short horizon step parameter and the long horizon step parameter are determined. When the dominant frequency band energy estimate reaches its corresponding preset threshold or the phase change rate reaches its corresponding preset threshold, the short horizon step parameter is increased while the long horizon step parameter is kept not lower than the preset lower limit. When the dominant frequency band energy estimate is lower than its corresponding preset threshold and the phase change rate is lower than its corresponding preset threshold, the short horizon step parameter is decreased while the long horizon step parameter is increased within the allowable range of the stabilization target parameter. Output the cost weight vector, short horizon step number parameter, long horizon step number parameter, and initial optimization value.
7. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S6 specifically refers to: The current prediction initial state is determined based on the dominant frequency phase augmented state vector, and a discrete-time prediction model and prediction time domain are established using hierarchical Koopman mapping based on the cost weight vector, short horizon step parameters, long horizon step parameters and optimization initial values. Within the short horizon, an anti-phase feedforward term is generated based on the dominant frequency parameter and phase parameter. The sum of the anti-phase feedforward term and the control command to be optimized is used as the equivalent input of the prediction model. Limits are applied to the amplitude and rate of change of the anti-phase feedforward term to avoid excitation overshoot. The optimization objective of the learning model predictive control is constructed by superimposing the dominant frequency band energy suppression term and the control command change rate penalty term in the short horizon and the tilt error term and energy consumption term in the long horizon using a cost weight vector. The tilt error term is used to determine the target tilt angle with reference to the stable target parameter, and the dominant frequency band energy suppression term is calculated based on the energy of the flexible mode subspace state component in the frequency band corresponding to the dominant frequency parameter. By controlling the obstacle function, angle limit constraints, angular velocity limit constraints, travel limit constraints, cable or structural strain safety boundary constraints, and vibration energy envelope upper bound and monotonically decreasing constraints are gradually applied within the prediction step, so that the corresponding safety function remains non-negative in the prediction time domain and does not increase at a predetermined convergence rate. Under the above objectives and constraints, a hot start solution is performed using optimized initial values to obtain a control command sequence that meets the requirements of the control barrier function; Output control command sequence.
8. The method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S7 specifically refers to: Under a unified clock, the start time of the current control cycle is selected. Based on the first control instruction in the control instruction sequence, the execution control signal is generated through the actuator interface. The amplitude limit, rate of change limit, and action time limit of the execution control signal are implemented to meet the requirements of travel limit and stabilization target parameters. The execution control signal is applied to the vehicle-mounted mast actuator, and the multimodal response after the action is collected. The multimodal response includes at least the base inertial measurement unit data, mast multi-point accelerometer data, mast multi-point strain gauge data, mast top visual displacement data, and actuator status data. The actuator status data includes at least the actuator position, actuator speed, drive current, and supply voltage. The calibration and unified coordinate system established by S1 is reused to perform time synchronization, coordinate transformation and basic filtering on the multimodal response, and to perform integrity checks and outlier removal to generate execution response data. Output execution response data.
9. A method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S8 specifically refers to: Under a unified clock and coordinate system, a preset time window corresponding to the current control cycle is selected to extract the execution response data; The current tilt angle is estimated based on the base inertial measurement unit data and the visual displacement data of the mast top in the execution response data. The target tilt angle is determined with reference to the leveling target parameters. The instantaneous tilt angle error and the mean square tilt angle error within the time window are calculated to form leveling error feedback. A bandpass filter is constructed around the dominant frequency parameter to set the frequency band range. The mast multi-point accelerometer data and mast multi-point strain gauge data in the execution response data are bandpass filtered to calculate the dominant frequency band energy estimate. The energy envelope of the bandpass signal is obtained by Hilbert transform. The upper bound deviation of the energy envelope is calculated based on the upper limit of the dominant frequency band energy in the steady-state target parameter. At the same time, the monotonic descent of the energy envelope within the time window is statistically analyzed to form frequency domain energy feedback. The steady-state error feedback and the frequency domain energy feedback are combined in a predetermined order to form a closed-loop feedback index vector; Output closed-loop feedback index vector.
10. A method for stabilizing a vehicle-mounted mast based on multimodal adaptive reinforcement according to claim 1, characterized in that, S9 specifically refers to: Based on the closed-loop feedback index vector extraction of steady-state error feedback and frequency domain energy feedback, an online adaptive loss function is constructed. The loss function is formed by weighting the mean square tilt angle error, the deviation of the dominant frequency band energy estimate relative to the steady-state target parameter, the upper bound deviation of the energy envelope, and the monotonic descent penalty with non-negative weights. Online updates are performed on the parameters of the hierarchical Koopman map, which include a one-step prediction matrix for the rigid subspace, a one-step prediction matrix for the flexible modal subspace, and lifting transformation parameters. The loss function is minimized by recursive least squares with a forgetting factor or small-step gradient descent. Specifically, the recursive least squares with a forgetting factor is used for the one-step prediction matrix of the flexible modal subspace to track the drift of the dominant frequency parameter, and the small-step gradient descent is used for the one-step prediction matrix of the rigid subspace to maintain a steady state. Projection-based constraints are applied to the updated flexible modal subspace related parameters to satisfy mode orthogonality constraints and energy consistency constraints, and the parameter change rate is limited by spectral radius constraints and regularization to ensure the stability of one-step prediction. Online updates are performed on the parameters of the enhanced collaboration module, which include weight scheduling strategy parameters and hot-start generation strategy parameters. The negative of the aforementioned loss function is used as the reward signal, and the weight scheduling strategy parameters are updated and improved using policy gradient or approximate value, so that the generated cost weight vector, short horizon step number parameter and long horizon step number parameter improve the closed-loop performance within their respective preset ranges. The parameters of the hot start generation strategy are updated in a supervised manner, so that the control instruction part of the optimized initial value gradually approaches the time-shifted version of the control instruction sequence of the previous control cycle in the internal cache of the module and the rate of change is bounded. When the frequency domain energy feedback shows that the upper bound deviation of the energy envelope is positive or the monotonic descent is below the threshold, the weights of the dominant frequency band energy term and the tilt error term in the cost weight vector are increased and the reduction of the short horizon step parameter is prohibited. When both the steady-state error feedback and the frequency domain energy feedback are below the threshold, the learning rate is reduced to suppress overfitting. The updated hierarchical Koopman mapping parameters and the updated enhanced co-module parameters are output for use in the next control cycle.
Citation Information
Patent Citations
Axis-invariant-based inverse kinematics modeling and solving method for multi-axis robot
EP3838501A1
Improvements in or realating to apparatus and method for copying information onto disk media
GB9321943D0