A method for capturing and locking marine equipment based on reinforcement decision-making algorithms
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-14
AI Technical Summary
[0008]本发明的一个目的在于提出一种基于强化决策算法的海上装备捕获锁紧方法,针对现有技术中捕获锁紧过程包含自由空间靠近、导向接触、半约束滑移、完全约束锁紧等多阶段且阶段切换条件复杂,基于阈值规则的状态机在边界工况下易误判进而导致卡滞、冲击或反复解锁重试,以及控制目标权重和切换参数难以随海况与接触不确定性自适应调整的问题,提出了结合门控多模型动力学预测、事件风险预测、强化学习阶段候选与目标函数权重输出、以及带迟滞与最小驻留时间的切换保护的技术方案;其中在上一控制周期确认阶段模式约束下由门控多模型动力学预测模型输出动力学子模型权重、状态预测量与卡滞风险指标、反弹风险指标和锁紧可达性指标,由强化学习决策模型输出候选阶段模式及滚动优化控制目标函数权重参数,经切换保护模块确认阶段模式后由滚动优化控制器生成位姿调整控制指令与锁紧执行机构控制指令
[0045]1、通过门控多模型动力学预测模型对自由空间靠近、导向接触、半约束滑移以及完全约束锁紧等不同阶段的动力学进行分段建模,并输出动力学子模型权重及状态预测量,使控制器能够在接触不确定性和海况扰动下获得更符合当前阶段的状态演化预测,提高阶段推进与控制计算的准确性和鲁棒性;
Smart Images

Figure CN122569063A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology for marine equipment, and in particular to a method for capturing and locking marine equipment based on a reinforced decision-making algorithm. Background Technology
[0002] Marine equipment capture and locking technology is widely used in docking, recovery, and securing operations between ships and target equipment, such as the capture and locking of floating targets by marine recovery devices and the automatic insertion and locking of docking mechanisms on work platforms. Existing technologies typically acquire information such as relative posture, relative velocity, contact force, or contact torque through position and attitude measurement devices, force sensors, and actuator state feedback units, and employ position control, force control, or a hybrid position-force control to achieve process control such as approach, guided contact, sliding alignment, and locking. To achieve multi-stage process management, engineering commonly employs state machines based on thresholds and logical rules for stage identification and switching, and some solutions introduce rolling optimization control to improve tracking performance and constraint satisfaction capabilities.
[0003] Existing technologies still have shortcomings under actual sea state disturbances, contact uncertainties, and boundary conditions, mainly including:
[0004] 1. The phase switching relies heavily on fixed thresholds and rule logic, which can easily lead to misjudgment when faced with sensor noise, wave excitation and sudden changes in contact state, resulting in jamming, impact or repeated unlocking and retrying.
[0005] 2. Lack of predictive ability for contact process dynamics and events makes it difficult to assess jamming risk, rebound risk and locking accessibility in a timely manner, resulting in insufficient robustness of control strategies under critical conditions;
[0006] 3. The weights of the control objective function and the switching protection parameters are usually preset constants, which are difficult to adaptively adjust with changes in sea state, target model and contact parameters. This can easily lead to problems of overly conservative or overly aggressive control, affecting the success rate and safety of locking.
[0007] Therefore, a method for capturing and locking marine equipment that can overcome the shortcomings of the existing technology is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0008] One objective of this invention is to propose a method for capturing and locking marine equipment based on a reinforcement decision algorithm. Addressing the problems in existing technologies where the capture and locking process involves multiple stages such as free-space approach, guided contact, semi-constrained sliding, and fully constrained locking, with complex stage switching conditions, and where threshold-based state machines are prone to misjudgment under boundary conditions leading to jamming, impact, or repeated unlocking retries, as well as the difficulty in adaptively adjusting control target weights and switching parameters with sea conditions and contact uncertainties, this invention proposes a technical solution combining gated multi-model dynamics prediction, event risk prediction, reinforcement learning stage candidates and objective function weight output, and switching protection with hysteresis and minimum dwell time. Specifically, under the constraint of the stage mode confirmation in the previous control cycle, the gated multi-model dynamics prediction model outputs the dynamic sub-model weights, state prediction quantities, jamming risk indicators, rebound risk indicators, and locking reachability indicators. The reinforcement learning decision model outputs candidate stage modes and rolling optimization control objective function weight parameters. After the switching protection module confirms the stage mode, the rolling optimization controller generates pose adjustment control commands and locking actuator control commands. This invention has the technical effects of reducing the risk of accidental switching and jamming, improving the stability and locking success rate of multi-stage propulsion, and enhancing the safety of the capture and locking process.
[0009] This invention provides a method for capturing and locking marine equipment based on a reinforcement decision algorithm, comprising:
[0010] S1. Collect and capture the interaction state data between the locking device and the target equipment. The interaction state data includes relative pose information, relative motion information, contact force and / or contact torque information, and the state of the locking actuator. S2. Preprocess the interaction state data to calculate the current state features characterizing the current interaction state. S3. Under the confirmation phase mode constraints of the previous control cycle, input the current state features into the gated multi-model dynamics prediction model for prediction, obtaining the dynamics sub-model weights, state predictions, and event risk indicators. S4. Based on the current state features, dynamics sub-model weights, state predictions, and event risk indicators, the reinforcement learning decision model outputs candidate phases. The process involves several steps: S5, inputting candidate stage modes and event risk indicators into the switching protection module for switching release determination, and suppressing stage mode switching based on preset hysteresis thresholds and preset minimum dwell time to determine the confirmation stage mode and corresponding confirmation objective function weight parameters for this control cycle; S6, selecting the dynamic sub-model corresponding to the confirmation stage mode as the rolling prediction model based on the current state characteristics, confirmation stage mode, and confirmation objective function weight parameters, and generating control commands for the capture and locking device by the rolling optimization controller to complete the stage advancement and stage mode switching of the capture and locking process.
[0011] Optionally, S1 includes:
[0012] During each control cycle, the position and attitude measurement device acquires the relative position, relative attitude, and relative velocity of the capture and locking device relative to the target equipment.
[0013] The contact force or contact torque of the capture and locking device is obtained by a force sensor installed on the capture and locking device;
[0014] The state of the locking actuator is obtained by the stroke sensor or state feedback unit of the locking actuator, and the state of the locking actuator includes the position state, speed state and drive state of the locking actuator.
[0015] Optionally, S2 includes:
[0016] The relative position, relative attitude, relative speed, contact force or contact torque, and the state of the locking actuator are synchronized in time so that the relative position, relative attitude, relative speed, contact force or contact torque, and the state of the locking actuator correspond to the same control cycle.
[0017] The data that has completed time synchronization processing is subjected to coordinate unification processing, and the relative position, relative attitude, relative velocity, and contact force or contact torque are transformed to a preset unified coordinate system;
[0018] The data that has undergone coordinate unification processing is subjected to time-series filtering to obtain the filtered relative position, filtered relative attitude, filtered relative velocity, filtered contact force or contact torque, and filtered locking actuator status.
[0019] The current state characteristics are calculated from the filtered relative position, filtered relative attitude, filtered relative speed, filtered contact force or contact torque, and filtered locking actuator state. The current state characteristics include centering deviation and contact state indication.
[0020] Optionally, S3 includes:
[0021] Under the mode constraints of the confirmation phase of the previous control cycle, the current state characteristics are fed into the ingress control multi-model dynamic prediction model;
[0022] The gating function of the gated multi-model dynamics prediction model outputs the weights of the dynamic sub-models corresponding to the free space approach mode, guided contact mode, semi-constrained sliding mode and fully constrained locking mode respectively according to the current state characteristics, and makes the weights of the dynamic sub-models satisfy the normalization constraint.
[0023] Each dynamic sub-model outputs its own next-time state prediction based on the current state characteristics.
[0024] The next time step state prediction is obtained by weighting and fusing the predicted states of each next time step according to the weights of the dynamic sub-models.
[0025] The event prediction module of the gated multi-model dynamics prediction model outputs the event risk index based on the current state characteristics and the next moment state prediction. The event risk index includes a jamming risk index, a rebound risk index, and a locking reachability index.
[0026] Optionally, S4 includes:
[0027] The current state features are combined with the weights of the dynamic sub-model, the predicted state at the next moment, and the event risk index to form decision features, and the decision features are fed into the reinforcement learning decision model for inference calculation.
[0028] The reinforcement learning decision model outputs candidate stage patterns within the stage pattern set and outputs objective function weight parameters for rolling optimization control, wherein the objective function weight parameters include centering deviation term weight, contact force or contact torque term weight, and relative velocity term weight.
[0029] Optionally, S5 includes:
[0030] The candidate phase pattern and event risk indicators are input into the handover protection module, and the handover discrimination model of the handover protection module generates a handover permission flag.
[0031] The switching permission flag is determined according to the following conditions: when the candidate stage mode is different from the confirmation stage mode of the previous control cycle, the stage dwell time count value reaches the preset minimum dwell time, and the event risk indicators meet the following conditions: the jamming risk indicator is not greater than the jamming switching threshold, the rebound risk indicator is not greater than the rebound switching threshold, and the locking reachability indicator is not less than the locking switching threshold.
[0032] Meanwhile, the switching of phase modes is suppressed according to the preset hysteresis threshold. Specifically, when the candidate phase mode meets the locking reachability index not less than the locking entry threshold, it is allowed to switch from the semi-constrained sliding mode to the fully constrained locking mode. When the candidate phase mode meets the locking reachability index not greater than the locking exit threshold, it is allowed to switch from the fully constrained locking mode to the semi-constrained sliding mode. The locking entry threshold is greater than the locking exit threshold.
[0033] When the allow flag is switched to allow, the confirmation phase mode of this control cycle is updated to the candidate phase mode, and the objective function weight parameters are determined as the confirmation objective function weight parameters of this control cycle.
[0034] When the allow flag is switched to disallow, the confirmation phase mode of the current control cycle is maintained as the confirmation phase mode of the previous control cycle, and the confirmation objective function weight parameters of the previous control cycle are determined as the confirmation objective function weight parameters of the current control cycle.
[0035] The stage dwell time count is updated based on whether the confirmation stage mode of the current control cycle has changed compared to the confirmation stage mode of the previous control cycle. Specifically, the stage dwell time count is incremented when there is no change, and the stage dwell time count is cleared when there is a change.
[0036] Optionally, S6 includes:
[0037] The rolling optimization controller reads the current state features, the confirmation phase mode, and the weight parameters of the confirmation objective function, and determines the dynamic sub-model corresponding to the confirmation phase mode as the rolling prediction model.
[0038] In the rolling time domain, the state evolution of the capture and locking device is predicted based on the rolling prediction model, an optimization objective function is constructed based on the weight parameters of the confirmation objective function, and the control sequence is optimized and solved under the control constraints corresponding to the confirmation stage mode.
[0039] The first control quantity of the control sequence obtained by optimization is determined as the control command, which includes a pose adjustment control command and a locking actuator control command;
[0040] The control command is then sent to the capture and locking device for execution, so as to drive the capture and locking device to complete the phase advancement and phase mode switching of the capture and locking process according to the confirmation phase mode.
[0041] Optionally, in S3, the kinetic sub-model includes identifiable parameters related to contact stiffness, contact damping, and / or friction, and the identifiable parameters are updated online based on the current state characteristics and the contact force and / or contact torque information to correct the predictions of the kinetic sub-model.
[0042] Optionally, the preset minimum dwell time and / or jamming switching threshold, rebound switching threshold, locking switching threshold, locking entry threshold, and locking exit threshold are selected from multiple preset parameter groups, which are indexed according to sea state level, target equipment model, and / or the prediction uncertainty index.
[0043] Optionally, before determining the first control quantity of the control sequence obtained by optimization as the control command, the first control quantity is further subjected to safety projection or safety filtering to ensure that the control command meets the preset safety boundary constraints of contact force, contact torque and relative velocity.
[0044] The beneficial effects of this invention are:
[0045] 1. By using a gated multi-model dynamic prediction model, the dynamics of different stages such as free space approach, guided contact, semi-constrained slip and fully constrained locking are segmented and modeled, and the weights and state predictions of the dynamic sub-models are output. This enables the controller to obtain state evolution predictions that are more consistent with the current stage under contact uncertainty and sea state disturbance, thereby improving the accuracy and robustness of stage propulsion and control calculations.
[0046] 2. The event prediction module outputs jamming risk indicators, rebound risk indicators, and locking reachability indicators, and inputs them together with the candidate stage mode into the switching protection module. Combined with the preset hysteresis threshold and minimum dwell time, the switching release is determined, which effectively suppresses erroneous switching and jitter switching under boundary conditions, reduces the probability of jamming, impact and repeated unlocking retries, and improves the stability and safety of the locking process.
[0047] 3. The reinforcement learning decision model outputs candidate stage modes and objective function weight parameters of rolling optimization control within the stage mode set, realizing adaptive adjustment of stage modes and control objective weights, and linking with the rolling optimization controller to generate pose adjustment control commands and locking actuator control commands, thereby improving capture and locking efficiency and locking success rate while ensuring safety constraints. Attached Figure Description
[0048] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0049] Figure 1 This is a flowchart of a method for capturing and locking marine equipment based on a reinforcement decision algorithm proposed in this invention. Detailed Implementation
[0050] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0051] refer to Figure 1 A method for capturing and locking maritime equipment based on a reinforcement decision-making algorithm, comprising:
[0052] S1. Collect and capture the interaction state data between the locking device and the target equipment. The interaction state data includes relative pose information, relative motion information, contact force and / or contact torque information, and the state of the locking actuator. S2. Preprocess the interaction state data to calculate the current state features characterizing the current interaction state. S3. Under the confirmation phase mode constraints of the previous control cycle, input the current state features into the gated multi-model dynamics prediction model for prediction, obtaining the dynamics sub-model weights, state predictions, and event risk indicators. S4. Based on the current state features, dynamics sub-model weights, state predictions, and event risk indicators, the reinforcement learning decision model outputs candidate phases. The process involves several steps: S5, inputting candidate stage modes and event risk indicators into the switching protection module for switching release determination, and suppressing stage mode switching based on preset hysteresis thresholds and preset minimum dwell time to determine the confirmation stage mode and corresponding confirmation objective function weight parameters for this control cycle; S6, selecting the dynamic sub-model corresponding to the confirmation stage mode as the rolling prediction model based on the current state characteristics, confirmation stage mode, and confirmation objective function weight parameters, and generating control commands for the capture and locking device by the rolling optimization controller to complete the stage advancement and stage mode switching of the capture and locking process.
[0053] In this specific embodiment, S1 includes:
[0054] The controller uses a fixed control cycle Run and control cycle number Identifier At each sampling time, within each control cycle, the data acquisition unit on the locking device side simultaneously reads the outputs of the position and attitude measurement device, the force sensor, and the status feedback unit of the locking actuator, and writes a timestamp generated by the same master clock to each output data. The master clock uses the IEEE 1588 precision time protocol to perform hardware time synchronization with each sensor interface board to ensure consistent time reference across devices.
[0055] The position and attitude measurement device consists of a binocular camera fixedly mounted on the capture and locking device, a visual marker fixedly mounted on the target equipment, and an embedded computing unit. The embedded computing unit stores the camera's intrinsic parameters, binocular extrinsic parameters, and the geometric parameters of the visual marker in the target equipment's coordinate system. Within each control cycle, it performs image acquisition, feature extraction, marker pose calculation, and rigid body transformation calculation, thereby outputting the relative position and attitude of the capture and locking device relative to the target equipment. Simultaneously, it calculates the pose difference and time interval between adjacent control cycles. Calculate and output the relative velocity, which includes the relative linear velocity and the relative angular velocity, and write it into the data frame of this control cycle along with the relative position and relative attitude as relative pose information and relative motion information;
[0056] The force sensor is fixedly installed at the capture contact interface of the capture and locking device and adopts a six-dimensional force and torque sensor structure. In each control cycle, it outputs the contact force and contact torque. The contact force is a three-dimensional vector composed of force components along the three axes of the sensor's measurement coordinate system, and its unit is... The contact torque is a three-dimensional vector composed of torque components around the three axes of the sensor measurement coordinate system, and its unit is _____. The data acquisition unit records the sensor range status word during reading to indicate whether saturation has occurred, thus ensuring that invalid measurements can be eliminated in subsequent processing.
[0057] The locking actuator is an electric cylinder type locking mechanism, and its state feedback unit includes a stroke sensor, a speed estimation module, and a driver feedback interface. The stroke sensor outputs the position state of the locking actuator in each control cycle, denoted by m. The speed estimation module performs numerical differentiation on the position state and outputs the speed state of the locking actuator, denoted by m. This means that the driver feedback interface outputs the drive state of the locking actuator in each control cycle and represents it in terms of drive current or equivalent drive force command to characterize the actuator's output capability and saturation boundary;
[0058] In this embodiment, the aforementioned interactive state data is used in the control cycle. The collected results are uniformly encapsulated into an interaction state data vector. And satisfy:
[0059] ;
[0060] in Indicates the first The interaction state data vector of each control cycle Represents a relative position vector. Represents a relative attitude quaternion vector. It represents a relative velocity vector and is composed of relative linear velocity and relative angular velocity. Represents the contact force vector. Represents the contact torque vector. This represents the state vector of the locking actuator and is composed of position states. Speed state With drive state Composed of splicing elements, superscript This indicates transpose and is used to concatenate the components in a column vector manner. Along with timestamps Write to the circular buffer and send to step The preprocessing module is used for subsequent synchronization, coordinate unification and filtering.
[0061] In this specific embodiment, S2 includes:
[0062] The preprocessing module receives the interaction state data vector. and its timestamp Then, time synchronization processing, coordinate unification processing, and temporal filtering processing are completed sequentially, and the current state features used to characterize the current interaction state are calculated. ;
[0063] Time synchronization processing employs a "master clock alignment and resampling" mechanism, with the preprocessing module using the master clock time... As the alignment reference for this control cycle, timestamps falling within the interval are retrieved from the circular buffers of each sensor. Inside and closest One piece of data is used as the sensor in the control cycle. The synchronization value is set such that when a sensor has no valid data in the specified range, the synchronization value of the sensor in the previous control cycle is kept as the synchronization value of the current control cycle, and the holding event is written into the data validity flag for subsequent modules to determine.
[0064] Coordinate unification processing adopts a unified coordinate system The unified coordinate system The locking docking reference surface is fixed to the target equipment, and its origin is defined as the center of the locking hole or locking pin target. of The shaft is defined as the desired approach direction of the capture and locking device and shaft and The shaft is tensioned to form a docking reference plane, thereby uniformly expressing the relative position, relative attitude, relative speed, contact force, contact torque, and the state of the locking actuator after time synchronization. In this process, relative position and relative velocity are transformed by rotation and translation through coordinate transformation parameters pre-calibrated and fixed in the controller. Relative attitude is represented by quaternions in a unified coordinate system, and reference frame transformation is completed through quaternion multiplication. Contact force and contact torque are transformed from the sensor measurement coordinate system to a pre-calibrated fixed rotation relationship. Furthermore, when there is a known force sensor installation eccentricity, the torque term corresponding to this eccentricity is compensated to the contact torque to ensure that the force and torque are expressed about the same reference point.
[0065] Timing filtering processes apply first-order discrete low-pass filtering to each channel signal after coordinate unification to suppress the impact of sea state disturbances and measurement noise on stage switching boundaries. The filter uses fixed coefficients. Furthermore, the relative position, relative attitude, relative speed, contact force, contact torque, and locking actuator state are executed independently, and their update relationships are as follows:
[0066] ;
[0067] in Indicates control cycle The original signal vector after coordinate unification for a certain type. This indicates that this type of signal vector is in the control period The filtered output, This indicates that this type of signal vector is in the control period The filtered output, This represents the filter coefficients and is used to set the filter time constant. Indicates the control cycle number;
[0068] After filtering, the preprocessing module uses the filtered relative positions as a basis. Filtered relative attitude Filtered relative velocity Filtered contact force Filtered contact torque and the status of the locking actuator after filtering Calculate the current state features The centering deviation is determined by The lateral alignment deviation, axial clearance deviation, and attitude mismatch angle together constitute the deviation. The lateral alignment deviation is taken as... exist shaft and The Euclidean norm of the axial component, and the axial clearance deviation are taken as... exist The absolute value in the axial direction, the attitude mismatch angle is determined by... The minimum rotation angle relative to the unit quaternion is obtained and restricted to the interval. Internally, it ensures uniqueness;
[0069] Contact status indicators use discrete level variables Represent and write , among which when Time setting Represents the free space state, when And lock the position status of the actuator Less than the locking start position threshold Time setting Indicates the guiding contact state, when and Time setting This indicates a constrained contact state, while also specifying the direction of the contact force. of The axis projection symbol is written as a normal contact direction mark. To distinguish between compaction and pull-out trends, the current state characteristics are ultimately formed, including the centering deviation and contact state indication. The data validity flag is then output to the gated multi-model dynamics prediction model in step S3.
[0070] In this specific embodiment, S3 includes:
[0071] Gated multi-model dynamics prediction model in each control cycle Receive current state features and the confirmation phase mode of the previous control cycle and in Under the constraints of the pattern, the state of the next moment is predicted in one step and the event risk index is output.
[0072] The set of phase patterns is defined as follows Where F represents free space approach mode, C represents guide contact mode, S represents semi-constrained sliding mode, and L represents fully constrained locking mode;
[0073] The gating function is implemented using a two-layer fully connected neural network, and its parameters are fixed at the factory when the controller is shipped. The network input is... The vector after channel-wise linear normalization has 32 neurons in the first hidden layer with ReLU activation function, and the number of neurons in the output layer is... It also outputs the unnormalized score for each mode. ;
[0074] To achieve the "mode constraint confirmation phase of the previous control cycle", the gating function introduces mode transition prior during softmax normalization. The mode transition prior is determined by a fixed transition matrix. Given in pattern order A permutation whose non-zero elements take values of , The remaining elements are set to 0 to prevent cross-level jumps, thus ensuring that the gating weights are only allocated within the range of "maintaining the current mode or transferring to an adjacent mode";
[0075] The dynamic sub-model contains four one-beat predictors. Each predictor with The next time step state prediction for the input-output corresponding pattern The F sub-model employs a contactless assumption and primarily maintains relative velocity to perform a one-step extrapolation of relative pose, while the C sub-model operates in a unified coordinate system. Introducing linear contact compliance in the normal direction and employing fixed contact stiffness With fixed contact damping Predicting the relative velocity change caused by normal contact and using fixed tangential damping in the tangential direction. To suppress high-frequency slip vibration, the S-submodel introduces Coulomb friction while retaining the normal compliance of the C-submodel and adopts a fixed friction coefficient. To predict the tangential slip trend, the L sub-model tightens the relative pose constraints and adopts a fixed contact stiffness. With fixed contact damping Predict small deformations under full constraints and use the state of the locking actuator as a constraint convergence variable in one-step prediction;
[0076] The gated multi-model dynamics prediction model outputs the weights of the dynamics sub-models based on the gating function and performs weighted fusion of the predictions of each sub-model. The calculation relationship is as follows:
[0077] ;
[0078] ;
[0079] in Indicates control cycle Time mode The weights of the dynamic sub-model, Indicates the gating function for the mode The unnormalized score of the output This indicates the mode during the confirmation phase of the previous control cycle. Transition to pattern under constraints The prior coefficients and by the matrix Indexing yields, Represents an exponential function. This represents the pattern index for summing over all patterns. Represents a set of phase patterns. Representation pattern The next moment state prediction, Representation pattern The dynamic sub-model predictor, Indicates the current state characteristics. This represents the next-time state prediction obtained through weighted fusion;
[0080] Event prediction module with and The concatenated vector is used as input and implemented using a two-layer fully connected neural network with three outputs. The parameters of this network are fixed at the factory. The first hidden layer has 32 neurons and uses ReLU activation. The output layer uses a sigmoid mapping to restrict the three outputs to a range. Internally, this allows us to obtain indicators of stagnation risk. Rebound risk indicators With locking accessibility index ,in This is used to characterize the degree of risk of significant obstruction of relative motion due to geometric interference or frictional self-locking at the next moment. Used to characterize the risk of separation or rebound due to reversed normal velocity after contact. Used to characterize the degree to which a fully constrained locking mode can be entered under the current centering deviation and locking actuator state constraints;
[0081] Output dynamics sub-model weights Next-moment state prediction and event risk indicators .
[0082] In this specific embodiment, S4 includes:
[0083] Reinforcement learning decision-making models in each control cycle Receive current state features and the weights of the dynamic sub-model Next-moment state prediction Event risk indicators The above quantities are then assembled in a fixed order to form decision features. Subsequently Linear normalization is performed channel by channel to obtain the network input vector. The channel mean vector and channel standard deviation vector required for normalization are stored as constants in the controller and are consistent with the statistical caliber used during training, thereby ensuring that the data distribution is consistent between online inference and offline training.
[0084] Reinforcement learning decision-making models employ policy networks with deterministic reasoning. Implementation and network parameters After offline training, the policy network is solidified. It consists of a shared feature extraction backbone and dual output heads. The shared feature extraction backbone is a two-layer fully connected network with 64 neurons per layer and uses the ReLU activation function. The first output head is a stage-mode output head with an output dimension of [missing information]. and respectively correspond The free space approach mode F, guided contact mode C, semi-constrained sliding mode S, and fully constrained locking mode L are defined. The second output head is the objective function weight output head with an output dimension of 3, and it corresponds to the weight of the deviation term in the middle. Weight of contact force or contact torque term Weights of relative velocity terms ,in The centering deviation term is used in the weighted rolling optimization control objective function. Used as a softening or penalty term for contact force or contact torque constraints in the weighted rolling optimization control objective function. The relative velocity suppression term is used in the weighted rolling optimization control objective function;
[0085] The relationship between the output of the policy network and the candidate stage modes and the weight parameters of the objective function is as follows:
[0086] ;
[0087] ;
[0088] ;
[0089] cen, con, vel ;
[0090] in Indicates control cycle The decision feature vector, Represents the feature vector of the current state. Representing the patterns respectively The weights of the dynamic sub-model, This represents the next-time state prediction vector obtained through weighted fusion. Indicates the risk of stagnation. Indicators indicating the risk of a rebound Indicates the lockability index. Indicates to The normalized network input vector, The parameter is The policy network, This represents the unnormalized score vector of the stage mode output header. This represents the original output vector of the objective function weight output head, and its components correspond to... Describes the softmax normalization function and Representation pattern The normalized score, Indicates the candidate phase mode. This represents the pattern index corresponding to the maximum value. This represents the sigmoid function. Indicates the first Each objective function weight parameter and They represent the first The lower and upper bounds of each weight parameter are fixed and take values of... , ;
[0091] Candidate phase mode With the objective function weight parameter vector The output to step S5 is used to switch the release confirmation of the protection module.
[0092] In this specific embodiment, S5 includes:
[0093] The switching protection module is used in each control cycle. Candidate phase mode With the objective function weight parameter vector Simultaneously receive event risk indicators and stagnation risk indicators Rebound risk indicators With locking accessibility index And read the confirmation phase mode of the previous control cycle. Confirmation of the objective function weight parameter vector from the previous control cycle and the stage dwell time count value ,in and The values all belong to the stage pattern set. Corresponding to the free space approach mode F, guided contact mode C, semi-constrained sliding mode S, and fully constrained locking mode L respectively, the switching protection module uses a deterministic switching discrimination model to generate a switching permission flag. and And respectively indicate "not allowed" and "allowed";
[0094] when Time setting And directly determine the confirmation phase mode of this control cycle. And confirm the objective function weight parameter vector This avoids introducing unnecessary weight fluctuations when there are no changes during the candidate stage; when When switching the discrimination model, both the minimum residence time constraint and the risk threshold constraint are checked simultaneously. The minimum residence time is defined as the number of control cycles. And require Risk threshold constraints use fixed thresholds. and And require and and When the above constraints are simultaneously satisfied, And Updated to And will Updated to Set when any constraint is not satisfied and maintain and ;
[0095] Meanwhile, to suppress the jitter switching between the semi-constrained slip mode S and the fully constrained locking mode L near the critical reachability, the switching protection module controls the switching of the involved... The candidate switching implements a hysteresis threshold mechanism and overrides the locking reachability threshold criterion, wherein the locking entry threshold is set to... The lockout exit threshold is set as follows:
[0096] And satisfy ,when and Only when and and and Simultaneously satisfying the condition is required for placement. And allows switching from S to L, when and Only when and and and Simultaneously satisfying the condition is required for placement. It also allows switching from L to S, thus maintaining the stability of the phased mode when the lockout accessibility metric fluctuates around a single threshold;
[0097] After completion and After confirmation, the switching protection module confirms the phase mode according to the current control cycle. Compared to the confirmation phase mode of the previous control cycle Has the dwell time count changed during the update phase? Its update relationship is as follows:
[0098] ;
[0099] in Indicates control cycle The stage dwell time count value, Indicates control cycle The stage dwell time count value, Indicates control cycle The confirmation phase mode, Indicates control cycle The confirmation phase model will ultimately and Together with the switching permission flag Output to steps This is used to select the prediction model and construct the objective function for the rolling optimization controller.
[0100] In this specific embodiment, S6 includes:
[0101] Rolling optimization controller in control cycle Read current state characteristics Read confirmation phase mode With confirmation of the objective function weight parameter vector and select from the model library One-to-one dynamic sub-model as rolling prediction model ;
[0102] in The relative position, relative attitude, relative speed, contact force, contact torque, and locking actuator state and its derived quantities are combined and include centering deviation and contact state indication, thus ensuring that the rolling prediction includes both geometric centering state and contact interaction state.
[0103] The rolling prediction model Let be a discrete-time one-step state transition function with a sampling period of . Its internal parameters and steps During this stage, the contact stiffness, contact damping, and friction coefficient remain consistent and fixed, and are set according to a unified coordinate system within the model. The evolution of pose and velocity is coupled with the evolution of contact force and contact torque in the calculation, so that the predicted state includes the predicted quantities of contact force and contact torque consistent with the optimization constraints.
[0104] The scroll optimization controller in the scroll time domain length Within the prediction window, a constrained discrete-time optimization problem is constructed and solved, with the decision variable being the control sequence:
[0105] ;
[0106] in Indicates during the control cycle Next for the future The control amount applied by the tap, Indicates in The displacement increment control quantity below, Indicates in Small angle incremental control of the attitude. This indicates a locking actuator control command with a value that is a normalized drive command and mapped to a motor current command by the driver.
[0107] The optimization problem is in Solve under the corresponding control constraints, and the control constraints are in sets. and It means that, among them Apply single-shot amplitude constraints uniformly to all stages. And apply range constraints to the position state of the locking actuator. To prevent the mechanical travel from exceeding the limit;
[0108] Approach mode to free space Apply contact force prediction norm constraints With relative velocity norm constraint , for guided contact mode Apply contact force prediction norm constraints Contact torque prediction norm constraint Relative velocity constraint with normal direction Semi-constrained slip mode Apply contact force prediction norm constraints Tangential relative velocity constraint Locking actuator command suppression constraint To ensure that this stage only performs centering slippage without premature locking, a fully constrained locking mode is used. Applying relative velocity norm constraints Restriction of closing direction of locking actuator To ensure that the locking action proceeds in one direction only;
[0109] Based on this, the rolling optimization controller solves the following optimization problem and obtains the optimal control sequence. :
[0110] ;
[0111] in Indicates during the control cycle Information about the future The predicted state of the shot, The initial predicted state is represented by a feature equal to the current state. Indicates the future number Control of the amount of shooting, Indicates the future number Control of the amount of shooting and when The actual control quantity issued in the previous control cycle is used to achieve continuous smoothing. Indicates by The obtained centering deviation vector includes lateral centering deviation and attitude mismatch. Indicates by The contact interaction vector obtained by analysis and the contact force With contact torque It is assembled in a fixed order. Indicates by The relative velocity vector obtained from analysis, Represents the L2 norm, Represents the infinite norm, This indicates the weight of the centering deviation term. Indicates the weight of the contact force or contact torque term. Indicates the weight of the relative velocity term. This indicates that the control increment smoothing weights have fixed values. ;
[0112] The optimization problem is solved using sequential quadratic programming, with two linearization iterations performed in each control cycle. The linearization point is the sequence after shifting the optimal control sequence of the previous control cycle by one cycle as the initial value, and is warm-started in the controller memory. The quadratic programming subproblem is solved using the active set method, with a maximum number of iterations set to 30 and a first-order optimality residual threshold set to [value missing]. Thus ensuring A feasible control sequence is obtained within the system;
[0113] The rolling optimization controller will select the optimal control sequence. The first control quantity The control command for this control cycle is then broken down to obtain the pose adjustment control command. Control commands for locking actuator Subsequently, the pose adjustment control command is sent to the pose adjustment actuator of the capture and locking device via the real-time bus, and the locking actuator control command is sent to the locking actuator driver to drive the capture and locking device according to the confirmation phase mode. Implement phased progression and phase mode switching in the capture and locking process.
[0114] In this specific implementation, each dynamic sub-model Includes identifiable parameters related to contact stiffness and contact damping, and in each control cycle Based on steps Forming current state features The filtered relative position used at that time Filtered relative velocity Contact force after filtering Contact torque after filtering The identifiable parameters are updated online to correct the prediction for the next period;
[0115] Unified coordinate system of The axis is defined as the desired approach direction, and the normal contact amount is defined on that axis. The controller then... Read Axial components ,from Read Axial components ,from Read Axial components And construct the normal compressive displacement With normal compression velocity ,in Indicates control cycle The normal compressive displacement is denoted by m. Indicates control cycle normal compression velocity and with express, Indicates the relative position after filtering. Components on the axis, Indicates the relative velocity after filtering. Components on the axis;
[0116] Simultaneously, the normal contact force scalar is defined as ,in Indicates control cycle The normal contact force scalar is denoted by N. Indicates the contact force after filtering. Components on the axis;
[0117] Online updates only apply to contact status indicators. and and and Execution time, among which This indicates that the output from step S2 is written to... Contact status indication quantity, This represents the L2 norm of the filtered contact torque vector. This represents the filtered contact torque vector;
[0118] The controller confirms the mode of the previous control cycle. Maintain an independent set of parameter estimates for this model. With covariance matrix ,in Indicates control cycle And the mode is The estimated normal contact stiffness at that time, and the unit is Indicates control cycle And the mode is The estimated normal contact damping value at that time, and the unit is , This indicates the value corresponding to the parameter estimate. Covariance matrix;
[0119] The parameter estimation is set to [value] during system power-on initialization. , , ,and ,in express identity matrix;
[0120] When the update conditions are met, a forgetting factor is used. The recursive least squares method for The update is performed, and the update relationship is as follows:
[0121]
[0122]
[0123]
[0124] ;
[0125] in Represents the regression vector. Represents the recursive gain vector. Indicates the forgetting factor, This represents the covariance matrix of the previous control cycle. This represents the parameter estimation vector for the previous control cycle;
[0126] After the update is completed, and A physically feasible region projection is performed to suppress noise-induced divergence, where the stiffness projection interval is fixed. And the damping projection range is fixed as And write the projected parameters into the mode. Corresponding dynamic sub-model This replaces the fixed contact stiffness and fixed contact damping in the sub-model in subsequent predictions, thereby using online identification results to correct the accuracy of the gated multi-model dynamic prediction model in calculating the state evolution and event risk indicators of the contact stage.
[0127] In this specific embodiment, the preset minimum dwell time used in step S5 and the threshold for stuck switching Rebound switching threshold Locking switching threshold Locking entry threshold With lockout threshold Instead of using a single constant, parameters are selected by index from multiple preset parameter sets built into the controller's non-volatile memory and applied in each control cycle. Write to the parameter register of the switching protection module;
[0128] The preset parameter set is defined as follows Each group contains six parameters. ,in , And each group satisfies To ensure that the delay is valid;
[0129] The index selection uses three types of index quantities simultaneously: sea state grade, target equipment model, and prediction uncertainty index, where the sea state grade is denoted as... And the significant wave height output by the sea state sensor The average obtained by averaging over a 60-second sliding time window and then discretized is as follows: when when when ;
[0130] The target equipment model is designated as Furthermore, the target equipment identification code in the mission configuration file is directly read and fixed as... or and adopt defined rules to Integrating into environmental levels, specifically when Environmental level ,when Environmental level To reflect the impact of the target equipment model on contact uncertainty and operational risk;
[0131] Prediction uncertainty is indicated as Furthermore, it is calculated from the sub-model weights and sub-model prediction discrepancies output by the gated multi-model dynamics prediction model, and its calculation formula is as follows:
[0132] ;
[0133] in Indicates control cycle The prediction uncertainty index The index of the stage pattern and the value belonging to the stage pattern set represent the control cycle. Time mode The weights of the dynamic sub-model, Representation pattern The next moment state prediction, This represents the next-time state prediction obtained by weighted fusion of sub-model weights. Represents the L2 norm;
[0134] The controller will With fixed threshold The comparison forms the uncertainty level ,when Time to take Indicates low uncertainty, when Time to take Indicates high uncertainty; the parameter set selection rule is based on the index triplet. The only certainty is when Time selection This set of parameters is then written into the switching protection module for determining the minimum dwell time, risk threshold, and hysteresis in this control cycle. Time selection The set of parameters is then written into the switching protection module for determining the minimum dwell time, risk threshold, and hysteresis in this control cycle. This allows the switching protection parameters to undergo deterministic adaptive changes with sea state level, target equipment type, and prediction uncertainty index, while remaining consistent with the switching permission flag generation logic in step S5.
[0135] In this specific embodiment, in step S6, the first control quantity of the control sequence obtained by the rolling optimization controller is... Before being identified as a control command, the controller performs... Perform security filtering and output security control quantities. ,in Indicates control cycle The first optimal control quantity. Representing a unified coordinate system The displacement increment control quantity below, Representing a unified coordinate system Small angle incremental control of the attitude. This indicates a control command for locking the actuator;
[0136] Security filtering processing follows the model library in step S3 and the confirmation phase mode. Corresponding dynamic sub-model As a one-beat forward predictor, and based on the current state features Using the candidate control variable as input, the next-step predicted state is calculated, and the next-step predicted contact force vector is obtained by parsing the next-step predicted state. Next beat prediction of contact torque vector Predicting the relative velocity vector with the next beat and with fixed safety boundary parameters It was checked, among which Indicates the safety boundary of contact force. Indicates the safety boundary of the contact torque. Indicates the relative velocity safety boundary;
[0137] when Not greater than and Not greater than and Not greater than At that time, the safety filter processing output And use it as the control command for this control cycle;
[0138] When any safety boundary is triggered, the safety filtering process uses line segment projection along the control increment direction and a binary search to determine the maximum release coefficient. First, the actual control quantity issued in the previous control cycle is recorded as the reference control quantity. initialize the search interval to and Then perform 12 binary search iterations and take the result each time. Construct test control variables and pass Perform a one-shot prediction check of the safety boundary, and when the check meets the safety boundary, set... When the check does not meet the safety boundary, After the iteration is complete, take The safety control quantity is output according to the following relationship:
[0139] ;
[0140] in Indicates control cycle Safety control quantity, Indicates control cycle Safety control quantity, This represents the maximum release coefficient and its value range is [value range missing]. , This represents the first optimal control output from the rolling optimization controller;
[0141] After the binary search iteration ends At that time, safety filtration will Fixed as Furthermore, the control command component of the locking actuator is set to zero to suppress locking advancement, and the control quantity is re-solved after updating the state characteristics and risk indicators in the next control cycle. Decomposed into pose adjustment control commands Control commands for locking actuator And send it to the capture and locking device for execution.
[0142] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0143] This invention addresses the challenges of complex stage transitions and prone to misjudgment of boundary conditions in the maritime capture and locking process, which involves multiple stages of interaction, including free-space approach, guided contact, semi-constrained slippage, and fully constrained locking. It constructs a closed-loop system comprising world model prediction and risk assessment, enhanced decision-making to provide candidate stages and control weights, switching protection to suppress erroneous switching, and rolling optimization to generate control commands. Specifically, a gated multi-model dynamics prediction model outputs corresponding state predictions and event risk indicators at different interaction stages. This enables the controller to not only control based on current observations but also to proactively assess jamming, rebound, and locking accessibility using prediction results. The reinforcement learning decision-making model combines prediction information to output candidate stage patterns and objective function weight parameters. This allows rolling optimization control to execute with more matched control objectives and constraints at different stages, improving the stability of stage progression and locking success rate, while reducing the probability of repeated unlocking retries.
[0144] This invention addresses the aforementioned technical problems by improving multi-stage hybrid dynamics in several ways: First, it employs a gated multi-model structure and performs predictions under the constraint of the confirmed stage mode in the previous control cycle, making the dynamic predictions more stable at stage boundaries and reducing mode drift caused by noise or disturbances. Second, it adds an event prediction module to the world model, explicitly outputting indicators such as jamming risk, bounce risk, and lockout reachability, providing quantifiable safety and feasibility data for stage switching. Third, it introduces a switching protection mechanism including a hysteresis threshold and minimum dwell time to confirm the release of candidate stages output by reinforcement learning, thus suppressing jittery and erroneous switching. Through these structural improvements, this invention can more reliably complete stage switching and lockout process control under complex sea conditions and contact uncertainties, thereby achieving a safer, more stable, and higher-success-rate capture and lockout technology.
Claims
1. A method for capturing and locking marine equipment based on a reinforcement decision-making algorithm, comprising: S1. Collect and capture the interaction state data between the locking device and the target equipment. The interaction state data includes relative pose information, relative motion information, contact force and / or contact torque information, and the state of the locking actuator. S2. Preprocess the interaction state data to calculate the current state features characterizing the current interaction state. S3. Under the constraint of the confirmation phase mode of the previous control cycle, input the current state features into the gated multi-model dynamics prediction model for prediction, obtaining the dynamics sub-model weights, state predictions, and event risk indicators. S4. Based on the current state features, dynamics sub-model weights, state predictions, and event risk indicators, the reinforcement learning decision model outputs candidate phase modes and objective function weight parameters for rolling optimization control. S5. Input the candidate stage mode and event risk index into the switching protection module for switching release determination, and suppress the stage mode switching according to the preset hysteresis threshold and preset minimum dwell time to determine the confirmation stage mode and the corresponding confirmation objective function weight parameters of this control cycle; S6. Based on the current state characteristics, confirmation stage mode and confirmation objective function weight parameters, select the dynamic sub-model corresponding to the confirmation stage mode as the rolling prediction model, and generate control commands for the capture and locking device by the rolling optimization controller and send them to the capture and locking device for execution to complete the stage advancement and stage mode switching of the capture and locking process.
2. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S1 includes: During each control cycle, the position and attitude measurement device acquires the relative position, relative attitude, and relative velocity of the capture and locking device relative to the target equipment. The contact force or contact torque of the capture and locking device is obtained by a force sensor installed on the capture and locking device; The state of the locking actuator is obtained by the stroke sensor or state feedback unit of the locking actuator, and the state of the locking actuator includes the position state, speed state and drive state of the locking actuator.
3. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S2 includes: The relative position, relative attitude, relative speed, contact force or contact torque, and the state of the locking actuator are synchronized in time so that the relative position, relative attitude, relative speed, contact force or contact torque, and the state of the locking actuator correspond to the same control cycle. The data that has completed time synchronization processing is subjected to coordinate unification processing, and the relative position, relative attitude, relative velocity, and contact force or contact torque are transformed to a preset unified coordinate system; The data that has undergone coordinate unification processing is subjected to time-series filtering to obtain the filtered relative position, filtered relative attitude, filtered relative velocity, filtered contact force or contact torque, and filtered locking actuator status. The current state characteristics are calculated from the filtered relative position, filtered relative attitude, filtered relative speed, filtered contact force or contact torque, and filtered locking actuator state. The current state characteristics include centering deviation and contact state indication.
4. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S3 includes: Under the mode constraints of the confirmation phase of the previous control cycle, the current state characteristics are fed into the ingress control multi-model dynamic prediction model; The gating function of the gated multi-model dynamics prediction model outputs the weights of the dynamic sub-models corresponding to the free space approach mode, guided contact mode, semi-constrained sliding mode and fully constrained locking mode respectively according to the current state characteristics, and makes the weights of the dynamic sub-models satisfy the normalization constraint. Each dynamic sub-model outputs its own next-time state prediction based on the current state characteristics. The next time step state prediction is obtained by weighting and fusing the predicted states of each next time step according to the weights of the dynamic sub-models. The event prediction module of the gated multi-model dynamics prediction model outputs the event risk index based on the current state characteristics and the next moment state prediction. The event risk index includes a jamming risk index, a rebound risk index, and a locking reachability index.
5. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S4 includes: The current state features are combined with the weights of the dynamic sub-model, the predicted state at the next moment, and the event risk index to form decision features, and the decision features are fed into the reinforcement learning decision model for inference calculation. The reinforcement learning decision model outputs candidate stage patterns within the stage pattern set and outputs objective function weight parameters for rolling optimization control, wherein the objective function weight parameters include centering deviation term weight, contact force or contact torque term weight, and relative velocity term weight.
6. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S5 includes: The candidate phase pattern and event risk indicators are input into the handover protection module, and the handover discrimination model of the handover protection module generates a handover permission flag. The switching permission flag is determined according to the following conditions: when the candidate stage mode is different from the confirmation stage mode of the previous control cycle, the stage dwell time count value reaches the preset minimum dwell time, and the event risk indicators meet the following conditions: the jamming risk indicator is not greater than the jamming switching threshold, the rebound risk indicator is not greater than the rebound switching threshold, and the locking reachability indicator is not less than the locking switching threshold. Meanwhile, the switching of phase modes is suppressed according to the preset hysteresis threshold. Specifically, when the candidate phase mode meets the locking reachability index not less than the locking entry threshold, it is allowed to switch from the semi-constrained sliding mode to the fully constrained locking mode. When the candidate phase mode meets the locking reachability index not greater than the locking exit threshold, it is allowed to switch from the fully constrained locking mode to the semi-constrained sliding mode. The locking entry threshold is greater than the locking exit threshold. When the allow flag is switched to allow, the confirmation phase mode of this control cycle is updated to the candidate phase mode, and the objective function weight parameters are determined as the confirmation objective function weight parameters of this control cycle. When the allow flag is switched to disallow, the confirmation phase mode of the current control cycle is maintained as the confirmation phase mode of the previous control cycle, and the confirmation objective function weight parameters of the previous control cycle are determined as the confirmation objective function weight parameters of the current control cycle. The stage dwell time count is updated based on whether the confirmation stage mode of the current control cycle has changed compared to the confirmation stage mode of the previous control cycle. Specifically, the stage dwell time count is incremented when there is no change, and the stage dwell time count is cleared when there is a change.
7. The maritime equipment capture and locking method based on a reinforcement decision algorithm according to claim 1, S6 includes: The rolling optimization controller reads the current state features, the confirmation phase mode, and the weight parameters of the confirmation objective function, and determines the dynamic sub-model corresponding to the confirmation phase mode as the rolling prediction model. In the rolling time domain, the state evolution of the capture and locking device is predicted based on the rolling prediction model, an optimization objective function is constructed based on the weight parameters of the confirmation objective function, and the control sequence is optimized and solved under the control constraints corresponding to the confirmation stage mode. The first control quantity of the control sequence obtained by optimization is determined as the control command, which includes a pose adjustment control command and a locking actuator control command; The control command is then sent to the capture and locking device for execution, so as to drive the capture and locking device to complete the phase advancement and phase mode switching of the capture and locking process according to the confirmation phase mode.
8. A method for capturing and locking marine equipment based on a reinforcement decision-making algorithm according to claim 1, characterized in that, In S3, the kinetic sub-model includes identifiable parameters related to contact stiffness, contact damping, and / or friction, and the identifiable parameters are updated online based on the current state characteristics and the contact force and / or contact torque information to correct the predictions of the kinetic sub-model.
9. A method for capturing and locking marine equipment based on a reinforcement decision-making algorithm according to claim 6, characterized in that, The preset minimum dwell time and / or jamming switching threshold, rebound switching threshold, locking switching threshold, locking entry threshold, and locking exit threshold are selected from multiple preset parameter groups. The preset parameter groups are selected by indexing according to sea state level, target equipment model, and / or the prediction uncertainty index.
10. A method for capturing and locking marine equipment based on a reinforcement decision-making algorithm according to claim 7, characterized in that, Before determining the first control quantity of the control sequence obtained by optimization as the control command, the method further includes performing safety projection or safety filtering on the first control quantity to ensure that the control command meets the preset safety boundary constraints of contact force, contact torque and relative velocity.