A human-robot co-adaptation underwater robot operation decision method and system
Patent Information
- Application Number
- CN202610921295.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-25
AI Technical Summary
首先,操作人员发出的多模态控制指令(如手柄操作、语音命令、手势指示等)难以被完整、高效地编码与传输,传统视频流压缩方法在极度有限的带宽下无法兼顾语义信息的完整保留与传输效率,极易导致远端系统对操作意图的理解出现偏差
(1)本发明的人机协同决策方法通过多模态意图语义压缩技术,显著降低深海低带宽声学信道的通信负担,实现操作者复杂指令的精准理解与高效传输,克服传统视频流压缩导致的语义失真问题;
Smart Images

Figure CN122469858B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater robot collaborative operation technology, and in particular to a human-robot collaborative underwater robot operation decision-making method and system. Background Technology
[0002] Currently, remotely operated vehicles (ROVs), as core equipment for deep-sea operations, are widely used in fields such as subsea oil and gas pipeline inspection, deep-sea mineral exploration, underwater structure maintenance, and marine scientific research. However, existing ROV systems primarily rely on fiber optic cables or umbilical cables to connect to surface mother ships for real-time, high-bandwidth data transmission and command interaction. While this wired communication method provides millisecond-level low latency and high reliability data links, the physical cable length strictly limits the effective operating radius of the ROV to a few kilometers. Furthermore, the deployment, recovery, and maintenance of the system all require the support of large support vessels, resulting in extremely high overall operating costs.
[0003] To overcome the limitations of operational radius, the untethered autonomous operation solution adopts underwater acoustic communication technology. Underwater acoustic communication utilizes the properties of sound waves propagating in water, theoretically enabling communication distances of tens or even hundreds of kilometers, making large-scale independent operation of underwater robots possible. However, underwater acoustic channels have inherent limitations such as narrow bandwidth (typically ≤10kbps) and significant propagation delays (up to the second level).
[0004] Under the aforementioned limited communication conditions of low bandwidth and high latency, existing technologies face a series of severe challenges. First, multimodal control commands issued by operators (such as handle operations, voice commands, and gesture instructions) are difficult to encode and transmit completely and efficiently. Traditional video stream compression methods, under extremely limited bandwidth, cannot simultaneously ensure the complete preservation of semantic information and transmission efficiency, which can easily lead to deviations in the remote system's understanding of the operator's intentions. Second, significant signal latency causes severe spatiotemporal misalignment between the underwater robot's status feedback and the operator's control commands. This makes the traditional master-slave direct teleoperation mode highly susceptible to trajectory tracking instability or even loss of control under dynamic and uncertain deep-sea current disturbances. Although existing digital twin technologies can predict system states to some extent, they are mostly unidirectional and static simulation models, lacking bidirectional predictive compensation and robust command stream reconstruction capabilities, making it difficult to effectively address the sparsity and non-uniformity of the command stream caused by time-varying channels. Furthermore, most existing human-machine collaborative control systems employ fixed permission allocation strategies, failing to dynamically adjust to real-time changes in channel quality and sudden environmental disturbances. This results in a high rate of human-machine command conflicts, poor system robustness, and a significant reduction in the success rate of complex tasks. More seriously, existing systems lack a layered and progressive security protection mechanism, failing to quickly degrade to an intrinsically safe physical backup control mode under extreme abnormal conditions, posing a serious security vulnerability.
[0005] In summary, there is an urgent need for a flexible decision-making method that can adapt to restricted communication in the deep sea, so as to achieve precise alignment of intents, dynamic optimization of permissions, and closed-loop stability assurance, thereby overcoming the risks of collaborative failure such as inaccurate state synchronization, distortion of intent transmission, and control conflicts. Summary of the Invention
[0006] This invention provides a human-robot collaborative underwater robot operation decision-making method and system to achieve precise intent alignment, dynamic permission optimization, and closed-loop stability assurance, thereby overcoming the risks of collaborative failure such as inaccurate state synchronization, distorted intent transmission, and control conflicts.
[0007] The technical solution of the present invention is as follows: A human-robot collaborative underwater robot operation decision-making method includes the following steps: (1) Establish a communication-task effectiveness quantification model for multimodal perception channels and generate a modality selection map to dynamically recommend the optimal combination of perception modes; use a multimodal big language model to align and fuse the natural language instructions from the operator's end with the scene information and parse them into structured task instructions. (2) Construct a bidirectional predictive digital twin, perform forward prediction of the delayed robot state at the operator end, simulate the received sparse task instructions at the robot end, reprogram and smooth interpolate the task instructions, and generate a continuous and feasible human task flow. (3) The received sparse task instructions are parsed into robot autonomous decision-making flow through a large language model; control instructions are generated by fusing human task flow and machine autonomous decision-making flow according to dynamic autonomous weights; the final execution mode is selected by adopting a three-layer progressive safety strategy based on the control instructions and the real-time status of the system; a nonlinear model with time-varying delay is established for the entire closed-loop system, and the key parameter boundary conditions to ensure system stability are derived using Lyapunov-Krasovsky functional theory.
[0008] Preferably, in step (1), the communication-task effectiveness quantification model of the multimodal sensing channel includes an information density function. and task performance function : ; ; in, For the original data bandwidth requirements, To determine the intent ambiguity after semantic parsing of modal signals. For the current task phase, This represents the current system state.
[0009] Information density function Quantization of perception modality The amount of effective task information that can be conveyed by a unit of data, the task performance function Quantization of perception modality In specific task phases and system status The contribution of information density function to the final mission success rate. and task performance function Calibration is performed through human factors experiments and simulations.
[0010] Preferably, in step (1), a multimodal large language model is used to align and fuse the natural language instructions from the operator's end with the scene information, parsing them into structured task instructions, including: (1-i) Convert the asynchronous raw signals of each sensing modality into low-level features that can be used for fusion; (1-ii) Bind the instructions in the natural language commands to the eye-tracked object entities; (1-iii) Align the timestamps of each sensing modality signal with the twin scene snapshot; (1-iv) By projecting aligned multi-sensory modal information into the same semantic space through a multimodal large language model, structured task instructions are output.
[0011] The task instructions include core actions, target objects, constraints, and success criteria.
[0012] Step (1-i) includes: Extract time-frequency features (Mel spectrogram or Mel frequency cepstral coefficients) from audio signals as low-level features of the audio signals; Convolutional neural networks are used to extract feature maps or feature vectors from visual signals (cameras, eye trackers) as low-level features of the visual signals. The temporal feature stream of six-DOF pose, velocity, or force / torque data of the handle / pose signal is used as a low-level feature of the handle / pose signal. The digital twin scene information is structured into text descriptions or attribute graphs as low-level features of the scene information.
[0013] Multimodal large language models evaluate the contribution of each perceptual modality through an internal attention mechanism, which serves as the internal confidence level for each perceptual modality.
[0014] To improve robustness, preferably, step (1) further includes a spatiotemporal attention fusion network guided by the confidence of each perceptual modality, which performs weighted fusion and short-term intent prediction on the multi-perceptual modal signals to offset transmission delay; including: Align the low-level feature sequences of each perception modality on the time axis; Construct a dynamic weight generator to generate dynamic weights for each perceptual modality. : ; in, For a moment Perception modality Low-level features; It is a state of historical integration; This is the current task phase; This represents the current system state. For sensing modes The confidence prior; It is a multilayer perceptron; The fusion intent feature is obtained by weighting and summing the low-level features and dynamic weights of each perception modality. : ; in, This represents a learnable mapping function.
[0015] Utilizing a Transformer-based predictor, based on fused intent features Historical sequence predicts future moments Intentional characteristics .
[0016] Preferably, semantic-level compression is performed on the uplink and downlink data during transmission, including: Downlink compression at the operator end: The parsed task instructions are encoded into high-level skill primitive IDs and their parameters, or incremental motion instructions relative to the digital twin, and transmitted using a joint source-channel coding strategy; Robot-side uplink compression: Utilizes a detection and segmentation model to extract and transmit the target object's category ID, six-DOF pose parameters, and appearance feature descriptor.
[0017] Preferably, in step (2), constructing a bidirectional predictive digital twin includes: On the operator's end, the digital twin embeds a rigid-flexible coupled dynamics model, which uses forward simulation to predict the robot's state from the delayed moment to the current real time, while local simulation of operation commands is used to predict the robot's future state. On the robot side, a local digital twin copy is maintained, delayed task instructions are received and simulated for execution, and the simulation results are compared with real-time sensor data for deviation detection and instruction reconstruction reference.
[0018] Preferably, in step (2), the task instructions are reprogrammed and smoothed interpolated to generate a continuous and feasible human task flow, including: when a new task instruction is received, the task instruction is time-stamped and semantically decoded with the predicted state of the local digital twin to restore the operator's intention trajectory; within the sliding time window, the intention trajectory is reprogrammed and smoothed interpolated using robot dynamics constraints and job continuity requirements to generate a continuous, feasible human task flow that conforms to the operator's original intention.
[0019] Preferably, step (3) includes: (3-1) A two-stage planning generation robot autonomous decision-making flow validated by a large language model and knowledge graph; (3-2) Real-time assessment of operator intent confidence, channel quality, environmental disturbance intensity and machine autonomy confidence to calculate dynamic autonomy weights, and based on the dynamic autonomy weights, fusion of human task flow and machine autonomous decision flow to generate control commands; (3-3) Implement a three-layer progressive security protection system, including intelligent hybrid execution, model protection execution, and physical backup execution; (3-4) Establish a nonlinear model with time-varying delay for the entire closed-loop system, and derive the key parameter boundary conditions to ensure the stability of the system using Lyapunov-Krasovsky functional theory.
[0020] Preferably, step (3-1) includes: The large language model, after domain fine-tuning, parses the task instructions, calls upon internal common sense and domain knowledge to output a draft skill sequence; the draft skill sequence is then fed into a differentiable knowledge graph system, where entity relationships, physical constraints, and a skill prerequisite-effect model are used to verify logical consistency, security, and feasibility, and a graph neural network is used to predict the skill success rate. The system then calls upon a digital twin to deduce and outputs the machine's autonomous decision-making flow.
[0021] Preferably, in step (3-2), the dynamic autonomous weights The calculation formula is: ; in, ~ Indicates the adjustable weighting coefficient; For bias terms; The function maps the weighted sum to the interval (0,1); Confidence level of operator intent; For channel quality; Intensity of environmental disturbance; Confidence level of machine autonomy.
[0022] Operator Intent Confidence The consistency of the fusion weights output by the multimodal fusion network and the internal confidence of the multimodal large language model output are comprehensively measured.
[0023] Operator Intent Confidence The calculation formula is: ; in, This represents the dynamic weight vector for each sensing mode; for The information entropy of the distribution is used to quantify the uniformity of the weight distribution. The higher the entropy, the more uniform the weight distribution of each mode, indicating higher consistency of the multimodal signal and higher confidence in the intended signal. The internal confidence level of the MLLM output; These are weighting coefficients; The sigmoid function maps the result to the (0,1) interval.
[0024] Channel quality The calculation formula is: ; in, For round-trip time; Packet loss rate; and These are the weighting coefficients.
[0025] Environmental disturbance intensity Evaluation is based on sensor data from the robot's local inertial measurement unit (IMU), Doppler log (DVL), etc., and is quantified, for example, by calculating the deviation between the resultant external force and the desired control force, or by directly measuring the velocity variance of the flow field disturbance.
[0026] Confidence of machine autonomy The calculation formula is: ; in, For digital twins or MPC models at the current moment t The predicted state; The actual state measured by the robot's sensors; This is the sensitivity coefficient; Map the error to the interval (0,1].
[0027] Preferably, in step (3-2), the generation of control instructions based on the fusion of human task flow and machine autonomous decision-making flow according to dynamic autonomous weights includes: ; in, For control commands; For human task flow; This is for machine-driven autonomous decision-making processes.
[0028] Preferably, step (3-3) includes: Define the comprehensive uncertainty index : ; in, ~ These are weighting coefficients; when At that time, directly execute control commands. ; when When necessary, it automatically switches to model predictive control; when At this time, switch to physical safety control; , This is the state threshold.
[0029] Preferably, step (3-4) includes: Establish a nonlinear model with time-varying delay for the entire closed-loop system: ; in, It is a time-varying round-trip delay in communication; , For the robot's generalized coordinates, For the robot's speed, This refers to the robot's internal decision-making state. Indicates system state The derivative; To achieve uniform eventual bounded stability, a Lyapunov-Krasovsky functional is constructed. The stability condition is transformed into a linear matrix inequality. : ; in, The symbolic expression for a linear matrix inequality (LMI) is given by the Lyapunov-Krasovsky functional. After differentiating along the system trajectory, a matrix function is derived through inequality scaling. This represents the theoretical upper limit of the allowable delay. This represents the maximum rate of change in round-trip communication delay. This represents the theoretical upper limit of the value of autonomy. Represents dynamic autonomous weights The maximum rate of change, It is a structure The weight matrix.
[0030] Solve the above linear matrix inequality to find the maximum allowable range of the key parameters that ensure system stability.
[0031] Based on the same inventive concept, the present invention also provides a system for executing the aforementioned human-robot collaborative underwater robot operation decision-making method, comprising: Multimodal Intent Perception and Task Parsing Module: Establishes a communication-task effectiveness quantification model for the multimodal perception channel and generates a modality selection map, dynamically recommending the optimal combination of perception modalities; utilizes a multimodal large language model to align and fuse the natural language instructions from the operator's end with scene information, parsing them into structured task instructions; Remote operation state synchronization and robust command transmission module: Construct a bidirectional predictive digital twin, perform forward prediction of delayed robot state at the operator end, simulate the received sparse task commands at the robot end, reprogram and smooth interpolate the task commands, and generate a continuous and feasible human task flow. The hierarchical hybrid intelligent elastic decision-making and collaborative stability assurance module parses the received sparse task instructions into a robot autonomous decision-making flow through a large language model; it generates control instructions by fusing human task flow and machine autonomous decision-making flow based on dynamic autonomous weights; it selects the final execution mode using a three-layer progressive safety strategy based on the control instructions and the real-time system status; and it establishes a time-varying nonlinear model for the entire closed-loop system, deriving the key parameter boundary conditions to ensure system stability using Lyapunov-Krasovsky functional theory.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The human-machine collaborative decision-making method of the present invention significantly reduces the communication burden of the low-bandwidth acoustic channel in the deep sea through multimodal intent semantic compression technology, realizes the accurate understanding and efficient transmission of complex instructions of the operator, and overcomes the semantic distortion problem caused by traditional video stream compression. (2) The dynamic permission allocation and hierarchical security architecture of the present invention autonomously adjusts the human-machine control weights based on real-time channel quality and environmental disturbance intensity, effectively suppressing instruction conflicts under sudden flow field disturbances and improving the collaborative robustness of complex tasks; (3) The strong time delay stability guarantee mechanism of the present invention breaks through the instability bottleneck of the traditional master-slave control mode under second-level delay, and realizes reliable collaboration of the "human-machine-environment" closed-loop system in deep-sea cableless operation. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the decision-making architecture for human-robot collaborative underwater robot operations under limited communication conditions. Detailed Implementation
[0034] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not limit it in any way.
[0035] This invention provides a human-robot collaborative underwater robot operation decision-making method and system under limited communication conditions. The method achieves efficient collaborative operation between the underwater robot and the operator in a low-bandwidth, high-latency acoustic channel through multimodal intent compression, dynamic permission allocation and closed-loop stability guarantee mechanism, which significantly improves the success rate of complex tasks and the robustness of the system.
[0036] The technical solution of this invention comprises the following three core parts: 1. Multimodal Intent Perception and Task Parsing Module: This module first establishes a communication-task effectiveness quantification model for the multimodal perception channel, defining information density functions and task effectiveness functions for each modality to form a dynamic modality selection map, understanding the operator's intent with optimal communication cost. Secondly, it constructs an intent parsing hub centered on a fine-tuned Multimodal Large Language Model (MLLM), which aligns and fuses natural language commands with multimodal information such as eye movements and scene context, directly outputting a structured, executable task description (including core actions, target objects, constraints, and success criteria). To improve robustness, a spatiotemporal attention fusion network based on MLLM confidence guidance is further proposed to perform weighted fusion and short-term intent prediction on asynchronous multimodal signals. Finally, this module implements semantic-level communication compression: the robot transmits semantic features extracted by a lightweight model (such as object pose and feature vectors) uplink, while the operator transmits high-level skill primitives or incremental commands parsed by MLLM downlink, employing a joint source-channel coding strategy.
[0037] 2. Remote Operation State Synchronization and Robust Command Transmission Module: The core of this module lies in constructing a bidirectional predictive digital twin. On the operator's side, the digital twin uses an embedded dynamics model to predict the delayed robot state, providing near real-time feedback to the operator, and locally simulates operation commands to predict the robot's future state. On the robot's side, a lightweight local twin receives delayed semantic commands and performs rapid simulation, comparing the results with real-time sensor data for deviation detection and command reconstruction reference. To address the sparsity and non-uniformity of the command stream caused by time-varying channels, a robust command stream reconstruction engine is designed. This engine timestamps and semantically decodes the received command packets and the predicted state from the local twin, reconstructing the operator's intended trajectory. Based on robot dynamics constraints and operational continuity, it reprograms and smoothly interpolates this sparse trajectory within a sliding window, generating a continuous, feasible, and original reference command stream, providing clear input for subsequent decisions.
[0038] 3. Layered Hybrid Intelligent Resilient Decision-Making and Collaborative Stability Assurance Module: This module adopts a three-layer architecture to ensure secure and resilient decision-making and execution.
[0039] Task planning layer: Employs a two-stage reliable planner consisting of initial planning using an LLM, verification using a knowledge graph, and refinement. The domain fine-tuning LLM parses natural language instructions into a draft skill sequence, which is then checked for logical consistency, security, and feasibility using a differentiable knowledge graph. Graph neural networks (GNNs) are then used for success rate prediction and digital twin inference, ultimately outputting a verifiable formal task plan.
[0040] Hybrid Decision Layer: A dynamic weighted hybrid decision-maker is designed to evaluate operator intent confidence, channel quality, environmental disturbance intensity, and machine autonomy confidence in real time, and calculate dynamic autonomous weights. Based on these weights, the human task flow from the planning layer and the machine's local autonomous decision flow are integrated in real time to generate final control commands, achieving flexible integration and smooth switching of human and machine intelligence.
[0041] Safety Execution Layer: Three progressive safety protection layers are set up: Mode 1 (Intelligent Hybrid Execution) directly executes hybrid instructions; Mode 2 (Model Protection Execution) switches to protective controllers such as Model Predictive Control (MPC) when uncertainty is high; Mode 3 (Physical Backup Execution) switches to model-independent compliant impedance control when extreme anomalies occur to ensure intrinsic safety.
[0042] Collaborative stability analysis: A nonlinear model with time-varying delays and parameter uncertainties is established for the entire "human-machine-intelligent agent" closed-loop system. Aiming for uniform eventual bounded stability, stability analysis is conducted using Lyapunov-Krasovsky functional (LKF) theory. Linear matrix inequality conditions for key parameters ensuring system stability (such as maximum allowable delay and autonomous weight boundaries) are derived, providing a theoretical basis for system parameter design and safety monitoring. Verification and calibration are achieved through simulation and hardware-in-the-loop experiments.
[0043] like Figure 1 As shown, the specific steps include: (1) Multimodal intent perception and task parsing First, we conduct communication-task performance modeling and comparative analysis of multimodal sensing channels. The operator is considered a source of intent signals, and intent sensing devices, including but not limited to controllers, virtual reality (VR) glasses, microphones, eye trackers, and cameras, are treated as sensing channels with different characteristics. Based on this, we establish and quantitatively evaluate the channel capacity and task adaptability of the main sensing modalities.
[0044] Specifically, for each mode Define two key functions: information density function With task performance function Based on this model, a mode selection map is generated to dynamically recommend the optimal combination of sensing modes for different task stages and channel conditions. The function is defined as follows: A) Information density function : Quantify the amount of effective task information that a unit of data in this modality can convey, where, For the original data bandwidth requirements, The intention ambiguity after semantic parsing of the modal signal.
[0045] B) Task performance function Quantify the modality at a specific task stage and system status The contribution of the final mission success rate to the overall success rate. It is the stage identifier of the current task (such as "search", "approximation", "operation"). This is a snapshot of the current system state (e.g., robot pose, remaining battery power, environmental visibility, etc.). This function needs to be calibrated through human factors experiments and simulations. Based on this model, a modal selection map is generated offline to represent the system at different task stages. Different channel conditions and system states The system dynamically recommends the optimal combination of perception modalities to achieve a balance between communication cost and task performance.
[0046] Secondly, low-level features of each modality are obtained, and high-level semantic alignment and parsing are performed based on a multimodal large language model (MLLM).
[0047] Low-level feature extraction: The asynchronous raw signals of each modality need to be converted into low-level features that can be used for fusion. Specifically: Speech / audio signals: Low-level features are extracted by time-frequency features such as Mel-spectrogram or Mel-frequency cepstral coefficients (MFCC).
[0048] Visual signals (cameras, eye trackers): feature maps or feature vectors of images are extracted through a lightweight convolutional neural network (CNN) front end.
[0049] Handle / pose signal: Directly use its six degrees of freedom (6-DoF) pose, velocity or force / torque data as a temporal feature stream.
[0050] Digital twin scenario information: structured into machine-readable text descriptions or attribute graphs.
[0051] MLLM Cross-Modal Semantic Alignment and Parsing: Constructing an integrated intent understanding hub centered on MLLM. Operator-issued natural language commands (such as "Check that red valve") via voice, text, or graphical interface will be aligned and fused with gaze point information from eye trackers and the current scene context from the digital twin. (1) Referential resolution: MLLM binds referentials in natural language (such as "the red one") to specific object entities in the eye-tracking gaze point or scene context (such as "valve_003").
[0052] (2) Spatiotemporal correlation: Align the timestamps of each modal signal with the snapshot of the digital twin scene to ensure the consistency of semantic context.
[0053] (3) Joint Embedding and Parsing: The aligned multimodal information (text embedding, visual coordinates, scene description) is projected onto the same high-dimensional semantic space through the input layer of MLLM to form a joint embedding representation. Based on this, the fine-tuned MLLM outputs a structured robot task description, including: 1) core actions; 2) target object; 3) constraints; 4) success criteria. At the same time, MLLM outputs an evaluation of the contribution of its internal attention mechanism to each input modality as the internal confidence of each modality.
[0054] Next, based on MLLM confidence guidance, spatiotemporal attention fusion and short-term intent prediction are performed. To address the issue of single-modal signal quality fluctuations and offset communication delays, a spatiotemporal attention fusion network based on MLLM confidence guidance is proposed, whose input is the low-level feature sequences of the aforementioned modalities. Regarding the time alignment strategy, interpolation and other methods are first used to align the asynchronous modal features on the time axis. For dynamic weight fusion, a lightweight network is designed as a dynamic weight generator, typically a multilayer perceptron (MLP) with several hidden layers. Its input includes the modal features at the current time step. Historical integration status Current task phase System status The network uses the internal confidence scores of each modality obtained from MLLM as prior guidance. Dynamic weights are generated for each modality. : ; in, A probability distribution function is a mathematical function that maps any real vector to a probability distribution. Its core function is to normalize the probability of multi-class classification. That is, the aforementioned lightweight MLP network, The confidence prior is provided for MLLM. The weight generation comprehensively considers signal quality, task performance, and semantic confidence.
[0055] In feature fusion and prediction, weighted summation is used to obtain the fusion intention feature. , This represents a learnable mapping function (typically a single-layer or multi-layer neural network) used to map the raw low-level features of each perceptual modality. Mapping to a shared feature space of uniform dimension allows for weighted summation of features from different modalities. Subsequently, a Transformer-based short-term intent predictor is used, based on... Historical sequence predicts the future Intentional characteristics of time This partially offsets the transmission delay of downlink commands.
[0056] Finally, task-oriented semantic communication compression is implemented. To adapt to low-bandwidth channels, the fused and predicted intent information is efficiently compressed and transmitted.
[0057] A) Robot-side uplink compression: The robot does not transmit raw video / point clouds. Instead, it utilizes a lightweight object detection and instance segmentation model (such as the YOLO series or a lightweight variant of Mask R-CNN) deployed on the robot to extract and transmit only key task information, i.e., semantic features. These include: the target object's category ID, 6-DOF pose parameters, and lightweight appearance feature descriptors (such as NetVLAD descriptors) for matching or recognition. For images, a neural image compression technique is employed that aims to preserve the accuracy of these semantic features for reconstruction, rather than pursuing pixel-level fidelity.
[0058] B) Downlink Compression at the Operation Terminal: After being parsed by MLLM, operator commands are encoded into high-level skill primitive IDs and their parameters, or incremental motion commands relative to the digital twin. A joint source-channel coding strategy is adopted to provide strong error correction protection for key data (such as skill IDs) and to perform lossy compression coding on continuous parameters.
[0059] (2) Remote operation status synchronization and robust command transmission This section aims to address the critical challenge of maintaining consistent state perception among humans, machines, and the environment over unreliable, high-delay channels, and providing a high-quality, continuous, and secure input command stream for the synthesis of final control commands. Its core function is to receive and process raw information, providing directly usable data for decision-making and execution.
[0060] A) Cross-domain state synchronization based on digital twins and model predictions Digital twins are crucial for maintaining state synchronization; this study aims to establish a bidirectional predictive digital twin. At the operator's end, the twin receives robot state feedback contaminated with delays and packet loss. The twin embeds a pre-validated rigid-flexible coupling system dynamics model, using forward simulation to predict and display the robot's state from... up to the current real time This allows the operator to interact with a near real-time virtual agent. Simultaneously, the operator's commands within the twin are recorded and locally simulated, forming a predicted future state for the robot. On the robot's end, a lightweight local twin copy is also maintained. It receives delayed operator semantic commands and rapidly simulates their execution within the twin, comparing the simulation results with real-time data from local sensors. This serves two purposes: firstly, to detect deviations between the commands and the expected physical reality; and secondly, to provide a benchmark reference for subsequent command flow reconstruction.
[0061] B) Robust command flow reconstruction and preprocessing under time-varying channels Due to channel latency and packet loss, the command stream received by the robot from the operator is sparse and non-uniform. The core task of this section is to design a command stream reconstruction and preprocessing engine. When a new operator command packet is received, the engine first performs time-stamp alignment and semantic decoding with the local twin's predicted state to reconstruct the operator's intention trajectory. Subsequently, within a sliding time window, utilizing the robot's dynamic constraints and job continuity requirements, the sparse intention trajectory is reprogrammed and smoothly interpolated to generate a continuous, feasible reference command stream that conforms to the operator's original intention. This process does not involve autonomous decision-making; its core is to interpret and smooth the operator's intentions, providing a clear human instruction input for subsequent hybrid decision-making layers.
[0062] (3) Hierarchical Hybrid Intelligent Elastic Decision-Making and Collaborative Stability Analysis The architecture is divided into three logical layers from top to bottom: task planning, hybrid decision-making, and secure execution, and is guaranteed by a rigorous stability theory.
[0063] A) LLM-enhanced interpretable task planning and decomposition layer As the highest level of decision-making, it is responsible for transforming the operator's ambiguous high-level instructions into verifiable and executable sequences of machine skills. Its core is to build a reliable planner that goes through two stages: initial planning of a Large Language Model (LLM) and validation and refinement of the knowledge graph.
[0064] i) Semantic parsing and skill draft generation: When the system receives a semantically compressed task instruction from the operator, the Large Language Model (LLM), deployed on the operator's end and fine-tuned with domain data (such as underwater operation manuals and historical task records), first performs semantic parsing on the instruction. The LLM will then call upon its internal common sense and domain knowledge acquired through retrieval-enhanced generation techniques to output a structured skill sequence draft.
[0065] ii) Knowledge Graph Validation and Executability: The draft was then submitted to a domain-specific differentiable knowledge graph system for validation and refinement. The system was constructed as follows: First, core entities (e.g., robots, robotic arms, targets, environment) and relationships (e.g., "located," "operable," "collision risk") within the underwater operations domain were defined, and the preconditions and post-capture states of skills were formally described (e.g., the prerequisite for the skill "grab" is "the end effector of the robotic arm is within 10cm above the target," and the post-capture state is "the target is held"). Second, this structured knowledge was represented as graph data, and differentiable reasoning was achieved using graph neural networks (GNNs). The knowledge graph, utilizing its structured entity relationships, physical constraints, and predefined skill prerequisite-effect models, performed logical consistency checks, security verification (e.g., "the robotic arm is locked" must be satisfied before "grab"), and feasibility assessments on the draft. The GNN in the graph can quickly predict the success rate of each skill in the sequence and can invoke digital twins for millisecond-level deduction. Finally, an interpretable and verifiable formal work plan (a skill sequence) is output as a blueprint to guide subsequent execution. This interpretable and verifiable formal work plan is then transformed into a continuous, time-varying sequence of control instructions through real-time simulation using a digital twin; this is the machine's local autonomous decision-making flow. .
[0066] B) Dynamic Weighted Hybrid Decision-Making and Command Fusion Layer This layer is the core of real-time decision-making, responsible for dynamically fusing human task flows from the planning layer based on current communication quality, environmental disturbances, and confidence levels of capabilities of both parties. With machine local autonomous decision-making flow .
[0067] The core of decision-making is calculating a dynamic autonomous weight. This weight is calculated in real time by an evaluator whose input incorporates the following factors: i) Operator Intent Confidence Based on the consistency of fusion weights output by the multimodal fusion network in step (1) (i.e., dynamic weight vector) The value is comprehensively measured by considering the distribution characteristics of the multimodal signals (speech, eye-tracking, and handheld devices) and the internal confidence of the MLLM output. For example, the value is high when the fusion weights of the speech, eye-tracking, and handheld devices are evenly distributed and the MLLM confidence is high; the value decreases when a certain modality signal is missing or conflicting, causing the fusion weights to concentrate or the MLLM confidence is low. The calculation formula is as follows: ; in, The modal dynamic weight vectors output by the spatiotemporal attention fusion network in step (1); The information entropy of the weight distribution is used to quantify the uniformity of the weight distribution. The larger the entropy, the more uniform the weight distribution of each mode, indicating higher consistency of the multimodal signal and higher confidence in the intended signal. The internal confidence level of the MLLM output; These are weighting coefficients; The sigmoid function maps the result to the (0,1) interval.
[0068] ii) Acoustic channel quality Round-trip delay monitored in real time by the communication module and packet loss rate The calculation shows that: ; in, and These are the weighting coefficients.
[0069] iii) Environmental disturbance intensity Evaluation is based on sensor data from the robot's local inertial measurement unit (IMU), Doppler log (DVL), etc., and is quantified, for example, by calculating the deviation between the resultant external force and the desired control force, or by directly measuring the velocity variance of the flow field disturbance.
[0070] iv) Confidence in machine autonomy The matching degree of the current state is evaluated based on the robot's local decision-making module (such as the Model Predictive Controller, MPC), and calculated according to the following formula: ; in, For digital twins or MPC models at the current moment t The predicted state, This represents the actual state measured by the robot's sensors. This is the sensitivity coefficient, which controls the rate at which the confidence level decreases as the error increases. The error is mapped to the (0,1] interval; the larger the error, the lower the confidence level.
[0071] The evaluator can be modeled as a lightweight network or a rule-based system, with an example of a rule-based dynamic autonomous weight. The calculation is as follows: ; in, ~ This represents the adjustable weighting coefficient. For bias terms, The function maps the weighted sum to the interval (0,1). Its output... This represents the appropriate weight given to granting machines autonomous decision-making in the current context. When communication is excellent and the operator's intentions are clear, Approaching 0; when communication is interrupted, the environment changes drastically, or operator instructions are not feasible, Approaching 1.
[0072] Obtaining dynamic weights Then, the decision-making layer executes the final instruction fusion to generate control instructions that directly drive the robot. Integration follows these principles: ; This step is the final stage of control command generation. It combines pre-processed, interpretable human instructions with environmentally adaptable autonomous machine instructions, according to the current optimal decision-making strategy, into a single, executable control command. To ensure smoothness, The changes themselves need to be filtered to avoid jitter caused by the switching of control.
[0073] C) Layered security protection and execution layer This layer serves as the final barrier to ensure the absolute security of the physical system. It employs a three-tiered, progressive security strategy, selecting the final execution mode based on the mixed instructions output from the upper layers and the real-time system status. A comprehensive uncertainty index is defined. : ; in, ~ These are the weighting coefficients.
[0074] Mode 1 (Intelligent Hybrid Execution): When That is, when the uncertainty is low, the instructions output by the hybrid decision layer are directly executed.
[0075] Mode 2 (Model Protection Execution): When When the uncertainty is moderate, the system automatically switches to a model-based protection controller (such as Model Predictive Control, MPC). This controller is based on a simplified dynamic model with high confidence, and its primary goal is to ensure system stability and safety.
[0076] Mode 3 (Physical Guarantee Execution): When In other words, when there is high uncertainty, the system switches to the lowest level of physical safety control (such as impedance control with fixed gain). This mode does not rely on any environmental model; it only ensures that the physical interaction between the robot and the environment is compliant and intrinsically safe, buying time for human intervention. In the above formula, , This is the state threshold.
[0077] D) Stability analysis of cooperative systems This section aims to establish and verify the intrinsic stability theory of the "human-machine-intelligent agent" hybrid closed-loop system, which is the cornerstone for ensuring safe and reliable collaborative operation under extreme communication constraints and dynamic environments. The stability analysis will start from rigorous mathematical definitions and, through three steps—system modeling, theoretical derivation, and experimental verification—will provide quantifiable theoretical basis for the parameter design, security boundaries, and online monitoring of the entire collaborative architecture.
[0078] i) Closed-loop system modeling The entire system is abstracted as a nonlinear closed-loop control model with time-varying state delays and parameter uncertainties. The state of the underwater robot system is defined. This includes robot generalized coordinates. ,speed and internal decision-making status Closed-loop dynamics can be formally expressed as: ; in, It is a time-varying round-trip delay in communication, a function The structure is determined by robot dynamics, instruction reconfiguration logic, and hybrid decision rate. Indicates system state The derivative of .
[0079] ii) Stability analysis and verification Given the system's complexity and the presence of continuous environmental perturbations, uniform eventual bounded stability is chosen as the stability objective. Using Lyapunov-Krasovsky functional (LKF) theory as the primary analytical tool, a functional dependent on both the current and historical states is constructed. Differentiating it along the system trajectory, and using model equations, inequality scaling, etc., the stability conditions are transformed into a set of linear matrix inequalities with respect to the system parameters: ; in, This represents a symbolic expression for a linear matrix inequality (LMI), derived from the Lyapunov-Krasovsky functional. After differentiating along the system trajectory, a matrix function is derived through inequality scaling. This represents the theoretical upper limit of the allowable delay. This represents the maximum rate of change in round-trip communication delay. This represents the theoretical upper limit of the value of autonomy. Represents dynamic autonomous weights The maximum rate of change, It is a structure The weight matrix is given. By solving the above inequalities, the maximum allowable range of the key parameters that ensure system stability can be obtained.
[0080] Finally, verification was conducted in a high-fidelity digital twin simulation environment, actively injecting time-varying delays, packet losses, and disturbances of varying amplitudes and frequencies, and recording the system state. The stability boundary was visually verified using phase diagrams and Lyapunov functional evolution trajectories. Parameters were intentionally set to exceed the theoretical stability boundary, and instability phenomena were observed to verify the effectiveness of the theory. In hardware-in-the-loop testing, a realistic acoustic channel model was applied using a communication simulator. The performance degradation of the system near the critical stability parameters was measured, such as increased terminal jitter amplitude and task error divergence, allowing for comparison and calibration between the theoretical boundary and the actual engineering boundary.
[0081] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A human-robot co-adaptation method for underwater robotic operation decision-making, characterized in that, Includes the following steps: (1) Establish a communication-task effectiveness quantification model for multimodal perception channels and generate a modality selection map to dynamically recommend the optimal combination of perception modalities; use a multimodal large language model to align and fuse the natural language instructions from the operator's end with scene information, and parse them into structured task instructions; also includes a spatiotemporal attention fusion network guided by the confidence of each perception modality to perform weighted fusion and short-term intent prediction on multimodal signals to offset transmission delay; including: Align the low-level feature sequences of each perception modality on the time axis; Constructing a dynamic weight generator to generate dynamic weights for each perception modality : ; wherein, is a time instant is a perception modality of low-level features; is a history fusion state; is a current task phase; is a current system state; is a confidence prior for a perception modality is a multi-layer perceptron; The fusion intention feature is obtained based on low-level features of each perception mode and dynamic weight weighted summation : ; denotes a learnable mapping function; Utilizing a Transformer-based predictor, based on fused intent features Historical sequence predicts future moments Intentional characteristics ; (2) Construct a bidirectional predictive digital twin, perform forward prediction of the delayed robot state at the operator end, simulate the received sparse task instructions at the robot end, reprogram and smooth interpolate the task instructions, and generate a continuous and feasible human task flow. (3) The received sparse task instructions are parsed into robot autonomous decision-making flow through a large language model; control instructions are generated by fusing human task flow and machine autonomous decision-making flow according to dynamic autonomous weights; the final execution mode is selected by adopting a three-layer progressive safety strategy based on the control instructions and the real-time status of the system; a nonlinear model with time-varying delay is established for the entire closed-loop system, and the key parameter boundary conditions to ensure system stability are derived using Lyapunov-Krasovsky functional theory.
2. The human-robot collaborative underwater robot operation decision-making method according to claim 1, characterized in that, In step (1), the communication-task effectiveness quantification model of the multimodal sensing channel includes the information density function. and task performance function : ; ; in, For the original data bandwidth requirements, To determine the intent ambiguity after semantic parsing of modal signals. For the current task phase, This represents the current system state.
3. The human-machine collaborative underwater robot operation decision-making method according to claim 1, characterized in that, In step (2), the task instructions are reprogrammed and smoothed by interpolation to generate a continuous and feasible human task flow. This includes: when a new task instruction is received, the task instruction is time-stamped and semantically decoded with the predicted state of the local digital twin to restore the operator's intention trajectory; within the sliding time window, the intention trajectory is reprogrammed and smoothed by using robot dynamics constraints and job continuity requirements to generate a continuous, feasible human task flow that conforms to the operator's original intention.
4. The human-robot collaborative underwater robot operation decision-making method according to claim 1, characterized in that, Step (3) includes: (3-1) A two-stage planning generation robot autonomous decision-making flow validated by a large language model and knowledge graph; (3-2) Real-time assessment of operator intent confidence, channel quality, environmental disturbance intensity and machine autonomy confidence to calculate dynamic autonomy weights, and based on the dynamic autonomy weights, fusion of human task flow and machine autonomous decision flow to generate control commands; (3-3) Implement a three-layer progressive security protection system, including intelligent hybrid execution, model protection execution, and physical backup execution; (3-4) Establish a nonlinear model with time-varying delay for the entire closed-loop system, and derive the key parameter boundary conditions to ensure the stability of the system using Lyapunov-Krasovsky functional theory.
5. The human-machine collaborative underwater robot operation decision-making method according to claim 4, characterized in that, Step (3-1) includes: The large language model, after domain fine-tuning, parses the task instructions, calls upon internal common sense and domain knowledge to output a draft skill sequence; the draft skill sequence is then fed into a differentiable knowledge graph system, where entity relationships, physical constraints, and a skill prerequisite-effect model are used to verify logical consistency, security, and feasibility, and a graph neural network is used to predict the skill success rate. The system then calls upon a digital twin to deduce and outputs the machine's autonomous decision-making flow.
6. The human-robot collaborative underwater robot operation decision-making method according to claim 4, characterized in that, In step (3-2), dynamic autonomous weights The calculation formula is: ; in, ~ Indicates the adjustable weighting coefficient; For bias terms; The function maps the weighted sum to the interval (0,1); Confidence level of operator intent; For channel quality; Intensity of environmental disturbance; Confidence level of machine autonomy.
7. The human-machine collaborative underwater robot operation decision-making method according to claim 4, characterized in that, Step (3-3) includes: Define the comprehensive uncertainty index : ; in, ~ These are weighting coefficients; when At that time, directly execute control commands. ; when When necessary, it automatically switches to model predictive control; when At this time, switch to physical safety control; , This is the state threshold.
8. The human-machine collaborative underwater robot operation decision-making method according to claim 4, characterized in that, Steps (3-4) include: Establish a nonlinear model with time-varying delay for the entire closed-loop system: ; in, It is a time-varying round-trip delay in communication; , For the robot's generalized coordinates, For the robot's speed, This refers to the robot's internal decision-making state. System status The derivative; To achieve uniform eventual bounded stability, a Lyapunov-Krasovsky functional is constructed. The stability condition is transformed into a linear matrix inequality. : ; in, This represents the theoretical upper bound of the allowable delay. This represents the maximum rate of change in round-trip communication delay. This represents the theoretical upper limit of the value of autonomy; Represents dynamic autonomous weights The maximum rate of change; It is a structure The weight matrix; Solve the linear matrix inequalities to find the maximum allowable range of the key parameters that ensure system stability.
9. A system for implementing the human-robot collaborative underwater robot operation decision-making method according to any one of claims 1-8, characterized in that, include: Multimodal Intent Perception and Task Parsing Module: Establishes a communication-task effectiveness quantification model for the multimodal perception channel and generates a modality selection map, dynamically recommending the optimal combination of perception modalities; utilizes a multimodal large language model to align and fuse the natural language instructions from the operator's end with scene information, parsing them into structured task instructions; Remote operation state synchronization and robust command transmission module: Construct a bidirectional predictive digital twin, perform forward prediction of delayed robot state at the operator end, simulate the received sparse task commands at the robot end, reprogram and smooth interpolate the task commands, and generate a continuous and feasible human task flow. The hierarchical hybrid intelligent elastic decision-making and collaborative stability assurance module parses the received sparse task instructions into a robot autonomous decision-making flow through a large language model; it generates control instructions by fusing human task flow and machine autonomous decision-making flow based on dynamic autonomous weights; it selects the final execution mode using a three-layer progressive safety strategy based on the control instructions and the real-time system status; and it establishes a time-varying nonlinear model for the entire closed-loop system, deriving the key parameter boundary conditions to ensure system stability using Lyapunov-Krasovsky functional theory.
Citation Information
Patent Citations
Robot multi-modal fusion autonomous decision-making method and system based on large language model
CN121351007A
Cleaning robot autonomous operation method and system based on multi-sensor fusion and reinforcement learning
CN121455151A